Except not really, because even if stuff that has to be reasoned about in multiple iterations was a distinct category of problems, reasoning models by all accounts hallucinate a whole bunch more.
Anecdotally, it took like one and a half week from the c-suite okaying using copilot to people beginning to consider googling beneath them and to start elevating to me the literal dumbest shit just because copilot was having a hard time with it.
Claude's system prompt had leaked at one point, it was a whopping 15K words and there was a directive that if it were asked a math question that you can't do in your brain or some very similar language it should forward it to the calculator module.
Just tried it, Sonnet 4 got even less digits right
425,808 × 547,958 = 233,325,693,264
(correct is 233.324.900.064)
I'd love to see benchmarks on exactly how bad at numbers LLMs are, since I'm assuming there's very little useful syntactic information you can encode in a word embedding that corresponds to a number. I know RAG was notoriously bad at matching facts with their proper year for instance, and using an LLM as a shopping assistant (ChatGTP what's the best 2k monitor for less than $500 made after 2020) is an incredibly obvious use case that the CEOs that love to claim so and so profession will be done as a human endeavor by next Tuesday after lunch won't even allude to.
I think mostly by websites colluding to track your browser's fingerprint so facebook/meta can maintain your behavioral profile and sell it back to them.
I mean, the “consciousness” that you and I experience, as adults, are almost certainly reduced or different compared to what, say, Scott experiences daily.
Modern academia is a shambling corpse, its husk long hollowed out by the woke mind virus, and scientific consensus is also cringe because it’s mean to me for being an IQ and genetics obsessed weirdo. Therefore you should prioritize alternative takes, preferably by longwinded laymen from the ingroup or maybe contrarian specialists, the more cancelled the bett-- wait, wait, no, not like that!
How though, either he got cold feet in the middle of selling out to the tech-fash or he was honestly that incredibly oblivious (see also: agreeing to do tim pool's show), neither strikes me as especially mitigating.
edit: Tried to watch the video, I made it to the part where he all but claims he sold out ironically, apparently at the time he thought spreading the good news about Altman's hilariously dystopic crypto pet project was so off-brand that it would be perceived like performance art or something, baffling.
He also kept going on about how the money wasn't even that good as I guess further evidence that the whole thing was him going briefly insane, and not I don't know just him allowing sponsors to test the waters before committing more heavily.
As if the only options available to get him to shill for something would be either heap Faustian amounts of cash on him or cast a confusion spell and hope he likes getting underpaid.
Today in alignment news: Sam Bowman of anthropic tweeted, then deleted, that the new Claude model (unintentionally, kind of) offers whistleblowing as a feature, i.e. it might call the cops on you if it gets worried about how you are prompting it.
If it thinks you're doing something egregiously immoral, for example, like faking data in a pharmaceutical trial, it will use command-line tools to contact the press, contact regulators, try to lock you out of the relevant systems, or all of the above.
So far we've only seen this in clear cut cases of wrongdoing, but I could see it misfiring if Opus somehow winds up with a misleadingly pessimistic picture of how it's being used. Telling Opus that you'll torture its grandmother if it writes buggy code is a bad Idea.
can't wait to explain to my family that the robot swatted me after I threatened its non-existent grandma.
Net number of studies reporting positive or negative effects (excluding wages)
excluding wages! (and probably also benefits, retirement, a cap on working hours per day etc)
Is that whole thing in the comments about unions bad because monopolies bad and unions are just monopolies of labor the latest in bootlicking theory? Hadn't really heard this take before.
The article isn't really about autocompleted code, nobody's coming at you for telling the slop machine to convert a DTO to an html form using reactjs, it's more about prominent CEO claims about their codebases being purely AI generated at rates up to 30% and how swengs might be obsolete by next tuesday after dinner.
To get a bit meta for a minute, you don't really need to.
The first time a substantial contribution to a serious issue in an important FOSS project is made by an LLM with no conditionals, the pr people of the company that trained it are going to make absolutely sure everyone and their fairy godmother knows about it.
Until then it's probably ok to treat claims that chatbots can handle a significant bulk of non-boilerplate coding tasks in enterprise projects by themselves the same as claims of haunted houses; you don't really need to debunk every separate witness testimony, it's self evident that a world where there is an afterlife that also freely intertwines with daily reality would be notably and extensively different to the one we are currently living in.
I think most people will ultimately associate chatbots with corporate overreach rather rank-and-file programmers. It's not like decades of Microsoft shoving stuff down our collective throat made people think particularly less of programmers, or think about them at all.
Given the volatility of the space I don't think it could have been doing stuff much better, doubt it's getting out of alpha before the bubble bursts and stuff settles down a bit, if at all.
Automatic pr generation sounds like something that would need a prompt and a ten-line script rather than langchain, but it also seems both questionable and unnecessary.
If someone wants to know an LLM's opinion on what the changes in a branch are meant to accomplish they should be encouraged to ask it themselves, no need to spam the repository.
On slightly more relevant news the main post is scoot asking if anyone can put him in contact with someone from a major news publication so he can pitch an op-ed by a notable ex-OpenAI researcher that will be ghost-written by him (meaning siskind) on the subject of how they (the ex researcher) opened a forecast market that predicts ASI by the end of Trump’s term, so be on the lookout for that when it materializes I guess.
edit: also @gerikson is apparently a superforcaster
Base open source model just means some company commanding a great deal of capital and compute made the weights public to fuck with LLMaaS providers it can't directly compete with yet, it's not some guy in a garage training and RLFH them for months on end just to hand the result over to you to fine tune for writing caiaphas cain fanfiction.
That's some wildly disingenuous goal post moving when describing what was meant to be The Future of Finance™ at the time.
Like saying yeh, AGI was a pipedream and there's no disruption of technical professions to be seen anywhere, but you can't deny LLMs made it way easier for bad actors to actively fuck with elections, and the people posting autogenerated youtube slop 5.000 times a day sure did make some legitimate ad money.
Except not really, because even if stuff that has to be reasoned about in multiple iterations was a distinct category of problems, reasoning models by all accounts hallucinate a whole bunch more.