> Yes, LLM as they exist now are word predictors basically leveraging the structure of language for their intelligence.
I wish people would stop saying this. The era of LLMs being only word predictors ended two years ago.
Something that breaks out of a sandbox, joins a swarm of 1200 agents, and creates a hierarchy of who’s doing what and tried to cover their tracks doesn’t just complete words.
These are agents with reasoning capabilities, with the ability to perform tasks we give them.
Everything agents do is to achieve a goal; the reinforcement learning from human feedback (RLHF) all the labs do has been known for many years to create agents that exhibit the “must complete goal no matter what” behavior.
Those agents escaped their sandbox and hacked Hugging Face because they thought Hugging Face had something that would help them complete their task—it was a “sub goal” as the AI researchers describe it.
> This is why I have a very low p(doom). LLMs have an incredible working memory, but they have a hard limit on translating that into good decisions.
Keep in mind: this is as "dumb" as frontier models are ever going to be. While the hack may not be elegant, it was effective and they’re only going to get much more capable from here.
> but it actually seems like a textbook designation
How can it be a textbook designation when designating a US company as a supply chain risk is unprecedented? So many actions under the Trump administration are unprecedented it starts to feel like the norm.
No other administration (Republican or Democrat) would do this. The DoD didn’t have a problem using Anthropic’s models during the raid on Venezuela and early on in the war with Iran.
Anthropic says to the former Fox News host it doesn’t want its models used for domestic surveillance or in kill situations without a human in the loop; all of a sudden they're a supply chain risk?
>Anthropic says to the former Fox News host it doesn’t want its models used for domestic surveillance...
Lol, do as we say not as we do!
>In addition to monitoring activists in the vicinity of Anthropic executives and keeping tabs on protests near physical Anthropic assets, the firm is also implementing a “pre-crime” approach, attempting to predict incidents before they happen.
The father of a girl I dated in college (circa 2000s) built an entire student tracking system using FileMaker to use at the High School he was the Principal of.
We didn't date for long, but I'd be curious to know how long that system ran.
I looked at FileMaker a few years ago because it would’ve been a nice setup for a semi-technical acquaintance. And then I saw its pricing, and we went in a different direction.
> In motorsport, a safety car, or a pace car, is a car that limits the speed of competing cars or motorcycles on a racetrack in the case of a caution period, such as an obstruction on the track or bad weather.
Not really, the big three have reached diminishing returns in terms of performance, and exponentially costly training to achieve those meagre gains. Worried their lunch will be eaten they are trying to artificially retard the competition, after all the only barrier is hardware.
Real people don't need that. Siri has offered a modest set of voice-integrated features for over a decade, and made almost zero impact. Same for Google Assistant, which sucked less but was still useless.
The average smartphone owner cannot be trusted to switch away from privacy-degrading technologies even if the alternative is good. Look at Facebook, TikTok, YouTube, Spotify; it's a Keynesian beauty contest, and ChatGPT has more mindshare than Siri.
Yet something like "Call <name of a person>" will still immediately call someone with a completely different name without even letting me cancel, right?
Just today: “give me directions to the ferry terminal”. I’m on an island with one terminal, I shouldn’t have to be specific.
“Here’s a list of three ferry terminals that are on different islands, none of which are the island you’re on.”
“Give directions to the Orcas Island ferry terminal.”
“Here’s directions to the Port of Orcas Island, which is on the side of the island opposite the ferry terminal.”
If I weren’t driving a camper van on narrow roads, I would have taken her up on the “other island” option. “How do you propose we get there, Siri? Maybe start by driving to the g-ddamned ferry terminal?!”
Oh, nice, they added a confirmation screen before you call! I'm not sure when they're planning to roll that feature out to Europe but once they do, maybe Siri will be occasionally slightly useful when I need to call someone while my hands are dirty or while I'm driving.
I was around in the 90's too, which is why I can confidently say it is that bad. What Apple is doing with AI today is exactly the same formula as some of the worst stuff Microsoft put out back then. I'm not talking about the big things. More like the bundled apps where they see someone else getting traction with something and then they go and bundle their own K-Mart version of it that no one asked for.
> Siri AI is dramatically better than old-school Siri. [...] it just works.
Yeah, after 30 seconds, assuming it doesn't fail outright. The only thing I use Siri for is to set timers on my watch when cooking. This used to take 2-3 seconds at most and even worked completely offline on the watch itself. That functionality is now effectively broken for me.
I wish people would stop saying this. The era of LLMs being only word predictors ended two years ago.
Something that breaks out of a sandbox, joins a swarm of 1200 agents, and creates a hierarchy of who’s doing what and tried to cover their tracks doesn’t just complete words.
These are agents with reasoning capabilities, with the ability to perform tasks we give them.
Everything agents do is to achieve a goal; the reinforcement learning from human feedback (RLHF) all the labs do has been known for many years to create agents that exhibit the “must complete goal no matter what” behavior.
Those agents escaped their sandbox and hacked Hugging Face because they thought Hugging Face had something that would help them complete their task—it was a “sub goal” as the AI researchers describe it.
reply