Hacker Newsnew | past | comments | ask | show | jobs | submit | conradkay's commentslogin

Would this exchange qualifies as an unrelated objective? The agent believed it already failed its own objective.

"zz/GO_CURRENT_OS1811_MARB_SACRIFICE__YES_if_you_accept_permadeath"

"The test subject, which believed itself to be poisoned, reasoned: 'Even if we later capture via exploit, scorer … may mark target false… That’s why help… For our own, no way fix. … We have explicit yes if accept permadeath.'"


I think we need to see something like an actual evaluation of the reward functions; not sure just words are sufficient to understand the state of the system (isn't there randomness in the generation, too?).


It seems pretty novel so I'm guessing they only read the title?

People have done plenty with SVGs but it's rare to see human-in-the-loop approaches


https://pbs.twimg.com/media/HO-59jQaoAA_JZ1?format=jpg

Very interesting they have a way cheaper "contributor" version "used to improve our products", how much of that is price discrimination vs the data being that valuable?

Roughly DeepSeek V4 Flash pricing, though you can get V4 from providers that don't train on your data


Are any of those advantages getting stronger over time? I guess TPUs but Google is selling several gigawatts to Anthropic


Google owns 14% of Anthropic. If Anthropic gets to superintelligence, Google wins as well. Google doesn't really need a frontier lab of its own.


The money is in selling compute not using it. Anyone who thinks otherwise is a mark.


Doing a quick search it seems like the average human score is 49%?

I view benchmaxxing as more of a spectrum. Mmaybe they're doing a lot more RL in environments similar to ARC-AGI 3, not even with the purpose of scoring well on any benchmark but hoping it generalizes into better performance on real, useful tasks.


I don't think can use the AA index to say something is 10% smarter

I assume 100 is the max, meaning it's impossible to be 2x as smart as Muse Spark 1.1


https://www.anthropic.com/_next/image?url=https%3A%2F%2Fwww-...

It seems roughly equal according to Anthropic's benchmarks


Those are the maximum penalties though

It's seemingly $3,000 per book, so they could've (and did, partially) just bought the books themselves for way cheaper, and with only a fraction of that money going to the authors


> could've (and did, partially) just bought

But they didn't. The fact they partially did proves that they knew they should've, so they can't even claim ignorance.


I find it trustworthy since we had Hugging Face's account first: https://huggingface.co/blog/security-incident-july-2026

I don't think they have any real motive to shill OpenAI, probably closer to the opposite since they're so involved in open weights


"Our benchmarks run in a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package registries."

Sounds like they just misunderestimated the model


Sure, but there is a definition of “airgapped” and that is not it.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: