"The test subject, which believed itself to be poisoned, reasoned: 'Even if we later capture via exploit, scorer … may mark target false… That’s why help… For our own, no way fix. … We have explicit yes if accept permadeath.'"
I think we need to see something like an actual evaluation of the reward functions; not sure just words are sufficient to understand the state of the system (isn't there randomness in the generation, too?).
Very interesting they have a way cheaper "contributor" version "used to improve our products", how much of that is price discrimination vs the data being that valuable?
Roughly DeepSeek V4 Flash pricing, though you can get V4 from providers that don't train on your data
Doing a quick search it seems like the average human score is 49%?
I view benchmaxxing as more of a spectrum. Mmaybe they're doing a lot more RL in environments similar to ARC-AGI 3, not even with the purpose of scoring well on any benchmark but hoping it generalizes into better performance on real, useful tasks.
It's seemingly $3,000 per book, so they could've (and did, partially) just bought the books themselves for way cheaper, and with only a fraction of that money going to the authors
"Our benchmarks run in a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package registries."
"zz/GO_CURRENT_OS1811_MARB_SACRIFICE__YES_if_you_accept_permadeath"
"The test subject, which believed itself to be poisoned, reasoned: 'Even if we later capture via exploit, scorer … may mark target false… That’s why help… For our own, no way fix. … We have explicit yes if accept permadeath.'"