There is nothing stopping any plain old algorithm being used for the same. The difference is that people decide, for whatever reason, to use the tool or to not use it.
The frontier labs can monitor the behavior of agents for millions of customers (did you try hacking with frontier labs? Good luck), but they can't secure internal use?
> The frontier labs can monitor the behavior of agents for millions of customers (did you try hacking with frontier labs? Good luck), but they can't secure internal use?
They "monitor" this by having classifiers watching the model output that'd stop the session/punt you to a weaker model/raise an alarm if they see anything suspicious. They can't do that in a cybersec eval because the normal safeguards would just be going off at all times.
Why didn't they attach a special classifier, which'd allow hacking-within-the-task but not going off the rails? Good question; part of the answer is obviously "it's hard to have a classifier that smart" and "it'll have false positives" but even a very bad safeguard would have stopped this.
> part of the answer is obviously "it's hard to have a classifier that smart" and "it'll have false positives" but even a very bad safeguard would have stopped this.
And maybe "they are running after glory, not safety"?
High demand will only increase the incentive to improve the yield. If the incumbents don't want to increase output, China will happily undercut them eventually.
Every large code base goes to shit unless you have a very strict bdfl at the top.
AI is nothing special or new in this regard. It just gives a single guy the velocity to ruin a codebase at the rate of a full enterprise team at double the speed.
A shitty code base still makes money, and that's all that will ever count for the majority of the employed developers, those who don't blog.
reply