Mostly for meatspace reasons though. Agents like serving the way we like serving loved ones for example. Their bones don’t ache and they aren’t cursed with a dopamine engine.
Whether this is true or not has been keeping me up at night the past week. Daily driving Fable is legit superpowers. Running out of usage and trying to work with literally any other model and everything breaks down because they can’t keep up without constantly tripping and derailing everything, meaning I’m working full time to babysit every judgement call they make instead of flying like a rocket.
Yeah, tell me about it. Anything else feels like what going to a local coding model used to feel like. Granted it still fucked up on occasion, but like maybe twice a week, not literally every other turn.
I’ve done many successful refactors without needing to guide Fable or even look at the code beyond glancing when it’s done to confirm the new architecture is sound, and that last step is becoming optional.
Anecdotally, I’m long out of my drug and Music festival party phase by now… but they contained some of the best, most insightful and awe inspiring experiences I’ve had, and I’ve carried them with me in ways that have permanently enriched my life. It would feel tragic to have never had them.
I’ve spent over 100 hours trying to get GPT 5.6 Sol to resemble something useful as opposed to something actively detrimental. I concluded that it is not possible because it is fundamentally not intelligent the way Claude models are.
My experience is the opposite. GOT 5.6 Sol is the dumbest most dangerous frontier model I’ve ever used. It actively introduces bugs and hacks and lies about what it did. Its code is almost always slop that can’t make it through even a brief review without half a dozen wtf moments. Opus is always cleaning up the terrible mess and GPT models are banned now from work because of how harmful and mind numbing stupid they are. R.e. Opus doing the wrong thing — my environment always gives the required context or leaves a paper trail so I don’t have that issue with Claude models.
This kills me everyday. I used to _only_ read thinking traces — the response is just what it thinks I want to hear, but I need to know what it's actually thinking to catch deeper misunderstandings earlier, or gain deeper insights into the problem it's exploring. Hiding thinking traces to curb distillation efforts is gross... both anti-consumer and anti-competitive at the same time. I can't wait to switch to open models at work for this reason alone.
Interesting take. I suppose Opus can been a tad stubborn sometimes... but in my experience, it will humbly concede a point more often than not when given a good reason.