Altman may not be an altruist but the overall direction the organization takes is what matters in the end. And as long as that organization retains even a semblance of a non-profit (despite having a revenue centric approach to tackle competition in the field), it is still a preferable organization to Anthropic and Amodei.
This maneuver requires you to anticipate all the edge cases or error messages beforehand which is practically not possible in many situations. The moment something unanticipated happens or the model changes its processing logic, the tool call system stops working just like any other deterministic program or tool.
> This maneuver requires you to anticipate all the edge cases or error messages beforehand which is practically not possible in many situations. The moment something unanticipated happens or the model changes its processing logic, the tool call system stops working just like any other deterministic program or tool.
Not all; error messages are part of UX design, and the user error message should always give an error that indicates what the user can do to fix the problem.
If you cannot open a file for writing, don't just return "error: cannot open MyFile.txt", return "MyFile.txt: permission denied" (so user can request additional permissions from whoever), "MyFile.txt: no space left on device" (so user can free up some space), "Myfile.txt: file exists and is a directory" (So user can retry with a different name, or remove the directory, etc).
I think what is happening now is that, with so many of the agent-using pool of devs having never shipped to end-users before, they are surprised that their "program" (the tool) is being used wrong by the end-user (the LLM).
Those of us with battle-scars already expect the user to use it wrong and have learned that it's easier to tell the user how to fix the problem than to ask the user to read the manual/do it the correct way.
So much this. I tell my juniors: To a beginner programmer, errors are 'the end'. They feel they did their best, it is not their fault and that is the error message they print. Experienced programmers know the user struggle, for them an error message is 'a beginning'. The first step of the user striving to solve the problem. They gave that command and they did not give it to fail. They (the users) still want to teach their goal.
Pro tip: Don't just print the return code, also print the call and it's arguments that failed, even without a stack trace.
Just add a --verbose flag that shows the stacktrace when there is an error. Then add a footer message when an error appears in non-verbose mode that invites the user/agent to use --verbose to get the full picture.
It obviously may end up in thousands of tokens burned through though (you can also fix that adding different levels of verbosity), but hopefully errors are not common.
I'm a Senior Freelance Programmer, I can see many of my past and present clients moving towards the exact path you described. I keep warning them during meetings that Claude model isn't sustainable for long, eventually the VCs will come for their revenues and Claude will be forced to close their access to all but the most enterprisey ones with deep pockets. The mere electricity cost for that kind of high level reasoning and abstraction can't be subsidized forever. However, there are other forces which pull them towards Claude and AI workflows. Most of the clients are in a "wait and watch" mode right now, using LLM assistance for code generation but not fully depending on them.
Before LLMs came, there used to be the technical debt to deal with in a project, now there is also the added cognitive debt which is way more subtle and impactful long-term. If your source of truth isn't source code but a prompt (or even a series of prompts with branches) and the executor of prompts is a non-deterministic agent, I think you've already lost the battle there.
You're standing on the shore, and your clients are having fun in the water. The tide is going up, and you're screaming at your clients "come back! it's not safe". And so, they show you the face. You appear to them like the boring guy who's not fun to hang out with.
Eventually the tide is high, there is strong current, and they are being swept away further and further from the shore and they are panicking : "pyeri! help us! please!"
People (the non tech people, the MBA people) don't want to hear what you, the tech guy has to say. You're the not fun guy. Stay in touch until they do need you and say : you were right. That's the day you charge them a dear price for the service.
AI is still at the bait stage of rollout. They subsidize it, they want you to get hooked onto it to the point where you cannot do without it. Then only, they start to charge.
I used google code assist for around 9 months. It was free. I would ask it questions from time to time, to help to fix bugs, and to avoid to spend an hour browsing SO. Now, it's around $30 per month.
They are losing too much too fast atm, they have reached the stage where they have to start to charge. Another one of their strategies is : IPO. Once they (openai/anthropic) are listed on the nasdaq, you will pay whether you want it or not (via your exposure to the nasdaq/S&p500 with your etfs).
> Stay in touch until they do need you and say : you were right. That's the day you charge them a dear price for the service.
That assumes OP will still be in business if that day even happens. Odds are that cheap AI access will be around for longer than freelancers can remain solvent.
Using today's model prices as a rebuttal is a very weak argument.
Two years ago, SOTA was gpt-o1, and it was much more expensive than Fable. Now, for $4,699, you can easily run a much smarter Qwen3.6-35B locally with DGX Spark.
Think about where we are. This is an era where a new SOTA arrives every two months. It took LLMs only about 18 months to go from chain-of-thought reasoning to disproving the unit-distance conjecture. chatGPT itself is only three and a half years old.
DeepSeek V4, released two months ago, is almost as cheap as the electricity costed, has the ability to being absolutely a top-tier model in 2025 standards.
You ignore that Claude are not alone, tech progresses and reduce costs, and there are always the Chinese alternatives which are becoming sufficiently better over time.
The electricity cost per unit of machine “reasoning” is vastly less than the cost of salary for human reasoning. That’s a weak argument. You should focus on the second part… LLMs (at least today’s) don’t build simple solutions, and the complexity they introduce has a cost.
Definitely. One of my fav techniques is to ask an LLM to simplify something by 90% or 99% if it looks overly complicated. Or asking the output to be critically reviewed by another agent too.
As a Freelance Programmer, are you even getting consistent clients at decent rates? If so, how are you getting clients consistently and how do you convince businesses that you are better than AI?
Might not be a very popular advice or even what you're looking for, but I think Anthropic banning you (or any user for that matter) is a blessing in disguise. The sheer amount of compute resources consumed by their high reasoning "pondering..." and "bloviating..." tokens isn't sustainable at scale. Eventually, they must ban everyone but those with deep and infinite pockets in order for this model to be sustainable and turn revenues.
Computers and LLMs are great at automation of low-level human cognitive tasks like memory, decisions and loops, etc. but struggle enormously with high cognitive tasks like reasoning, deep logic, nuance, etc. Not that it can't be done (Claude platform is proof that it can) - but the cost and scaling advantage in this realm belongs to the human brain, not the LLM.
This is as unhelpful as it gets. The genie is out of the bottle - people can create unimaginable things with claude within hours. What would take months/years can be done in a day or a week.
Being banned on those platforms is a real setback for many users. One might argue that openai etc. are valid alternatives, but when they dropped fable (and perhaps reinstate?), not being able to use it simply means others can do more/better.
I don't know if it's unhelpful if you consider broader time scales. In 2027 or 2028, Anthropic and OpenAI might decide to stop subsidizing LLM usage and charge enough to profit. Would you pay $100 for a bug fix? They're already headed in this direction. Fable was 10x cost of Opus 4.6. They can't keep burning cash forever and we're reaching the peak of what can be justifiably spent on training cycles. Speaking of training, did OpenAI just give up?
> What would take months/years can be done in a day or a week.
Hahahaha, this reads like pure unadulterated marketing. I sincerely hope you're getting paid for these things at least, it would be sad for you to be this way without even getting anything in return.
Oh! Are you a 100% human programmer willing to write in your contracts that you accept full liability for all impacts of any bugs in your code? Do you buy professional insurance, or just write completely perfect code?
This is idiotic. It’s like telling someone they should be grateful the electric company has banned them because artificial lighting will mess with their body clock.
If K2 or GLM 5.2 are on par with Opus 4.8 I'll eat my hat. They're good, but they're not that good. Deepseek V4 Pro has been better than Sonnet for me, but the only model that comes close to or surpasses Opus 4.8 is GPT-5.5.
GLM 5.2 is far better than deepseek V4. Seriously feels like I’m talking to a Claude model. Also burns tokens like one, so there is that. Deepseek is unbeatable on price/quality.
Honestly just give it time. This stuff moves so fast next month the conversation will be different. For folks who don’t like the ID privacy issues, use Deepseek et al and it should be able to get the job done even if the experience takes a bit more wrangling.
The problem with the ID verification is that they can pair introspective conversations with ID. Either that bothers people or it doesn’t.
Main point: we can’t fret about current state models because the ID verification has future implications. Models will change and competition will catch up. Do what feels right in the long run not whether TODAYS model is better at Anthropic.
Both Anthropic and OpenAI don't want to continue training models indefinitely.
Anthropic CEO has expressed potentially slowing down on model training. There is little return for billions of dollars burnt for 1-2% increase on various benchmarks. These companies profit via inference.
Not to mention, the whole Fable being banned by the US Gov is a scary prospect for future models. What is the point of spending billions if its going to get blocked?
Of course this can't go on forever. Especially not on LLMs. But are we really close to the limits of what these LLMs can do? I'm not sure we are.
The difference between GPT-5/Opus 4 and GPT-5.5/Opus 4.8 is striking. For software development anyway, there's no comparison. And all this has happened in a year.
My assumption is there will be another 2-3 years of improvements ahead of us on LLMs alone. Through hardware upgrades, larger training runs, better data quality, better algorithms, etc.
Of course, by then these models will be quite expensive. Will my company pay for it? I don't know. I'm sure some people will though.
Using even double the total tokens and taking, what, 2-3x the time?, still seems worth it if prices are 5x+ cheaper (which OpenRouter [1] claims is the case).
On NeuralWatt for my personal projects at home (not affiliated, just a happy customer), I get so much more mileage out of GLM than I get out of Claude at work, specifically because it's priced as a hammer I can pound any nail-shaped-object with, not a delicacy I need to carefully budget-analyze to try to figure out if it's worth burning my monthly spend limits on this task.
At some point, there will come a saturation point for that "Opportunity cost FOMO train ride", and I think we are already past that point. Mythos class models are a whole different beasts and cutting edge on reasoning but not much use for the problem domains most developers are trying to solve.
The present Sonnet/Opus versions (~4.8) will likely be what everyone in the enterprise might end up using eventually. And even though local models aren't there yet, there are budget alternatives from the families of DeepSeek, Kimi, GPT, MiniMax, etc. available through APIs of NVidida, OpenRouter, Groq, etc. which are very much Sonnet grade.
Personally, I don't think we're at that point yet. While I do think model improvement is starting to plateau (reaching a local ceiling), I'm not convinced local models are as good as sonnet/opus yet. The gap is still too much. But I'm excited for those models to reach those levels.
It became a thing due to the experience of self-reflection. The million dollar question is how come humans (and few other organisms) are able to self reflect on their biology, life situations, logic, math and even consciousness itself? However complex and sophisticated a machine's brain is, be it biological or mechanical (AI/AGI), no known laws of science allows it to self-reflect. This is famously called "the hard problem of consciousness" in philosophy which remains unresolved to this day.
Self-reflecting may not be the distinct enough feature. Any physical/chemical/electrical reaction can be termed as self-reflecting, as it reflects on what just happened and then responds with an effect. AI is already able to reflect on it's outputs and refine them, and distinguish between the user and it's own identity. Living things have evolved senses and long-term memory to help them with faster macro-responses beyond the usual physical reactions.
When a ball hits a bat, the ball also has a short-term memory and sense in the forms of how the inter-molecular forces detect and respond to the event of getting too close to the molecules of the bat and react with a repelling force. A more evolved form would be your consciousness.
Further, a lot of living things on earth might not have self-awareness.
It doesn't reflect itself, we only see the UI of a complex process, not the real thing. We don't understand what happened in our brains any better for being able to feel conscious. We can only be conscious of what is cost effective and cost necessary to feel, in order to persist and survive. Animals for example and primitive humans could reproduce without understanding reproduction mechanisms, just the operational side.
> However complex and sophisticated a machine's brain is, be it biological or mechanical (AI/AGI), no known laws of science allows it to self-reflect.
...What? If a human brain can, that's quite literally proof a machine can? We are made out of matter that obeys the laws of physics.
And we have never made a machine remotely complicated enough to mimick a human brain. LLMs are the closest we've ever come, and they're not even close. Nor are the even made in a way condusive to doing so (focused on generating requested output, not postulating randomly). So as to the mechanical machine specifically, nothing exists to even be capable of being observed to make such a claim!
It was totally a rigged referendum, that bus hoarding propaganda somehow worked and the masses fell for the lies. But a great number didn't and it was barely lost by a few percentage points. David Cameron should have stayed and battled it out instead of resigning.
David Cameron may be the biggest idiot in almost 300 years of British Prime Ministers, and there's been a few beauties. He really could have done a better job with the whole referendum business. Up/Down only. No "by a majority of...". No requirement for a second referendum on the terms of the disengagement agreement. Asking him to stay around would have been like asking the driver that drove you over the cliff to drive you home. And then there was that other beauty, Boris!
The other misses opportunity was to require a majority from all the UK's constituent countries.
One argument against Scottish independence vote just 2 (!) years prior was that Scotland would lose membership in the EU. And they really like this one it seems, with 62% voting to remain.
But since they're only 8% of the population and apparently don't count, they instead got Brexited against their will. Similar with Northern Ireland and, I suppose, Good Friday agreement.
Wales voted to leave, for some reason. Maybe they hoped one of the weekly hospital builds would happen there.
I think it should have been at least a two stage process. First vote do you want to leave? Second vote after figuring out the details of hard/soft etc. go back to the voters with - so this is the deal, do you want it? The second likely would have been a no.
Surely Liz Truss has to be the worst example of a Prime Minister? Known as the Iron Weather Vane (c.f. Thatcher's Iron Lady) and she lasted less time than a lettuce.
> various economic analyses estimate that the broader, ongoing cost of Brexit to the UK economy ranges between £100 billion and £140 billion due to reduced trade and investment
>Britain's national debt is growing at a faster pace than any country in the world except Botswana,