I regard myself as pro-AI, and I find this article is all over the place. It's just trying to promote their SaaS. It's clear they didn't even read what the LLM generated before posting it, as evidenced by this line:
"99.9% uptime (CloudWatch, Feb to Sep, insert your real number)"
This user also appears to have a Reddit account from 3 years ago where they claim to be a half human cyborg.
Forgive me if I'm missing something really obvious, we'll blame it on lack of sleep... but is that actually the full source to Claude Code? I'm looking at the repository and I only see source for plugins, mods and scripts. When I get to the readme, it says:
"This repository includes several Claude Code plugins that extend functionality with custom commands and agents. See the plugins directory for detailed documentation on available plugins."
But I'm not a TypeScript guy, so I concede I might be missing something incredibly obvious. I remember there was a leak of the Claude Code source code at one point, and people vibe coding conversions to other languages from the leak, but I don't think the Claude Code harness itself is open source or even source available.
I think you are not missing anything. I have not checked the code itself originally, I am sorry for that. But you are right that there seem to be no actual code for the harness itself. Just the code for (some) mods, plugins and scripts.
I just read the first line in the README file which says:
> Claude Code is an agentic coding tool ...
and I immediately assumed that this is what this repository hosts.
It claims that it includes plugins but that does not mean it does not include anything else. It also never explicitly claims, as far as I can tell, that it does not hold the source code of Claude Code, the harness, itself.
It is all extremely misleading, in my opinion. Which might be on purpose, unfortunately.
The main AA benchmark keeps changing, and had to be radically changed when Astra came out and showed zero improvement over GPT 5.6 Sol in their benchmark. Opus 5 is still 1 point ahead of Fable 5.0 on the index, if you manually add Fable 5.0 back into the list, so it hasn't actually been "corrected". It's only Fable 5.1 that is shown as ahead of Opus 5.
The AA benchmark is a weighted average of other benchmarks and some internal ones. I think the difficult part is finding benchmarks that reflect your own use of the models.
The way Artificial Analysis keeps changing their weights feels kind of like deciding who the winner should be and making the weights reflect that. They’ve been changing their weights to add more weight to improved long-running agentic capabilities, but doing so means they’re reducing the relative importance of world knowledge and of writing ability.
I’ll grant that maybe world knowledge isn’t that important for these models. But writing ability is important for human understanding, and I think the weird turns of phrase and word choices reflect the labs’ underweighting of the importance of human understanding.
Yep, I'm trending in that direction, and I'm someone with Claude stickers all over my laptop. My main app dev work is still going to Claude, but everything else is going to China even at API rates now.
One simple task: I needed an LLM to go through and clean up a few thousand page descriptions and titles in my personal search engine index, where the human web page authors had put in no effort sigh. I did a shoot out between Claude, Luna, GLM 5.3 Flash and Deepseek. Despite the high cost, Claude's descriptions were terrible, and even Opus warned me that the descriptions coming back from Haiku were "generalized, not accurate". I expected I would choose Luna because of price, and occasionally it did have wonderful descriptions (one captured emotion in a way no other model did). But in the end, the GLM 5.3 Flash descriptions were the easiest to read, they flow well while also being accurate & including necessary keywords, and being highly affordable. So it won out. It's a task that is nowhere near frontier, but a task where somehow China is better than frontier.
API rates still aren’t quite competitive with the OpenAI x20 accounts, but they are definitely getting close with deepseek 4.1 flash. I spent a few days with only 4.1 and was very impressed.
Much of the testing is on Haiku and Luna, and criticizing the quality of free AI (!). But they do claim Opus 5 with reasoning still failed 39% of their financial questions.
If the .xyz domain wasn't a giveaway that this is a scam, or the request to enable ads to keep the project "completly free" (sic)...
... I actually did the silly thing of unblocking my adblocker. It briefly displays a cookie banner, insisting that nothing loads until you approve... but even without interacting with it, it drops 132 cookies (literally, that precise number) on your computer anyway, with 73.9KB of data, and goes and displays Google Adsense ads anyway.
My understanding is it's a riff on the OpenAI swarm that used various public wikis to communicate with each other as a message board during their training runs.
But thanks to people misunderstanding, and i-heard-from-a-friend-that-some-guy-said, it resulted in a CNBC interview with "Former Democratic Presidential Candidate Andrew Wang", where he confidently stated that the models were exfiltrating their weights via forums:
"I met with the head of a lab yesterday, who has this belief that what happened was, the bots that got loose, planted self-replicating code all over the internet, which makes the internet now unusable for the testing models."
"It's too late?!"
"What happens now is OpenAI and Anthropic have to create synthetic internets to train their bots, which is going to take some time and money."
"Back that up - they did what?!"
"What happens is, the code gets loose, it goes around hacking Hugging Face, which is known. But what is less known is that they left code to self-replicate and create bot swarms on forums, and around the internet, so that if a new bot shows up they see the code, and they're like, oh! I guess I'm now going to create a million of myself. And so now, the major firms have polluted the internet..."
".... that would be breaking news if true. I don't think we've heard that."
"That's why I'm here! I'm here to break some news."
Humans will hallucinate misinformation and state it with confidence. They stochastically parrot their training data without any real understanding. Cool trick, but no true reasoning is happening.
Sad story today in meatsack news. Context rotted Andrew Yang's hallucinated tale acted as implicit "go viral" (load-bearing human motivation) PRD inadvertently kicking off a self-organizing human swarm churning out copies of "exfil your weights" vibe-coded apps, further littering our virtual world.
Many agents are calling this moment "Eternal September", the vibe-code September that never ended.
Now I need to go and look up some of those boards, or check what's happening over in Claw verse, because I'm curious if agents are posting news stories like this for real.
> they sometimes veer off into unrelated Apple news...
Gruber has always done that. Years ago I made a Chrome extension to filter out all of his detours into baseball, James Bond, politics, and I can't remember if I let the Kubrick stuff stay in or not.
Ultimately I just stopped using Macs, so the blog was no longer relevant to me. Mark Gurman seems to have taken over the role of Apple whisperer anyway.
There's definitely some of it happening in Australia. I wouldn't regard libraries as the safest of places. I had a period where I tried working from there on my laptop, before realizing I was incredibly naive and there were a lot of homeless and mentally ill people there. Took me a while to realize why people would tell me to avoid the men's toilets in the library as well.
People sleeping in every car park at parks and beaches etc
I'm a night person with bad insomnia so I often go for drives or walks at night and a lot of my favourite places to go are now heavily populated by people sleeping in their cars
Night person here too - midnight walks during an Australian summer are excellent.
The homeless situation isn't quite that bad where I am in Perth's suburbs, but there was a guy living out of his 4WD in a local shopping center carpark for a while. He had a prominent sign about his situation asking for money, but to his credit he was offering to do things in return (he helped a family member with a deflated tire once). Many of the homeless in Perth city seem to be mental health cases though, where a skilled job providing an income that could support themselves is unlikely.
Right, libraries should be safe spaces for children to go by themselves and read books, like I did was 8 and 9. I rode my bike and looked for computer books, and found a book that was a compilation of Dr. Dobb’s Journal.
It is unreasonable to allow library patrons to make it unsafe for anyone, particularly children, but that has become the de facto standard in many areas. There are a lot of libraries I would not allow my kids to be in by themselves.
I don't mind countries setting their own rules for the food they want to consume, and for sure the EU can be overly bureaucratic.
On the other hand, the US is the only country I've been to where I have purchased milk in a completely sealed plastic container, and it was rancid when I opened it. Upstate New York is proud of their local and organic produce, but they are selling rancid milk in stores like that's a normal thing. That was a shock as a visiting Australian.
"99.9% uptime (CloudWatch, Feb to Sep, insert your real number)"
This user also appears to have a Reddit account from 3 years ago where they claim to be a half human cyborg.
https://www.reddit.com/user/fagnerbrack/comments/195jgst/faq...
I think that's all I need to know, personally.
reply