Hacker Newsnew | past | comments | ask | show | jobs | submit | pietz's commentslogin

I find the claims from OpenAI somehow more relatable and reasonable.

- They threw compute on a problem another team/company was rumored to have solved to see what their secret model could do.

- The texts I read do make it seem like OpenAI wanted to talk and share credit generously.

- Imagine working on a frontier math problem with someone at Anthropic and not only do you use Codex but also through a non-business account that allows training on your data.

- Timeline-wise, if they mainly used GPT 5.6 it's unlikely any meaningful data made it into an model that's being internally validated right now.


It’s fishy though that they heard one of seven problems was about to be solved and threw perhaps 15 million bucks at the right one.

[Edit: they said "two of": "On Tuesday, September 1, we heard rumors that two Millennium Prize problems had been resolved. .. we launched an effort ... on all open Millennium Prize problems".]


Why in the world is that fishy?

Isn't that exactly what almost everyone would do given that they wanted to see how capable their model is and the tense competition they have with Anthropic right now? Stealing impressive headlines from your competitor is pure gold.


If they were willing to spend 105 million it, perhaps. But that would be surprising. It’s fishy because they spent something like 15 million on the right one.

Again, that is the point. There is rumours that this one thing would be solvable, so they focus on this, spending 15m on one specific thing, instead of 105m on many. How is that fishy?

I understood the rumors (as described by OpenAI) to be that “one of” the prizes was solvable. If they heard NS specifically was solvable and aren’t saying that, they’re intentionally obfuscating that.

Because that’s the one Anthropic was rumored to have solved.

"On Tuesday, September 1, we heard rumors that two Millennium Prize problems had been resolved. Inspired by these rumors and by the step change in performance of our internal model, we launched an effort to evaluate it on all open Millennium Prize problems and a few other high-impact problems."

That doesn't reject my claim. They just didn't name them in this post. It feels like, you're going through great lengths reading something into this.

No, sorry I was reading it correctly but they addressed the issue in the post itself. They make it clear they were aiming at all 7 problems looking for the "two" that were rumored to be solved. But they didn't go hog on NS until they made progress. See sibling comments. I misread it the first time to be a claim that they "heard one of 7 human intractable problems are solved and spent 15 million on the right one".

Didn't they say that they launched an effort to evaluate their model on all open Millennium Prize problems, and then narrowed down to Navier-Stokes as the most promising after seeing results on simplified versions of the problems?

Yes, you're right. It seems like they claim to have started broadly and narrowed it down based on some progress. That leads to different questions but does answer my initial "fishy" point.

Isn't that *exactly* the type of solution you'd expect from AI?

Move 37 comes to mind.


Only if you understand nothing about the difference between LLMs and AlphaGo.

In case someone is asking: THIS is what a launch article should be like. 10/10.

Mission accomplished. That's both cool and fast.

> That's both cool and fast.

and probably a barely modified knock-off of some github project that it trained on


You're so upset that you have to invent an imaginary hypothesis to make yourself feel better.

Yes, very imaginary to think that the code comes from pretrained data and copy pasting whole blocks. It's not like this is exactly how LLMs work.

Correct, that's not how LLMs work [0].

[0] - https://arxiv.org/abs/1706.03762


Are you a frequent user of LLMs? That "copy pasting whole blocks" mental model doesn't hold up to regular usage, in my opinion.

They probably saw that report years ago of copilot dumping out the fast inverse sqrt function, and assume that's all they can do. From experience most anti-LLM people have either never used them, or used them back in the 3.5-4 era and then never again, though you might have even more experience with those people than I do. :P

they're wrong but they are right that this isn't interesting

The bar could not be any lower these days I guess

I know everyone is benchmaxxing but this one feels one step too far. Doesn't DeepSWE have both public and private tasks? I'd love to see the diff here.

It looks more like Google execs losing their mind and pressuring researchers to put DeepSWE directly into the training set.


That's not being debated here. The initial reported numbers were false and this was simply pointed out. You're changing the subject.

Opus 5 medium has the same score as 3.8 flash on artificial analysis intelligence index.

Are you implying Google or Artificial Analysis are reporting false numbers? What's your source?


BTW you're comparing 3.8 flash high to opus 5 medium. 3.8 flash medium scores lower.

Flash models are on the order of 1/10th the size of Opus models, so some flex in the thinking level is fair.

When comparing closed models, the only thing that actually matters to anyone using them is some mix of cost and speed. Considering how much memory a server is using, when evaluating models that you'll never have access to in order to host yourself, doesn't really make sense.

Your comment is really strange, why are you defensive towards WarmWash when gemini flash 3.8 high is both 6 times faster and costs less, while having the same intelligence score as claude opus 5 medium?

>Considering how much memory a server is using, when evaluating models that you'll never have access to in order to host yourself, doesn't really make sense.

This entire sentence makes no sense given what is being discussed.

https://artificialanalysis.ai/models/gemini-3-8-flash

https://artificialanalysis.ai/models/claude-opus-5-medium


I was being pragmatic. These are closed models on closed systems that you cannot hope to host. They are only available as black boxes available over web APIs served by their owners. Within that black box perspective, that we're force to have, the size of the model is, quite literally, just how much memory that server is using.

intelligence/model size is not a useful metric for a black box user.

intelligence/cost and intelligence/speed is a useful metric for a black box user.

Yes, it's cool, but as a black box user, the amount of memory a model is using on a server that I do not own has exactly zero practical use to me.

Cheers!


Flash is just a name with no defined or consistent meaning even within labs, let alone between them. Considering both are closed weight, there is no way to truly assess how big the size delta between the two is. Then again, who cares about size, performance and end-to-end speed+cost are what matters along with task adherence, task assessment and so on.

Model size also can not be inferred by tokens/sec for a multitude of reasons, but to showcase two examples, Opus 5 and Sonnet 5, as well as Gemini 3.1 Pro Preview and 3.1 Flash have each very comparable output speeds when using the same deployment as a basis for comparison, despite it being very likely that within their generation, the former are larger than the latter. Feel the need to mention this, as I unfortunately stumble upon so many poorly reasoned, speculative hype post trying to infer model size via utterly unreliable metrics, not based in actual data.

It’s like comments below arguing about the reasoning levels not normalized to some metric (like cost, output token amount or duration) but just the labels or high, max, medium, etc. Those mean almost nothing even when comparing models based on the same pretrain (just compare GPT-5.4 to GPT-5.2), they mean less than nothing comparing different labs releases.


It's not totally a mystery

https://arxiv.org/html/2604.24827v1

The short of it is by using hard facts knowledge that is difficult to compress, and then quizzing models on these facts and calibrating against a bunch of open models, you can kind of feel out the size of closed models.


I really like that one, but it kinda highlights what I could have far better explained. Their 90% PI is three times in both directions. Between 3T and 24T for GPT-5.5.

That’s a massively wide, inaccurate and at best barely informative range, demonstrating that even the most well thought out method will yield little usable information.

Additionally, I got some private evaluation taking a similar approach towards gauging models in topics I’ve found either over or underfitted by labs. If we just used that to rank models (not get a potential size range but just a rough order) Thinking Machines Inkling would need to be lager than Fable 5.


> [...] shows an intelligence score of 59, the same as Opus 5 medium!

Nothing here is false, you are simply confused. You either didn't read what they wrote in its entirety or decided to reinterpret what they did write.


"Beating opus" is the false part, no?

Stop lying. mattlondon said "gemini-3-8-flash shows an intelligence score of 59" which is undeniably correct. You can't say that number is false. You're literally lying.

All you had to do is go hover your mouse over "Models" in the top bar, hover over Claude Opus 5 and and click on medium: https://imgur.com/mlRCrt1

When you do that you arrive on this page: https://artificialanalysis.ai/models/claude-opus-5-medium

The gemini flash page for reference: https://artificialanalysis.ai/models/gemini-3-8-flash

You have to be an incredibly dishonest person to see a 59 on both pages and say "the initial reported numbers were false and this was simply pointed out. You're changing the subject".


The irony of this article being fully AI generated...

Anyway, it's over for Perplexity. They never had a great a product and the only reason for using them, was when they offered Pro accounts for free. Many people joined. Me included. But with a "meh" product and the general AI business not being very sticky, they lost quite harshly.

I thought they might be able to make money as a search api/index, but this article closed the book.


They were way above their main competition at the time Google decided to ignore the entire open web but they were still focused on searching it.

Since then, they decided to change focus into answering questions, and didn't maintain the quality of search results.


I'm not convinced an unstructured collection of memory files is the way to go at all.

If you look at how agents navigate source code, they do not look at directory names, and drill down into the ones with plausible names, instead the grep the whole repo for plausible keywords.

Of course, ideally your data would be structured, but the agents will mostly be grepping anyway, and maybe look at sibbling files.


its the same way they crawl websites, its horribly inefficient

With tiny models surpassing huge, 6 months old models on benchmarks, does anybody have some smart words to share on how these still "feel" different?

Artificial Analysis ranks GPT 5.6 Luna similar to GPT 5.4, but that never matches my real world experience. AA seems to do a good job making a single number as representative as possible but there is still so much benchmarks don't communicate.


I agree. They are definitely good - no issues with instruction following for example - but they miss the "intelligence" larger models have.

For implementation tasks, where I have the problem already defined and researched, or just simple task, I'd definitely use something like Luna xhigh or max. If the task is vague, or involves planning, I'd rather use Sol medium, even though it's theoretically worse on benchmarks.


Could do a "Big model for architecture and planning and smaller model (or local model) for implementation" sort of thing


One term of art that's emerged for this "feel" is "big model smell", first coined by @aidan_mclau. [0]

To my surprise I couldn't find any proper explainers of the term in a quick search, despite grokking it after seeing it in various contexts on Twitter, but Fable 5 offered a useful analogy: "A student who memorized worked solutions and one who understands the subject score the same on the test; you can only tell them apart by asking a question the test didn't. Real-world use is nothing but those questions, which is why a single AA number feels right and wrong at the same time."

In other words, big model smell is related to the underlying ability to "understand" when tasks are underspecified or out-of-distribution. This ability can be mimicked to parity by smaller, distilled models according to the density of the training data for particular tasks, but neural scaling laws still hold for generalized reasoning ability.

More recently with these smaller models, there's a separate but related "RL-fried" phenomenon, where they rely on CoT to "grind toward a checkable answer even in contexts (open dialogue, taste, ambiguity) where there is no checkable answer, and you get the tell: over-hedged, over-structured, relentlessly on-task, deaf to the subtext."

There are some other insights and caveats in the (short) conversation that I feel you may appreciate reading. [1]

[0] https://x.com/aidan_mclau/status/1807843014104211855 [1] https://claude.ai/share/d511a348-7c36-432f-a6d5-9deab2802615


Appreciate you taking the time. That fable analogy is well put. Almost obvious once you know it.


I wonder why they even released this. It's worse than the pixel 10 in some aspects, which in itself didn't feel like a big step forward.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: