Hacker Newsnew | past | comments | ask | show | jobs | submit | hmate9's commentslogin

Napkin math if we assume gpt 6 astra on max is >$15 million (just for output tokens) for those wondering.

over 5 days, you couldn't achieve that level of testing and communication with humans on such a complex problem in that amount of time.

some might go so far as to call this a country of geniuses in a data center.


In a way, I think you have it backwards.

Two mathematicians, through insight and thought, wrote out the proof over 1-2 years.

It took OpenAI a cost of $15m and with 10,000 subagents; that's around 60-120 mathematician's salaries ($250k-125k salary) for 1 year.

And, given now the cloud that OpenAI may have just "interpolated" (aka stole) the result, it's even more of a bear case for AI.


Bear case? I’m sorry?

70x uplift is a bear case?


Where did you get the human figure?

> cost of $15m

The retail price is not the cost.

Not to mention that the exponential plummeting cost of tokens means that that $15 million will be a "pocket change" within a decade or less: https://a16z.com/llmflation-llm-inference-cost/


> The retail price is not the cost.

True. The cost is probably much higher, since they are still subsidizing as part of the first phase of the enshitification playbook.


Yeah but they at least they got to steal $1 million from that nasty math prof who didn't want to remove his co-author.

They said in the post that they are NOT claiming the prize

This could fund 10 top income mathematicians for 8 years (based on https://careers.usnews.com/best-jobs/mathematician/salary ). Imagine what kinds of results we'd have to transform the foundations of science if we were giving brilliant minds this kind of funding to do nothing but research for most of decade....

Instead, we get slop proofs that are technically correct as PR stunts to enable corrupt kleptocrats, and most likely will drive research into culs-de-sac.


It is more expensive per task than 5.6-sol high: https://artificialanalysis.ai/models/gemini-3-8-flash#price-...

Cheaper at medium level while still being same score as Sol medium

Huh, according to some of those charts, it's both dumber, and more expensive to run against their benchmarking tasks than Fable??? Seems crazy to me.

Perhaps the model is able to evaluate that it's not done, and to keep pressing on in the face of mounting failures, until it eventually arrives at a solution. Where Fable can skip that.

Sol is still underrated imo, especially for the current discounted price

This model is not for builders and engineers. DeepSWE score of 49% is behind gpt 5.4 and muse spark. It's clearly intended to be an efficient model for google gemini usage.

What is interesting is how this is announced before any Gemini Pro progress. From the outside it seems as though Google cannot keep up with other frontier models.


Presumably 3.6 Flash is primarily meant to serve their own needs for the Gemini chat app, voice app (which Sergey Brin says he uses a lot in the car on the way to work), and for their search "AI Assistant".

Flash 3.6 is certainly capable of many coding tasks, but clearly a model of this size of not trying to compete at the frontier as a software development tool, and not clear why Google really need to complete there other than for PR-related AI bragging rights.

I don't know how a frontier model like GPT 5.6 or Fable could have done better (I have no need that justifies paying for them), but yesterday I used the free Gemini chat app (i.e. Flash 3.6) to discuss and explain this poorly written recent AI paper to me, and honestly couldn't ask for much more.

https://alignment.openai.com/measuring-reward-seeking/


Who tested it on DeepSWE?

Edit: Oh it's in the other link

https://blog.google/innovation-and-ai/models-and-research/ge...


Everyone wants to announce as late as possible (i.e., last) to chart the highest. Google is in a position financially to take a hit for these last few months.


Thank you guys for the feedback. I'll fix valid words being rejected, having space as a keyboard shortcut for shuffle, and maybe a "zen" mode without a timer.


What is this prediction based on?


Based on my conjecture that Anthropic is ahead on AI research, and that OpenAI doesn't know how to make Fable-class models.


Fable is allegedly a massive model (estimates between 6-10+ trillion, with a few hundred billion active). If 5.6 is just an incremental upgrade over 5.5 (at the same model size) then it won't be able to fully compete with Fable just yet.


I suspect the same just based on their versioning scheme fwiw.


solid


You can pay the most if you can get the most value out of it


No, you pay the most if you believe that you might get the most value out of it.

Moreover, the AI companies have not bought anything with their own money, but with the money of naive investors who believe that their money will be used by the AI companies to buy things out of which they will be able to get the most value.

So for now, this is strictly only speculation, which has driven the prices sky high. It remains to be seen who will really get any value (besides Micron, NVIDIA and the like, who have got good money for their products).


Not defending this, and it's far from ideal, but also credit card details were already pretty much confirming identities already.


This is incredibly bullish for china and open source models


I can’t help but feel like there’s something here that will matter for future LLMs.

The bidirectionality could be a big deal: being able to refine a sentence with both left and right context feels closer to how editing/thinking actually works than committing to each token forever.

Maybe the current models aren’t good enough yet, but the direction feels important.


I have google ai pro plan and tried antigravity with 3.5 flash but it used up all my quota in two prompts. If that is not a bug then it is seriously unusable.


Yesterday, or the day before, Google lowered the AI Pro quota from 33x standard usage to 4x.

From the talk on the Gemini subreddit it's severely lower than before. I'm likely canceling my AI Pro.

The update also broke the app for me. Editing a message crashes the app every time. I'm on a Pixel lol


The crunch is real.

- The model is appox 3.3x cost. - The model is realistically almost 5x cost due to token usage - Google has TPUs to run this on (yet the cost) - Google has a lot more security and backup cash compared to all other AI companies, likely even combined (yet the cost)

We can continue moving the goal posts, but it seems we're at a bit of a wall. Costs are increasing, intelligence is improving, but the cost is rising drastically.

You'd think Google of all companies in the mix would be able to sustain lower costs with how integrated they are with TPU, Deepmind and effectively unlimited budget.


It's an experience anyone who used Google BigQuery would be familiar with: start with an amazing engineering product, and keep continuously degrading the value users get out of a fixed dollar spend. It's like Google doesn't understand that lock-in doesn't work when customers can easily switch to Claude or GPT.


The way they're charging for failed generations is brutal.

Checked my 5 hour quota, it was 0%, got this for multiple attempts:

I'm getting more image requests than usual, so I can't create that for you right now. Please try again later.

or

Can you ask me again later? I'm being asked to create more images than usual, so I can't do that for you right now.

Went back and found they took 34% of my quota for the privilege of repeating that same error.

I think the "Usage Limits" screen is new so who knows how long they've been counting errors against our quota. I guess I should be grateful it's now visible.


I'm seeing this too.

API price for gemini-3.5-flash is 3x gemini-3-flash-preview so they might be throttling it 3x sooner. They should either drop API prices or not advertise AI Pro as supporting Antigravity.

https://ai.google.dev/gemini-api/docs/pricing#gemini-3.5-fla...


The web version went from 100 Pro Prompts per day to...12 per 5 hours lol. I just did 3 back and forth not even technical planning for an infra project and I am ~25% thorough. Insane.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: