Hacker Newsnew | past | comments | ask | show | jobs | submit | gundmc's commentslogin

There are numerous benchmarks that measure cost per task, which factors out tokens entirely. Gemini 3.8 flash is significantly lower than Sol on basically all of them

https://artificialanalysis.ai/#cost-tabs

That said, Luna is the undisputed king here at the moment and is what I use as my workhorse model.


>There are numerous benchmarks that measure cost per task, which factors out tokens entirely. Gemini 3.8 flash is significantly lower than Sol on basically all of them https://artificialanalysis.ai/#cost-tabs

Not sure if you read your own link but Sol 56 high ranks smack between Gemini 3.8 flash medium and high. Gemini 3.8 flash comes in as more expensive per task than Sol 56 high according to artificial analysis.

Luna high is literally 30X cheaper than Gemini 3.8 flash high.

You can limit the model viewer and they're getting better at testing multiple effort levels now: https://artificialanalysis.ai/?models=gpt-5-6-sol-medium%2Cg...

One reason is clear: Sol uses dramatically fewer output tokens than Gemini 38 flash https://artificialanalysis.ai/?models=gemini-3-8-flash%2Cgem...


I open the link and I see Flash 3.8 high at 0.58 and Sol at 0.95. I don't understand why you say that "Sol 56 high ranks smack between Gemini 3.8 flash medium and high" but that is clearly wrong.

On cost per intelligence task, Gemini38flash and Sol56 trade back and forth on cost depending on effort level. https://i.imgur.com/zPaWPXx.png As seen in this image, literally: Sol56 high ranks in between Gemini 38 medium and high. The image proves it.

I also included Sol56 xhigh, which ranks above even Gemini38 high.


I don't know if my code is just "complex", but I find that Luna on max ignores the surrounding style and completely ignores logical consequences of a change, like just writing `del arg1, del arg2, ...` instead of dropping it from the surrounding code. All LLMs make questionable decisions at times, but Luna requires so much guidance that it's faster to just type it out yourself. What kind of routine tasks can one accomplish with such a model?

Do you have code formatters, linters and static analysis?

I can get extremely dumb models to get our code style correct because of those guard rails and a specific style document.


It's so funny how many people diverge on the same model.

Ps. For the last week I diverged to Luna too, still need to check 3.8 flash.

But 3.6 flash was my go-to model 3 weeks ago and before it was deepseek flash/pro for a while.

None of the claude models seemed cost effective though.


Yes, I'm surprised there isn't more conversation around this being a way of the administration lashing out at Anthropic like they tried to do with the supply chain risk maneuver.


> Meta is going to surpass Google this year in revenue. I agree on the diversification front though

Meta might surpass Google on _digital advertising revenue_.

Google's overall revenue is still ~2x Meta's


In 20 years I expect basically all of these to move to web-based interfaces and away from thick clients. You're already seeing graphics heavy use cases like CAD do this (Onshape has been hugely popular and is cloud native on Linux). Even behemoths like SAP are increasingly web enabled through fiori.


it would be awesome to see less windows-only software. i am all for it.


Mythbusters made a version of this in an unaired segment of their 2006 episode about passing gas https://youtu.be/RHcDP_Yew-g?si=T7AONGdXPd4d_gM3


> in an unaired segment

checks out


The 5080 is 16GB VRAM, not system memory. I don't think you can get 24-32GB VRAM in a $500 box


Not on the same scale as AI, but my first ever AirBnB host still owns harley.com. He made his money writing "The Yellow Pages of the Internet" physical books and had turned down numerous lucrative offers from Harley Davidson.

Really fascinating and quirky guy as you can probably infer from the site.


Similarly, the guy who owned nissan.com never sold out and continues to spite Nissan Motors even in death.

https://nissan.com/

You've got to actually use a trademark-adjacent domain in good faith though, otherwise you might get the rug pulled from under you.

https://www.roadandtrack.com/news/a69634055/75-million-dolla...


Yes, I've been waiting for a real breakthrough with regard to 3D parametric models and I don't think think this is it. The proprietary nature of the major players (Creo, Solidworks, NX, etc) is a major drag. Sure there's STP, but there's too much design intent and feature loss there. I don't think OpenSCAD has the critical mass of mindshare or training data at this point, but maybe it's the best chance to force a change.


In this case it's not a direct duplicate of the link, but another post on the same topic. I find there is value in reading through the previous discussion that I may not have otherwise seen.


Gemini CLI is open source too, though I think the consensus is it's a distant third behind Claude Code and Codex


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: