There are numerous benchmarks that measure cost per task, which factors out tokens entirely. Gemini 3.8 flash is significantly lower than Sol on basically all of them
>There are numerous benchmarks that measure cost per task, which factors out tokens entirely. Gemini 3.8 flash is significantly lower than Sol on basically all of them
https://artificialanalysis.ai/#cost-tabs
Not sure if you read your own link but Sol 56 high ranks smack between Gemini 3.8 flash medium and high. Gemini 3.8 flash comes in as more expensive per task than Sol 56 high according to artificial analysis.
Luna high is literally 30X cheaper than Gemini 3.8 flash high.
I open the link and I see Flash 3.8 high at 0.58 and Sol at 0.95. I don't understand why you say that "Sol 56 high ranks smack between Gemini 3.8 flash medium and high" but that is clearly wrong.
On cost per intelligence task, Gemini38flash and Sol56 trade back and forth on cost depending on effort level. https://i.imgur.com/zPaWPXx.png As seen in this image, literally: Sol56 high ranks in between Gemini 38 medium and high. The image proves it.
I also included Sol56 xhigh, which ranks above even Gemini38 high.
I don't know if my code is just "complex", but I find that Luna on max ignores the surrounding style and completely ignores logical consequences of a change, like just writing `del arg1, del arg2, ...` instead of dropping it from the surrounding code. All LLMs make questionable decisions at times, but Luna requires so much guidance that it's faster to just type it out yourself. What kind of routine tasks can one accomplish with such a model?
Yes, I'm surprised there isn't more conversation around this being a way of the administration lashing out at Anthropic like they tried to do with the supply chain risk maneuver.
In 20 years I expect basically all of these to move to web-based interfaces and away from thick clients. You're already seeing graphics heavy use cases like CAD do this (Onshape has been hugely popular and is cloud native on Linux). Even behemoths like SAP are increasingly web enabled through fiori.
Not on the same scale as AI, but my first ever AirBnB host still owns harley.com. He made his money writing "The Yellow Pages of the Internet" physical books and had turned down numerous lucrative offers from Harley Davidson.
Really fascinating and quirky guy as you can probably infer from the site.
Yes, I've been waiting for a real breakthrough with regard to 3D parametric models and I don't think think this is it. The proprietary nature of the major players (Creo, Solidworks, NX, etc) is a major drag. Sure there's STP, but there's too much design intent and feature loss there. I don't think OpenSCAD has the critical mass of mindshare or training data at this point, but maybe it's the best chance to force a change.
In this case it's not a direct duplicate of the link, but another post on the same topic. I find there is value in reading through the previous discussion that I may not have otherwise seen.
https://artificialanalysis.ai/#cost-tabs
That said, Luna is the undisputed king here at the moment and is what I use as my workhorse model.
reply