Have a 5090, and yes it's very fast. But it's like the worst ADHD team member and requires constant supervision and review from larger models. It's context size on-card is good for super, suuuuuper shallow precision work.
The gb10/spark on top of it, that thing can refactor enormous monorepo architecture. The time it takes the 5090 to compact, reiterate and execute a plan is often the same time as the gb10.
could you not be like that? provide alternative language or go away. If youre offended, say so and be real. Noncommittal posits of personal preference are linguistic mosquitos of communication. on the flip side, how dare you disenfranchise a legitimate adhd perspective. one that i would say is entirely valid as someone functionally crippled by such plight. If you truly are offended, perhaps there is some truth you are reacting to preventing you from truly responding in good faith. words are lame like that ya? mine are as nauseating as your flyby ego droppings.
> could you not be like that? provide alternative language or go away. If youre offended, say so and be real.
Ok, as somebody with ADHD I find it offensive because I don't need constant supervision, implying people with ADHD need constant supervision is belittling and just plain wrong. So, I will call out an offensive trope if I see it.
> If you truly are offended, perhaps there is some truth you are reacting to preventing you from truly responding in good faith
No, because if there was some truth to it, I wouldn't be offended. Perhaps stop with the amateur psychology? You're not very good at it.
Exactly; when I first got my RTX 5070 Ti (16gb, to game with!!!, upgrading from VEGA56), I loaded then-latest Qwen3.6 (~30B, cannot remember exactly). My only prior LLM experience was with models <8gb, primarily llama3.1.
My technical-expert twin played around with these LLMs, for about an hour, and then correctly reasoned "it's able to be WRONG, faster."
This seems apt. My next LLM machine will be closer to 96gb+ vRAM.
Personally I think the price is way too high right now. It’s a power hungry gaming GPU. The efficient single card equivalent would be a 4500 Blackwell which launched at about $3500. Or you could get a 9700 32GB or an Arc B70 for well under $2k, today. You only buy a 5090 if you want absolute speed.
32GB is still not that much. I would rather get a Spark and have the RAM to experiment with larger LLMs, even if it was slow.
How does this relate to 4500 vs 5090? I'm just pointing out that 5090 likely has twice the performance of the 4500 and likely maintains that at 2x watts if you want.
You didn't specify in your earlier post, so I wasn't sure exactly which comparison you were making. But yeah, the perf/watt actually looks the same for those, so the cost per token evens out. It is nice not having to manage 400+W though. I like the 4000 for that reason, it's effectively a 3090 that runs at half the TDP.
A) the macos value add is enormous if you have any investment in the ecosystem, B) for me at least a GPU is completely useless for anything but being a token generator.
> No thanks to the "macos value add" that forces you to use Metal while Valve customers frolick in Protonland.
Crossover works on macos, too. So does moltenvk, so does vanilla wine, etc etc. You can run most games without a hitch these days (allegedly, according to /r/macgaming). But I don't play video games so a GPU would probably be better off in some kid's computer.
A GPU would be better-off attached to your Mac in an eGPU enclosure. There is not a single Apple Silicon GPU on the market that leads the industry in prefill, decode or power efficiency.
But of course, Apple doesn't allow that as part of their ecosystem. It's really a privilege to have MoltenVK perform worse than the fanmade HoneyKrisp driver. It's valuable when Apple refuses to sign AArch64 CUDA drivers for macOS. It's exciting to pay Crossover to support half of the library Proton offers for free.
Clearly, I'm some sort of ingrate that selfishly demands the best things, without considering how to accommodate the poor trillion-dollar megacorporation.
A 5090 has a 1.79TB/s memory bandwidth. Qwen 3.8 27B NVFP4 is 22GB. You cannot generate tokens faster than the weights can traverse the GPU memory, so that makes max generation speed without MTP to be 81T/s. Say MTP is giving you 0.5 acceptance rate (very good), that is 1.5 * 81 is 121T/s. Even with a perfect acceptance rate you would only get 162T/s.
Off the top of my head, I'm guessing we're missing sparse attention. But I'll run your challenge through and see where the gaps are. I promise I'm telling the truth :)
brother my point is they don't care about their product offerings outside of their phones. this post/thread is about one of their product offerings which is not a phone which is inferior to their competitors'. simple.
What in that thread is particularly impressive or noteworthy? According to the benchmarks I've seen, M6 raster performance is actually less efficient than M5 in many scenarios.
This 1000%. Data centres don't equate to medium sized labs and businesses. A stack of Macs is up and running without digging trenches, an electrician on staff and a department of PhDs to justify the spend.
It's likely that a stack of Macs will draw more power for slower prefill/decode than equivalently priced Nvidia GPUs. If power efficient inference is the goal, Macs are a non-starter.
I've had the opposite experience where we're refactoring the gnarliest shit anyone's ever seen because AI can actually understand it well enough to decompose, test, refactor, etc.
reply