It could be but proposing a standard and getting it adopted by everyone else is its own monopolistic behavior. Everyone here shits the bed when google tries to do it for ad measurement and decloaking
Edit: Now I think about it, this might be the cheapest way to run the DeepSeek V4 Flash 0731 on a dedicated inference server at original weights. I haven’t run mixed load benchmarks but I guess it’s possible to generate $3-$4 worth of tokens per hour and still maintain a usable per-user throughput.
There’s also input tokens. For many agentic use cases input/output ratio can be 2:1 or even 4:1. Non-quantised DeepSeek V4 Flash costs $0.14 / $0.28 on most inference providers with ZDR.
When self-hosting the model, cache hits are basically free. RAM cache also helps with hit rate (prefixes may be around for hours instead of minutes).
To be clear this project doesn’t aim to achieve the best inference economics per token. MI300X doesn’t have native MXFP4 so it’s not even the right platform for the model. That’s why very few deployment recipes are available.
It’s interesting to me because MI300X is quite accessible to a small team with budget for just 1-2 GPUs. DeepSeek V4 Flash otherwise wouldn’t even fit on 2x H100s.
We can run several coding agents during the day and batch inference jobs overnight and serve the entire team with guaranteed privacy, without compromising precision or speed.
In fact we found that many inference providers are quantising the weights or even KV cache, and due to the low prices they serve at massive batches, resulting in unstable throughput. I ran GSM8K as a quick validation test and this deployment is “better” than the OpenRouter endpoint in a statistically significant way (I wouldn’t name the provider here). I will run some follow up benchmarks and update the repo when I find some time.
830t/s is burst aggregate. ~500 is sustained and it's for 8 concurrent users. Meaning for $1.99/hour if you serve 8 users it's 8*$0.54, not just $0.54.
You shouldn't rent one out if you're just serving it for yourself, but from a financial standpoint if you sell to users you can take a 100% margin.
Bro 500 Aggregate. so that's 500 * 60 * 60 = 1.8M output which is .5$ at best... Not including pre-fill and stuff.
This is not the real margins, even if you are selling to 8 users it's 90 tps per median stream. So assuming that .6-.7$
This is not even remotely worth it.
You need to 3x this tps(~1500 tps) to be worth it, and that's what most providers are doing, at 20-30 users at 50-60 tps with better optimized batch processing and kernels you can make some profit.
This is exactly what I came to say. The price of Flash is so cheap that trying to run it locally or with your own hardware is pointless. I was using it about a month ago to program some stuff and ran it for 4 days non-stop and it cost me about $2.
> trying to run it locally or with your own hardware is pointless.
Serving local models has advantages other than price. If you work in restricted industries, or have a strong need to protect your IP, or if you just value privacy more than cost, you now have options.
If you don't do any attention steering, custom decoding or meddle with the weights maybe. Services are worthless unless all you do is write positive prompts.
As others have mentioned, there's the privacy factor as well.
Agentic workloads are somewhere around 1%/0.5%/98.5% input/output/cached tokens. Cached tokens are pretty much free for inference providers (if they implement sparse and compressed attention properly) and throughput for input tokens is much higher.
Lets assume that you've got 2 million input tokens, 1 million output tokens and 98.5 million cached tokens to process. That would cost 2 * $0.14 + 1 * $0.28 + 98.5 * $0.0028 = $0.8358 with DeepSeek API pricing.
For comparison, it would take 2M / 8000 + 1M / 800 = 1500 seconds to process this amount of tokens with the linked framework, which is about $0.83 when we assume $2/hr for one MI300X.
However, other inference providers have 10 times higher prices for cached tokens, which results in a comfortable margin.
And we should not discount that DeepSeek also gets paid in data, which is probably more valuable to them.
And I believe that this framework still has some room for optimization for generation with high batch sizes.
Your math is a bit funny if you're assuming the 1/0.5/98.5 ratios: you doubled input and output tokens but not cached. If you double cached tokens to match your original ratio it works out to around $1.11, and if you 10x the cached token cost it's around $6.08.
Based on your $0.83 estimate, the margin isn't great. This is within shooting distance of "at cost" which is probably pretty close to what DeepSeek is operating with, ignoring the value of the data they're collecting of course.
> And I believe that this framework still has some room for optimization for generation with high batch sizes.
If that optimization can bring this scenario closer to $0.50 then it gets pretty compelling, otherwise I'm not confident.
Oh, I messed up. Half-way through, I thought it would be a good idea to double the numbers so I don't have to deal with half millions, but forgot to also double the 98.5. Unfortunately, I can not edit it anymore.
I think the margins of DeepSeek may be a bit better than with this vibe-coded framework here, since they had the liberty of optimizing their models for their own hardware.
At the time, open frameworks were not anywhere close to achieving that number. Not sure whether they caught up. The software wizards at DeepSeek are quite skilled.
> should not discount that DeepSeek also gets paid in data, which is probably more valuable to them
That's agentic feedback loops for training, right? Any more detail on this, such as how they actually tell whether that data is good or not? That seems like a very hard problem, and like the value of that data is low compared to just building their own, controlled RL gyms.
Agents usually start with ingesting the existing code base, and DeepSeek can use those code bases for pretraining. And they will have filters on top of that to throw out garbage.
I am not sure how they are using the data for post-training, but there probably are ways to get signal out of it, e.g. sentiment analysis when the user begins cursing at the agent, or checking whether the user continued another session with the generated code, or started a new session with the same starting point as before, i.e. they git-stashed.
Generally, you can train on data that is quite bad (e.g. the entire internet). It will still work, but take much longer compared to clean data.
> e.g. sentiment analysis when the user begins cursing at the agent, or checking whether the user continued another session with the generated code, or started a new session with the same starting point as before, i.e. they git-stashed.
Thank you for elaborating, that's already useful. Anywhere I can learn more about this? I'm very interested in it!
Definitely not. Inference is not as expensive to operate as many people seem to assume. The frontier labs are probably making a lot of money from selling tokens. It’s covering all of the R&D costs like salaries, collecting training material, and running the large training operations that costs a lot of money.
They claimed that OpenAI and Anthropic have positive gross margins. I don't think there are many credible claims saying that's not the case (at least for API usage)?
Use nvidia hardware instead and use a larger cluster serving many more users concurrently. Easily 10x–20x higher token rate per GPU with public solutions like dynamo and sglang.
I think Deepseek is selling roughly at cost (perhaps a slight premium). They don’t guarantee that they don’t train on the submitted prompts, so I suspect they are mining the data. Mining for what? Well, who knows. Best case, mining to make Deepseek better. That said, I use Deepseek all the time. It has done a whole lot of ‘ls’ commands on my system, though.
If you have 2x DGX Spark it will run quite nicely. They cost only $8000 or so and use less power so you may be able to rent them cheaper than the MI300X.
As other calculated, even with a Mi300 you could not saturae it enough with one stream to break even with the DS API, so I think renting sparks would make it even harder cause they are considerably slower.
The MI350p exists and should run a decent quant (say, the ~96GB antirez mix) well, but you can get two rtx pro 6000s for one of these, or 8x (actually more) r9700 + probably the gear to run them, etc.
Otherwise, you can probably buy one of these second hand from somewhere (SXM A100s are available that way) and run it in an adapter board.
You can get one on ebay for like 20k, but it comes without the backplane and i dont think there is a pcie card adaptor from china like the ones for h200.
To be fair the development of GPUs have stalled over the years. If they kept up with the progress instead of focusing on enterprise market, likely 256GB consumer GPU would be a norm today.
Basically right from Lisa Su's speech: "AMD is essentially taking one of its MI350X accelerators and cutting it in half, resulting in a card with half as many compute resources, half as much memory, and perhaps most importantly, a bit over half of the power consumption"
The big question is whether the demand will stay if the subsidized pricing ends. That's what the bubble talk is about. Right now all the players compete for market share and don't care about the losses (hence the debt). But what happens if no one wants to lend them anymore?
There is no evidence they are losing money on inference, though?
Also if they are keeping the price low because they want to gain market share and reduce the competitiveness of Chinese models they won't be able to raise prices without providers serving open models (at cost + low margin) severely undercutting them.
There's no evidence they're making money, and we already know from the projected datacenter capacity in a few years that there will be for more supply than demand, so the major providers will have to repay that debt. Even if they are making money on inference, it's nowhere near enough to cover the bill. It's a losing proposition either way, especially with Chinese models now in play.
There is. Specifically the pricing for open models from third party providers on OpenRouter since inference is a "commodity" at this point. Unless we think that Opus/GPT-5.6 are many times less efficient than GLM 5.2 or Kimi Openai and Anthropic are making money from inference.
> it's nowhere near enough to cover the bill
Obviously it does not cover R&D, marketing and other spending but nobody has ever claimed that here.
I see your point, but Anthropic and OpenAI are not only less efficient, they have much higher operating expenses because of their massive AWS/GCP spending (due to not owning it) and have more capital expenditures than everybody else. It's not comparable to third-party providers making some profit on the margin.
They will be instantly bought out by companies, not individuals. The consumer bubble won’t pop for quite a while yet. Production also won’t ramp up while lack of real competition keeps the demand high.
Thing is, GPUs will always be on demand, look at their history, initially for gaming, then for hash cracking, then 3D rendering, then for crypto mining, and now AI training and fine tuning. When AI bubble bursts, there will be another bubble taking over.
The only solution is more companies making high end units, only competition will make it better for consumers.
Well, if no body can say when it will pop, then can we really say it's a bubble and it's overvalued?
I can tell you it will pop in 10 years and when it pops, it will still be 20x bigger than in 2026. Does that even make any sense?
People said AI bubble will pop soon in 2024 and that it was overvalued. Turns out, many AI stocks 10x, 20x since 2024. Actual usage has gone exponential as well. Anthropic revenue went from $100m ARR at start of 2024 to $80b ARR today.
> Well, if no body can say when it will pop, then can we really say it's a bubble and it's overvalued?
Well, given the literal trillions being spent, the only ways this pays off are:
1. AI replaces a non-trivial fraction of human employees.
2. Someone builds a Culture Mind, and humans become (hopefully) pampered pets of AIs we don't understand. Seems unlikely, but it would arguably count as a payoff even if it made money meaningless.
Or maybe the AIs don't want pets, and you get SkyNet. Which definitely doesn't care about paying off anyone's investments.
When you look at various news articles about investors, yeah, there are definitely a lot of rich people who think that they're going to automate all human labor or just bring about the Singularity. Possibly with them in charge of the rest of us. If you don't make these kinds of wild assumptions, then yeah, this is looking like one of the biggest bubbles ever.
Did you read your own article? It's well less than $3 trillion.
Now let's build the model out more. What is the projected revenue, backlog, improvements in existing big tech businesses such as AI helping Meta's ad business?
1. AI replaces a non-trivial fraction of human employees.
I mean, it's pretty clear that's going to happen. How could it not?
The only question is whether resources are optimally allocated at the moment to prepare for this. That seems unlikely at best. So yes, there is probably a bubble, and if so, then yes, it will pop, and then life will go on, with resources better allocated. Just like when the dot-com bubble popped.
> Across all tested generations, divergent paths serialize linearly with the number of paths k, following T(k)≈sk with no super-linear reconvergence penalty. Warp execution efficiency falls as 32/k, the penalty is independent of occupancy, and predication removes the serialization cost
There is more! I started to figure out how ADC/DAC chips actually work. I believe our CM6533 and its competitor the CM108 both use delta-sigma. The output isn't a staircase of those samples but a rapidly switching (MHz rates) 1-bit modulator (think: PCM) output, and a final analog filter smooths it into a continuous waveform.
Many countries now assume you have a phone. For example getting UK visa requires a smartphone. I don't think going without a phone is feasible nowadays.
Another question is if going with a burner phone that has just sim card and bank card, sufficient. But then you need appleid/google account on the device, and this again links back to your phone number, and it's not easy in practice to have proper clean device.
You can simply use an Apple account for your primary device and a Google account for your burner (or vice versa). Or set up a secondary Apple or Google account. Or use a device/OS that doesn’t require one of these accounts!
I hinted there how the NS chain of lookups works from . to your domain. The point is that we wanted to be able to move name servers around the ip addresss, but that wouldn't work for many domains. So - in some contexts moving IP's rapidly is possible, in some it's not. Fun.