This is the notice from DeepSeek regarding their API:
We will adjust the pricing for the Flash series effective from 12:00 Beijing Time on September 10, 2026. During off-peak hours, the unit price will be $0.003 for input cache hits, $0.15 for input cache misses, and $0.6 for output. Peak-hour prices will be double the off-peak rates. Please plan your usage accordingly.
-------------------------------------------------------
Hoje em sites como openrouter o valor é de $0.16 output .
Direct 1:1:1 comparison, for V4.1 Flash - V4 Pro 0813 - V4 Flash 0731
Input cache hits (per 1m tokens) - $0.003 Vs. $0.022 Vs. $0.007
Input cache miss (per 1m tokens) - $0.15 Vs. $0.66 Vs. $0.22
Output (per 1m tokens) - $0.6 Vs. $1.98 Vs. $0.66
This is taken from https://api-docs.deepseek.com/quick_start/pricing, and it's comparing only off-peak hours pricing. It looks like V4.1 Flash is cheaper than the current 0731 flash model, and much cheaper than the current V4 Pro model.
> healthy profit margin (as near as we can tell from the outside)
Ugh we still don't know if this is true and it's nearly impossible to calculate without a full understanding of the real CAPEX cycle. Stop spreading these rumors until we know for sure.
SemiAnalysis estimates their profit margin to be 70%. To be losing money on inference implies that their costs are almost 4X higher than SemiAnalysis has calculated. That's not credible.
And which they could not charge anyone for. Unless these were extra resources that would otherwise go unused it cost them the amount they could have charged for them. Normally I would expect most businesses to make reasonable tradeoffs when it comes to how to allocate resources. I’m not convinced that any of the AI providers should be given that benefit of the doubt.
Whether on net they turn a profit as company overall is neither here nor there.. My point is that they are selling API tokens at a profit (or if being pedantic, then at a price higher than the cost to serve them ignoring research costs). And that that price is got a healthy margin which they don't charge themselves.
I would like to imagine
accounting inference revenue on trained model and the depreciation cost for training that specific model must already been capitalized + compute to serve would be a net positive margin business. Ongoing training must rather be for future models.
But again once future models arrive they would render older models useless, so the asset must be depreciating really fast.
Would love someone to throw light on revenue and cost recognition at the unit level for this.
It's like having new solar panels installed every week. Sure you're "profitable" on the $0.20/kWh you're selling your "free" energy at when you ignore the cost of the solar panels you're buying every week.
Regardless the profit margin as a talking point seems to be bad as AI as a tech might never be reversed whether anthropic failed or succeeded. Indeed it's imperative we subsidize AI companies and tech to make them explore more solutions to scientific problems which has a downstream effect on human flourishing.
Like? I feel breakthroughs that can be found via AI might help us more in the long term where even previously non AI fields can be helped by AI. So you have specific non AI research in mind that we're underinvesting in? Because the USA is already spending crazy anyway for healthcare and I don't feel like funding is the issue but better incentives, reforms etc
But US also spends too much on education as well. The issue doesn't seem to be funding but the educational reform like in mississippi, where they increased student performance without increasing their budget too much. That's why you see bad k12 educational outcomes compared to the budget spent in blue states. It's all about efficiency. Give AIa chance in few years as I feel it can make great strides.. it's hard to imagine that chatgpt released in 2022 and look at the progress in just few years as it just changed software engineering field entirely.. i expect similar kinda progress where of course humans will still be making breakthroughs but it'll be accelerated with the help of AI.
Spending on health insurance is spending on health care.. Americans want free healthcare but no tax bump so health insurance is a compromise.. when even just ACA was passed and premiums increased, democrats got destroyed at midterms so Americans might be living in la la land.
You see funding of chatgpt as a panacea for progress.
I see funding of chatgpt as one of small part of a history where governments and industry fund basic science and moonshot programs, not to generate revenue, but to explore what is possible.
LLM funding is not aimed at improving our understanding of the world, it's aimed at making people reliant so that they may extract wealth through subscriptions for shareholders.
Americans don't get good healthcare and education because that's what they vote for, in elections and wallets. I am hopeful that that changes, but we shall see.
Why can't it both? Of course they are not gonna do it just because it improves the world and understanding but because there's an incentive to align money with progress. Even the vaccines initially were distributed to get monetary gains and as the government started subsiding it as well, it became cheaper to produce.. that's basic capitalism and markets and regulations 101, no human is that selfless to give it out for free and they shouldn't because it's their investment in time, money, effort etc. but we should strive to align the greed aspects with good outcomes.
No Americans get fat and don't have a personal responsibility to maintain their health.. no amount of free healthcare is gonna change that.. they vote for free healthcare, see their taxes raise, then vote against cz they don't see tradeoffs in life.. it's better to maintain better habits than rely on govt to subsidize bad behaviour. There should be some basic coverage for poor people but not too much to sustain irresponsibly
Funding for basic research is being slashed by the current administration. Our society is underinvesting in basic scientific research. And, AI will not fill the gap.
It's just because of this administration but future admins can revert it back and even then, i would expect the fund receivers themselves will eventually use AI so.
This the correct way to look at it. Just as the person spending 5 years working on this will have learnt many things which will be useful after this problem is solved, you have to factor in the training cost (sure it's only done "once", but that is the same for the person too once they jump on the next problem).
The model wouldn't not be able to solve this without all the training leading up to the actual execution, so counting only the tokens of the execution doesn't give the full picture.
ChatGPT has raised over 120 billion USD in funding.
For argument let’s just say we paid all the mathematicians 200k in salary from graduation till retirement. Say 40 years. That’s about 8 million. Let’s round that up to USD 10 million. We can see the future and pay to raise all the baby mathematicians.
For 100 billion that’s 10000 mathematician lifetimes. For 1 AI company _so far_.
Well, it's taken 4 billion years for life to evolve into humans to be able to do math. That's a lot of resources, right?
Likewise, LLMs also needed the same amount of evolution.
My point is that it's silly to make these comparisons on resources. A single SOTA trained LLM isn't just doing advanced math research. It's used by hundreds of millions or even billions daily for various tasks. It's just a tool humans invented.
You are misinformed on token amounts. As one example, z.ai [0] gives you somewhere in the ballpark of 150-300 Mtok/week, so 1-2x your amounts. Plus the model has a 1M window, and will be smarter.
Fair context. Still, those prices are at least somewhat subsidized by you selling your plaintext session data which is a non-starter for me. Especially when dealing with code or data that is under NDA and not mine to sell in exchange for cheaper tokens.
Also, when I want 1M context I have 118GB of usable vram on my Strix Halo, or I can combine 4 r9700s and have 128gb and can run 1M context models like laguna or deepseek, but in practice lots of smaller sessions is better for my workflow in most cases.
Qwen 3.8 27b is smart enough that all I want is to speed-max and paralell-max on that.
Sure privacy (or legality) considerations are valid, and depending on the subscription or the API you use, you may or may not get that.
But we are not talking about the same product anymore. Your $12k homelab does not provide the same product as a $20/month subscription (to say nothing of a $200/month one). And the trade-offs of the homely may be worth it (or necessary) for you, but it may not be worth it (or even be feasible) for others.
Well at the very least lets agree that if you are going to use cloud models, there are much better options than OpenAI and Anthropic prices for most people. One does not have to give Sam or Dario money.
Personally though, I would sooner trade my car for GPUs than let a third party be in control of the tools I use to do my job.
So instead of paying $20 or maybe 100$ a month for the equivalent amount of output, I can spend $10k on GPUs (plus a few more thousands for related hardware), then for ongoing electricity, and then have space to store these things.
You're right people don't need the subscription, they just need a far more privileged life. One were dropping $12k instead of $20-100 a month is an equivalent financial strain.
I don't think that math will work out. If you are okay with using only 7M tokens per day, and 7M from a small model, then you don't have very demanding needs. If you don't have demanding needs, I think you're unlikely to be willing to space $1200 ( plus $800 for the rest of the system) to run a local model, when for $10/month you can get an open code subscription (or api access) which won't train on your data (or something of open router).
I know that 7M a day is a pittance for my use cases, and given that you have multiple cards, it wasn't enough for your use case either.
Sure, I just add more cards as I need more throughput, and can combine up to four cards when I need 1M context on a smarter model, but in practice since 3.8 27b came out it is all I use.
Also, unless your sessions are end to end encrypted to a secure enclave, then they are living in plain text -somewhere- and privacy policies tend to change when money is left on the table, if blackhats do not get to the data and sell it first.
I really like the 3D version, but I strongly believe you need to consider the number of tokens required to complete a task, it heavily impacts the results for certain models that rely heavily on test time compute (Glm-5.3-flash is the newest example).
I think this is the most import graph on their page. I wish they would let us filter by intelligence, or pass rate, and then see this graph. This is the tradeoff that actually matters, cost/token or tok/s can be very misleading (take glm-5.3-flash as an example).
Update: After some backlash, Google has clarified that they will only ban your "Antigravity and/or Gemini CLI accounts," not your Google account. How very generous! Keep being tone-deaf then...
I disagree. The y-axis is some arbitrary intelligence score that we use as a proxy for performance on whatever our specific task happens to be. So it doesn't matter if a model is a 0, 1, or 20 along this axis, they are all useless for the tasks I want.
And as the complexity of your score increases, the cutoff goes up. We can quibble about where your personal cutoff is, but it aint 0.
reply