The best DeepSeek alternatives, priced on one workload
DeepSeek repriced its API at 16:00 UTC on 16 August 2026, and the number in the headlines is the wrong one. The widely quoted +1,100% is real, but it describes a single line: V4-Pro input on a cache hit, from $0.003625 to $0.044 per 1M tokens. On a realistic cached agent workload the increase is about 2x off-peak and 4x at peak. What genuinely changed is that DeepSeek is no longer automatically the cheapest option, so this page prices one fixed workload against every vendor's three real rates: input on a cache miss, input on a cache hit, and output.
Pick by the property you were actually buying from DeepSeek:
- An unconditional licence → Z.ai: GLM-5.2 is plain MIT, no revenue cap, but its API flagship has no public weights.
- The lowest bill → MiniMax: $4.56 on the workload below against DeepSeek's $6.95 off-peak, on non-commercial weights.
- Data residency or an EU contract → Qwen Model Studio with six regional endpoints, or Mistral, the only vendor here publishing a training opt-out on every plan.
- Staying is often correct: nobody here is both cheaper and more open, and DeepSeek's own weights are still MIT.
Why teams look elsewhere
What actually changed on 16 August 2026
DeepSeek did not get worse at its job: V4-Pro is a stronger model than what it replaced, the cache still costs nothing to write, and the weights are still MIT. What changed is the price and, more quietly, the reference point the price is quoted against.
The anchor swapped, not just the number
On 3 August DeepSeek's own pricing footnote said peak hours would cost 2x the regular prices. The announcement that shipped on 13 August says off-peak is half of peak. Same 2:1 ratio, opposite reference point, and the old regular price is gone from the page. The shipped off-peak rate is roughly twice the old flat rate, and the final text never states a before-and-after figure.
The reward for prompt reuse was cut about 4x
V4-Pro's cache-hit rate used to be one one-hundred-twentieth of the miss rate. It is now one thirtieth. Caching still helps enormously and still carries no write surcharge and no hourly storage fee, which remains a real advantage over Anthropic and Google, but the discount that made DeepSeek unbeatable on repeated prompts is a quarter of what it was.
Off-peak is a timezone lottery
Peak is 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday: Beijing business hours. A US Eastern team never works during peak and pays roughly 2x. A team on Central European Summer Time has its 09:00 to 12:00 inside peak and pays double for its morning. There is also no batch endpoint at all, so off-peak scheduling is the only volume lever DeepSeek offers.
Two things still not written down
How reasoning tokens are billed is documented nowhere, while thinking mode is on by default at high effort, and in tool-calling agents the reasoning content must be resent every turn or the API returns 400. There is also no published free tier despite the deduction rules referencing a granted balance, and the privacy policy lists training on your inputs as a purpose with no published opt-out.
The shortlist
7 DeepSeek alternatives worth evaluating
The ranking answers one question: if DeepSeek's API stopped being the right choice tomorrow, in what order would you actually try these? Licence terms on the weights weigh heavily, because that is what most DeepSeek buyers came for, and every price below is the same normalized workload from the table further down. Every pick lists a real weakness.
The only vendor here whose weights carry a real OSI licence with no revenue cap, no user count and no attribution trigger: GLM-5.2 is tagged MIT on Hugging Face, with roughly 2.7M downloads. That is precisely the property most people bought DeepSeek for. Weakness, and it is a big one: it costs $18.56 on the workload below against DeepSeek's $6.95 off-peak, and the model its API actually sells as flagship, GLM-5.3, has no public weights at all while carrying the same rate card. First place on licence, and on nothing else.
The lowest real bill of any vendor here, first-party or reseller: $4.56 normalized, 34% below DeepSeek off-peak, at $0.30 input on a miss, $0.06 on a hit and $1.20 output, with no hourly cache-storage fee. Weakness: the weights behind that price are not usable commercially, since M3 requires prior written consent for any commercial self-host, and the only MiniMax model that was ever Apache-2.0, M1, has been delisted from the pricing, models and rate-limit pages while its weights remain on Hugging Face. There is no free LLM tier and no trial credits.
The residency answer, and the only vendor publishing two distinct cache mechanisms with distinct rates: six regional endpoints, in Beijing, Hong Kong, Singapore, Tokyo, Frankfurt and US Virginia, across three protocols, with an explicit cache priced at 10% of the input rate. Its self-host gate, a $50M MaaS threshold, is the loosest of any non-MIT licence on this page. Weakness: $23.20 normalized, more than 3x DeepSeek off-peak, and batch cannot be combined with the cache discount.
The EU-law option, and the only vendor here that publishes a training opt-out on every plan including Free, alongside a standing dollar-denominated allowance of $10 per month in API credits. Mistral Large 3 is genuinely Apache-2.0 and lands at $5.80 normalized, which undercuts DeepSeek off-peak. Weakness: the flagship you are actually pushed toward, Medium 3.5, is $23.40 and its weights are barred outright above $20M monthly revenue, so the cheap row and the open row are not the row Mistral sells hardest.
Ranked here rather than higher because of what its own, admirably clear three-rate card says: $15.00 per 1M output tokens, which is 3.8x DeepSeek V4-Pro at peak, on a model that always reasons with reasoning_effort defaulting to max. That combination lands at $46.80 normalized, the most expensive open-weight API here. Weakness beyond price: the flagship is excluded from the 60%-of-standard batch tier its three older siblings get, and the classic V1 series is labelled for full platform sunset on 31 August 2026.
The cheapest way to buy DeepSeek's flagship from anyone other than DeepSeek, and the only party in this market that discloses deployment precision at all: $11.82 normalized, priced at exactly 0.85x DeepSeek's peak on all three lines, with a per-endpoint quantization field that reveals fp8 or fp4 at eight of fifteen endpoints serving these weights. Weakness: it is a router, so unless you pin a provider the rate is a routed average and the endpoint you get may be the quantized one.
Earns the last slot on transparency rather than price: it publishes DeepSeek V4-Pro-0813 by exact model ID at $1.32 input, $0.13 cached and $3.96 output, prints the cache as a rate rather than hiding it behind a toggle, and is the only provider here with a published DeepSeek fine-tuning table. Weakness: $15.28 normalized, 2.2x DeepSeek off-peak for the same weights, and the deployment precision is still not stated on its own page.
Deliberately excluded, because each exclusion is itself a finding. Groq serves no DeepSeek V4 model at all and publishes no statically reachable per-token price for anything: its docs pricing path returns 404 and its model list renders a loading placeholder. Fireworks AI advertises the 0813 flagship on its marketing page while its docs table does not list it, and runs two unlabelled price columns per model. Nebius publishes no cache-read rate for any model, which prices these weights at $42.00 on the workload below. DeepInfra is the cheapest reseller at $12.00 and caps output at 16,384 tokens against DeepSeek's 384,000. Sail Research and BaseTen serve at fp4, a quarter of the published weight precision, at full-precision headline rates. OpenAI, Anthropic and Google are kept as a price ceiling in the table only: none publishes weights for the models being priced.
Side by side
DeepSeek alternatives compared on one workload
Prices as of August 2026, from each vendor's own pricing page, USD per 1M tokens. The workload is fixed for every row: 1,000 calls per month, 20,000 input tokens each at an 80% cache-hit rate, 2,000 output tokens each, which works out to 16 x cache-hit rate + 4 x cache-miss rate + 2 x output rate and totals 20M input and 2M output tokens. That formula is deliberately generous, because it charges nothing for writing a cache entry, which Anthropic bills at 1.25x to 2x and Google bills per hour. DeepSeek's own rows are the baseline, including the pre-increase price.
| Vendor and model | Monthly cost, same workload | Cache hit / miss / output | Weights licence | Free allowance | Region or jurisdiction |
|---|---|---|---|---|---|
| DeepSeek V4-Pro, before 16 Aug | $3.54 | 0.003625 / 0.435 / 0.87 | MIT, no threshold | None published | PRC law, single endpoint |
| DeepSeek V4-Pro, off-peak | $6.95 | 0.022 / 0.66 / 1.98 | MIT, no threshold | None published | PRC law, single endpoint |
| DeepSeek V4-Pro, peak | $13.90 | 0.044 / 1.32 / 3.96 | MIT, no threshold | None published | PRC law, single endpoint |
| DeepSeek V4-Flash, off-peak | $2.31 | 0.007 / 0.22 / 0.66 | MIT, no threshold | None published | PRC law, single endpoint |
| Z.ai GLM-5.3 | $18.56 | published card 1.40 in / 4.40 out | ✓ MIT on GLM-5.2, ✗ none for 5.3 | Not published as a dollar amount | PRC |
| MiniMax M2.x | $4.56 | 0.06 / 0.30 / 1.20 | ✗ non-commercial, written consent | None for LLMs | PRC, Singapore endpoint |
| Alibaba Qwen, Model Studio | $23.20 | cache at 10% of input rate | Gate at $50M MaaS revenue | Trial quota per model | 6 regions, incl. Frankfurt |
| Mistral Large 3 | $5.80 | Apache-2.0 model, three rates published | ✓ Apache-2.0 | $10 per month in credits | EU, opt-out on all plans |
| Mistral Medium 3.5 | $23.40 | flagship card | Barred above $20M monthly revenue | $10 per month in credits | EU, opt-out on all plans |
| Moonshot Kimi K3 | $46.80 | output 15.00, always reasons | Gate at $20M MaaS revenue | Not published | PRC |
| OpenRouter, DeepSeek V4-Pro | $11.82 | 0.85x DeepSeek peak on all three lines | MIT, DeepSeek's own weights | None | Routed, precision disclosed |
| Together AI, DeepSeek V4-Pro | $15.28 | 0.13 / 1.32 / 3.96 | MIT, DeepSeek's own weights | None | US, precision not stated |
| OpenAI gpt-5-nano | $1.08 | 0.005 / 0.05 / 0.40 | ✗ none for this model | None | US, regions on enterprise |
| Anthropic Claude Haiku 4.5 | $15.60 | 0.10 / 1.00 / 5.00 | ✗ none | None | US, EU and other regions |
| Google Gemini 3.5 Flash-Lite | $6.68 | 0.03 / 0.30 / 2.50, plus $1.00 per hour cache storage | ✗ none | Free tier on AI Studio | Global, regions available |
Numbers that do not fit in cells: the widely quoted +1,114% is the V4-Pro cache-hit input line at peak, $0.003625 to $0.044. At a 0% cache-hit rate the same DeepSeek V4-Pro workload costs $17.16 off-peak instead of $6.95, a 2.5x swing from one variable DeepSeek explicitly calls best-effort. DeepSeek publishes no batch endpoint, so no batch discount exists at any price; Moonshot discounts batch to 60% of standard but excludes its flagship, and Alibaba forbids combining batch with the cache discount. Concurrency on DeepSeek is 500 for V4-Pro and 2,500 for V4-Flash, account-wide, with free capacity expansion. Prices and licences in this category change monthly; check each vendor before committing. Compiled 24 August 2026.
Official pages: DeepSeek pricing · DeepSeek change log · Z.ai · MiniMax · Alibaba Model Studio · Mistral · Moonshot · OpenRouter · Together AI
A fair call
When DeepSeek is still the right choice
The honest version of this page has to say that a 2x increase on a very low base is still a low price, and that the thing most people came to DeepSeek for was never only the price.
DeepSeek is still right if…
- You want the weights behind the API. DeepSeek-V4-Pro-0813 and V4-Flash-0731 are the exact versions the API serves, published under plain MIT with no threshold. Nobody cheaper is more open.
- Your team works US hours. Every US business hour falls in off-peak, so your effective increase is about 2x rather than 4x, and weekends are off-peak everywhere.
- Your prompts repeat. The cache is automatic, needs no code change, has no write surcharge and no hourly storage fee, which is still not true of Anthropic or Google.
- You want to distil the outputs. DeepSeek's terms grant you rights in outputs and explicitly permit training other models, which OpenAI's and Anthropic's terms do not.
Look elsewhere if…
- Data must not leave a region. DeepSeek stores and processes in the PRC with no residency option: Qwen Model Studio or Mistral are the answers.
- You need a contractual training opt-out. DeepSeek's privacy policy lists model training as a purpose with no published opt-out; Mistral publishes one on every plan.
- Price is the whole decision. MiniMax at $4.56 and OpenAI gpt-5-nano at $1.08 both undercut DeepSeek off-peak on this workload.
- You depend on batch discounts or a documented reasoning-token billing rule: DeepSeek publishes neither.
Common questions
Common questions about DeepSeek alternatives
What is the best DeepSeek alternative in 2026?
It depends which property you were buying DeepSeek for. If it was the MIT licence on the weights, Z.ai is the only vendor here shipping plain MIT with no revenue or user threshold, on GLM-5.2. If it was the price, MiniMax is the cheapest credible API in the comparison at about $4.56 on a fixed 20M-input, 2M-output workload, against $6.95 for DeepSeek V4-Pro off-peak, but its M3 weights are non-commercial. If it was data residency, Alibaba's Model Studio publishes six regional endpoints and Mistral is the EU-jurisdiction option with a training opt-out on every plan. Nobody on this list is both cheaper than DeepSeek and more open than DeepSeek.
Did DeepSeek really raise its API prices by 1,100%?
That figure is arithmetically correct and it describes one line: V4-Pro input on a cache hit went from $0.003625 to $0.044 per 1M tokens at peak, which is 12.14x. It is the smallest item on almost any real invoice. Priced on a realistic cached agent workload of 20M input tokens at an 80% cache-hit rate and 2M output tokens, DeepSeek V4-Pro went from $3.54 per month to $6.95 off-peak or $13.90 at peak, so about 2x and about 4x. V4-Flash went from $1.16 to $2.31 or $4.62. Those are the numbers to plan with.
When are DeepSeek's peak hours, and does off-peak actually help?
Peak is 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday through Friday, which is Beijing business hours. For a US Eastern team every business hour falls in off-peak, so the effective increase is roughly 2x. For a team in Central European Summer Time, peak covers 08:00 to 12:00 local, so a European working morning is billed at double. Same published price, different bill, decided by timezone. Weekends are entirely off-peak.
Is DeepSeek still open source after the price increase?
Yes, and this is the part that did not change. The weights the API serves, DeepSeek-V4-Pro-0813 at 1.7T parameters and DeepSeek-V4-Flash-0731 at 304B, are both published on Hugging Face under a plain MIT licence with no revenue threshold, no user cap and no attribution trigger. Very few teams can realistically host 1.7T parameters, so the licence matters more as an exit option than as a deployment plan, but it is genuinely more permissive than Qwen's $50M MaaS gate, Kimi K3's $20M gate, Mistral Medium 3.5's $20M monthly ceiling or MiniMax M3, which needs prior written consent for any commercial self-host.
Is buying DeepSeek's weights from another provider cheaper?
No. Nine third-party providers converge on $1.32 input and $3.96 output per 1M tokens, which is exactly DeepSeek's own peak rate, so on the same workload a reseller costs 1.7x to 6.0x DeepSeek off-peak. The differences that matter are the ones not printed on their pricing pages: fp8 or fp4 quantisation at eight of fifteen endpoints in OpenRouter's index, a 16,384-token output cap at DeepInfra against DeepSeek's own 384,000, a 20x cache-read rate at SiliconFlow, no published cache rate at all at Nebius, and Groq serving no DeepSeek V4 model whatsoever while appearing in every listicle.
How does DeepSeek bill reasoning tokens?
It does not say. Neither the pricing page, nor the thinking-mode guide, nor the token-usage documentation states how chain-of-thought tokens are billed, while every frontier lab states it explicitly. That silence matters because thinking mode is enabled by default at high effort on V4, and because in a tool-calling agent DeepSeek requires the intermediate reasoning content to be passed back in every subsequent turn or the API returns a 400 error, which means the chain of thought is re-billed as input for the rest of the conversation.