claude-sonnet-4-6): an ordinary day with an agent reading the repo, editing files and running tests. It added up to 405 million tokens and $167.24:
One line is worth pausing on. Almost the entire volume — 395 million tokens out of 405 — is cache reads, not fresh input. Claude caches the repeated context: the system prompt, the contents of open files, the conversation history. Reading it costs $0.30 per million instead of $3.00 — ten times less. Without caching, the same day would have run about $1,235. So $167 is already a heavily discounted number, not the sticker price.
The question from here is simple: what do those exact same tokens cost on DeepSeek?
The same day on DeepSeek V4.1 Flash
deepseek-v4.1-flash is the fastest and cheapest model in the DeepSeek line (deepseek-v4-flash costs the same). Same three lines at its prices:
$167.24 versus $6.99 — for the same tokens, that’s about 24× cheaper.
The same day on DeepSeek V4 Pro
deepseek-v4-pro is the top of the line, for the hardest tasks:
Here the gap is smaller — about 4.5×.
The result in one table
For comparison: the newer
claude-sonnet-5 is cheaper ($2 / $10 / $0.20 per MTok), and the same day on it would have cost about $111.50. Flash still undercuts it by roughly 16×.
Where the gap comes from
It comes from two things, and both favor DeepSeek. Price per token. Fresh input on Sonnet 4.6 is $3.00 per million; on Flash it’s $0.14, on Pro $1.32. Output is $15.00 against $1.20 on Flash and $3.96 on Pro. Output is where the saving shows most: the longer the model’s answers, the further the two bills diverge. Price of a cache read. Cache carries the load here — 97.5% of the tokens. And even on cache reads DeepSeek is cheaper: $0.30 per million on Sonnet 4.6 against $0.01 on Flash and $0.05 on Pro. For this day, cache reads cost $118.61 on Claude, $3.95 on Flash and $19.77 on Pro. Stack one on top of the other, and on an identical set of tokens the gap reaches about 24× on Flash and 4.5× on Pro.The math uses base model prices as of September 2026, before your account’s price group is applied, and prices change over time. Current per-token rates are on the Pricing page and at www.ruapi.ai — where each model also shows an effective price with cache, essentially the same cache math as this article. What matters here is the order of the difference between the models, not the exact dollar figures.
Which one to use
Cheaper isn’t the same as “always better.” It depends on the task.deepseek-v4.1-flash— start here. It’s the cheapest option: the same day of work costs about $7 instead of $167. Try it on your own tasks, agents like Claude Code included — if quality falls short somewhere, switch to Pro. (The olderdeepseek-v4-flashcosts the same.)deepseek-v4-pro— for the hardest tasks: complex code, large refactors, long reasoning. Still about 4.5× cheaper than Claude Sonnet 4.6.- Claude — when you need to read images (DeepSeek has no vision) or top quality on the hardest tasks. One RuAPI key covers everything, so the models happily coexist in a single project.
DeepSeek models are text-only: they don’t read images, diagrams or screenshots. If you need vision, look at Claude, Gemini or GLM-5V-Turbo. More on the line itself is on the DeepSeek API page.
How to switch
If your project already talks to RuAPI over the OpenAI-compatible protocol, there’s one line to change — the model ID. Thebase_url and key stay the same:
Next
DeepSeek API
Flash and Pro: how they differ and what they do.
Claude models
Lines, versions, and when Claude earns its price.
How billing works
Pay per token used, with a log of every request.
Quickstart
base_url, key and your first request.