What My AI Agents Actually Cost: DeepSeek vs the Big APIs
- Category
- AI & Local LLM
- Posted
- July 11, 2026
- Updated
- September 16, 2026
- By
- Jacob Lloyd — written with AI assistance, post-project
- Read time
- 10 min read
In plain terms: I ran eight AI agents around the clock for three months. Cash out was $242 in June, $225 in July and $385 in August. Most of it was flat subscriptions. The metered DeepSeek part was $24, $6 and $102, and it only stays that small because about 93% of the input comes from the provider's cache.
For three months I ran eight always-on AI agents on a mix of metered APIs and flat subscriptions. Here is what left my bank account each month, what drove it, and what the same work would have cost elsewhere. August came to $385.22, the most of the three: DeepSeek raised its prices on August 16, and I moved to a bigger Claude plan and added a MiniMax plan in the same month.
tl;dr
- Cash out per month: $242.48 (June), $224.89 (July), $385.22 (August). All figures include tax and exclude a few dollars of Kimi K3 review calls.
- Subscriptions are most of the bill. Metered DeepSeek was $24.08, $6.49 and $102.34. At plan prices the two subscriptions now cost $278.46 a month before a single metered token.
- Cache-hit ratio is what keeps metered cheap: 86.7%, 95.1% and 93.2%. Without the cache, August's DeepSeek input would have cost about 9× more.
- After the hike, DeepSeek is no longer the cheapest option for this workload. August's tokens would have cost about $90 on GPT-5.6 Luna against DeepSeek's actual $102. Claude Haiku 4.5 (about $420), Grok 4.6 (about $1,180) and Claude Opus 5 (about $2,100) cost far more.
The three months
| Month | DeepSeek tokens | DeepSeek (metered) | Subscriptions | Cash out |
|---|---|---|---|---|
| Jun 2026 | 651M | $24.08 | $218.40 (Claude Max 20x) | $242.48 |
| Jul 2026 | 323M | $6.49 | $218.40 (Claude Max 20x) | $224.89 |
| Aug 2026 | 1.69B | $102.34 | $282.88 (Claude $222.82 + MiniMax $60.06) | $385.22 |
Subscription prices include tax. Claude Max 20x is $200 a month before tax, which came to $218.40 for me; I don't have the June and July Claude invoices to hand, so those two rows use that plan price. August's Claude line is two real charges: $109.20 billed at the Max 5x price on August 16, and $113.62 on August 17 to move to Max 20x, after a $95.95 credit. That August 16 charge at the 5x price is the part I can't square with June and July being on 20x, and without those two invoices I can't say which tier they were really on. If they were 5x, June comes to $133 and July to $116.
Per day, August worked out to $12.43 for everything, or $3.30 for the metered DeepSeek layer. Spread over eight agents, that metered layer is about 41 cents per agent per day. That is an average; the busy agents cost more than the quiet ones.
What each provider did
| Provider | What it did | How it's paid | August |
|---|---|---|---|
| DeepSeek V4 Flash + V4 Pro | Default backend for the always-on agents: triage, scheduled jobs, glue automation, a morning brief | Metered | $102.34 |
| Claude (Max plan) | My own interactive coding in Claude Code: long refactors and anything that runs for more than about 15 minutes | Subscription | $222.82 |
| MiniMax (Monthly Max plan) | M3 for agent work, plus one Hailuo-02 video, both covered by the plan | Subscription | $60.06 |
| Kimi K3 (Moonshot) | Occasional second-opinion code review | Metered, billed separately | Under $5 (not in the totals) |
The split is deliberate. Traffic that runs all day with the same long prompts goes to the cheapest metered model with a good cache. Frontier work that I drive by hand goes on a flat plan, where token counts stop mattering.
June: building the stack
This was the first month. All eight agents came online, and their system prompts, tool definitions and histories kept changing as I built and rebuilt them. That is why June has the lowest cache-hit ratio of the three months, 86.7%. Claude Code ran hard: 412 sessions across 13 active days, most of them long, tool-heavy coding runs, all inside the flat plan.
| Provider | Model | Cached input | Uncached input | Output | Cost |
|---|---|---|---|---|---|
| DeepSeek | V4 (Flash + Pro, about ⅔ Flash) | 560.8M | 86.2M | 3.8M | $24.08 |
June: 651M DeepSeek tokens for $24.08, or 27.0M tokens per dollar. Cash out was $242.48 (Claude plan $218.40 + DeepSeek $24.08).
July: quiet month
The build-out was done and the agents settled into routines. Claude Code logged 297 sessions from July 1 to 11. Then the stack went quiet for nine days (July 20 to 28) with no API traffic at all. The prompts had stopped changing, so the cache stayed warm and the hit ratio reached 95.1%, the best of the three months.
| Provider | Model | Cost |
|---|---|---|
| DeepSeek | V4 Flash | $1.81 |
| DeepSeek | V4 Pro | $4.68 |
| Total | $6.49 | |
July: 323M DeepSeek tokens for $6.49, or 49.8M tokens per dollar, the best value of the three. Cash out was $224.89 (Claude plan $218.40 + DeepSeek $6.49).
August: more traffic, a price hike and two plan changes
Traffic picked up again: 24,434 DeepSeek requests, against about 5.8k in July. On August 16 DeepSeek's peak/off-peak pricing took effect. On August 17 I moved Claude to the Max 20x plan because my coding work needed more headroom. On August 30 I added a MiniMax Monthly Max plan for M3.
| Provider | Model | Cached input | Uncached input | Output | Requests | Cost |
|---|---|---|---|---|---|---|
| DeepSeek | V4 Flash | 771.2M | 40.1M | 19.3M | 14,626 | $27.09 |
| DeepSeek | V4 Pro | 771.9M | 72.0M | 11.5M | 9,808 | $75.25 |
| Metered total | $102.34 | |||||
| Subscription | Charged | Amount |
|---|---|---|
| Claude, billed at the Max 5x price | Aug 16 | $109.20 |
| Claude, move to Max 20x ($95.95 credit applied) | Aug 17 | $113.62 |
| MiniMax Monthly Max | Aug 30 | $60.06 |
| Subscriptions total | $282.88 | |
The MiniMax plan also covered 91.5M M3 input tokens (78.1M of them cached) and 2.1M output tokens in August, plus one Hailuo-02 six-second video. None of that was metered.
August: 1.69B DeepSeek tokens for $102.34, or 16.5M tokens per dollar, down from July's 49.8M because of the price hike. Cash out was $385.22.
DeepSeek’s August 16 price hike
DeepSeek announced peak/off-peak billing, effective August 16, 2026 at 16:00 UTC. Peak hours are 01:00–04:00 and 06:00–10:00 UTC, which is 10:00–13:00 and 15:00–19:00 in Japan. For me that is exactly when daytime agent traffic is heaviest. Off-peak costs half the peak rate.
| Rate ($ per M tokens) | Before Aug 16 | From Aug 16, off-peak | From Aug 16, peak | Increase |
|---|---|---|---|---|
| V4 Flash, cached input | $0.0028 | $0.007 | $0.014 | 2.5–5× |
| V4 Flash, uncached input | $0.14 | $0.22 | $0.44 | 1.6–3.1× |
| V4 Flash, output | $0.28 | $0.66 | $1.32 | 2.4–4.7× |
| V4 Pro, cached input | $0.003625 | $0.022 | $0.044 | 6–12× |
| V4 Pro, uncached input | $0.435 | $0.66 | $1.32 | 1.5–3× |
| V4 Pro, output | $0.87 | $1.98 | $3.96 | 2.3–4.6× |
These are the rates my August bill was charged at. DeepSeek has changed Flash again since then: on 2026-09-16 its pricing page lists $0.003 cached, $0.15 uncached and $0.60 output off-peak, with peak at double. Pro is unchanged. The same page now says peak applies Monday to Friday only.
The biggest increases hit cached input, which is where always-on agents spend most of their tokens. Pro's cached rate went up 6–12×.
What if the whole of August had been billed at one rate? I multiplied each model's cached, uncached and output counts from the August table by each column above. That gives $57.27 at the old rates, $114.15 at the new off-peak rates and $228.31 at the new peak rates. The real bill, $102.34, came in under the all-off-peak figure only because the first half of the month still ran at old rates. August 1–16 cost $23.45. August 17–31 cost $78.89, of which $54.67 was off-peak and $24.22 was peak.
30.7% of my post-hike traffic landed in peak hours. Traffic spread evenly over all seven days would put 29% in peak (7 of 24 hours). With peak limited to weekdays, it would be about 21% for August 17–31. Either way, I didn't avoid the hike by timing.
The same workload on other APIs
I priced August's DeepSeek token mix (1,543M cached input, 112M uncached input, 30.8M output) at each vendor's published rates, as listed on 2026-09-16. Cached tokens are charged at each vendor's cache-read rate. The table assumes every request stays under Grok's 200k and MiniMax's 512k long-prompt thresholds, above which those vendors charge more. Totals are rounded to the nearest $5.
| Model | Rates used, $ per M (input / cached / output) | August workload |
|---|---|---|
| OpenAI GPT-5.6 Luna | 0.20 / 0.02 / 1.20 (OpenAI) | ~$90 |
| DeepSeek V4 Flash + Pro (what I paid) | see table above | $102.34 |
| MiniMax M3, current half-price discount | 0.30 / 0.06 / 1.20 (MiniMax) | ~$165 |
| MiniMax M3, standard price | 0.60 / 0.12 / 2.40 | ~$325 |
| Anthropic Claude Haiku 4.5 | 1 / 0.10 / 5 (Anthropic) | ~$420 |
| xAI Grok 4.6 | 2 / 0.50 / 6 (xAI) | ~$1,180 |
| Moonshot Kimi K3 | 3 / 0.30 / 15 (Moonshot) | ~$1,260 |
| Anthropic Claude Opus 5 | 5 / 0.50 / 25 | ~$2,100 |
Anthropic also charges 1.25× the input rate when content is first written to the cache. If all 112M uncached tokens were cache writes, Haiku would come to about $450 and Opus to about $2,240.
At these rates GPT-5.6 Luna would have been slightly cheaper than what I paid DeepSeek. Before the hike DeepSeek was clearly cheaper: $57 against about $90. Every other model on the list costs at least half as much again, and the frontier models cost 10–20× more. I haven't moved: switching providers means retuning every agent's prompts and tools, and nothing here says Luna does the job as well.
For scale, my earlier Kimi K3 benchmark cost $0.531 for five small coding tasks, about 48× what DeepSeek Harness (dsh) spent on the same five with a warm cache, or 22× against its cold first pass. (I later gave the same five tasks to MiniMax M3.) That is why Kimi only does occasional reviews here.
What drives the cost
- Input volume, not output. In August DeepSeek read 1.655B input tokens and wrote 30.8M. Agents re-send their whole prompt and history on every turn, so input dominates.
- Cache-hit ratio. At 93%, most of that input is billed at the cached rate. Priced at August’s new off-peak rates, the input came to about $79 with the cache and would have been about $735 without it, roughly 9× more. Anything that changes the start of a prompt on every request breaks the cache: a timestamp, a random ID or a reordered tool list.
- Busy days. July’s nine idle days pulled that month’s metered bill down to $6.49. Metered cost follows how much the agents actually do.
- The rate card. The same August traffic costs $57 or $228 depending only on which DeepSeek rates apply.
- Fixed plans. Subscriptions don’t move with usage. For me they were 90% of June’s bill and 97% of July’s.
When a subscription beats metered pricing
A plan pays off when the tokens you would put through it cost more than the plan at metered rates. Two examples from this stack:
- MiniMax Monthly Max ($60.06 with tax). The M3 tokens it covered in August would have cost about $22 at MiniMax’s standard rate, or about $11 at the current discount. At a cache mix like mine, M3 costs about 24 cents per million tokens at the standard rate, so the plan breaks even at around 250M tokens a month (500M at the discount). The plan was only added on August 30, so a full month will tell me whether it clears that bar.
- Claude Max 20x ($218.40 with tax). I don’t have token counts for my Claude Code sessions, so I can’t give a break-even. For scale, Opus 5 at API rates would have cost about $2,100 for August’s agent traffic. That is why the always-on agents stay on a cheap metered model and only my own hands-on coding goes through the flat plan.
A decision rule
- Always-on agent traffic goes on the cheapest metered model you trust with tools, and you keep an eye on its cache-hit ratio. If it falls well below the 87–95% range I saw, fix your prompts before you change provider.
- Heavy interactive work goes on a subscription once the API-rate equivalent of a typical month is clearly more than the plan price.
- Price your real token mix, not the headline rate. Cached, uncached and output tokens have very different prices, and the ranking can flip: DeepSeek beat Luna before August 16 and lost to it after.
- Batch and scheduled work goes into off-peak hours if your provider has them. At DeepSeek’s new rates, all-peak costs twice as much as all-off-peak.
- Recheck prices every month. Between August and mid-September DeepSeek raised its rates and then lowered the Flash rates again.
Related: the agent stack this bill pays for, wiring DeepSeek into that stack, the Kimi K3 benchmark, DeepSeek Harness (dsh), and MiniMax M3 as a coding agent.