What My AI Agents Actually Cost: DeepSeek vs the Big APIs

Posted
July 11, 2026
Updated
September 16, 2026
By
Jacob Lloyd — written with AI assistance, post-project
Read time
10 min read

In plain terms: I ran eight AI agents around the clock for three months. Cash out was $242 in June, $225 in July and $385 in August. Most of it was flat subscriptions. The metered DeepSeek part was $24, $6 and $102, and it only stays that small because about 93% of the input comes from the provider's cache.

For three months I ran eight always-on AI agents on a mix of metered APIs and flat subscriptions. Here is what left my bank account each month, what drove it, and what the same work would have cost elsewhere. August came to $385.22, the most of the three: DeepSeek raised its prices on August 16, and I moved to a bigger Claude plan and added a MiniMax plan in the same month.

tl;dr

  • Cash out per month: $242.48 (June), $224.89 (July), $385.22 (August). All figures include tax and exclude a few dollars of Kimi K3 review calls.
  • Subscriptions are most of the bill. Metered DeepSeek was $24.08, $6.49 and $102.34. At plan prices the two subscriptions now cost $278.46 a month before a single metered token.
  • Cache-hit ratio is what keeps metered cheap: 86.7%, 95.1% and 93.2%. Without the cache, August's DeepSeek input would have cost about 9× more.
  • After the hike, DeepSeek is no longer the cheapest option for this workload. August's tokens would have cost about $90 on GPT-5.6 Luna against DeepSeek's actual $102. Claude Haiku 4.5 (about $420), Grok 4.6 (about $1,180) and Claude Opus 5 (about $2,100) cost far more.

The three months

MonthDeepSeek tokensDeepSeek (metered)SubscriptionsCash out
Jun 2026651M$24.08$218.40 (Claude Max 20x)$242.48
Jul 2026323M$6.49$218.40 (Claude Max 20x)$224.89
Aug 20261.69B$102.34$282.88 (Claude $222.82 + MiniMax $60.06)$385.22

Subscription prices include tax. Claude Max 20x is $200 a month before tax, which came to $218.40 for me; I don't have the June and July Claude invoices to hand, so those two rows use that plan price. August's Claude line is two real charges: $109.20 billed at the Max 5x price on August 16, and $113.62 on August 17 to move to Max 20x, after a $95.95 credit. That August 16 charge at the 5x price is the part I can't square with June and July being on 20x, and without those two invoices I can't say which tier they were really on. If they were 5x, June comes to $133 and July to $116.

$400 $300 $200 $100 $0 June July August $218 $24 $242 $218 $6 $225 $223 $60 $102 $385 DeepSeek metered MiniMax plan Claude Max plan
Monthly cash out, stacked by provider. The Claude plan is the base of every bar. July's DeepSeek slice is tiny because the stack sat idle for nine days. August's grew with more traffic and the new rates. Kimi K3 (under $5 a month) is not shown.

Per day, August worked out to $12.43 for everything, or $3.30 for the metered DeepSeek layer. Spread over eight agents, that metered layer is about 41 cents per agent per day. That is an average; the busy agents cost more than the quiet ones.

What each provider did

ProviderWhat it didHow it's paidAugust
DeepSeek V4 Flash + V4 ProDefault backend for the always-on agents: triage, scheduled jobs, glue automation, a morning briefMetered$102.34
Claude (Max plan)My own interactive coding in Claude Code: long refactors and anything that runs for more than about 15 minutesSubscription$222.82
MiniMax (Monthly Max plan)M3 for agent work, plus one Hailuo-02 video, both covered by the planSubscription$60.06
Kimi K3 (Moonshot)Occasional second-opinion code reviewMetered, billed separatelyUnder $5 (not in the totals)

The split is deliberate. Traffic that runs all day with the same long prompts goes to the cheapest metered model with a good cache. Frontier work that I drive by hand goes on a flat plan, where token counts stop mattering.

June: building the stack

This was the first month. All eight agents came online, and their system prompts, tool definitions and histories kept changing as I built and rebuilt them. That is why June has the lowest cache-hit ratio of the three months, 86.7%. Claude Code ran hard: 412 sessions across 13 active days, most of them long, tool-heavy coding runs, all inside the flat plan.

ProviderModelCached inputUncached inputOutputCost
DeepSeekV4 (Flash + Pro, about ⅔ Flash)560.8M86.2M3.8M$24.08

June: 651M DeepSeek tokens for $24.08, or 27.0M tokens per dollar. Cash out was $242.48 (Claude plan $218.40 + DeepSeek $24.08).

July: quiet month

The build-out was done and the agents settled into routines. Claude Code logged 297 sessions from July 1 to 11. Then the stack went quiet for nine days (July 20 to 28) with no API traffic at all. The prompts had stopped changing, so the cache stayed warm and the hit ratio reached 95.1%, the best of the three months.

ProviderModelCost
DeepSeekV4 Flash$1.81
DeepSeekV4 Pro$4.68
Total$6.49

July: 323M DeepSeek tokens for $6.49, or 49.8M tokens per dollar, the best value of the three. Cash out was $224.89 (Claude plan $218.40 + DeepSeek $6.49).

August: more traffic, a price hike and two plan changes

Traffic picked up again: 24,434 DeepSeek requests, against about 5.8k in July. On August 16 DeepSeek's peak/off-peak pricing took effect. On August 17 I moved Claude to the Max 20x plan because my coding work needed more headroom. On August 30 I added a MiniMax Monthly Max plan for M3.

ProviderModelCached inputUncached inputOutputRequestsCost
DeepSeekV4 Flash771.2M40.1M19.3M14,626$27.09
DeepSeekV4 Pro771.9M72.0M11.5M9,808$75.25
Metered total$102.34
SubscriptionChargedAmount
Claude, billed at the Max 5x priceAug 16$109.20
Claude, move to Max 20x ($95.95 credit applied)Aug 17$113.62
MiniMax Monthly MaxAug 30$60.06
Subscriptions total$282.88

The MiniMax plan also covered 91.5M M3 input tokens (78.1M of them cached) and 2.1M output tokens in August, plus one Hailuo-02 six-second video. None of that was metered.

August: 1.69B DeepSeek tokens for $102.34, or 16.5M tokens per dollar, down from July's 49.8M because of the price hike. Cash out was $385.22.

DeepSeek’s August 16 price hike

DeepSeek announced peak/off-peak billing, effective August 16, 2026 at 16:00 UTC. Peak hours are 01:00–04:00 and 06:00–10:00 UTC, which is 10:00–13:00 and 15:00–19:00 in Japan. For me that is exactly when daytime agent traffic is heaviest. Off-peak costs half the peak rate.

Rate ($ per M tokens)Before Aug 16From Aug 16, off-peakFrom Aug 16, peakIncrease
V4 Flash, cached input$0.0028$0.007$0.0142.5–5×
V4 Flash, uncached input$0.14$0.22$0.441.6–3.1×
V4 Flash, output$0.28$0.66$1.322.4–4.7×
V4 Pro, cached input$0.003625$0.022$0.0446–12×
V4 Pro, uncached input$0.435$0.66$1.321.5–3×
V4 Pro, output$0.87$1.98$3.962.3–4.6×

These are the rates my August bill was charged at. DeepSeek has changed Flash again since then: on 2026-09-16 its pricing page lists $0.003 cached, $0.15 uncached and $0.60 output off-peak, with peak at double. Pro is unchanged. The same page now says peak applies Monday to Friday only.

The biggest increases hit cached input, which is where always-on agents spend most of their tokens. Pro's cached rate went up 6–12×.

What if the whole of August had been billed at one rate? I multiplied each model's cached, uncached and output counts from the August table by each column above. That gives $57.27 at the old rates, $114.15 at the new off-peak rates and $228.31 at the new peak rates. The real bill, $102.34, came in under the all-off-peak figure only because the first half of the month still ran at old rates. August 1–16 cost $23.45. August 17–31 cost $78.89, of which $54.67 was off-peak and $24.22 was peak.

30.7% of my post-hike traffic landed in peak hours. Traffic spread evenly over all seven days would put 29% in peak (7 of 24 hours). With peak limited to weekdays, it would be about 21% for August 17–31. Either way, I didn't avoid the hike by timing.

The same workload on other APIs

I priced August's DeepSeek token mix (1,543M cached input, 112M uncached input, 30.8M output) at each vendor's published rates, as listed on 2026-09-16. Cached tokens are charged at each vendor's cache-read rate. The table assumes every request stays under Grok's 200k and MiniMax's 512k long-prompt thresholds, above which those vendors charge more. Totals are rounded to the nearest $5.

ModelRates used, $ per M (input / cached / output)August workload
OpenAI GPT-5.6 Luna0.20 / 0.02 / 1.20 (OpenAI)~$90
DeepSeek V4 Flash + Pro (what I paid)see table above$102.34
MiniMax M3, current half-price discount0.30 / 0.06 / 1.20 (MiniMax)~$165
MiniMax M3, standard price0.60 / 0.12 / 2.40~$325
Anthropic Claude Haiku 4.51 / 0.10 / 5 (Anthropic)~$420
xAI Grok 4.62 / 0.50 / 6 (xAI)~$1,180
Moonshot Kimi K33 / 0.30 / 15 (Moonshot)~$1,260
Anthropic Claude Opus 55 / 0.50 / 25~$2,100

Anthropic also charges 1.25× the input rate when content is first written to the cache. If all 112M uncached tokens were cache writes, Haiku would come to about $450 and Opus to about $2,240.

At these rates GPT-5.6 Luna would have been slightly cheaper than what I paid DeepSeek. Before the hike DeepSeek was clearly cheaper: $57 against about $90. Every other model on the list costs at least half as much again, and the frontier models cost 10–20× more. I haven't moved: switching providers means retuning every agent's prompts and tools, and nothing here says Luna does the job as well.

For scale, my earlier Kimi K3 benchmark cost $0.531 for five small coding tasks, about 48× what DeepSeek Harness (dsh) spent on the same five with a warm cache, or 22× against its cold first pass. (I later gave the same five tasks to MiniMax M3.) That is why Kimi only does occasional reviews here.

What drives the cost

  • Input volume, not output. In August DeepSeek read 1.655B input tokens and wrote 30.8M. Agents re-send their whole prompt and history on every turn, so input dominates.
  • Cache-hit ratio. At 93%, most of that input is billed at the cached rate. Priced at August’s new off-peak rates, the input came to about $79 with the cache and would have been about $735 without it, roughly 9× more. Anything that changes the start of a prompt on every request breaks the cache: a timestamp, a random ID or a reordered tool list.
  • Busy days. July’s nine idle days pulled that month’s metered bill down to $6.49. Metered cost follows how much the agents actually do.
  • The rate card. The same August traffic costs $57 or $228 depending only on which DeepSeek rates apply.
  • Fixed plans. Subscriptions don’t move with usage. For me they were 90% of June’s bill and 97% of July’s.

When a subscription beats metered pricing

A plan pays off when the tokens you would put through it cost more than the plan at metered rates. Two examples from this stack:

  • MiniMax Monthly Max ($60.06 with tax). The M3 tokens it covered in August would have cost about $22 at MiniMax’s standard rate, or about $11 at the current discount. At a cache mix like mine, M3 costs about 24 cents per million tokens at the standard rate, so the plan breaks even at around 250M tokens a month (500M at the discount). The plan was only added on August 30, so a full month will tell me whether it clears that bar.
  • Claude Max 20x ($218.40 with tax). I don’t have token counts for my Claude Code sessions, so I can’t give a break-even. For scale, Opus 5 at API rates would have cost about $2,100 for August’s agent traffic. That is why the always-on agents stay on a cheap metered model and only my own hands-on coding goes through the flat plan.

A decision rule

  1. Always-on agent traffic goes on the cheapest metered model you trust with tools, and you keep an eye on its cache-hit ratio. If it falls well below the 87–95% range I saw, fix your prompts before you change provider.
  2. Heavy interactive work goes on a subscription once the API-rate equivalent of a typical month is clearly more than the plan price.
  3. Price your real token mix, not the headline rate. Cached, uncached and output tokens have very different prices, and the ranking can flip: DeepSeek beat Luna before August 16 and lost to it after.
  4. Batch and scheduled work goes into off-peak hours if your provider has them. At DeepSeek’s new rates, all-peak costs twice as much as all-off-peak.
  5. Recheck prices every month. Between August and mid-September DeepSeek raised its rates and then lowered the Flash rates again.

Related: the agent stack this bill pays for, wiring DeepSeek into that stack, the Kimi K3 benchmark, DeepSeek Harness (dsh), and MiniMax M3 as a coding agent.


← More AI & Local LLM