What My AI Agents Actually Cost: DeepSeek vs the Big APIs

Posted
July 11, 2026
Updated
August 14, 2026
By
Jacob Lloyd β€” written with AI assistance, post-project
Read time
8 min read

In plain terms: A look at what running AI assistants around the clock actually costs. Two months of heavy use came to about $31 on a budget AI service β€” just $6.49 for July. That service is raising its prices on August 16, so the same month of use will soon cost roughly two to four times more, and OpenAI's cheapest new tier becomes the least expensive option for this kind of workload. The trick that keeps any of it cheap is caching β€” most of the text sent to the AI is repeated, and repeated text is billed at a steep discount.

My home agent stack β€” eight named agents running 24/7 on DeepSeek's API β€” has now finished two full months. I pulled the July usage exports and did the math again: what did it really cost, and what would the identical workload cost on the big APIs?

tl;dr

  • July 2026, full month: $6.49 total ($1.81 flash + $4.68 pro) across 22 active days. June was $24.08; the two months together, $30.57.
  • Volume: 323M tokens across ~5.8k requests in July; ~974M tokens across the two months.
  • The kicker: 95% of July's input tokens were cache hits β€” that ratio, not the sticker price, is what makes always-on agents affordable.
  • Same workload elsewhere (estimated): July: roughly $135–$300 on the flagships, $12–$85 on their cheap tiers β€” and OpenAI's new GPT-5.6 Luna is the cheapest non-DeepSeek option at ~$12 (June: $33–$915).
  • Heads-up (added 2026-08-14): DeepSeek raises prices on August 16 with peak/off-peak billing. My July workload would have cost ~$14–$27 at the new rates instead of $6.49 β€” which makes GPT-5.6 Luna, at ~$12, the new cheapest option. Old-vs-new breakdown below.

The real numbers

July’s full-month volume: 305.5M cached-input tokens, 15.6M cache-miss input tokens, 1.9M output tokens β€” 323M tokens, for $6.49 ($1.81 on flash, $4.68 on pro). June was the heavy month at ~560.8M cache hits, ~86.2M cache misses, and ~3.8M output tokens for $24.08. Together that’s ~866M cached, ~102M missed, ~5.7M output β€” about 974M tokens and $30.57 over two months. July also had a nine-day quiet stretch (the 20th–28th) with zero API traffic; across the 22 days it did run, the average was $0.30/day.

That 95% cache-hit rate (up from 87% in June) isn’t luck β€” it’s the natural shape of agent traffic. Every turn re-sends the same system prompt, tool definitions, and conversation history, and providers bill that repeated prefix at a tiny fraction of the normal input price. DeepSeek charges $0.0028–$0.003625 per million cached tokens (through August 15 β€” see the next section); without caching, July’s input bill alone would have been ~14–19Γ— higher.

DeepSeek’s August 16 price increase: old vs new

Twelve days after this post’s last update, DeepSeek warned developers of a significant price increase, and the pricing page now has the details: effective August 16, 2026 at 16:00 UTC, the API moves to peak/off-peak billing, with off-peak at half the peak rate. Peak hours are 01:00–04:00 and 06:00–10:00 UTC β€” in Japan that’s 10:00–13:00 and 15:00–19:00 JST, i.e. exactly when a household’s daytime agent traffic is heaviest.

Here’s every rate, old vs new:

Rate ($/M tokens)OldNew off-peakNew peakIncrease
v4 flash β€” cached input$0.0028$0.007$0.0142.5–5Γ—
v4 flash β€” input (miss)$0.14$0.22$0.441.6–3.1Γ—
v4 flash β€” output$0.28$0.66$1.322.4–4.7Γ—
v4 pro β€” cached input$0.003625$0.022$0.0446–12Γ—
v4 pro β€” input (miss)$0.435$0.66$1.321.5–3Γ—
v4 pro β€” output$0.87$1.98$3.962.3–4.6Γ—

The steepest hikes land exactly where always-on agents live: cached input. Pro’s cache-hit rate rises 6–12Γ—, flash’s 2.5–5Γ— β€” the very line item that made this stack cheap.

What my actual months would have cost at the new rates β€” same token mixes as above, same model split (July ran roughly 50/50 flash/pro by volume; June about two-thirds flash), with three timing scenarios since the new billing depends on when traffic runs:

MonthActual (old rates)New: all off-peakNew: spread evenly 24/7New: all peak
July 2026$6.49~$14~$18~$27
June 2026$24.08~$43~$55~$85

Call it 2–4Γ— depending on when your traffic runs. Still cheap in absolute terms β€” the whole stack stays under $30/month even in the worst case β€” but it rearranges the comparison table below: at ~$12, GPT-5.6 Luna now undercuts even DeepSeek’s best-case all-off-peak price (~$14) for this exact workload. For the first time since this stack existed, the cheapest way to run it isn’t DeepSeek. Whether that’s worth a migration is a different question β€” the increase moves my bill by about ten dollars a month, and shifting cron-driven jobs into off-peak windows (peak is only 7 of 24 hours) claws a chunk of it back.

Same workload, other providers

Apples-to-apples estimate: each month’s exact token mix priced on the same current published rates (all verified 2026-08-14 β€” DeepSeek, Anthropic, OpenAI, xAI, Moonshot). Cache hits billed at each provider’s cached-input rate; misses at the normal input rate (Anthropic charges a 1.25Γ— cache-write premium on misses, which is included). The DeepSeek row shows the pre-August-16 rates my bills were actually charged at β€” the new rates are in the section above.

Provider / modelCached in $/MInput $/MOutput $/MJune workloadJuly workload
DeepSeek v4 flash+pro (actual, old rates)$0.0028–0.0036$0.14–0.44$0.28–0.87$24.08$6.49
DeepSeek v4 flash+pro (new rates, from Aug 16)$0.007–0.044$0.22–1.32$0.66–3.96~$43–85~$14–27
OpenAI GPT-5.6 Luna$0.02$0.20$1.20~$33~$12
Anthropic Claude Haiku 4.5$0.10$1.00 (+write)$5.00~$185~$60
xAI Grok 4.3$0.20$1.25$2.50~$230~$85
OpenAI GPT-5.6 Terra$0.20$2.00$12.00~$330~$115
xAI Grok 4.5$0.30$2.00$6.00~$365~$135
Moonshot Kimi K3 (flagship)$0.30$3.00$15.00~$485~$165
xAI Grok 4.6 (flagship, Aug 12)$0.50$2.00$6.00~$475~$195
OpenAI GPT-5.6 Sol (flagship)$0.50$5.00$30.00~$825~$290
Anthropic Claude Opus 5 / 4.8 (flagship)$0.50$5.00 (+write)$25.00~$915~$300

Assumptions: numbers rounded to the nearest ~$5; June mix (560.8M hit / 86.2M miss / 3.8M out) and July mix (305.5M hit / 15.6M miss / 1.9M out); June estimates priced at today's rates (GPT-5.6 only launched July 9; Grok 4.6 on August 12); assumes each provider would achieve the same cache-hit ratio (realistic β€” the traffic shape is identical); ignores batch discounts, xAI's long-context surcharge, and β€” except for the DeepSeek new-rates row β€” time-of-day pricing. Two roster changes since the last update: xAI shipped Grok 4.6 as its new flagship and cut Grok 4.5's cached rate from $0.50 to $0.30, and Anthropic launched Claude Opus 5 at the same rates as Opus 4.8, so that row covers both.

The July 30 OpenAI price cuts moved the low end of this table more than the top, and DeepSeek's August 16 increase finishes the job. Luna's cached-input rate ($0.02/M) is roughly 6–7Γ— DeepSeek's old cached rate, which is why the same 95%-cache-hit workload landed at ~$12 on Luna against $6.49 on DeepSeek. Against the new DeepSeek rates ($0.007–0.044/M cached), Luna's ~$12 beats DeepSeek's ~$14–$27 outright β€” the caching discount that kept the gap in DeepSeek's favor is exactly what got repriced.

The honest part: quality differs

This is not “DeepSeek equals Claude for 4% of the price.” The frontier models β€” Claude Opus, GPT-5.6, Grok 4.5 β€” are noticeably stronger on complex, long-horizon, agentic work. But routine agent chatter, message triage, drafts, scheduled jobs, and glue automation don’t need frontier intelligence, and that’s ~95% of what an always-on stack does. That’s exactly the split my setup exploits: cheap tokens for the busywork, frontier brains for the hard stuff β€” the same division of labor from when the agent stack broke and Claude Code fixed it.

The frontier half of my usage runs through Claude Code on the Max subscription β€” a flat monthly fee, so the heavy coding and debugging work never touches per-token billing at all. The DeepSeek dollars cover only the 24/7 agents.

The other bill: Claude Code on Max

For completeness, the frontier side, measured from this box’s local session history: June 2026: 412 Claude Code sessions across 13 active days. July 1–11: 297 sessions across 8 active days β€” about 709 sessions and ~314 MB of transcripts in six weeks. These are long, tool-heavy agentic coding runs, not one-off chats.

All of it is covered flat by the Claude Max subscription at $200/month. Subscription usage isn’t metered, so there’s no exact API-equivalent figure β€” but hundreds of long Opus-class agent sessions at list rates ($5/$25 per million tokens) would plausibly land in the four figures monthly. The honest total for the whole stack:

  • ~$200/mo flat β€” Max subscription, all the frontier coding and debugging
  • ~$15/mo metered β€” DeepSeek, the 24/7 agents ($30.57 across June + July; call it ~$30–45/mo once the August 16 rates land)

Gotchas

  • Protect your cache-hit ratio. A timestamp or random ID near the top of a system prompt invalidates the cached prefix on every request β€” at these ratios that’s a 10–30Γ— bill multiplier.
  • Output tokens barely matter here. 5.7M output vs ~968M input across the two months (1.9M vs 321M in July alone); agents read far more than they write.
  • Prices move fast β€” in both directions. OpenAI cut GPT-5.6 Terra 20% and Luna 80% on July 30; xAI shipped Grok 4.6 on August 12 and dropped Grok 4.5’s cached rate to $0.30; and DeepSeek β€” the whole premise of this post β€” raises prices on August 16. Recheck the pricing pages before planning around any number above.
  • Time-of-day billing rewards scheduling. Under DeepSeek’s new peak/off-peak split, peak is only 7 of 24 hours (01:00–04:00 and 06:00–10:00 UTC) and costs double. Cron jobs, nightly consolidation, and batch work can simply be scheduled off-peak; it’s the interactive daytime traffic that pays full freight.

← More AI & Local LLM