DEEPSEEK-V4-FLASH-VISION-EXP -- THE CHEAP TIER GETS EYES, AND A FREE FILES API LANDS WITH IT (2026-08-21, vendor-primary): DeepSeek put **DeepSeek-V4-Flash-Vision-Exp** live on the API Platform, an **experimental** multimodal model that DeepSeek says **matches V4-Flash on text** (agents, reasoning, world knowledge) while adding image input. Call it with `model='deepseek-v4-flash-vision-exp'`. **The billing detail is the one that decides whether this is actually cheap: images are tokenized at up to 384 tokens each and billed at V4-Flash pricing** -- so vision is not a separate premium rate, it is ordinary Flash tokens, which is unusually favourable. Supports **Chat Completions, Messages and Responses**, and mixed text+image input via base64, external URL, or the new Files API. **Files API is live and free to use**: upload an image once and reference it by `file_id` across requests instead of re-uploading. **DeepSeek Harness 0.1.1** shipped the same day with out-of-the-box support. **TREAT THE HEADLINE BENCHMARK CLAIM AS VENDOR-STATED AND UNVERIFIED: DeepSeek says that on multimodal agent benchmarks the model 'makes a major leap over V4-Flash, bringing multimodal agent performance close to Opus-4.8'.** No benchmark table, no named suite, no third-party confirmation -- and 'close to' a competitor's flagship is precisely the sort of claim that needs independent replication before it belongs in a ranking. **Also note the 'Exp' suffix: DeepSeek is labelling this experimental, so it is not a stability guarantee.**
2026-08-21•DeepSeek (api-docs.deepseek.com/news/news260821, on-page '2026/08/21') -- fetched 2026-08-28 via curl
THE PEAK/OFF-PEAK PRICE RISE IS NOW EXECUTED AND LIVE -- CONFIRMED ON THE VENDOR RATE CARD (2026-08-16 16:00 UTC, re-verified 2026-08-17): the change we flagged on 8/13 as scheduled has landed. **DeepSeek's Models & Pricing page no longer shows the old $0.435/$0.87 standing rate at all** -- the peak/off-peak table is now the only rate card published, which is the confirmation that matters. **Live V4-Pro rates per 1M:** off-peak **$0.022 cache-hit / $0.66 cache-miss input / $1.98 output**; peak **$0.044 / $1.32 / $3.96**. **Live V4-Flash rates:** off-peak **$0.007 / $0.22 / $0.66**; peak **$0.014 / $0.44 / $1.32**. **Peak hours, verbatim from the page footnote: 'Peak hours are 01:00 - 04:00 and 06:00 - 10:00 UTC (all other hours are off-peak),'** with 'off-peak rates are half of the peak rates.' That is 7 peak hours in 24. **THE CORRECTION THIS PAGE HAS BEEN MAKING SINCE 8/13 NOW STANDS CONFIRMED AGAINST A LIVE RATE CARD, SO DO NOT LET A LATER SOURCE REVERSE IT:** the 'off-peak is half of peak' wording invites reporting this as a discount scheme, and it is not one. Measured against the $0.435/$0.87 price that was in force until 8/16, **V4-Pro off-peak input is up ~52% and off-peak output up ~128%**; at peak, input roughly triples and output is up ~355%. V4-Flash off-peak output is up ~136% from $0.28. **There is still no hour of the day at which the new pricing beats the old.** The practical consequence for anyone who budgeted DeepSeek as the cheap frontier option: the headline gap to Western frontier models has narrowed sharply, and scheduling work into off-peak windows now only avoids a surcharge on top of an already-raised base -- it does not recover the old price. Concurrency limits published alongside are unchanged at 500 (V4-Pro) and 2500 (V4-Flash)
2026-08-16•DeepSeek Models & Pricing (api-docs.deepseek.com/quick_start/pricing) -- live rate card fetched 2026-08-17 via Firecrawl render, since the pricing table is client-rendered and does not appear in the raw HTML
V4-PRO IS FINALLY GA (2026-08-13, vendor change log) -- AND THE ACCOMPANYING 'PEAK/OFF-PEAK' PRICING IS A PRICE **RISE**, NOT A DISCOUNT. THIS IS THE MOST MISREPORTABLE ITEM ON THIS PAGE, SO READ THE NUMBERS. **(1) The GA itself.** DeepSeek's change log states 'The GA release of DeepSeek-V4-Pro has been rolled out on the APP, Web, and API,' model version **DeepSeek-V4-Pro-0813**, calling convention unchanged (`deepseek-v4-pro`). This closes the watch item open since 7/31, when only V4-Flash graduated and DeepSeek said Pro would 'follow soon.' **Vendor-published GA benchmarks:** HLE (without/with tools) **42.7/60.0**, Terminal Bench 2.1 **87.9**, NL2Repo **61.5**, Cybergym **83.3**, DeepSWE **62.7**, Toolathlon-Verified **74.1**, Agents' Last Exam **25.7**, AutomationBench public **31.8**, plus internal sets DSBench-FullStack **71.1** and DSBench-Hard **67.2**. All first-party. **(2) The pricing change -- effective 16:00 UTC on 2026-08-16.** DeepSeek is moving to peak/off-peak billing with 'off-peak prices set at half the peak-hour prices.' **Peak hours are 01:00-04:00 and 06:00-10:00 UTC**; every other hour is off-peak (so ~7 of 24 hours are peak). New V4-Pro rates per 1M: **off-peak $0.022 cache-hit / $0.66 cache-miss input / $1.98 output**; **peak $0.044 / $1.32 / $3.96**. New V4-Flash rates: **off-peak $0.007 / $0.22 / $0.66**; **peak $0.014 / $0.44 / $1.32**. **NOW COMPARE TO WHAT YOU PAY TODAY:** V4-Pro is currently **$0.435 input / $0.87 output**, and V4-Flash **$0.14 / $0.28**. So **even the cheapest new off-peak rate is above the current standing price** -- V4-Pro off-peak input rises ~52% and output ~128%; at peak, input roughly triples and output rises ~355%. V4-Flash off-peak output rises ~136%. **There is no time of day at which the new pricing is cheaper than the old.** Framing it as a discount scheme (which the 'off-peak is half of peak' wording invites) is wrong; the correct read is a substantial across-the-board increase with a time-of-day surcharge layered on top. **This also supersedes our own long-standing note that the 75%-off V4-Pro rate had become the permanent standing price** -- that was true from 2026-05-26 until this change, and it ends on 8/16. **(3) Other GA changes:** thinking effort is now three levels (**low / high / max**) on both Pro and Flash, and the API 'natively supports the OpenAI Responses API format and is specifically adapted for Codex' with a one-click config script
2026-08-13•DeepSeek API change log (api-docs.deepseek.com/updates, entry dated 2026-08-13) and Models & Pricing (api-docs.deepseek.com/quick_start/pricing, carrying both current and 2026-08-16 rate tables) -- both fetched 2026-08-13
**[SUPERSEDED 2026-08-13 -- V4-Pro reached GA on that date; see the 2026-08-13 entry above. Kept because it documents that the aggregators calling V4 'GA' in July were wrong at the time.]** V4-FLASH OFFICIAL RELEASE / PUBLIC BETA -- AND V4-PRO IS STILL NOT GA (2026-07-31, vendor changelog; CORRECTS WIDESPREAD AGGREGATOR REPORTING): DeepSeek's own change log states, verbatim: 'The official release of the DeepSeek-V4-Flash API is now in public beta... The DeepSeek-V4-Pro API and the APP/WEB models are unchanged. **The official release of DeepSeek-V4-Pro will follow soon.**' So only FLASH graduated -- aggregators claiming 'DeepSeek V4 went GA in mid/late July' are wrong, and V4-Pro remains Preview. No API change needed: set the model name to `deepseek-v4-flash`. **Vendor-published V4-Flash-0731 benchmarks** (agent-focused, stated as 'far exceeding V4-Pro-Preview'): Terminal Bench 2.1 **82.7**, NL2Repo **54.2**, Cybergym **76.7**, DeepSWE **54.4**, Toolathlon verified **70.3**, Agent Last Exam **25.2**, Automation Bench (public) **25.1**, plus internal sets DSBench-FullStack 68.7 and DSBench-Hard 59.6. Caveats DeepSeek states itself: code-agent numbers were run with the **DeepSeek Harness minimal mode (which it says is still 'to be released soon')** at max effort, topp=0.95, temperature=1.0 -- so they are not straightforwardly reproducible yet, and all figures are first-party. Architecture note: **V4-Flash-0731 keeps the same architecture and size as V4-Flash-Preview and was only re-post-trained.** It also now **natively supports the Responses API format and is specifically adapted for Codex**
2026-07-31•DeepSeek API change log (api-docs.deepseek.com/updates, fetched 2026-08-03)
CORRECTION -- THE 2x PEAK-HOUR PRICING HAS NOT ACTUALLY STARTED (re-verified 2026-08-03): our earlier entry described time-of-day pricing as shipping alongside the mid-July V4 release. It has not. DeepSeek's pricing page still frames it in the future tense -- it 'will soon adopt a peak/off-peak pricing policy', with peak hours 09:00-12:00 and 14:00-18:00 Beijing time billed at 2x, and explicitly: '**The effective date will be subject to the official announcement.**' No such announcement has been published as of 2026-08-03. Treat 2x peak pricing as ANNOUNCED-BUT-NOT-IN-EFFECT; current rates are the standing ones shown in the pricing table above
2026-08-03•DeepSeek pricing docs (api-docs.deepseek.com/quick_start/pricing, re-checked 2026-08-03)
LEGACY API ALIASES RETIRE 2026-07-24 (vendor-primary, HARD deadline): **`deepseek-chat` and `deepseek-reasoner` will be fully retired and inaccessible after July 24, 2026, 15:59 UTC.** The aliases currently route to deepseek-v4-flash (non-thinking/thinking respectively). Migration is a one-line change: keep base_url, update `model` to `deepseek-v4-pro` or `deepseek-v4-flash` -- but note the gotcha that `deepseek-reasoner` maps to FLASH-tier thinking, not V4-Pro, so 'upgrading' to Pro changes both cost and behavior. Any production code still pinned to the legacy aliases breaks on the 24th
2026-07-24•DeepSeek API docs (api-docs.deepseek.com/news/news260424)
Issues here are sourced from our editorial sweeps, not real-time telemetry. Newer issues may exist.