skip to content
The Weighted Average

Models & Open Source

600 Screenshots Cost 5 Cents on DeepSeek's Vision API

DeepSeek's V4-Flash-Vision caps any image at 384 tokens and allows 600 per request, putting a maxed-out visual agent call at about $0.05 off-peak.

A close-up of a camera lens against a blurry background
A close-up of a camera lens against a blurry background. Photograph by Agence Olloweb

DeepSeek’s new multimodal endpoint prices vision as a rounding error. V4-Flash-Vision-Exp bills at most 384 tokens per image regardless of source resolution and accepts up to 600 images in a single request, according to DeepSeek’s vision API documentation. At the model’s off-peak cache-miss input rate of $0.22 per million tokens, published on DeepSeek’s models and pricing page, a fully loaded 600-image call costs 230,400 tokens, or about $0.051 in image input — five cents to put an entire screenshot corpus in front of a model. At peak hours, which DeepSeek defines as 01:00 –04:00 and 06:00 –10:00 UTC on weekdays, the same request costs $0.101.

Compare that to the price band the rest of the market sets for text alone. OpenAI’s cheapest listed frontier tier, gpt-5.6-luna, charges $0.20 per million short-context input tokens on the OpenAI API pricing page, while its flagship gpt-5.6-sol charges $4.00. DeepSeek’s vision variant charges the same rate as its text sibling — vision is not a premium SKU here, it is a free feature on an already cheap model. That pricing decision, not the benchmark score, is the thing worth reacting to.

Why the token cap is the product decision

Most vision pricing punishes resolution. DeepSeek’s does the opposite: the model normalizes images to roughly 800×800 pixels depending on aspect ratio before processing, an optional detail field downscales to 512×512 when fine detail is unnecessary, and the billing ceiling holds at 384 tokens per image either way. Maximum edge length is 8,192 pixels, dropping to 4,096 once a request carries 15 or more images. For a screenshot-driven agent — the dominant visual workload in engineering — the practical effect is that a 4K capture and a thumbnail cost the same.

The second cost lever is the new Files API, which DeepSeek says is free and lets a developer upload an image once and reference it by ID across many requests, with a 64 MiB per-file limit against 32 MiB for publicly-accessible URLs. Agent loops re-send the same context constantly; paying zero to keep a reference image resident removes an entire category of bandwidth waste. Combined with cached-input pricing at $0.007 per million tokens off-peak, the marginal cost of a stable visual context approaches nothing.

On capability, DeepSeek claims the vision variant approaches and sometimes beats Opus 4.8 on its own internal multimodal agent benchmarks, per The Decoder’s report on the release and the company’s launch note on X. Those are vendor benchmarks on vendor tasks, and the “Exp” in the model name is doing work: this is an experimental endpoint, not a supported production tier. The model also plugs into OpenAI’s Chat Completions and Responses APIs plus Anthropic’s Messages endpoint, so the integration cost of a bake-off is close to zero.

Who should test it, and what breaks the case

One more number is worth computing before the trial. At 384 tokens per image, a browser agent that captures one screenshot per step and runs 50 steps per task consumes 19,200 image tokens per task, or $0.0042 off-peak in visual input. A million such tasks costs about $4,200 in image tokens. That is the scale at which vision stops being a feature flag and becomes an architecture: when looking at the screen is cheaper than parsing the DOM, agents will look at the screen.

The teams with the clearest reason to run an evaluation this week are the ones burning frontier tokens on screenshot triage: browser agents reading rendered pages, CI systems diffing visual regressions, document pipelines extracting text from scans, and support tooling classifying user-submitted images. Those are high-volume, low-nuance workloads where a $0.051 ceiling per 600 images changes what is affordable to attempt. Route the hard visual reasoning to a frontier model and the bulk classification here, then measure the escape rate — the same routing discipline this paper applied when DeepSeek’s V4 Pro introduced a time-of-day price clock and off-peak scheduling became a real budget lever.

Three things break the case. The experimental label means no stability guarantee: DeepSeek reserves the right to adjust prices, and the peak/off-peak split already doubles cost during eight weekday hours, so a workload pinned to European business hours pays the higher rate by default. Data residency is the second constraint — this is a Chinese API endpoint, and for regulated workloads that is a compliance decision before it is a pricing one, which is why the grey-market Claude proxy economy covered in today’s edition matters as a cautionary structure rather than a template. Third, vendor-reported benchmark parity with Opus 4.8 is not third-party parity; until an independent multimodal agent evaluation lands, treat the capability claim as a hypothesis your own eval set should test.

The verdict: run a 500-image bake-off against your current vision provider this week, score accuracy on your real task rather than a benchmark, and price both at your actual peak-hour distribution. The compute economics are moving fast in both directions right now — server hardware is getting more expensive while inference gets cheaper, the tension at the center of today’s lead on Nvidia’s memory passthrough. Evidence that would change the verdict: an independent benchmark showing a materially worse escape rate than a frontier model, or DeepSeek retiring the experimental endpoint before it stabilizes.

Sources