DeepSeek Image Input: How to Use the New V4 Flash Vision Model

Date:

DeepSeek has added image understanding to its fastest and cheapest model. The new deepseek-v4-flash-vision-exp model accepts images alongside text through the same API endpoint developers already use for V4-Flash, with no price premium. For teams building AI agents that need to read screenshots, analyze charts, or process documents, the wait for a low-cost multimodal option from DeepSeek is over.

Announced on August 21, 2026, the model is labeled experimental, but it is production-ready for API use today. DeepSeek says it matches the text-only V4-Flash on reasoning, agent performance, and world knowledge while adding substantial visual understanding. The launch also includes a free Files API for image storage and reuse, plus out-of-the-box support in the newly released DeepSeek Harness 0.1.1.

Here is everything developers need to know about the new model, from pricing and token mechanics to the three input methods and practical agent use cases.

What Is DeepSeek V4 Flash Vision?

DeepSeek V4 Flash Vision is a multimodal extension of the existing DeepSeek V4 Flash model. While V4 Flash launched as a text-and-code model in April 2026, the vision variant adds the ability to accept image inputs in the same API call as text.

The model is not a separate product or a wrapper around a separate vision service. It is a single model that processes mixed text-and-image requests natively. When you send an image alongside a text prompt, the model tokenizes the image internally and reasons over both modalities together.

Key characteristics:

  • Model name: deepseek-v4-flash-vision-exp
  • Architecture: 284B total parameters / 13B active (mixture-of-experts)
  • Context window: 1M tokens
  • Image billing: Up to 384 tokens per image, at V4-Flash text pricing
  • Supported APIs: Chat Completions (OpenAI-compatible), Messages (Anthropic-compatible), Responses API
  • Status: Experimental (the “Exp” suffix)

The “Exp” label means DeepSeek considers this a preview release. The company may update or replace the model identifier as it gathers production feedback. For developers, this means the API is available now, but the model name could change in future updates.

DeepSeek’s V4 family has been evolving rapidly. The text-only V4-Flash uses a mixture-of-experts architecture with 284 billion total parameters but only 13 billion active per token, which is what makes it both fast and cheap. The vision variant inherits this architecture and its 1 million token context window, meaning you can send long conversations with many images without running into context limits.

Why Developers Have Been Waiting for This

Vision capability in DeepSeek models has been one of the most requested features in the developer community. When DeepSeek V4 launched in April 2026, it was explicitly text-only. Developers building agents on the V4 platform had to route image inputs through a separate vision model (often Gemini or Qwen) before passing the extracted text to DeepSeek. The image-generation side of the market has moved even faster: Microsoft’s MAI-Image-2.6 Preview, released the same month, took No. 1 on the Artificial Analysis image-editing board. OpenAI’s GPT Image 2.5 Flare and Sunburst, launched in September 2026, added six quality tiers and native transparent backgrounds to the competitive landscape while charging no API premium for editing inside the same model. That added latency, complexity, and cost to every multimodal workflow.

The demand was not about novelty. It was about economics. AI agents make dozens or even hundreds of model calls per task. A coding agent might inspect a screenshot, debug a UI regression, read an error log from an image, or verify a visual output. Each of those steps requires vision. If every image step routes to a model that charges 10x or 100x more per token, the cost compounds fast.

As one analysis from PicEditor noted, the significance of V4 Flash Vision is not that “a new model dropped.” It is that DeepSeek bolted vision onto the model family it already serves at commodity pricing. For agent workflows that make many calls per task, the per-call cost of adding vision determines whether the entire automation is economically viable.

DeepSeek also timed the release to coincide with two ecosystem updates:

  • DeepSeek Harness 0.1.1: The agent framework now supports deepseek-v4-flash-vision-exp out of the box. Agents built on Harness can send mixed text-and-image input without custom configuration.
  • Files API: A free image storage service that lets developers upload an image once and reference it by file_id across multiple requests, avoiding re-upload overhead.

The combination of a vision-capable model, a free file storage layer, and agent framework support creates a complete multimodal stack that did not exist in the DeepSeek ecosystem a week ago.

DeepSeek’s aggressive pricing strategy is not new. The company has consistently undercut Western competitors, a pattern explored in our analysis of why Chinese AI models are so much cheaper than OpenAI and Anthropic. Vision at text-model pricing is the latest extension of that strategy, and one Z.ai is now matching: its late-August 2026 release of GLM-5.3-Flash (the open-weight model previously known as Ox Alpha) entered the market at $0.075 per million input tokens with multimodal input included at no premium.

Benchmark Results: How Vision-Exp Compares

DeepSeek published an official 11-benchmark comparison alongside the launch, split into seven text-based agent evaluations and four multimodal-specific ones. The results are more nuanced than the headline “close to Opus 4.8” claim suggests.

Text-based agent evaluation

On text-only tasks, Vision-Exp matches or slightly improves on the text-only V4-Flash-0731 it descends from, confirming that adding vision did not degrade text performance:

Benchmark Vision-Exp V4-Flash-0731 Opus 4.8
Terminal Bench 2.1 83.9 82.7 85.0
NL2Repo 57.7 54.2 69.7
Cybergym 75.3 76.7 78.3
DeepSWE 59.3 54.4 58.0
Toolathlon-Verified 75.9 70.3 76.2
DSBench-Hard 63.6 59.6 71.7
AutomationBench 25.7 25.1 27.2

Vision-Exp beats Opus 4.8 on one text benchmark (DeepSWE, 59.3 vs 58.0) and comes within 1 to 2 points on Terminal Bench, Toolathlon, and AutomationBench. The gaps are wider on NL2Repo (12 points) and DSBench-Hard (8 points).

Multimodal agent evaluation

The multimodal benchmarks are where the vision capability shows its impact. Vision-Exp beats Opus 4.8 on two of four and ties closely on the other two:

Benchmark Vision-Exp V4-Flash-0731 Opus 4.8
ApexBench (Pass@1) 36.5 26.2 39.4
Agents’ Last Exam 27.3 25.2 25.7
Chartography 64.3 — 65.0
ZeroBench (Pass@5) 35.0 — 34.0

The jump from text-only V4-Flash to Vision-Exp is substantial on multimodal tasks: ApexBench improves by 10.3 points (26.2 to 36.5) and DeepSWE by 4.9 points (54.4 to 59.3). Vision-Exp wins Agents’ Last Exam and ZeroBench outright, while trailing ApexBench by 3 points and Chartography by less than 1 point.

Two important caveats apply to all of these numbers. First, the benchmarks were run using DeepSeek’s internal Harness Minimal Mode with specific settings (max reasoning tier, top_p=0.95, temperature=1.0). Independent third-party verification has not yet been published. Second, the comparison target is Claude Opus 4.8, not the newer Opus 5, so the competitive landscape may have shifted since these were run.

Pricing: Image Understanding at Text-Model Rates

DeepSeek is billing vision requests at the same rate as text-only V4-Flash calls. Images are converted into tokens (up to 384 per image) and added to the text token count for billing. There is no multimodal surcharge.

Here is the current V4-Flash pricing structure, which applies to both the text-only and vision models:

Pricing Tier Input (per 1M tokens) Output (per 1M tokens)
Off-peak, cache miss $0.22 $0.66
Off-peak, cache hit $0.0044 $0.66
Peak, cache miss $0.44 $1.32
Peak, cache hit $0.0088 $1.32

Peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC. Off-peak rates apply at all other times.

DeepSeek recently restructured its V4 pricing to add peak and off-peak tiers, replacing the original flat rates. Those rates changed again on September 10 with the V4.1 Flash launch and price cut, which also routes Pro traffic to the cheaper Flash model. Even with the increase, V4-Flash remains significantly cheaper than comparable Western models. A broader comparison of LLM API pricing across all major providers shows DeepSeek consistently at the bottom of the cost curve.

How does this compare to other vision-capable models? The gap is substantial:

Model Input (per 1M tokens) Output (per 1M tokens) KV-cache tokens per ~800×800 image
DeepSeek V4 Flash $0.22 (off-peak) $0.66 ~90
GPT-5.6 Luna $0.20 $1.20 ~765
Gemini 3.6 Flash $1.50 $7.50 ~1,100
Claude Sonnet 4.6 $3.00 $15.00 ~870
Claude Opus 4.8 $5.00 $25.00 ~870

Image token counts are approximate and vary by resolution. Competitor rates from official pricing pages as of mid-2026.

The KV-cache column is the key number. DeepSeek encodes an 800×800 image into roughly 90 tokens, while Claude uses approximately 870 and Gemini around 1,100. Combined with already-low text pricing, the per-image cost through DeepSeek’s vision API is roughly 1/10th to 1/170th of competing providers, depending on which model you compare against.

Visual representation of cost comparison between AI vision models showing progressive price scaling
Visual comparison: DeepSeek’s per-image cost scales to a fraction of competing vision models (Credit: Intelligent Living)

For agent workflows that process many images per session, this difference is not marginal. It is the difference between a viable production pipeline and one that breaks the budget after a few hundred requests.

How Image Tokenization Works

When you send an image to deepseek-v4-flash-vision-exp, the model does not store the raw pixels. It converts the image into tokens through an automatic resize-and-encode pipeline:

  1. Small images (below roughly 384×384 pixels) are scaled up while preserving aspect ratio.
  2. Larger images are scaled down so the total pixel count approximates an 800×800 image.
  3. After resizing, each image consumes up to 384 tokens, regardless of original size.

This means a 2000×2000 photo and a 5000×5000 photo cost the same in tokens after processing. There is no benefit to pre-shrinking images before sending them, unless you want to reduce upload bandwidth.

When a request contains multiple images, each image is tokenized independently under the same rule. The 384-token cap applies per image, not per request.

Diagram showing how an image is converted into tokens through automatic resizing and encoding
How DeepSeek tokenizes images: automatic resizing ensures consistent token costs regardless of original resolution (Credit: Intelligent Living)

Three Ways to Send Images to the API

DeepSeek supports three input methods, all using the standard content-array format familiar from OpenAI’s API. The base URL for all methods is https://api.deepseek.com.

1. Base64-encoded image (inline)

Encode the image as a base64 string and embed it directly in the request. This is the simplest option for local files and small images. The encoded data counts toward the 48 MiB request body limit.

import base64
from openai import OpenAI

client = OpenAI(
    api_key="<DeepSeek API Key>",
    base_url="https://api.deepseek.com"
)

with open("image.jpg", "rb") as f:
    b64 = base64.b64encode(f.read()).decode("utf-8")

response = client.chat.completions.create(
    model="deepseek-v4-flash-vision-exp",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "What is in this image?"},
            {"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{b64}"}}
        ]
    }]
)

print(response.choices[0].message.content)

2. External image URL

Pass a publicly accessible HTTP(S) link and the model downloads the image for you. The URL must be at most 8,192 characters, the image file may be at most 32 MiB, and the download must complete within 60 seconds.

response = client.chat.completions.create(
    model="deepseek-v4-flash-vision-exp",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "Describe this image."},
            {"type": "image_url", "image_url": {"url": "https://example.com/image.jpg"}}
        ]
    }]
)

3. Files API reference

Upload an image once with the Files API, then reference its file_id in subsequent requests. This is the best option for images used across multiple requests or images larger than 32 MiB. Files API images may be up to 64 MiB each.

response = client.chat.completions.create(
    model="deepseek-v4-flash-vision-exp",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "What is in this image?"},
            {"type": "file", "file_id": "file-api-xxxxxxxxxxxxxxxx"}
        ]
    }]
)

The model also supports the Anthropic-compatible /messages endpoint (base URL: https://api.deepseek.com/anthropic) and the OpenAI Responses API, where images are carried in input_image content parts.

Supported Formats and Technical Limits

The model accepts four image formats. Format detection is based on actual file content, not the filename or declared MIME type.

Specification Value
Supported formats JPEG, PNG, GIF, WebP
Max single image size (base64 / external URL) 32 MiB
Max single image size (Files API) 64 MiB
Max images per request 600
Max total image size per request 64 MiB (200 MiB with Files API images)
Max image dimension 8192 px per side (drops to 4096 px with 15+ images)
External URL length limit 8,192 characters
Request body size limit 48 MiB

A few restrictions to note:

  • Images are accepted in user messages only. Sending images in system or assistant messages returns a 400 error.
  • Only deepseek-v4-flash-vision-exp accepts images. Sending images to other DeepSeek models returns a 400 error.
  • The model supports a detail field for image_url inputs: low (downscaled to 512×512, cheaper), high / original (keeps original resolution), and auto (currently defaults to original).

Agent Use Cases: Where Vision Changes the Game

The most impactful applications of V4 Flash Vision are not simple “describe this photo” queries. They are agent workflows where visual perception is one step in a multi-step automation.

AI agent observing multiple screens including a web browser, code editor, document viewer, and chart display
V4 Flash Vision enables agents to observe and act on visual information across multiple interfaces (Credit: Intelligent Living)

Before this release, agents built on the DeepSeek platform were blind to visual information. A coding agent could edit code but could not inspect a screenshot of a UI bug. A research agent could parse text documents but could not read charts embedded in PDFs. A QA agent could run test scripts but could not verify visual output.

V4 Flash Vision closes these gaps at a price point that makes multi-step visual reasoning affordable. Here are concrete use cases developers are already targeting:

  • UI debugging and visual regression testing: Agents can inspect screenshots of web interfaces, compare them to expected designs, and flag layout issues without routing to a separate vision model.
  • Document and chart extraction: Reading text from screenshots, extracting data from charts, parsing invoices and receipts. The model’s low per-image cost makes high-volume document processing economically viable.
  • Browser automation: Agents that navigate web pages can now “see” the current page state, identify buttons and form fields, and decide the next action based on visual context.
  • Code output verification: After generating code that produces a visual output (a chart, a UI component, a rendered page), the agent can verify the result by examining a screenshot.
  • Multi-image analysis: With support for up to 600 images per request, the model can compare product images, analyze image sequences, or process batches of screenshots in a single call.

The key insight is that agent economics compound. An agent might make 20 to 50 model calls per task. If each image-inclusive call costs 10x more than a text-only call (as it does with Claude or GPT vision pricing), the total task cost multiplies rapidly. At V4-Flash rates, the marginal cost of adding vision to an agent step is negligible.

This shift mirrors a broader industry trend. Alibaba’s Qwen3-VL models demonstrated that smaller vision models can match larger ones on many tasks, and DeepSeek’s approach takes that further by pricing vision at text-model rates.

DeepSeek’s Files API and Harness Integration

Alongside the vision model, DeepSeek has launched two complementary services:

Files API

The Files API is free to use. It allows developers to upload images (or other files) once and reference them by file_id in subsequent API calls. This is useful in three scenarios:

  • When a single request would exceed the 48 MiB inline body limit.
  • When the image is larger than 32 MiB (the per-image limit for base64 and external URLs). The Files API supports up to 64 MiB per file.
  • When the same image is used across multiple requests, avoid re-uploading each time.

Files API limits: up to 25 GiB storage per user, up to 10,000 files, supporting JPEG, PNG, GIF, and WebP formats.

DeepSeek Harness 0.1.1

DeepSeek Harness is the company’s agent framework. Version 0.1.1, released the same day as the vision model, adds native support for deepseek-v4-flash-vision-exp. Agents built on Harness can now send mixed text-and-image input through standard /goal and /plan commands without custom model adapters.

For models that do not support image input, Harness includes a fallback mechanism that uses OCR, color statistics, and pixel analysis to decompose images into structured text before passing them to a text-only model. This means even agents using older DeepSeek models can process images through the Harness framework, albeit with lower fidelity than native vision.

Experimental Status and What It Means

The “Exp” suffix in the model name signals that this is an experimental release. DeepSeek has a history of using experimental labels for initial API launches before graduating models to stable identifiers. Developers should be aware of a few implications:

  • The model name may change. DeepSeek could rename or version the endpoint as it iterates on the model. Build your integration with the expectation that deepseek-v4-flash-vision-exp might be replaced by a stable identifier in a future update.
  • Performance may improve. Experimental models often receive quality updates without a name change. The image understanding you get today may get better over time.
  • The model is API-only. Unlike some DeepSeek models, the vision variant is not available as open weights for self-hosting. Community efforts to add vision to self-hosted V4 Flash (such as the DGX Spark plugin) exist but are separate from this official release.

Despite the experimental label, the model is available in production API endpoints with the same rate limits and SLA as other V4 models. Developers do not need to request special access or join a waitlist.

Frequently Asked Questions

Does DeepSeek V4 Flash have vision?

Yes. As of August 21, 2026, the deepseek-v4-flash-vision-exp model accepts image inputs alongside text through the DeepSeek API. The original text-only deepseek-v4-flash model does not accept images. You must use the -vision-exp variant specifically.

How much does image input cost with DeepSeek V4 Flash Vision?

Images are billed at the same V4-Flash token rate as text. Each image consumes up to 384 tokens after automatic resizing. At off-peak rates ($0.22 per million input tokens), a single image costs approximately $0.000084. At peak rates, approximately $0.000168. The cache-hit discount applies to image tokens as well.

Can DeepSeek V4 Flash Vision compete with GPT and Claude on visual reasoning?

DeepSeek’s published benchmarks show a mixed picture. The model beats Claude Opus 4.8 on 3 of 11 evaluations (DeepSWE, Agents’ Last Exam, and ZeroBench), ties closely on several others, but trails by 8 to 12 points on NL2Repo and DSBench-Hard. On multimodal tasks specifically, it wins 2 of 4 benchmarks. These are self-reported results using DeepSeek’s internal evaluation setup, and independent verification has not yet been published. The comparison is against Opus 4.8, not the newer Opus 5. For practical agent workflows like screenshot reading, document processing, and chart analysis, the cost advantage makes it the more practical choice regardless of benchmark margins.

Is the DeepSeek vision model available through OpenRouter or other providers?

Currently, deepseek-v4-flash-vision-exp is available directly through the DeepSeek API at api.deepseek.com. Third-party providers may add support over time, but the official API is the primary access point.

What is the difference between DeepSeek V4 Flash and V4 Flash Vision?

V4 Flash is a text-only model. V4 Flash Vision adds image input support while maintaining the same text capabilities (reasoning, agent performance, and world knowledge). The vision model is billed at the same rate and uses the same API format, with the addition of image content blocks in user messages.

The Cheapest Multimodal API Just Arrived

DeepSeek V4 Flash Vision marks the moment when multimodal AI became accessible at text-model prices. The model is experimental, but the API is live, the pricing is real, and the agent ecosystem is already integrating it. DeepSeek’s own benchmarks show it beating Claude Opus 4.8 on 3 of 11 evaluations while costing a fraction as much, though the wider gaps on tasks like NL2Repo and DSBench-Hard suggest it is not yet a universal replacement for frontier models.

For developers who have been routing image inputs through expensive frontier models, the math is straightforward. At $0.22 per million input tokens with images capped at 384 tokens each, the cost of adding vision to an agent step drops from cents to fractions of a cent. For workflows that process hundreds of images per session, the savings compound quickly.

The launch also signals DeepSeek’s broader strategy. By pairing the vision model with a free Files API and agent framework support in Harness 0.1.1, the company is building an integrated multimodal stack rather than shipping a standalone model. As the experimental label matures into a stable release, developers can expect the ecosystem around V4 Flash Vision to deepen further.

Aaron Jackson
Aaron Jackson
With a decade of hands-on experience in publishing and social media, and a B.Eng in Robotics from UWE, I'm passionate about turning challenges into opportunities. My focus is on creating solutions rather than merely highlighting problems.

Share post:

Popular

Enhanced Geothermal Systems Hit the Grid: Inside the 2026 Breakthrough

For decades, geothermal power was confined to a handful...

MIT’s Pressurized Wind Tunnel Unlocks a Wind Farm Efficiency Boost

Wind farms routinely fall short of the performance their...

The Indie Creator’s Tech Stack: How to Produce a Monetizable AI Short Film in 48 Hours

The creator economy is undergoing a seismic shift, and...

Microbial Fuel Cells: Bacteria Cleans Wastewater and Generates Electricity

Every day, the world spends a fortune in electricity...