AI Disruption

AI Disruption

DeepSeek Multimodal Model deepseek-v4-flash-vision-exp Released

DeepSeek launches vision-enabled V4-Flash multimodal model with Files API, low-cost token pricing, and Agent-ready image understanding.

Meng Li's avatar
Meng Li
Aug 22, 2026
∙ Paid

“AI Disruption” Publication 10,500 Subscriptions 20% Discount Offer Link.


DeepSeek Finally “Opens Its Eyes.”

Just now, DeepSeek launched a multimodal model designed for the Agent era: deepseek-v4-flash-vision-exp.

The model is already available on the DeepSeek API platform, and users can access it by configuring the appropriate settings.

According to the official introduction, this is an experimental model. Building on the text capabilities of DeepSeek-V4-Flash, it adds visual understanding abilities, enabling it to process image inputs and further improve performance on multimodal Agent tasks.

Image
DeepSeek Harness: Claude Code & Codex Become Sub-Agents

DeepSeek Harness: Claude Code & Codex Become Sub-Agents

Meng Li
·
Aug 20
Read full story

At the same time, DeepSeek Harness version 0.1.1 has added native support for this model. Developers can directly call deepseek-v4-flash-vision-exp to integrate image understanding capabilities into existing Agent workflows.

This means DeepSeek’s multimodal capabilities are beginning to move beyond simple “image Q&A” toward more complex task execution scenarios.

User's avatar

Continue reading this post for free, courtesy of Meng Li.

Or purchase a paid subscription.
© 2026 Meng Li · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture