“AI Disruption” Publication 10,500 Subscriptions 20% Discount Offer Link.
DeepSeek Finally “Opens Its Eyes.”
Just now, DeepSeek launched a multimodal model designed for the Agent era: deepseek-v4-flash-vision-exp.
The model is already available on the DeepSeek API platform, and users can access it by configuring the appropriate settings.
According to the official introduction, this is an experimental model. Building on the text capabilities of DeepSeek-V4-Flash, it adds visual understanding abilities, enabling it to process image inputs and further improve performance on multimodal Agent tasks.
At the same time, DeepSeek Harness version 0.1.1 has added native support for this model. Developers can directly call deepseek-v4-flash-vision-exp to integrate image understanding capabilities into existing Agent workflows.
This means DeepSeek’s multimodal capabilities are beginning to move beyond simple “image Q&A” toward more complex task execution scenarios.





