DeepSeek Releases Experimental Vision Understanding Model, Multimodal Agent Capability Approaches Opus-4.8
DeepSeek launched its experimental vision understanding model, V4-Flash-Vision-Exp, on August 21, adding image understanding capabilities while maintaining existing text performance, with visual agent benchmark results approaching Opus-4.8. It…
DeepSeek launched its experimental vision understanding model, V4-Flash-Vision-Exp, on August 21, adding image understanding capabilities while maintaining existing text performance, with visual agent benchmark results approaching Opus-4.8. It also introduced a free Files API and DeepSeek Harness 0.1.1, supporting developers in handling mixed image-text tasks. The API retains V4-Flash pricing, with images converted to a maximum of 384 tokens. This update expands multimodal agent application scenarios.
Original: https://wallstreetcn.com/articles/3780000
insigtX content is informational and educational, not investment advice.