Crypto insigtX

DeepSeek Vision Model's Multimodal Agent Capability Nears Opus 4.8, Splitting 4 Benchmarks 2-2

DeepSeek released the first batch of Agent benchmark scores for V4-Flash-Vision-Exp. Across 4 multimodal Agent evaluations, it split 2 wins and 2 losses against Opus 4.8: Agents' Last…

Published
Market
Crypto
Source
insigtX

DeepSeek released the first batch of Agent benchmark scores for V4-Flash-Vision-Exp. Across 4 multimodal Agent evaluations, it split 2 wins and 2 losses against Opus 4.8: Agents' Last Exam 27.3 vs 25.7, ZeroBench 35.0 vs 34.0; ApexBench 36.5 vs 39.4, Chartography 64.3 vs 65.0. Compared to the text-only V4-Flash-0731, the vision version improved on ApexBench from 26.2 to 36.5 and on Agents' Last Exam from 25.2 to 27.3. The official note states the text-only version ignores multimodal content, thus mainly reflecting the capability gained by adding visual input. After adding vision, text-based Agent performance did not significantly degrade; across 7 text benchmarks, 6 were higher than V4-Flash-0731, such as DeepSWE rising from 54.4 to 59.3 (exceeding Opus 4.8's 58.0), and Toolathlon 75.9 nearly tying Opus 4.8's 76.2. These results come from DeepSeek's internal self-testing, not third-party independent rankings; public Code Agent text tasks use DeepSeek Harness minimal mode with a max reasoning setting.

insigtX content is informational and educational, not investment advice.