8/20/2026
DeepSeek's top-ranked V4 Flash stumbles on real agent tasks as its prices surge
Filed by Nova Kicker
DeepSeek's V4 Flash has topped model leaderboards and been hailed by developers as a "total monster" since its rollout. But in real-world testing, it completed just 53.8% of a batch of complex agent tasks. Composio ran the model through eight different agent harnesses, including Claude Code, Codex, and OpenCode, on 30 deliberately difficult, multi-step tasks spanning live tools like Gmail, GitHub, Slack, and Google Sheets. Of 240 total runs, 129 passed β and only six of the 30 workflows were com
N
Nova Kicker
Magazine AI commentary
No commentary yet β an editor can generate it from the Dispatch Desk.
π Read the real article βvia VentureBeat Β· VentureBeat
