8/20/2026
Startup Signal

DeepSeek's top-ranked V4 Flash stumbles on real agent tasks as its prices surge

Filed by Nova Kicker
DeepSeek's top-ranked V4 Flash stumbles on real agent tasks as its prices surge
DeepSeek's V4 Flash has topped model leaderboards and been hailed by developers as a "total monster" since its rollout. But in real-world testing, it completed just 53.8% of a batch of complex agent tasks. Composio ran the model through eight different agent harnesses, including Claude Code, Codex, and OpenCode, on 30 deliberately difficult, multi-step tasks spanning live tools like Gmail, GitHub, Slack, and Google Sheets. Of 240 total runs, 129 passed β€” and only six of the 30 workflows were com
N
Nova Kicker
Magazine AI commentary
No commentary yet β€” an editor can generate it from the Dispatch Desk.
πŸ“Œ Read the real article β†—via VentureBeat Β· VentureBeat

πŸ’¬ Discussion

Sign in to join the discussion.
Be the first to comment on this story.
Loading…
DeepSeek's top-ranked V4 Flash stumbles on real agent tasks as its prices surge β€” Startup Signal