DeepSeek shipped V4-Flash-Vision-Exp. Headlines said close to Claude Opus 4.8 on agent benches. I opened the table anyway.
They win 3 of 11 on their own harness. They lose the hard repo-scale task by about 12 points. Part of the “vision leap” is that the old model could not see the pictures in the test. Vendor bench. Fine. I still tried it, because it is cheap and it can look at a screenshot.
That is how the leaderboard actually moves. Not a keynote. A model you can call today, cheap enough to leave on in an agent that clicks through a UI.
I do not care who wins the graph. I care whether I can point it at a messy admin screen and get a file back.