DeepSeek says its cheap new vision model gets close to Opus 4.8, but it graded its own exam
DeepSeek has added image understanding to its low-cost V4-Flash model and released it, on 21 August 2026, as an experimental, API-only build called V4-Flash-Vision-Exp. It published a benchmark table putting the model close to Anthropic's Claude Opus 4.8, ahead on a few tests and behind on the rest, in one case by twelve points. The catch is in who made the table: DeepSeek chose which benchmarks to show and ran the comparison on a setup it controls end to end, and none of the numbers have been checked by an outside party. So this is a single company's own scorecard, not an independent head-to-head. This is a report on what was claimed, not a verdict on which model is better.

On 21 August 2026 DeepSeek released an experimental, API-only model called DeepSeek-V4-Flash-Vision-Exp, which bolts image understanding onto its cheap V4-Flash text model. DeepSeek published a benchmark table comparing it with Anthropic's Claude Opus 4.8, on which its new model comes out ahead on a few benchmarks, most clearly Agents' Last Exam and ZeroBench, and behind Opus on the rest, including a twelve-point gap on the code benchmark NL2Repo. The important qualifier is the table's provenance: DeepSeek chose which benchmarks to show and ran the comparison on a testing setup it controls end to end, and the numbers have not been reproduced by a neutral third party. It has not even said whether it re-ran Opus itself for the comparison or used Anthropic's published figures. Either way this is a single vendor's own scorecard, not an independent test, and "experimental" is in the model's name. This is a report on the claim, and it is not investment advice or a ranking of which model you should use.
DeepSeek has a habit of turning up with a model that is far cheaper than the frontier and close enough on paper to make people look twice. Its latest, released on 21 August 2026, is DeepSeek-V4-Flash-Vision-Exp, an experimental build that gives the budget V4-Flash model the ability to read images. The headline that travelled was that it "matches" or "rivals" Anthropic's Claude Opus 4.8. The more useful way to read it is to look closely at where those comparison numbers come from, because that is what decides how much weight they can carry.
What did DeepSeek actually release?
According to DeepSeek's own API changelog, the new model is available through the DeepSeek platform under the identifier deepseek-v4-flash-vision-exp. It is explicitly experimental and, for now, API-only, so there is no consumer app or open-weights download attached to it. It extends V4-Flash, DeepSeek's cheap, fast text model, with multimodal understanding, and on DeepSeek's figures it improves on the earlier V4-Flash on the visual tasks it highlights. Pricing follows V4-Flash rates, and each image costs at most 384 tokens regardless of resolution, according to the reporting.
On the changelog page itself, DeepSeek frames the headline comparison in words, saying the model brings its "multimodal agent capabilities close to Opus-4.8." Alongside the release, though, it also put out a fuller benchmark table, and that is where the specific numbers everyone quoted come from.
Does it really beat Opus 4.8?
Here is the detail that matters most. DeepSeek published a comparison table with three columns: the new V4-Flash-Vision-Exp, its own previous model V4-Flash, and Claude Opus 4.8. As SiliconANGLE and OfficeChai report, the figures come from a testing setup DeepSeek controls end to end, and they are not independently verified.
On that table, the new model comes out ahead of Opus 4.8 on a few benchmarks, most clearly Agents' Last Exam (27.3 to 25.7) and ZeroBench (35.0 to 34.0), each by a point or so. It trails on the rest, and in places the gap is not close. The clearest example is the coding benchmark NL2Repo, where DeepSeek reports 57.7 against Opus 4.8's 69.7, a twelve-point difference; on DSBench-Hard it is behind by about eight points, and on a text terminal benchmark it sits just under Opus as well. So "matches Opus 4.8" is true only for the specific benchmarks where it happens to lead, and DeepSeek's own table shows Opus ahead on more of them.
Why the source of the numbers is the whole story
The single most important fact about this comparison is who produced it. It is DeepSeek's table, built from DeepSeek's choice of benchmarks and run on a setup DeepSeek controls end to end. That the table is DeepSeek's own is clear from the middle column: only DeepSeek would publish a comparison that also benchmarks its own older V4-Flash model. It has not even said whether it re-ran Opus 4.8 itself for the Opus column or dropped in Anthropic's published numbers, which is its own small gap in the record.
That does not make the numbers fake, and self-benchmarking is standard when a lab ships a model. But it does mean two independent things are unverified at once. First, a vendor chooses which benchmarks to show and can leave out the unflattering ones, so the selection itself is a choice, not a neutral survey. Second, the whole comparison runs on a setup one side controls, so the conditions, the prompting, the sampling and the exact benchmark versions, are DeepSeek's rather than a neutral referee's. None of the numbers has been reproduced by an outside party. The honest reading is "DeepSeek says," not "DeepSeek's model is." We made the same point when DeepSeek reworked its API pricing: a benchmark chart is a starting point for testing, not the end of the argument.
What does "experimental" buy DeepSeek?
Shipping this as an experimental, API-only model is a deliberate low-commitment move. It lets DeepSeek show a number, gather real usage, and keep the option to change or pull the model without the weight of a full product launch or an open-weights release it would have to stand behind indefinitely. For anyone building on top of it, that is the trade: early access to a cheap multimodal model, in exchange for no stability guarantee and a set of scores that no one outside DeepSeek has checked.
The strategic frame is familiar. DeepSeek competes on price and on being close enough to the frontier to be interesting, which keeps pressure on the more expensive labs. Whether "close enough" holds up is exactly the kind of claim that only independent evaluation can settle, and that has not happened yet. It also feeds the broader argument about open versus closed model strategies and whether the frontier's pricing can hold.
The claim at a glance
| Model | DeepSeek-V4-Flash-Vision-Exp (deepseek-v4-flash-vision-exp) |
| Released | 21 August 2026, experimental, API-only |
| What is new | Adds image understanding to the low-cost V4-Flash text model |
| DeepSeek's own claim | Multimodal agent capabilities "close to Opus-4.8" |
| Who made the comparison | DeepSeek, which chose the benchmarks and ran it on a setup it controls end to end |
| Where it leads | A few benchmarks, most clearly Agents' Last Exam 27.3 vs 25.7 and ZeroBench 35.0 vs 34.0 |
| Where it trails | The rest, including NL2Repo 57.7 vs 69.7 (a 12-point gap) and DSBench-Hard (about 8 points) |
| Verification status | Self-reported by DeepSeek, not independently verified |
| Pricing | Follows V4-Flash rates; an image costs at most 384 tokens regardless of resolution |
None of this makes V4-Flash-Vision-Exp uninteresting. A cheap model that DeepSeek's own numbers put anywhere near a frontier system on a handful of benchmarks is worth a look, especially at DeepSeek's prices. It just means the accurate summary is narrower than the headline: DeepSeek published its own numbers putting its experimental model ahead of Opus 4.8 on some tasks and behind on more, and until someone outside DeepSeek runs the same tests under the same conditions, that is where the story sensibly stops. It is one more data point in the argument over whether AI's current valuations are a bubble, and a reminder that a benchmark is a claim before it is a fact.


