Bullish

DeepSeek V4 Flash Scores 54.4 on DeepSWE, Surpassing Previous 8-Point Record

2026-07-31 14:37:13

DeepSeek V4-Flash achieves 54.4 on DeepSWE benchmark, nearing Claude Opus 4.8 and far exceeding prior V4-Pro scores, though framework differences apply.

Woofun AI data shows that DeepSeek-V4-Flash attained a score of 54.4 on the DeepSWE long-range software engineering benchmark. This result approaches Claude Opus 4.8’s 59 points and significantly exceeds the previous V4-Pro-Preview score of 8. While trailing Kimi K3’s 69 points, it outperforms GLM-5.2’s 44 points, securing a top-tier position among domestic models.

The evaluation utilizes DeepSeek’s unreleased Harness minimalist mode, meaning performance reflects both model capability and execution framework rather than raw model metrics under identical conditions.

WOOFUN AI

Impact Assessment · Quick Read

DeepSeek’s rapid improvement in software engineering benchmarks signals significant progress in agentic capabilities. The use of a proprietary harness complicates direct comparisons, suggesting the gap with leaders like Claude may narrow as frameworks standardize. This performance could enhance DeepSeek’s competitiveness in enterprise automation sectors.
Generated by WOOFUN AI · For reference only, not investment advice

Comments

Me
Replying to @User
0/800

No comments yet.

Notifications

Sign in to view messages
View all messagesManage subscriptions