Correct-patch rate with and without fix instructions
50%
60%
70%
80%
90%
100%
CORRECT PATCHES (%)
FRONTIER MODEL
baseline
fix instructions
76.9%
90.4%
Claude Opus 5
67.3%
84.6%
GPT-5.6 Sol
71.2%
76.9%
Kimi K3
53.8%
63.5%
DeepSeek V4 Pro
Figure 1.
Every patch was judged against a verified human fix and scored as correct (1) or incorrect (0).