Correct-patch rate with and without fix instructions

50%60%70%80%90%100%CORRECT PATCHES (%)FRONTIER MODELbaselinefix instructions76.9%90.4%Claude Opus 567.3%84.6%GPT-5.6 Sol71.2%76.9%Kimi K353.8%63.5%DeepSeek V4 Pro
Figure 1. Every patch was judged against a verified human fix and scored as correct (1) or incorrect (0).