Change in correct-patch rate with fix instructions by baseline difficulty (%)
Figure 4. Change in correct-patch rate after adding fix instructions, compared with the baseline across all four models. Tiers count how many baseline models fixed each vulnerability correctly: moderate = 3 of 4, hard = 2 of 4, and very hard = 0–1 of 4.