We audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and found it no longer reliably measures frontier coding capability.
We find the eval to be saturated at a ~70% noise ceiling, and are retracting our previous recommendation that the research community u…↗