bookmarks
4 bookmarks

We’re demonstrating how frontier models have continued to improve in realistic mental health conversations with MentalHealthBench. This new open benchmark was built with input from more than 80 mental health clinicians. We’re releasing it openly so other researchers can examine…↗

We audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and found it no longer reliably measures frontier coding capability. We find the eval to be saturated at a ~70% noise ceiling, and are retracting our previous recommendation that the research community u…↗

Take a look at your favourite skill. Go on, take a look. Check for lines like: - "Make the commit message very detailed" - "Be thorough" - "Make the implementation easy to read" What do these lines have in common? They're no-ops. They do nothing to change the agent's behavio…↗