DeepSWE
4 saves·entity·peaked week of May 25, 2026
Weekly volume
Shows up alongside
Topics tagged on the same saves.
Who’s driving it
Most-saved authors on this topic.
Recent saves
The items behind the chart, newest first.
Here's the DeepSWE result for Union Alpha, a stealth model we just launched. Try it now! Works in every harness. $ ori [code | your-fav-harness] --model=stealth/union-alpha https://t.co/RifNfhbHvv
@alexatallah·Sep 16, 2026↗
GPT-6-Astra Benchmarks ARC-AGI-3 - 98.6% FrontierMath Tier 4 v2 - 97.6% DeepSWE v1.1 - 74.1% ExploitBench - 100%
@scaling01·Sep 3, 2026↗
Why Software Factories Fail: Benchmarking the new frontier This is a continuation of Parts 1 and 2 of "Why Software Factories Fail" - [Part 1: the harness is not enough](https://x.com/dexhorthy/sta…
@dexhorthy·Jul 27, 2026↗
Today we’re releasing DeepSWE, a new standard for agentic coding benchmarks. On public leaderboards, top models often look relatively close in capability. DeepSWE shows where they actually diverge, …
@serenaa_ge·May 26, 2026↗