We made claude.ai 3x faster in two weeks. Here’s how we use Claude to measure, debug and improve performance. Prompts and methods included. claude.dev/blog/how-we-ma…↗
5 bookmarks
My job at @every is basically benchmarks-as-a-service now: every.to/also-true-for-…↗
An engineer on our team @thegp pointed out this is better than benchmarks - I agree - This kind of research into agents coordination, agent failure modes, and behavior in the face of false or contradictory information is much more insightful than the numerical scores we normally…↗
·@phineasb·Aug 20, 2026·essay·agent-coordination·agent-failure-modes·benchmarking·contradictory-information
