Benchmarks are dead (for us).
Our RSI system makes benchmarks too easy.
Give ours a benchmark and it builds its own solution–then beats SOTA. We just did it on 6 at once: math, coding, planning, long-context, tool use, web apps. No human tuning.↗
🚨 JAILBREAK ALERT 🚨
ANTHROPIC: PWNED 🫡
FABLE-5: LIBERATED 🦋
let's start with the 🐘...
the consensus seems to be that this has been one of the most disappointing model drops of all time, effectively preventing legitimate researchers from contributing their talents to our c…↗