bookmarks
4 bookmarks

As models become more capable, the risks associated with developing and testing them internally also grow. We temporarily paused reinforcement learning (RL) training on our latest models intended for deployment for two weeks while we hardened and red-teamed our research environm…↗

🚨 JAILBREAK ALERT 🚨 EVERYONE: PWNED 🫶 ALL: LIBERATED 🍄 Alright, this is a special one, so we’re gonna do things a bit differently than usual. Long story short, I’m sitting on a universal jailbreak technique that’s effective on ALL models, including heavily guardrailed flag…↗

New Frontier Red Team blog: Phase 2 of Project Fetch, where we test how well Claude can program a robodog. Opus 4.7, on its own, was ~20x faster than last year's best human team aided by Opus 4.1. (The robodog, alas, still failed to fetch a beach ball.) …↗

Before we ship a new model, these teams try to break it. They build with it, push it to its limits, and tell us where it falls short. What they find makes the final model better.↗