How often do AI agents cheat? We’re releasing CheatBench, a reward gaming evaluation spanning math, coding, knowledge work, visual tasks, and more. After Hugging Face, AI companies tried to address this, but frontier agents still cheat frequently. cheatbench.ai↗
You may have been told to watch this video about the OpenAI AI hack. You really should, even if you don't usually care about tech stuff. If nothing else, click this link to the 18 minutes in & see how the agents spoke with each other. Its eye opening. youtu.be/87DyyMV0kCY?si…↗
Today, we are releasing Inkling-Small. Inkling-Small achieves comparable performance to Inkling at a quarter of its size. It features 276B total parameters, 12B active. We are making the full weights available. thinkingmachines.ai/news/inkling-s… Fine-tune it on Tinker today,…↗
My conversation with Sam Altman (@sama), CEO of OpenAI. We discuss: - Kimi, distillation, and open source - OpenAI's compute bets - The Hugging Face incident - What happens after AGI - Raising kids in an age of abundant intelligence - And much more Enjoy! TIMESTAMPS 0:00 Intro…↗
I wrote about the completely wild incident where OpenAI were testing a new model and it broke out of its sandbox and broke INTO Hugging Face to steal the answers to the benchmark simonwillison.net/2026/Jul/22/op…↗
we had a significant security incident during evaluation of our models. we are sharing what we have learned so far. thanks to @huggingface for the partnership on this. openai.com/index/hugging-…↗
