We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our …
I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we've had at OpenAI in recent weeks. Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We'll have more to share soon.
We're publishing our most detailed threat intelligence report to date. It covers how people tried to misuse Claude—for cyberattacks, influence operations, surveillance, biology, and building weapons—and how we found and stopped them. We disrupted every operation in the report…
So Jacob Coxon, who dramatically resigned from Anthropic yesterday, worked there for a grand total of six weeks. He started with them in July. All of his socials appeared yest. It has all the signs of a highly coordinated op through doomer mega donors and the corporate media.
Super Smash Bros Melee in MR. Where the players fight on your own table as a platform. Even hang off the table! On the Meta Quest. https://t.co/gCGBuDCJPP
GPT-Live-1 is now available in the API. Bring ChatGPT’s natural back-and-forth to your app, with voice agents that listen while they speak and work with the models and harness you choose. https://t.co/gIl1gwsBDV
Now everyone can put data to work. We’re introducing a new Data agent in ChatGPT Work so you can turn your company’s data into answers, interactive dashboards, and action—just by asking. Just add the Data Plugin in ChatGPT Work, connect to the data sources and context you alre…
I can’t stress enough how little an idea matters compared to the agency of the people executing the idea. I have had the privilege of knowing and sometimes even working with some of the most successful people (by various metrics). The difference between mediocre and excellent…
Introducing SWE-2, our closest model yet to the frontier. On leading evals, it scores on par with recent frontier models – at up to 70% lower cost. We scaled RL to multiple trillions of parameters, with a refined recipe that pushes the Pareto curve on both capabilities & c…
Your pet has a job now. Pets help you keep track of your chats while you’re away from the desktop app. Now, you can start a new chat right from your pet, too. Not a pet person? Meet Mini: the same shortcuts and updates, taking up less space on your screen. Pick your pet—or go…
New in Claude Code: claude plugin eval See what value your plugin is adding, or if it needs more work. You can create test cases, run your plugin or skill against those test cases, score those runs, then run each case again without the plugin to see the differences. https://t.…
AI Engineering Skills Map: Shaping the build When you’re skilled at AI Engineering, your best work won’t be merely implementing a product that someone else spec’ed out. Instead, you will actively shape the build. Before modern AI tools accelerated and expanded what a single de…
Go from idea to a working agent faster with the Agents API. Build and run cloud agents with the Codex harness, fully managed by OpenAI. We handle orchestration, long-running sessions, and context management. You focus on what makes your agent unique. Available in public beta.…
This is Knap. It's a new language I created that turns data into Markdown. The syntax should feel familiar and comes with wonderfully pleasant features to modify and format plain text. Knap is open source. Over a million people already use Knap directly or indirectly because it…
Introducing Super Smash Royale, my best creation yet Play Battle Royale in the browser with 45 characters from Melee + Brawl Built in a few days with GPT-6 Astra, @ElevenLabs, @MeshyAI, @Blender and @threejs Solo, multiplayer, and controller support smashroyale dot io https:…
this was wild amounts of disinformation / fear mongering / the stupidest interview ive ever seen: 1) ai did NOT hack huggingface on its own "independent volition". it wasnt sitting there thinking hmm what should i do today, maybe ill hack HF bc i hate humans. No, 10841 *was pro…
@rcx86 If only I had thought to train it on StackOverflow instead I think it would have clicked for me that LLMs can be promptable, general purpose Q&A engines. Ah well.
We just rolled out CursorBench 4.0! It includes new tasks for how well models follow instructions, work on challenging projects over time, and is more difficult than before (so all models score lower). https://t.co/cYVgRBEFWE
🌊 SYSTEM PROMPT LEAK 🌊 Got the full system prompts and tools for GPT-6 Astra! 🚀 This MASSIVE dump comes in at >330k characters for the prompts and >1.1M for the tools 🤯 Hope you enjoy! 🤗 https://t.co/qRGW8pFb1x Won't all fit in a tweet, of course, but here's the first …
Here you go, link in the next post https://t.co/HIyZvKM4l8
Ranked by when the bookmark was saved during this London date range. Interactions are likes, reposts, replies, and quotes combined. Metrics are X's latest public counts; older items can be stale once they leave the recent-bookmarks sync window.