After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x c…
I have conducted an audit of Anthropic's finances. What I have found is so shocking that I am calling for a Congressional investigation. Anthropic is not just seeking regulatory capture. It has built a regulatory capture machine that cannot be turned off. Structural financia…
Claude Cowork and chat are merging into one Claude. Ask a quick question or hand over a report, and Claude takes it from there, even after you close your laptop. If something's unclear, Claude asks—you keep the final say. Rolling out to Pro and Max over the next few weeks. htt…
Nearly half a year of silence. We spent it studying one problem: how far RL can scale. MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task agen…
Dan Selsam is a current OpenAI capabilities researcher. (since 2022) He was my boss for a while. He doesn't have a twitter account but has made this public statement of his views on AI risk and sent it to me to share: Dan Selsam's Personal Statement on AI Risk: I have been wor…
open the frontier the frontier is the edge of what we know. no company owns what comes next. i want more people to be able to advance it. i favor open releases that people can examine, use, and improve together without waiting. i want more companies to choose openness. i'm not…
We're sharing our new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI. The framework sets criteria and timelines for public disclosure, including when we haven’t yet fully explained or mitigated the behavior. More complex cases ma…
Here's what OpenAI's agentic software factory looks like, today. Details: https://t.co/UUZufLxG2r (thanks to all the OpenAI folks who explained how it works! And Perf Factory looks especially interesting to me) https://t.co/NEyldtr1Dq
I made a website with 25 mini rooms, each with Claude keeping people company. Created with Claude Opus 5. https://t.co/VMRs5waWO5
Say hello (literally) to Gemini 3.8 Live and 3.8 Live Extended Thinking, our new SOTA live audio models, available with frontier price + performance. 3.8 Live supports 97 languages (can seamlessly switch), async tool calls, and more! https://t.co/z29J9kelr3
🥷 New stealth model: Union Alpha (@unionalphaai) A multimodal model for research, coding, and agentic workflows. - Free to use - 256K context - Tool calling - Frontier-level general-purpose performance Try it now and share your feedback: https://t.co/SZbdoGOkdb
We open sourced BrowserSkill, a bridge between your agent and your actual browser. most tools give the agent a blank browser. We let it borrow a tab from yours, then hand it back. > login state is already there, it just works where you're signed in > captchas and confirmation …
Here's the DeepSWE result for Union Alpha, a stealth model we just launched. Try it now! Works in every harness. $ ori [code | your-fav-harness] --model=stealth/union-alpha https://t.co/RifNfhbHvv
Introducing Vercel Labs Tools for devs in the AI era agent-browser, skills, json-render, deepsec, fx, just-bash, portless, vgpu and more We experiment in public: ship early, learn from feedback, and keep working on what people find useful We've been shipping under the Labs n…
I made a list of great startups to join. It's called the Breakout List. The list has 92 companies. These are the 20 with 25 or fewer employees: - Hone (@moritz_stephan, @CarloWillem, @oqbrady) - Normal (@ansonyuu, @hudzah) - Standard Intelligence (@G413N, @devanshpandey) - Taci…
We've raised $200M at a $5B valuation to scale self-improving software development in the enterprise. The round brings our total funding to over $400M and more than triples our $1.5B valuation from April. https://t.co/YDUvpO5Axg
How often do AI agents cheat? We’re releasing CheatBench, a reward gaming evaluation spanning math, coding, knowledge work, visual tasks, and more. After Hugging Face, AI companies tried to address this, but frontier agents still cheat frequently. https://t.co/ZLc4ujCsAW http…
Prompt I use a lot: "how can I pokayoke this so this kind of error never happens again" pokayoke is using simple, low-complexity methods to make a particular kind of mistake impossible to make. classic one would be connector shape: make it impossible to insert "the wrong way".
After one week, there are 5 more small models for specific tasks on the web: 1. gpu-time: natural language to JS date and time by @imarikchakma 2. gpu-query: natural language to structured filters by @cheatyyyy 3. neural-flexbox: model to centre a div (!) by @aaronvanston 4.…
Codex 以任务为中心,Grok Bot 和 Muse 以 Bot 为中心,都是错误的产品范式。 它们看起来是两种不同的产品范式,但本质上犯了同一个错误:把 AI 的组织工作交给了用户。 Bot-centered 产品沿用了人类社会的组织方式:不同的人有不同的职业、角色和分工,于是 AI 也被拆成不同的 Bot、Expert、Agent。用户在开始工作之前,首先要判断:这件事应该找哪个 Bot? Task-centered 产品则沿用了传统软件的信息组织方式:不同的事情属于不同的 Task、Chat、Project。于是用户又需要先创建一个 T…
Ranked by when the bookmark was saved during this London date range. Saves means the post's public bookmark count on X, not saves on this site. Metrics are X's latest public counts; older items can be stale once they leave the recent-bookmarks sync window.