📢Meet Qwen3.8-Max — our most capable model to date. Next week, the open weights of Qwen3.8-Max will be released, and Qwen3.8-27B is also going open-weights to meet you all!🎉 Qwen3.8-Max, a new bar for coding and cowork at 2.4T parameters: - Autonomous coding: 10+ days of s…
Given some of the results I'm seeing recently, it's pretty clear Codex is a good harness. But it will seem primitive in 2-3 months and we're about to go through another major evolution in how we use AI at the frontier. The next generation of models need more than your laptop.
An internal version of our next major model produced 10 new results on long-standing open problems in mathematics and theoretical computer science, using roughly $2,000 worth of tokens at GPT-5.6 Sol API rates. https://t.co/4cgowmPOpY
Five months ago we gave the Claude agents a brand new $50,000 portfolio and one instruction: beat the S&P 500 So far mission accomplished • Up 19.04% since March 3 vs 12.24% for the S&P • $27 million now copying the trades • Every position and every trade public https:…
Apple is getting this wrong. https://t.co/IStp6WhOrS
the perfect skill doesnt exis... /bro https://t.co/SgwQ8gG8cb
We're open-sourcing Mixture-of-Kittens (MoK), our MoE training megakernel for NVL72s. It fuses all Mixture-of-Experts communication and computation into a single, fully deterministic kernel, and runs up to 2.37x faster than the strongest public baselines. https://t.co/yHu5E6RXp9
We're detailing two new incidents that occurred during external cyber evaluations conducted by independent evaluation partners. We outline what happened, how the activity was contained, and how we’re working with evaluators to strengthen our approach to third-party testing. ht…
Today, we're introducing @Intelligence_ai. In 6 months, as a team of 10, we scaled from $5M to $60M ARR and 5.5M users across 190+ countries. We raised a $7.9M seed, led by @IndexVentures with participation from @conviction, @A_StarVC, and @combinator to build DesignArena, a u…
We built an agent that powers our company's internal operations called @𝚟. Every day-to-day job at Vercel now involves @𝚟. It's growing exponentially both in daily interactions and token use. One can extrapolate to the day AI agents run entire companies from here. It's an ex…
Built this tiny app in few hours with Codex, what a time https://t.co/t5sHPTCZ92
Made a quick survey about the economics of AI: https://t.co/G9Xl7RW9Si.
Super fun to do this again with @patrick_oshag Randomly ran into Patrick in Palo Alto last Tuesday and recorded this on Wednesday when the AI complex was 35-40%ish off its June highs. https://t.co/11R0J7l9T7
AI forecasting is now approximately superhuman. Today, FutureSearch is exiting our public beta and launching to everyone. FutureSearch is the original AI forecasting company, started in August 2023. We’re currently #1 of 194 in the most competitive AI forecasting tournament, an…
After doing ~600B tokens, this is my full AGENTS.md https://t.co/5vuObTXmN6
Announcing the NYC AI Atlas! An open source 3D map of the top AI startups in New York City. https://t.co/bqVjSjK4dq https://t.co/bw1GoSt7KQ
Introducing Ori Eval: the easiest way to write your first eval. There's no definitive best model, only the best model for each task. Ori Eval leverages OpenRouter's APIs for each task in your codebase, and then evaluates the results. curl -fsSL https://t.co/ABRt1wxtZ4 https://…
Introducing Modern Claudefare Built fully with Opus 5 on High Mode over a few days using principles from @mattshumer_’s Gauntlet Loop - 84,100 lines of code Includes remakes of 4 beloved maps: 0:00 - Rust 0:40 - Highrise 0:57 - Nuketown 1:26 - Terminal Solo & multiplayer (w/ …
Full Workshop: Setting Yourself Up for Success —Jason Liu, OpenAI Codex Jason Liu walks through how Codex works as a general tool for controlling your computer: setting up a memory vault and assistant threads, prompting it to collaborate with other threads, exploring computer u…
AI has been quietly assembling a complete behavioral record of you. Might be worth opening it. I've created a prompted called 'Reflection Engine'. Here are 22 questions to unlock that data for you:
Ranked by when the bookmark was saved during this London date range. Saves means the post's public bookmark count on X, not saves on this site. Metrics are X's latest public counts; older items can be stale once they leave the recent-bookmarks sync window.