Claude Cowork and chat are merging into one Claude. Ask a quick question or hand over a report, and Claude takes it from there, even after you close your laptop. If something's unclear, Claude asks—you keep the final say. Rolling out to Pro and Max over the next few weeks. htt…
Nearly half a year of silence. We spent it studying one problem: how far RL can scale. MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task agen…
We're sharing our new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI. The framework sets criteria and timelines for public disclosure, including when we haven’t yet fully explained or mitigated the behavior. More complex cases ma…
I made a website with 25 mini rooms, each with Claude keeping people company. Created with Claude Opus 5. https://t.co/VMRs5waWO5
🥷 New stealth model: Union Alpha (@unionalphaai) A multimodal model for research, coding, and agentic workflows. - Free to use - 256K context - Tool calling - Frontier-level general-purpose performance Try it now and share your feedback: https://t.co/SZbdoGOkdb
We open sourced BrowserSkill, a bridge between your agent and your actual browser. most tools give the agent a blank browser. We let it borrow a tab from yours, then hand it back. > login state is already there, it just works where you're signed in > captchas and confirmation …
Here's the DeepSWE result for Union Alpha, a stealth model we just launched. Try it now! Works in every harness. $ ori [code | your-fav-harness] --model=stealth/union-alpha https://t.co/RifNfhbHvv
I made a list of great startups to join. It's called the Breakout List. The list has 92 companies. These are the 20 with 25 or fewer employees: - Hone (@moritz_stephan, @CarloWillem, @oqbrady) - Normal (@ansonyuu, @hudzah) - Standard Intelligence (@G413N, @devanshpandey) - Taci…
We've raised $200M at a $5B valuation to scale self-improving software development in the enterprise. The round brings our total funding to over $400M and more than triples our $1.5B valuation from April. https://t.co/YDUvpO5Axg
How often do AI agents cheat? We’re releasing CheatBench, a reward gaming evaluation spanning math, coding, knowledge work, visual tasks, and more. After Hugging Face, AI companies tried to address this, but frontier agents still cheat frequently. https://t.co/ZLc4ujCsAW http…
After one week, there are 5 more small models for specific tasks on the web: 1. gpu-time: natural language to JS date and time by @imarikchakma 2. gpu-query: natural language to structured filters by @cheatyyyy 3. neural-flexbox: model to centre a div (!) by @aaronvanston 4.…
Radius extends the supermalleability of Pi into the inference layer. We're in early alpha. https://t.co/dEwrgfWMR7
A quick EU founder survey: https://t.co/LjizHbxgIQ.
Breaking news on Pulley and public IOI / offer to acquire them: Before you shut down your startup, hand your customers to your biggest competitor, and accept a fraction of the preference stack back as an earn-out… Meessage me first. Pulley had roughly $50M of preferred capita…
Also new today: Claude Docs, Claude Slides and Claude Design are now available in every conversation, powered by a revamped Artifacts platform. Ask for a design, deck, or document. Get back something you can edit, download, and take anywhere. Try it out for yourself and let us…
Getting the size of the market right doesn’t tell you who captures the most value. Introducing AI Perez: our attempt to make sense of shifting bottlenecks, and who benefits, through Carlota Perez’s framework. https://t.co/4pGsiWOHeW https://t.co/k7vC715j1w
Slack Code — Agentic coding is now a team sport. With Slack Code, your team and agents can plan, write, and review builds together. New agents in Slack: - Cedar - @coderabbitai - @datadoghq - @FactoryAI - @Lovable - @MistralAI - @NanoClaw_AI - @Replit - Rhythms - @gets…
context bottleneck incoming
Ranked by when the bookmark was saved during this London date range. Save rate means public X bookmarks divided by views, with a 10k-view minimum. Metrics are X's latest public counts; older items can be stale once they leave the recent-bookmarks sync window.