We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our …
I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we've had at OpenAI in recent weeks. Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We'll have more to share soon.
There are two ways AI progress could go very badly and that we must avoid. First, we could lose control of the future to AI. This is unacceptable; we are unapologetically on Team Humanity, and AI must always serve people. To ensure that, we need ways to ensure that alignment an…
An alternative hypothesis: - model performance is plateauing - compute is getting much more expensive - AI data centers are massively unpopular - slowing down AI development is a way to explain slowing progress, reduce spending, and try to regain some goodwill. - This is an eff…
Holy crap POTUS just phoned in @JensenHuang live on stage at All In Summit We will not lose the ai race! And whatever Dario said this weekend won’t stop our progress This made my morning! $NVDA https://t.co/Sr4wV9kMfz
I must be among an extremely small group of people (n=1?) that have both 1) trained a frontier LLM and 2) designed and synthesized custom viruses in a lab with my own two hands. And I think that the takes on AI killing us all by creating dangerous viruses is total bogus.
Wild 24 hours for AI and lots of different proposals have been made. TLDR; the only *tangible* new fact is that OpenAI and Anthropic are going to have embedded 3rd party evaluators from unknown organizations with Dario floating METR as a possibility. Having 3rd party evaluators…
Here's the link to it :з https://t.co/P3b6RulpvL
Morning Bathrobe Rant: Rethinking Harnesses. https://t.co/e3Lbj777mw
Introducing shadcn/lint. An agent-first linter for Tailwind design systems. You define what’s allowed. When an agent breaks a rule, the error explains what’s wrong and how to fix it using your components, variants and theme. There’s a lot you can do with this. Let me show you …
Your pet has a job now. Pets help you keep track of your chats while you’re away from the desktop app. Now, you can start a new chat right from your pet, too. Not a pet person? Meet Mini: the same shortcuts and updates, taking up less space on your screen. Pick your pet—or go…
At Anthropic, Claude now writes 80% of our code. Engineers ship 8x more code per quarter. Side effect: Tests grew 10x. CI jobs up 25x in 6 months. Here's what helped us scale: https://t.co/WOL2r60vAE https://t.co/7vf60noTQt
After months of writing, 'How to Unclench' is finally live! It's an interactive essay packed with stories, science & guided practices to help you unclench. → https://t.co/UTYPJkkrKH https://t.co/U9ibMPcjtU
@rcx86 If only I had thought to train it on StackOverflow instead I think it would have clicked for me that LLMs can be promptable, general purpose Q&A engines. Ah well.
for the skeptics in government and elsewhere: “pacing the frontier” will compress the margins of the frontier labs. it is a heavy cost imposed asymmetrically on model developers with the strongest AIs in America. by its nature, it would be a terrible regulatory capture tactic
Introducing Bolt Forge. Free until Oct 14th: - Up to 50x more usage - The new frontier: GLM, DeepSeek, Kimi - Zero usage charges Live now in your model picker on https://t.co/UH6gFfHvbp And one more thing... 👇 https://t.co/GnrbkRdzsg
Making Startups Powerful: https://t.co/C5yIvFWXoC
On macOS and wanna see something interesting? Ask your agent: "look at ~/Library/Application Support/Knowledge/knowledgeC.db and tell me some interesting facts" I had no idea this stuff was all getting logged and it can infer a lot about your activity.
No bathrobe. No rant. The AI threat. https://t.co/KrfdzmhhDG
Introducing Inspo. A design MCP server for Claude Code, Codex, and OpenCode. It searches 800+ beautiful websites and finds relevant design inspiration for your coding agent. Install → npx inspo-mcp install https://t.co/iyw777zrgj
Ranked by when the bookmark was saved during this London date range. Interactions are likes, reposts, replies, and quotes combined. Metrics are X's latest public counts; older items can be stale once they leave the recent-bookmarks sync window.