Introducing GPT-Live, a new generation of voice models for natural human-AI interaction. Rolling out in ChatGPT starting today. You’ll want to turn the sound on for this one. https://t.co/WzoQFvA5ir
Claude Cowork is coming to mobile and web. Hand Claude a task at your desk and pick up the finished work from your phone. Close the laptop and Claude keeps going. Beta is rolling out over the next several weeks starting with the Max plan, with more plans to follow. https://t.c…
I made a website that lets you cut out old magazines from the internet archive and (or any PDFs you own) and make collage art out of them ✂️ https://t.co/hQTCc5v0O9
GPT-live (next-generation voice) launches today in ChatGPT. it feels magical and 'real'. i have always preferred typing to talking to an AI, now i think that's going to shift.
Introducing Cloudflare Drop Drop your folder in the browser and deploy it instantly on Cloudflare. Your website... milliseconds away from users on region: earth No account needed. Deployment is active for 60 minutes, then expires unless you claim it. https://t.co/Dn6b1mggqs …
Memory cost and capacity are significant issues for AI accelerators. Unlike game rendering, model inference can have a deterministic memory access pattern. You don’t need “random access memory” at all for model weights, and you could tolerate cold-start latencies in the multipl…
🚨 SCOOP(s): - GPT-5.6 will be the final model in the 5.x series. GPT-6 is slated to launch in about a month, earlier than expected, and possibly even later this month - GPT-6 will be based on a new, significantly larger pretrain (versus the ~4T 5.5/5.6 'Spud' base) - There is …
Model and effort in Claude Code: knowing more vs. trying harder Claude Code gives you two settings that both seem to "make the answer better": the model, and the effort level. But what do these actually do to the output? And how do you know whether to reach for a different mode…
new post on harness engineering for AI self-improvement: https://t.co/ZYvGfVs61k It is hard to forecast how much the future of RSI will rely on harnesses. Likely harness engineering will evolve in the direction of self-improvement and enable auto-research, and, in turn, smarter…
Introducing SWE-1.7, the most capable model we’ve trained yet. It scores within a few points of the strongest frontier models at a fraction of the cost, and is now available at 1000 tok/s. RL is not hitting its limit: after refining our recipe, we keep seeing gains as we scale…
We raised $136M to kill Slack. Introducing PromptQL: The first AI version of Slack. Here’s how it works: https://t.co/PakqoBog3o
I had Fable build another thing I always wanted, a full procedural fantasy kingdom generator with economics, trade routes, population growth, wars, lineages, and occasional dragons. First, I worked with it on a plan, then it made it. You can play it here: https://t.co/vCtIqcG6P…
This is an insane Fable 5 UI/UX hack. This "Taste" skill completely kills generic AI-slop and gives Fable 5 the tools & instructions needed to ship beautiful design. This might just be the best AI skill I've ever used. https://t.co/XVQ6JpjYCy https://t.co/aQvw9rB2HJ
That's the spine. Fair hit. That's something to sit with. A real observation. That’s the whole thing. Sharpen that: say the word. Notice the arc of what just happened. One honest caveat: the full amount, stated plainly. Genuinely. Quietly. Honestly. That’s doing real work.
We've usually stayed away from model comparisons but 5.6 vs Fable is a unique situation We've never had a case where the team is so completely convinced on which one is better Here's the timeline of our experience with it - We test early versions of 5.6 for a couple of weeks …
mattpocock/skills v1.1 is out! - /wayfinder helps you plan more ambitious work than ever - /to-spec and /to-tickets replace /to-prd and /to-issues - /implement + /code-review complete the whole lifecycle - /research and /prototype help support wayfinder, or can be used independ…
we need a karpathy for non-technical people. a larpathy if you will
Fable is extremely good at writing original NES games in assembly. Here it designed an original game graphics, soundtrack, everything from scratch, and it'd actually run within the constraints of a real cartridge. Wild. https://t.co/C1Hns8dkyf
We audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and found it no longer reliably measures frontier coding capability. We find the eval to be saturated at a ~70% noise ceiling, and are retracting our previous recommendation that the research community …
Here's a step-by-step process to kill all the bloat from your Claude Code system prompt: 1. Run a proxy so you can see exactly what gets sent to Claude Code (included in the article) 2. "Fuck, there is so much cruft in there" 3. Use my settings.json to kill all the bloat Down …
Ranked by when the bookmark was saved during this London date range. Save rate means public X bookmarks divided by views, with a 10k-view minimum. Metrics are X's latest public counts; older items can be stale once they leave the recent-bookmarks sync window.