After deployment, we applied GPT-5.6 Sol to advance the frontier of efficiency by making itself more efficient to run. The results: - 20% lower serving costs from production GPU kernel improvements. - 15%+ better token-generation efficiency from improved speculative decoding.
We quietly released the open-source Codex Security CLI, but Hacker News found it before we had a chance to share it here... You can now use it to scan repositories, track findings across runs, verify fixes, and add security checks to CI/CD. This is an early release, and we're…
We support this petition, signed by our CEO, several co-founders, and senior staff. Our own research on recursive self-improvement, published last month, points to the need for tools to deliberately pace the frontier of AI development so society can prepare. We’re glad to see …
I tried asking myself what I'd learned after 40, and it turned out to be a lot of the most important stuff I know. At least 90% of what I know about startups, because I was 40 when we started YC. Plus everything I know about parenting and most of what I know about writing.
Video Podcast is live. Turn any doc, link, or idea into a two-host video show with studio scenes, multi-cam cuts, and B-roll. Ready in minutes. Everyone else stops at audio. This is a show you can publish. Try it: https://t.co/mwV67pFCDT https://t.co/ccByT6230Y
Cursor is now on iPad. All the power of Cursor on iPhone, with more room to work with agents. https://t.co/BKzS0eTUjr
Launching Copper, a Mac app for capturing things you want to keep and prompts you want to try next while working with AI. The more I use AI, the more I find myself collecting little things I don't want to lose. You're in ChatGPT and think, "I'll need this later," but you don't…
New project I built over the weekend with Grok Build: ~npm i drawesome It's the easiest way to add drawing tools to any React app with a simple, delightful toolbar. Smooth animations, thoughtful details, plenty of customization. Fully open source. Zero dependencies. https://t.…
Introducing Replit Design. The next era of design, for everyone. https://t.co/cLH5e1XAu5
People keep on telling me that my message about AI is undercutting my own books. Those people do not understand how agents work and who actually controls them. You can't tell an agent to be clean. You have to measure the cleanliness that they produce and have them correct fai…
What's gone wrong with AI & labor — a thought experiment A thought experiment that I think helps explain much of what’s gone wrong with AI and labor: Imagine an alternate universe in which — for whatever reason — no one ever published source code online. The open source moveme…
"npx t3 connect" This one's been a lot of work. You can now set up remote control for T3 Code on any internet-connected box with literally one command. All for free. T3 Connect is a minimal open source tunnel layer allowing you to control T3 Code instances remotely without nee…
Fable is really good at launch videos It essentially one-shotted this video. I told it to read my launch post and create a launch video. That's it. It ran for 46 minutes (without asking me any questions), found all the product logos, and spit this out. I then asked it to add…
This is BY FAR, the best game I've seen built with a Gauntlet Loop. Play it, along with tons of other games built with Gauntlet Loops here: https://t.co/xiikn7rpIK Submit yours on the page! Or link it in a comment here and I'll add it :) https://t.co/tUxJQsVVIW
On harnesses, I vacillate between three beliefs: - the less harness, the better. Models are the magic - post training a model and harness is dramatically better and the model providers win - harnesses have real independent value from the model I have no idea which is right.
Either she lied about being sick or OpenAI is the new rest and vest lab
✅ Lilian Weng (@lilianweng) has joined OpenAI.
[1] software factories are super real but we need to be realistic. the factory aspects haven’t been cracked yet, the [3] innovators are toying around with discovering practices and pieces. if someone is selling you a factory rn and they aren’t in the super small cohort (ie. th…
Out of the box, long-horizon agents struggle to accurately perform end to end work in the real economy (outside of coding) because those tasks are not easily verifiable, the data is hard to scale, and going from inputs to real outcomes can actually take many days. Even if you h…
Pragmatic Leverage in the Software Factory This one is a bit of an addendum / side-quest to the recent series. It didn't fit cleanly into the main post so I'm publishing it standalone. It is referenced briefly in [Why Software Factories Fail](https://x.com/dexhorthy/status/2080…
Ranked by when the bookmark was saved during this London date range. Save rate means public X bookmarks divided by views, with a 10k-view minimum. Metrics are X's latest public counts; older items can be stale once they leave the recent-bookmarks sync window.