bookmarks
3,946 bookmarks

We’re demonstrating how frontier models have continued to improve in realistic mental health conversations with MentalHealthBench. This new open benchmark was built with input from more than 80 mental health clinicians. We’re releasing it openly so other researchers can examine…↗

Announcing: the @a16z Ops Engineering Fellowship ⚙️ Because: operations (along with everything else??) is becoming something best approached with an engineering mindset. We've been tracking this space closely, and are seeing a new kind of builder emerging inside companies... Th…↗

We heard you loud and clear. ChatGPT Voice can now: - Use plugins like your email, calendar, and Slack. - Be powered by GPT-6 Astra, Sol, and Luna. - Be used in ChatGPT Work on web and mobile, so you can create docs, decks, sites, and spreadsheets or tackle complex tasks in th…↗

Today at ZapConnect, we debut Next Gen Zaps. One of the most important product moments in @Zapier's history. With Next Gen Zaps, customers can deploy agent-built Zaps in seconds. Simply describe your problem to your agent and it will build you a Zap that runs automatically, ev…↗

·@wadefoster·Sep 23, 2026·launch·chatgpt·claude·next-gen-zaps·zapconnect

Everyone is tired of reading AI slop. Anthropic says Opus 5.5 writes more naturally and actually follows instructions, so we ran an eval. We tested it against the new versions of OpenAI's GPT-6 Luna and Sol to see whether models can solve a problem AND write a decent explanation…↗

holy shit, opus 5.5 is kind of insane at animation. i didn't write a single line of code. it wrote the story, drew every frame, and made the music. no image assets at all, it's all JS. the story is called "small print": claude gets a pile of requests every day. it circles the hu…↗

·@Voxyz_ai·Sep 23, 2026·demo·claude·hyperframes·javascript·opus-5-5

I've asked Claude Opus 5.5 to create 100 HTML files Rules were simple: - look stunning - zero repetitive designs - go full creative mode CLAUDE OPUS 5.5 IS THE BEST MODEL I HAVE EVER TESTED !!! The previous best results on this experiment was of Fable 5.1, but Opus 5.5 is NEXT…↗

Opus 5.5 is way, way, way better than Opus 5. Sorry about that model, please try this one.↗

·@__nmca__·Sep 23, 2026·model-release·opus·opus-5·opus-5-5

We tested Claude's Opus 5.5 for a week. Some tips: 1. Set a time goal. It won't stop on its own, and it'll eat your weekly limit. Say "10 minutes, produce X." 2. Name the deliverable. We asked for a run of show and got training materials and handouts but it forgot to give out a…↗

i love prompting for a full walkthrough of a flow i.e here i had it show me all of the steps involved in oauth and found redundant steps, layout shift, and some weird ui state agents don't have to produce slop, they can also help you make great software↗

·@RhysSullivan·Sep 22, 2026·workflow·agent-software·oauth·prompting·ui-state

I asked Opus 5.5 to try a bunch of redesigns of my personal website using my MAX sub, using workflows to iterate and critique. I was really pleased with how it hit the tone I was looking for. Then I asked it to make a trailer with all of its iterations.↗

Claude Opus 5.5 is the best LLM for Operations we've benchmarked. It scored 40% on @Zapier's AutomationBench: 657 of the hardest real business workflows we have, across 6 departments. It excelled at Ops work where the model has to go find the procedure and use several tools ac…↗

I was invited to test Opus 5.5. The differences at the high end are hard to notice unless you're being ambitious. They reward ambition though. If you can't tell the difference here, or if you're intentionally micro-prompting a model, I think the fast models (composer, Gemini f…↗

GPT-6 Sol and Luna are out. Not only are they a very significant improvement across the board, but also in writing and general "you know when you try it" quality. We are also permanently reducing the API price by 50% making both of them viable for a ton of new usecases and maki…↗

·@thsottiaux·Sep 22, 2026·model-release·api-pricing·gpt-6·luna·plus

Opus 5.5 is here! 🎉 It's honestly such a nice model to work with. Feels like Fable, but ~30% faster and ~40% cheaper per task than Opus 5. Go try it out!↗

·@lydiahallie·Sep 22, 2026·model-release·cost-reduction·fable·model·opus-5-5

opus 5.5 = personality of opus 4.6 that we all *desperately* wanted back + the intelligence & taste of fable 5.1. and a 25% usage bump! i honestly don’t really see a reason to use another model right now? s-tier release.↗

·@mckaywrigley·Sep 22, 2026·model-release·fable-5-1·opus-4-6·opus-5-5

GPT-6 Sol and Luna just landed in Astra’s orbit. Both launch today with API prices 50% lower than GPT-5.6. Build with Sol. Scale with Luna. To production and beyond.↗

·@OpenAIDevs·Sep 22, 2026·launch·api-pricing·astra·gpt-6·luna

I asked Opus 5.5 to make a game inspired by the discovery of the first computer, the Antikythera mechanism. Everything you see and hear is procedurally generated in real time, from code by Opus 5.5. It's a single 3 MB HTML file. Playable Claude artifact below.↗

Please welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe. GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale. We’ve also made caching and inference more efficient, and…↗

·@OpenAI·Sep 22, 2026·model-release·api-pricing·caching·gpt-6-luna·gpt-6-sol

As part of our efforts to pace the frontier, we’re committed to supporting independent assessments with deep levels of access across training, evaluation, and deployment. That access should enable third party assessors to challenge our assumptions, identify risks we may have mis…↗

One more thing: we’re increasing five-hour usage limits on Pro, Max, and Team plans. We’re also providing subscription users a rate limit reset, which you can save and use whenever you choose.↗

·@claudeai·Sep 22, 2026·news·claude·max·pro·rate-limit-reset