bookmarks
9 bookmarks

Everyone is tired of reading AI slop. Anthropic says Opus 5.5 writes more naturally and actually follows instructions, so we ran an eval. We tested it against the new versions of OpenAI's GPT-6 Luna and Sol to see whether models can solve a problem AND write a decent explanation…↗

As part of our efforts to pace the frontier, we’re committed to supporting independent assessments with deep levels of access across training, evaluation, and deployment. That access should enable third party assessors to challenge our assumptions, identify risks we may have mis…↗

Before we ship a new model, these teams try to break it. They build with it, push it to its limits, and tell us where it falls short. What they find makes the final model better.↗

We’re announcing: VibeBench, a new benchmark for what actually matters — how models feel when used on real work by experienced software engineers. But, we need your help. Here’s how it works: 1. An initial cohort of 1000 qualified software engineers (join: …↗

As a fun Saturday vibe code project and following up on this tweet earlier, I hacked up an **llm-council** web app. It looks exactly like ChatGPT except each user query is 1) dispatched to multiple models on your council using OpenRouter, e.g. currently: "openai/gpt-5.1", "googl…↗