bookmarks
2 bookmarks

GPT 6 Astra is here. We ran the numbers on AutomationBench: It's the highest score we've ever recorded. Clean sweep across every domain. Scores 41.4% at Max effort. For context, no model had cleared 40% before today (GPT-5.6-Sol scored 28.8%) 𝗕𝗲𝘀𝘁 𝗳𝗶𝘁 𝗳𝗼𝗿: reconcilia…↗

We built an AI benchmark that measures real work. Today we're releasing it to everyone. AI evals tell you whether a model can do complex reasoning or generate code. Useful, but usually not the question our customers ask. They want to know: can this model find the right CRM reco…↗