9 bookmarks
lol did nobody at Anthropic stop for a second and wonder why the numbers looked this absurd before posting the “victory”-tweet? openai.com/index/how-two-…↗
We’re introducing GeneBench-Pro, a research-level benchmark for a harder kind of AI progress: how well agents can navigate messy biological data, choose the right analysis path, and make judgment calls that real computational research depends on. openai.com/index/introduc…↗


