bookmarks
9 bookmarks

Today we're announcing Base Labs, a dedicated research organization focused on advancing open-source AI. We believe in a healthy, open frontier model ecosystem. To enable this, we are working on: - Blue-sky research on continual learning, the science of RL, and how models learn…↗

Introducing Supabase Evals. Our benchmark for how well AI coding agents build with Supabase. We run agents like Claude Code, Codex, and Open Code against real tasks and score what they do.↗

·@supabase·Aug 1, 2026·tool·benchmark·claude-code·codex·supabase

GPT-5.6 Sol has been used to solve open problems in mathematics. So why was it struggling with ARC-AGI-3, a benchmark of 2D puzzle games? We investigated. The harness was not letting it remember what it had learned. We found that enabling two API settings tripled our scores wi…↗

·@OpenAI·Jul 30, 2026·benchmark·2d-puzzle·api-settings·arc-agi-3·benchmark

We audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and found it no longer reliably measures frontier coding capability. We find the eval to be saturated at a ~70% noise ceiling, and are retracting our previous recommendation that the research community u…↗

Introducing LifeSciBench, a benchmark for measuring and improving how well AI supports real-world life science research. Developed with 173 scientists from biotechnology and pharmaceutical research, LifeSciBench includes 750 expert-authored tasks across seven biological research…↗

We’re announcing: VibeBench, a new benchmark for what actually matters — how models feel when used on real work by experienced software engineers. But, we need your help. Here’s how it works: 1. An initial cohort of 1000 qualified software engineers (join: …↗