evals
21 saves·entity·peaked week of Jun 29, 2026
Weekly volume
Shows up alongside
Topics tagged on the same saves.
Who’s driving it
Most-saved authors on this topic.
Recent saves
The items behind the chart, newest first.
Using Claude Code: Spending your effort One of the best parts of our newest Claude models is how they respond to effort without breaking the prompt cache in Claude Code, but I’ve received a lot of q…
@trq212·Sep 25, 2026↗
Introducing SWE-2, our closest model yet to the frontier. On leading evals, it scores on par with recent frontier models – at up to 70% lower cost. We scaled RL to multiple trillions of parameters,…
@cognition·Sep 10, 2026↗
Your AGENTS.md is holding you back [Context engineering](https://posthog.com/newsletter/context-engineering?utm_source=posthog-newsletter&utm_medium=post&utm_campaign=agents-md) used to focus on add…
@posthog·Aug 31, 2026↗
Own Your Intelligence: A How-To Guide By @sonyatweetybird, @gradypb, and @w1tness1ngh The race for the AI application layer is not only about UI, workflows, or GTM... it is a fight for the intellig…
@sonyatweetybird·Aug 17, 2026↗
Introducing Ori Eval: the easiest way to write your first eval. There's no definitive best model, only the best model for each task. Ori Eval leverages OpenRouter's APIs for each task in your codeba…
@OpenRouter·Aug 3, 2026↗
Introducing Supabase Evals. Our benchmark for how well AI coding agents build with Supabase. We run agents like Claude Code, Codex, and Open Code against real tasks and score what they do. https://t…
@supabase·Aug 1, 2026↗
New content strategy: be the Robinhood of AI Take concepts/products/best practices from Token Street (SF/X/engineering bubble) & translate for Main Street (non-technical knowledge workers/execs). W…
@businessbarista·Jul 26, 2026↗
I interviewed @trq212 and @_catwu from the Claude Code team at @aiDotEngineer a couple of weeks ago - the video is now out, so I've published an annotated transcript of our conversation https://t.co/…
@simonw·Jul 21, 2026↗
You just hired a million bad employees. AI was supposed to replace human labor. It did the opposite. For the first time in history, humans are cheaper than software. And AI is creating more jobs…
@gsivulka·Jul 15, 2026↗
https://t.co/xv6csf1SbV
@satyanadella·Jul 13, 2026↗
https://t.co/Es6BXNRoYK
@thealexker·Jul 3, 2026↗
Pi's Edit Tool Thread on pi's edit tool since that came up in discussions at @aiDotEngineer, in part in relation to some people seeing edit failures even on SOTA anthropic models. pi's edit tool is…
@mitsuhiko·Jul 3, 2026↗
Want to try GLM 5.2 in production but worried how it might change your product? Don’t worry, we got you: 1. Install Inference Gateway (https://t.co/4bm9zThXUH) 2. Keep sending traffic to your curre…
@samhogan·Jun 30, 2026↗
Evals: the strategic IP that will define the next era of AI We've spoken to hundreds of execs in the past few months, and we're hearing a clear refrain: "AI isn't delivering ROI yet, but we're all i…
@GarrettLord·Jun 21, 2026↗
Google's free 5-day AI Agents course is back, and this time it's all about vibe coding with agents. The last one had 1.5M learners. Here's what the course covers: → Day 1: Agents + vibe coding. Bui…
@petergyang·Jun 2, 2026↗
🚀Introducing Motus Tracing: open-source observability for AI agents. Without traces, an agent is a black box that burns tokens. Yet most agent observability and tracing stacks today live behind acc…
@JiaZhihao·May 17, 2026↗
we built the first sane way to debug your agent locally. you can see your traces. codex/claude code can too. this lets them write evals and test your agents automatically. best part: it's complet…
@benhylak·May 15, 2026↗
Strong Opinions, Loosely Held on Agent + Harness Engineering: 1. You can outperform any default harness+model (including codex & claude code) on pretty much any Task by engineering the harness aroun…
@Vtrivedy10·May 7, 2026↗
https://t.co/3E2XKw18IW
@garrytan·Apr 22, 2026↗
New command: agent-browser skills Cached skills go stale when the CLI updates. Now the CLI serves skill content at runtime. One thin skill installed via: npx skills add vercel-labs/agent-browser E…
@ctatedev·Apr 12, 2026↗
Mastra is a TypeScript framework for building AI-powered applications and agents, including chatbots and evaluation workflows, with a modern Next.js/React stack.
@mastra-ai·Apr 7, 2026↗