I got 96.2% on ARC-AGI-3 with Opus 5, and 99.3% pass@2. The program is basically Claude Code + Opus 5 (high), one action command, and filesystem logs. Almost nothing ARC specific. github.com/jerber/arc-code↗
@beffjezos Grok 4.7 is significantly better than 4.6 and should be ready in 3 to 4 weeks. Initial training is complete and now we’re adding a massive amount of SpaceX company data in supplemental training. This will be something special.↗
@cognition Grok 4.7 will exceed all current models. That said, Anthropic is a great company and will probably release improved models soon. However, the SpaceX training corpus is so awesome & unique that I would be shocked if any model is better at real-world engineering than…↗
‼️huge ssi news. ilya is about to take his first tentative steps out of the age of research and back into the age of scale. it’s time to smell what ssi is cooking. ssi have built a small reasoning engine that can compete with much larger training runs because his data is bette…↗
We built a software factory for AI SDK. Each step is an agent, and humans merge changes. Four weeks in: ▪️ The factory authors up to 35% of merged PRs ▪️ It closed 70% of issues in July ▪️ Open bugs are down 25% vercel.com/blog/building-…↗
Two months in after relaunch, really proud of the @digg crew: @addison and @JustinMezzell @maubrowncow -- 100+ running tasks/agents/judges, automated clustering, sentiment analysis, image creation, newsletters, everything... traffic is up 20%+ MoM, @grok models have been great at…↗
Today, we're announcing @ref_tools open beta. Ref is your team’s shared space to decide what agents build before any code gets written. We’ve raised $4M from @villageglobal, @Daybreak_Fund, @tmrohan, @RitualVC, @tonigemayel @barrald @wesmckinn and @bentossell. Every engineerin…↗
Given the state of AI we must internalize: Everything hackable will get hacked. Two forces are at work: ▪︎ Near-frontier open weight models hack without refusal ▪︎ Frontier models help you defend yourself How we're approaching this new reality ↓ vercel.com/blog/everythin…↗
I made an Open Source version of the Grok Bot complete with all the features It does not need any subscriptions at all and uses the existing subscriptions you already have It can spin up virtual machines from @asciidotdev It uses @trycua for computer use It also has plugin s…↗
Is your AGENTS.md helping AI coding agents do their best work? A well-structured AGENTS.md gives agents the right context without overwhelming them. @mattpocockuk shares why keeping instructions focused, using progressive disclosure, and organizing documentation can improve agen…↗
Had @RyanGreenblatt on to discuss/debate recursive self-improvement. This might be the most important question in the world right now - whether within a year or so of achieving human level intelligence, you slingshot towards having 10s of billions of superintelligences, each of…↗
understanding progress on RSI is one of the most important things we can do to chart the future. @danrobinson and I made a game and explorer based on economic models of AI research, to make the inputs and constraints that define RSI more intuitive. paradigm.xyz/research/rsi/?…↗
I made a new skill: /interface-review It reviews your work across multiple categories like UI, typography, layout, color, writing and accessibility and gives you a detailed analysis of the findings. github.com/jakubkrehel/sk…↗








