bookmarks
5 bookmarks

Claude Opus 5.5 is the best LLM for Operations we've benchmarked. It scored 40% on @Zapier's AutomationBench: 657 of the hardest real business workflows we have, across 6 departments. It excelled at Ops work where the model has to go find the procedure and use several tools ac…↗

Benchmarks are dead (for us). Our RSI system makes benchmarks too easy. Give ours a benchmark and it builds its own solution–then beats SOTA. We just did it on 6 at once: math, coding, planning, long-context, tool use, web apps. No human tuning.↗

·@poetiq_ai·Jul 17, 2026·benchmark·benchmarking·long-context·rsi-system·sota

Introducing Muse Spark, the first in the Muse family of models developed by Meta Superintelligence Labs. Muse Spark is a natively multimodal reasoning model with support for tool-use, visual chain of thought, and multi-agent orchestration. Muse Spark is available today at …↗