top of page
All Posts


DecisionBench: Measuring the Agent Handoff, Not Just the Answer
A benchmark for emergent delegation in long-horizon agentic workflows. We introduce DecisionBench, a benchmark for emergent delegation in long-horizon agentic workflows. It measures not just whether a task gets solved, but whether an agent hands its subtasks to the right peer model along the way. We characterize the benchmark with a five-condition reference sweep across an 11-model, 7-vendor pool, covering 23,375 task instances, and release the substrate, the annotation layer
May 229 min read


Introducing IntelligenceArena
A new way to measure intelligence in the age of AI agents Artificial intelligence has entered a new phase. Models are no longer evaluated in isolation, and systems are no longer defined by a single benchmark score. Today’s AI landscape is composed of agents, workflows, and continuously evolving tools that operate in real-world environments. Yet the way we evaluate these systems has not kept pace. Most benchmarks remain static, capturing capability at a fixed moment while igno
Apr 254 min read
bottom of page
