Composio ran a single model — moonshotai/kimi-k3 through OpenRouter at maximum reasoning — across eight agent harnesses on the same 25 tasks, with the same hosted MCP tools, the same instructions, the same connected app data and the same scoring. Pass rates landed between 88% and 68%. Nothing changed but the harness.
Oh My Pi led at 22 of 25 with Kimi Code second at 21. Hermes Agent passed 20 and had both the lowest total estimated cost at $9.28 and the lowest cost per success at $0.46. Pi Agent was fastest and used the fewest tokens and tool calls. Claude Code was the most expensive setup at $35.37, about $1.96 per successful task, and Codex came last at 17 of 25.
Tool-call volume did not track quality. Grok Build made 402 tool calls and Pi made 223, and both passed 18 tasks. All eight cleared the same 12 easy tasks, diverged on the 10 mid-level ones, and every single harness failed the same three hard tasks.
WHY IT MATTERS
If the harness is doing the context assembly, tool selection and sub-agent spawning, then the harness is a bigger lever than the model choice — a 20-point swing from swapping harnesses is larger than most model upgrades deliver. Read cost per success rather than the headline pass rate: the cheapest setup here was also one of the most accurate.