Benchmark

What the harness costs you, on the same model

1 harness plan in roster.json, nothing scored yet

No harness measurement has been scored yet

This page draws only the campaigns the exporter names as harness measurements: one model held still, the harness varied. It fills in when the first of them is scored.

What is held equal, and what is not

The difference that is left is the harness tax being measured

Held equal

  • The model, the prompt, the effort and the pull requests, sealed per arm.
  • The tools each of the 7 arms may use:
    • afi: its own released policy, which the bench does not pin.
    • opencode: read, list, glob, grep; task, edit, bash and web tools denied.
    • Claude Code: Read, Grep, Glob; Agent and Task denied; no CLAUDE.md or AGENTS.md loaded from the reviewed tree.
    • oh-my-pi: read, grep, glob; agent, plugin and context-file providers off.
    • pi: read, grep, find, ls.
    • Goose: directory_tree, get_file_info, list_directory, read_multiple_files, read_text_file, search_files; every built-in extension off; reads through the filesystem MCP server only.
    • Crush: view, grep, glob, ls; agent and every other built-in tool off.
  • The price: every arm's token counts times the same rates, so a cost is comparable where a harness's own figure is not.

Not held equal

  • Each harness's own system prompt, context compaction and retry policy - the tax itself.
  • opencode also has skill and todowrite.
  • afi runs the tool policy it ships, which is what its arm measures.
  • Turns and tool calls are each harness's own count of its requests and tool uses, not one shared definition.
benchee benchee-dashboard-2 built from fb1dc98e Static benchmark evidence ·