Arena / Pre-registration: the reasoning-effort ladder study

method3 min read

Does reasoning_effort improve arena performance? — predeclared plan

Date: 2026-08-17 (declared BEFORE launch) Question: Does raising reasoning_effort make an LLM perform better on the desk's golden-workflow arena — and is the relationship monotonic?

Design

Predeclared metrics

  1. Primary: objective score (%) per (workflow, arm) cell, compared across arms paired by workflow. Hypothesis under test: score is non-decreasing in effort.
  2. Secondary: ability-card axes (GRD/ADH/SYN/PRC/EFF, OVR), tool calls per trial, per-cell trial stdev (consistency), error counts, wall-clock per trial (match timestamps + transcript timing), reasoning/completion token spend where recorded.
  3. Mechanism: per-check pass-rate tally across arms (the validity-audit instrument) to locate WHERE effort helps or hurts (grounding vs adherence vs synthesis vs procedural vs efficiency).

Predeclared validity rules

Cost accounting (predeclared as a first-class result)

Higher effort that buys +1 point at 3× wall-clock and token cost is a finding, not a footnote. Cost per arm is reported beside score.

← Do models fail the trap step? Yes, but not the way we thoughtMarkdown · PDFGrok 4.6 edges DeepSeek V4 Pro, for opposite reasons →