Provisional
2 cards · measured, never ranked
Runs measured outside a contested field. An ability card is an absolute measurement — passed/total per axis, EFF against each workflow's own par — so it stays meaningful with no opponent. A rank does not: with no field, a first place measures nothing. Nothing here carries one, and nothing here reaches the leaderboard.
Run #138 — the same model at low, the effort arm in full
86OVRgemini-3-8-flashlowPlaymaker91GRD92ADH99SYN94PRC47EFF—CON
71–98 across 6 workflows · 1 trial
ops-settlement-day98risk-limit-breach-day97trader-rfq-booking-day87risk-manager-control-day82high-board-portfolio-review-day80confirmation-desk-day71
Run #134 — Gemini 3.8 Flash, the first board after the thought-signature fix
79OVRgemini-3-8-flashxhighPlaymaker91GRD97ADH99SYN95PRC4EFF89CON
71–84 across 6 workflows
risk-limit-breach-day84ops-settlement-day83risk-manager-control-day83trader-rfq-booking-day81high-board-portfolio-review-day72confirmation-desk-day71
Run #138 — the same model at low, the effort arm in full 2026-09-04 · report
Gemini 3.8 Flash across all six golden workflows at low, one trial each — the paired arm for run #134 above, same model, same route, same harness, same commit, with reasoning effort as the ONLY variable. Read the two cards together: this one is why run #134 should not be read as a model regression. Dropping from xhigh to low cuts tool calls 53% (161.6 to 75.3 per workflow) and costs 1.8 objective points (95.6 to 93.8) — and OVR RISES from 79 to 86, because EFF goes from 3.7 to 46.8. Four of the six workflows score IDENTICALLY at both efforts, including both perfect 100.0s, so the accuracy that xhigh buys is confined to two workflows and is worth less than the calls it spends. The effort effect on searching is NOT uniform: on five workflows filesystem search all but vanishes at low, but on the vision workflow it barely moves (151 calls against 159), so that workflow's rummaging is inherent rather than effort-driven. ONE TRIAL per cell, so CON is not measured rather than perfect, and the figures carry real single-sample noise — ops-settlement-day was measured twice at low, in run #137 and here, and scored 93.2 then 100.0.
Run #134 — Gemini 3.8 Flash, the first board after the thought-signature fix 2026-09-02 · report
Gemini 3.8 Flash across all SIX golden workflows at xhigh, published as cards only: one model has no field, so a rank of #1 of 1 would measure nothing, and six workflows cannot be folded into a single leaderboard row. Two trials per cell EXCEPT confirmation-desk-day, which carries ONE — its second trial was lost to a ZenMux 402 quota exhaustion and recorded invalid/infra_blank rather than scored, so that workflow's CON is not measured rather than perfect. Zero truncation and zero malformed tool calls across all six arms, measured rather than assumed. This is the FIRST board to run after the Gemini thought-signature fix. Gemini 3 routes return an encrypted signature beside every tool call and require it echoed back on the assistant turn carrying that call; ChatOpenAI does not preserve that field by design, so the second model turn of every tool loop was refused with a 400 and this run initially failed on all six workflows. The control is what settles the reading: gemini-3.7-flash, which has four published boards on this same route, fails identically once the field is stripped. So the enforcement is new, not the model, and no earlier board's numbers move. The finding is correctness bought at a price that outweighs it. Against the gemini-3.7-flash arm of run #129 — same route, same xhigh effort, the five workflows they share — 3.8 is 2.8 objective points BETTER (97.1 against 94.3, ahead on four of the five, with a perfect 100.0 on two workflows). It spends 84% MORE tool calls to get there (96 against 52), and EFF collapses from 31 to 4. The efficiency loss more than eats the correctness gain: mean OVR over those same five workflows is 80.6 against 3.7's 83.0, so AT THIS EFFORT the newer model finishes 2.4 points behind its predecessor while being the more accurate of the two. EFF is 0 on four of the six workflows, and the vision workflow alone took 490 tool calls. CORRECTED 2026-09-04 — read this card as a measurement of xhigh, not of the model. A follow-up arm (run #137) pinned the same model to low on ops-settlement-day: it cut filesystem search calls by 94%, from 36 to 2, and total calls by 56%, posting OVR 89 there against 83 at xhigh and 85 for its predecessor — the best cell measured on that workflow. The over-execution this card reports is substantially an effort setting rather than a model trait, and the desk default is low. A low arm across all six workflows has not been measured, so no low card is published.