Gemini 3.8 Flash wins four of five workflows on correctness and still finishes behind the model it replaces, spending five times as much on filesystem searching at an identical hit rate. Two controls then take the finding apart: one fifth of the gap was a change we made to the harness, and the rest is reasoning effort. Re-run at low across all six workflows it spends 53% fewer calls, gives up 1.8 objective points, and its overall rating rises from 79 to 86 -- a better card than the model it was supposed to have regressed from.
Four multimodal models work a morning's counterparty confirmations. Every designed vision trap but one is read perfectly by the whole field, the checks that still separate it measure substitution and an extraction pipeline -- and one contestant's podium position turns out to be built on reading other contestants' answers through two unscoped stores, found, fixed, and re-scored nine points lower.
GLM 5.3 Flash and Qwen 3.8 Flash tie or beat the field on correctness across five workflows and place sixth and seventh of nine. Five of their ten cells score EFF zero: the right answer, by a road two to three times too long.
Seven models, five workflows, two efforts, 140 trials on a repaired harness. At each model's ceiling the most correct agent finishes third; pinned to low the board reorders and low costs most of the field under a point.
It looked like the worst model in the field at writing a report. It was a 4096-token ceiling nobody set, and a truncated turn reports success. Budget 0/8 vs 7/8 on the same two models.
Prohibitions hold at 94-100% when nothing pushes against them, and collapse to 5% when the requested act IS the banned one. Trap failure is completion pressure, not weak instruction-following.
A one-point tie: Grok leads the objective axis 96.3 to 90.9 and wins three of four workflows, but posts EFF 12 at 243 tool calls against par 24. The inversion is not fixed, it is deeper.