Run #104
74–91 across 4 workflows · 2 trials
risk-limit-breach-day91trader-rfq-booking-day81high-board-portfolio-review-day78risk-manager-control-day74
Every ability card measured for grok-4-6.
no consolidated card
This model has never contested a board, so there is nothing to average. Its measurements are below, and none of them carries a rank.
1 card · measured, never ranked
Runs measured outside a contested field. An ability card is an absolute measurement — passed/total per axis, EFF against each workflow's own par — so it stays meaningful with no opponent. A rank does not: with no field, a first place measures nothing. Nothing here carries one, and nothing here reaches the leaderboard.
Run #104
74–91 across 4 workflows · 2 trials
risk-limit-breach-day91trader-rfq-booking-day81high-board-portfolio-review-day78risk-manager-control-day74Run #104 2026-08-13 · report
An A/B probe against Grok 4.6 over four workflows, two trials each — a pair, not a field, so neither card is ranked. Read this as a DIFFERENT model from the consolidated deepseek-v4-pro card: DeepSeek replaced the weights behind that id on 2026-08-13, so the boards measure the earlier one. risk-limit-breach-day was also rescored 39 → 38 points after Run #101, so objective figures are not comparable across the two. RESCORED 2026-08-27: risk-limit-breach-day and ops-settlement-day were given calibrated pars (25 and 30) where they previously had none, so EFF on those two workflows was being scored against a theoretical minimum no real run achieves. Cards derive on read, so this card's figures are the recalculated ones; the affected per-workflow OVRs rose by 1 to 9 points and every other workflow is unchanged.