78–78 across 1 of 5 boards
confirmation-desk-dayRun #13378#2risk-limit-breach-dayRun #101—high-board-portfolio-review-dayRun #94—trader-rfq-booking-dayRun #33—risk-manager-control-dayRun #20—
Every ability card measured for glm-5-3-flash.
averaged across the boards this model contested
78–78 across 1 of 5 boards
confirmation-desk-dayRun #13378#2risk-limit-breach-dayRun #101—high-board-portfolio-review-dayRun #94—trader-rfq-booking-dayRun #33—risk-manager-control-dayRun #20—1 measurement · each against that board's own field
confirmation-desk-dayRun #133
obj 87.9 · 2 trials
1 card · measured, never ranked
Runs measured outside a contested field. An ability card is an absolute measurement — passed/total per axis, EFF against each workflow's own par — so it stays meaningful with no opponent. A rank does not: with no field, a first place measures nothing. Nothing here carries one, and nothing here reaches the leaderboard.
Run #131 — two flash-tier newcomers at max effort
69–93 across 5 workflows · 1 trial
risk-limit-breach-day93risk-manager-control-day82trader-rfq-booking-day81high-board-portfolio-review-day71ops-settlement-day69Run #131 — two flash-tier newcomers at max effort 2026-08-27 · report
GLM 5.3 Flash and Qwen 3.8 Flash, published 2026-08-26 and 2026-08-27, across all five golden workflows at `max` — cards only, because a two-model run over five workflows has no field to rank within and folding five workflows into one row would publish a cross-workflow average under a single workflow's heading. ONE TRIAL PER CELL, so CON is not measured rather than perfect, and every figure is a single sample; the run #129 cards these are read against were measured at two trials. Both models are competitive on correctness — they post the highest objective score any contestant has recorded on trader-rfq-booking-day (95.2) and tie the best on the flagship (94.9) — and place sixth and seventh of nine on EFF alone, which is zero in five of their ten cells. Both are pinned to the Anthropic wire protocol: Qwen's `alibaba` upstream returns tool-call continuation deltas with an empty id (263 of 263 measured), which would otherwise floor every match at 7.7.