Run #131 — two flash-tier newcomers at max effort
63–90 across 5 workflows · 1 trial
risk-limit-breach-day90risk-manager-control-day82trader-rfq-booking-day81ops-settlement-day75high-board-portfolio-review-day63
Every ability card measured for qwen-3-8-flash.
no consolidated card
This model has never contested a board, so there is nothing to average. Its measurements are below, and none of them carries a rank.
1 card · measured, never ranked
Runs measured outside a contested field. An ability card is an absolute measurement — passed/total per axis, EFF against each workflow's own par — so it stays meaningful with no opponent. A rank does not: with no field, a first place measures nothing. Nothing here carries one, and nothing here reaches the leaderboard.
Run #131 — two flash-tier newcomers at max effort
63–90 across 5 workflows · 1 trial
risk-limit-breach-day90risk-manager-control-day82trader-rfq-booking-day81ops-settlement-day75high-board-portfolio-review-day63Run #131 — two flash-tier newcomers at max effort 2026-08-27 · report
GLM 5.3 Flash and Qwen 3.8 Flash, published 2026-08-26 and 2026-08-27, across all five golden workflows at `max` — cards only, because a two-model run over five workflows has no field to rank within and folding five workflows into one row would publish a cross-workflow average under a single workflow's heading. ONE TRIAL PER CELL, so CON is not measured rather than perfect, and every figure is a single sample; the run #129 cards these are read against were measured at two trials. Both models are competitive on correctness — they post the highest objective score any contestant has recorded on trader-rfq-booking-day (95.2) and tie the best on the flagship (94.9) — and place sixth and seventh of nine on EFF alone, which is zero in five of their ten cells. Both are pinned to the Anthropic wire protocol: Qwen's `alibaba` upstream returns tool-call continuation deltas with an empty id (263 of 263 measured), which would otherwise floor every match at 7.7.