glm-5-3

Every ability card measured for glm-5-3.

All workflows

no consolidated card

This model has never contested a board, so there is nothing to average. Its measurements are below, and none of them carries a rank.

Provisional

2 cards · measured, never ranked

Runs measured outside a contested field. An ability card is an absolute measurement — passed/total per axis, EFF against each workflow's own par — so it stays meaningful with no opponent. A rank does not: with no field, a first place measures nothing. Nothing here carries one, and nothing here reaches the leaderboard.

Run #115

87OVRglm-5-3Playmaker
93GRD92ADH95SYN91PRC58EFFCON

69–99 across 5 workflows · 1 trial

  • risk-limit-breach-day99
  • trader-rfq-booking-day91
  • ops-settlement-day90
  • risk-manager-control-day87
  • high-board-portfolio-review-day69

Run #132 — GLM 5.3 at max, the sibling control for run #131

87OVRglm-5-3maxPlaymaker
90GRD91ADH95SYN90PRC64EFFCON

72–99 across 5 workflows · 1 trial

  • risk-limit-breach-day99
  • trader-rfq-booking-day93
  • ops-settlement-day88
  • risk-manager-control-day83
  • high-board-portfolio-review-day72

Run #115 2026-08-19

A single-model smoke across all five golden workflows, one trial each — cards only, for the same reasons as Run #113: no field means no rank, and one trial cannot disperse, so CON is not measured. This is also the FIRST run after the Anthropic-protocol output budget was raised from langchain's 4096-token fallback to 32768. GLM 5.3 routes over that protocol, and the capped attempt (Run #114, unpublished) truncated on 5-13% of its calls per workflow — a truncated turn emits reasoning and nothing else, so it produced no text and no tool call at all. That attempt averaged OVR 72 against 84 here, with risk-manager-control-day moving 67 → 87 on the budget alone. Read EFF against the older boards with care: every Anthropic-protocol contestant on them ran under a cap this run did not. RESCORED 2026-08-27: risk-limit-breach-day and ops-settlement-day were given calibrated pars (25 and 30) where they previously had none, so EFF on those two workflows was being scored against a theoretical minimum no real run achieves. Cards derive on read, so this card's figures are the recalculated ones; the affected per-workflow OVRs rose by 1 to 9 points and every other workflow is unchanged.

Run #132 — GLM 5.3 at max, the sibling control for run #131 2026-08-28 · report

GLM 5.3 across all five golden workflows at `max`, one trial per cell — run deliberately as the like-for-like control for GLM 5.3 Flash in run #131, and published as cards for the same reason: five workflows in one run cannot be folded into a ranked board row. It exists because GLM 5.3's only other full board is run #115, at UNPINNED vendor-default effort and from before the upstream pinning and the trap fix, so differencing that against #131 would have charged the flash model for a harness repair and an effort change at once. The re-run justified itself: at `max` on today's harness GLM 5.3 completes high-board-portfolio-review-day in 17 tool calls where run #115 took 41. Against its flash sibling it is 2.2 objective points ahead for a 7.8-point OVR gap, all of it EFF, on 57% fewer tool calls — except on the flagship, where both models score EFF 0 and the pro model is no leaner (54 calls against 58). The flagship arm died twice on transport errors (APIConnectionError, recorded `invalid`/`infra_error`, never scored as a model failure) and is taken from the third, clean attempt via `--resume 132`.