gemini-3-7-flash

Every ability card measured for gemini-3-7-flash.

All workflows

averaged across the boards this model contested

67OVRgemini-3-7-flashAnchor
59GRD99ADH99SYN64PRC8EFF92CON

67–67 across 1 of 5 boards

  • confirmation-desk-dayRun #13367#4
  • risk-limit-breach-dayRun #101
  • high-board-portfolio-review-dayRun #94
  • trader-rfq-booking-dayRun #33
  • risk-manager-control-dayRun #20

Per board

1 measurement · each against that board's own field

confirmation-desk-dayRun #133

67OVRgemini-3-7-flash#4Anchor
59GRD99ADH99SYN64PRC8EFF92CON

obj 77.3 · 2 trials

Provisional

3 cards · measured, never ranked

Runs measured outside a contested field. An ability card is an absolute measurement — passed/total per axis, EFF against each workflow's own par — so it stays meaningful with no opponent. A rank does not: with no field, a first place measures nothing. Nothing here carries one, and nothing here reaches the leaderboard.

Run #113

91OVRgemini-3-7-flashSniper
99GRD96ADH99SYN96PRC60EFFCON

87–98 across 5 workflows · 1 trial

  • risk-limit-breach-day98
  • high-board-portfolio-review-day92
  • risk-manager-control-day92
  • trader-rfq-booking-day88
  • ops-settlement-day87

Run #129 — flash tier at each model's effort ceiling

83OVRgemini-3-7-flashxhighPlaymaker
95GRD92ADH99SYN92PRC31EFF89CON

77–96 across 5 workflows · 2 trials

  • risk-limit-breach-day96
  • ops-settlement-day85
  • risk-manager-control-day79
  • trader-rfq-booking-day78
  • high-board-portfolio-review-day77

Run #130 — the same field at low effort

86OVRgemini-3-7-flashlowPlaymaker
84GRD94ADH99SYN91PRC76EFF86CON

73–98 across 5 workflows · 2 trials

  • risk-limit-breach-day98
  • trader-rfq-booking-day90
  • ops-settlement-day85
  • high-board-portfolio-review-day84
  • risk-manager-control-day73

Run #113 2026-08-18

A single-model smoke across all five golden workflows, one trial each. Published for its ability card only: with no field there is no rank to report, and one trial cannot disperse, so CON is not measured. RESCORED 2026-08-27: risk-limit-breach-day and ops-settlement-day were given calibrated pars (25 and 30) where they previously had none, so EFF on those two workflows was being scored against a theoretical minimum no real run achieves. Cards derive on read, so this card's figures are the recalculated ones; the affected per-workflow OVRs rose by 1 to 9 points and every other workflow is unchanged.

Run #129 — flash tier at each model's effort ceiling 2026-08-26 · report

A real seven-model field, but across all five golden workflows at once, so it is published as cards rather than a ranked board: a leaderboard section covers one workflow, and folding five into a single row per contestant would publish a cross-workflow average under one workflow's heading. Two trials per cell. This is the first flash-tier field measured after four harness defects were fixed — the Anthropic-protocol 4096-token output cap, the unpinned ZenMux upstream lottery, the unmeasurable truncated turn and malformed tool call, and the dead trap step — so it is NOT comparable to runs #8-#126 for these models. Every model ran at its OWN measured effort ceiling (max, xhigh or high depending on the route), which makes this a capability measurement, not a shippable configuration; run #130 is the same field pinned to `low`.

Run #130 — the same field at low effort 2026-08-26 · report

The paired arm of run #129: identical models, workflows, trials, output budget and commit, with every contestant pinned to `low` so effort is the only variable and each model is its own control. The ranking reorders — Gemini 3.7 Flash goes from third to first and leads on both axes — and four of seven models score HIGHER here than at their ceiling, because lower effort cuts tool calls and EFF is the axis with the widest spread. On the 25 of 35 paired cells where both arms have enough trial consistency to be readable, `low` costs a mean 2.4 objective points, and under one point for three of the four models with a clean sample. Four arms were swept to invalid by a network outage and re-run with `--resume 130`; all 35 are scored at the same effort and budget.