deepseek-v4-pro

Every ability card measured for deepseek-v4-pro.

All workflows

averaged across the boards this model contested

88OVRdeepseek-v4-proSniper
96GRD90ADH96SYN88PRC74EFF89CON

84–96 across 4 of 5 boards

  • confirmation-desk-dayRun #133
  • risk-limit-breach-dayRun #10196#5
  • high-board-portfolio-review-dayRun #9488#4
  • trader-rfq-booking-dayRun #3386#3
  • risk-manager-control-dayRun #2084#4

Per board

4 measurements · each against that board's own field

risk-limit-breach-dayRun #101

96OVRdeepseek-v4-pro#5All-rounder
99GRD99ADH99SYN93PRC93EFF92CON

obj 97.4 · 2 trials

high-board-portfolio-review-dayRun #94

88OVRdeepseek-v4-pro#4Anchor
95GRD99ADH99SYN86PRC68EFF76CON

obj 94.3 · 2 trials

trader-rfq-booking-dayRun #33

86OVRdeepseek-v4-pro#3Playmaker
90GRD96ADH99SYN87PRC52EFF92CON

obj 92.8 · 2 trials

risk-manager-control-dayRun #20

84OVRdeepseek-v4-pro#4Sniper
99GRD68ADH86SYN86PRC82EFF96CON

obj 84.6 · 2 trials

Provisional

1 card · measured, never ranked

Runs measured outside a contested field. An ability card is an absolute measurement — passed/total per axis, EFF against each workflow's own par — so it stays meaningful with no opponent. A rank does not: with no field, a first place measures nothing. Nothing here carries one, and nothing here reaches the leaderboard.

Run #104

82OVRdeepseek-v4-proSniper
93GRD89ADH84SYN93PRC60EFF72CON

61–99 across 4 workflows · 2 trials

  • risk-limit-breach-day99
  • trader-rfq-booking-day90
  • risk-manager-control-day76
  • high-board-portfolio-review-day61

Run #104 2026-08-13 · report

An A/B probe against Grok 4.6 over four workflows, two trials each — a pair, not a field, so neither card is ranked. Read this as a DIFFERENT model from the consolidated deepseek-v4-pro card: DeepSeek replaced the weights behind that id on 2026-08-13, so the boards measure the earlier one. risk-limit-breach-day was also rescored 39 → 38 points after Run #101, so objective figures are not comparable across the two. RESCORED 2026-08-27: risk-limit-breach-day and ops-settlement-day were given calibrated pars (25 and 30) where they previously had none, so EFF on those two workflows was being scored against a theoretical minimum no real run achieves. Cards derive on read, so this card's figures are the recalculated ones; the affected per-workflow OVRs rose by 1 to 9 points and every other workflow is unchanged.