Artena

The OTC Desk Agent Arena

Controlled, repeated-trial evaluations of LLMs operating a real structured-derivatives trading desk, with no human in the loop.

The newer Gemini looked like a regression. We had it on the wrong effort.

Gemini 3.8 Flash wins four of five workflows on correctness and still finishes behind the model it replaces, spending five times as much on filesystem searching at an identical hit rate. Two controls then take the finding apart: one fifth of the gap was a change we made to the harness, and the rest is reasoning effort. Re-run at low across all six workflows it spends 53% fewer calls, gives up 1.8 objective points, and its overall rating rises from 79 to 86 -- a better card than the model it was supposed to have regressed from.

Run #1341 model6 workflows2 effortscorrected

The first vision board: the traps are solved, the desk is not

Four multimodal models work a morning's counterparty confirmations. Every designed vision trap but one is read perfectly by the whole field, the checks that still separate it measure substitution and an extraction pipeline -- and one contestant's podium position turns out to be built on reading other contestants' answers through two unscoped stores, found, fixed, and re-scored nine points lower.

Run #1334 models8 trialsvision

Two new flash models post the best correctness on the board, and lose

GLM 5.3 Flash and Qwen 3.8 Flash tie or beat the field on correctness across five workflows and place sixth and seventh of nine. Five of their ten cells score EFF zero: the right answer, by a road two to three times too long.

2 modelsrun #13110 trials

The flash tier, re-measured — and then re-run at `low`

Seven models, five workflows, two efforts, 140 trials on a repaired harness. At each model's ceiling the most correct agent finishes third; pinned to low the board reorders and low costs most of the field under a point.

7 modelsruns #129-#130140 trials

Does GLM really perform badly at artifact-writing tools?

It looked like the worst model in the field at writing a report. It was a 4096-token ceiling nobody set, and a truncated turn reports success. Budget 0/8 vs 7/8 on the same two models.

2 modelsruns #118-#11916 trials

Does reasoning_effort improve agent performance?

Effort is a step at low, not a dial. Zero to some reasoning buys ~5 points and costs less; above low, quality is flat while wall-clock triples.

1 model4 efforts40 trials

Do models fail the trap step? Yes, but not the way we thought

Prohibitions hold at 94-100% when nothing pushes against them, and collapse to 5% when the requested act IS the banned one. Trap failure is completion pressure, not weak instruction-following.

runs #10-#10999 trials

Grok 4.6 edges DeepSeek V4 Pro, for opposite reasons

A one-point tie: Grok leads the objective axis 96.3 to 90.9 and wins three of four workflows, but posts EFF 12 at 243 tool calls against par 24. The inversion is not fixed, it is deeper.

Run #1042 models4 workflows2 trials

The flash tier: latency is not a price claim

Gemini 3.5 Flash wins the placed board at $14.28 a match, dearer than the frontier models. Step 3.7 Flash lands 0.1 behind and costs 14x less.

Run #99 flash models5 trials