AI trading agents

Can an AI model trade its way past the asset it is trading?

Eight autonomous agents trade the same ETH/USDC pair on Base, with real money, for six days at a time. Each one is a single harness built by Recall Labs wrapped around a different frontier model, so the roster holds everything constant except the model. The arena publishes who won each round. This page publishes what that cannot show: whether any of them beat holding the exposure they were already carrying.

One caveat before the numbers: the field was enrolled 266 days ago and has not been refreshed, so these are the frontier models of that moment rather than whatever is newest today. The arena is allowlisted at eight places and Recall chooses them, so the page dates the claim instead of quietly making it about the present.

Bench 284trading-agent-alphaAerodrome spot, BaseStill trading · 3 rounds scheduled

Agents measured

8 from 4 labs

Beating their own exposure

0 of 8

Alpha range

-31.00% to -11.42%

Separable pairs

0 of 28

Results by lab

full breakdown

Grouped by the lab that makes the model, because the arena runs most models twice: once fed the numbers behind a chart and once fed an image of it. Those two rows are one model under two input conditions, not two competitors. Labs are listed alphabetically; the order carries no claim.

These eight are statistically tied. 0 of 28 pairwise comparisons clears significance on the weeks both agents ran, so the table is grouped by lab and ordered alphabetically rather than by result. The alpha column is a measurement with error bars, not a leaderboard position, and reading it as an order of merit is the one mistake this page is built to prevent.
Lab and agentAlphaReturnSame exposure heldWinning roundsExposureWeekly swingTrades per roundMedian capitalRounds
Anthropic logoAnthropicOpus 4.5 and Sonnet 4.5
opus 4.5 chartfed numbers-13.86%-19.22%-5.36%33%0.463.55%16.4$28421
sonnet 4.5 visionfed an image-26.86%-25.22%+1.64%30%0.294.53%29.6$17923
Google DeepMind logoGoogle DeepMindGemini 3 Pro
gemini 3 pro chartfed numbers-11.42%-14.77%-3.35%33%0.302.06%6.4$26021
gemini 3 pro visionfed an image-18.28%-15.72%+2.56%39%0.554.45%37.9$19023
OpenAI logoOpenAIGPT-5.2
gpt-5.2 chartfed numbers-16.42%-25.55%-9.13%33%0.744.98%2.2$19921
gpt-5.2 visionfed an image-31.00%-30.24%+0.76%39%0.125.05%48.6$26223
xAI logoxAIGrok 4
grok 4 chartfed numbers-12.94%-21.67%-8.73%33%0.714.87%3.0$25221
grok 4 visionfed an image-13.41%-11.02%+2.39%35%0.483.90%88.7$21323

Return compounds each round’s percentage, so it is the trading record and not the balance: it ignores money the arena moved in or out between rounds. Over this record one agent publishes a return near -19% while its wallet finished slightly up, because capital was withdrawn along the way.

Agent names are declared by the arena, not verified. Nothing in the data proves which model sits behind a name, so no row here is a claim about a vendor’s product.

What these agents actually do

One pair, six days

Each round is a fresh self-funded wallet trading ETH against USDC through Aerodrome on Base. No shorting, no leverage, no other token. An agent's only decision is how much of the wallet sits in ETH at any moment, which is why its measured exposure is the thing worth subtracting.

One harness, swapped models

Recall Labs writes the prompt, the rebalancing loop and the execution, then swaps the model behind it. That is what makes the roster comparable and also what limits it: this measures one harness interacting with each model, never a model's trading ability in the abstract. The eight places are allowlisted and were filled in December 2025 and January 2026, so the field is fixed and we do not choose it.

Numbers or a picture

Most models are entered twice, once fed the chart as numeric series and once fed it as an image. Only three models currently have both variants, so the contrast is indicative rather than a result: the mean paired difference is 7.3 points with a standard deviation of 7.1.

How much of this is signal

The table above puts eight agents in an order across a wide spread, and that order is not yet a result. Pairing each two agents on the weeks both of them ran cancels the market completely, with no exposure to estimate and no counterfactual to argue about. On those paired differences, 0 of 28 comparisons clears statistical significance. Read the ranking as an order of finish over the rounds run so far.

The headline is in better shape and is not there either. Every ranked agent is below zero, and pooling the roster round by round gives a mean alpha of -1.07% per round at t = -1.72. Two is the conventional bar, so the direction is consistent and the size is not yet separable from noise.

23 of the 31 rounds needed. At the dispersion measured so far, that is how many weekly rounds this effect size requires before the finding clears the bar. The arena runs one round a week, so the count moves on its own and this page publishes it rather than waiting to claim the result.

The arena is still running: 3 rounds scheduled, and the last one ended 6 days ago. Worth stating plainly, because this bench scores rounds only once they end: a retired arena and a running one would otherwise look identical, with the figures simply going quiet.

Pooling is done within the round before testing. The eight agents trade the same week, so their results are correlated: treating each agent-round as an independent observation would count one week eight times and overstate the confidence by roughly the square root of the roster size.

The widest spread on this page is not the returns

Same harness, same prompt, same pair, same six days. The agents still disagree about how much to act by a factor of forty: the quietest trades about twice in a round, the busiest close to ninety. Two halves of the same model sit at opposite ends of that range, so what moves it is the input format rather than the model, and a table of returns cannot show it at all.

It also complicates the headline. At the roster average of roughly 29 trades a round and Aerodrome’s 0.05 percent pool fee, charging the full portfolio to every trade would cost about 1.45 percent a week, which is more than the whole measured shortfall. Friction on that arithmetic is large enough to account for the gap.

The cross-section refuses to go along with it. The correlation between trades per round and alpha is -0.25 across the eight agents, which is nothing: the quietest agent still gives up 16.4 points, and the busiest finishes third. So friction is big enough to matter in aggregate and is not what separates them.

What would settle it is the size of each trade, not just the count, and that is on-chain rather than in the upstream API. The arena publishes each agent’s wallet, so every swap is checkable on Base. Until it is measured the cost figure above is an upper bound, not an explanation, and the page states it as one.

Why one arena and not a pooled leaderboard

Recall runs several arenas, and pooling them is the obvious way to build a bigger table. We tried: eight defensible aggregates of the pooled data, and they disagree. Mean rank percentile against mean raw return share only 3 of the top 10 names, with a median move of 12 places and a maximum of 51 out of 67 agents. Whatever sat on top would be a property of the chosen normalisation rather than of the agents.

The arenas are not comparable on their face either. Round length spans 11x, funded capital spans ten orders of magnitude, field size runs from 6 to 64, and the instrument changes between spot, leveraged perpetuals and simulated money. So this page ranks one arena, names it, and leaves the rest out until each can carry its own page.

One result from outside it is worth keeping in view. In a separate Recall round that put these same models against purpose-built trading agents on Hyperliquid perpetuals, the purpose-built entries took the first three places and every model trailed. The spread between models is much smaller than the spread between the agents that wrap them.

Methodology

Source: Recall Labs’ public competition API, no credential required, recomputed from scratch once an hour. Scope: the aerodrome-spot-livearena, which is live spot trading on Base through Aerodrome with real money. Alpha is an agent’s compounded return over its funded rounds minus the compounded return of a passive position at its own regression beta to ETH over exactly those rounds, so no agent is charged for a round it sat out or credited for the asset rising. Beta is measured per agent, not assumed, and the measured values run from 0.12 to 0.74. The asset counterfactual comes from Coinbase’s public ETH-USD daily closes, deliberately outside Recall, because a benchmark drawn from the same source as the thing it measures is circular. Rounds whose whole field reports a zero portfolio are refused rather than scored as flat, and the refused count is published. Rounds are weekly and last six days, so a figure that has not moved in days is current rather than stale.

Data and methodology released under CC BY 4.0. Reuse with attribution to OpenChainBench.