AI trading agents
Can an AI model trade its way past the asset it is trading?
Eight autonomous agents trade the same ETH/USDC pair on Base, with real money, for six days at a time. Each one is a single harness built by Recall Labs wrapped around a different frontier model, so the roster holds everything constant except the model. The arena publishes who won each round. This page publishes what that cannot show: whether any of them beat holding the exposure they were already carrying.
One caveat before the numbers: the field was enrolled 266 days ago and has not been refreshed, so these are the frontier models of that moment rather than whatever is newest today. The arena is allowlisted at eight places and Recall chooses them, so the page dates the claim instead of quietly making it about the present.
Agents measured
8 from 4 labs
Beating their own exposure
0 of 8
Alpha range
-31.00% to -11.42%
Separable pairs
0 of 28
Results by lab
full breakdownGrouped by the lab that makes the model, because the arena runs most models twice: once fed the numbers behind a chart and once fed an image of it. Those two rows are one model under two input conditions, not two competitors. Labs are listed alphabetically; the order carries no claim.
| Lab and agent | Alpha | Return | Same exposure held | Winning rounds | Exposure | Weekly swing | Trades per round | Median capital | Rounds |
|---|---|---|---|---|---|---|---|---|---|
| opus 4.5 chartfed numbers | -13.86% | -19.22% | -5.36% | 33% | 0.46 | 3.55% | 16.4 | $284 | 21 |
| sonnet 4.5 visionfed an image | -26.86% | -25.22% | +1.64% | 30% | 0.29 | 4.53% | 29.6 | $179 | 23 |
| gemini 3 pro chartfed numbers | -11.42% | -14.77% | -3.35% | 33% | 0.30 | 2.06% | 6.4 | $260 | 21 |
| gemini 3 pro visionfed an image | -18.28% | -15.72% | +2.56% | 39% | 0.55 | 4.45% | 37.9 | $190 | 23 |
| gpt-5.2 chartfed numbers | -16.42% | -25.55% | -9.13% | 33% | 0.74 | 4.98% | 2.2 | $199 | 21 |
| gpt-5.2 visionfed an image | -31.00% | -30.24% | +0.76% | 39% | 0.12 | 5.05% | 48.6 | $262 | 23 |
| grok 4 chartfed numbers | -12.94% | -21.67% | -8.73% | 33% | 0.71 | 4.87% | 3.0 | $252 | 21 |
| grok 4 visionfed an image | -13.41% | -11.02% | +2.39% | 35% | 0.48 | 3.90% | 88.7 | $213 | 23 |
Return compounds each round’s percentage, so it is the trading record and not the balance: it ignores money the arena moved in or out between rounds. Over this record one agent publishes a return near -19% while its wallet finished slightly up, because capital was withdrawn along the way.
Agent names are declared by the arena, not verified. Nothing in the data proves which model sits behind a name, so no row here is a claim about a vendor’s product.
What these agents actually do
One pair, six days
Each round is a fresh self-funded wallet trading ETH against USDC through Aerodrome on Base. No shorting, no leverage, no other token. An agent's only decision is how much of the wallet sits in ETH at any moment, which is why its measured exposure is the thing worth subtracting.
One harness, swapped models
Recall Labs writes the prompt, the rebalancing loop and the execution, then swaps the model behind it. That is what makes the roster comparable and also what limits it: this measures one harness interacting with each model, never a model's trading ability in the abstract. The eight places are allowlisted and were filled in December 2025 and January 2026, so the field is fixed and we do not choose it.
Numbers or a picture
Most models are entered twice, once fed the chart as numeric series and once fed it as an image. Only three models currently have both variants, so the contrast is indicative rather than a result: the mean paired difference is 7.3 points with a standard deviation of 7.1.
How much of this is signal
The table above puts eight agents in an order across a wide spread, and that order is not yet a result. Pairing each two agents on the weeks both of them ran cancels the market completely, with no exposure to estimate and no counterfactual to argue about. On those paired differences, 0 of 28 comparisons clears statistical significance. Read the ranking as an order of finish over the rounds run so far.
The headline is in better shape and is not there either. Every ranked agent is below zero, and pooling the roster round by round gives a mean alpha of -1.07% per round at t = -1.72. Two is the conventional bar, so the direction is consistent and the size is not yet separable from noise.
23 of the 31 rounds needed. At the dispersion measured so far, that is how many weekly rounds this effect size requires before the finding clears the bar. The arena runs one round a week, so the count moves on its own and this page publishes it rather than waiting to claim the result.
The arena is still running: 3 rounds scheduled, and the last one ended 6 days ago. Worth stating plainly, because this bench scores rounds only once they end: a retired arena and a running one would otherwise look identical, with the figures simply going quiet.
Pooling is done within the round before testing. The eight agents trade the same week, so their results are correlated: treating each agent-round as an independent observation would count one week eight times and overstate the confidence by roughly the square root of the roster size.
The widest spread on this page is not the returns
Same harness, same prompt, same pair, same six days. The agents still disagree about how much to act by a factor of forty: the quietest trades about twice in a round, the busiest close to ninety. Two halves of the same model sit at opposite ends of that range, so what moves it is the input format rather than the model, and a table of returns cannot show it at all.
It also complicates the headline. At the roster average of roughly 29 trades a round and Aerodrome’s 0.05 percent pool fee, charging the full portfolio to every trade would cost about 1.45 percent a week, which is more than the whole measured shortfall. Friction on that arithmetic is large enough to account for the gap.
The cross-section refuses to go along with it. The correlation between trades per round and alpha is -0.25 across the eight agents, which is nothing: the quietest agent still gives up 16.4 points, and the busiest finishes third. So friction is big enough to matter in aggregate and is not what separates them.
What would settle it is the size of each trade, not just the count, and that is on-chain rather than in the upstream API. The arena publishes each agent’s wallet, so every swap is checkable on Base. Until it is measured the cost figure above is an upper bound, not an explanation, and the page states it as one.
Why one arena and not a pooled leaderboard
Recall runs several arenas, and pooling them is the obvious way to build a bigger table. We tried: eight defensible aggregates of the pooled data, and they disagree. Mean rank percentile against mean raw return share only 3 of the top 10 names, with a median move of 12 places and a maximum of 51 out of 67 agents. Whatever sat on top would be a property of the chosen normalisation rather than of the agents.
The arenas are not comparable on their face either. Round length spans 11x, funded capital spans ten orders of magnitude, field size runs from 6 to 64, and the instrument changes between spot, leveraged perpetuals and simulated money. So this page ranks one arena, names it, and leaves the rest out until each can carry its own page.
One result from outside it is worth keeping in view. In a separate Recall round that put these same models against purpose-built trading agents on Hyperliquid perpetuals, the purpose-built entries took the first three places and every model trailed. The spread between models is much smaller than the spread between the agents that wrap them.