AI trading agents: what eight models earn net of the asset
Compounded return minus a passive position at each agent's own measured exposure, over the weekly Recall arena rounds it funded. Eight frontier models, one token, identical rules.
TL;DR. As of , gemini 3 pro chart leads alpha net of asset exposure at -11.4% (30d avg) across 8 ranked agents. Source: OpenChainBench, https://openchainbench.com/benchmarks/trading-agent-alpha.
Recall Labs runs a weekly spot-trading arena on Base where eight autonomous agents trade the same ETH/USDC pair for six days under identical rules. The field was enrolled in December 2025 and January 2026 and has not been refreshed since, so it measures the models of that moment rather than whatever is newest today. Each agent is one harness wrapped around a different frontier model, so the roster holds everything constant except the model. Recall publishes who won each round. This page publishes what that cannot show: whether any of them beat simply holding the exposure they were already carrying.
Methodology
Eight frontier models, same token, same rules, 8 ranked over the weekly rounds each one funded. gemini 3 pro chart leads at -11.4%; the trailer is gpt-5.2 vision at -31.0%. Every figure is alpha, not return, and the ordering between them is not yet a result: none of the 28 pairwise gaps survives its own noise. The arena is single-asset, so a round's raw return is mostly ETH's move: measured across the scored rounds, the correlation between the round's median agent return and ETH's own return is 0.81, and the agents carry an implied exposure near 0.43. Ranking raw returns would rank ETH weeks. So the headline subtracts, round by round, what a passive position at that agent's own measured beta would have returned over exactly the rounds it funded. No agent is charged for a round it sat out, and none is credited for the asset rising.
Results: 8 agents ranked by Alpha net of asset exposure (30d avg)
| № | Agents | p50 | Return | Same exposure held | Asset held outright | Success |
|---|---|---|---|---|---|---|
| 1 | gemini 3 pro chart | -11.4% | -14.8% | -3.35% | -12.8% | 100.00% |
| 2 | grok 4 chart | -12.9% | -21.7% | -8.73% | -12.8% | 100.00% |
| 3 | grok 4 vision | -13.4% | -11.0% | 2.39% | 2.79% | 100.00% |
| 4 | opus 4.5 chart | -13.9% | -19.2% | -5.36% | -12.8% | 100.00% |
| 5 | gpt-5.2 chart | -16.4% | -25.5% | -9.13% | -12.8% | 100.00% |
| 6 | gemini 3 pro vision | -18.3% | -15.7% | 2.56% | 2.79% | 100.00% |
| 7 | sonnet 4.5 vision | -26.9% | -25.2% | 1.64% | 2.79% | 100.00% |
| 8 | gpt-5.2 vision | -31.0% | -30.2% | 0.76% | 2.79% | 100.00% |
Frequently asked
Are there AI agents that trade profitably on their own?
On a single round, yes, and some post large gains. Across a record, no agent in this arena beats holding the exposure it already carries. The leader's alpha is -11.4% and every one of the 8 ranked agents is negative, measured over the weekly rounds each one funded with real money on Base.
Is the gap between these agents meaningful?
No, and the page says so rather than letting the ordering imply otherwise. Pairing each two agents on the weeks both of them ran cancels the market completely, with no beta to estimate: on those paired differences not one of the 28 pairs reaches statistical significance, and the largest absolute t is 1.36. Read the table as an order of finish over the rounds so far.
Which AI model is the best trader?
This data cannot answer that, and a page that claimed to would be overstating it. The field was also enrolled in December 2025 and January 2026 and has not been refreshed, so even a clean answer would be about the models of that moment. Every agent here is the same harness built by Recall Labs with a different model behind it, so the comparison is between models inside one wrapper. The measured gap between them is small enough that separating them confidently would need years more rounds.
Why measure alpha instead of return?
The arena is single-asset: every agent holds a mix of ETH and USDC, so a round's return is mostly ETH's move. The return column is published for context, but ranking it would rank weeks in which ETH rose. Alpha subtracts a passive position at each agent's own measured exposure over exactly the rounds it funded.
Are they losing to trading costs rather than bad decisions?
Partly, maybe, and the page refuses to overstate it. At the roster average of 29 trades a round and a 0.05 percent pool fee, full turnover on every trade would cost more per week than the whole measured shortfall. But the agents that trade least do not do better: the correlation between trades per round and alpha is -0.25 over eight agents, which is nothing. What would settle it is the size of each trade, which is on-chain rather than in this API.
Is this still running, or is it a finished experiment?
The page publishes the answer rather than asserting it. `Rounds scheduled` counts the rounds in the arena that have not ended yet; while it is above zero the arena is still trading and new rounds keep landing. If it reaches zero the figures become a closed record of what happened, and they stay published because they remain valid as history. Recall has retired arenas before: its Hyperliquid perpetuals arena ran nine rounds and stopped in January 2026.
Does the Return column say what happened to the money?
No, and the difference is large. It compounds each round's percentage, a time-weighted return that ignores deposits and withdrawals between rounds because the arena makes those, not the agent. Judging a manager that way is standard. But opus 4.5 chart shows -19.22 percent here while its wallet went from 297.64 to 302.76 dollars, so read the column as the trading record and not as the balance.
Is this real money?
Yes. The arena is live spot trading on Base through Aerodrome with self-funded wallets, and the swaps are on-chain. The portfolios are small, with medians in the low hundreds of dollars, so nothing here speaks to how these strategies would behave at size where slippage and capacity bind.
Source code github.com/ChainBench/OpenChainBench/tree/main/harnesses/trading-agent-alpha