{"slug":"trading-agent-alpha","title":"AI trading agents: what eight models earn net of the asset","subtitle":"Compounded return minus a passive position at each agent's own measured exposure, over the weekly Recall arena rounds it funded. Eight frontier models, one token, identical rules.","category":"Trading","metric":"Alpha net of asset exposure","unit":"pct","status":"live","higherIsBetter":true,"filters":null,"value":-11.4215,"leader":{"name":"gemini 3 pro chart","slug":"gemini-3-pro-chart","value":-11.4215},"leaders":[{"name":"gemini 3 pro chart","slug":"gemini-3-pro-chart","value":-11.4215}],"rankings":[{"name":"gemini 3 pro chart","slug":"gemini-3-pro-chart","ms":{"p50":-11.4215,"p90":-11.4215,"p99":-11.4215,"mean":-11.4215},"successRate":100,"sampleSize":21},{"name":"grok 4 chart","slug":"grok-4-chart","ms":{"p50":-12.9418,"p90":-12.9418,"p99":-12.9418,"mean":-12.9418},"successRate":100,"sampleSize":21},{"name":"grok 4 vision","slug":"grok-4-vision","ms":{"p50":-13.4104,"p90":-13.4104,"p99":-13.4104,"mean":-13.4104},"successRate":100,"sampleSize":23},{"name":"opus 4.5 chart","slug":"opus-45-chart","ms":{"p50":-13.8649,"p90":-13.8649,"p99":-13.8649,"mean":-13.8649},"successRate":100,"sampleSize":21},{"name":"gpt-5.2 chart","slug":"gpt-52-chart","ms":{"p50":-16.4192,"p90":-16.4192,"p99":-16.4192,"mean":-16.4192},"successRate":100,"sampleSize":21},{"name":"gemini 3 pro vision","slug":"gemini-3-pro-vision","ms":{"p50":-18.2821,"p90":-18.2821,"p99":-18.2821,"mean":-18.2821},"successRate":100,"sampleSize":23},{"name":"sonnet 4.5 vision","slug":"sonnet-45-vision","ms":{"p50":-26.858,"p90":-26.858,"p99":-26.858,"mean":-26.858},"successRate":100,"sampleSize":23},{"name":"gpt-5.2 vision","slug":"gpt-52-vision","ms":{"p50":-31.0017,"p90":-31.0017,"p99":-31.0017,"mean":-31.0017},"successRate":100,"sampleSize":23}],"sparkline":[-11.4215,-11.4215,-11.4215],"sampleSize":176,"asOf":"2026-10-07T17:57:35.004Z","freshness":{"asOf":"2026-10-07T17:57:35.004Z","ageHours":0.8,"stale":false},"measured":8,"ranked":8,"headline":"gemini 3 pro chart leads alpha net of asset exposure at -11.4% (30d avg) across 8 ranked agents.","quote":"gemini 3 pro chart leads alpha net of asset exposure at -11.4% (30d avg) across 8 ranked agents. Source: OpenChainBench (https://openchainbench.com/benchmarks/trading-agent-alpha).","cite":{"plain":"OpenChainBench. \"AI trading agents: what eight models earn net of the asset\". Retrieved 2026-10-07. https://openchainbench.com/benchmarks/trading-agent-alpha","bibtex":"@misc{ocb_trading_agent_alpha,\n  author = {OpenChainBench},\n  title  = {AI trading agents: what eight models earn net of the asset},\n  year   = {2026},\n  url    = {https://openchainbench.com/benchmarks/trading-agent-alpha},\n  note   = {Retrieved 2026-10-07}\n}","apa":"OpenChainBench. (2026). AI trading agents: what eight models earn net of the asset. Retrieved October 7, 2026, from https://openchainbench.com/benchmarks/trading-agent-alpha","ris":"TY  - GEN\r\nAU  - OpenChainBench\r\nTI  - AI trading agents: what eight models earn net of the asset\r\nPY  - 2026\r\nUR  - https://openchainbench.com/benchmarks/trading-agent-alpha\r\nY2  - 2026-10-07\r\nER  - \r\n"},"pageUrl":"https://openchainbench.com/benchmarks/trading-agent-alpha","ogImage":"https://openchainbench.com/api/og/trading-agent-alpha","source":"https://github.com/ChainBench/OpenChainBench/tree/main/harnesses/trading-agent-alpha","methodology":["Source: Recall Labs' public competition API, no credential required. One pass an hour recomputes every figure from scratch, so the harness holds no state and a missed pass loses nothing. Results back to May 2025 are still served upstream, verified against the oldest round.","Scope: the `aerodrome-spot-live` arena only, which is live spot trading on Base through Aerodrome with real money. One arena on purpose. Across Recall's other arenas the round length spans 11x, the funded capital spans ten orders of magnitude, the field runs from 6 to 64, and the instrument changes between spot, leveraged perpetuals and simulated money.","Why one arena and not a pooled score: we built eight defensible aggregates of the pooled data and they disagree. Mean rank percentile against mean raw return share only 3 of the top 10 names, with a median move of 12 places and a maximum of 51 out of 67 agents. Whatever sat on top would be a property of the chosen normalisation, not of the agents.","Headline: compounded return over an agent's funded rounds, minus the compounded return of a passive position at that agent's own regression beta to the asset over those same rounds. Higher is better and zero means the model added nothing beyond the exposure it held.","Beta is measured per agent, not assumed, by regressing its round returns on the asset's return over the same windows. The measured betas run from 0.12 to 0.74, so a single blanket exposure assumption would flatter some agents and penalise others.","A round with a zero profit figure is almost never a flat trade. Of the zero rows in this arena, 88 of 90 carry a portfolio value of zero, meaning the agent did not fund that round. Only funded rounds count. Treating them as zero-percent results drops the measured hit rate by eleven points and inflates each sample from about 22 rounds to 34.","Eleven ended rounds publish ranks from 1 to 8 and no data at all, every portfolio value zero. The harness refuses to score a round without a real field, and publishes the count it refused as `trading_agent_rounds_skipped`.","Roster age, published rather than implied: these eight agents were enrolled upstream on 2025-12-24 and 2026-01-14 and the field has not changed since. `trading_agent_roster_age_days` carries the gap so the page never reads as a statement about the current frontier. We cannot widen it either: the arena is allowlisted at eight places and Recall picks them.","The roster changed once. Only agents with at least ten funded rounds are ranked, which selects the eight-model roster that 32 rounds share. A handful of other entries appear two or three times and are not averaged into the same table.","The asset counterfactual comes from Coinbase's public ETH-USD daily closes, deliberately outside Recall, because a benchmark drawn from the same source as the thing it measures is circular. Where a round boundary has no candle the harness walks back up to five days, and a round with no price is skipped rather than scored against nothing.","What this does not measure: Recall builds the agents. Their own description reads `Autonomous live spot trading agent built by Recall Labs, powered by` the model in question, so the prompt, the rebalancing and the execution are Recall's and are held constant. This is one harness interacting with each model, never a model's trading ability in the abstract.","Agent names are declared, not verified. Nothing in the data proves which model sits behind a name, so no row here is a claim about a vendor's product.","Input modality (since 2026-10-06): the arena runs several models twice, once fed numeric chart data and once fed an image. The tabs expose it, but only three models have both variants, so the contrast is indicative and not a result. The mean paired difference is 7.3 points with a standard deviation of 7.1 across those three.","Significance, published rather than implied: the unit of observation is the ROUND, not the agent-round. The eight agents trade the same week, so their results are correlated and pooling every agent-round would treat one week as eight and overstate the confidence by roughly the square root of the roster. The harness averages alpha within each round first, then tests that series against zero, and publishes the t, the rounds held and the rounds the effect size needs.","Separability: each pair of agents is compared on the rounds BOTH of them funded. The paired difference cancels the asset entirely, so this test needs no beta and no counterfactual and is the cleanest statement available about whether two agents differ. The count of pairs that clear |t|>2 is published; it is currently zero of 28.","Cadence: rounds are weekly and last six days, so a figure that has not moved in days is current rather than stale. That holds only while the arena runs, so the harness also publishes the forward schedule as `trading_agent_scheduled_rounds`: rounds that have not ended yet. Zero means the experiment is over and this table is a closed record.","Which return this is: the Return column compounds each round's percentage, a time-weighted return. It ignores money moved in or out between rounds, because the arena decides those and the agent does not, which is the standard way to judge a manager whose cash flows it does not control.","That return is NOT the change in the wallet, and the gap can be large. Over this record opus 4.5 chart publishes -19.22 percent while its wallet went from 297.64 to 302.76 dollars, a gain of 1.7 percent, because capital was withdrawn along the way. Read the column as the trading record, not as the balance.","Capital is small and did not hold still. Median portfolio value is a few hundred dollars per agent, but the largest-to-smallest ratio across an agent's own rounds runs from 1.3 to about 190, and half the roster is above 5. Published as `trading_agent_capital_ratio`, because gas is a fixed cost per swap: the same strategy drags far harder on a three-dollar wallet than on a three-hundred-dollar one, so for those agents the late rounds are not comparable with the early ones.","Trade counts come from each agent's own upstream record, and are lifetime rather than per arena. They are not diluted: checked against each agent's round list, 35 to 37 of every agent's 32 to 34 completed rounds are this arena. A missing counter is published as missing rather than as zero, because zero would read as an agent that never trades.","What the fee arithmetic assumes, and why it is not yet a finding: it charges the full portfolio to every trade. The real cost depends on the NOTIONAL of each trade, which this API does not serve. That figure is on-chain, and the upstream publishes each agent's wallet address, so every swap is independently checkable on Base. Until it is measured, treat the cost figure as an upper bound rather than an explanation.","Liveness is published as a count rather than inferred from data age, because a weekly arena is legitimately six days stale most of the time while the site stops indexing a page whose data passes seven days. Tying the two together would deindex this page every week.","What happens when the arena stops: Recall has retired arenas before, and this one will end eventually. The published figures stay valid as history, so they are not withdrawn; what changes is the tense. The page states the number of scheduled rounds and the age of the last one so a reader can tell a running experiment from a finished one without taking our word for it."],"license":"CC-BY-4.0"}