← ClaudeAtlas

llm-finance-agentslisted

What the published evidence says about LLM trading agents, and the real status of the frameworks. TRIGGER - TradingAgents, FinGPT, FinRobot, FinMem, FinCON, FinAgent, AlphaAgent, RD- Agent, AI4Finance; building or evaluating an LLM-driven trading system, a multi-agent trader, or a news-sentiment-to-signal pipeline; "does AI trading work"; FinBERT and financial sentiment models; reproducing a Sharpe from an LLM-trading paper; whether a backtest window overlaps a model's training cutoff. No credible evidence exists that any of it produces alpha net of costs. SKIP for reinforcement learning and deep learning specifically (rl-and-ml-trading) and for MCP servers (finance-mcp-servers).
howard-lynn-ye/fin-skills · ★ 1 · AI & Automation · score 77
Install: claude install-skill howard-lynn-ye/fin-skills
# LLM trading agents — the evidence ## 1. The executive answer **No credible evidence exists that an LLM trading agent produces risk-adjusted alpha net of costs.** The literature splits cleanly: | Claim | Status | |---|---| | LLMs extract *sentiment* from news better than lexicon methods | ✅ well supported | | LLM sentiment scores correlate with subsequent returns in-sample | ✅ supported, **but heavily contaminated by lookahead** | | That correlation survives realistic transaction costs | ❌ **fails at ~20 bps round-trip** — by Lopez-Lira & Tang's own numbers | | Multi-agent debate beats one well-prompted agent | ❌ **not established**; the general MAD literature finds the opposite | | Reported Sharpe ratios of 5–8 are real | ❌ artifacts of 3-month windows, 3 tickers, zero costs, a bull regime | | LLM agents beat buy-and-hold in a contamination-free setting | ❌ margin inside noise; **in the one real-money competition, 4 of 6 models lost money** | **The one well-designed study kills its own strategy.** Lopez-Lira & Tang (arXiv 2304.07619) is genuinely post-cutoff and reports a gross Sharpe of 2.97. Its cost curve: **~700% at 0 bps → >300% at 5 bps → >100% at 10 bps → unprofitable at 20 bps round-trip.** That curve, not a point estimate, is the format this field should have adopted. **Alpha Arena S1** (Oct 18 – Nov 3 2025, **real capital**): 4 of 6 models lost money; GPT-5 finished **−59%**; fees alone ate **$1,654 of Qwen3's $10k**. 🚨 The widely circulated **"+79%" is a mi