arXivML & AI Methods
LiveMACE: Process-Aware Evaluation of LLM Agent Capabilities in Evolving Markets
A benchmark evaluates frontier LLMs as live trading agents and finds that realized returns often diverge from capability-specific measurements, revealing outcome-capability gaps through decision traces.
Featured in No. 134 on 9 Oct 2026 · 1 day after release

- Released
- 8 Oct 2026
- First featured
- No. 134 · 9 Oct 2026
- Published in
- Not yet, as far as Semantic Scholar knows
- Fanfare
- 3 of 5
- Identifier
- arXiv:2610.09872
- Authors
- Jun Zhao et al.
Abstract
From arXiv (CC0).
Evaluating agents by outcomes alone can obscure the capabilities that produce them. This problem is especially pronounced in evolving environments, where outcomes reflect a closed-loop interaction between agent behavior and changing external conditions. We introduce LiveMACEBench, a process-aware benchmark that uses live financial markets as a naturally evolving testbed for persistent LLM agents. Five frontier LLMs operate along continuous trajectories under matched Tool Use, Persistent Memory, Rule Following, and Multi-Agent Collaboration configurations. We evaluate them through both realized outcomes and mechanism-specific diagnostics derived from complete decision traces. Across 30 days of live evaluation, we find a pronounced outcome-capability gap: realized returns often diverge from capability-specific measurements, and similar outcomes can arise from markedly different patterns of mechanism use. Trace-level diagnostics further expose distinct bottlenecks across capabilities, demonstrating that mechanism access, effective mechanism use, and downstream performance are not interchangeable measures of agent capability. LiveMACEBench makes this distinction measurable, turning live markets from a performance leaderboard into a diagnostic environment for agent capability
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).