---
title: LiveMACE: Process-Aware Evaluation of LLM Agent Capabilities in Evolving Markets
url: https://www.ml-quant.com/papers/arxiv/2610.09872/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-10-09
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2610.09872
source_url: https://arxiv.org/abs/2610.09872
featured: 2026-10-09
citations: unknown
topic: ML & AI Methods
---


# LiveMACE: Process-Aware Evaluation of LLM Agent Capabilities in Evolving Markets

A benchmark evaluates frontier LLMs as live trading agents and finds that realized returns often diverge from capability-specific measurements, revealing outcome-capability gaps through decision traces.

- Source: https://arxiv.org/abs/2610.09872
- Identifier: arXiv:2610.09872
- Released: 2026-10-08
- First featured: Quant Letter No. 134 (2026-10-09): https://www.ml-quant.com/issues/2026-10-09/
- Citations (Semantic Scholar): not tracked
- Published in: not yet
- Topic: ML & AI Methods
- Authors: Jun Zhao, Leiming Fu, Yanbo Wen, Yiding Wang, Xuantong Liu, Yang Shu, Yuyang Lu, Xuanran Xing, Jingqi Tong, Hao Xu, Qi Zhang, Xuanjing Huang

## Abstract (arXiv, CC0)

Evaluating agents by outcomes alone can obscure the capabilities that produce them. This problem is especially pronounced in evolving environments, where outcomes reflect a closed-loop interaction between agent behavior and changing external conditions. We introduce LiveMACEBench, a process-aware benchmark that uses live financial markets as a naturally evolving testbed for persistent LLM agents. Five frontier LLMs operate along continuous trajectories under matched Tool Use, Persistent Memory, Rule Following, and Multi-Agent Collaboration configurations. We evaluate them through both realized outcomes and mechanism-specific diagnostics derived from complete decision traces. Across 30 days of live evaluation, we find a pronounced outcome-capability gap: realized returns often diverge from capability-specific measurements, and similar outcomes can arise from markedly different patterns of mechanism use. Trace-level diagnostics further expose distinct bottlenecks across capabilities, demonstrating that mechanism access, effective mechanism use, and downstream performance are not interchangeable measures of agent capability. LiveMACEBench makes this distinction measurable, turning live markets from a performance leaderboard into a diagnostic environment for agent capability

## Related

- [Can Generative AI agents behave like humans? Evidence from laboratory market experiments](https://www.ml-quant.com/papers/arxiv/2505.07457/): Large Language Models (LLMs) have potential in mimicking human behavior in economic markets, but need more research for improved diversity and accuracy.
- [Does Autonomous Quant Research Improve? Leakage, Evidence Scarcity and Forward Tests in an LLM Factor-Discovery Loop](https://www.ml-quant.com/papers/ssrn/7555699/): An autonomous agent evolved 940 factors over 17 days and finds that reusing backtest data inflates edge by a quarter to a third and in-sample improvement predicts worse performance.
- [Can AI Make Money in Crypto? Measuring the Gap from Backtests to Real Markets](https://www.ml-quant.com/papers/arxiv/2609.34510/): A benchmark compares machine learning, reinforcement learning, and large language model trading methods across historical backtests, paper trading, and live markets to measure the gap.
- [LiveOption: Evaluating LLM Agents in Structured Option Trading with Nonlinear Payoffs](https://www.ml-quant.com/papers/arxiv/2609.33470/): An evaluation framework tests language model agents on structured option trading tasks, finding current systems underperform in most real-world scenarios.
- [Propose, Don't Judge: An Anytime-Valid Referee for LLM Agents That Mine Investment Factors](https://www.ml-quant.com/papers/arxiv/2609.27051/): Proposes a statistical referee that judges investment factors proposed by language-model agents using out-of-sample market outcomes, ensuring false-discovery control at any stopping time.
- [AlphaDiverse: Post-Training Local Quantitative Research Agents for Diverse Exploration in Alpha Factor Mining](https://www.ml-quant.com/papers/arxiv/2609.29014/): Proposes a multi-agent system with post-training that automates alpha factor mining locally, using diverse research paths and joint optimization to broaden exploration while maintaining prediction quality.
