ML-QuantSubscribe

arXivML & AI Methods

LiveMACE: Process-Aware Evaluation of LLM Agent Capabilities in Evolving Markets

A benchmark evaluates frontier LLMs as live trading agents and finds that realized returns often diverge from capability-specific measurements, revealing outcome-capability gaps through decision traces.

Featured in No. 134 on 9 Oct 2026 · 1 day after release

LiveMACEBench overview
Figure 1 : LiveMACEBench overview. Agents interact with a shared, continuously evolving live-market environment along persistent trajectories. Evaluation combines realized outcomes with mechanism-specific process diagnostics to distinguish task performance from how effectively agents use each mecha…
Released
8 Oct 2026
First featured
No. 134 · 9 Oct 2026
Published in
Not yet, as far as Semantic Scholar knows
Fanfare
3 of 5
Identifier
arXiv:2610.09872
Authors
Jun Zhao et al.

Abstract

From arXiv (CC0).

Evaluating agents by outcomes alone can obscure the capabilities that produce them. This problem is especially pronounced in evolving environments, where outcomes reflect a closed-loop interaction between agent behavior and changing external conditions. We introduce LiveMACEBench, a process-aware benchmark that uses live financial markets as a naturally evolving testbed for persistent LLM agents. Five frontier LLMs operate along continuous trajectories under matched Tool Use, Persistent Memory, Rule Following, and Multi-Agent Collaboration configurations. We evaluate them through both realized outcomes and mechanism-specific diagnostics derived from complete decision traces. Across 30 days of live evaluation, we find a pronounced outcome-capability gap: realized returns often diverge from capability-specific measurements, and similar outcomes can arise from markedly different patterns of mechanism use. Trace-level diagnostics further expose distinct bottlenecks across capabilities, demonstrating that mechanism access, effective mechanism use, and downstream performance are not interchangeable measures of agent capability. LiveMACEBench makes this distinction measurable, turning live markets from a performance leaderboard into a diagnostic environment for agent capability

Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).

    Type to search. Try rough volatility, LLM agents or FinGPT.

    ↑↓ move↵ openesc closeFull search page