ML-QuantSubscribe

arXivPortfolio & Allocation

MemTrial: Learning When to Trust Memory in LLM Portfolio Agents

MemTrial uses factorial design to isolate what each memory contributes to portfolio decisions by crediting experiences with their marginal effect rather than shared market moves.

Featured in No. 134 on 9 Oct 2026 · on release day

Comparison of existing agent approach versus MemTrial's memory trust mechanism with factorial design
Figure 1. (a) Existing agents credit experiences with market-driven outcomes and always use these credits. (b) MemTrial measures what each experience changes on the same date and acts on its values only after they have predicted unseen dates; otherwise it stays near 1/ N . Two rows of icons. Row (a…
Released
9 Oct 2026
First featured
No. 134 · 9 Oct 2026
Published in
Not yet, as far as Semantic Scholar knows
Fanfare
3 of 5
Identifier
arXiv:2610.11732
Authors
Guanghao Wu et al.

Abstract

From arXiv (CC0).

Large language model (LLM) agents for portfolio management learn from experience: they credit each experience in their memory with the outcome of the decisions that used it. In financial markets, however, this outcome mostly reflects the market move shared by all decisions on that date, so the credit tracks the market rather than the experience, and these agents often do worse than simply holding the equal-weight (1/$N$) portfolio. We ask how an agent can credit an experience with what it changes, and answer it by putting memory on trial: drafts of the same decision with and without an experience face the same market, so the outcome they share cancels in their difference. Our agent, MemTrial, drafts each decision with eight combinations of its retrieved experiences, chosen by a fractional factorial design, and credits each experience with its Banzhaf value, the average of these differences. As each date occurs once and each draft is a noisy LLM sample, these credits are noisy and may not hold on new dates. MemTrial therefore pools them across dates and similar experiences with a hierarchical Bayesian model, acts on them only after they have predicted unseen dates, and otherwise stays anchored at a conservative reference such as 1/$N$. On four benchmarks, MemTrial not only benefits from experiences that matter (the best of 15 methods on a semi-synthetic benchmark with known experience quality) but also limits its losses when its values do not hold (at most 2.2\% below 1/$N$ on PortBench and InvestorBench, against 15--38\% for the best experience-learning agent). Averaged over five settings, it improves the utility of the best experience-learning agent by 21.2\%, and with eight LLMs it beats every LLM-based baseline on InvestorBench.

Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).

    Type to search. Try rough volatility, LLM agents or FinGPT.

    ↑↓ move↵ openesc closeFull search page