ML-QuantSubscribe

arXivML & AI Methods

Propose, Don't Judge: An Anytime-Valid Referee for LLM Agents That Mine Investment Factors

Proposes a statistical referee that judges investment factors proposed by language-model agents using out-of-sample market outcomes, ensuring false-discovery control at any stopping time.

Featured in No. 132 on 25 Sep 2026 · 1 day after release · 0 citations today

Persistence and monetisation by family
Figure 4: Persistence and monetisation by family. Left: the decay curve IC(k) of a ranking known at t-1 against the return on day t+k-1 , with the viability threshold \delta . Right: what a single-factor quintile sleeve nets per year at 15 bp per side when re-formed every h trading days. Reversal n…
Released
24 Sep 2026
First featured
No. 132 · 25 Sep 2026
Citations (Semantic Scholar)
0
Influential citations
0
Published in
Not yet, as far as Semantic Scholar knows
Fanfare
4 of 5
Identifier
arXiv:2609.27051
Authors
Bo Qu et al.

Abstract

From arXiv (CC0).

Language-model agents now run the whole of quantitative factor research: they propose investment factors, backtest them, select the survivors and retire them. We ask which of those jobs an agent should keep. Our answer is governed self-evolution: the agent may propose, and a frozen statistical referee that the agent cannot touch must judge. The referee scores each candidate only on market outcomes revealed after submission, by betting, so its false-discovery guarantee holds at every stopping time for any proposal policy. We cross three proposers (a script, a bandit and a language model) with this referee and with three deliberately leaky ones, in a synthetic world with planted truth, a probe-authoring environment and a ten-year walk-forward on the CSI 500. Who judges sets the number of false admissions: the frozen referee admits 5-11 times fewer sub-threshold factors than the leaky referees under a scripted proposer, and no proposer closes that gap. Who proposes sets the yield: the language model beats the script, matches the bandit, and adds the one capability a bandit lacks, writing its own diagnostic probes. The certificate's price is time: an admitted true factor waits about 500 trading days, and the certified portfolio's Sharpe ratio therefore trails an ungated one. Judging belongs to the procedure; proposing and instrument-making belong to the agent.

Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).

    Type to search. Try rough volatility, LLM agents or FinGPT.

    ↑↓ move↵ openesc closeFull search page