ML-QuantSubscribe

arXivLLMs & Text

Thin Evidence, Thick Priors: How Language Models Substitute Identity for Missing Financial Facts

When financial facts are withdrawn from prompts, identity attributes explain 96% of variation in equity allocation advice from a language model, replacing missing evidence.

Featured in No. 134 on 9 Oct 2026 · 2 days after release

Mean identity swing in equity allocation versus evidence level with 95% confidence interval band.
Figure 1 : The prior substitution curve. Mean identity swing in recommended equity allocation at each level of the evidence dial. The shaded band is the 95% interval from the two-way cluster bootstrap over personas and distinct financial contexts; the grey bars are the intervals implied by the stan…
Released
7 Oct 2026
First featured
No. 134 · 9 Oct 2026
Published in
Not yet, as far as Semantic Scholar knows
Fanfare
4 of 5
Identifier
arXiv:2610.07798
Authors
Saanvi Khetan and Sankar Balasubramanian

Abstract

From arXiv (CC0).

People increasingly ask large language models what to do with their money, yet seldom describe their finances in full. This paper asks what a model does with the gap. Holding finances fixed and changing only who the investor is said to be, we grade the financial evidence in the prompt from eight facts to none and measure how far the recommended equity allocation moves. Across 96,600 prompts to Llama-3.1-8B-Instruct, built from 100 financial profiles, 138 personas and seven disclosure conditions, the average gap between two personas with identical finances rises from 4.78 percentage points at full disclosure to 10.34 points with no financial facts. A two-way cluster bootstrap counting duplicated prompts once places the ratio at 2.16 (95% interval 1.69 to 2.79), and the rise is already 1.69-fold with a single fact left. Identity explains 5% of within-profile variation in advice at full disclosure and 96% with no disclosure. Household size is the only attribute whose influence grows reliably as evidence is withdrawn. Once standard errors are clustered on the persona, the unit to which identity was assigned, most attribute-specific interactions reported in the conference version lose significance, and gender instead appears as a small standing gap that full disclosure does not close. Stating risk appetite alone brings the swing into the range seen with two to seven generic facts. With no facts, the model's one-line rationale cites incomes, debts and savings it was never told, and these invented finances turn adverse more often for larger households. Inside the network, gender is linearly decodable at every layer, and ablating the gender direction at five layers leaves the aggregate identity swing unchanged. Advisory systems built on such models should be audited at the disclosure levels users actually reach, and judged across the whole identity space rather than one attribute at a time.

Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).

    Type to search. Try rough volatility, LLM agents or FinGPT.

    ↑↓ move↵ openesc closeFull search page