ML-QuantSubscribe

arXivML & AI Methods

Same Text, Different Numbers: The Divergence of LLM-Based Measures

Cross-model rank correlations for LLM-extracted sentiment, clarity and risk average only 0.52, and model choice significantly alters coefficient signs and significance in downstream analysis.

Featured in No. 133 on 2 Oct 2026 · 4 days after release

Figure OA.D.1 : Self-reported confidence: distribution and conditional agreement
Figure OA.D.1 : Self-reported confidence: distribution and conditional agreement. Panel (a): pooled distribution of the confidence field across the seven providers and thirteen constructs (a single invalid response with confidence >1 is excluded). Panel (b): mean pairwise Spearman correlation of sc…
Released
28 Sep 2026
First featured
No. 133 · 2 Oct 2026
Published in
Not yet, as far as Semantic Scholar knows
Fanfare
3 of 5
Identifier
arXiv:2609.31013
Authors
Hamid Boustanifar and Sasan Mansouri

Abstract

From arXiv (CC0).

Researchers increasingly use generative large language models (LLMs) to convert corporate text into empirical variables. We examine the extent to which LLM-based textual measures are invariant to model choice using thirteen measures, including sentiment, management clarity, uncertainty, answer specificity, and climate and political risk. Seven LLMs from different providers score earnings call transcripts of S&P 500 companies on these constructs. Cross-model rank correlations average only 0.52, and transcript-level differences common across providers account for only 34% of total score variation. Cross-model disagreement does not predict subsequent analyst or market disagreement, consistent with a substantial model-specific component rather than common ambiguity in the underlying disclosure. Model choice significantly affects downstream inference, with coefficient magnitudes, signs, and statistical significance varying substantially across models. Averaging across providers makes transcript rankings more stable for most constructs, but score levels remain sensitive to the models included in the ensemble. LLM-generated variables should therefore be treated as model-contingent measurements and validated across providers.

Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).

    Type to search. Try rough volatility, LLM agents or FinGPT.

    ↑↓ move↵ openesc closeFull search page