---
title: Same Text, Different Numbers: The Divergence of LLM-Based Measures
url: https://www.ml-quant.com/papers/arxiv/2609.31013/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-10-02
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2609.31013
source_url: https://arxiv.org/abs/2609.31013
featured: 2026-10-02
citations: unknown
topic: ML & AI Methods
---


# Same Text, Different Numbers: The Divergence of LLM-Based Measures

Cross-model rank correlations for LLM-extracted sentiment, clarity and risk average only 0.52, and model choice significantly alters coefficient signs and significance in downstream analysis.

- Source: https://arxiv.org/abs/2609.31013
- Identifier: arXiv:2609.31013
- Released: 2026-09-28
- First featured: Quant Letter No. 133 (2026-10-02): https://www.ml-quant.com/issues/2026-10-02/
- Citations (Semantic Scholar): not tracked
- Published in: not yet
- Topic: ML & AI Methods
- Authors: Hamid Boustanifar, Sasan Mansouri

## Abstract (arXiv, CC0)

Researchers increasingly use generative large language models (LLMs) to convert corporate text into empirical variables. We examine the extent to which LLM-based textual measures are invariant to model choice using thirteen measures, including sentiment, management clarity, uncertainty, answer specificity, and climate and political risk. Seven LLMs from different providers score earnings call transcripts of S&P 500 companies on these constructs. Cross-model rank correlations average only 0.52, and transcript-level differences common across providers account for only 34% of total score variation. Cross-model disagreement does not predict subsequent analyst or market disagreement, consistent with a substantial model-specific component rather than common ambiguity in the underlying disclosure. Model choice significantly affects downstream inference, with coefficient magnitudes, signs, and statistical significance varying substantially across models. Averaging across providers makes transcript rankings more stable for most constructs, but score levels remain sensitive to the models included in the ensemble. LLM-generated variables should therefore be treated as model-contingent measurements and validated across providers.

## Related

- [The Impact of the ESG Factor on Industrial Performance An Analysis using Machine Learning Techniques](https://www.ml-quant.com/papers/ssrn/4910370/): The study applies machine learning to explore the link between ESG performance and corporate earnings, using data from over 850 European and US firms from 2007-2021.
- [A Yield-Based Asset Ratio to Boost Minimum Investment Returns](https://www.ml-quant.com/papers/ssrn/4746302/): The article proposes a strategy to improve weak medium-term returns in retirement portfolios by adjusting the stock percentage based on the earnings yield of stock and the current yield of bonds, with caution needed when stock prices exceed sustainable levels.
- [AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments](https://www.ml-quant.com/papers/arxiv/2405.07960/): AI Evaluation in Clinical Environments: The paper introduces AgentClinic, a benchmark for assessing large language models in simulated clinical environments, highlighting the significant impact of biases on diagnostic accuracy and patient interactions.
- [Learning Performance-Improving Code Edits](https://www.ml-quant.com/papers/arxiv/2302.07867/): The research presents a framework for optimizing programs using large language models, achieving a mean speedup of 6.86, outperforming average individual programmers.
- [Generative agent-based modeling with actions grounded in physical, social, or digital space using Concordia](https://www.ml-quant.com/papers/arxiv/2312.03664/): Concordia is a library designed to help build and operate Generative Agent-Based Models (GABMs), using Large Language Models (LLMs) to simulate physical or digital environments.
- [Jamba-1.5: Hybrid Transformer-Mamba Models at Scale](https://www.ml-quant.com/papers/arxiv/2408.12570/): Transformer-Mamba Models: Jamba-1.5 is a new large language model with enhanced conversational and instruction-following capabilities, featuring a unique quantization technique for cost-effective inference.
