---
title: Resolving the Missing Financial Data Crisis: A Generative AI Pipeline for SEC 10-K Extraction
url: https://www.ml-quant.com/papers/arxiv/2609.35864/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-10-02
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2609.35864
source_url: https://arxiv.org/abs/2609.35864
featured: 2026-10-02
citations: unknown
topic: LLMs & Text
---


# Resolving the Missing Financial Data Crisis: A Generative AI Pipeline for SEC 10-K Extraction

Large language models outperform traditional methods at extracting missing financial data from SEC filings, with accuracy improving as model size matches document complexity.

- Source: https://arxiv.org/abs/2609.35864
- Identifier: arXiv:2609.35864
- Released: 2026-09-30
- First featured: Quant Letter No. 133 (2026-10-02): https://www.ml-quant.com/issues/2026-10-02/
- Citations (Semantic Scholar): not tracked
- Published in: not yet
- Topic: LLMs & Text
- Authors: Prisha Nair, Roee Shraga

## Abstract (arXiv, CC0)

SEC 10-K filings contain substantial financial information that is not consistently captured in structured datasets, creating a missing-data problem affecting over 70% of firms and half of total market capitalization. This can disproportionately bias quantitative analysis against smaller firms, which may be excluded due to limited available data. Traditional financial extraction methods such as Regular Expressions (Regex) and BERT, have been widely used. However, they are highly brittle when parsing complex SEC 10-K filings, which leads to data that is existent in the files being lost since these methods do not consider that a data attribute could be located in a different section or a footnote. This study evaluates several Large Language Models (LLMs), including Llama-3 8B, Qwen-2.5 14B, and Llama-3.3 70B, to figure out individual model strengths and weaknesses when extracting specific attributes from SEC 10-K text. The extraction quality was evaluated across four financial variables of varying structural complexity: Cash and Cash Equivalents (tabular), Short-Term Debt (hybrid), Credit Facilities (narrative), and Research and Development (hybrid). Results show that while smaller models like Llama-3 8B experience performance degradation under complex negative prompting, aligning parameter scale with document complexity yields high zero-shot accuracy. Qwen-2.5 14B excels as a tabular specialist with an 83.33% F1 score on Cash, whereas Llama-3.3 70B effectively navigates dense narrative footnotes, achieving a 76.92% F1 score on R&D. This scalable framework addresses critical information gaps in quantitative finance datasets and eliminates missing-data bias through a more thorough analysis of the SEC 10-K files.

## Related

- [ChatGPT, Generative AI, and Investment Advisory](https://www.ml-quant.com/papers/ssrn/4519182/): The article shows that AI like ChatGPT can generate portfolio recommendations based on policy announcements, potentially outperforming markets unlike traditional textual analysis.
- [LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models](https://www.ml-quant.com/papers/arxiv/2305.13655/): Enhancing Text-to-Image Models: The study suggests a two-stage process using a pretrained language model to improve image generation accuracy in diffusion models, enabling multi-round scene specification in various languages.
- [Planning In Natural Language Improves LLM Search For Code Generation](https://www.ml-quant.com/papers/arxiv/2409.03733/): PLANSEARCH is a new search algorithm that creates diverse solutions for natural language problems, outperforming traditional methods in various benchmarks.
- [Predicting Liquidity-Aware Bond Yields using Causal GANs and Deep Reinforcement Learning with LLM Evaluation](https://www.ml-quant.com/papers/arxiv/2502.17011/): The paper introduces a new method for predicting bond yields using Causal Generative Adversarial Networks and reinforcement learning, which improves forecasting performance by generating synthetic bond yield data.
- [Generative AI Impact on Labor Market: Analyzing ChatGPT's Demand in Job Advertisements](https://www.ml-quant.com/papers/arxiv/2412.07042/): A study reveals the growing demand for ChatGPT-related skills in the U.S. labor market, identifying five key skill sets and emphasizing the widespread use of Generative AI in various sectors.
- [AmbigNLG: Addressing Task Ambiguity in Instruction for NLG](https://www.ml-quant.com/papers/arxiv/2402.17717/): Task Ambiguity in NLG: AmbigNLG, a new task and dataset, tackles task ambiguity in instructions for Natural Language Generation, improving the alignment of generated text with user expectations and boosting the performance of Large Language Models.
