---
title: s1: Simple test-time scaling
url: https://www.ml-quant.com/papers/arxiv/2501.19393/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2501.19393
source_url: https://arxiv.org/pdf/2501.19393
featured: 2025-02-05
citations: 1462
topic: LLMs & Text
---


# s1: Simple test-time scaling

The research presents a method called budget forcing, which uses a small dataset to achieve test-time scaling and improved reasoning performance in language modeling, particularly in competition math questions.

- Source: https://arxiv.org/pdf/2501.19393
- Identifier: arXiv:2501.19393
- Released: 2025-01-31
- First featured: Quant Letter No. 84 (2025-02-05): https://www.ml-quant.com/issues/2025-02-05/
- Citations (Semantic Scholar): 1462
- Published in: Conference on Empirical Methods in Natural Language Processing
- Topic: LLMs & Text

## Related

- [Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters](https://www.ml-quant.com/papers/arxiv/2408.03314/): The research investigates enhancing Large Language Models' (LLMs) performance using more test-time computation, suggesting a compute-optimal scaling strategy based on prompt difficulty.
- [Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling](https://www.ml-quant.com/papers/arxiv/2412.05271/): The paper presents InternVL 2.5, a sophisticated multimodal large language model that performs well on various benchmarks, exceeding 70% on the MMMU benchmark.
- [To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning](https://www.ml-quant.com/papers/arxiv/2409.12183/): The chain-of-thought method is most beneficial for tasks involving math or logic in large language models, indicating a need for new methods that utilize intermediate computation across various applications.
- [Chain-of-Thought Reasoning Without Prompting](https://www.ml-quant.com/papers/arxiv/2402.10200/): Altering the decoding process in large language models can enhance their reasoning abilities and performance without needing specific prompts.
- [Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models](https://www.ml-quant.com/papers/arxiv/2501.09686/): The article discusses advancements in Large Language Models (LLMs) reasoning, emphasizing the use of reinforcement learning and thought simulation for complex reasoning, and the potential of scaling during training and testing.
- [GenSim2: Scaling Robot Data Generation with Multi-modal and Reasoning LLMs](https://www.ml-quant.com/papers/arxiv/2410.03645/): GenSim2 is a scalable framework for robotic simulation that uses large language models for task creation and a multi-task language-conditioned policy architecture to learn from demonstrations, improving policy performance and enabling zero-shot transfer.
