---
title: Featured papers
url: https://www.ml-quant.com/papers/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
---


# Papers featured by ML-Quant

6,396 papers. Per venue: arXiv 2071, SSRN 2741, RePEc 748, Machine learning 836.

## Most cited

- [Mamba: Linear-Time Sequence Modeling with Selective State Spaces](https://www.ml-quant.com/papers/arxiv/2312.00752/): 9205 citations. Sequence Modeling: Mamba, a neural network architecture that doesn't use attention or MLP blocks, provides faster inference and better performance in language, audio, and genomics than Transformers.
- [DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models](https://www.ml-quant.com/papers/arxiv/2402.03300/): 9003 citations. Advancing Math Reasoning in Language Models: DeepSeekMath7B is a new language model that uses web data and Group Relative Policy Optimization for advanced mathematical reasoning, scoring high on the MATH benchmark.
- [Mistral 7B](https://www.ml-quant.com/papers/arxiv/2310.06825/): 3920 citations. Superior Language Model: Mistral 7B v0.1 is a language model with 7 billion parameters that excels in reasoning, mathematics, and code generation, and has a version specifically designed to follow instructions.
- [Depth Anything V2](https://www.ml-quant.com/papers/arxiv/2406.09414/): 2228 citations. Depth Anything V2 is a new model for monocular depth estimation, using synthetic and large-scale pseudo-labeled real images for faster, more accurate results and setting a new evaluation benchmark.
- [MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark](https://www.ml-quant.com/papers/arxiv/2406.01574/): 2199 citations. MMLU-Pro, an improved dataset, expands the Massive Multitask Language Understanding benchmark by adding tougher questions and more choices, serving as a better benchmark to monitor progress in the field.
- [Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters](https://www.ml-quant.com/papers/arxiv/2408.03314/): 2189 citations. The research investigates enhancing Large Language Models' (LLMs) performance using more test-time computation, suggesting a compute-optimal scaling strategy based on prompt difficulty.
- [Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality](https://www.ml-quant.com/papers/arxiv/2405.21060/): 1953 citations. The research identifies a link between state-space models and Transformers in deep learning, leading to the creation of a faster language modeling architecture, Mamba-2.
- [Octo: An Open-Source Generalist Robot Policy](https://www.ml-quant.com/papers/arxiv/2405.12213/): 1880 citations. Octo is a large transformer-based policy for robotic manipulation, trained on a vast dataset, that can be instructed via language or images and adapted to new domains.
- [AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration](https://www.ml-quant.com/papers/arxiv/2306.00978/): 1800 citations. The study suggests Activation-aware Weight Quantization (AWQ), a hardware-friendly method for quantizing large language models that reduces error and improves performance on various benchmarks.
- [Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling](https://www.ml-quant.com/papers/arxiv/2412.05271/): 1775 citations. The paper presents InternVL 2.5, a sophisticated multimodal large language model that performs well on various benchmarks, exceeding 70% on the MMMU benchmark.
- [Qwen2.5-Coder Technical Report](https://www.ml-quant.com/papers/arxiv/2409.12186/): 1558 citations. The report unveils the Qwen2.5-Coder series, an improvement from its predecessor, showcasing remarkable code generation abilities and achieving top-tier performance in various code-related tasks.
- [s1: Simple test-time scaling](https://www.ml-quant.com/papers/arxiv/2501.19393/): 1462 citations. The research presents a method called budget forcing, which uses a small dataset to achieve test-time scaling and improved reasoning performance in language modeling, particularly in competition math questions.
- [DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model](https://www.ml-quant.com/papers/arxiv/2405.04434/): 1459 citations. MoE Language Model: DeepSeek-V2, a language model with 236B parameters, offers enhanced performance and cost efficiency compared to its predecessor, ranking high among open-source models.
- [Mastering Diverse Domains through World Models](https://www.ml-quant.com/papers/arxiv/2301.04104/): 1418 citations. Algorithm Mastery: DreamerV3, a universal algorithm, excels in over 150 varied tasks, including diamond collection in Minecraft without human input, expanding the scope of reinforcement learning.
- [MemGPT: Towards LLMs as Operating Systems](https://www.ml-quant.com/papers/arxiv/2310.08560/): 1373 citations. Extended Context in LLMs: MemGPT is a system that manages different memory levels, providing extended context within large language models' limited context windows, enhancing document analysis and multi-session chat performance.
- [SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models](https://www.ml-quant.com/papers/arxiv/2303.08896/): 1236 citations. Hallucination Detection for LLMs: The paper presents SelfCheckGPT, a new approach for fact-checking black-box model responses without an external database, proving its superior ability to detect and rank factual and non-factual sentences.
- [SimPO: Simple Preference Optimization with a Reference-Free Reward](https://www.ml-quant.com/papers/arxiv/2405.14734/): 1173 citations. Simple Preference Optimization: SimPO improves reinforcement learning from human feedback by using the average log probability of a sequence as the implicit reward, enhancing training stability and computational efficiency.
- [Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction](https://www.ml-quant.com/papers/arxiv/2404.02905/): 1169 citations. The article discusses Visual AutoRegressive modeling (VAR), a new image learning method that outperforms diffusion transformers in terms of speed, image quality, and scalability.
- [A Simple and Effective Pruning Approach for Large Language Models](https://www.ml-quant.com/papers/arxiv/2306.11695/): 981 citations. Wanda, a new method, efficiently prunes weights in Large Language Models without retraining, offering a more efficient approach to inducing sparsity in pretrained models.
- [PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis](https://www.ml-quant.com/papers/arxiv/2310.00426/): 967 citations. PIXART-$\alpha$, a Transformer-based text-to-image model, generates high-quality images at a low cost, reducing CO2 emissions and offering a cost-effective solution for the AIGC community.
- [TinyLlama: An Open-Source Small Language Model](https://www.ml-quant.com/papers/arxiv/2401.02385/): 902 citations. Small Open-Source Language Model: The article presents TinyLlama, a compact 1.1B language model that performs remarkably well in various tasks despite its small size, having been pretrained on around 1 trillion tokens.
- [TÜLU 3: Pushing Frontiers in Open Language Model Post-Training](https://www.ml-quant.com/papers/arxiv/2411.15124/): 888 citations. Open Language Model Post-Training: The Tulu 3 model, a top-tier post-trained language model, is introduced, outperforming other models and providing a detailed guide for its use and adaptation.
- [Simple and Effective Masked Diffusion Language Models](https://www.ml-quant.com/papers/arxiv/2406.07524/): 850 citations. The performance of diffusion models in language modeling has been enhanced by using an effective training recipe and a simplified objective, setting a new standard among diffusion models.
- [Training Large Language Models to Reason in a Continuous Latent Space](https://www.ml-quant.com/papers/arxiv/2412.06769/): 742 citations. The article presents Coconut, a new approach that uses the last hidden state of large language models for reasoning in an unrestricted latent space, proving its effectiveness in enhancing the LLM on multiple reasoning tasks.
- [Fourier Neural Operator with Learned Deformations for PDEs on General Geometries](https://www.ml-quant.com/papers/arxiv/2207.05209/): 727 citations. The study introduces geo-FNO, a new framework for solving partial differential equations on any geometry, proving to be faster and more accurate than standard and machine learning-based solvers.

## Venues

- [arXiv](https://www.ml-quant.com/papers/arxiv/index.md)
- [SSRN](https://www.ml-quant.com/papers/ssrn/index.md)
- [RePEc](https://www.ml-quant.com/papers/repec/index.md)
- [Machine learning](https://www.ml-quant.com/papers/ml/index.md)
