ML-Quant
Papers
6,396 papers featured in the letter since May 2023, each with our one-sentence summary and, where Semantic Scholar knows it, what became of it.
- arXiv
- 2,071
- SSRN
- 2,741
- RePEc
- 748
- Machine learning
- 836
By venue
arXiv
Quantitative-finance and ML-for-finance preprints from arXiv.
- Financial Tail Risk Beyond Lipschitz Continuity via Semi-Discrete Optimal Transport
- Propose, Don't Judge: An Anytime-Valid Referee for LLM Agents That Mine Investment Factors
- The Impossible Trinity of Time-Series Validation: A Conservation Law among Training Sufficiency, Test Coverage, and Temporal Causality
2,071 paperssince 24 May 2023
SSRN
Working papers in finance and economics from SSRN.
- Artificial intelligence and financial markets
- Label alchemy: Target engineering for improved stock selection
- Algorithmic Collusion by Reinforcement-Learning Pricing Agents: Simulation Evidence and Implications for Financial Markets and Competition Law
2,741 paperssince 24 May 2023
RePEc
Economics working papers from RePEc's NEP field reports.
- Assessing the Benefits of Optimized Agentic AI Systems for Asset Pricing
- Stablecoins Meet the Mundell–Fleming Trilemma
- Skewness Risk Premia and the Cross-Section of Currency Returns
748 paperssince 24 May 2023
Machine learning
The general machine-learning papers the letter carried in 2023-25.
- Fairness in Survival Analysis: A Novel Conditional Mutual Information Augmentation Approach
- A comparison of translation performance between DeepL and Supertext
- Revisiting Expected Possession Value in Football: Introducing a Benchmark, U-Net Architecture, and Reward and Risk for Passes
836 paperssince 16 Oct 2023
Most cited
Featured papers with the most citations today (Semantic Scholar).
- 5 Jun 20249,205cites
Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Sequence Modeling: Mamba, a neural network architecture that doesn't use attention or MLP blocks, provides faster inference and better performance in language, audio, and genomics than Transformers.
Machine learningML & AI Methods
- 7 Feb 20249,003cites
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Advancing Math Reasoning in Language Models: DeepSeekMath7B is a new language model that uses web data and Group Relative Policy Optimization for advanced mathematical reasoning, scoring high on the MATH benchmark.
Machine learningLLMs & Text
- 16 Oct 20233,920cites
Mistral 7B
Superior Language Model: Mistral 7B v0.1 is a language model with 7 billion parameters that excels in reasoning, mathematics, and code generation, and has a version specifically designed to follow instructions.
Machine learningLLMs & Text
- 20 Jun 20242,228cites
Depth Anything V2
Depth Anything V2 is a new model for monocular depth estimation, using synthetic and large-scale pseudo-labeled real images for faster, more accurate results and setting a new evaluation benchmark.
Machine learningOtherIn Neural Information Processing Systems
- 9 Oct 20242,199cites
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
MMLU-Pro, an improved dataset, expands the Massive Multitask Language Understanding benchmark by adding tougher questions and more choices, serving as a better benchmark to monitor progress in the field.
Machine learningOtherIn Neural Information Processing Systems
- 7 Aug 20242,189cites
Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
The research investigates enhancing Large Language Models' (LLMs) performance using more test-time computation, suggesting a compute-optimal scaling strategy based on prompt difficulty.
Machine learningLLMs & Text
- 5 Jun 20241,953cites
Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
The research identifies a link between state-space models and Transformers in deep learning, leading to the creation of a faster language modeling architecture, Mamba-2.
Machine learningML & AI MethodsIn International Conference on Machine Learning
- 22 May 20241,880cites
Octo: An Open-Source Generalist Robot Policy
Octo is a large transformer-based policy for robotic manipulation, trained on a vast dataset, that can be instructed via language or images and adapted to new domains.
Machine learningML & AI MethodsIn Robotics: Science and Systems Conference
- 24 Jul 20241,800cites
AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration
The study suggests Activation-aware Weight Quantization (AWQ), a hardware-friendly method for quantizing large language models that reduces error and improves performance on various benchmarks.
Machine learningLLMs & TextIn GetMobile: Mobile Computing and Communications
- 12 Dec 20241,775cites
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
The paper presents InternVL 2.5, a sophisticated multimodal large language model that performs well on various benchmarks, exceeding 70% on the MMMU benchmark.
Machine learningLLMs & Text
- 25 Sep 20241,558cites
Qwen2.5-Coder Technical Report
The report unveils the Qwen2.5-Coder series, an improvement from its predecessor, showcasing remarkable code generation abilities and achieving top-tier performance in various code-related tasks.
Machine learningOtherFeatured 2×
- 5 Feb 20251,462cites
s1: Simple test-time scaling
The research presents a method called budget forcing, which uses a small dataset to achieve test-time scaling and improved reasoning performance in language modeling, particularly in competition math questions.
Machine learningLLMs & TextIn Conference on Empirical Methods in Natural Language Processing
- 22 May 20241,459cites
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
MoE Language Model: DeepSeek-V2, a language model with 236B parameters, offers enhanced performance and cost efficiency compared to its predecessor, ranking high among open-source models.
Machine learningLLMs & Text
- 24 Apr 20241,418cites
Mastering Diverse Domains through World Models
Algorithm Mastery: DreamerV3, a universal algorithm, excels in over 150 varied tasks, including diamond collection in Minecraft without human input, expanding the scope of reinforcement learning.
Machine learningML & AI Methods
- 16 Oct 20231,373cites
MemGPT: Towards LLMs as Operating Systems
Extended Context in LLMs: MemGPT is a system that manages different memory levels, providing extended context within large language models' limited context windows, enhancing document analysis and multi-session chat performance.
Machine learningLLMs & TextFeatured 2×
- 16 Oct 20231,236cites
SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models
Hallucination Detection for LLMs: The paper presents SelfCheckGPT, a new approach for fact-checking black-box model responses without an external database, proving its superior ability to detect and rank factual and non-factual sentences.
Machine learningLLMs & TextIn Conference on Empirical Methods in Natural Language ProcessingFeatured 2×
- 10 Jul 20241,173cites
SimPO: Simple Preference Optimization with a Reference-Free Reward
Simple Preference Optimization: SimPO improves reinforcement learning from human feedback by using the average log probability of a sequence as the implicit reward, enhancing training stability and computational efficiency.
Machine learningML & AI MethodsIn Neural Information Processing Systems
- 12 Jun 20241,169cites
Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
The article discusses Visual AutoRegressive modeling (VAR), a new image learning method that outperforms diffusion transformers in terms of speed, image quality, and scalability.
Machine learningML & AI MethodsIn Neural Information Processing Systems
- 8 May 2024981cites
A Simple and Effective Pruning Approach for Large Language Models
Wanda, a new method, efficiently prunes weights in Large Language Models without retraining, offering a more efficient approach to inducing sparsity in pretrained models.
Machine learningLLMs & TextIn International Conference on Learning Representations
- 3 Jan 2024967cites
PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
PIXART-$\alpha$, a Transformer-based text-to-image model, generates high-quality images at a low cost, reducing CO2 emissions and offering a cost-effective solution for the AIGC community.
Machine learningML & AI MethodsIn International Conference on Learning Representations
- 9 Jan 2024902cites
TinyLlama: An Open-Source Small Language Model
Small Open-Source Language Model: The article presents TinyLlama, a compact 1.1B language model that performs remarkably well in various tasks despite its small size, having been pretrained on around 1 trillion tokens.
Machine learningLLMs & Text
- 5 Feb 2025888cites
TÜLU 3: Pushing Frontiers in Open Language Model Post-Training
Open Language Model Post-Training: The Tulu 3 model, a top-tier post-trained language model, is introduced, outperforming other models and providing a detailed guide for its use and adaptation.
Machine learningLLMs & Text
- 12 Jun 2024850cites
Simple and Effective Masked Diffusion Language Models
The performance of diffusion models in language modeling has been enhanced by using an effective training recipe and a simplified objective, setting a new standard among diffusion models.
Machine learningLLMs & TextIn Neural Information Processing Systems
- 12 Dec 2024742cites
Training Large Language Models to Reason in a Continuous Latent Space
The article presents Coconut, a new approach that uses the last hidden state of large language models for reasoning in an unrestricted latent space, proving its effectiveness in enhancing the LLM on multiple reasoning tasks.
Machine learningLLMs & Text
- 8 May 2024727cites
Fourier Neural Operator with Learned Deformations for PDEs on General Geometries
The study introduces geo-FNO, a new framework for solving partial differential equations on any geometry, proving to be faster and more accurate than standard and machine learning-based solvers.
Machine learningML & AI MethodsIn J. Mach. Learn. Res.
Latest featured
- 25 Sep 20263fanfare
LLM-Based Semantic Surprises in FOMC Communication: Asset Prices and Financial-Market Stress
Semantic surprises extracted from Federal Reserve statements predict subsequent financial-stress dynamics and reduce forecast error by up to 23%, particularly when initial stress is high or during recessions.
SSRNLLMs & Text
- 25 Sep 20262fanfare
Banking-System Heterogeneity and Monetary Policy Transmission in the Euro Area: High-Frequency Shocks, Local Projections, and Regime Dependence
A 100-basis-point contractionary monetary shock lowers inflation and sales across 20 euro-area economies, with transmission strength varying by bank asset-risk exposure and assets-to-GDP ratio rather than a simple weak-strong taxonomy.
- 25 Sep 20263fanfare
Firm-Specific Price Delay and Momentum
Momentum profits concentrate among firms with high price delay, a measure of information friction, directly supporting theories that gradual information incorporation drives momentum.
- 25 Sep 20263fanfare
Monetary policy transmission by securitising banks
Banks engaged in securitization contract lending more sharply after monetary tightening because their investor base demands higher returns and cuts risk exposure when rates rise.
- 25 Sep 20264fanfare
Artificial intelligence and financial markets
A survey examines how AI transforms information production, intermediation, and market structure, with implications for efficiency, competition and financial stability.
SSRNML & AI Methods
- 25 Sep 20263fanfare
Forward Guidance and the Dynamics of Bank Credit: The Bank Balance-Sheet Channel of Monetary News
High-frequency analysis reveals contractionary forward guidance immediately cuts bank lending, while expansionary guidance produces weak stimulus, driven by binding capital constraints.
- 25 Sep 20262fanfare
Crossing the Zero Lower Bound: Negative Interest Rates and Corporate Valuation
Comparing firms across the ECB's 2014 negative rate adoption shows treated European firms had higher valuations but reduced leverage, suggesting cash-flow and discount-rate channels dominate tax-shield effects.
- 25 Sep 20263fanfare
Incentives at Play: Fee-Induced Volume on a Regulated Perpetual Futures Venue
Analysis of Kalshi's regulated Bitcoin and Ethereum futures reveals that 39-48% of notional trades are mechanical fixed-size orders that vanish when fees are charged, indicating costless artificial volume rather than legitimate trading.
- 25 Sep 20263fanfare
Speculative Leverage and Factor Momentum
Factor momentum strategies earn 49 basis points per month extra return following quarters of rapid margin-debt growth, a predictability that persists after publication and reflects limits to arbitrage correction.
- 25 Sep 20263fanfare
Welcome to the Factor Zoo: Where Mutual Fund Alpha Hides
Using factor selection, the study finds mean active alpha of plus 9 basis points monthly for mutual funds, reversing the no-alpha conclusion when benchmarks are tailored to each fund.
- 25 Sep 20263fanfare
Hedge Fund Trading and Sovereign Bond Yield Sensitivity
Leveraged hedge fund positions amplify sovereign bond yield sensitivity to monetary shocks by over a quarter through directional rebalancing, with effects scaling to position intensity.
- 25 Sep 20263fanfare
Tail-Risk Forecasting with General Cubic Distributions
A cubic quantile framework forecasts Value-at-Risk and Expected Shortfall more reliably than GARCH benchmarks across eight equity indices without requiring a parametric density.
- 25 Sep 20263fanfare
Sell, Hold Out, or Accept: The Creditor's Trilemma in Distressed Debt Exchanges
Analysis of 284 distressed exchanges from 2009-2022 reveals over 50% of firms face subsequent default, with large illiquid creditors trapped in a prisoner's dilemma explaining high acceptance rates.
- 25 Sep 20264fanfare
Algorithmic Collusion by Reinforcement-Learning Pricing Agents: Simulation Evidence and Implications for Financial Markets and Competition Law
Q-learning pricing agents in simulated duopolies reach supracompetitive outcomes with no communication, achieving collusion indices of 0.778 and 40% profit gains over competitive benchmarks.
- 25 Sep 20263fanfare
Execution-Aware Alpha Mining: Teaching LLM Factor Agents to Account for Trading Costs
The paper builds a closed-loop system where an LLM proposes equity factors penalized for execution costs and shows that accounting for trading costs dramatically improves net performance.
SSRNML & AI Methods
- 25 Sep 20263fanfare
Prices or implied volatilities? Choosing the loss function in machine learning option pricing
The paper compares machine learning option pricing trained on pricing errors versus implied-volatility errors using 8.67 million S&P 500 index-option observations from 1997 through 2025.
- 25 Sep 20262fanfare
Signature-Based Structural Models and Applications in Credit Markets
The study develops a time-varying signature asset model for structural credit that improves calibration across CDS maturities and equity option prices, especially for high-yield firms.
- 25 Sep 20262fanfare
MartingaleONet: Physics-Constrained Operator Learning for Real-Time Option Pricing and Volatility Calibration
A deep operator network maps volatility surfaces to option prices under the Heston model 15,000 times faster than finite-difference methods while reducing dynamic hedging variance by over 59% under transaction costs.
- 25 Sep 20263fanfare
Expectations and the Term Structure of Interest Rates
Decomposing yield sensitivity without assuming rational expectations reveals that expectations rather than risk premia drive short- and medium-term bond yields, with systematic inconsistencies across horizons.
- 25 Sep 20264fanfare
Label alchemy: Target engineering for improved stock selection
Reshaping the prediction target through location, scale and shape transformations raises long-short Sharpe from 0.68 to 1.69, with label choice mattering more than model choice.
SSRNML & AI Methods
- 25 Sep 20262fanfare
State-dependent global banking systemic risk: An integrated framework of network connectedness, tail risk, and global financial conditions
Combining quantile-connectedness, tail-risk measures, and network analysis, the research shows tail connectedness exceeds median levels and lower-tail effects persist longer, with the VIX alone reliably predicting next-week systemic risk.
- 25 Sep 20263fanfare
Hedge Fund Performance and Interest Rate Conditions: Evidence from Regulatory Data
Using SEC filings from 2013-2021, the paper finds hedge fund returns show heterogeneous sensitivity to interest rates, with effects varying by strategy, leverage, and derivative exposure.
- 25 Sep 20263fanfare
Memorisation or Alpha? Detecting Look-Ahead Contamination in Cross-Sectional Equity Signals
Testing whether a large language model ranks stocks by forecasting or memory, the study finds a significant information-coefficient gap of 0.185 inside versus outside its training window, suggesting substantial look-ahead contamination.
SSRNML & AI Methods
- 25 Sep 20262fanfare
The Low Return Channel of Negative Interest Rates in Bank Lending
Japan's 2016 negative-rate policy reduced lending from low-profitability banks holding reserves, consistent with lower expected returns on bank assets rather than deposit-side stress.
- 25 Sep 20262fanfare
Data-Driven Minimax-Regret Portfolio Optimization under Tail-Risk Ambiguity
The research proposes a data-driven portfolio method that blends tail-risk models and projects onto valid mixtures, providing bounds on Expected Shortfall regret without Wasserstein assumptions.
- 25 Sep 20263fanfare
Industry Information and Equity Return Predictability
Using production, employment, and sales data across 426 industries, the research shows that upstream industry signals predict aggregate monthly stock returns with 23.8% out-of-sample R-squared.
- 25 Sep 20263fanfare
Beta Recall, Alpha Recall, and a Contamination Detector that Needs No Labels * Measuring Training-data Leakage in LLM Equity Signals
The study measures recall versus forecasting in an LLM's stock rankings by comparing cross-sectional information coefficients inside and outside the training window.
SSRNML & AI Methods
- 25 Sep 20262fanfare
Fedspeak, LLM-Derived Signals, and High-Frequency Trading
Semantic and tonal shifts across sequential Federal Reserve communications generate significant intraday price movements and abnormal volume, revealing incomplete information absorption at initial announcement.
- 25 Sep 20264fanfare
From D&I to D&I: European Capital Markets' Regime Shift from Diversity and Inclusion to Defence and Infrastructure
European defence stocks repriced sharply starting November 2021, two to three months before Russia's invasion, delivering 26% alpha and reflecting release of ESG-exclusion constraints.
- 25 Sep 20263fanfare
Settlement Risk and Currency Markets
Hungary's 2015 adoption of payment-versus-payment settlement reduced currency excess returns by ten basis points, demonstrating settlement risk is a priced friction limiting arbitrage.
- 25 Sep 20263fanfare
Elastic in cash, inelastic in repo: Hedge funds in the treasury and repo markets
Using German sovereign bond repo data, the research shows hedge funds are price-elastic in cash markets but highly inelastic in repo, inheriting elasticity from their cash-market counterparties.
- 25 Sep 20262fanfare
Collateral policy surprises
Expansionary central bank collateral policy surprises reduce bank default risk and volatility while compressing government bond spreads, transmitting effects distinctly from asset purchases.
- 25 Sep 20263fanfare
Rate Risk and Rate Insurance
Stock returns are dampened by rate insurance: falling rates cushion payoff risk in bad times while rising rates in good times hedge duration exposure.
- 25 Sep 20263fanfare
Credit Card Banking
Analysis of 550 million US credit card accounts shows that despite high charge-off rates, card lenders earn 1.5% alpha and 6.8% return on assets through pricing power and non-interest income.
- 25 Sep 20262fanfare
Common Risk Factors in the Returns on Stocks, Bonds (and Options), Redux
The research identifies common risk factors spanning stocks, corporate bonds, and options linked to economic indicators, revealing significant market segmentation and cross-asset hedging opportunities.
- 25 Sep 20263fanfare
Bank Runs With and Without Bank Failure
A database of 3,984 historical US bank runs shows runs are more likely in weak banks but often occur in strong banks; failures concentrate in fundamentally weak institutions.
- 25 Sep 20264fanfare
Assessing the Benefits of Optimized Agentic AI Systems for Asset Pricing
Optimized AI systems analyzing earnings call transcripts double explained variation in stock returns versus standard benchmarks while improving interpretability through human-readable decision rules.
RePEcML & AI Methods
- 25 Sep 20263fanfare
Prices and Monetary Policy: The Role of Financial Constraints
Swedish data reveals that financially constrained firms adjust prices less to monetary shocks, materially dampening aggregate inflation response to policy changes.
- 25 Sep 20264fanfare
Stablecoins Meet the Mundell–Fleming Trilemma
Wallet-level stablecoin data shows crisis countries experience inflows during banking restrictions; this endogenizes capital mobility and tightens monetary policy constraints.
RePEcCrypto & DeFi
- 25 Sep 20263fanfare
Pricing Risk Globally: Intermediary Constraints, the Dollar, and the Global Financial Cycle
A two-country model shows that uncertainty shocks tighten intermediary constraints, widening credit spreads, appreciating the dollar, and raising currency risk premia globally.