ML-QuantSubscribe

arXivDerivatives & Volatility

HAN-Mamba: Hierarchical Selective State Space Networks for Multi-Scale Financial Volatility Forecasting

Replaces transformers with selective state space encoders in a hierarchical architecture for realized volatility forecasting, reducing error to 0.1927 on the Optiver benchmark with 33 percent fewer parameters than the attention variant.

Featured in No. 134 on 9 Oct 2026 · 1 day after release

Context-length scaling
Figure 2: Context-length scaling. The attention encoders saturate and then regress as L_{s} grows, while the selective state space encoders improve monotonically with diminishing returns
Released
8 Oct 2026
First featured
No. 134 · 9 Oct 2026
Published in
Not yet, as far as Semantic Scholar knows
Fanfare
2 of 5
Identifier
arXiv:2610.10323
Authors
Mihai Bogdan Deaconu and Ioan Daniel Pop

Abstract

From arXiv (CC0).

Short-horizon realized volatility forecasting requires the integration of market information that evolves at incompatible temporal resolutions, from second-level order book dynamics to weekly regime drift. Our conference work introduced HAN-T, a hierarchical architecture in which scale-specific Transformer encoders process short, mid, and long-horizon streams and a learned attention fuser weighs their contributions. This article replaces the quadratic attention encoders with selective state space (Mamba) encoders while retaining attention only in the fuser, where the input is a three-token set rather than a long sequence. The resulting hybrid, HAN-Mamba, summarizes each stream through a recurrent state whose input-dependent gating matches two structural properties of volatility: persistent but decaying memory and abrupt regime shifts. On the Optiver Realized Volatility Prediction benchmark under time-aware five-fold cross-validation, HAN-Mamba improves mean RMSPE over HAN-T (0.1942 vs. 0.1965) with 33% fewer parameters. Its linear-time encoders further allow the high-frequency context to be extended from 60 to 240 buckets, reducing error to 0.1927 where the attention variant saturates, and support constant-time streaming updates at inference. Ablations attribute the gains to the encoder swap, confirm that the hierarchical prior transfers across sequence-model families, and show that the permutation-invariant attention fuser remains the correct mechanism for cross-scale integration.

Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).

    Type to search. Try rough volatility, LLM agents or FinGPT.

    ↑↓ move↵ openesc closeFull search page