---
title: KiT: A Foundation Model for Financial Time-Series Forecasting using DiffusionTransformers
url: https://www.ml-quant.com/papers/arxiv/2609.34507/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-10-02
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2609.34507
source_url: https://arxiv.org/abs/2609.34507
featured: 2026-10-02
citations: unknown
topic: ML & AI Methods
---


# KiT: A Foundation Model for Financial Time-Series Forecasting using DiffusionTransformers

Diffusion transformer foundation model reformulates candlestick prediction as conditional trajectory generation, achieving 0.057 mean return RankIC across markets and timescales.

- Source: https://arxiv.org/abs/2609.34507
- Identifier: arXiv:2609.34507
- Released: 2026-09-29
- First featured: Quant Letter No. 133 (2026-10-02): https://www.ml-quant.com/issues/2026-10-02/
- Citations (Semantic Scholar): not tracked
- Published in: not yet
- Topic: ML & AI Methods
- Authors: Boyu Zhang, Haorui Li

## Abstract (arXiv, CC0)

Financial candlestick forecasting is fundamental to quantitative investment, yet it remains exceptionally challenging due to extremely low signal-to-noise ratios and vast heterogeneity across markets and instruments. Existing approaches have largely attempted to introduce deep learning to capture hidden temporal features, but most adopt an auto-regressive formulation, which leads to error accumulation during inference. Meanwhile, general-purpose time-series foundation models are not tailored to the unique structure of k-line data and yield unsatisfactory performance on downstream candlestick forecasting tasks. To tackle these problems, we introduce KiT, a K-line Diffusion Transformer foundation model, and reformulate future prediction as conditional path generation via flow matching: given a historical context window, the model generates an ensemble of plausible future OHLCV trajectories. We pre-train KiT at multiple parameter scales on billions of candlestick bars spanning multiple markets and timescales. Across three markets and seven resolutions, KiT attains a mean return RankIC of 0.057 and a mean volatility RankIC of 0.66, leading at every timescale and outperforming both task-specific financial forecasters and general time-series foundation models. Code will be available at: https://github.com/Luciferbobo/KiT.

## Related

- [Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality](https://www.ml-quant.com/papers/arxiv/2405.21060/): The research identifies a link between state-space models and Transformers in deep learning, leading to the creation of a faster language modeling architecture, Mamba-2.
- [Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models](https://www.ml-quant.com/papers/arxiv/2501.01423/): The paper proposes a new model, VA-VAE, that aligns the latent space with pre-trained vision foundation models, enabling faster convergence of Diffusion Transformers in high-dimensional latent spaces and achieving top performance on ImageNet 256x256 generation.
- [Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models](https://www.ml-quant.com/papers/arxiv/2411.04996/): The Mixture-of-Transformers (MoT) is a sparse multi-modal transformer architecture that reduces pretraining costs and allows modality-specific processing with global self-attention.
- [ECG-FM: An Open Electrocardiogram Foundation Model](https://www.ml-quant.com/papers/arxiv/2408.05178/): ECG-FM, a transformer-based model for ECG analysis, shows strong performance in predicting cardiac conditions, having been pretrained on 2.5 million samples.
- [An Advanced Ensemble Deep Learning Framework for Stock Price Prediction Using VAE, Transformer, and LSTM Model](https://www.ml-quant.com/papers/arxiv/2503.22192/): The research introduces a combined deep learning framework for stock price prediction, demonstrating its effectiveness and reliability in predicting stock price movements.
- [Bytes Are All You Need: Transformers Operating Directly On File Bytes](https://www.ml-quant.com/papers/arxiv/2306.00238/): File Byte Transformers: ByteFormer is a deep learning model that enhances image classification accuracy by 5% and can perform audio classification without specific preprocessing, showcasing its versatility.
