---
title: EverMine: Dissecting the Self-Evolution of Research Capabilities in Long-Horizon Alpha Research
url: https://www.ml-quant.com/papers/arxiv/2609.33524/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-10-02
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2609.33524
source_url: https://arxiv.org/abs/2609.33524
featured: 2026-10-02
citations: unknown
topic: ML & AI Methods
---


# EverMine: Dissecting the Self-Evolution of Research Capabilities in Long-Horizon Alpha Research

A framework tests whether self-evolving research capabilities improve alpha discovery over time, finding no consistent gains from accumulated experience and tools.

- Source: https://arxiv.org/abs/2609.33524
- Identifier: arXiv:2609.33524
- Released: 2026-09-29
- First featured: Quant Letter No. 133 (2026-10-02): https://www.ml-quant.com/issues/2026-10-02/
- Citations (Semantic Scholar): not tracked
- Published in: not yet
- Topic: ML & AI Methods
- Authors: Siyuan Li, Jiangfeng Zhang, Rui Yao, Weihua Qiu, Mingyang Xu, Zixuan Yuan

## Abstract (arXiv, CC0)

Self-evolving agents aim to turn research feedback into reusable skills, tools, and research rules. Whether these accumulated capabilities continue to improve later research requires controlled evaluation. Long-horizon alpha discovery provides a state-dependent setting: once a new factor enters the portfolio, the predictive information already covered changes, so the value of the same candidate or experience may change over time. We introduce EverMine, an empirical framework for studying self-evolving research capabilities in long-horizon alpha discovery. EverMine decomposes the research state into history (Hist), the current factor portfolio (Frontier), and reusable capabilities (Cap). Under matched resource limits, we compare complete runs with fixed or evolving Cap, and replace Cap while holding Hist and Frontier fixed to estimate the conditional value of accumulated capabilities. We also combine full trajectories with historical-state replay to examine how experience-based decisions affect candidate selection and portfolio outcomes. Across 18 long-horizon trajectories, end-to-end comparisons show no consistent gain from Cap evolution. Across 48 continuation branches from shared Hist and Frontier states, accumulated Cap also does not consistently outperform the initial Cap. Parameter tuning of existing factor structures can still improve the portfolio. In an exploratory replay of two screening batches from one Evolving trajectory, some screened-out candidates have positive marginal value at the original state, yet submitting all screened-out candidates sequentially slightly lowers final portfolio IC in both batches. These results show that candidate value depends on the evolving portfolio and submission order, and motivate evaluating self-evolving research capabilities through end-to-end outcomes, conditional capability value, and the consequences of experience-based decisions.

## Related

- [Mamba: Linear-Time Sequence Modeling with Selective State Spaces](https://www.ml-quant.com/papers/arxiv/2312.00752/): Sequence Modeling: Mamba, a neural network architecture that doesn't use attention or MLP blocks, provides faster inference and better performance in language, audio, and genomics than Transformers.
- [Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality](https://www.ml-quant.com/papers/arxiv/2405.21060/): The research identifies a link between state-space models and Transformers in deep learning, leading to the creation of a faster language modeling architecture, Mamba-2.
- [Octo: An Open-Source Generalist Robot Policy](https://www.ml-quant.com/papers/arxiv/2405.12213/): Octo is a large transformer-based policy for robotic manipulation, trained on a vast dataset, that can be instructed via language or images and adapted to new domains.
- [Mastering Diverse Domains through World Models](https://www.ml-quant.com/papers/arxiv/2301.04104/): Algorithm Mastery: DreamerV3, a universal algorithm, excels in over 150 varied tasks, including diamond collection in Minecraft without human input, expanding the scope of reinforcement learning.
- [SimPO: Simple Preference Optimization with a Reference-Free Reward](https://www.ml-quant.com/papers/arxiv/2405.14734/): Simple Preference Optimization: SimPO improves reinforcement learning from human feedback by using the average log probability of a sequence as the implicit reward, enhancing training stability and computational efficiency.
- [Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction](https://www.ml-quant.com/papers/arxiv/2404.02905/): The article discusses Visual AutoRegressive modeling (VAR), a new image learning method that outperforms diffusion transformers in terms of speed, image quality, and scalability.
