---
title: Fine-tuning can cripple your foundation model; preserving features may be the solution
url: https://www.ml-quant.com/papers/arxiv/2308.13320/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2308.13320
source_url: https://arxiv.org/abs/2308.13320
featured: 2024-07-03
citations: 102
topic: ML & AI Methods
---


# Fine-tuning can cripple your foundation model; preserving features may be the solution

Concept forgetting in AI models can be significantly reduced by a new fine-tuning method called LDIFS, which helps retain pre-trained knowledge while working on different tasks.

- Source: https://arxiv.org/abs/2308.13320
- Identifier: arXiv:2308.13320
- Released: 2023-08-25
- First featured: Quant Letter No. 55 (2024-07-03): https://www.ml-quant.com/issues/2024-07-03/
- Citations (Semantic Scholar): 102
- Published in: Trans. Mach. Learn. Res.
- Topic: ML & AI Methods

## Related

- [Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models](https://www.ml-quant.com/papers/arxiv/2501.01423/): The paper proposes a new model, VA-VAE, that aligns the latent space with pre-trained vision foundation models, enabling faster convergence of Diffusion Transformers in high-dimensional latent spaces and achieving top performance on ImageNet 256x256 generation.
- [Adjoint Matching: Fine-tuning Flow and Diffusion Generative Models with Memoryless Stochastic Optimal Control](https://www.ml-quant.com/papers/arxiv/2409.08861/): The study presents Adjoint Matching, a new algorithm that enhances dynamical generative models by refining reward fine-tuning, leading to improved consistency, realism, and adaptability to unseen human preference reward models.
- [Continual Diffusion: Continual Customization of Text-to-Image Diffusion with C-LoRA](https://www.ml-quant.com/papers/arxiv/2304.06027/): CLoRA, a new method, prevents catastrophic forgetting in text-to-image models when introducing new concepts, achieving top performance in continual learning settings for image classification.
- [Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained Transformers](https://www.ml-quant.com/papers/arxiv/2409.20537/): The paper presents Heterogeneous Pre-trained Transformers (HPT), a method for training robotic models across various tasks, improving the performance of fine-tuned policies by over 20% on unseen tasks in both simulated and real-world environments.
- [Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models](https://www.ml-quant.com/papers/arxiv/2411.04996/): The Mixture-of-Transformers (MoT) is a sparse multi-modal transformer architecture that reduces pretraining costs and allows modality-specific processing with global self-attention.
- [ECG-FM: An Open Electrocardiogram Foundation Model](https://www.ml-quant.com/papers/arxiv/2408.05178/): ECG-FM, a transformer-based model for ECG analysis, shows strong performance in predicting cardiac conditions, having been pretrained on 2.5 million samples.
