ML-QuantSubscribe

Machine learningML & AI Methods

Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models

The paper proposes a new model, VA-VAE, that aligns the latent space with pre-trained vision foundation models, enabling faster convergence of Diffusion Transformers in high-dimensional latent spaces and achieving top performance on ImageNet 256x256 generation.

Featured in No. 81 on 8 Jan 2025 · 6 days after release · 395 citations today · published in 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

Released
2 Jan 2025
First featured
No. 81 · 8 Jan 2025
Citations (Semantic Scholar)
395
Influential citations
79
Published in
2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
Shares when featured
14
Identifier
arXiv:2501.01423

Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).

    Type to search. Try rough volatility, LLM agents or FinGPT.

    ↑↓ move↵ openesc closeFull search page