ML-QuantSubscribe

Machine learningML & AI Methods

Can We Remove the Square-Root in Adaptive Gradient Methods? A Second-Order Perspective

The research explores the performance of adaptive gradient optimizers without the square root, showing they maintain performance on transformers and improve on convolutional architectures.

Featured in No. 64 on 5 Sep 2024 · · 27 citations today · published in International Conference on Machine Learning

Released
5 Feb 2024
First featured
No. 64 · 5 Sep 2024
Citations (Semantic Scholar)
27
Influential citations
2
Published in
International Conference on Machine Learning
Shares when featured
45
Identifier
arXiv:2402.03496

Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).

    Type to search. Try rough volatility, LLM agents or FinGPT.

    ↑↓ move↵ openesc closeFull search page