Machine learningOther
u-μP: The Unit-Scaled Maximal Update Parametrization
The u-$\mu$P scheme merges Maximal Update Parametrization and Unit Scaling techniques to make model hyperparameters size-independent and easier to train in low-precision, leading to more efficient models that function immediately in FP8.
Featured in No. 59 on 31 Jul 2024 · 7 days after release · 31 citations today
- Released
- 24 Jul 2024
- First featured
- No. 59 · 31 Jul 2024
- Citations (Semantic Scholar)
- 31
- Influential citations
- 0
- Published in
- Not yet, as far as Semantic Scholar knows
- Shares when featured
- 93
- Identifier
- arXiv:2407.17465
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).