ML-QuantSubscribe

Machine learningLLMs & Text

Do Large Language Model Benchmarks Test Reliability?

The article highlights the need for reliable large language models, criticizes current benchmarks for their inadequacy, and suggests the use of platinum benchmarks to reduce label errors and ambiguity.

Featured in No. 92 on 9 Apr 2025 · · 54 citations today

Released
5 Feb 2025
First featured
No. 92 · 9 Apr 2025
Citations (Semantic Scholar)
54
Influential citations
6
Published in
Not yet, as far as Semantic Scholar knows
Shares when featured
73
Identifier
arXiv:2502.03461

Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).

    Type to search. Try rough volatility, LLM agents or FinGPT.

    ↑↓ move↵ openesc closeFull search page