ML-QuantSubscribe

Machine learningLLMs & Text

BLINK: Multimodal Large Language Models Can See but Not Perceive

Multimodal LLMs Visual Perception Benchmark: The authors present Blink, a benchmark for multimodal language models that tests visual perception abilities, showing that current models struggle with these tasks.

Featured in No. 46 on 24 Apr 2024 · 6 days after release · 603 citations today · published in European Conference on Computer Vision

Released
18 Apr 2024
First featured
No. 46 · 24 Apr 2024
Citations (Semantic Scholar)
603
Influential citations
78
Published in
European Conference on Computer Vision
Shares when featured
21
Identifier
arXiv:2404.12390

Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).

    Type to search. Try rough volatility, LLM agents or FinGPT.

    ↑↓ move↵ openesc closeFull search page