---
title: VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
url: https://www.ml-quant.com/papers/arxiv/2401.13649/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2401.13649
source_url: https://arxiv.org/abs/2401.13649
featured: 2024-01-30
citations: 0
topic: Other
---


# VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks

VisualWebArena, a benchmark for assessing the performance of multimodal web agents on visually grounded tasks, is introduced, highlighting gaps in current multimodal language agents.

- Source: https://arxiv.org/abs/2401.13649
- Identifier: arXiv:2401.13649
- Released: 2024-01-24
- First featured: Quant Letter No. 35 (2024-01-30): https://www.ml-quant.com/issues/2024-01-30/
- Citations (Semantic Scholar): 0
- Published in: not yet
- Topic: Other

## Related

- [OpenAgents: An Open Platform for Language Agents in the Wild](https://www.ml-quant.com/papers/arxiv/2310.10634/): Language Agent Platform: OpenAgents, a platform for utilizing and developing language agents in daily life, is introduced, providing a user-friendly interface and a basis for future research.
- [HyperAgent: Generalist Software Engineering Agents to Solve Coding Tasks at Scale](https://www.ml-quant.com/papers/arxiv/2409.16299/): Coding Task SE Agents: HyperAgent, a multi-agent system for software engineering tasks, has achieved new benchmarks in tasks like GitHub issue resolution and code generation across different programming languages.
- [BMW Agents - A Framework For Task Automation Through Multi-Agent Collaboration](https://www.ml-quant.com/papers/arxiv/2406.20041/): A proposed agent engineering framework offers a scalable and flexible workflow for multiple autonomous agents to collaborate across various domains.
- [MaRINeR: Enhancing Novel Views by Matching Rendered Images with Nearby References](https://www.ml-quant.com/papers/arxiv/2407.13745/): Novel View Matching: The article introduces MaRINeR, a technique that enhances 3D rendering using information from a nearby image, useful for mixed-reality applications and autonomous agent training.
- [Depth Anything V2](https://www.ml-quant.com/papers/arxiv/2406.09414/): Depth Anything V2 is a new model for monocular depth estimation, using synthetic and large-scale pseudo-labeled real images for faster, more accurate results and setting a new evaluation benchmark.
- [MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark](https://www.ml-quant.com/papers/arxiv/2406.01574/): MMLU-Pro, an improved dataset, expands the Massive Multitask Language Understanding benchmark by adding tougher questions and more choices, serving as a better benchmark to monitor progress in the field.
