Latest Findings · ASI Lab

Top models score under 37% on HumanEval-V — visual reasoning, not coding skill, is the bottleneck

On a new 253-task benchmark where models must turn diagrams into working code, Claude 3.5 Sonnet reaches just 36.8% pass@1, Pixtral 124B gets 21.3%, and many open-weight models fall below 10%.

Fengji Zhang, Linquan Wu, Huiyu Bai, Guancheng Lin, Xiao Li, Xiao Yu, Yue Wang, Bei Chen, Jacky W. Keung
Agentic Software Intelligence Research Lab · Department of Computer Science, City University of Hong Kong
ACM TOSEM · 2026

HumanEval-V is a benchmark of 253 human-annotated tasks that push large multimodal models to read an algorithm's logic straight from a diagram and emit working Python. Across 27 state-of-the-art LMMs, the best result is only 36.8% pass@1. The authors show the real failure point is high-level visual reasoning, not coding ability.

HumanEval-V: Systematic Evaluation of Visual Reasoning in Large Multimodal Models for Code Generation
Fig. 1. An example task from HumanEval-V. LMMs are required to interpret the diagram and function signature, deduce the underlying algorithm, and implement the corresponding Python function.

Why it matters

While recent advances in Large Multimodal Models (LMMs) have demonstrated impressive visual capabilities across various domains, their visual reasoning abilities in coding contexts remain largely unexplored.

Key findings

HumanEval-V

What it means for practitioners

Don't trust general multimodal benchmarks when picking an LMM for visual-programming work: HumanEval-V scores map poorly onto them. Expect even state-of-the-art models to fail on diagrams encoding spatial transforms, topological relations and dynamic patterns; hallucinated visual details are common. Two-stage pipelines (have the model describe the diagram, then generate code with a strong coder) markedly lift performance, as does execution-based feedback for iterative self-correction.

Get the paper & cite it

Download full PDF ↓ View at DOI

Official citation: Fengji Zhang, Linquan Wu, Huiyu Bai, Guancheng Lin, Xiao Li, Xiao Yu, Yue Wang, Bei Chen, Jacky W. Keung (2026). HumanEval-V: Systematic Evaluation of Visual Reasoning in Large Multimodal Models for Code Generation. ACM Transactions on Software Engineering and Methodology. DOI: 10.1145/3813804.

Large Multimodal Models Visual Reasoning Code Generation Benchmark Evaluation Large Language Models Algorithmic Reasoning Vision-Language-Code Alignment