CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition
Lai Wei, Chengqi Li, Jiapeng Li, Ruina Hu et al. · arXiv · 2026
Abstract
Learning from context is usually measured with text. In practice the context a model must work from is often a figure, a table, a converted financial report, or a map. This benchmark separates the failure into three questions — can the model locate the relevant context, apply new information from it, and absorb genuinely new knowledge — across 3,443 instances. The best of six recent multimodal models scores under 0.29.
Why it matters
A useful corrective if you plan to feed charts and scanned documents to a multimodal model. That headline number says the capability is far weaker than the demos imply.
https://arxiv.org/abs/2607.25294