Counting the number of items in a visual scene remains a fundamental yet challenging task in computer vision. Traditional approaches rely on domain-specific counting architectures trained on datasets with predefined object categories. However, recent progress in developing large-scale multimodal vision-language models (VLMs) suggests that these domain-general architectures may offer a flexible alternative for open-set object counting. In this study, we systematically compare the performance of state-of-the-art specialized counting architectures against VLMs across two established counting datasets and a novel benchmark designed for fine-grained control over visual properties of test images. Our findings show that most VLMs can approximately enumerate the number of items in a visual scene, matching or even surpassing the performance of specialized computer vision architectures. Notably, enumeration accuracy significantly improves when VLMs are prompted to generate intermediate representations - such as object locations and verbal labels - prior to counting. Nevertheless, none of the models can reliably count the number of objects in complex visual scenes, showing that further research is still needed to create AI systems capable of robust counting in realistic environments. We made our code publicly available.
Assessing the Visual Enumeration Abilities of Specialized Counting Architectures and Vision-Language Models
Hou K.;Mi J.;Zorzi M.;Ballan L.;Testolin A.
2026
Abstract
Counting the number of items in a visual scene remains a fundamental yet challenging task in computer vision. Traditional approaches rely on domain-specific counting architectures trained on datasets with predefined object categories. However, recent progress in developing large-scale multimodal vision-language models (VLMs) suggests that these domain-general architectures may offer a flexible alternative for open-set object counting. In this study, we systematically compare the performance of state-of-the-art specialized counting architectures against VLMs across two established counting datasets and a novel benchmark designed for fine-grained control over visual properties of test images. Our findings show that most VLMs can approximately enumerate the number of items in a visual scene, matching or even surpassing the performance of specialized computer vision architectures. Notably, enumeration accuracy significantly improves when VLMs are prompted to generate intermediate representations - such as object locations and verbal labels - prior to counting. Nevertheless, none of the models can reliably count the number of objects in complex visual scenes, showing that further research is still needed to create AI systems capable of robust counting in realistic environments. We made our code publicly available.Pubblicazioni consigliate
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.




