Have you ever wondered if artificial intelligence really sees the world the same way we do? While AI models, known as Visual Language Models (or VLMs, if you like specific names), are performing well in visual tasks that need a high level of understanding—kind of like college-level exams—they often stumble when it comes to seeing simple things, like the direction or position of an object. It’s a bit like having a genius student who can solve complex math problems but can’t tell left from right. This research paints a picture of how AI ‘vision’ might differ from ours.
In this intriguing study, scientists used a toolkit from neuropsychology—a field focused on the connection between brain functions and behaviors—to test three top-performing AI models. They used 51 tests from clinical and experimental setups to see how these models ‘see’ compared to how healthy adults do. Amazingly, while these models pass with flying colors when identifying objects, they fail in understanding basic visual concepts like the orientation or continuity of objects. These shortcomings are considered quite significant in humans and highlight a big difference in how artificial and human intelligence process visual information.
Imagine the impact of this discovery: while your phone’s AI might be great at recognizing your face to unlock your phone, it might not really understand what’s on your phone screen beyond that recognition. This means that as AI continues to evolve, researchers will need to focus on teaching these models to see the world as we do, developing an intuitive understanding of basic visual elements. This research could lead to creating AI that integrates smoothly into our daily lives, with a more human-like perception and understanding of the world around us.
Surprisingly, AI models excel at recognizing complex objects but can falter when it comes to understanding simple visual concepts such as which way an object is facing.
FAQs
Why do Visual Language Models struggle with basic visual concepts?
Visual Language Models can identify complex objects but often miss simpler visual cues like orientation and position because these foundational concepts don’t require explicit training for humans but are not naturally acquired by AI.
How were VLMs tested for their visual capabilities?
Researchers used 51 tests from clinical and experimental sources, designed to examine visual processing, to compare VLMs to a normative performance seen in healthy adults.
What implications does this research have for AI development?
This study suggests that AI needs to improve in basic visual understanding, indicating a need for developers to focus on teaching AI not just complex recognition, but also intuitive visual perception.
Can AI models become as visually adept as humans?
While AI models excel in many tasks, achieving human-like visual perception would require advancements in training these systems to intuitively understand basic visual concepts as humans do.
How might this research affect future AI applications?
The findings could lead to the development of AI that integrates better into everyday tasks by evolving in its ability to perceive and interpret the world more like humans.
Background
Visual Language Models, or VLMs, are AI systems designed to understand and process visual information like images and video content. They are often tested on their ability to perform tasks that require understanding complex visual scenes or objects. Neuropsychology offers methods to assess how these models perform by examining aspects of visual perception that are fundamental to humans, such as recognizing positions, orientations, and the continuity of objects.
History
The exploration of artificial intelligence in visual processing has evolved significantly over the years. Early AI systems struggled with basic image recognition, but advancements led to the development of sophisticated models capable of interpreting complex visual data. However, recent studies like this one reveal gaps in AI’s visual understanding, indicating that even state-of-the-art models do not fully replicate human vision. This research builds on past work by demonstrating that AI’s success in complex tasks doesn’t equate to a full understanding of simpler visual concepts, a crucial step toward more human-like AI vision systems.
Based on “Visual Language Models show widespread visual deficits on neuropsychological tests” by Gene Tangtartharakul, Katherine R. Storrs, available on arXiv (arxiv.org/abs/2504.10786), used under CC BY 4.0 (creativecommons.org/licenses/by/4.0/).





































































