Study Reveals AI’s Inability to Understand Human Social Interactions in Motion
A new study conducted by Leyla Isik, a professor of cognitive science at Johns Hopkins University, highlights a critical limitation in current artificial intelligence models: their inability to accurately interpret human social interactions in dynamic scenes. Despite their prowess in recognizing objects and faces in still images, over 350 AI models specializing in video, image, or language analysis struggled to interpret social behavior in three-second video clips. Human participants, on the other hand, consistently evaluated these scenes with high agreement. Even language models performed slightly better when aided by human-written descriptions but still fell short of matching human understanding. Researchers attribute this limitation to the way AI neural networks are structured, focusing on static imagery rather than dynamic events that engage different areas of the human brain. This gap, termed a ‘blind spot’ by the study’s authors, could severely hamper the use of AI in real-world environments like autonomous driving, where understanding human intentions is crucial. The study underlines that despite massive progress, AI still lacks the nuanced comprehension required to interact seamlessly with human society.
