Humans possess an innate ability to interpret the intentions behind emotional expressions and physical actions. This skill enables us to approach someone with welcoming gestures or withdraw from someone displaying aggressive body language. While artificial intelligence (AI) has made significant progress in recognizing emotions, it still struggles to interpret the underlying intentions behind those emotions. A recent study, involving participants from Japan and Taiwan, aims to close this gap in understanding.
Consider a situation where an individual approaches you from afar. Without seeing their face or hearing their voice, you must quickly determine if they are a friend or a potential threat. Humans adeptly decipher nuanced body movements, allowing for quick assessments that can be vital for survival. In contrast, most AI systems have concentrated on detecting basic emotions, like joy or sadness, and have overlooked the critical aspect of social intent—the cues we send to others. For a robot or AI application, understanding whether someone poses a danger is far more crucial than merely recognizing their emotional state.
Researchers have established a new standard for assessing "embodied social intention," shedding light on how we express threats and highlighting a significant "alignment gap" between human perception and AI interpretation. This research is being presented at the 20th IEEE International Conference on Automatic Face and Gesture Recognition (FG2026) in Kyoto, under the title "Friend or Foe? Benchmarking Human Perception and ST-GCN Decoding of Embodied Social Intention."
As part of this study, Tohoku University researchers documented 160 motion-capture performances from 80 individuals across Japan and Taiwan. The performers acted out friendly or hostile gestures directed towards an "imaginary alien" unfamiliar with human customs and language, relying solely on nonverbal communication methods.
Examples of friendly gestures included bending forward to convey politeness and opening arms as a sign of goodwill. In contrast, hostile interactions were represented through aggressive actions, such as throwing objects to distance the alien.
Additionally, 77 observers from Japan, Taiwan, and China assessed the videos, categorizing the movements as friendly or hostile. Notably, Taiwanese performers used large, dynamic gestures to convey hostility, making it easily recognizable. Their vigorous movements communicated their intentions effectively to the viewers. However, Japanese performances exhibited more restrained and subtle hostile movements, containing significantly less energy—ten times less than the Taiwanese actions—resulting in a higher recognition rate among Japanese viewers (76% accuracy) compared to their Taiwanese (69%) and Chinese counterparts (65%).
When the researchers applied an AI model (ST-GCN) to the data, they uncovered a crucial shortfall. While the AI achieved a 69% accuracy rate in interpreting the videos, it did not demonstrate human-like reasoning. Observers from different cultures displayed consistent agreement on the intent behind the movements (with correlations exceeding 0.79), yet the AI’s correlations with human perceptions were alarmingly low at just 0.26.
Humans employ a cognitive technique known as "inverse planning," allowing us to deduce the mental objectives underlying specific movements. Conversely, the AI merely identified physical patterns, neglecting the deeper social implications of subtle or passive-aggressive body language. For instance, a person standing still with arms tightly crossed and slightly turned away is perceived by humans as a signal to back off, while the AI might interpret this as non-threatening due to minimal movement.
The identified "alignment gap" poses potential risks for human-machine interactions. AI systems that can recognize high-energy threats but overlook subtle cues of low-energy hostility might fail to mitigate tension in nuanced situations. Closing this divide will necessitate developing AI that not only recognizes behaviors accurately but also aligns closely with human social cognition, allowing it to interpret both movements and their intended meanings effectively.

