When a large language model (LLM) responds to inquiries, is it engaging in genuine reasoning like humans, or merely generating text that mimics rational thought? This distinction is crucial, as it influences our trust in AI, the level of oversight required, and its implications for the real world.
Melanie Mitchell from the Santa Fe Institute suggests that our current methods for assessing machine cognition are inadequate, positing that AI represents a form of “alien intelligence” that functions on non-human cognitive processes. On a recent episode of The Joy of Why, Mitchell discussed with Stephen Strogatz how psychological techniques, originally designed to study infant and animal cognition, can be adapted to evaluate AI. She proposed six principles for improving the assessment of machine cognition, covering topics from the difficulties of deciphering AI processes to novel AI applications in mathematics. One intriguing comparison she makes is to a math-performing horse from the early 20th century, highlighting the complexities of intelligence assessment.
In their conversation, recorded on July 23, 2026, Mitchell reflected on the rapid advancements in AI since she last appeared on the show five years ago, shortly before the emergence of ChatGPT. While she expressed astonishment at how models trained on vast datasets of human language and images have reached their current capabilities, she also remarked on the polarized reactions within the AI community and society at large—divisions between those who see AI as a potential danger and those who view it as a positive development.
Mitchell, a cognitive scientist and computer scientist, advocates for considering AI from the perspective of fields like developmental psychology—examining how young children acquire intelligence, for instance. This cross-disciplinary approach offers insights valuable for understanding AI, which operates under different cognitive mechanisms than humans, despite being trained on human-generated data.
The podcast's discussion extended to the challenges of interpreting AI’s workings and recent mathematical breakthroughs made possible by AI, including the solution to a longstanding problem proposed by hungarian mathematician Paul Erdős. Mitchell noted that while past AI models struggled with adaptability beyond their training environments, large language models may exhibit enhanced ability in transferring knowledge across different domains due to their comprehensive training datasets. Yet, she cautioned that we still lack clear understanding of what constitutes successful knowledge transfer.
As the conversation progressed, Mitchell highlighted the need for robust experimental frameworks in AI research, citing examples where AI appeared to outperform human understanding without truly grasping the underlying concepts. Moreover, she encouraged thoughtful evaluation of AI systems, advocating for controlled experiments akin to those practiced in psychology to ensure reliable results in AI assessments.
One illustrative historical anecdote was the story of Clever Hans, a horse believed to perform arithmetic tasks. It was later revealed that Hans was simply responding to subtle cues from his human trainers rather than demonstrating genuine mathematical ability. This example served to underscore the importance of rigorous testing in order to discern genuine understanding from coincidental performance in AI systems.
Moving to current AI capabilities, Mitchell espoused the need for better experimental design in AI assessments, cautioning against drawing conclusions based solely on benchmark results. This includes a discussion on the ongoing fusion and separation of AI from traditional cognitive science domains, as well as a reflection on how models are evaluated based on task performance versus deeper understanding.
Mitchell also expressed hope for the future development of AI interpretability. She suggested that applying mechanistic interpretability—deeper analysis of AI operations using techniques modeled after neuroscience—could facilitate a more profound grasp of AI limitations and capabilities.
The episode closes with broader reflections on the role of human intelligence and creativity in an AI-driven future. While acknowledging the potential for AI to overshadow human contributions in certain domains, both Mitchell and Strogatz emphasize the enduring value of human insight, curiosity, and the nuanced understanding that comes from lived experience—attributes that machines cannot replicate.
Listeners can tune into The Joy of Why on platforms like Apple Podcasts, Spotify, and TuneIn, or stream it directly from Quanta for further explorations into complex subjects in mathematics and science.




