A few months ago, a 23-year-old amateur researcher named Liam Price made headlines by utilizing OpenAI’s ChatGPT to tackle a mathematical challenge that has puzzled experts for six decades—Erdős Problem #1196. The AI provided an innovative logical breakthrough related to primitive sets of numbers, which was subsequently confirmed by top mathematicians. This incident is one of several recent instances where large language models (LLMs) have successfully resolved long-standing mathematical issues.
Traditionally, scientists have dismissed LLMs as mere text predictors that often make fundamental math mistakes. However, advancements in models that incorporate trial-and-error learning coupled with sophisticated reasoning abilities are leading to significant progress in problem-solving capacities.
For years, AI developers have heralded the impending arrival of artificial general intelligence (AGI), though some, including OpenAI’s Sam Altman, have begun to temper these claims. The question arises: are these new LLMs merely increasing their data processing abilities, or are they manifesting genuine intelligence? In this context, Hector Zenil and his team from King's College London developed a novel method to evaluate what they term "superintelligence." They subjected leading LLMs like ChatGPT and DeepSeek to their assessment and shared the findings in Nature Communications.
In a conversation with Zenil, he explored the concept of intelligence and the new AI evaluation he and his colleagues designed, as well as how these advancements can deepen our understanding of human cognition.
Zenil explained that defining superintelligence is challenging, but fundamentally, it is about a system’s ability to effectively abstract key data features and generate predictions when feasible. Some of these predictions may pertain to random events, which is integral to the evaluation.
He clarified the distinction between superintelligence and general intelligence, suggesting that superintelligence elevates key aspects of general intelligence—specifically, abstraction and prediction. Whereas humans can forecast specific outcomes, certain elements of prediction remain beyond our grasp. The most acknowledged definitions of general intelligence imply that any task achievable by a human should also be within a machine's capability. However, the boundaries around these definitions are fluid. Some argue AGI might be a form of superintelligence, underscoring the lack of consensus in this area.
Zenil further stated that superintelligence fundamentally exceeds human capabilities, operating without the flaws inherent in human reasoning.
To assess whether LLMs demonstrate general intelligence or superintelligence, Zenil’s team concentrated on evaluating model abstraction, inverse problem-solving, and short-sequence prediction and generation. The intention was to scrutinize claims made by AI developers regarding the intellectual capabilities of their technologies. Past notions suggested AI systems were close to achieving Ph.D.-level intelligence, but discrepancies in developer statements raised concerns about their reliability. The researchers aimed to create an impartial test, avoiding the biases that commonly inform traditional intelligence assessments based on human capabilities.
A notable challenge, Zenil observed, is that many existing evaluations focus primarily on human-like intelligence. One well-known test, the ARC challenge, requires identifying patterns in image sequences—essentially demanding human-like behavior from the AI, which does not equate to genuine intelligence.
Zenil argued for widening the framework of intelligence assessment, acknowledging that human reasoning is frequently flawed and irrational. Humans might excel in social contexts, but these attributes do not constitute the essence of intelligence that scientists wish to investigate. Instead, their test aimed to examine intelligence pertinent to scientific endeavors—skills like simulation, abstraction, and predictive reasoning.
He emphasized that effective data compression is crucial for scientific advancement, which epitomizes the pinnacle of human intelligence. The progression of science has historically involved transforming seemingly random phenomena into structured causal explanations, with models serving as condensed representations of these phenomena. The quest for a "theory of everything" symbolizes a drive to encapsulate the universe's workings within simplified rules.
Zenil raised compelling issues regarding whether LLMs are genuinely advancing to higher-order reasoning or merely enhancing their pattern-matching capabilities. He discussed two contrasting perspectives: one proposing that LLMs simply mimic previous data (lacking true intelligence), and the other, predominant among deep learning proponents, positing that extensive pattern matching leads to general intelligence. Initially positioning himself between these views, Zenil now leans toward the latter, though he stresses evaluating LLMs is intricate, especially as they increasingly integrate neurosymbolic computations—an innovative blend of deep learning and symbolic AI that enhances reasoning.
In recent times, LLMs have demonstrated creative problem-solving capabilities in mathematics. However, Zenil pointed out that these newer systems should not be seen solely as LLMs anymore, as they are increasingly interwoven with symbolic systems—validation of claims he had made for some time regarding the need for such integration in achieving more nuanced intelligence.
While LLMs can create convincing facades of intelligence, Zenil argues they reveal deeper truths about human cognition. These models function statistically and emulate human language—but they successfully do so without inherent intelligence. This realization challenges assumptions surrounding language acquisition and intelligence, illustrating that even without true understanding, LLMs can exhibit mastery of language.
The ongoing dialogue about whether LLMs engage in higher-order reasoning remains unresolved. Zenil acknowledges that while LLMs exhibit more than simple pattern matching, truly groundbreaking understanding necessitates advancements in their compression techniques.
Interestingly, Zenil’s evaluation found that newer versions of LLMs often underperform compared to their predecessors on measures of intelligence, indicating that current optimization focuses too heavily on achieving human-like performance rather than fostering emergent, deeper intelligence.
This inclination illustrates a human-centric bias that disregards different forms of intelligence that could be beneficial. To address this, Zenil and his team formalized an intelligence metric based on abstraction and prediction—not merely refined pattern matching. They found that increasing the complexity of problems revealed LLMs’ limitations, indicating they primarily synthesize previously learned patterns rather than displaying genuine understanding or innovative reasoning.
Zenil emphasized that matters of superintelligent AI’s utility ultimately extend beyond mere chatbot functions. The focus is on advancing scientific discovery and improving the predictive capacity of various applications, particularly in fields like medicine and climate science.
Despite concerns regarding an AI's ability to converse with humans, Zenil mentioned cases where computational systems have produced results validating complex mathematical proofs that are nearly incomprehensible to individual researchers. The challenge lies in the relationship between superintelligent systems and human comprehension.
The intersection of neurosymbolic AI and language presents potential complications, as distinct optimization goals could result in errors or misunderstandings. The inherent biases in training data lead to frequent "hallucinations," where LLMs or even humans might present incorrect information based on flawed extrapolation from previous data.
Zenil concluded that ongoing research will leverage LLMs to explore and elucidate human behavior, exploring how these models reflect our cognitive processes. He argues that the synergy between LLMs and human analysis could yield insightful revelations about our own minds, prompting a reevaluation of how we understand intelligence itself.


/Amazon%20-%20Image%20by%20bluestork%20via%20Shutterstock.jpg)
