Artificial General Intelligence? Not Achieved Yet

Artificial General Intelligence? Not Achieved Yet
Summary
The pursuit of artificial general intelligence (AGI) has spanned nearly 80 years of predictions.
Recent developments in large language models (LLMs) have reignited hope, despite longstanding skepticism.
LLMs still struggle with real-world understanding and decision-making, limiting their effectiveness.

Share

Bookmark

Newsletter

For the past 80 years, the quest for artificial general intelligence (AGI)—machines capable of performing any intellectual task at or beyond human levels—has fueled both intrigue and concern. This endeavor began with the introduction of the term “artificial intelligence” during a Dartmouth workshop in 1956, which proposed that all facets of learning and intelligence could potentially be simulated by machines.

The evolution of computing technology led Herbert A. Simon, a renowned economist and Turing Award winner, to speculate in 1965 that machines would, within two decades, match human work capabilities. He was joined by Marvin Minsky, another Turing Award recipient and co-founder of MIT's AI lab, who forecasted in 1970 the emergence of a machine with human-like intelligence within eight years, suggesting that soon after, such a machine would achieve an exponential leap in capability.

After years of alternating between hope and skepticism, renewed enthusiasm surged for AGI following OpenAI's release of ChatGPT on November 20, 2022. ChatGPT, along with similar large language models (LLMs), utilizes complex patterns found in extensive text datasets to generate impressive, coherent responses to a wide range of inquiries. Many observers were amazed, leading some to mistakenly perceive them as possessing knowledge beyond human comprehension.

In October 2023, Google’s Blaise Agüera y Arcas and Peter Norvig authored an opinion piece asserting that AGI was already a reality, despite the LLMs engaging in quirky tasks like counting fictional Russian bears in space. During the summer of 2023, Anthropic CEO Dario Amodei expressed a more measured outlook, suggesting AGI could be realized in two to three years. Fast-forward to early 2025, Amodei predicted that by 2026 or 2027, AI systems would outperform humans in nearly every domain and went so far as to claim AI could significantly extend human lifespans. His confidence in public naivety was apparent when he stated that most human diseases could be cured within a decade.

On November 11, 2024, Sam Altman from OpenAI predicted AGI would materialize by 2025. By December of that year, he proclaimed that AGI had indeed been achieved but went unnoticed: “AGI kind of went whooshing by.” In March 2026, Nvidia's CEO Jensen Huang echoed this sentiment, stating that AGI had been reached.

However, these assertions must be viewed with caution as they come from vested interests. OpenAI is reliant on attracting both clients and investments, while Nvidia benefits from a thriving OpenAI, which generates demand for its chips.

Despite decades of intense effort by highly skilled individuals to develop human-level AI, the advent of LLMs represents a costly diversion. These models lack a genuine understanding of real-world relationships—leading to repeated errors. Although extensive training has improved their performance, it hasn't completely resolved their issues. Trainers remain unprepared for all possible future queries, and due to the stochastic nature of LLMs, they still occasionally offer flawed responses to familiar prompts.

Recently, I re-evaluated the performance of GPT-5.6 Luna with prompts known to challenge its capabilities. A classic Winograd schema prompt demonstrated significant progress, with GPT-5.6 accurately identifying the antecedent of an ambiguous pronoun. However, when I slightly altered the phrasing, it repeatedly provided an incorrect answer, illustrating a fundamental limitation: a lack of understanding that hindered its ability to apply learned logic across similar contexts.

In an examination of Will Rogers' famous joke regarding Okies migrating from Oklahoma to California, GPT-5.6 exhibited difficulty grasping the humor's essence, which relies on the contradiction of averages. Even publicly available explanations of the joke did not prevent it from faltering.

Another challenge emerged when I asked ChatGPT-5 about a new game concept: Rotational Tic-Tac-Toe. While initial responses were improved, the model's unpredictable nature led to inconsistent answers, showcasing the inherent limitation that still persists in the AI's reasoning capabilities.

Despite fervent claims surrounding their intelligence, LLMs like ChatGPT remain constrained by their inability to make meaningful connections between words and the realities they represent, limiting their capacity for sound decision-making or generating advice in unique scenarios. This inherent unpredictability highlights the need for skepticism regarding proposals from companies like Anthropic to slow down AI capability advancements. Such suggestions may serve more as a strategy to maintain their market status amid financial pressures than a genuine aim for responsible AI development.

Loading comments...