Every few weeks, a new AI development emerges that would have seemed unimaginable just a couple of years ago.
One model has outperformed some of the brightest young mathematicians globally. Another aids in scientific exploration. Meanwhile, there are models that can write software that would typically take experts days or even weeks to complete.
Today, there are also AI systems capable of operating independently for extended periods before needing human assistance.
While I’m not declaring that we’ve achieved artificial general intelligence (AGI) based solely on these individual advancements, it’s essential to stop viewing them in isolation.
By examining these breakthroughs together, we can better understand how close we are to reaching AGI.
Beyond Simple Interactions
Earlier this year, I posited that the fundamental components of AGI were finally aligning.
Initially, AI began learning to answer questions. Following that, it developed the ability to reason through intricate problems. Recently, we’ve seen AI manage to work autonomously for much longer durations.
Reflecting on the past six months, I now believe that we’ve witnessed the addition of several pivotal advancements.
Let’s consider mathematics, for instance.
Last summer, an advanced variant of Google’s Gemini Deep Think secured a gold medal score at the International Mathematical Olympiad—an event often regarded as one of the most challenging math competitions for high school students.
Typically, only around 8% of participants achieve this level of recognition in an average year.
The AI managed to solve five out of six problems, earning 35 points out of a potential 42, graded by the same judges who assess human contenders.
This was a remarkable feat.
Moreover, Google has indicated that newer iterations of Deep Think are now assisting mathematicians and scientists in addressing real research challenges in mathematics, physics, and computer science.
In an internal assessment, the system achieved around 90% accuracy on advanced mathematical proofs.
This, to me, is a significant indicator of the AGI timeline.
Solving known problems is one aspect, but aiding in unexplored challenges brings us much closer to true general intelligence.
Anthropic recently showcased a similar significant achievement.
Their research teams tasked Claude AI agents with an open-ended machine-learning task devoid of any pre-established answers.
The agents were required to generate original ideas, write code, conduct experiments, analyze their findings, and refine their methods until they reached a viable solution.
Anthropic reported that the top-performing AI team surpassed seasoned human researchers working on the same challenge.
This exemplifies the essence of general intelligence: addressing unknown issues, experimenting, and evolving one’s approach.
We’re observing similar trends in the realm of software engineering.
Anthropic tasked Claude with developing a new C compiler capable of compiling the Linux kernel, one of the most intricate open-source software initiatives ever undertaken.
Across two weeks, Claude engaged in nearly 2,000 coding sessions, processed around 2 billion input tokens, generated 140 million output tokens, and ultimately crafted about 100,000 lines of code.
And it succeeded.
Claude produced software capable of running one of the world’s most complex operating systems across various computer processors.
Even more astonishing is the fact that this endeavor cost less than $20,000 in API usage.
That’s an incredibly low figure considering the scale of the project, which could easily require hundreds of thousands of dollars in engineering salaries for a similar undertaking.
In my opinion, this marks another advancement on the road to AGI.
To think as a human does, one must handle substantial, unfamiliar projects, maintain focus over prolonged periods, and resolve challenges along the way.
On that note, AI has shown significant improvements in maintaining concentration.
As I highlighted earlier, the METR research group evaluates how long an AI can engage with a genuine problem before requiring human intervention.
And the trend is impressive.
METER shows that the amount of work completed autonomously by leading AI models has roughly doubled every seven months.
Anthropic has observed a parallel trend.
After analyzing millions of Claude coding sessions, they found that some of the longest independent coding sessions nearly doubled within three months, increasing from under 25 minutes to over 45 minutes before the AI required assistance.
If this trajectory continues, those 45 minutes might extend to four hours, eventually reaching an entire workday.
This reinforces my belief that viewing artificial general intelligence as a conclusive endpoint is misguided.
There won’t be a definitive announcement proclaiming the arrival of AGI.
Instead, we will see a steady accumulation of advancements.
My Perspective
In my view, the developments we've witnessed over the past six months signal that we are inching closer to AGI.
Intelligence encompasses more than just knowing answers; it involves applying that knowledge to achieve meaningful outcomes.
And today’s AI is becoming remarkably adept at this.
Nonetheless, a critical question arises.
Given AI's increasing capabilities, why are businesses not reaping more substantial benefits?
We’ll delve into this question further in tomorrow’s edition.




