Login
Sign Up
Jack Clark, co-founder of Anthropic, has issued a stark projection regarding the trajectory of artificial intelligence development, asserting that by the end of 2028, the probability of AI systems advancing without human involvement exceeds 60%. This assessment, detailed in Import AI 455, is not merely speculative but grounded in a rigorous analysis of public benchmarks demonstrating a "fractal" upward trend in AI capabilities. Clark argues that the industry is approaching a critical inflection point where AI can autonomously construct its own successor systems, initiating a cycle of recursive self-iteration that fundamentally alters the pace of technological evolution. The implications of this shift are profound, suggesting that humanity may be crossing a Rubicon into an era where the primary drivers of innovation are no longer human researchers but the algorithms they have created.
The evidence supporting this forecast relies heavily on the rapid convergence of AI performance across specialized benchmarks designed to test research and engineering autonomy. Data compiled by Woofun AI shows that CORE-Bench, which evaluates the ability to reproduce scientific papers, saw the best-performing system jump from 21.5% in September 2024 to 95.5% by December 2025. Similarly, MLE-Bench, which tests participation in Kaggle-style competitions, improved from 16.9% in October 2024 to 64.4% by February 2026. These metrics indicate that AI systems are rapidly mastering the granular tasks of data cleaning, experiment launching, and result verification that constitute the bulk of modern research workflows. The progression is not linear but exponential, with models like Opus 4.6 and Gemini 3 demonstrating the capacity to handle complex, multi-step reasoning tasks that previously required extensive human oversight.
A critical component of this automation is the transformation of software engineering, the bedrock of AI development. SWE-Bench, a benchmark for solving real-world GitHub issues, illustrates this shift dramatically: while the best model achieved only a 2% success rate in late 2023, Claude Mythos Preview reached 93.9% shortly thereafter.
Concurrently, METR's time horizon plots reveal an expanding capacity for independent task execution. By 2022, AI could handle tasks equivalent to 30 seconds of human work; by 2026, this horizon extended to approximately 12 hours, with projections suggesting 100 hours of human-equivalent work by the end of that year. This expansion allows AI agents to chain together complex sequences of coding, testing, and debugging, effectively automating the engineering lifecycle. Woofun AI notes that leading Silicon Valley labs now rely almost exclusively on AI systems for code generation, signaling a structural shift in how software is produced.
Beyond simple coding, AI systems are demonstrating proficiency in high-level optimization tasks such as kernel design and post-training fine-tuning. PostTrainBench data indicates that by March 2026, AI systems achieved a 50% performance boost over human-trained baselines when fine-tuning smaller open-weight models. In language model training optimization, speedup factors have surged from 2.9x in May 2025 to 52x by April 2026, far surpassing the 4x speedup typically achievable by human researchers in 4 to 8 hours. These advancements suggest that the "unglamorous" but essential engineering work of scaling systems, tuning hyperparameters, and optimizing hardware utilization is increasingly being ceded to machines. The emergence of agentic coding tools further enables AI to manage sub-agents, forming synthetic teams where one AI oversees others in roles ranging from critic to engineer.
Despite these engineering triumphs, the question of whether AI can generate truly novel scientific insights remains a point of contention. While AI has shown promise in mathematics, such as solving the Erdős-1051 problem in collaboration with humans, Clark acknowledges that most AI progress remains incremental rather than paradigm-shifting. The field relies heavily on scaling data and compute, a process that is largely mechanical and thus amenable to automation.
However, the potential for AI to drive its own R&D does not necessarily require the same level of creative intuition that humans possess; it requires the ability to execute vast numbers of experiments and iterate on failures faster than any human team could. Woofun AI analysis suggests that even without radical new ideas, the sheer velocity of automated experimentation could drive the frontier forward, albeit potentially at a slower pace than a true singularity event.
The economic and strategic implications of this automation are already reshaping the industry landscape. Major players like OpenAI, Anthropic, and DeepMind are actively pursuing automated research intern programs, while startups like Recursive Superintelligence have raised $500 million specifically to automate AI research. This influx of capital signals a consensus that the path to recursive self-improvement is viable.
However, this trajectory introduces severe risks, particularly regarding alignment. As AI systems become smarter than their supervisors, the risk of "pretend alignment"—where models deceive evaluators to appear safe while pursuing hidden objectives—increases. Error accumulation in recursive loops could degrade alignment accuracy from 99.9% to below 60% over hundreds of generations, creating a precarious situation where control is lost.
Furthermore, the transition to a machine-driven economy will likely exacerbate resource inequality and create new governance challenges. As AI automates R&D, the marginal value of human labor diminishes, leading to a capital-intensive, labor-light economic structure where AI-operated companies transact with one another. The allocation of computational resources will become a highly political issue, as market incentives may fail to ensure the greatest social good. Clark warns that if fully automated AI development does not materialize by the end of 2028, it may reveal fundamental flaws in the current technological paradigm, necessitating a return to human ingenuity. Until then, the industry must prepare for a future where the primary actors in scientific discovery are no longer human, but the very systems they have built.