📊 Full opportunity report: Engineering Is Automated. Research Is the Residual. on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
AI systems now automate most core engineering tasks in AI R&D, reaching near saturation on several benchmarks. However, research—particularly creative and strategic aspects—remains less automated, leaving a residual human role.
Recent developments in AI capabilities demonstrate that systems can now automate the core engineering tasks involved in AI research and development, reaching near-complete saturation on multiple benchmarks. Meanwhile, the more creative, strategic aspects of research remain less automated, with human involvement still essential. This shift has significant implications for the future of AI innovation and the role of human researchers.
Several key benchmarks measuring AI proficiency in core AI R&D skills have shown rapid progress, approaching or reaching saturation levels. For example, the CORE-Bench, which tests the ability to reproduce research papers, improved from 21.5% in September 2024 to 95.5% by December 2025, with some experts declaring it ‘solved.’ Similarly, the MLE-Bench, assessing performance on Kaggle competitions, advanced from 16.9% in October 2024 to 64.4% in February 2026, surpassing the original design expectations. These trends indicate that AI systems can now handle complex, previously friction-laden engineering tasks with minimal human intervention.
Conversely, research that involves creativity, strategic planning, and hypothesis generation—such as designing new experiments or formulating novel theories—remains less amenable to automation. Thorsten Meyer notes that while engineering tasks are increasingly automated, the residual research component may itself be a form of engineering at scale, suggesting that the boundary between engineering and research is blurring. The key open question is whether future AI systems will also automate the more abstract, innovative aspects of research.
Engineering is automated.
Research is the residual.
Six skill benchmarks. Edison’s framing. The question Clark leaves open is whether research is just engineering at scale.
Jack Clark’s Import AI #455 catalogs six benchmarks measuring AI capability on AI R&D tasks and concludes “AI can today automate vast swatches, perhaps the entirety, of AI engineering.” The residual question is research. The structural read on the residual: it may not be a permanent moat.
Six skills. One trajectory.
Clark catalogs six benchmarks measuring AI capability on AI R&D-relevant tasks. Each individual benchmark could be noise. Six benchmarks moving together is a curve. The pattern is the cascade observed across the broader Clark series — visible here in the specific R&D-skill domain.
Three data points. Mixed signal.
Clark provides three data points on the creative-spark question. Yes-evidence: Erdős-1051, centaur math discovery, sporadic Move-37-style moments. No-evidence: low yield, framing dependence, absence of acceleration. The mixed signal is the honest read.
The data supports two readings. Pessimistic: rare moments suggest creative insight is qualitatively distinct from engineering work. Optimistic: rare moments are an artifact of low-volume exploration; more shots on goal yields more discoveries. Both readings are consistent with Clark’s “vast swatches, perhaps the entirety” claim. They differ on the residual.
Five dimensions Clark gestures at but leaves underdeveloped.
Clark’s section is rigorous on the empirical evidence. Five strategic dimensions matter for the institutional response that the Clark series synthesis argues is structurally inadequate.
Two readings. Different equilibria.
The structural question Clark leaves open: is research a permanent moat that bounds automated AI R&D, or is it engineering at scale that dissolves with more shots on goal? Both readings are consistent with the current data. They differ by orders of magnitude in consequences.
Productivity multiplier years
Recursive loop operational
Five audiences. Asymmetric cost of being wrong.
The institutional response should not bet on inspiration being a permanent moat. If the distinction holds, capacity built is still useful. If it closes, capacity is necessary. Asymmetric cost-of-being-wrong points toward building now.
IN INDUSTRY
IN ACADEMIA
POLICYMAKERS
INVESTORS
EVERYONE ELSE
Engineering is automated. The residual is the question. The institutional response should not bet on inspiration being a permanent moat.
Implications for AI Development and Human Roles
The near-complete automation of AI engineering tasks suggests a potential shift in how AI research is conducted, with human researchers focusing more on strategic and creative aspects. This could accelerate development cycles, reduce costs, and democratize AI innovation, but also raises concerns about the future role of human insight and oversight. The transition may reshape the landscape of AI R&D, emphasizing the need to understand what aspects of research remain inherently human.
As an affiliate, we earn on qualifying purchases.
Progress in AI Benchmarks Indicates Rapid Capability Growth
Over the past 18 months, multiple independent benchmarks—covering research reproduction, Kaggle competition performance, and kernel design—have shown rapid progress toward saturation. The CORE-Bench, which measures the ability to reproduce research papers, has seen a 4.4× improvement, with some experts calling it ‘solved.’ The MLE-Bench, assessing practical ML competition skills, has also advanced significantly, indicating that AI systems are approaching human-level proficiency in engineering tasks critical to AI R&D. These developments suggest a structural shift in AI capabilities, driven by advances in foundational models and automation techniques.
“The residual research component may itself be a form of engineering at scale, suggesting that the boundary between engineering and research is blurring.”
— Thorsten Meyer
As an affiliate, we earn on qualifying purchases.
Unclear Extent of Automation in Creative Research Tasks
It remains uncertain how much of the creative, strategic, and hypothesis-driven aspects of research can be automated. While engineering tasks are nearing full automation, the capacity of AI to generate novel theories, design experiments, and make strategic decisions is still developing. Experts differ on whether future AI systems will fully automate these aspects or if they will remain inherently human domains for the foreseeable future.
AI research paper reproduction tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Milestones in AI R&D Automation and Human Oversight
In the coming months, researchers will likely focus on expanding automation to more abstract research tasks, testing the limits of current models. Monitoring developments in AI’s creative capabilities and the integration of automated research workflows will be critical. Additionally, institutions may reevaluate the division of labor between AI systems and human researchers, potentially leading to new standards and practices in AI R&D.
As an affiliate, we earn on qualifying purchases.
Key Questions
What are the main benchmarks indicating AI automation progress?
Key benchmarks include CORE-Bench, measuring research reproduction; MLE-Bench, assessing Kaggle competition performance; and various kernel design metrics. All show rapid improvement toward saturation, indicating increasing automation of engineering tasks.
Does automation mean AI can now do all research independently?
Not yet. While engineering and reproducibility tasks are approaching full automation, creative, strategic, and hypothesis-driven research still rely heavily on human insight. The extent of future automation in these areas remains uncertain.
How might this shift impact human researchers?
Human researchers may shift focus toward high-level strategy, theory development, and oversight, as routine engineering tasks become automated. This could accelerate innovation but also requires new skills and roles.
What are the risks of over-relying on automated AI in research?
Potential risks include reduced diversity of approaches, overfitting to current models, and loss of human intuition in hypothesis generation. Ensuring oversight and maintaining human judgment will be essential.
Source: ThorstenMeyerAI.com