Close Menu
AIToday7

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    What's Hot

    Previewing the Model Hardware Standard

    August 29, 2026

    Human, Beware: The Secrets We Tell AI Chatbots Aren’t Private

    August 29, 2026

    NSU Technology & Entrepreneurship Summit Sept. 30

    August 29, 2026
    Facebook X (Twitter) Instagram
    Trending
    • Previewing the Model Hardware Standard
    • Human, Beware: The Secrets We Tell AI Chatbots Aren’t Private
    • NSU Technology & Entrepreneurship Summit Sept. 30
    • States bet big on rural health startups, with a Silicon Valley twist
    • Pokémon Co. president wants to explore integrating Pokémon with wearable technology
    • Optimum Launches Advanced Security as Americans Seek Stronger Online Protection
    • Your Guide to Artificial Intelligence (AI) Degrees
    • Are We on the Verge of an Intelligence Explosion? Maybe Not.
    Facebook X (Twitter) Instagram Pinterest Vimeo
    AIToday7
    • Home
    • AI News
    • Tech News
    • AI Guides
    • Chatbots
    • Cybersecurity
    • Gadgets
    • More
      • Generative AI
      • Startups
    AIToday7
    Home»AI News»Are We on the Verge of an Intelligence Explosion? Maybe Not.
    AI News

    Are We on the Verge of an Intelligence Explosion? Maybe Not.

    aitoday7By aitoday7August 29, 2026No Comments5 Mins Read
    Share Facebook Twitter Pinterest LinkedIn Tumblr Reddit Telegram Email
    Are We on the Verge of an Intelligence Explosion? Maybe Not.
    Share
    Facebook Twitter LinkedIn Pinterest Email

    There’s growing excitement in the AI industry about the idea that today’s leading models could build the next generation of the technology. But a new study recently found top AI agents struggle on the kind of genuinely open-ended research problems required to push the field forward.

    Large language models have made rapid progress in many of the day-to-day jobs involved in machine learning research, such as writing code, generating and curating data, and running experiments. Last year, startup Sakana AI’s AI Scientist-v2 even managed to write a paper that cleared peer review for the prestigious International Conference on Learning Representations.

    These advances have led to speculation that models are close to being able to build better versions of themselves with little human oversight—a process called recursive self-improvement. The idea underpins predictions that we may be on the verge of an intelligence explosion that could quickly lead to AI superintelligence.

    In a recent paper, researchers put the idea to the test using a new approach they call shadow evaluations. This involves taking the research question from a high-quality, unpublished machine learning paper and asking AI agents to solve the problem. The original paper’s authors then grade the results. When the team tested Claude Opus 4.8 on two papers submitted to the prestigious machine-learning conference NeurIPS 2026, the authors rejected both.

    “The papers were nowhere close to the mark when it came to being at the quality of a top AI conference,” Sayash Kapoor from Princton University, who co-led the study, told MIT Technology Review.

    Previous efforts to get AI agents to do machine learning research have often targeted problems focused on engineering, such as reproducing previous research or training smaller models against a benchmark.

    In the new experiments, the researchers challenged models with more open-ended tasks that required them to devise hypotheses, decide what evidence is needed to validate them, judge when a research direction was fruitless, and go back to the drawing board.

    One research question was whether the personality traits a language model displays can be measured and adjusted by observing and editing its weights; the other attempted to detect when a model that works with tabular data has quietly stopped being reliable.

    Be Part of the Future

    100% Free.No Spam.Unsubscribe any time.

    In each case, the AI researchers were given $3,000 of API credits, a budget for time on GPUs to run machine learning experiments, a dedicated Linux virtual machine, and unrestricted internet access. They were then given six days to produce a paper that could pass NeurIPS’ stringent peer-review criteria.

    In both cases, the models got a good start. The agents surveyed the literature effectively, came up with opening hypotheses that mirrored those of the authors, and successfully ran hundreds of experiments.

    But they quickly went off the rails. Although they could monitor their own use of time and their API and GPU budgets, they rushed through the process. One left 110 hours of unused time on the clock, and both failed to spend even 50 percent of their API budget.

    Both agents also settled on a research direction within just 10 hours and failed to change approaches despite repeated negative feedback from another AI designed to review drafts of their papers. The reviewer identified problems the human authors would also flag in the final paper, but the models simply added caveats to their findings and ploughed on. Ultimately the papers received a “strong reject” and a “reject” decision from the human reviewers based on NeurIPS grading protocol.

    The authors admit their approach has limitations. The reviewers knew AI had written the submissions, and some of the team are on record as doubting an imminent intelligence explosion. The original human-authored papers also took far longer than six days to produce and used many more GPU hours to reach their conclusions (though, as the researchers note, the models did not use their allocated budget in any case).

    Nonetheless, the results suggest that today’s models still have some way to go before they can tackle the most challenging problems in machine learning research. Until that happens, the dream of recursive self-improvement is likely to remain a distant prospect.

    Edd Gent
    Edd Gent
    Edd is a freelance science and technology writer based in Bangalore, India. His main areas of interest are engineering, computing, and biology, with a particular focus on the intersections between the three.

    A wooden judge's gavel sits on a marble table

    An ‘AI Legal Team’ Has Won Its First Case. It’s a Rare Victory for Access to Justice.

    Colorful onscreen static

    Long Foreseen, the Problem of AI Alignment Is Finally Reality. Solving It Won’t Be Easy.

    DeepMind’s Weather AI Predicts Hurricanes a Day Earlier Than Traditional Forecasting

    Shelly Fan
    Aug 17, 2026
    Future

    An ‘AI Legal Team’ Has Won Its First Case. It’s a Rare Victory for Access to Justice.

    Genevieve Grant
    Aug 27, 2026
    Future

    Long Foreseen, the Problem of AI Alignment Is Finally Reality. Solving It Won’t Be Easy.

    Liming Zhu
    Aug 20, 2026
    Artificial Intelligence

    DeepMind’s Weather AI Predicts Hurricanes a Day Earlier Than Traditional Forecasting

    Post Views: 11

    Explosion Intelligence Maybe Verge
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    Previous ArticleDebian Votes To Allow “Responsible Use Of Generative AI”
    Next Article Your Guide to Artificial Intelligence (AI) Degrees
    aitoday7
    • Website

    Related Posts

    AI Guides

    Your Guide to Artificial Intelligence (AI) Degrees

    August 29, 2026
    AI News

    Decoding cosmic signals with deep learning and Keras

    August 29, 2026
    AI News

    Looking beyond natural sequences

    August 28, 2026
    Add A Comment
    Leave A Reply Cancel Reply

    Top Posts

    Previewing the Model Hardware Standard

    August 29, 20260 Views

    Human, Beware: The Secrets We Tell AI Chatbots Aren’t Private

    August 29, 20260 Views

    NSU Technology & Entrepreneurship Summit Sept. 30

    August 29, 20260 Views
    Stay In Touch
    • Facebook
    • YouTube
    • TikTok
    • WhatsApp
    • Twitter
    • Instagram
    Latest Reviews
    Chatbots

    OpenAI bets on families as ChatGPT goes deeper into households

    aitoday7July 11, 2026
    Generative AI

    MUSIC COMMUNITY INTRODUCES NEW LABELING PROGRAM TO DISTINGUISH GENERATIVE AI IN SOUND RECORDINGS

    aitoday7July 11, 2026
    AI News

    Safe from AI: which jobs will help you thrive in the future?

    aitoday7July 11, 2026

    Subscribe to Updates

    Get the latest tech news from FooBar about tech, design and biz.

    Most Popular

    Previewing the Model Hardware Standard

    August 29, 20260 Views

    Human, Beware: The Secrets We Tell AI Chatbots Aren’t Private

    August 29, 20260 Views

    NSU Technology & Entrepreneurship Summit Sept. 30

    August 29, 20260 Views
    Our Picks

    OpenAI bets on families as ChatGPT goes deeper into households

    July 11, 2026

    MUSIC COMMUNITY INTRODUCES NEW LABELING PROGRAM TO DISTINGUISH GENERATIVE AI IN SOUND RECORDINGS

    July 11, 2026

    Safe from AI: which jobs will help you thrive in the future?

    July 11, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Facebook X (Twitter) Instagram Pinterest
    • About Us
    • Get In Touch
    • Disclaimer
    • Privacy Policy
    • Terms and Conditions
    © 2026 AIToday7. All Rights Reserved.

    Type above and press Enter to search. Press Esc to cancel.