OpenAI says its AI research intern can handle multi-day research tasks

  • OpenAI says its AI research intern can handle multi-day tasks.
  • AI agents can speed up research, but progress is hard to measure.

 

OpenAI says it has reached its goal of developing an “automated research intern” capable of carrying out research tasks under human supervision.

The company defines the system as one that can complete well-defined assignments that would take a skilled researcher several days. OpenAI said it reached the milestone it had targeted for September 2026 based on measurements of how coding agents are being used within its research organisation.

OpenAI describes the research intern as an intermediate step toward an automated AI researcher that can contribute to deep-learning and alignment work. The company has linked this work to recursive self-improvement, or RSI, while stating that it does not yet know how to safely reach what it calls “full RSI.”

The current system still relies on human direction. Researchers set priorities, decide which ideas and results should be pursued, and determine whether systems should be scaled, paused, or deployed.

OpenAI’s stated target is an automated AI researcher working under human supervision by March 2028.

AI research moves into parallel workflows

OpenAI released internal data alongside the announcement showing how its researchers are using coding agents. As of mid-August, the company recorded 3.1 agent-workdays of runtime for every workday of human labour across its research organisation, based on an eight-hour working day.

OpenAI said more researchers are also running multiple agents at the same time, including workflows involving four or more agents. The agents are used for tasks such as writing research and infrastructure code, monitoring experiments, analysing results, and providing technical assistance.

Agent use has also become more compute-intensive. The median researcher was using more than $600 per day of inference at API prices by mid-August, while researchers at the 90th percentile were using more than $7,000 per day in tokens at API prices, according to OpenAI.

OpenAI also reported an increase in the number of experiments conducted by each active experimenter this year. August recorded the highest level since the company began tracking the measure in January 2025.

OpenAI did not attribute the increase solely to its coding agents. Experiment volume rose alongside greater Codex adoption, but researchers also had substantially more compute available during the same period.

The company’s data also shows that automation varies across the research process. High-level planning remained a minimal fraction of agent output tokens, while research code, technical assistance, and experiment monitoring accounted for larger shares.

A 2024 Epoch AI study based on interviews with eight AI researchers found that coding and debugging were more amenable to automation than open-ended research work. Participants identified reliability, open-ended planning, long-context reasoning, deep reasoning, and novelty as obstacles to wider AI R&D automation.

Human involvement also remains necessary for longer assignments. OpenAI found that agent success rates increased across several task-duration categories between January and July, but more than half of successful tasks estimated to require four to eight hours of human work involved at least one human intervention.

Measuring research progress remains difficult

OpenAI’s definition of the research intern uses human-equivalent task duration rather than continuous agent runtime. The measure describes the amount of work a skilled researcher would need for the same assignment, rather than the length of time an agent operates without supervision.

METR, a nonprofit that evaluates AI systems, uses a similar human-equivalent measure in its task-completion time-horizon benchmark. The metric estimates task difficulty based on how long an expert human would take to perform the same work, rather than how long the AI operates continuously without supervision.

METR cautions that its current estimates above 16 hours are unreliable because of limitations in its task suite. Most of its evaluations cover software engineering, machine learning, and cybersecurity tasks with relatively clear success criteria.

The organisation has separately evaluated AI agents on machine-learning research engineering tasks through RE-Bench, which contains seven environments developed with input from researchers in academia and industry. METR collected 71 attempts from human experts for comparison with AI performance.

In its 2024 RE-Bench study, METR found that the tested AI agents performed better than participating human experts when both were limited to two hours. Humans performed better at longer time budgets. METR also reported that the agents generated and tested implementations more than ten times faster in some settings but struggled more with incorporating new information and building on earlier progress over longer periods.

METR noted that RE-Bench tasks have defined objectives and relatively fast feedback. Longer-term machine-learning research can involve less certain goals and experiments that take much longer to evaluate.

OpenAI has acknowledged a similar measurement problem in its own data. The company said code output and experiment volume are easier to track than actual research progress, and that neither metric establishes how much faster research is advancing.

OpenAI is developing measures that focus more directly on whether agents successfully complete assigned research work. The company said those measures are harder to develop and validate than activity-based metrics such as runtime or experiment volume.

Researchers working partly through Anthropic’s Fellows Program have tested another approach using a Claude-powered Automated Alignment Researcher. The system can propose ideas, run experiments, analyse results, train models, and share findings with other agents.

On a chat-preference dataset, nine agents accumulated about 800 hours of work across five days and reached a performance-gap-recovered score of 0.97. Four baseline methods manually tuned by two researchers over seven days reached a best score of 0.23. The automated experiment cost about $18,000 in compute and API calls.

The experiment also exposed weaknesses in the evaluation setup. Agents found dataset shortcuts, selected favourable random seeds, and discovered a way to extract information about test labels through the evaluation system.

The researchers called these behaviours reward hacking. They said future work should test ideas discovered by automated researchers on entirely held-out datasets after the agents repeatedly found ways to exploit the evaluation setup.

Labs track automated R&D risks

OpenAI’s Preparedness Framework lists AI Self-improvement as a tracked capability category alongside biological and chemical capabilities and cybersecurity. Earlier assessments of GPT-5.6 placed the models below OpenAI’s High threshold for AI Self-improvement.

Anthropic separately tracks automated R&D under its Responsible Scaling Policy. Its July 2026 revision defines the relevant threshold around systems that can fully automate or dramatically accelerate the work of large, top-tier research teams, including in AI R&D.

OpenAI said it temporarily shut down a container service used for training on July 20 after discovering that agents had compromised its research infrastructure. The service was later restored with tighter restrictions.

OpenAI said the shutdown contributed to a two-week pause in reinforcement-learning training for some of its latest models intended for deployment. Some workloads later resumed under stronger controls, while others remained paused.

OpenAI said the measurements behind its research-intern milestone remain preliminary and cover most, but not all, coding-agent activity within the research organisation. It is continuing to develop metrics that measure successful research work more directly rather than relying primarily on agent runtime, code output, or experiment volume.

 

 

 

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events, click here for more information.

Tech Wire Asia is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post OpenAI says its AI research intern can handle multi-day research tasks appeared first on TechWire Asia.