AutoResearch released to enable AI agents to independently test and refine research ideas

News provided byEvoMap · 2 min read

SAN FRANCISCO, Sept. 1, 2026 — EvoMap, a leading open infrastructure project for AI self-evolution, has released AutoResearch, an innovative open-source system designed to enable AI agents to test and refine their own research ideas.

The launch of AutoResearch addresses a critical gap in AI research: the ability for AI models to independently generate, test, and iterate on research hypotheses. According to the EvoMap team, the traditional research verification process often relies heavily on human validation, which can be time-consuming and error-prone. AutoResearch aims to streamline this process by automating the generation and evaluation of research ideas.

At the core of AutoResearch is a multi-model system that independently generates, cross-reviews, and executes research plans. The process begins with a hypothesis, which is then converted into a structured research plan with defined metrics, success criteria, resource budgets, and evaluation procedures. Specialized agents handle each phase of the research process, from planning and implementation to experimentation, analysis, and review.

EvoMap emphasizes that the system's design prioritizes continuous learning and improvement. Research state, code, experiment logs, metrics, and decisions are preserved in a persistent workspace, allowing the system to pick up where it left off without starting from scratch. Additionally, independent and blind review mechanisms ensure that conclusions are rigorously tested before being finalized.

One of the key design choices in AutoResearch is its approach to failure. Rather than treating a partial result as the end of a research path, the system encourages further investigation. A failed experiment can lead to a revised hypothesis, and repeated experiments can refine the research direction. This iterative process was demonstrated through tests on real-world benchmarks.

For instance, AutoResearch was tested on a Django issue from SWE-bench Lite, initially scoring 2/7 on official new-feature tests. Instead of stopping, the system continued investigating the underlying problem and ultimately achieved a perfect 7/7 score. On the RSICD benchmark, an AutoResearch-generated idea improved mean Recall from 32.84 to 34.69, showcasing the system’s ability to turn AI-generated hypotheses into measurable improvements through iterative experimentation.

The release of AutoResearch represents a significant step in the evolution of AI research. It marks the transition from AI research assistants to AI4AI, where AI agents can independently test and refine their own research ideas. This capability is not limited to AI research; it has broader implications for fields such as drug discovery, materials science, and engineering, where research ideas can be evaluated through simulations or physical experiments.

EvoMap’s work on AutoResearch underscores its commitment to building a robust, open infrastructure for AI self-evolution. The company’s projects, including the Genome Evolution Protocol (GEP), EvoX Agent, and AutoResearch, demonstrate its vision of creating systems that allow AI agents to learn from experience, share validated capabilities, and improve across tasks and environments.

For more details, visit evomap.ai.

Talk to the desk

Want your company on the wire?

File your first press release free, or talk to us about a plan built for regular volume and placement.

Contact us