AI Can Publish at Top AI Conferences Now!
Today we are introducing Apex Research, and the results are hard to overstate.
Over the past several months, our team has been building an autonomous research system capable of conducting end-to-end machine learning research with minimal human intervention.
To find out how good it really is, we refused to hide behind internal benchmarks. We took it straight to one of the toughest arenas in the field.
We submitted AI-generated manuscripts with only limited non-expert human intervention during the process to ACL Rolling Review, the peer-review engine behind the world's top AI conferences, under the exact same double-blind process used for human researchers. Reviewers evaluated the papers without ever knowing they were written by AI.
The results exceeded our expectations.
Several papers achieved review scores above historical acceptance thresholds, while overall performance approached that of human-authored submissions. In accordance with our experimental protocol, all submissions will be voluntarily withdrawn before final publication.
Beyond the technical results, we hope this work encourages broader discussion across the research community.
As AI systems become increasingly capable of generating original scientific work, important questions remain:
How should AI contributions be disclosed?
How should authorship standards evolve?
How should the research community evaluate scientific contributions in the age of AI?
Today we are releasing our full technical report, including our methodology, evaluation protocol, quantitative results, representative reviewer comments, and a discussion of limitations.
We look forward to engaging with researchers, conference organizers, and the broader community on these questions.
Technical Report: https://t.co/b6EAKhzO8e
Follow our page to stay up to date with future product releases, technical reports, and research breakthroughs.
Interested in experiencing our products? Join the waitlist for early access:
https://t.co/lNJABFjwhs.
We also invite you to join the conversation, what do you think these results mean for the future of scientific research? Share your thoughts in the comments.
π Posters in Hall A:
β’ MARS β Modular Agent with Reflective Search (#905)
β’ PaperBanana β Automating Academic Illustration (#2100)
In Seoul for ICML? Come say hi π
At #ICML2026 in Seoul on July 6, our Google Cloud AI Research team is presenting our work on autonomous research agents. I'll present ScientistOne β an end-to-end AI scientist that takes an idea all the way to a verifiable paper. A talk, a workshop, and 2 posters π§΅π
π οΈ Expo Workshop β "Agentic Forecasting & Multi-Agent Ecosystems" Jul 6, 4β7pm KST Β· Hall D2 Forecasting, agent security, and multi-agent orchestration, with the CAIR team. https://t.co/3JcV2W6oTf
Can you trust an AI-written research paper? We checked 75 of them.
The answer is: NOT YET.
So we built ScientistOne β every claim traces to evidence. The only system to lead on all integrity checks, while matching SOTA performance.
Check out papers: https://t.co/Alxoqlrxdt
[Results]
Integrity: 0 hallucinated refs, 12/12 score verification, 14/15 method-code alignment, ~98% claim provenance rate.
Performance: On par with top discovery agents on ADRS, Gold medals on MLE-Bench, and near-SOTA on Parameter Golf.
[Our solution - ScientistOne]
It maintains evidence chains throughout the pipeline.
Problem Investigator reads full-text PDFs.
Discovery agent runs a parallel explore-exploit search tree.
Claim Verifier checks every claim against its source before the final paper is produced.
[Problem] We built CoE Audit β 4 integrity checks applied to 5 systems.
Every baseline had systematic failures: hallucinated references up to 21%, score reproduction as low as 42%
No existing evaluation checks whether AI-generated papers are actually grounded in evidence.
[Results]
Integrity: 0 hallucinated refs, 12/12 score verification, 14/15 method-code alignment, ~98% claim provenance rate.
Performance: On par with top discovery agents on ADRS, Gold medals on MLE-Bench, and near-SOTA on Parameter Golf.
[Our solution]
ScientistOne maintains evidence chains throughout the pipeline.
A Problem Investigator reads full-text PDFs.
Discovery agent runs a parallel explore-exploit search tree.
Claim Verifier checks every claim against its source before the final paper is produced.
[Problem] We built CoE Audit β 4 integrity checks applied to 5 systems.
Every baseline had systematic failures: hallucinated references up to 21%, score reproduction as low as 42%
No existing evaluation checks whether AI-generated papers are actually grounded in evidence.
π’ Deadline update for the ACM CAIS workshop AI Discovery in the Wild: submissions are now due May 7 AoE.
π€ Encouraging submissions on LLM/agent systems, infrastructure, optimization, evaluation, and real-world deployment challenges.
Concurrent submissions are welcome.
CFP π
π’ Deadline update for the ACM CAIS workshop AI Discovery in the Wild: submissions are now due May 7 AoE.
π€ Encouraging submissions on LLM/agent systems, infrastructure, optimization, evaluation, and real-world deployment challenges.
Concurrent submissions are welcome.
CFP π
Very excited to co-organize the "AI Agents for Discovery in the Wild" workshop at CAIS 2026, May 26, San Jose!
Bridging the gap between AI agents and real-world problemsπ
Non-archival! NeurIPS dual-submissions welcome!
ποΈDeadline: May 4th
πWebsite: https://t.co/oAg0Sp59sI