1/ Excited to share our new work: “Accelerating Scientific Research with Gemini in the Real-World.”
We extend the AI Co-Scientist into an execution-grounded research system spanning ideation → experimentation → paper writing, and validate it across materials science, biology, and computer science. 🧵
8/ What I find most exciting is the direction this points toward: AI systems that don’t just reason about science, but can iteratively interact with experiments, code, measurements, and human experts.
7/ We evaluated these mechanisms in a double-blind study with 30 domain experts and 450 independent reviews.
The reliability modules reduced severe result hallucinations and plagiarism relative to the ablated baseline, while the safety system refused 98.7% of hazardous prompts.
6/ An important part of autonomous science is not just capability, but scientific reliability.
We therefore added mechanisms that explicitly penalize hallucination/plagiarism and verify manuscript claims against raw experimental execution logs before finalization.
5/ In computer science, we pushed autonomy further.
Given only a research directive, Co-Scientist autonomously designed and iterated an inference-time scaling architecture, Agent_H, which outperformed six frontier models on length-adjusted HealthBench Hard & Professional.
4/ In biology, Co-Scientist built a vision-based system to predict emergent E. coli swarming morphologies at unseen inducer concentrations from sparse experimental images.
The predictions showed quantitative concordance with subsequent wet-lab measurements across most evaluated morphology metrics.
3/ In materials science, Co-Scientist helped design CVD synthesis protocols tailored to a custom lab setup.
This included a safer precursor route for a 2D layered material with characteristics consistent with Ti₃C₂Tₓ MXene, as well as single-attempt growth of monolayer MoS₂, MoSe₂, and WS₂.
2/ A central question for AI-for-science is:
Can an AI system move beyond hypothesis generation to interface with real-world experiments and empirical feedback?
Here, we explore that spectrum, from human-executed wet-lab experiments to fully autonomous computational research.
Medical training begins in textbooks, but clinical competence is forged through practice. We introduce ResidencyRL: a multi-turn online RL method to train health agents in simulation, and demonstrate significant gains on top of frontier systems. https://t.co/CGw4EO7qh7
Introducing MicroVQA: A Multimodal Reasoning Benchmark for Microscopy-Based Scientific Research #CVPR2025
✅ 1k multimodal reasoning VQAs testing MLLMs for science
🧑🔬 Biology researchers manually created the questions
🤖 RefineBot: a method for fixing QA language shortcuts
🧵
🚨Large video-language models LLaVA-Video can do single-video tasks.
But can they compare videos?
Imagine you’re learning a sports skill like kicking: can an AI tell how your kick differs from an expert video?
🚀 Introducing "Video Action Differencing" (VidDiff), ICLR 2025
🧵
Biomedical datasets are often confined to specific domains, missing valuable insights from adjacent fields. To bridge this gap, we present BIOMEDICA: an open-source framework to extract and serialize PMC-OA.
📄Paper: https://t.co/zkb2m5yeal
🌐Website: https://t.co/atIKEfOpkv
Among the most impressive aspect of the Llama 3.1 release is the accompanying research paper! Close to 100 pages of deep knowledge-sharing on LLMs like we havn't seen very often recently
What a treat!
It covers everything, pretrainining data, filtering, annealing, synthetic data, scaling laws, infrastructures, parallelism, training recipees, post-training adaptation, tool-use, benchmarking, inference strategies, quantization, vision, speech, videos...
Mind-blown! Maybe the single paper you can read today to join the field of LLM from zero right to the frontier
Read it here and feel the open-science https://t.co/nANpZtiP0s