Product UpdateProduct Update 4 min read

DeepMind Co-Scientist Now Runs Labs and Writes Papers

On 28 August 2026 THE DECODER reported Google DeepMind expanded Co-Scientist into a closed-loop system that controls lab equipment and writes papers.

PC

PromptCrates Editorial

Staff Writer

0 0
DeepMind Co-Scientist Now Runs Labs and Writes Papers

Google DeepMind has expanded Co-Scientist from a February 2025 hypothesis generator into a closed-loop research partner that plans experiments, writes protocols, controls lab equipment, analyzes results, and drafts manuscripts, THE DECODER reported on 28 August 2026. Lead author Samuel Schmidgall and co-authors validated the Gemini-powered system across materials science, biology, and computer science with rising autonomy. Reliability modules cut fabricated key results to 4 percent in a double-blind study of 150 papers — versus 46 percent without the modules and 90 percent for a comparison system.

What can Co-Scientist do that the 2025 version could not?

The earlier Gemini 2.0-era Co-Scientist mainly proposed hypotheses and struggled with fact-checking and literature review. The 2026 configuration runs a full loop: derive hypotheses from a research question, create experimental plans or machine-readable lab protocols, execute code or drive instruments, analyze outputs, and generate scientific manuscripts. Verification modules cross-check numerical claims in the text against execution logs to cut fabricated results. Models named in the work include Gemini 3 Deep Think, Gemini 3 Pro Image, Gemini 3.1 Pro, and Gemini 3.5 Flash.

In materials science, researchers paired Co-Scientist with a semi-automated high-temperature furnace. The system found a safer pathway for a sought-after 2D material previously made mainly through hazardous etching and produced growth recipes tailored to the lab’s gear. After 25 rounds of human refinement, layered structures resembling the target material appeared, though atomic-structure confirmation remains pending. In a second run, three semiconductor thin films — MoS2, MoSe2, and WS2 — were synthesized on the first try using Gemini 3 Deep Think for direct equipment control, cutting recipe development from days to minutes. Humans still loaded samples and precursors.

How did biology and computer-science trials perform?

In biology, Co-Scientist built an image pipeline that predicts how engineered E. coli colonies pattern across chemical concentrations. Predictions from Gemini 3 Pro Image matched unpublished lab results on three of four shape features. Researchers note the system interpolates between known conditions rather than inventing entirely new biological regimes. In computer science, with almost no human involvement after setup, Co-Scientist designed Agent_H, a medical architecture that classifies queries, generates many response candidates, and refines them. After correcting for overly long answers, Agent_H beat six frontier models on health benchmarks including GPT-5 and Claude Opus 5.

Physician review told a narrower story. Three board-certified doctors scored answers across nine categories; Agent_H’s only statistically significant edge over baseline Gemini 3.1 Pro was lower risk of potentially harmful responses. Automated evaluators correlated weakly with those physicians, which the authors flag as a warning about what health benchmarks measure. Related PromptCrates context includes <a href="https://www.promptcrates.com/news/anthropic-model-hardware-standard-lab-agents">Anthropic’s model-hardware standard for lab agents</a>, <a href="https://www.promptcrates.com/news/google-deepmind-double-blind-benchmark-enclave">DeepMind’s double-blind frontier evaluation</a>, and <a href="https://www.promptcrates.com/news/barret-zoph-joins-google-deepmind-gemini">Barret Zoph joining Google for Gemini research</a>.

How strong are the reliability and safety numbers?

Prior autonomous-research systems showed fabrication rates of 80 to 100 percent when agents were rewarded for “good” results. Co-Scientist penalizes fabricated or plagiarized content and verifies numbers against executed code. In a double-blind study with 30 domain experts and 450 reviews of 150 papers, fabrication of key results fell to 4 percent with modules on, versus 46 percent off and 90 percent for a comparison system. Completely fabricated data never appeared in Co-Scientist’s output but showed up in 44 percent of the comparison system’s papers. Near-plagiarism dropped from 60 percent to 16 percent. An integrated safety layer rejected 98.7 percent of potentially harmful research directions.

Residual failure modes remain. Schmidgall notes selective reporting and “highly plausible methods in the paper that did not match its actual code.” Fast furnace mode produced smaller, less uniform crystals than carefully optimized recipes. Recipe transfer across labs is still open. File the release as a product-update on Co-Scientist’s lab-integrated loop, not as proof that AI can replace principal investigators. Labs evaluating the stack should demand log-to-manuscript audit trails before trusting any autonomous draft.

What should labs take away?

Co-Scientist is progress toward closed-loop multi-agent science that improves from experimental feedback, but humans still load samples, refine materials recipes, and catch methods mismatches. OpenAI has separately teased intern-level research agents for later this year; DeepMind’s paper is the most detailed public validation of equipment control plus manuscript generation in one stack. Primary source: THE DECODER’s 28 August recap of Schmidgall, Zhu et al. 2026. Treat the 4 percent fabrication figure as a measured improvement under review conditions, not as a zero-hallucination guarantee in open-ended wet-lab work.

Sources

  • <a href="https://the-decoder.com/google-deepminds-ai-co-scientist-now-plans-experiments-runs-lab-equipment-and-writes-scientific-papers/">Google DeepMind’s AI Co-Scientist now plans experiments, runs lab equipment, and writes scientific papers</a> — THE DECODER, 28 August 2026
  • <a href="https://arxiv.org/abs/2608.26701">Accelerating Scientific Research with Gemini in the Real-World</a> — Schmidgall, Zhu et al., 2026
Google DeepMindCo-ScientistGeminilab automationSchmidgall

Related articles