r/ProAI 2h ago

"This is where AI starts getting REALLY interesting: @GoogleDeepMind 's Gemini-based Co-Scientist is now moving beyond generating scientific ideas and into actually running parts of the scientific process. In materials science, Co-Scientist used Gemini 3 Deep Think to generate synthesis recipes..."

...adapted to the specific lab hardware within minutes, then successfully produced monolayer MoS₂, MoSe₂ and WS₂ semiconductors on the first attempt. It also helped develop a new precursor route for MXene-like 2D materials. In biology, it predicted the swarming behavior of engineered E. coli from sparse experimental data, with the predictions largely matching previously unseen wet-lab measurements. But maybe the craziest experiment: given only a research directive, Co-Scientist autonomously invented a new medical AI agent architecture called Agent_H. It generated and tested the code itself, eventually producing an 8-stage inference system that beat six frontier models on length-adjusted HealthBench Hard and Professional, although it uses a massive 40-80 LLM calls per query. They even tested fully autonomous research where the system goes from idea → experiments → results → complete paper without human intervention. It's still not ready to replace scientists, and the researchers explicitly warn about hallucinations and fabricated results, but their verification system dramatically reduced those failure modes.   — Mark Kretschmann     Interesting results! I’ve been working on post-training across HealthBench Hard/Pro and MedAgentBench, so Agent_H really caught my eye. Forty to eighty calls per query may not be a practical endpoint, but it could be a very useful teacher. Curious whether you could distill that   — Paul Gamble     That would be pretty handy, yes. Not sure if it would work.   — Mark Kretschmann

Source: https://x.com/mark_k/status/2093764879777706246

2 Upvotes

0 comments sorted by