Sponsored by

THE BRIEFING

If this were a general tech newsletter, today’s Google story would mostly be about who is leaving and what it means for Gemini.

But here at BAIO, we’re more interested in destinations than departures.

Four former Google superstars have now formed a company around automating science. And the biggest star of them all, Demis Hassabis, is giving up day-to-day control of DeepMind to spend more time on AGI, science and Isomorphic Labs.

Scientific discovery is becoming one of the problems some of AI’s most accomplished people want to spend the next part of their careers trying to solve.

In this issue we also have two new blueprints for autonomous labs, Jura Bio’s claimed scaling law for biological AI, Insilico’s new drug-discovery benchmark, Goodfire taking variant interpretation from DNA to proteins, and a new trick for Boltz-2.

Let’s dive in.

The Surprising Kindle Hack 15 Million Readers Already Know About

Every day, the ebook editions of huge New York Times bestsellers are discounted to just $2. While most are unaware, readers with the inside scoop are filling their libraries with these deals.

Readers are finding these savings through a free service called BookBub.

It sends you deals on bestselling ebooks – starting at $2. No subscriptions. No monthly fees. Just books you actually want to read, at prices that don’t make you do a double-take.

You can expect:
• Personalized picks based on your reading taste
• New deals every day to find new favorites
• Flexibility to buy from all major ebook retailers

Given a single ebook can cost $15 or more – and avid readers can spend over $100 on ebooks a year – these savings add up fast.

BookBub’s deals are only available for a limited time. Sign up to see what’s on sale today.

SHOULD I STAY OR SHOULD I GO
What is going on at Google DeepMind?

Demis Hassabis, pictured in a video released by Google DeepMind earlier this year. Credit: Google DeepMind/YouTube

“He doesn't care about chatbots. He wants to use AI to cure cancer.”

That is how one “longtime Googler” described Demis Hassabis to Business Insider after Google announced one of the biggest changes in DeepMind’s history.

Hassabis is handing over day-to-day control of Google DeepMind, the company he co-founded in 2010, and moving into what Google calls a “new strategic role” as Chair of Google DeepMind and Chief Scientist of Alphabet.

Koray Kavukcuoglu, who has been DeepMind’s CTO and is Google’s Chief AI Architect, takes over the operational job as Senior Vice President of Google DeepMind, reporting directly to Sundar Pichai.

Hassabis remains closely involved, but his own list of priorities looks rather different.

In a message published on Google’s blog he says he wants “time and space to focus on the big picture.” He will work with Pichai on strategic and global AGI issues, and lean further into Isomorphic Labs, the Alphabet drug-discovery company he also leads.

There is a perfectly reasonable interpretation of all this. Running the organization building Gemini for hundreds of millions of users is a big operational job. Hassabis is a Nobel-winning scientist who has spent much of his career trying to build AGI and use it to solve scientific problems. Google may simply be moving him toward the work he most wants to do.

Business Insider, which spoke to six current and former Googlers, reports something close to that. Semafor separately cites “two people familiar with Hassabis’ thinking” that he ”had been drifting away from the day-to-day responsibilities of running the company’s Gemini AI models and its consumer AI strategy […] increasingly shifting them to Koray Kavukcuoglu.”

But this is also a delicate moment at Google DeepMind.

Just a week ago, BAIO reported that the dedicated AlphaFold team had been dispersed. Some researchers moved to other scientific projects or Isomorphic Labs; others have left. AlphaFold co-creator John Jumper went to Anthropic earlier this summer.

And on the same day as Hassabis’s new role was announced, Google revealed that longtime Chief Scientist of Google DeepMind and Google Research Jeff Dean and Google Senior Fellow Sanjay Ghemawat were leaving to start Discovery Loop. They are joined by Google DeepMind Vice President Oriol Vinyals and longtime Google AI researcher Quoc Le in building a new company dedicated to automating scientific discovery (more on that below).

Which brings us to the spicier version of what may have been going on behind the scenes.

Pathfounders, run by former longtime TechCrunch editor Mike Butcher, reports that “industry sources” say Hassabis actually wanted to leave alongside Dean. According to its sources, Google management feared that losing both men at once would cause its share price to collapse, and persuaded Hassabis to move into the chair role instead. Alphabet shares fell about 4% on the day.

That account remains uncorroborated, however. Google’s version is that Pichai and Hassabis had been discussing a role focused on AGI for some time. We have no evidence that Hassabis is preparing to leave.

Intriguingly, Hassabis’ new role fits either account. If Google genuinely believes he is best used as its long-horizon AGI and science strategist, this is the kind of job it might create. If Google was trying to retain a founder who no longer wanted to run its increasingly product-focused AI operation, it might create much the same job.

Why it matters
 

This looks like a good thing. Hassabis has never seemed particularly interested in spending his time on the NanoBananas of the world. He has spent years inside Google balancing his scientific ambitions against the demands of running one of the world’s most important AI organizations - work that now falls to someone else. Whether Hassabis stays at Google for the long haul or eventually leaves, one thing seems safe to bet on: he will spend as much of his time as possible using AI to accelerate science, not least the life sciences.

Poll
 

This space could put you in front of a highly engaged audience of scientists, founders, investors and operators tracking the AI × biology frontier.

If this space was yours, what would you want them to understand, try, read, join or build with you?

BAIO sponsorships can be tailored around what you’re trying to achieve. Get in touch here.

THE TRAVELING WILBURYS OF SCIENCE
Four Google legends bet on automating discovery

From left: Oriol Vinyals, Sanjay Ghemawat, Jeff Dean, and Quoc Le. Credit: Discovery Loop

New startup Discovery Loop has no disclosed funding amount or valuation. It has no public biology program, wet lab, product or scientific result.

Venture capitalists love to say that the founders matter more than the idea. It can be a bit of a cliché. In this case, it’s not. As Vinod Khosla told Wired, “with this team, I wouldn’t need to know what they were doing before I backed them.”

So before we get to what Discovery Loop wants to build, let’s look at who is building it.

Jeff Dean and Sanjay Ghemawat joined Google in its early years and helped build some of the computing infrastructure that made the company possible. Dean later co-founded Google Brain and became Chief Scientist of Google DeepMind and Google Research.

Quoc Le was a founding member of Google Brain, and helped pioneer sequence-to-sequence learning and automated machine-learning research. Oriol Vinyals was a Google DeepMind vice president and technical lead on Gemini, with major contributions to sequence-to-sequence models and AlphaStar.

Between the four of them, you are looking at people who helped build both the infrastructure modern AI runs on and several generations of the AI itself.

“Our relative advantage isn't just our technical ability; it is the unprecedented scale of the systems we have previously built. We possess true full-stack depth that spans chips, hardware infrastructure, software infrastructure, ML models, and products”, the startup writes.

And now they want to apply that talent to the process of discovery.

The idea behind Discovery Loop is to automate the repeated cycle researchers use to make progress: come up with something worth trying, build and run the experiment, measure what happened, then use the result to decide what to try next. The company wants AI to run many such loops in parallel rather than waiting for humans to move through them sequentially.

If you’ve been following BAIO for a while, you know by now that Discovery Loop is far from the only player trying to close this loop. Just in the past five months we’ve told you about:

☑️ OpenAI connecting GPT-5 to Ginkgo’s automated lab and running 36,000 experiments.

☑️ Medra building a 38,000-square-foot autonomous laboratory around robots and closed-loop experimentation.

☑️ Lila Sciences’ “AI Science Factories” where models can call software, instruments, robots and people.

☑️ FutureHouse’s Robin proposing biology experiments, had humans run them, then analyzing the results and deciding what to investigate next.

Discovery Loop is starting closer to home. Before turning its system on biology or materials, it wants to automate machine-learning research itself - then use those capabilities to improve its own technology. That puts it adjacent to the recursive-self-improvement work underway at OpenAI, Anthropic and other frontier labs, where increasingly capable AI systems help build the next generation of AI. Discovery Loop’s stated ambition is broader: build an experimental engine for any scientific or engineering problem with a measurable outcome.

The New York Times reports the new company is backed by seed funding from Radical Ventures, Khosla Ventures and several Silicon Valley investors, including Google’s parent company, Alphabet.

Why it matters
 

Google desperately tried to keep these people. Dean told Wired that Sundar Pichai spent several meetings trying to persuade the four founders to stay, but they left anyway. It’s a big blow for Google. For science, it looks like a huge win. Four people who helped build Google and modern AI have decided that automating discovery is the problem worth spending the next part of their careers on.

HARNESSING THE AGENTIC LAB
What would a truly autonomous lab actually require?

Credit: Gemini

Above, we told you about Discovery Loop’s plan to automate scientific discovery. One of its first orders of business should probably be reading two new papers from researchers at Stanford, Princeton, Carnegie Mellon, Genentech, Medra and elsewhere.

Discovery Loop is starting with machine-learning research and hopes to eventually take the same idea into fields such as medicine. These papers begin closer to the physical laboratory and ask: what would actually have to change before we could seriously call a lab autonomous?

They contain plenty of mid 2020’s vocabulary - agentic laboratories, world models, physical AI. But underneath the buzzwords are two fairly sober conclusions.

First, we are not where the rhetoric sometimes suggests. Today’s systems can automate parts of an experiment or close relatively narrow loops, but broadly autonomous biological laboratories do not exist.

Second, the authors have a surprisingly concrete picture of what is missing.

Biology is an especially difficult place to build this. There are roughly 20,000 human genes, hundreds of thousands of disease-associated variants and tens of thousands of cell types and states, creating combinations no laboratory could simply test one by one. Not to mention that the laboratory itself is a complex environment.

Models, databases, protocols, robots and instruments are still largely separate pieces. Humans connect them: translating an idea into a protocol, moving information between systems, noticing when something looks wrong, adapting when an experiment fails and deciding what to try next. This manual way of coordinating experiments does not scale to biology’s enormous and complex search spaces.

The first paper, led by Mengdi Wang and Le Cong, argues that an agentic lab needs a layer that ties those pieces together. The authors call it an “agentic harnessing layer.” The idea is to keep the scientist’s goal, the AI’s plan, the available tools, what the instruments are doing, what has already happened and how certain the system is about all of it connected as the experiment unfolds.

For that to work, however, the AI needs something else: an up-to-date picture of the laboratory itself.

The paper calls this a laboratory world model. It would keep track of things such as samples, reagents, instruments, protocols, previous results and human interventions - using that information both to understand the current state of the lab and to predict what might happen if the system tries something next.

The second paper, led by Jian Ma and Genentech’s Aviv Regev, makes clear why that’s important.

Imagine a robot goes through all the programmed motions for pipetting a sample. The software records the step as completed. But the liquid was unusually viscous, partially clogged the tip and too little actually reached the well. The robot did what it was told; the experiment did not get what it needed.

The authors call this the “reality gap”: the difference between the experiment the computer thinks happened and what physically happened at the bench.

Closing it means instruments have to tell the system much more about their state. Cameras, pressure readings, calibration data, sample tracking and other signals could reveal those clogged tips, contamination, swapped samples, degraded reagents or drifting measurements before corrupted data gets fed back into the AI and influences the next experiment.

“Done responsibly, these systems have the potential to make scientific discovery more programmable, reproducible, adaptive, and scalable while enabling scientists to focus on higher-level scientific reasoning and discovery”, Le Cong writes on X.

Both borrow the levels familiar from self-driving cars. Current systems mostly sit toward the lower and middle rungs. The Regev/Ma paper says systems such as GPT-5/Ginkgo come closer to Level 4 within narrow, standardized workflows, but that fully demonstrated domain-level autonomy remains an open frontier.

The Wang/Cong paper goes further at the top of its ladder. Its Level 5 laboratory would learn from accumulated experiments, failures and experience, then use that to improve its own ability to do science - explicitly bringing recursive self-improvement into the autonomous-lab vision.

Which rather neatly closes our own little loop, bringing us back to the challenge Discovery Loop has set for itself.

Why it matters
 

These papers suggest the autonomous-lab idea has matured beyond a vague ambition. Researchers can now describe, in fairly concrete engineering terms, what is missing between today’s AI agents, robots and automated instruments and a laboratory that can genuinely run scientific discovery loops on its own. The destination is becoming specific enough to design toward - and to measure progress against.

AD

Stop making AI decisions in the dark.

Leadership is asking: where is AI delivering value for us and where is it creating risk? Right now, most teams have no idea.

With Harmonic Security’s Usage Explorer, you get a complete picture of how your organization actually uses AI, automatically categorized into custom use cases with complete tool-level granularity.

THE TRAINING DATA PROBLEM
Jura says it found a scaling law for biological AI

Credit: Jura

Modern AI has a useful property biology has mostly lacked: give a model more training data and its performance often improves in a predictable way. Jura Bio says it has now produced that kind of scaling inside the laboratory.

“We believe it's the first rigorously characterized data scaling law in biological AI,” writes Jura founder and CEO, Elizabeth Wood, on X.

BAIO covered Jura in March, when the company reported manufacturing roughly 10 quadrillion designed antibody sequences and screening 209 million of them against 100 targets.

Now Jura has trained transformer models on progressively larger slices of that experimental data. Across training sets spanning nearly three orders of magnitude in size, prediction error fell smoothly along a power-law curve - the kind of predictable improvement that has helped drive scaling in mainstream AI.

More data also improved the model’s ability to predict antibodies that would bind targets it had never encountered during training. But Jura also gave it a more difficult assignment: could the model tell which target an antibody was most likely to bind when several were possible? Performance improved with more data, but the company says the model still struggled with that specificity problem.

As promising as it sounds, there are also important limits to be aware of. This is company-authored evidence rather than a peer-reviewed paper, and the scaling law has been shown in one particular antibody-target system. We do not yet know whether the same relationship holds across other biological problems or independent laboratories.

Why it matters
 

AI for biology may not have to depend on whatever training data nature and decades of research happen to provide. If Jura’s result generalizes, labs could deliberately generate the examples AI needs - designing large numbers of biological candidates, testing them against chosen targets, measuring the outcomes and using those measurements as training data. That would make training data something researchers can manufacture, not just collect.

BENCHMARKS
Insilico opens its drug-discovery exams to outside AI

Credit: Insilico Medicine

Insilico Medicine has opened a new way for outside AI models to sit its drug-discovery exams.

The company calls it Drug Discovery and Development (DDD) Benchmark as a Service. Organizations can connect a model through a standard API; Insilico runs the tests, grades the outputs against expert reference answers and returns a score report.

DDD has two parts. Drug Discovery Foundations contains more than 300 evaluations spanning disease biology, molecular properties, retrosynthesis, structure-based design and clinical development. Some use proprietary data, while others use public datasets that Insilico says it has cleaned to reduce the risk that test questions appeared in model training.

The second suite, Drug Candidate Essentials, is closer to a simulated drug program. It asks models to make a sequence of decisions from finding a promising molecule through choosing a preclinical candidate, with reference answers anchored in Insilico’s own drug programs.

The setup is designed partly around a persistent benchmark problem: models can score artificially well if test questions have leaked into their training data. Insilico tries to limit that in two ways - by using proprietary test sets the models are unlikely to have seen, and by decontaminating the public datasets it includes.

The proprietary material creates a different problem, however. Outsiders cannot inspect those test sets in the same way they can a public benchmark, while Insilico is simultaneously an AI drug developer, the benchmark designer and the grader.

DDD wasn’t Insilico’s only announcement this week. In late July, we told you about Target Z, its new non-opioid pain program. Now comes Target Y: ISM9077, an AI-designed inhibitor aimed at another undisclosed target, this time in eye disease. Insilico says the molecule showed efficacy in preclinical models of dry age-related macular degeneration, uveitis and dry eye, and could potentially be developed as either an oral drug or an eye drop. It is the company’s 32nd nominated preclinical candidate since 2021.

Why it matters
 

To quote CEO Alex Zhavoronkov: ”As we advance toward pharmaceutical superintelligence, measuring genuine capability, not memorized answers, is how we ensure AI delivers the highest-quality, differentiated medicines and helps extend healthy, productive longevity for people everywhere.”

SNEAK PEEK
Goodfire takes variant interpretation from DNA to protein

Credit: Goodfire

In April, BAIO covered EVEE, which looked inside Arc Institute’s Evo 2 DNA model to predict whether genetic variants were harmful and generate hypotheses about why.

Now Goodfire has released MAPS - the Mechanistic Atlas of Protein Sequences. It applies a similar idea to ESM-C, a 6-billion-parameter protein model, and provides predictions for 2.1 million missense variants. A missense variant changes one amino acid in a protein.

MAPS looks at how a mutation changes ESM-C’s internal representation of the protein. Goodfire then uses lightweight probes to predict whether the mutation is harmful and which properties of the protein appear to have been disrupted.

One example is a PAX3 mutation called R56L. PAX3 is a protein involved in development. MAPS predicts that the mutation disrupts the part of PAX3 that binds to DNA. Earlier laboratory work found that R56L does indeed weaken PAX3’s DNA-binding function. Goodfire researcher Ryo Yamamoto says EVEE, looking at the same mutation from the DNA side, did not surface that mechanism.

Why it matters
 

If MAPS-type explanations hold up, mechanistic interpretability could turn a variant score into something more useful: a hypothesis telling scientists what to test next.

THE EDGE

Boltz-Perturb is a free modification for Boltz-2 that tries to make the model explore more ways a small molecule could fit into a protein’s binding pocket. Instead of retraining Boltz-2, it adds controlled noise while the model is making a prediction, nudging it toward alternative binding poses. Across a 57-target benchmark, the authors report that combining Boltz-Perturb’s different sampling strategies found successful poses for 21 targets, compared with 11 for ordinary Boltz-2.

ON OUR RADAR

Until next time,
Peter at BAIO