In partnership with

THE BRIEFING

What comes to mind when you think of Anthropic? For Derya Unutmaz, it is biosecurity fearmongering, regulatory capture and restrictions on scientific research.

That is where we begin in the first official episode of The Eval, BAIO’s new free video series.

Most often, The Eval will be a podcast where my friend Derya and I talk about all things AI × bio. Sometimes we’ll bring in other guests, and occasionally it may take a different form altogether.

In the episode, Derya pulls apart Dario Amodei’s argument against frontier open-weight models. We discuss Anthropic’s control over scientific access, whether closed models may create the concentration of power the company says it fears - and why Derya thinks the real threat may be authoritarian control over AI itself.

Oh, and if you have ideas for what we should talk about, or guests we should invite - feel free to reply to this email. Maybe you are the guest.

Alright, in today’s issue we have Oncoformer looking for cancer in routine clinical data, Atlas Discovery trying to predict which patients will respond to drugs before a trial, and six lessons from a proposed test for medical AI superintelligence.

Let’s dive in.

How owning AI deployment expands your career

Across product, ops, and CX teams, a new kind of role is taking shape: the person responsible for making AI actually work, day to day. In this roundtable, three people living this shift share what it's really like: Simone Santiago Broad (Yoco), Yelva Espinoza (Zumba Fitness), and Fin's Dave Lynch. You'll hear how they carved out these roles, what the job looks like across industries, the skills they'd hire for, and the challenges they're tackling right now.

Watch the full conversation on demand.

THE AI WAR ON CANCER
Oncoformer looks for cancer in the data hospitals already have

Credit: Gemini

Cancer may leave a slow trail through routine medical records before anyone knows to look for a tumor.

Researchers have built Oncoformer, a multimodal transformer model that follows changes in ordinary laboratory tests, vital signs and - when available - chest X-rays across repeated hospital visits. t was developed on longitudinal health records from 2.81 million people treated at two hospitals in Wenzhou, China, then tested on an independent cohort of 862,247 people from four other Chinese hospitals. The findings are published in Cell.

The researchers specifically tested whether Oncoformer could identify people who would be diagnosed with cancer up to a year later. To stop the model from seeing clues generated during the diagnostic work-up, they excluded all data recorded during the final 90 days before diagnosis. They then compared people who received a cancer diagnosis within the tested period with age- and sex-matched people who remained cancer-free. Given one person from each group, Oncoformer assigned the future cancer patient the higher risk score about 87% of the time. It performed better when it had more recorded clinical visits and a longer medical history to learn from, suggesting that it was detecting gradual physiological changes rather than one obvious abnormal result.

The researchers also tested Oncoformer prospectively in 3,025 asymptomatic people already considered at high risk for colorectal cancer. Everyone received a model score and then underwent colonoscopy, which found 60 cancers. Oncoformer caught 54 of them, but also raised false alarms for about 390 cancer-free people. So it was good at finding cancers, but much less convincing as a practical screening tool: only about one in eight people it flagged actually had cancer. If used to decide who should receive colonoscopy, that threshold would send many cancer-free people for an invasive follow-up.

The same model was also used to estimate tumour stage, predict treatment response and separate patients by recurrence risk. Those analyses were retrospective, and the authors have not shown that acting on any of the predictions improves care. Oncoformer was developed entirely on data from Chinese hospitals, where more than 99% of participants were Asian. Testing in UK Biobank suggests that some of what the model learned can transfer to a predominantly White European population, but broader prospective validation is still needed.

BAIO has covered several cancer AIs that extract more information from material hospitals already possess: GigaTIME inferred expensive molecular imaging from routine pathology slides, while another model estimated a $3,500 breast cancer genomic test from the same kind of slide. Oncoformer applies that approach earlier in the patient journey. Instead of extracting more information from a scan or biopsy after cancer is already suspected, it asks whether years of routine lab results, vital signs and chest X-rays can identify who should receive the scan, biopsy or molecular test in the first place.

Why it matters
 

Oncoformer could turn tests and images already collected during routine care into a continuously updated triage layer - pointing to who should receive the colonoscopy, scan or molecular test next. The prospective pilot makes that idea plausible. Whether it is useful will depend on how many cancer-free people it sends for unnecessary follow-up.

This space could put you in front of a highly engaged audience of scientists, founders, investors and operators tracking the AI × biology frontier.

If this space was yours, what would you want them to understand, try, read, join or build with you?

BAIO sponsorships can be tailored around what you’re trying to achieve. Get in touch here.

CLINICAL TRIALS
Atlas Discovery launches to predict human drug response before trials

The Atlas co-founders: Christian Gensbigler, Shaamil Karim, and Sanjukta Bhattacharya. Credit: Atlas Discovery/X

Clinical trials are still where drug developers learn the answer that matters most: does this treatment work in people, and for whom?

Atlas Discovery, a three-person YC startup launching from stealth, thinks AI could change how they are designed. The company is building foundation models intended to predict how individual patients will respond to drugs before treatment begins.

Atlas says its models learn broad biological patterns from abundant preclinical data, then use much scarcer clinical data to predict human response. The goal is to help drug developers choose which medicines to advance, find the patients most likely to benefit and potentially revisit drugs that failed in the wrong population.

Its first public case study revisits UNIFI, a Phase III trial of ustekinumab for moderate-to-severe ulcerative colitis. Before treatment, 358 participants had colon biopsies taken and the gene activity in their tissue measured. Eight weeks later, 56 had responded to the drug and 302 had not.

Atlas passed the pretreatment measurements through a biological foundation model, then trained a second model to distinguish responders from non-responders. Given one responder and one non-responder, it assigned the responder the higher score 76% of the time.

The company then asked what such a predictor might change in a future trial. Under its statistical assumptions, enrolling more people predicted to respond could reduce the relevant comparison from 320 to 91 participants in each arm while retaining the same statistical power - 458 fewer people overall.

There are several limitations here. The model was trained and tested retrospectively within the same trial, with no independent study showing that it can identify responders in advance. It also saw only patients who received ustekinumab. So while it can identify who improved after treatment, it cannot yet tell whether the drug caused that improvement or whether some patients would have improved anyway.

We have recently covered two other attempts to change how clinical trials are designed. Warpspeed tries to predict whether a trial’s design and assumed treatment effect will hold up. Formation Bio uses patient data to rescue stalled drugs and find the patients most likely to benefit. Atlas goes after the same bottleneck from another angle: use biology measured before treatment to predict who should enter the next trial.

Why it matters
 

Clinical trials are often treated as an unavoidable law of drug development. In reality, they are blunt statistical instruments - but also the best method we currently have for learning whether a drug works, for whom and by how much. AI could eventually let companies enter trials with sharper hypotheses and run smaller, better-targeted studies. Atlas is nowhere near proving that yet, but it joins Warpspeed and Formation Bio in trying to redesign the trial rather than merely endure it.

6 THINGS I LEARNED
A proposed superintelligence test for medical AI

Credit: Gemini

A large group of clinicians and AI researchers - including people from OpenAI, Anthropic, Google, Microsoft, Meta and Amazon - has published a Nature Medicine Comment asking how we will know when medical AI becomes superintelligent.

Here are my 6 takeaways from the paper:

1. They take medical superintelligence seriously

The authors point to randomized studies where language models have outperformed expert clinicians on clinical reasoning, alongside the rapid improvement of AI agents that can carry out longer tasks.

They’re not arguing that medical superintelligence has arrived. But the possibility has become credible enough that medicine needs to define it before companies define it for us.

2. Beating the average doctor would not count

A system should have to outperform the best clinicians - or teams of specialists - across a wide range of meaningful clinical tasks.

Medicine is often collective. Difficult cases involve consultations, tumour boards and repeated reassessment, not one doctor answering one question alone.

3. The wrong benchmark can manufacture a superhuman result

The paper highlights an autonomous prescription system promoted with 99.2% agreement between AI and clinicians. But that score came from urgent-care telehealth cases, not the chronic prescription renewals the system would actually handle.

The reverse can happen too: a narrow multiple-choice test can make a useful system look worse than it is.

4. Real medicine happens over time

Most medical benchmarks ask for one diagnosis, test or treatment. Actual care may require searching years of records, separating current information from outdated plans, ordering tests, revising conclusions and coordinating with nurses, pharmacists and specialists.

A system that excels at isolated diagnostic questions may still fail at managing a patient whose condition, test results and treatment plan change over six months.

5. Humans may be bad guardians of a superior medical AI

Keeping a doctor in the loop sounds reassuring. But the authors point out that people are unreliable supervisors of automated systems, especially when reviewing large numbers of decisions over long periods.

And if the AI eventually becomes better than the human reviewer, checking every individual answer stops being a meaningful safety test. The evidence would have to shift toward outcomes: did patients benefit, were they harmed, what did care cost, and did AI-supported treatment outperform ordinary care in randomized trials?

6. They have already started building the test

The authors have begun turning their proposal into the Medical AI Superintelligence Test, or MAST. Its public preview tests diagnostic and management reasoning, safety, medical imaging and agentic clinical tasks, including work inside electronic health records.

THE EDGE

You probably already use AI for individual research tasks. Reactorfield wants to help you turn those isolated prompts into workflows that run repeatedly around your own scientific work.

The four-week program is for academic and industry scientists, engineers, founders and research teams. Participants choose a real part of their own work to automate, then build it through weekly sprints with one-to-one support, a reusable library of agent skills and prompts, and sessions with people from frontier AI labs and science-AI companies. Selected teams can also receive hands-on help connecting agents directly to workflows in their own labs.

Reactorfield expects serious build time each week, and the goal is to leave with a working research workflow rather than another certificate. AI expertise is not required. Applications close August 16, and the program runs online from August 31 to September 28, with some in-person components and graduation events in San Francisco and Boston.

ON OUR RADAR

Until next time,
Peter at BAIO

Keep Reading