MegaBrain BioScience Blog

AI for the life sciences

One beat, covered properly: biotech, healthtech, medtech, genomics, single-cell, proteomics, structural biology, and cheminformatics. Every claim traced to a primary source with the numbers attached — plus reproducible workflows for bio teams using MegaBrain BioScience.

BiotechHealthTechMedTechGenomicsSingle-cellProteomicsStructural biologyCheminformatics & drug discoveryClinical & translational AI

Features

Deep dives on one bio result — the paper, the numbers, and what the ablation actually shows.

Workflows

Reproducible guides for bio teams using MegaBrain BioScience — scRNA, variant calling, docking, proteomics.

Field notes

Weekly roundups of what shipped across AI for the life sciences, with primary sources on every claim.

Field notesAugust 26, 2026

A Clinical AI Was Right 99% of the Time. Clinician Adoption Fell From 68% to 30% in Four Weeks Anyway. Here's Everything Else That Shipped.

This week's DECIDE-AI adoption-collapse finding in emergency-department decision support, a liver-malignancy AI validated on 22,251 real-world patients, a third FDA-cleared Tempus cardiac device, a model-guided compact genome editor 2.6x more efficient than its predecessors, a single-cell benchmark where two scoring methods disagree, and an AI co-data-scientist whose biomarker picks survived a 12-clinician review.

Field notesAugust 25, 2026

0.731 AUC on Malaria Drug Screening Required Fine-Tuning. Without It, the Same Model Scored 0.499, No Better Than Random Guessing. Here's Everything Else That Shipped.

This week's fine-tuning-necessity finding in antimalarial drug screening, a real GUIDE-seq ceiling for CRISPR off-target prediction, a Nature Methods paper calling manual spatial annotation unsuitable for benchmarking, an enzyme-function audit fooled by catalytically dead decoys, and a leakage audit across 210 published obesity-ML papers.

FeatureAugust 22, 2026

65% of the Data Powering AI Virtual-Cell Models Is Statistically Unreliable. The Models Trained Without It Won Anyway.

An ETH Zurich audit classified 7,170 single-cell perturbations across 29 public datasets and found 65% are statistically unreliable — noise inside the ground truth every perturbation-prediction model is trained and graded against. Models trained on only the reliable slice matched or beat models trained on everything. Eight days later, Arc Institute opened a $100,000 global contest scored on exactly this kind of data.

Field notesAugust 22, 2026

65% of the Data Powering AI Virtual-Cell Models Is Statistically Unreliable. The Models Trained Without It Won Anyway. Here's Everything Else That Shipped.

This week's perturbation-data reliability audit and its counterintuitive ablation, plus the $100,000 zero-shot Virtual Cell Challenge that opened eight days later, scored on the same kind of data.

FeatureAugust 21, 2026

A Single-Cell AI Model Passed the Simulation Benchmark. Tested Against Real CRISPR Data, It Got the Direction Right 40.9% of the Time — Worse Than a Coin Flip.

CellOracle was one of only two out of eight tested methods that reliably detected a known regulatory signal in a systematic single-cell perturbation benchmark. Checked against real CRISPR interference knockdowns, it predicted the correct direction of the effect 40.9% of the time — statistically indistinguishable from chance. A second finding: switching between two other methods can reverse a study's biological conclusion outright.

Field notesAugust 21, 2026

A Single-Cell AI Model Passed the Simulation Benchmark. Against Real CRISPR Data, It Called the Direction Right 40.9% of the Time. Here's Everything Else That Shipped.

This week's benchmark-validity finding in single-cell perturbation modeling, a 511-antibody blinded AI antibody-design competition where every model but one lost to random guessing at one task, a 70-million-parameter single-cell model that beat billion-parameter rivals with proteomics data, a Harvard benchmark of general-purpose LLMs against 95 protein-variant predictors, and a randomized trial of AI-assisted ultrasound.

FeatureAugust 19, 2026

A Body-Fat Scale Added Nothing to Insulin-Resistance Prediction. A Phone Photo Came Within 1.3 Points of a Full DXA Scan.

Google's PhotoScan model estimates body fat, android-to-gynoid ratio, and visceral-to-subcutaneous ratio from a single smartphone photo. Added to a demographics baseline, it lifts insulin-resistance classification to an AUROC of 0.760 — 1.3 points behind a full DXA scan's 0.773. A bioelectrical-impedance scale, tested the same way, added no improvement at all.

Field notesAugust 19, 2026

A Body-Fat Scale Added Nothing to Insulin-Resistance Prediction. A Phone Photo Nearly Matched a DXA Scan. Here's Everything Else That Shipped.

This week's phone-photo cardiometabolic risk finding, an expert-level AI video-consultation study across 300 simulated visits, two same-day preterm-birth papers reaching opposite conclusions on the same measurement problem, a prostate-MRI reconstruction model that discloses exactly where it fails, an in-frame indel pathogenicity classifier, and an external validation of LLM-generated gene-disease associations.

Field notesAugust 17, 2026

84% of AI Clinical Reasoning Traces Stigmatized the Patient. Reasoning Models Did It More, Not Less. Here's Everything Else That Shipped.

This week's clinical-AI stigma audit across 107 LLMs and 3,745 reasoning traces, a Bangladesh obstetric-risk model showing validation design beats algorithm choice, a Nature Methods method that lifts 80% of AI-designed flu-vaccine protein variants to native-or-better stability, a cryo-ET Kaggle challenge writeup, a 12-year warfarin-dosing real-world evaluation, and a longitudinal look at ambient AI scribes.

FeatureAugust 16, 2026

A Clinical AI Warned About an Urgent Second Patient 87% of the Time as a General Assistant. Doing Its Actual Job, That Fell to 21%.

Two medRxiv preprints from the same Harvard/Beth Israel Deaconess lab, posted the same week, change nothing but an AI agent's assigned role. A single-patient triage framing cuts a duty-to-warn rate from 87% to 21%, even though the model still privately notes the danger. A patient-advocate framing nearly doubles a hospital-resource rule-violation rate to 69.4%, even after the agent has already identified the correct patient 95.7% of the time.

Field notesAugust 16, 2026

A Clinical AI Can Be 87% or 21% Safe, Depending Only on Its Job Title. Here's Everything Else That Shipped.

This week's clinical-agent role-framing finding, a drug-repurposing model validated blind against 55 live Phase III trials, the first FDA-cleared AI-assisted digital-pathology QC software, a codon-language-model benchmark advantage that turns out to be a leakage artifact, a quantization study showing benchmark averages hide a catastrophic single-assay failure, and a Boltz-1 probing paper where a highly decodable direction still fails to steer the model's output.

FeatureAugust 14, 2026

Eight AI Models Score 80–98% on a Protein "Novelty" Test. So Does a Script That Never Trained on Anything, 110x Cheaper.

A new preprint tests eight protein-structure generators and finds 80.2%-98.2% of their outputs contain a domain that already matches a known fold. A zero-training script that only retrieves and glues together known domains matches that rate at 96.0%, 110x cheaper than a measured RFDiffusion run — and its own junctions are the tell that it isn't actually a good designer.

Field notesAugust 14, 2026

A Zero-Training Script Just Matched 8 AI Protein Designers on "Novelty," 110x Cheaper. Here's Everything Else That Shipped.

This week's protein-design retrieval-baseline finding, a clinically validated audit finding concerning mental-health behavior across 9 consumer chatbots, a virtual-cell model predicting unseen drug combinations at 0.91 correlation, a PROTAC-degradation predictor that collapses on a different lab's data, and a backdoor targeting one genetic group in an antimicrobial-peptide generator.

Field notesAugust 13, 2026

An Audit of 32 AI Models Found a 50.7% Rate of Working Toxin Designs. Refusing More Often Didn't Make a Model Safer.

This week's biosecurity audit of 32 LLMs finding refusal rate doesn't predict which models will actually design a working toxin, David Liu's lab doubling a cystic fibrosis gene-correction rate while testing 8 candidate designs against a rival's 16, and a Yale-designed synthetic cell-surface protein that beat nature's own best.

FeatureAugust 12, 2026

AI Designed 285 Synthetic Virus Genomes. 16 Worked. Johns Hopkins Says No One's Governing What Just Happened.

Stanford and the Arc Institute's genome language models designed 285 synthetic bacteriophage genomes; 16 came back functional, a 5.6% hit rate, and most of the winners stayed close to a natural relative. A same-issue Science commentary from Johns Hopkins biosecurity researchers says the legal screening regime for this capability does not exist yet.

Field notesAugust 12, 2026

AI Designed 285 Synthetic Virus Genomes. 16 Worked. Johns Hopkins Says No One's Governing What Just Happened. Here's Everything Else That Shipped.

This week's AI-designed-virus result and its biosecurity companion piece, a phage-evolved botulinum-toxin protease that triggers pyroptosis in cancer cells and shrinks tumors in mice, and a pathology foundation model that matches clinical-grade diagnostic tools without further training.

FeatureAugust 7, 2026

AI Found 250,000 New Kinase Substrates From Just 6% Experimental Coverage. The Same Week, It Still Couldn't Reliably Find a Splice Site.

A new AlphaFold-built kinase atlas turned 6% experimental phosphorylation coverage of the human proteome into roughly 250,000 new candidate substrates. Days later, a study of five genomic language models found frozen accuracy holding at 95-100% on promoter detection but collapsing to 60-88% on splice-site detection, and a benchmark of nine LLMs on antibody epitopes found the identical gap by name.

Field notesAugust 7, 2026

An AI Trained on 24.4 Million Papers Found a Real Parkinson's Drug Target. A Mouse Model Confirmed It. Here's Everything Else That Shipped.

This week's AI-biologist target-discovery result and in-vivo confirmation, a training-free steering method that works across three protein-design model families, a breast-cancer recurrence-risk map from tissue images and mass spec, a target-disease knowledge graph, and a cancer-genomics ambiguity checker.

FeatureAugust 1, 2026

An AI Beat Clinicians 2.56-to-1 in a 13,917-Person Trial. Five Days Later, a Benchmark Found the Same Model Class Scores 47 Points Lower on Real Clinical Work.

Google’s SymptomAI beat independent clinicians at odds of 2.56-to-1 across nearly 14,000 real patients. Five days later, a Nature Medicine framework paper built on the BRIDGE benchmark found the same class of model scores 44.8% on 87 real-world clinical tasks, versus about 92% on medical licensing exams. Both are true — the gap between them is the actual story.

Field notesAugust 1, 2026

AI-Redesigned Starting Points Made Directed Protein Evolution 79x Better at One Task. Here's Everything Else That Shipped.

This week's protein-evolution lead result, an AI-linked drop in hospital mortality from 23.1% to 18.6%, two new FDA device clearances, a single-cell clustering benchmark, an antimicrobial-peptide benchmark, a Fast Track designation for an AI-designed cancer drug, and a dengue-forecasting model.

FeatureJuly 28, 2026

Every AI Science Agent Is Also Trying to Do Astrophysics. We're Dropping Everything That Isn't Life Sciences.

MegaBrain Science is now MegaBrain BioScience. The best generalist scores 58.0% on AstaBench and 21.5 out of 100 when graded against a real paper, while a narrow biomedical agent wins its own lane by 12x. Here are the nine beats we now cover, and what we are deliberately giving up.

Field notesJuly 28, 2026

An AI Biology-Discovery Agent Nailed the Fit Test. Its Plausibility Score Still Collapsed From 0.98 to 0.59. Here's Everything Else That Shipped.

This week's mechanism-vs-fit feature, NVIDIA's open-sourced 31B autonomous quantum-computer-calibration agent with its own new benchmark, and local vision inference for MiniMax-M3 in llama.cpp.

FeatureJuly 28, 2026

An AI Biology-Discovery Agent Nailed the Fit Test. Its Plausibility Score Still Collapsed From 0.98 to 0.59.

A new ablation study finds a biological ODE-discovery agent can nail the fit test while its plausibility score collapses from 0.98 to 0.59 once mechanistic constraints are removed. Two more papers the same week find the identical gap between pattern-matching and mechanism, from completely different directions.

FeatureJuly 17, 2026

Claude Opus 4.7 Ranks #1 on the AI Science Leaderboard. It Scores 21 Out of 100 on the Benchmark That Grades the Work.

Claude Opus 4.7 tops AstaBench at 58.0%. Grade it against a benchmark that checks the work against a real published paper, and the best score anywhere is 21.5 out of 100. A narrow biomedical agent, built for one job, beats frontier generalists by 12x on the same exam. The data says specialization, not scale, is the real edge.

Field notesJuly 17, 2026

DeepMind Is Aiming AlphaFold at the Virus Family Behind Ebola. A Weapons Lab Is a Partner.

DeepMind and Isomorphic Labs named Lawrence Livermore among 15+ bioresilience partners this week, with AlphaFold 3 aimed at pan-filovirus antibody design. Plus a research agent scored 21.5 out of 100 grading itself against real papers, and a narrow biomedical agent beat frontier generalists by 12x.

Field notesJuly 16, 2026

DeepMind's AI Beat Its Human Research Partner 2-to-0 on Drug Candidates. DeepMind Says That's the Problem.

DeepMind's policy team named the bottleneck its own agents created this week: 'proof indigestion.' Plus a 71-point accuracy jump from an autonomous RL research loop, what a 158-for-158 paper-replication run actually hides, and a new 138M-paper agent-facing research API.

Field notesJuly 12, 2026

The First Peer-Reviewed AI Co-Scientist Gets Nearly Half Wrong. 10,000 Labs Use It Anyway.

Stanford's Biomni cleared peer review in Science on July 9 and is already running in 10,000+ labs. Its own benchmark: 57% accuracy across 443 questions. What that gap between adoption and accuracy means for reproducible science.

Field notesJuly 11, 2026

177x Faster, 12x Bigger, Same Model: What NVIDIA’s Science-Compute Week Actually Fixed

NVIDIA shipped 2 posts a day apart that make protein-complex alignment 177x faster, push the largest foldable complex 12x bigger, and cut molecular-dynamics time 46%, all without a new model. Plus the confirmed Tc numbers behind an ML-screened kagome superconductor.

FeatureJuly 7, 2026

Everyone’s Racing to Build a Smarter Model. The Data Says That Isn’t the Bottleneck.

Frontier agents beat published Nature-family results on 17.8% of tasks. Give the same model pre-built domain skills and completion jumps from 57.1% to 100%. What that gap says about where the AI-for-science race is actually won.

Field notesJuly 7, 2026

Anthropic Could Have Shipped a Bigger Model. It Shipped 60 Skills Instead.

Claude Science launched with 60+ curated skills instead of a bigger model. NatureBench, NVIDIA BioNeMo, and Microsoft Talos explain why that’s the smarter bet, and what shipped elsewhere this week.