Research
  • Scaling laws
  • Representation learning
  • Disease monitoring

Scaling Laws Arrive in Liquid Biopsy

Few-Shot and Zero-Shot Disease Detection and Monitoring with Exai-1

Foundation models have transformed AI by learning general representations from large datasets and reusing them across tasks. We built Exai-1 to learn from cell-free RNA (cfRNA) in blood and test whether those patterns could help detect cancers with limited labeled data. Trained on a large cfRNA dataset we generated ourselves, Exai-1 uses self-supervised learning across RNA sequence, structure, and abundance to create the first multiscale cfRNA foundation model. It learned rich representations of both individual RNAs and whole blood samples, transferred across biofluids and assay chemistries, and supported few-shot learning when disease-specific data were scarce. We published these results in Nature Machine Intelligence, where Exai-1 was featured on the cover.

Since then, the evidence that it generalizes has kept growing. In a few-shot setting, the frozen Exai-1 representation enabled detection of head and neck cancer in blood collected up to four years before diagnosis. In a zero-shot setting, Exai-1 detected osteosarcoma without ever being trained on osteosarcoma samples or labels. Exai-1 shows what sample-efficient, bio-native AI can do: learn biology at scale, transfer across chemistries and cohorts, and adapt to new diseases from a handful of samples, or from none at all. It is the first model in an architecture we are now training on multimodal patient data spanning tissue, blood, and medical records.

Cover of Nature Machine Intelligence, December 2025, Volume 7, Number 12, titled Cell-free RNA profiling with a language model.
Exai-1 on the cover of Nature Machine Intelligence, December 2025 (Vol. 7, No. 12).

When biomedical AI begins to generalize

The defining achievement of modern AI is generalization. Beyond a certain scale and diversity of training, models begin to solve tasks they were never explicitly trained to solve. This transition famously transformed natural language processing. A general pretrained model could be adapted to new tasks with a small number of examples; sometimes with no examples at all. A classifier built for one disease usually needs many labeled samples. For rare cancers, those samples can take years to collect. With thousands of RNA measurements but only a few dozen cases, a model can learn quirks of the cohort instead of a reliable disease signal.

We built Exai-1 (featured on the cover of the December 2025 issue of Nature Machine Intelligence1) to ask whether the same thing could happen in human biology and in clinical settings. We started with cell-free RNA in blood—one of the richest signals in medicine, but also one of the noisiest. Cells release RNA fragments into the bloodstream, and the mix changes with cellular activity. Exai-1 learns patterns across these cfRNA measurements before it sees a new cancer cohort.

Most machine learning in medicine still follows the old paradigm: collect hundreds or thousands of labeled samples from one disease, select features, train a disease-specific classifier, and repeat the whole process for the next disease, usually wherever reimbursement makes the data collection worthwhile. That works for a few common cancers with large datasets.

Many diseases are simply too rare for that workflow. Large cohort studies may take years to assemble or remain permanently out of reach. With thousands of molecular features and only tens of cases, feature selection becomes unstable and disease-specific models overfit easily.

A foundation model should change that equation, and cfRNA gives it a rich source of biological information to learn from. Cells throughout the body continually release RNA fragments into the bloodstream through active secretion and other biological processes. The identities and abundances of those fragments reflect cellular state, creating molecular disease barcodes whose fingerprints can be measured directly in blood. A foundation model can learn the shared structure of those signals from a large, diverse training corpus, then transfer that representation to diseases with very little data of their own.

Our team was in a unique position to try this. Over several years, we had amassed the world's largest small-RNA database spanning tissue and blood. That data scale made self-supervised and semi-supervised training strategies possible: Exai-1 could first learn the structure of cfRNA directly from the data, then use targeted biological supervision to shape a representation that transfers across tasks.

The original Exai-1 paper showed that this works in ovarian cancer with only 48 labeled cases1. We have since tested the same model in other cancers. In head and neck cancer, Exai-1 enabled few-shot detection from blood collected a median of four years before diagnosis. In osteosarcoma, it detected a cancer that was completely absent from pretraining, without fine-tuning or fitting a new classifier.

Few-shot learning is the practical breakthrough because it turns a few dozen samples into a viable learning problem. Zero-shot generalization extends that sample efficiency to its limit: a new disease with no labeled examples at all.

The pretrained model stays fixed while a small classifier learns from limited labeled cases. In another test, we can use its existing cancer-detection output without training on the new disease at all. The ovarian, head and neck, and osteosarcoma studies below test both approaches.

A language model for the blood

Exai-1 is a multimodal, generative transformer trained on cell-free small RNA profiles. The overall dataset used for training Exai-1 contains ~13,000 plasma and serum samples. This translates into more than 306 billion RNA-abundance tokens. Where a language model learns that the meaning of a word depends on the words around it, Exai-1 learns that the meaning of an RNA measurement depends on the thousands of other RNAs observed with it in the same blood sample.

Each sample contains measurements from thousands of small non-coding RNAs, including our disease-annotated oncRNAs, miRNAs, tRNAs, yRNAs, and snoRNAs. Each RNA enters the model with a pretrained embedding that captures aspects of its sequence and structure. Exai-1 then combines that molecular representation with the RNA's measured abundance in blood.

The model uses self-attention to learn relationships among these RNAs and compresses the entire sample into a 32-dimensional latent representation. A generative decoder reconstructs masked measurements from that bottleneck. During pretraining, auxiliary objectives expose the model to cancer status, tissue of origin, assay version, and biofluid type.

Figure 1. The same frozen model serves both paths: a new classifier for few-shot tasks, or the pretrained cancer-status head for zero-shot detection.

The result is a representation that carries biological information while suppressing technical variation. It captures the structure of a cfRNA profile: which features move together, which patterns are associated with cancer, and which differences arise from sample collection or processing.

Exai-1 is small by current AI standards: 3.6 million parameters, trained on a single NVIDIA T4 in roughly 60 hours. It still marks the beginning of a scaling regime for biomedical AI, where more data and more compute should yield stronger representations, higher performance, and broader generalization. We're actively scaling Exai-2 now in a close collaboration with NVIDIA.

On held-out data, Exai-1 reconstructed masked cfRNA measurements with an R² of 0.89, compared with 0.57 for a dataset-average baseline. Across cancer-classification experiments, models trained on its 32-dimensional representation gained an average of 0.22 AUROC over models trained directly on the original 7,349 features.

That is the first capability: Exai-1 turns a noisy, high-dimensional blood sample into a compact representation that is easier to learn from.

Learning representations that survive distribution shift

AI models often look excellent until the input distribution changes. In blood-based assays, a deceptively simple example is the difference between plasma and serum. Both are derived from blood, but the collection process changes the observed cfRNA profile. A classifier can easily anchor on those technical differences, leaving it brittle when moved to a new biofluid, tube type, assay version, or cohort.

This source batching is pervasive in biological data. Collection site, patient population, tube type, storage, assay chemistry, and processing pipeline can each leave a signature stronger than the biology we want to measure. Models trained directly on raw features often learn the source of a sample because source is the easiest signal available.

We trained a conventional cancer classifier on plasma and evaluated it on serum collected from the same patients in Network's cfRNA dataset. In the raw feature space, AUROC fell from 0.74 on plasma to 0.56 on serum. The sample representation learned by the model failed to generalize. We then ran the same experiment using Exai-1's frozen embeddings. Performance was 0.77 on plasma and 0.74 on serum.

Figure 2. Retained signal is the clearest read: Exai-1 keeps 89% of its above-chance AUROC on serum; raw features keep 25%.

This is what useful representation learning looks like. The frozen Exai-1 latent space had already separated enough of the cancer signal from the biofluid signal for a plasma-trained classifier to generalize directly to serum.

The serum experiment provided early evidence that Exai-1 had learned a representation capable of surviving outside the exact conditions in which a downstream model was trained. By compressing source-specific variation while preserving shared biological structure, the model increased the effective signal-to-noise ratio available to the downstream classifier.

Ovarian cancer, held out of pretraining, recovered from 48 samples

Our next question was whether the representation could transfer to a disease that had been removed from pretraining. We excluded ovarian cancer entirely, pretrained Exai-1 on the remaining corpus, froze the model, and passed a downstream cohort of 48 ovarian cancer cases and matched controls through the encoder. We then trained a lightweight XGBoost classifier on the resulting embeddings.

Across 25 random seeds, training on the Exai-1 representation improved AUROC by 0.10 relative to training on the original cfRNA features. We could also sample from Exai-1's generative decoder to create synthetic profiles, which added a further 0.015 AUROC. The PCA control separated representation learning from dimensionality reduction. A matched 32-dimensional PCA produced only a small, borderline gain, while the pretrained Exai-1 representation improved AUROC by 0.101.

Figure 3. Adapted from Karimzadeh et al., Nature Machine Intelligence (2025), Fig. 3b1.

This gave us a reusable adaptation recipe:

  1. Keep the foundation model frozen.
  2. Embed a small cohort from a new cancer.
  3. Train a simple classifier on the embeddings.

The expensive representation learning has already happened. Adding a disease requires a small labeled cohort and a lightweight head. This is the source of Exai-1's sample efficiency: every new task inherits structure learned across thousands of samples, avoiding the need to rediscover that structure inside one rare-disease cohort.

That is the few-shot capability.

Head and neck cancer, detected a median of four years before diagnosis

The next test was harder in three ways: a new cancer, a new cohort, and samples collected years before clinical diagnosis.

In work with Dr. Hoang C.B. Nguyen at Massachusetts Eye and Ear, Harvard Medical School, we studied a prospective biobank of HPV-negative head and neck cancer2. The cohort contained 93 plasma samples: 26 from people who later developed oral cavity cancer and 67 controls matched for age, sex, and smoking history. The blood was collected a median of 50.2 months prior to diagnosis.

We kept Exai-1 frozen, embedded the 93 samples, and fit an XGBoost classifier on top using three-fold cross-validation. It reached an AUROC of 0.78, with 73.1% sensitivity at 77% specificity, and outperformed equivalent classifiers trained on raw cfRNA features or a matched PCA representation.

This remains few-shot learning because the downstream classifier saw labeled head and neck samples. But everything upstream of that small classifier came from a model that had never been pretrained on this cancer.

For an AI audience, the important result is the transfer. A frozen representation learned from other cancers and other contexts contained enough reusable structure to support detection in a 93-sample cohort. It did so in pre-diagnostic blood, where the biological signal is much weaker than it is at the time of treatment.

The ability to predict oral cavity cancer four years before diagnosis emerged when a small number of examples were projected into the representation Exai-1 had already learned.

Osteosarcoma: founder mode meets zero-shot AI

Sid Sijbrandij, the co-founder of GitLab, was diagnosed with osteosarcoma in 2022 after imaging revealed a six-centimeter mass extending from his T5 vertebra. He underwent surgery, radiation, and intensive chemotherapy. The cancer remained in remission for two years and then returned in late 2024.

By then, Sid had exhausted the standard treatment path and had no available clinical trial. He describes the change in his thinking as going into "Founder Mode" on his cancer4: engaging deeply with the details and actively coordinating every possible option.

He assembled a team of clinicians, scientists, and operators. They built a maximal diagnostic program spanning imaging, pathology, bulk and single-cell sequencing, immune profiling, organoids, and minimal residual disease assays from multiple providers. They generated multiple personalized treatments and pursued therapeutic hypotheses in parallel, using new measurements to decide what to try next5.

Sid's experience makes the rare-disease AI problem concrete. Osteosarcoma has too few patients and too little available molecular data to support the conventional cycle of large cohort collection, feature selection, and disease-specific model development. Working with Sid and his team, we asked whether Exai-1 could provide a far more sample-efficient starting point: could a frozen foundation model recognize osteosarcoma with zero osteosarcoma labels?

Osteosarcoma was completely absent from Exai-1's pretraining panel. Every model component remained frozen and every osteosarcoma sample remained outside model fitting. We evaluated 101 osteosarcoma plasma samples against 196 held-out controls using the existing cancer-detection output3.

The result was a zero-shot AUROC of 0.872 (against our in-house control, also held-out from Exai-1).

At an operating point set to 90% specificity, sensitivity was 75.2%. The model entered this evaluation with zero examples of osteosarcoma and zero disease-specific feature engineering. Its pretrained representation already contained enough general cancer biology to recognize the disease.

Figure 4. On osteosarcoma the comparison favors the baselines, which were trained on the cohort. Exai-1 still scores highest.

This is the AI capability we set out to build. The collaboration establishes zero-shot transfer in an initial cohort and gives rare-cancer development work a new starting point: a pretrained model that already separates cases from controls. Clinical validation remains ahead.

Results on held-out cancers
CancerLabels for fittingSetupCohortResult (AUROC)
Ovarianfew-shotKarimzadeh et al., Nat. Mach. Intell. 2025148labeled casesFrozen Exai-1 + XGBoost head
25 random seeds
Synthetic profiles: +0.015 (p = 0.07)
48 cases, matched controls
Excluded from pretraining
0.864median of 25 seeds 0.761 with raw features25 of 25 seeds improve, p = 9.1×10to the power of −10
Head and neckfew-shotNguyen et al., bioRxiv 2025226labeled casesFrozen Exai-1 + XGBoost head
3-fold cross-validation
Prospective HPV-negative biobank
26 cases, 67 matched controls
Blood drawn a median of 50 months before diagnosis
0.780cross-validated 0.620 with raw features73.1% sensitivity at 77% specificity
Osteosarcomazero-shotNetwork Bio internal analysis, with the Sijbrandij Foundation30labeled casesPretrained cancer-status head
applied as is, with no fitting
101 cases, 196 controls
Absent from pretraining, as were the controls
0.872zero-shot 0.813 raw features trained on it75.2% sensitivity at 90% specificity
  • Bars run from chance (0.5) to 1.0; the tick marks the raw-feature baseline. Ovarian values are medians over 25 random seeds.
  • The Labels for fitting column counts the labeled cases each model could learn from. Cross-validation trains on part of them in each fold.
  • Operating points and the ovarian p-value come from the cited sources. We trained and cross-validated the osteosarcoma baselines on osteosarcoma; Exai-1 was never trained on it.
Figure 5. The labeled cases used for fitting drop from 48 to none while Exai-1 stays frozen.

What comes next

Relative to its language counterparts, Exai-1 is a small model trained on one modality. Its results matter to us because of what they imply about scale. A 3.6 million parameter model, trained on cfRNA alone, learned a representation that transferred across biofluids, assay chemistries, cohorts, and cancers it had never seen. The next question is what the same architecture learns from more data and more modalities.

That is the question we are working on now. Exai joined Network Bio in June 2026, and the architecture behind Exai-1 is being trained on multimodal patient data: tissue, blood, and longitudinal medical records from a network of leading academic medical centers, across cancer, immunology, and cardiometabolic disease. Our view is that AI can now read human biology, and that data is the rate-limiting factor. Exai-1 is the first evidence for that view. The data foundation is how we test it at scale. The result of that work is Exai-2, a patient world model we are now developing.

Sources

  1. Karimzadeh, M. et al. A multimodal cell-free RNA language model for liquid biopsy applications. Nature Machine Intelligence 7, 1927–1938 (2025). DOI: 10.1038/s42256-025-01148-x PDF
  2. Nguyen, H.C.B. et al. Orphan non-coding RNAs drive tumorigenesis and enable pre-diagnostic detection in HPV-negative head and neck cancer. bioRxiv (2025). bioRxiv 2025.12.20.695703
  3. Osteosarcoma Phase 1 evaluation, Network Bio internal analysis, in collaboration with the Sijbrandij Foundation.
  4. Sijbrandij, S. I'm going Founder Mode on my cancer. sijbrandij.substack.com
  5. Hershberg, E. Going Founder Mode On Cancer. centuryofbio.com

Cite this post

Karimzadeh, M. and Goodarzi, H. (2026). Scaling Laws Arrive in Liquid Biopsy. Network Bio.