A pathology model trained on 2.3 million slides matches cleared products
Researchers have published PRISM2, a pathology foundation model that operates on the whole slide rather than the image patch. It was trained on 2.3 million whole-slide images paired with 14 million question-and-answer pairs derived from roughly 700,000 clinical pathology reports — using the diagnostic language pathologists already write as the supervision signal, instead of image self-supervision alone.
Asked yes-or-no questions with no additional training, PRISM2 matched the balanced accuracy of Paige Prostate and Paige Breast and outperformed Paige BLN — commercial products validated for clinical diagnostic use, with Paige Prostate holding FDA De Novo authorization in the US. The comparison ran on those products' own test datasets: 2,947 prostate cases, 1,691 breast and 753 breast lymph node. Earlier slide-level models, PRISM and TITAN, did not reach clinical-product performance even with linear probing. PRISM2 emits two embeddings — a general one and a diagnostic one taken from the hidden state of a 4-billion-parameter language model.
Diagnostic pathology is a bottleneck made of scarce trained human hours. A generalist model that reaches product-grade detection by being prompted, rather than retrained per task and per tissue, is the mechanism by which expert reading gets cheap enough to be routine instead of rationed.
The caveats belong in the room. The study is retrospective, on deidentified Memorial Sloan Kettering data licensed to Paige.AI, and both PRISM2 and its underlying tile model were trained only on slides scanned at MSK — the authors say robustness warrants further study. Biomarker gains were modest, and generating a full diagnostic report remains hard.
Source: Nature Medicine
MANY MINDED