LBPO.BCS01 · 生物信息与计算 · Late-Breaking
H-optimus-1:用于计算组织病理学的基础模型
H-optimus-1: A foundation model for computational histopathology
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
背景:
计算病理学已成功地将人工智能(AI)方法应用于多种场景,包括治疗反应、分子生物标志物和预后的预测。在这些方法中,基础模型(FM)因其能够同时应对多样化的下游用例而成为一种尤为有前景的解决方案。在本研究中,我们介绍H-optimus-1,这是一个组织学基础模型,在广泛的关键下游任务上实现了最先进的性能,包括生物标志物预测、突变状态分类和空间基因表达预测。
方法:
H-optimus-1是一个具有11亿参数的Vision Transformer(ViT),采用自监督学习进行训练。该模型在一个大规模专有数据集上进行了预训练,该数据集包含来自超过800,000名患者的100多万张全切片图像(WSI)切片。该数据集涵盖50多种器官;切片由3种扫描仪类型在超过4,000家临床中心进行数字化。
结果:
H-optimus-1在13项下游任务上进行了评估,涵盖切片级和图块级的15个数据集,包括HEST基准(Jaume等,2024),该基准评估模型在九种不同器官中从组织学图像预测空间基因表达的能力。在与现有开源和专有基础模型的对比中,H-optimus-1在这些任务中始终取得最高的平均性能。性能以切片级分类任务的AUROC、HEST上与基因表达的Pearson相关系数以及图块级分类任务的top-1准确率来衡量。
结论:
凭借一个大规模、高度多样化的预训练数据集,H-optimus-1取得了最先进的结果,并对从转移灶识别到突变和基因表达预测等关键下游任务展现出高度的泛化能力。
查看英文原文 English abstract
Background:
Computational pathology has successfully incorporated artificial intelligence (AI) methodologies for various applications, including the predictions of therapeutic response, molecular biomarkers, and prognosis. Among these methodologies, foundation models (FMs) have arisen as a particularly promising solution, due to their ability to tackle simultaneously a diverse set of downstream use-cases. In this work, we introduce H-optimus-1, a histology foundation model that achieves state-of-the-art performance on a broad range of key downstream tasks, including biomarker prediction, mutation status classification, and spatial gene expression prediction.
Methods:
H-optimus-1 is a 1.1 billion parameter Vision Transformer (ViT) trained with self-supervised learning. This model was pre-trained on an extensive proprietary dataset consisting of over 1 million whole-slide images (WSI) slides from more than 800,000 patients. This dataset covers over 50 organs; slides were digitized with 3 scanner types across over 4,000 clinical centers.
Results:
H-optimus-1 was evaluated on 13 downstream tasks encompassing 15 datasets at both the slide level and tile level, including the HEST benchmark (Jaume et al., 2024), which assesses a model's ability to predict spatial gene expression from histology images in nine different organs. Benchmarked against existing open-source and proprietary foundation models, H-optimus-1 consistently achieved the highest average performance across these tasks. Performance was measured as AUROC on slide-level classification tasks, Pearson correlation to gene expression on HEST, and top-1 accuracy on tile-level classification tasks.
Conclusions:
Leveraging a large, highly-diverse pre-training dataset, H-optimus-1 achieves state-of-the-art results and high generalizability to key downstream tasks, ranging from metastasis identification to mutation and gene expression prediction.
利益披露 Disclosure
M. Scalbert,
Bioptimus Employment.
C. Saillard,
Bioptimus Employment.
Owkin Stock Option.
T. Peeters,
Bioptimus Employment.
L. Gonzalez,
Bioptimus Employment.
D. Valter,
Bioptimus Employment.
F. Llinares-López,
Bioptimus Employment.
Google Stock.
Z. E. Mariet,
Bioptimus Employment.
Google Stock.
R. Jenatton,
Bioptimus Employment.