PO.BCS01.07 · 生物信息与计算
DIANNE:组织学差异属性的无分割定位
DIANNE: Segmentation-free localization of histology differential attributes
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
病理医生指导下对组织学图像中的区分为组织健康状况提供了洞见,推动着对疾病机制理解和临床决策的进步。利用人工智能的数字病理学在从组织学和空间组学图像中获取洞见方面日益重要。为训练计算模型,当前的数字病理学方法依赖于前期的人工标注,这既耗时又难以扩展。这种预标注过程也不适合研究新的空间行为,因为在这些情况下标注具有挑战性且数据需求不明确。
为解决这些问题,我们提出了DIANNE(差异图像标注环境),这是一种数字病理学方法,基于对深度学习成像特征的训练时正类混合增强(PCMA)来快速计算空间差异图像属性。DIANNE能够在标准工作站上于数秒内跨全切片图像(H&E或基于抗体的多重成像,如CODEX)定位差异属性,从而支持用于探索性研究的交互式应用。预测模型可以在交互式瓦片标注更改后实时重新训练,阐明组织切片内及切片间的重要生物学属性。
我们首先在静态组织学图像上测试DIANNE以进行肉瘤肿瘤检测。我们证明,仅用60张先前未标注的肿瘤切片即可开发出分类器,达到高精确率(0.86 ± 0.17)和召回率(0.73 ± 0.26),以及合理的特异性(0.58 ± 0.3)和低假阳性(0.0000 ± 0.0001)。这些指标与计算量更大的基于像素的方法相当,例如[Segmenter. Strudel R et al. arXiv 2105.05633, 2021](值为0.86 ± 0.15、0.86 ± 0.22、0.53 ± 0.28、0.0062 ± 0.0129),后者还需要区域标注、GPU和6小时的计算。此外,DIANNE具有交互式探索组织切片的独特能力。我们展示了从人类胰腺、胎盘和肾脏组织切片中检测生物结构,包括识别染色和成像伪影。对于肾小球检测,我们基于对2884个瓦片中仅40个阳性和60个阴性瓦片的交互式标注,展示了瓦片级召回率、特异性和假阳性分别为0.95、0.95、0.0002。359个真实肾小球中仅漏检9个,其中7个为部分裁剪。
DIANNE提供了多种空间推理工作流(交互式/静态、H&E/分子),以基于Jupyter widget的工具包实现。DIANNE仅需图像尺度的标注(而非像素级标注)即可高效训练基于基础模型的分类器。DIANNE快速、交互式的系统允许从静态和探索性输入进行实时训练,为发现空间数据集中的新行为提供了实用系统。
查看英文原文 English abstract
Pathologist-guided distinctions within histology images provide insights into tissue health, driving advances in understanding of disease mechanisms and clinical decision making. Digital pathology leveraging artificial intelligence is increasingly important to derive insights from histology and spatial omic images. To train computational models, current digital pathology methods rely on upfront manual annotations, which are time-consuming and difficult to scale. This pre-annotation process is also poorly suited for investigating novel spatial behaviors, where annotation is challenging and data requirements are unclear.
To address these issues, we present DIANNE (Differential Image Annotator Environment), a digital pathology approach for rapid computation of spatial differential image attributes based on train-time Positive Class Mixup Augmentation (PCMA) of deep learning imaging features. DIANNE enables localization of differential attributes across whole slide images (H&E or antibody-based multiplex imaging, e.g. CODEX) in seconds on standard workstations, enabling interactive applications for exploratory investigation. Predictive models can be re-trained in real-time after interactive tile annotation changes, clarifying important biological attributes within and across tissue slides.
We first test DIANNE on static histology images for sarcoma tumor detection. We demonstrate that classifiers can be developed with as few as 60 previously unannotated tumor slides to achieve high precision (0.86 ± 0.17) and recall (0.73 ± 0.26), as well as reasonable specificity (0.58 ± 0.3) and low false positives (0.0000 ± 0.0001). These metrics are comparable to more compute intensive pixel-based methods, e.g. [Segmenter. Strudel R et al. arXiv 2105.05633, 2021] (values 0.86 ± 0.15, 0.86 ± 0.22, 0.53 ± 0.28, 0.0062 ± 0.0129) which also required regional annotations, GPU and 6 hours of compute. Moreover, DIANNE has unique capabilities for interactive exploration of tissue slides. We show biological structure detection from human pancreatic, placenta and kidney tissue slides, including identification of staining and imaging artifacts. For kidney glomeruli detection we demonstrate tile-level recall, specificity, and false positives of 0.95, 0.95, 0.0002, respectively, based on interactive annotation of only 40 positive and 60 negative patches out of 2884 total patches. Just 9 out of 359 true glomeruli are missed, 7 of which were partially cropped.
DIANNE provides multiple spatial inference workflows (interactive/static, H&E/molecular) implemented in a Jupyter widget-based toolkit. DIANNE enables efficient training of foundation model-based classifiers with only image scale annotation, rather than at the pixel-level. DIANNE's rapid, interactive system allows real-time training from static and exploratory input, providing a practical system for discovery of novel behaviors in spatial datasets.
利益披露 Disclosure
S. Domanskyi, None..
J. C. Rubinstein, None..
T. Sheridan, None..
A. Thiesen, None..
J. Alcoforado Diniz, None..
R. Ramasamy, None..
D. S. Baker, None..
R. Sheldon, None..
Q. Wu, None..
G. Kuchel, None..
P. Robson, None..
J. H. Chuang, None.