PO.BCS01.11 · 生物信息与计算
一种轻量级自监督深度学习框架用于自动检测循环肿瘤细胞和癌症相关成纤维细胞
A lightweight self-supervised deep learning framework for automated detection of circulating tumor cells and cancer-associated fibroblasts
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
循环稀有细胞(CRC),包括循环肿瘤细胞(CTC)和循环癌症相关成纤维细胞(cCAF),是有价值的液体活检生物标志物,然而由于其极度稀有和形态异质性,其检测仍具挑战性。当前的识别方法主要依赖基于荧光的成像以及由训练有素的专家进行的耗时手动评估,这限制了高通量分析、可重复性和临床实施。此外,对主观视觉判断的强烈依赖使CRC判定高度依赖操作者,引入了显著的观察者间和观察者内变异性,并使跨中心的检测标准化复杂化。为克服这些限制,我们开发了一种自监督深度学习框架,能够使用最少的标注数据并减少对荧光信号的依赖,实现对CRC的稳健且可解释的检测。我们的方法采用两阶段训练策略:首先使用对比学习在大规模白细胞(WBC)数据集上预训练一个模型,使其能够从丰富的、形态相似的细胞中学习可泛化的形态表征。下一步,使用知识蒸馏将所学知识迁移到一个轻量级的学生模型中,随后在有限的CRC数据上对其进行微调。该蒸馏过程在保持检测准确性的同时显著降低了模型复杂度,从而实现适合临床工作流程的实时推断。在我们使用27例早期乳腺癌患者样本进行的实验中,常规的基于荧光的分析手动识别了CTC和cCAF。当应用于同一数据集时,所提出的框架实现了超过90%的CRC检测灵敏度和特异性,同时以最小的计算负担运行,并与手动专家评估显示出高度一致性。与常规的基于荧光的手动标注相比,我们的方法在速度、一致性和可扩展性方面提供了实质性提升,同时消除了专家依赖性评估固有的观察者间变异性。这些结果表明,自监督表征学习与知识蒸馏相结合,为自动化CRC检测提供了一种实用且临床可行的策略,在早期癌症诊断、纵向疾病监测和治疗反应评估方面具有潜在应用。
查看英文原文 English abstract
Circulating rare cells (CRCs), including circulating tumor cells (CTCs) and circulating cancer-associated fibroblasts (cCAFs), serve as valuable liquid biopsy biomarkers, yet their detection remains challenging due to extreme rarity and morphological heterogeneity. Current identification methods predominantly rely on fluorescence-based imaging and manual, time-consuming assessments by trained experts, which limit high-throughput analysis, reproducibility, and clinical implementation. Moreover, the strong dependence on subjective visual judgment makes CRC calling highly operator-dependent, introducing substantial inter- and intra-observer variability and complicating assay standardization across centers. To overcome these constraints, we developed a self-supervised deep learning framework that enables robust and interpretable detection of CRCs using minimal labeled data and with reduced dependence on fluorescence signals. Our approach employs a two-stage training strategy in which a model is first pretrained on large-scale white blood cell (WBC) datasets using contrastive learning, allowing it to learn generalizable morphological representations from abundant, morphologically similar cells. In the next step, knowledge distillation is used to transfer this learned knowledge into a lightweight student model that is subsequently fine-tuned on limited CRC data. This distillation process significantly reduces model complexity while preserving detection accuracy, thereby enabling real-time inference that is suitable for clinical workflows. In our experiments using samples from 27 patients with early-stage breast cancer, conventional fluorescence-based analyses manually identified both CTCs and cCAFs. When applied to the same dataset, the proposed framework achieved CRC detection sensitivity and specificity exceeding 90% while operating with minimal computational burden and showed high concordance with manual expert assessment. Compared with conventional fluorescence-based manual annotation, our approach offers substantial gains in speed, consistency, and scalability, while eliminating inter-observer variability inherent to expert-dependent assessments. These results suggest that self-supervised representation learning combined with knowledge distillation provides a practical and clinically viable strategy for automated CRC detection, with potential applications in early cancer diagnosis, longitudinal disease monitoring, and treatment response assessment.
利益披露 Disclosure
H. Woo, None..
S. Park, None.
J. Lee,
CTCELLS Inc. g., Board of Directors, non-salaried role).
I. Moon, None.
M. S. Kim,
CTCELLS Inc. g., Board of Directors, non-salaried role).