PO.BCS02.02 · 生物信息与计算
自监督学习在血小板RNA中定义了一个通用的恶性肿瘤生物标志物
Self-supervised learning defines a universal biomarker of malignancy in platelet RNA
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
背景:通过液体活检早期检测癌症对改善患者预后具有变革性潜力,然而其临床转化受限于机器学习模型泛化能力差的问题——这些模型往往过拟合于队列特异性的技术或生物学假象,并在独立数据集中失效。肿瘤教育的血小板(TEP)构成了一个丰富但尚未充分探索的系统性癌症信号来源:其RNA图谱整合了肿瘤-宿主相互作用,可能编码具有生物学普适性的疾病信息。依赖标注数据的传统监督方法尤其容易过拟合,使得TEP转录组图谱的大部分未被利用。本研究中,我们探究在无癌症标签下训练的自监督学习(SSL)能否从TEP转录组中提取出稳健且具有生物学可解释性、可跨队列泛化的癌症相关表征。
方法:我们将一个在2230万份转录组上预训练的SSL框架应用于2,134份TEP RNA-seq样本,这些样本涵盖八种癌症类型、来自四个独立队列并在多个测序中心产生。主要输出——单一的SSL衍生特征——在发现队列和外部验证队列中评估其泛癌种检测能力。其性能与在相同数据上训练的传统监督随机森林分类器进行比较。生物学可解释性通过转录本水平关联和通路富集分析进行评估。
结果:SSL特征在发现队列中实现了强劲的泛癌种性能(AUC 0.903),并在外部队列中稳定泛化(如胶质母细胞瘤AUC 0.785;非小细胞肺癌AUC 0.803)。直接比较中,监督分类器表现出有限的跨队列可迁移性(分别为AUC 0.553和0.711)。在各癌症类型中,SSL方法相较监督学习AUC中位提升0.23。在与筛查相关的99.9%特异性阈值下,SSL特征实现了47.9%的中位敏感性。在一个独立的结直肠癌队列中,它以AUC 0.819检测I期疾病,并在99.9%特异性下保持可测量的敏感性(24.1%)。生物学分析表明,SSL特征在所有队列中可重复地与参与上皮-间质转化和凝血通路的血小板转录本相关(FDR<1×10⁻⁶),包括COL6A3、FLNA、PF4和SPARC。
结论:从大规模转录组数据中学习的自监督表征能够从TEP中提取出可重复的、具有生物学可解释性的信号,该信号可跨队列和癌症类型泛化。该框架为开发用于早期癌症检测的稳健液体活检生物标志物提供了一个技术上和临床上均可扩展的策略。
查看英文原文 English abstract
Background: Early detection of cancer via liquid biopsy has transformative potential for patient outcomes, yet its clinical translation is limited by the poor generalizability of machine learning models, which often overfit to cohort-specific technical or biological artifacts and fail in independent datasets. Tumor-educated platelets (TEPs) constitute a rich but underexplored source of systemic cancer signals: their RNA profiles integrate tumor-host interactions and may encode biologically generalizable disease information. Conventional supervised approaches, which depend on labeled data, are particularly prone to overfitting, leaving much of the TEP transcriptomic landscape untapped. Here, we investigate whether self-supervised learning (SSL), trained without cancer labels, can extract a robust and biologically interpretable cancer-associated representation from TEP transcriptomes that generalizes across cohorts.
Methods: An SSL framework pretrained on 22.3 million transcriptomes was applied to 2,134 TEP RNA-seq samples spanning eight cancer types across four independent cohorts generated at multiple sequencing centers. The primary output-a single SSL-derived feature-was evaluated for pan-cancer detection in discovery and external validation cohorts. Performance was compared with a conventional supervised Random Forest classifier trained on the same data. Biological interpretability was assessed using transcript-level associations and pathway enrichment analysis.
Results: The SSL feature achieved strong pan-cancer performance in the discovery cohort (AUC 0.903) and consistently generalized across external cohorts (e.g., glioblastoma AUC 0.785; non-small cell lung cancer AUC 0.803). In direct comparison, the supervised classifier demonstrated limited cross-cohort transferability (AUC 0.553 and 0.711, respectively). Across cancer types, the SSL approach yielded a median 0.23 improvement in AUC over supervised learning. At a screening-relevant threshold of 99.9% specificity, the SSL feature achieved a median sensitivity of 47.9%. In an independent colorectal cancer cohort, it detected Stage I disease with an AUC of 0.819 and retained measurable sensitivity (24.1%) at 99.9% specificity. Biological analysis indicated that the SSL feature was reproducibly associated with platelet transcripts involved in epithelial-mesenchymal transition and coagulation pathways (FDR < 1×10⁻⁶), including COL6A3, FLNA, PF4, and SPARC, across all cohorts.
Conclusions: A self-supervised representation learned from large-scale transcriptomic data can extract a reproducible, biologically interpretable signal from TEPs that generalizes across cohorts and cancer types. This framework offers a technically and clinically scalable strategy for developing robust liquid biopsy biomarkers for early cancer detection.
利益披露 Disclosure
H. Shen, None..
Y. Bi, None..
J. Liu, None..
F. Shi, None..
Y. Yang, None..
M. Yang, None..
Y. Li, None..
K. Chen, None..
X. Li, None.