PO.MD01.01 · 分子诊断与数据

整合转录组数据以改进多组学肿瘤起源分类

Integrating transcriptomic data to improve multi-omic tumor-of-origin classification

海报缩略图:整合转录组数据以改进多组学肿瘤起源分类
编号 13 展板 13 时间 4/19 02:00–05:00 区域 Section 1 主讲 Pranav Gadde, No Degree
分会场 AACR Project GENIE: Predictive Models and AI
查看 PDF 下载 PDF 🔒 查看 / 下载完整 PDF 需登录并开通下载套餐 · 查看套餐 / 开通 AACR 官方页面

作者与单位 Authors & Affiliations

Pranav Gadde, Naga Mudda, Pratayanch Sav, Krithik Senthilkumar, Aadarsh Sivaraman, Krithik Mudda

Illinois Mathematics and Science Academy, Aurora, IL

摘要 Abstract

中文摘要
背景:原发灶不明的癌症(CUP)约占恶性肿瘤的2-5%,由于采用经验性治疗,预后较差。现有的基因组分类器准确率仅约60-75%。AACR Project GENIE联盟和数据库提供了大规模的真实世界基因组数据,但目前尚无与AACR GENIE数据整合的转录组特征。我们假设,在TCGA数据上开发一个经过验证的多组学框架,可以创建一种将转录组数据整合到GENIE中的方法,预期结果是提高肿瘤起源分类的精度。 方法:我们通过cBioPortal检查了来自TCGA PanCancer Atlas的1,556例实体瘤队列,涵盖四种癌症类型:COAD(n=76测试)、HNSC(n=103)、LUAD(n=102)和READ(n=31)。采用SMOTE训练了两个XGBoost分类器(80/20留出划分):(1)仅基因组基线分类器(肿瘤突变负荷、MSI评分、非整倍体评分)和(2)多组学分类器(相同的基因组特征加上20,506个RNA-Seq基因)。在留出测试集(n=312)上评估性能。 结果:多组学分类器达到93.0%的准确率,而仅基因组分类器为52.2%(提高40.7%)。各类别F1评分为:COAD=0.86,HNSC=1.00,LUAD=1.00,READ=0.58。基线分类器表现接近随机(COAD F1=0.50,READ F1=0.33)。特征重要性分析也确认了已知的谱系标志物为最具预测性的特征:KRT5(9.2%重要性,HNSC鳞状标志物)、SFTPB(6.1%重要性,LUAD肺表面活性物质)、GPA33(2.9%重要性,COAD/READ肠道标志物)、CDX1(2.7%重要性,肠道转录因子)、HOXB13、NAPSA和EVX2,从而证明了有意义的生物学模式识别。 结论:与基因组方法相比,多组学整合显著提高了肿瘤起源分类。其对成熟的组织特异性生物标志物的依赖提供了对诊断至关重要的生物学有效性。READ的表现较低(F1=0.58),原因是样本数量有限以及与COAD的相似性。总体而言,该框架在其他癌症类型中显示出良好的区分能力。这一TCGA概念验证建立了一个经过验证的流程,可利用AACR Project GENIE扩展至15-20种癌症类型,有助于推进CUP的临床诊断和精准肿瘤学。
查看英文原文 English abstract
Background: Cancers of Unknown Primary (CUP) make up approximately 2-5% of malignancies with a poor prognosis due to empirical therapy. Existing genomic classifiers only have approximately 60-75% accuracy. The AACR Project GENIE consortium and database provide large-scale and real-world genomic data, but there are currently no integrated transcriptomic features with AACR GENIE data. We hypothesized that developing a validated multi-omic framework on TCGA data could create an approach for integrating transcriptomic data into GENIE, with the expected outcome of improving the precision of the tumor-of-origin classification. Method: We examined a cohort of 1,556 solid tumors from the TCGA PanCancer Atlas, via cBioPortal, across four cancer types: COAD (n=76 test), HNSC (n=103), LUAD (n=102), and READ (n=31). Two XGBoost classifiers were trained (80/20 held-out split) with SMOTE: (1) Genomic-Only baseline classifier (Tumor Mutation Burden, MSI Score, Aneuploidy Score) and (2) Multi-Omic (same genomic features plus 20,506 RNA-Seq genes). Performance was evaluated on a held-out test set (n=312). Results: The Multi-Omic classifier achieved 93.0% accuracy compared to 52.2% for the Genomic-Only classifier (40.7% improvement). Per-class F1 scores were: COAD=0.86, HNSC=1.00, LUAD=1.00, and READ=0.58. The baseline classifier demonstrated near random performance (COAD F1=0.50 and READ F1=0.33). A feature importance analysis also confirmed known lineage markers as the top predictive features, KRT5 (9.2% importance, HNSC squamous marker), SFTPB (6.1% importance, LUAD lung surfactant), GPA33 (2.9% importance, COAD/READ intestinal marker), CDX1 (2.7% importance, intestinal transcription factor), HOXB13, NAPSA, and EVX2, therefore demonstrating meaningful biological pattern recognition. Conclusions: Multi-omic integration significantly increases tumor-of-origin classification compared to genomic methods. Its reliance on established, tissue-specific biomarkers provides biological validity critical for diagnostics. READ performance was lower (F1=0.58) due to the limited number of samples and the similarity to COAD. Overall, the framework showed good discrimination across other cancer types. This TCGA proof-of-concept establishes a validated pipeline for expanding to 15-20 cancer types using the AACR Project GENIE, helping advance clinical diagnosis of CUP and precision oncology.
利益披露 Disclosure
P. Gadde, None.. N. Mudda, None.. P. Sav, None.. K. Senthilkumar, None.. A. Sivaraman, None.. K. Mudda, None.

← 返回 AACR 2026 检索