PO.BCS02.06 · 生物信息与计算

血统-转录组关联:机器学习预测乳腺癌的化疗反应

The ancestry-transcriptome link: Machine learning predicts chemotherapy response in breast cancer

海报缩略图:血统-转录组关联:机器学习预测乳腺癌的化疗反应
编号 4206 展板 2 时间 4/21 09:00–12:00 区域 Section 5 主讲 H Guevara-Nieto, PhD
分会场 Machine Learning Approaches for Cancer Prediction
查看 PDF 下载 PDF 🔒 查看 / 下载完整 PDF 需登录并开通下载套餐 · 查看套餐 / 开通 AACR 官方页面

作者与单位 Authors & Affiliations

Michelle Guevara-Nieto1, María J. López-Munevar2, Carlos Orozco-Castaño3, Rafael Parra-Medina4, Laura Fejerman5, Valentina Zavala6, Jone Garai7, Jovanny Zabaleta8, Alba L. Combita-Rojas9, Liliana López-Kleine2

1Pathology Department, Universidad Nacional de Colombia, Bogota, Colombia,2Universidad Nacional de Colombia, Bogota, Colombia,3Instituto Nacional de Cancerología, Colombia, Bogotá, Colombia,4National Cancer Institute, Bogotá, Colombia,5UC Davis Comprehensive Cancer Center, Davis, CA,6Department of, University of California San Francisco, Santiago, Chile,7Stanley S. Scott Cancer Center, Louisiana State University Health Science Center, New Orleans, LA,8LSU Health New Orleans, New Orleans, LA,9Instituto Nacional de Cancerologia, Bogota, Colombia

摘要 Abstract

中文摘要
背景:乳腺癌对新辅助化疗(NAC)的耐药性在拉丁美洲仍是一项重大挑战,该地区有限的基因组代表性限制了精准肿瘤学的进展。理解遗传血统如何与转录组特征相互作用,可能揭示治疗反应的人群特异性预测因子。我们使用机器学习模型在哥伦比亚乳腺癌患者中研究了血统-转录组关联。 方法:我们分析了58名接受NAC治疗的局部晚期乳腺癌女性(29名有反应者,29名无反应者),涵盖五种分子亚型。RNA-seq鉴定出339个差异表达基因(DEG);在方差稳定化标准化后,保留了变异度最高的前10% DEG(n=34)。预测因子包括临床变量(肿瘤大小、TNM分期、T分期、N分期、分级、临床分期、治疗方案、年龄、BMI、绝经)、遗传血统比例(美洲印第安人-AMR、非洲-AFR、欧洲-EUR)和34个DEG。递归特征消除未能改善模型性能;因此,纳入了所有变量。训练了随机森林(500棵树)和XGBoost模型,并通过交叉验证进行超参数优化。 结果:XGBoost取得了最高性能(AUC = 0.90),使用的学习率为0.05、深度为12、子采样/列采样为90%。在各模型中,T分期、年龄和美洲印第安人血统基于增益、覆盖率和分裂频率一致地成为顶级预测因子。在转录组变量中,CACNA1D、CLEC3A、TFF1和TTK显示出最强的预测贡献。通过参数变化和重采样策略确认了模型的稳健性。 结论:将血统和转录组特征进行机器学习整合,能够准确预测哥伦比亚乳腺癌患者的NAC反应。美洲印第安人血统,连同关键临床变量和可重现的基因表达特征,影响了预测性能,凸显了人群特异性因素在治疗耐药性中的重要性。这一血统-转录组框架为推进代表性不足的拉丁美洲人群的精准肿瘤学提供了一种可扩展的、数据驱动的方法。
查看英文原文 English abstract
Background: Breast cancer resistance to neoadjuvant chemotherapy (NAC) remains a major challenge in Latin America, where limited genomic representation restricts precision-oncology advances. Understanding how genetic ancestry interacts with transcriptomic features may uncover population-specific predictors of treatment response. We investigated the ancestry-transcriptome link using machine-learning models in Colombian breast cancer patients. Methods: We analyzed 58 women with locally advanced breast cancer treated with NAC (29 responders, 29 non-responders) across five molecular subtypes. RNA-seq identified 339 differentially expressed genes (DEGs); the top 10% most variable DEGs (n=34) were retained following variance-stabilizing normalization. Predictors included clinical variables (tumor size, TNM stage, T-stage, N-stage, grade, clinical stage, treatment regimen, age, BMI, menopause), genetic ancestry fractions (Amerindian-AMR, African-AFR, European-EUR), and 34 DEGs. Recursive Feature Elimination did not improve model performance; therefore, all variables were included. Random Forest (500 trees) and XGBoost models were trained, with hyperparameter optimization via cross-validation. Results: XGBoost achieved the highest performance (AUC = 0.90) using a learning rate of 0.05, depth 12, and 90% subsample/colsample. Across models, T-stage, age, and Amerindian ancestry consistently emerged as top predictors based on gain, coverage, and split frequency. Among transcriptomic variables, CACNA1D, CLEC3A, TFF1, and TTK showed strongest predictive contribution. Model robustness was confirmed through parameter variation and resampling strategies. Conclusions: Machine-learning integration of ancestry and transcriptomic features accurately predicts NAC response in Colombian breast cancer patients. Amerindian ancestry, alongside key clinical variables and reproducible gene-expression signatures, influenced prediction performance, underscoring the importance of population-specific factors in treatment resistance. This ancestry-transcriptome framework provides a scalable, data-driven approach for advancing precision oncology in underrepresented Latin American populations.
利益披露 Disclosure
M. Guevara-Nieto, None.. M. J. López-Munevar, None.. C. Orozco-Castaño, None.. R. Parra-Medina, None.. L. Fejerman, None.. J. Garai, None.. J. Zabaleta, None.. A. L. Combita-Rojas, None.. L. López-Kleine, None.

← 返回 AACR 2026 检索