PO.BCS01.16 · 生物信息与计算

一种利用基因组特征对乳腺癌受体亚型进行分类的机器学习方法

A machine learning approach to classify breast cancer receptor subtype using genomic features

海报缩略图:一种利用基因组特征对乳腺癌受体亚型进行分类的机器学习方法
编号 2724 展板 17 时间 4/20 02:00–05:00 区域 Section 2 主讲 Sandro Satta, PhD
分会场 Integration of Clinical and Research Data
查看 PDF 下载 PDF 🔒 查看 / 下载完整 PDF 需登录并开通下载套餐 · 查看套餐 / 开通 AACR 官方页面

作者与单位 Authors & Affiliations

Sandro Satta, Philip Miller, Samuel Rivero-Hinojosa, Ekaterina Kalashnikova, Angel Rodriguez, Minetta C. Liu

Natera, Austin, TX

摘要 Abstract

中文摘要
乳腺癌患者的风险分层、治疗方案和预后目前依赖于对受体亚型的准确判定,即通过雌激素受体(ER)和孕激素受体(PR)的免疫组化(IHC)以及HER2表达的评估(IHC和/或通过原位杂交检测基因扩增)来确定。虽然基于IHC的亚型分型检测具有信息价值,但它们需要高质量的组织样本,且技术检测易受固定假象、抗体染色性能差异、半定量和主观结果判读的影响。在样本质量下降的情况下,基于IHC的亚型评估可能与基于基因表达的分类不一致,此时可能需要替代方法。本研究旨在开发一种机器学习分类器,能够利用基因组特征预测乳腺癌受体亚型,而无需依赖免疫组化或基因表达数据。 本研究纳入了19,559例原发性乳腺癌患者,通过Natera的专有真实世界数据库识别,并与临床理赔数据库相关联。激素受体(HR)和HER2亚型根据患者治疗编码确定。我们通过整合19,820个基因的体细胞突变,构建了一个生物学指导的特征集,使用来自Signatera™检测流程的全外显子测序(WES)数据。每个突变根据变异类别(SNV、插入、缺失)、超类(SNP/INDEL)、预测影响(VEP注释影响:MODIFIER至HIGH)和功能后果(如移码、终止获得、错义、同义)被赋予一个复合突变评分(mutation_score,范围1-12)。采用分层75/25训练-测试划分和超参数优化训练了一个随机森林分类器。 模型在训练队列中14,669例患者的特征上进行训练。在4,890例患者的测试队列中,该模型与通过用药理赔数据推断的HR/HER2状态达到80.3%的总体一致性,在四个主要亚型中表现均衡。各亚型指标为:HR+/HER2-,模型精确率0.935、召回率0.911、F1评分0.923;HR-/HER2+,精确率0.714、召回率0.753、F1评分0.783;HR+/HER2+,精确率0.748、召回率0.734、F1评分0.741;最后,TNBC亚型,精确率0.730、召回率0.816、F1评分0.770。 总体而言,该基因组分类器可准确地将乳腺癌分类为四种主要受体亚型之一。在针对临床报告的HR/HER2状态进行确切验证后,该分类器可用于指导对缺乏完整临床注释的去标识化基因组数据集的分析。
查看英文原文 English abstract
Risk stratification, treatment course, and prognosis for patients with breast cancer presently rely upon the accurate determination of receptor subtype, ascertained through immunohistochemistry (IHC) for estrogen receptor (ER) and progesterone receptor (PR), and evaluation of HER2 expression (IHC and/or gene amplification via in situ hybridization). While IHC-based subtyping assays are informative, they require high-quality tissue samples and the technical assays can be susceptible to fixation artifacts, variability in antibody staining performance, semi-quantitative and subjective result calling. In cases of diminished sample quality, IHC-based subtype assessment may not agree with gene expression-based classification, and alternative approaches may be needed. This study aimed to develop a machine learning classifier able to predict breast cancer receptor subtypes using genomic features, without relying on immunohistochemistry or gene expression data. This study included 19,559 patients with primary breast cancer, identified using Natera's proprietary real-world database, linked to a clinical claims database. Hormone receptor (HR) and HER2 subtype was determined from patient treatment codes. We developed a biologically-informed feature set by combining somatic mutations across 19,820 genes, using whole exome sequencing (WES) data from the SignateraTM testing workflow. Each mutation was assigned a composite mutation_score (range 1-12) based on variant class (SNV, insertion, deletion), superclass (SNP/INDEL), predicted impact (VEP annotation impact: MODIFIER to HIGH), and functional consequence (such as frameshift, stop-gain, missense, synonymous). A Random Forest classifier was trained with a stratified 75/25 train-test splitting and hyperparameter optimization. The model was trained on features from 14,669 patients in the training cohort. In a test cohort of 4,890 patients, the model achieved 80.3% overall agreement with HR/HER2 status as inferred through medication claims data, with balanced performance across four major subtypes. Per-subtype metrics were: for HR+/HER2-, the model showed a precision of 0.935, recall 0.911, and F1 score of 0.923; for HR-/HER2+, precision was 0.714, recall was 0.753, and F1 score was 0.783; for HR+/HER2+, precision was 0.748, recall was 0.734, and F1 score was 0.741; lastly, for the TNBC subtype, precision was 0.730, recall was 0.816, and F1 score was 0.770. Overall the genomic classifier accurately classifies breast cancer into one of the four major receptor subtypes. After definitive validation against clinically-reported HR/HER2 status, this classifier could be used to guide analyses of de-identified genomic datasets that lack complete clinical annotation.
利益披露 Disclosure
S. Satta, Natera, Inc. Employment, Stock, Stock Option. P. Miller, Natera, Inc. Employment, Stock, Stock Option. S. Rivero-Hinojosa, Natera, Inc. Employment. E. Kalashnikova, Natera, Inc. Employment, Stock, Stock Option. A. Rodriguez, Natera, Inc. Employment, Stock, Stock Option. M. C. Liu, Natera, Inc. Employment, Stock, Stock Option.

← 返回 AACR 2026 检索