PO.BCS01.15 · 生物信息与计算

使用基于序列的深度学习模型模拟转移性乳腺癌中碱基对水平的突变率

Modeling base-pair level mutation rate in metastatic breast cancer using a sequence-based deep learning model

海报缩略图:使用基于序列的深度学习模型模拟转移性乳腺癌中碱基对水平的突变率
编号 1500 展板 7 时间 4/20 09:00–12:00 区域 Section 6 主讲 Ariaki Dandawate, BS
分会场 Sequence Analysis
查看 PDF 下载 PDF 🔒 查看 / 下载完整 PDF 需登录并开通下载套餐 · 查看套餐 / 开通 AACR 官方页面

作者与单位 Authors & Affiliations

Ariaki Dandawate1, Christina Leslie2, Ekta Khurana3

1Physiology, Biophysics and Systems Biology, Weill Cornell Graduate School, New York, NY,2Memorial Sloan Kettering Cancer Center, New York, NY,3Weill Cornell Medicine, New York, NY

摘要 Abstract

中文摘要
通过利用基于序列的深度学习框架,我们旨在揭示突变在非编码调控区产生的机制,从而有可能发现转移性乳腺癌中的新型靶点。 围绕转移性乳腺癌的研究主要集中于分析编码突变以表征疾病进展。尽管已知非编码突变会影响转录因子结合和基因表达调控,但很少有非编码突变驱动因素被鉴定出来。此前发表的研究表明,转移性突变率与起源细胞中的开放染色质相关。然而,这项工作大多是在粗略的区域级尺度上完成的,仅鉴定出突变热点区域。在本研究中,我们旨在揭示转移性乳腺癌调控区中突变率升高的基因组位置,并鉴定其潜在的突变机制。 我们提出一种新型深度学习模型,以揭示转移性乳腺癌调控区中每碱基突变率的基于序列上下文的协变量。借助来自Hartwig Medical Foundation队列的超过500,000个突变,该神经网络在来自正常乳腺上皮调控区的序列上进行训练,并预测该区域每碱基的突变率分布。因此,该模型学习到序列特征如何在特定基因组位置改变突变可能性。 对模型显著性图的分析可鉴定出突变率高于预期的特定位点,这可能表明转录因子结合增加,超出了促转移调控程序的选择范围。这些发现提供了一种将非编码突变模式与转移基因组中的突变机制和调控效应相联系的方法。 未来工作旨在将该模型扩展到泛癌水平,揭示共享的和癌症特异性的非编码突变,这些突变有望揭示转移的模式。这一可扩展的序列模型框架相对于现有方法具有优势,尤其得益于其对转移性癌症基因组的碱基对水平分辨率建模。因此,它对于建模和揭示癌症调控机制的意义至关重要。
查看英文原文 English abstract
By leveraging a sequence-based deep learning framework, we seek to uncover mechanisms by which mutations arise in noncoding regulatory regions, potentially leading to the discovery of novel targets in metastatic breast cancer. Research around metastatic breast cancer has largely focused on analyzing coding mutations to characterize progression. Though noncoding mutations are known to affect transcription factor binding and regulation of gene expression, few noncoding mutation drivers have been identified. Previously published work has shown that metastatic mutation rate correlates with open chromatin in the cells-of-origin. However, this work has mostly been done at a coarse, region-level scale, identifying mutational hotspot regions. In this work, we aim to uncover genomic positions in regulatory regions of metastatic breast cancer with elevated mutation rates, and identify their potential mutation mechanism. We propose a novel deep learning model that uncovers the sequence context-based covariates of per-base mutation rate in regulatory regions of metastatic breast cancer. With access to over 500,000 mutations from the Hartwig Medical Foundation cohort, the neural network is trained on sequences from regulatory regions in normal breast epithelium, and predicts per-base mutation rate profiles for the region. As a result, the model learns how sequence features change mutation likelihood at particular genomic positions. Analysis of the saliency map of the model allows for identification of specific sites with higher-than-expected mutation rates, which is potentially indicative of increased transcription factor binding that extends beyond selection by pro-metastatic regulatory programs. These findings provide a method to connect noncoding mutation patterns to mutation mechanism and regulatory effects in metastatic genomes. Future work aims at expanding this model to a pan-cancer level, revealing shared and cancer-specific noncoding mutations that have potential to reveal patterns of metastasis. This scalable sequence model framework provides an advantage over existing methods particularly due to its base-pair level resolution modeling of the metastatic cancer genome. Thus, its implications for modeling and uncovering regulatory mechanisms of cancer is key.
利益披露 Disclosure
A. Dandawate, None.

← 返回 AACR 2026 检索