PO.BCS01.13 · 生物信息与计算

利用批量长读长测序对点突变进行时序分析以实现肿瘤克隆定相

Phasing tumor clones by timing point mutations using bulk long-read sequencing

海报缩略图:利用批量长读长测序对点突变进行时序分析以实现肿瘤克隆定相
编号 6912 展板 25 时间 4/22 09:00–12:00 区域 Section 4 主讲 Ataberk Donmez, BS;MS
分会场 New Algorithms and Computational Methods
查看 PDF 下载 PDF 🔒 查看 / 下载完整 PDF 需登录并开通下载套餐 · 查看套餐 / 开通 AACR 官方页面

作者与单位 Authors & Affiliations

Ataberk Donmez, Mikhail Kolmogorov

National Cancer Institute, Bethesda, MD

摘要 Abstract

中文摘要
癌细胞因体细胞畸变(如体细胞突变或结构变异)而与健康细胞不同。研究这些突变有助于我们更好地理解肿瘤演化,并为更有效的治疗策略开辟道路。癌细胞在不同细胞群之间表现出异质性,即使在单个肿瘤内部也是如此,而这些细胞群由其独特的体细胞突变所定义。给定一组体细胞突变——其中一些为不同克隆所共享,另一些则各自独有——确定哪些突变共同出现,可实现肿瘤克隆的定相。归根结底,从批量测序数据中对这些突变进行定相对于研究肿瘤演化至关重要。 SNV 和短 indel 是肿瘤细胞中最常见的体细胞突变,尽管这些突变可以通过短读长有效检测,但由于短读长的特性,连接邻近突变仍然具有挑战性。另一方面,长读长已成功用于将种系变异直接定相为兆碱基级别的相位块,能够连接远距离的 SNV,因此在体细胞变异定相方面显示出前景。肿瘤克隆的重建是一个多等位基因定相问题。获得所检测到的体细胞点突变(PM)的系统发育关系,是将这些突变定相为不同克隆的第一步,其中所构建系统发育树中的节点对应于可能的克隆单倍型。为此,我们提出基于成对 PM 在同时覆盖两个位点的读长中的共现情况、以及这些读长支持参考等位基因还是替代等位基因,来对成对 PM 进行相互时序分析。给定两个点突变以及一组覆盖这两个位置的读长,可以推断这两个突变是在同一分支中先后发生、共同发生,还是发生在不同分支中(即分歧)。由于成对 PM 的时序关系具有传递性,因此可以通过对未被读长连接的成对 PM 进行时序分析来获得更长的链。我们将这些关系表示为一个图,其中 PM 为节点,每种类型的关系用不同的边表示。 我们将该方法应用于 CASTLE 数据集(https://github.com/CASTLE-Panel/castle)中的 H2009 和 H1437 细胞系,使用常规及超长(100kb 以上)的 Oxford Nanopore 读长。尽管与典型的真实肿瘤样本相比,这些细胞系的异质性较低,但对体细胞突变进行定相也能够区分重复的染色体拷贝。对于 H1437 和 H2009 细胞系,我们共检测到 87,999 个和 162,334 个体细胞 SNP。用这些 SNP 构建的图包含的连通分量,中位跨度分别为 102Kb 和 55Kb(最大值:9.5Mb 和 6.4Mb)。对于这两个细胞系,连通分量的单倍型数量中位数均为 2(最大值分别为 56 和 80)。在检测到的这些单倍型中,两个细胞系由多个 SNP 组成的单倍型中位数均为 1(最大值分别为 23 和 46)。我们计划扩展该方法,以纳入对结构变异的时序分析。
查看英文原文 English abstract
Cancer cells differ from healthy cells due to somatic aberrations (e.g., somatic mutations or structural variations). Studying these mutations helps us better understand tumor evolution and opens the way to more effective treatment strategies. Cancer cells exhibit heterogeneity across different populations, even within a single tumor, and these populations are defined by their unique somatic mutations. Given a set of somatic mutations -some shared across different clones and some unique- determining which mutations co-occur enables phasing of tumor clones. Ultimately, phasing these mutations from bulk sequencing data is crucial for studying tumor evolution. SNVs and short indels are the most common somatic mutations in tumor cells and even though these mutations can be effectively detected with short reads, linking nearby mutations remains challenging due to the nature of these reads. On the other hand, long reads have been successfully used in direct phasing of germline variants into megabase-scale phase blocks by enabling linkage of distant SNVs and thus shows promise for phasing somatic variants.Reconstruction of tumor clones is a multi-allelic phasing problem. Obtaining a phylogeny of detected somatic point mutations (PM) is a first step in phasing these mutations into different clones where nodes in the constructed phylogeny correspond to possible clonal haplotypes. For this purpose, we propose timing pairs of PMs against each other based on their co-occurrences in the reads that cover both locations and whether these reads support the reference or alternate allele. Given two point mutations and a set of reads that cover these two positions; it is possible to deduce whether these two mutations occurred one after the other in the same branch, co-occurred together, or occurred in different branches (i.e., divergent). Since timing of PM pairs are transitive, it is possible to obtain longer chains by timing PM pairs that are not connected by reads. We represent these relationships as a graph where PMs are nodes and each type of relationship is represented by a different edge. We applied our approach to H2009 and H1437 cell lines from the CASTLE collection (https://github.com/CASTLE-Panel/castle) with regular and ultra-long (100kb+) Oxford Nanopore reads. Although these cell lines are less heterogeneous compared to typical real tumor samples, phasing somatic mutations can also distinguish duplicated chromosome copies. For H1437 and H2009 cell lines, we detected a total of 87,999 and 162,334 somatic SNPs. The graphs constructed with these SNPs contained connected components with median spans of 102Kb and 55Kb (maximum: 9.5Mb and 6.4Mb). For both cell lines, connected components had a median of 2 (maximum: 56 and 80) haplotypes. Of these detected haplotypes, median of 1 for both cell lines (maximum: 23 and 46) consists of multiple SNPs. We plan to extend our approach to include timing of structural variations.
利益披露 Disclosure
A. Donmez, None.. M. Kolmogorov, None.

← 返回 AACR 2026 检索