PO.MCB08.03 · 分子与细胞生物学

跨越8,000余个TCGA全基因组的结构变异连接处插入的起源

Origins of structural variant junctional insertions across >8,000 TCGA whole genomes

海报缩略图:跨越8,000余个TCGA全基因组的结构变异连接处插入的起源
编号 3247 展板 12 时间 4/20 02:00–05:00 区域 Section 22 主讲 Youyun Zheng, BA;BS
分会场 Genomic Profiling to Understand Cancer Biology
查看 PDF 下载 PDF 🔒 查看 / 下载完整 PDF 需登录并开通下载套餐 · 查看套餐 / 开通 AACR 官方页面

作者与单位 Authors & Affiliations

Youyun Zheng1, Gregory Raskind1, Sophie Webster1, Narmen Azazmeh1, Haruna Tomono1, Andrew Cherniack1, David Lehotzky1, Ron Solan1, Antonia Kowalewski1, Xavi Loinaz1, Hansol Park2, Vasuki N. Swamy1, David Heiman1, Samantha Van Seters1, Saveliy Belkin1, Sam Wiseman1, Chunyang Bao2, Luis A. Corchete Sanchez1, Zachary Everton1, Ryul Kim2, Beomki Lee2, Won-Chul Lee2, Chip Stewart1, Gengchao Wang1, Brian P. Danysh1, Young Seok Ju2, Esther Rheinbay1, Gad Getz3, Rameen Beroukhim1

1Cancer Program, Broad Institute, Cambridge, MA,2Inocras, San Diego, CA,3Massachusetts General Hospital, Charlestown, MA

摘要 Abstract

中文摘要
双链断裂修复在基因组中留下可识别的印记。其中最具特异性的是插入结构变异(SV)连接处的短序列——模板化插入,通常归因于聚合酶-θ介导的末端连接(TMEJ)。然而,基于精确字符串匹配的常见读出方法忽略了序列背景、与断裂点的距离以及候选模板的拓扑背景,从而夸大了假阳性并模糊了机制解读。我们提出一个可扩展的统计框架,通过针对经距离调整的全基因组经验零分布对比对得分分布进行建模,以受控的特异性来推断模板化插入。该方法明确考虑了不完美的复制,并将候选模板划分为四种断裂点邻近的构型,捕捉对机制有提示意义的位置和链关系。将该方法应用于大型体细胞和种系全基因组队列,揭示出模板使用跨越所有构型,但在不同细胞背景之间存在系统性差异。相当一部分事件反映了不完美的复制,与易错合成相一致,而其中一种构型显示出相对更高的表观保真度——提示在更广泛的TMEJ样图景中存在不同的生化途径。构型判定还对SV结构进行了分层:某些在简单重排中富集,而另一些则定位于聚集的、复杂的区域,表明局部拓扑与修复途径选择之间存在关联。除结构外,构型特异性的负荷还与DNA修复状态和选定的基因型相吻合:与同源重组缺陷一致的背景在特定构型中显示富集,而另一些则显示相反的方向性,这强调了"模板化插入"并非单一现象,而是一系列具有分歧决定因素的相关过程。为实现队列规模的分析,我们优化了核心比对以在单次运行中生成完整的得分矩阵,并将工作流程打包为容器化流程,实现了数量级的加速和可移植的可重现性。总之,这些结果确立了一种构型感知、统计上有原则的模板化插入读出方法,它对序列混杂因素稳健,并对机制有提示意义。在实践中,该框架提供了(i)用于研究人类样本中双链断裂修复的更锐利视角,(ii)修复状态生物标志物的线索,以及(iii)将SV拓扑与聚合酶使用联系起来的假说。由此,它旨在推动该领域从零散的序列勾勒转向关于涉及连接处插入的双链断裂修复的可重现、队列规模的推断。
查看英文原文 English abstract
Double-strand break repair leaves recognizable footprints in the genome. Among the most specific are short sequences inserted at structural-variant (SV) junctions-templated insertions often attributed to polymerase-θ-mediated end joining (TMEJ). Yet common readouts based on exact string matches overlook sequence background, distance from the break, and the topological context of candidate templates, inflating false positives and blurring mechanistic interpretation. We present a scalable statistical framework that infers templated insertions with controlled specificity by modeling alignment-score distributions against a distance-adjusted, genome-wide empirical null. The approach explicitly accommodates imperfect copying and partitions candidate templates into four breakpoint-proximal configurations, capturing positional and strand relationships that are informative of mechanisms. Applied to large somatic and germline whole-genome cohorts, the method reveals that template usage spans all configurations but differs systematically across cellular contexts. A notable fraction of events reflect imperfect copying, consistent with error-prone synthesis, whereas one configuration shows comparatively higher apparent fidelity-suggesting distinct biochemical routes within a broader TMEJ-like landscape. Configuration calls also stratify SV architecture: some are enriched in simple rearrangements while others localize to clustered, complex regions, indicating that local topology and repair pathway choice are linked. Beyond structure, configuration-specific burdens align with DNA-repair states and selected genotypes: contexts consistent with homologous-recombination deficiency show enrichment in particular configurations, while others display the opposite directionality, underscoring that “templated insertion” is not a single phenomenon but a family of related processes with diverging determinants. To enable cohort-scale analysis, we optimized the core alignment to produce full score matrices in a single pass and packaged the workflow into a containerized pipeline, yielding order-of-magnitude speedups and portable reproducibility. Together, these results establish a configuration-aware, statistically principled readout of templated insertions that is robust to sequence confounders and informative about mechanisms. Practically, the framework provides (i) a sharper lens for studying double-strand break repair in human samples, (ii) leads for repair-state biomarkers, and (iii) hypotheses connecting SV topology to polymerase usage. In doing so, it aims to move the field from anecdotal sequence sketches toward reproducible, cohort-scale inferences about double-strand break repair involving junctional insertions.
利益披露 Disclosure
Y. Zheng, None.. G. Raskind, None.. S. Webster, None.. N. Azazmeh, None.. H. Tomono, None. A. Cherniack, Bayer ). D. Lehotzky, None.. R. Solan, None.. A. Kowalewski, None.. X. Loinaz, None. H. Park, Inocras Employment. V. N. Swamy, None.. D. Heiman, None.. S. Van Seters, None.. S. Belkin, None.. S. Wiseman, None. C. Bao, Inocras Employment. L. A. Corchete Sanchez, None.. Z. Everton, None. R. Kim, Inocras Employment. B. Lee, Inocras Employment. W. Lee, Inocras Employment. C. Stewart, None.. G. Wang, None.. B. P. Danysh, None.. Y. Ju, None. E. Rheinbay, Inocras ). G. Getz, IBM ). R. Beroukhim, Karyoverse Stock. LOH Therapeutics Stock.

← 返回 AACR 2026 检索