PO.BCS01.17 · 生物信息与计算
利用细胞类型特异性核小体模式改进肺癌中游离DNA片段化模型
Using cell-type-specific nucleosome patterns to improve cell-free DNA fragmentation models in lung cancer
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
背景:游离DNA(cfDNA)分析已成为一种有前景的癌症检测工具。方法包括鉴定循环肿瘤DNA(ctDNA)、分析cfDNA甲基化,以及分析cfDNA片段(包括片段大小和位置)。我们此前证明片段大小多样性与基因表达强相关(Esfahani等,2022)。我们随后开发了cfDNA片段化过程的机制模型,以增强从cfDNA推断基因表达的能力(Liu等,2025)。
然而,仅分析cfDNA片段大小分布会遗漏关于片段位置的有用信息。由于片段位置源自核小体位置,我们纳入关于核小体定位的先验知识以改进cfDNA片段化模型。我们假设纳入细胞类型特异性核小体模式可推动cfDNA模型向细胞类型解卷积方向发展。
方法:为了将关于细胞类型特异性核小体组织的先验知识纳入我们的cfDNA片段化模型,我们使用来自(Valouev等,2011)的微球菌核酸酶测序(MNase-seq)数据,以及来自(Druliner等,2016)的MNase-转录起始位点序列捕获法(mTSS-seq)数据。对于每个转录起始位点(TSS),我们利用关于核小体组织的先验知识模拟定位良好的核小体。然后,根据基因的链信息和细胞类型生成额外的核小体。接着,我们生成切割位点,并生成来自不同细胞类型的模拟片段混合物。
结果:为了评估我们更新后cfDNA片段化模型的效能,我们计算了TSS窗口内以10 bp为间隔的覆盖度,并将其与cfDNA对照样本的覆盖度进行比较。我们还使用先前模型计算模拟cfDNA片段的覆盖度,并将其与cfDNA对照样本的覆盖度进行比较。我们纳入核小体组织先验知识的更新模型在模拟cfDNA片段化方面优于先前模型。
结论:我们证明整合细胞类型特异性核小体模式可改进cfDNA片段化模型。根据更多细胞类型的核小体组织信息,还可进一步改进。我们预期这种基于模型的方法将增强组织来源分类和癌症检测。
查看英文原文 English abstract
Background: Cell-free DNA (cfDNA) profiling has emerged as a promising tool in cancer detection. Methods include the identification of circulating tumor DNA (ctDNA), the profiling of cfDNA methylation, and the analysis of cfDNA fragments, including fragment size and position. We previously showed that fragment size diversity is strongly correlated with gene expression (Esfahani et al, 2022). We subsequently developed a mechanistic model for the cfDNA fragmentation process to enhance gene expression inference from cfDNA (Liu et al, 2025).
However, analysis of cfDNA fragment size distribution alone leaves out useful information about fragment position. Since fragment position arises from nucleosome position, we incorporate prior knowledge about nucleosome positioning to improve our cfDNA fragmentation models. We hypothesize that incorporating cell-type-specific nucleosome patterns can progress cfDNA models towards cell-type deconvolution.
Methods: In order to incorporate prior knowledge about cell-type-specific nucleosome organization into our cfDNA fragmentation model, we use Micrococcal Nuclease sequencing (MNase-seq) data from (Valouev et al, 2011) and MNase-Transcription Start Site Sequence Capture method (mTSS-seq) data from (Druliner et al, 2016). For each transcription start site (TSS), we leverage this prior knowledge about nucleosome organization to simulate well-positioned nucleosomes. Then, we generate additional nucleosomes given the gene's strand information and cell type. Next, we generate cuts, and generate a mixture of simulated fragments from various cell types.
Results: In order to evaluate the efficacy of our updated cfDNA fragmentation model, we calculate the coverage over 10 bp intervals within a window of the TSS, and compare it with the coverage for cfDNA control samples. We also calculate the coverage for simulated cfDNA fragments using our previous model, and compare it with the coverage for cfDNA control samples. Our updated model incorporating prior knowledge about nucleosome organization outperformed our previous model in simulating cfDNA fragmentation.
Conclusions: We show that integrating cell-type-specific nucleosome patterns improves cfDNA fragmentation models. Further improvements can be made given nucleosome organization information from additional cell types. We envision that this model-based approach will enhance tissue-of-origin classification and cancer detection.
利益披露 Disclosure
B. S. Liu, None..
M. Shahrokh Esfahani, None.