PO.BCS01.02 · 生物信息与计算
从FFPE存档组织检测非整倍体:一种针对FFPE-CUTAC数据的计算方法
Aneuploidy detection from FFPE archived tissues: A computational approach for FFPE-CUTAC data
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
一个多世纪以来,福尔马林固定石蜡包埋(FFPE)样本制备一直是生物材料长期保存的标准,但严重的分子降解长期以来一直是高质量基因组分析的技术障碍。我们最近引入了FFPE-Cleavage Under Targeted Accessible Chromatin(FFPE-CUTAC,靶向可及染色质下切割),使用抗RNA聚合酶II(RNAPII)抗体,作为一种灵敏、经济高效的FFPE转录分析替代方法。除转录外,这些图谱还提供了推断非整倍体(整条染色体臂的获得或缺失)的机会,这是癌症的一个关键标志。然而,FFPE-CUTAC数据的稀疏性带来了计算挑战,目前尚无现成的合用工具。
用于全基因组测序(WGS)数据的拷贝数变异(CNA)检测方法通常依赖于读段深度的倍数变化。然而,我们的分析显示FFPE CUTAC与匹配的WGS读段深度在整个基因组中呈中度相关(Pearson相关系数r = 0.525,Spearman相关系数ρ = 0.688)。因此,直接将标准的基于读段深度的CNA方法应用于FFPE-CUTAC可获得合理但欠优的恢复率。这些发现表明需要一种量身定制的方法,明确考虑FFPE-CUTAC数据固有的稀疏性和噪声。
我们的新方法为每个基因组区间估计读段深度基线,而非为整个样本估计单一基线,以更好地捕捉正常拷贝数水平。我们将基因组划分为1 Mb的区间,并使用GC含量作为基因组特征对这些区间进行分组。GC含量相似的区间共享同一读段深度基线,该基线定义为相应GC含量组内所有区间读段计数的平均值。我们使用30个脑膜瘤样本,对照来自同一患者队列的标准WGS参考数据,评估了我们策略的性能。我们的方法展示了95.4%的总体非整倍体检测准确率(1116/1170条染色体臂),并通过正确识别98.9%的所有完整臂确保了低假阳性率。相比之下,来自匹配FFPE RNA-seq分析的非整倍体检测较低,准确率为86.32%(1010/1170条染色体臂)。
我们建立了一种专为FFPE-CUTAC数据非整倍体分析而设计的策略。该方法解锁了从FFPE-CUTAC数据灵敏、经济地识别关键染色体臂变异的能力,减轻了生成额外昂贵WGS数据的需求。
查看英文原文 English abstract
For more than a century, Formalin-Fixed Paraffin-Embedded (FFPE) sample preparation has been the standard for long-term preservation of biological material, but severe molecular degradation has long been a technical barrier to high-quality genomic analysis. We recently introduced FFPE-Cleavage Under Targeted Accessible Chromatin (FFPE-CUTAC) with an antibody to RNA Polymerase II (RNAPII) as a sensitive, cost-effective alternative for profiling transcription in FFPEs. Beyond transcription, these profiles present an opportunity to infer aneuploidy (whole chromosome arm gain or loss), a key cancer hallmark. However, the sparse nature of FFPE-CUTAC data presents a computational challenge with no existing fit-for-purpose tools.
Copy number alterations (CNA) detection methods for whole genome sequencing (WGS) data typically rely on fold change in read depth. However, our analysis shows moderate correlation between FFPE CUTAC and matched WGS read depth across the genome(Pearson correlation coefficient r = 0.525, Spearman correlation coefficient ρ = 0.688). As a result, directly applying standard read-depth-based CNA methods to FFPE-CUTAC achieves a reasonable but suboptimal recovery rate. These findings indicate the need for a tailored approach that explicitly accounts for the sparsity and noise inherent in FFPE-CUTAC data.
Our new method estimates a read depth baseline for each genomic bin rather than a single baseline for the entire sample to better capture the normal copy number level. We partition the genome into 1 Mb bins and use GC content as a genomic feature to group these bins. Bins with similar GC content share the same read depth baseline, defined as the mean read count across all bins within the corresponding GC content group. We evaluate the performance of our strategy using 30 meningioma samples against a standard WGS reference data from the same patient cohort. Our method demonstrated a 95.4% overall aneuploidy detection accuracy (1116/1170 chromosome arms) and ensured a low false-positive rate by correctly identifying 98.9% of all intact arms. By comparison, aneuploidy detection from matching FFPE RNA-seq profiling was lower, with an accuracy of 86.32% (1010/1170 chromosome arms).
We established a strategy specifically designed for aneuploidy profiling from FFPE-CUTAC data. This method unlocks the ability to sensitively and affordably identify critical chromosome arm variations from the FFPE-CUTAC data, alleviating the need to generate additional costly WGS data.
利益披露 Disclosure
A. T. Parmar, None..
Y. Niu, None..
K. Ahmad, None..
S. Henikoff, None..
Y. Zheng, None.