PO.CL01.17 · 临床研究
靶向杂交捕获NGS cfDNA基因panel数据的片段组学分析产生一个正交的基因组维度,用于开发转移性乳腺癌患者的稳健预后模型
Fragmentomics analysis of targeted hybrid capture NGS cfDNA gene panel data yields an orthogonal genomic dimension for developing a robust prognostic model for metastatic breast cancer patients
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
引言:cfDNA片段分布分析,即片段组学,可产生有关基因转录活性的信息,但主要应用于WGS和WES数据集。靶向杂交捕获NGS cfDNA基因panel的广泛应用提供了丰富的数据来源,可用于挖掘正交的片段组学特征,以开发用于预后和预测反应的新型算法特征。在此,我们将一种为杂交捕获panel数据独特开发的片段组学计算方法应用于一项针对HER2阴性转移性乳腺癌患者的前瞻性研究,这些患者均接受了paclitaxel和bevacizumab治疗。
实验步骤:从NCBI(PRJNA745047)下载了基线(n=182)的Fastq文件及相关临床数据。将来自一个靶向48个乳腺癌相关基因和8个癌症基因启动子的定制panel的原始测序数据(平均深度约2600x)通过专有流程处理,得到一个由2936个特征组成的片段组学矩阵。样本按总生存期和关键临床变量平衡后分为训练-测试集(2/3-1/3)。在训练样本中使用片段组学矩阵特征,通过R语言的glmnet软件包建立总生存期(OS)和三阴性(TN)特征。使用log-rank检验和多变量Cox模型评估OS特征与OS之间的关联。使用Wilcoxon检验和AUC评估TN特征与TN状态之间的关联。
结果:在训练集中开发的OS特征使用了56个片段组学特征。该特征与OS之间的关联在低和高肿瘤含量分层(肿瘤含量使用可用的VAF数据定义)中在训练和测试数据集中方向一致(且每个名义log-rank检验p值 < 0.05)。在测试集中,将片段组学特征加入到已包含TN状态、肿瘤分级和转移部位参数的多变量OS模型中,显著改善了模型拟合(LRT p值 = 0.009),特征风险比为2.50(95% CI 1.25-4.99)。在训练集中开发的TN状态特征使用了56个片段组学特征,其中6个与OS特征重叠。在测试集中,该特征与TN状态显著相关(Wilcoxon p值 = 0.0052),AUC为0.75。
总结与结论:本研究证明了cfDNA杂交捕获片段组学数据在开发稳健预测因子方面的效用,即使在低肿瘤分数样本中也能预测总生存期和TN状态。尽管需要额外验证,这些特征可能代表有潜力的生物标志物,有助于临床决策。本研究还支持回顾性地使用任何cfDNA panel测序数据集进行片段组学分析以识别预测性生物标志物的可能性。
查看英文原文 English abstract
Introduction: Analysis of cfDNA fragment distributions, or fragmentomics, yields information about the transcriptional activity of genes but mostly has been applied to WGS and WES datasets. The widespread use of targeted hybrid capture NGS cfDNA gene panels provide a rich data source for mining orthogonal fragmentomic features for developing novel algorithmic signatures for prognosis and predicting response. Here, we applied an uniquely developed fragmentomic computational method for hybrid capture panel data to a prospective study of HER2 negative, metastatic breast cancer patients, who were all treated with paclitaxel and bevacizumab.
Experimental Procedures : Fastq files from baseline (n=182) and associated clinical data were downloaded from NCBI (PRJNA745047). Raw sequencing (avg. depth ~2600x) data from a custom panel targeting 48 breast cancer associated genes and 8 cancer gene promoters were processed through a proprietary pipeline to get a fragmentomics matrix comprised of 2936 features. Samples were split into a train-test set (2/3-1/3) balanced on overall survival and key clinical variables. Overall survival (OS) and triple negative (TN) signatures using fragmentomics matrix features in training samples with the glmnet package in R. Association between the OS signature and OS was evaluated using the log-rank test and multivariable cox models. Association between the TN signature and TN status was evaluated using the Wilcoxon test and AUC.
Results: An OS signature developed in the training set used 56 fragmentomics features. The association between the signature and OS was directionally uniform in low and high tumor content strata (tumor content defined using the available VAF data) in both training and test data sets (and each nominal log-rank test p-value < 0.05). In the test set, the addition of the fragmentomics signature to a multivariable OS model that already included parameters for TN status, tumor grade, and metastatic site significantly improved the model fit (LRT p-value = 0.009), with a signature hazard ratio of 2.50 (95% CI 1.25-4.99). The TN status signature developed in the training set used 56 fragmentomics features of which 6 overlapped with the OS signature. In the test set, the signature was markedly correlated with TN status (Wilcoxon p-value = 0.0052) with an AUC of 0.75.
Summary and Conclusions: This work demonstrates the utility of cfDNA hybrid capture fragmentomic data for developing robust predictors even in low tumor fraction samples of overall survival and TN status. Although additional validation is warranted, these signatures could represent potential biomarkers that could assist in clinical decision making. This work also supports the possibility of retrospectively using any cfDNA panel sequencing dataset for fragmentomics analysis to identify predictive biomarkers.
利益披露 Disclosure
G. M. Mayhew,
GeneCentric Employment.
J. H. Shepherd,
GeneCentric Employment.
J. Burdine,
GenceCentric Employment.
Y. Shibata,
GeneCentric Employment.
G. V. Milburn,
GeneCentric Employment.
M. V. Milburn,
GeneCentric Employment.
K. L. Pappan,
GeneCentric Employment.
J. M. Davison,
GeneCentric Employment.
K. Beebe,
GeneCentric Employment.