PO.BCS02.05 · 生物信息与计算

MethylFM:一种用于建模表观基因组调控动态的DNA甲基化基础模型

MethylFM: A DNA methylation foundation model for modeling epigenomic regulatory dynamics

海报缩略图:MethylFM:一种用于建模表观基因组调控动态的DNA甲基化基础模型
编号 5482 展板 18 时间 4/21 02:00–05:00 区域 Section 2 主讲 Limeng Pu, PhD
分会场 Deep Learning in Cancer
查看 PDF 下载 PDF 🔒 查看 / 下载完整 PDF 需登录并开通下载套餐 · 查看套餐 / 开通 AACR 官方页面

作者与单位 Authors & Affiliations

Limeng Pu, Xiang Chen

St. Jude Children's Research Hospital, Memphis, TN

摘要 Abstract

中文摘要
DNA甲基化调控基因表达、分化和疾病进程,是精准医学中计算建模的关键靶标。为了将高分辨率全基因组亚硫酸氢盐测序(WGBS)数据整合到统一的分析框架中,我们开发了MethylFM——一种基于transformer的基础模型,能够捕捉情境感知的甲基化模式并支持多种下游任务。我们在BLUEprint Epigenome数据集上进行训练,该数据集包含跨多种血细胞类型的单碱基WGBS图谱,并配有匹配的组蛋白修饰和转录组数据。通过聚焦于转录起始位点(TSS)周围±100个CpG位点,MethylFM靶向对基因表达和染色质动态至关重要的调控区域。该模型基于BERT风格的transformer,采用旋转位置嵌入和掩码值预测目标,学习能够同时捕捉甲基化中局部和长程依赖关系的稳健表征。MethylFM在三个下游应用中展示了其多功能性。在CpG水平的插补方面,它以较高的准确性从450k芯片数据重建了高分辨率WGBS图谱(R² > 0.6,MAE < 0.15)。在与METHimpute进行基准对比时,由于METHimpute运行耗时较大,评估仅限于两个样本;尽管如此,MethylFM取得了略高的准确性(R² = 0.518对比0.513),凸显了其精度和计算效率。在TSS水平的H3K27ac预测中,该模型达到R² = 0.614,与最先进的M2A(R² = 0.617)相当,突显了其直接从DNA甲基化推断启动子活性的能力。最后,基于预测的H3K27ac图谱的样本水平聚类准确地再现了造血谱系,超越了实验性H3K27ac数据(轮廓系数 = 0.47对比0.30),并接近RNA-seq衍生的聚类性能(轮廓系数 = 0.51),表明MethylFM捕捉到了具有生物学意义的表观遗传结构。总之,这些结果确立了MethylFM作为一种可泛化且高效的表观基因组建模框架,能够实现经济高效的甲基化插补、启动子活性预测和细胞身份表征,从而推进生物标志物发现和精准医学。
查看英文原文 English abstract
DNA methylation regulates gene expression, differentiation, and disease, making it a key target for computational modeling in precision medicine. To integrate high-resolution whole-genome bisulfite sequencing (WGBS) data into a unified analytical framework, we developed MethylFM, a transformer-based foundation model that captures context-aware methylation patterns and supports multiple downstream tasks.We trained on the BLUEprint Epigenome dataset, comprising single-base WGBS profiles across diverse blood cell types with matched histone modification and transcriptome data. By focusing on ±100 CpG sites around transcription start sites (TSS), MethylFM targets regulatory regions central to gene expression and chromatin dynamics. Built on a BERT-style transformer with rotary positional embeddings and a masked-value prediction objective, it learns robust representations capturing both local and long-range dependencies in methylation.MethylFM demonstrated versatility across three downstream applications. For CpG-level imputation, it reconstructed high-resolution WGBS profiles from 450k array data with strong accuracy (R² > 0.6, MAE < 0.15). When benchmarked against METHimpute, evaluation was restricted to two samples due to METHimpute's intensive runtime; nonetheless, MethylFM achieved slightly higher accuracy (R² = 0.518 vs. 0.513), underscoring both precision and computational efficiency. In TSS- level H3K27ac prediction, the model reached R² = 0.614, matching the state-of-the-art M2A (R² = 0.617) and highlighting its capacity to infer promoter activity directly from DNA methylation. Finally, sample-level clustering based on predicted H3K27ac profiles accurately recapitulated hematopoietic lineages, surpassing experimental H3K27ac data (Silhouette = 0.47 vs. 0.30) and approaching RNA-seq-derived clustering performance (Silhouette = 0.51), demonstrating that MethylFM captures biologically meaningful epigenetic structure. Together, these results establish MethylFM as a generalizable and efficient framework for epigenomic modeling, enabling cost-effective methylation imputation, promoter activity prediction, and cellular identity characterization to advance biomarker discovery and precision medicine.
利益披露 Disclosure
L. Pu, None.

← 返回 AACR 2026 检索