PO.BCS01.13 · 生物信息与计算
利用转录组学的测试时计算对儿童T细胞急性淋巴细胞白血病进行亚型分类
Test-time compute for subtype classification in pediatric T-cell acute lymphoblastic leukemia using transcriptomics
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
背景:T细胞急性淋巴细胞白血病(T-ALL)是一种高度侵袭性的血液系统恶性肿瘤,具有显著的分子和临床异质性。尽管目前的T-ALL亚型分类可能尚未被普遍纳入临床决策,但准确、快速分类的潜力对未来的预后分层和指导靶向治疗具有重大前景。传统的分类方法往往耗时且劳动强度大,凸显了对高效、数据驱动方法的需求。仅利用转录组数据,我们旨在通过开发一种新型方法来克服这些局限,以实现精确的T-ALL亚型分类。
方法:我们开发了一个从RNA-seq数据进行T-ALL亚型分类的机器学习流程,引入了测试时计算范式。首先,在一个大型发现队列(Polonen等,2024;1,112个样本,24,619个基因)上训练随机森林模型,通过为每个亚型选择前100个基因来识别最具预测性的特征。在预测时,该流程动态执行:(1)从训练和测试数据集中过滤预选特征(例如TARGET队列,264个样本,22,688个基因);(2)通过pycombat进行批次校正;(3)在处理后的训练集上重新训练随机森林分类器;(4)对处理后的测试数据进行预测。此测试时流程可通过Streamlit网页应用和命令行访问。
结果:我们使用F1评分来评估分类器的性能。在发现队列上进行初始训练的随机森林在分类TAL-like、TLX-like、NKX2-1、ETP-like和"其他"亚型时达到93%。从每个亚型选择前100个特征得到478个特征。使用这些选定特征对模型性能进行基准测试,得到96%的F1评分。随后我们将测试时流程应用于包含264个样本的独立TARGET队列。最终模型在5个亚型上达到83%的准确率。对临床相关亚型的性能最高:TAL-like(F1:0.96)、TLX-like(F1:0.94)和NKX2-1(F1:0.89)。"其他"类别得分中等(F1:0.60),而ETP-like在验证集中缺失。
结论:我们基于RNA-seq的模型通过利用测试时计算、整合动态批次校正和实时模型再训练以及一个精选的478基因特征集,提供了稳健、可扩展的T-ALL亚型分类。在独立队列上的强劲表现凸显了其作为精准肿瘤学快速、可靠工具的临床实用性。
查看英文原文 English abstract
Background: T-cell Acute Lymphoblastic Leukemia (T-ALL) represents a highly aggressive hematologic malignancy characterized by profound molecular and clinical heterogeneity. While current T-ALL subtype classifications may not yet be universally integrated into clinical decision-making, the potential for accurate and rapid classification holds significant promise for future prognostic stratification and guiding targeted therapies. Traditional classification methods can be time-consuming and labor-intensive, highlighting the need for efficient, data-driven approaches. Leveraging transcriptomic data alone, we aimed to overcome these limitations by developing a novel approach for precise T-ALL subtype classification.
Methods: We developed a machine learning pipeline for T-ALL subtype classification from RNA-seq data, introducing test-time compute paradigm. Initially, a Random Forest model was trained on a large discovery cohort (Polonen et al., 2024; 1,112 samples, 24,619 genes) to identify top predictive features by selecting the top 100 genes per subtype. At prediction time, the pipeline dynamically executes: (1) filtering preselected features from both training and test datasets (e.g., TARGET cohort, 264 samples, 22,688 genes); (2) batch correction via pycombat; (3) re-training of the Random Forest classifier on the processed training set; and (4) prediction on processed test data. This test-time pipeline is accessible via a Streamlit WebApp and Command Line.
Results: We used F1 score to evaluate the performance of our classifier. Random Forest on initial training with discovery cohort achieved 93% classifying TAL-like, TLX-like, NKX2-1, ETP-like and 'other' subtypes. Selecting the top 100 features from each subtype yielded 478 features. Benchmarking the performance of our model using these selected features resulted in a 96% F1 score. We then applied our test-time pipeline on the independent TARGET cohort with 264 samples. The final model achieved 83% accuracy across 5 subtypes. Performance was highest for clinically relevant subtypes: TAL-like (F1: 0.96), TLX-like (F1: 0.94), and NKX2-1 (F1: 0.89). The “other” category scored moderately (F1: 0.60), while ETP-like was absent in the validation set.
Conclusion: Our RNA-seq-based model delivers robust, scalable T-ALL subtype classification by leveraging test-time compute, integrating dynamic batch correction and real-time model retraining with a curated 478-gene feature set. Strong performance on an independent cohort highlights its clinical utility as a rapid, reliable tool for precision oncology.
利益披露 Disclosure
T. Mamidi, None..
I. Pushel, None..
B. Yoo, None..
M. S. Farooqi, None..
K. J. August, None.