PO.BCS01.03 · 生物信息与计算
HELP-TCR:一个用于T细胞受体库功能分析的统一且可解释的语言处理框架
HELP-TCR: A harmonized and explainable language processing framework for functional analysis of t-cell receptor repertoires
该海报暂无可下载的资料
AACR 官方页面
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
T细胞是抗肿瘤免疫的重要贡献者,但识别区分肿瘤反应性应答与背景免疫活性的特定T细胞受体(TCR)特征仍难以解析。肿瘤通过新抗原暴露、慢性刺激和免疫逃逸重塑TCR库,形成了具有生物学相关性但不易用现有计算方法量化的模式。当前的方法要么依赖于全局库汇总,要么依赖于可解释性有限的深度学习模型,这限制了其转化和临床应用。为弥补这一空白,我们开发了HELP-TCR,这是一个可解释的机器学习框架,将自然语言处理的概念应用于肿瘤相关TCR库的功能分析。HELP-TCR通过特征(单个氨基酸和/或氨基酸对)的位置特异性分布来表示TCR库,将序列转换为张量结构。为提高可重复性,我们采用了集成学习的原理,并提出了一系列预处理步骤,其中应用了一种共识分组方法来合并位置分布高度相似的特征,从而开辟了可解释降维的途径。经过修改(以适应任务维度)的深度学习架构实现了准确分类,而基于显著性图的事后分析则突出了对模型预测贡献最大的最具信息量的特征。使用来自非小细胞肺癌的TCR序列数据集,HELP-TCR展示了稳定且高的预测性能(AUC约0.96),优于DeepTCR(AUC 0.76)和TCR-BERT嵌入,后两者的类别可分性有限。事后显著性分析识别出有助于肿瘤-正常区分的位置氨基酸基序和残基对相互作用,提供了可能反映肿瘤反应性、免疫压力或肿瘤微环境内克隆重塑的机制性见解。HELP-TCR为在序列级信息具有相关性的情境下分析肿瘤相关TCR库提供了一个可解释且可重复的框架。其表征方式与TCR分析新兴的临床应用相契合,例如评估对免疫治疗的应答、识别新抗原反应性或肿瘤富集的克隆型,以及研究肿瘤组织内的免疫动态。HELP-TCR为生成可检验的假设、推进旨在改善患者分层和指导基于T细胞的治疗策略的转化工作提供了一个实用基础。
查看英文原文 English abstract
T cells are essential contributors to anti-tumor immunity, but identifying the specific T-cell receptor (TCR) features that differentiate tumor-reactive responses from background immune activity remain difficult to resolve. Tumors reshape TCR repertoires through neoantigen exposure, chronic stimulation, and immune escape, creating patterns that are biologically relevant not easily quantified with existing computational methods. Current approaches either rely on global repertoire summaries or on deep learning models with limited interpretability, which restricts their translational and clinical applications.To address this gap, we developed HELP-TCR, an explainable machine learning framework that applies concepts from natural language processing to the functional analysis of tumor-associated TCR repertoires. HELP-TCR represents TCR repertoires by the position-specific distributions of features (single amino acids and/or amino acid pairs), transforming sequences into tensor structures. To increase reproducibility, we adapt the principles of ensemble learning and proposed a series of preprocessing steps, among which a consensus grouping method is applied to merge the features with highly similar position-wise distributions opening the venue for explainable dimension reduction. A modified (to the dimensionality of the task) deep learning architecture enables accurate classification, while post-hoc analysis based on saliency map highlights the most informative features contributing to model predictions.Using a dataset of TCR sequences from non-small cell lung cancer, HELP-TCR demonstrated stable and high predictive performance (AUC ~0.96), outperforming DeepTCR (AUC 0.76) and TCR-BERT embeddings, which showed limited class separability. Post-hoc saliency analysis identified positional amino acid motifs and residue-pair interactions that contributed to tumor-normal discrimination, offering mechanistic insights potentially reflecting tumor reactivity, immune pressure, or clonal remodeling within the tumor microenvironment.HELP-TCR offers an interpretable and reproducible framework for analyzing tumor-associated TCR repertoires in settings where sequence-level information is relevant. Its representations align with emerging clinical applications of TCR profiling, such as assessing of response to immunotherapy, identification of neoantigen-reactive or tumor-enriched clonotypes and studying of immune dynamics within the tumor tissue. HELP-TCR provides a practical foundation for generating testable hypotheses and advancing translational efforts aimed at improving patient stratification and informing T-cell-based therapeutic strategies.
利益披露 Disclosure
Y. Kalesnik, None..
D. Krawczyk, None..
M. Pietrzak, None..
M. Seweryn, None.