LBPO.BCS01 · 生物信息与计算 · Late-Breaking
scCAP:一个面向本体的大语言模型框架,用于分层且标准化的单细胞类型注释
scCAP: An ontology-aware large language model framework for hierarchical and standardized single-cell type annotation
该海报暂无可下载的资料
AACR 官方页面
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
背景:在单细胞RNA测序分析中,准确的细胞类型注释仍具挑战性。现有工具(CellTypist、SCimilarity、SingleR、Geneformer)表现出各自不同的偏倚,且缺乏标准化的命名,阻碍了跨研究的可比性。此外,当前方法无法提供反映细胞类型本体的分层注释。
方法:我们开发了scCAP,这是一个元学习框架,采用面向本体的大语言模型(LLM)架构整合多种注释工具。scCAP采用HybridAttentionNet,结合了:(1)在语义嵌入上使用自注意力的知识流,以及(2)学习工具特异性可靠性模式的置信度流。一个关键创新是我们基于LLM的标签解读:对于每一次预测,LLM生成结构化描述,包括全称、标志基因和通路,从而实现自动缩写展开和聚类注释解读。这些描述被嵌入,以在异质命名体系之间进行标准化的语义比较。LLM还将标签分类到分层本体层级(L1:谱系;L2:细胞大类;L3:亚型),并依据Cell Ontology进行验证。训练期间的工具丢弃(dropout)正则化确保了即使个别工具不可用时也能保持稳健性能。此外,迭代式伪标注通过将高置信度注释作为训练信号,逐步优化预测结果。这种模块化架构允许在无需重新训练的情况下无缝添加/替换工具。
结果:我们在Cancer Cell Census Atlas和Human Lung Cell Atlas数据上,将scCAP与四种注释工具进行了评估。在Level 3层级,scCAP实现了0.898的平均语义相似度,优于SingleR(0.865)、CellTypist(0.861)、SCimilarity(0.857)和Geneformer(0.805)。在≥80%相似度阈值下,scCAP实现了87.7%的准确率,而CellTypist为73.3%。scCAP展现出一致的分层性能(L1:0.858;L2:0.874;L3:0.898),中位相似度更高,方差比单个工具更小。
结论:scCAP利用LLM实现面向本体的标签标准化和分层分类。其即插即用架构能够在纳入新兴工具的同时保持输出一致性,从而促进对肿瘤微环境异质性的元分析。
查看英文原文 English abstract
Background: Accurate cell type annotation remains challenging in single-cell RNA sequencing analysis. Existing tools (CellTypist, SCimilarity, SingleR, Geneformer) exhibit distinct biases and lack standardized nomenclature, hindering cross-study comparability. Furthermore, current methods fail to provide hierarchical annotations reflecting cell type ontologies.
Methods: We developed scCAP, a meta-learning framework integrating multiple annotation tools using an ontology-aware large language model (LLM) architecture. scCAP employs HybridAttentionNet combining: (1) a knowledge stream using self-attention on semantic embeddings, and (2) a confidence stream learning tool-specific reliability patterns. A key innovation is our LLM-based label interpretation: for each prediction, the LLM generates structured descriptions including full names, marker genes, and pathways, enabling automatic abbreviation expansion and cluster annotation interpretation. These descriptions are embedded for standardized semantic comparison across heterogeneous nomenclatures. The LLM also classifies labels into hierarchical ontology levels (L1: lineages; L2: cell classes; L3: subtypes), validated against Cell Ontology. Tool dropout regularization during training ensures robust performance even when individual tools are unavailable. Additionally, iterative pseudo-labeling progressively refines predictions by incorporating high-confidence annotations as training signals. The modular architecture allows seamless tool addition/replacement without retraining.
Results: We evaluated scCAP on Cancer Cell Census Atlas and Human Lung Cell Atlas data against four annotation tools. At Level 3, scCAP achieved 0.898 average semantic similarity, outperforming SingleR (0.865), CellTypist (0.861), SCimilarity (0.857), and Geneformer (0.805). At ≥80% similarity threshold, scCAP achieved 87.7% accuracy versus 73.3% for CellTypist. scCAP demonstrated consistent hierarchical performance (L1:0.858; L2:0.874; L3:0.898) with higher median similarity and tighter variance than individual tools.
Conclusions: scCAP leverages LLMs for ontology-aware label standardization and hierarchical classification. Its plug-and-play architecture enables incorporating emerging tools while maintaining consistent outputs, facilitating meta-analyses of tumor microenvironment heterogeneity.
利益披露 Disclosure
D. Shin, None..
S. Jang, None..
J. Lee, None..
J. Cho, None.