PO.BCS02.01 · 生物信息与计算
用于 RNA-Seq 的智能体 AI:从工作流程自动化到可操作的洞见
Agentic AI for RNA-Seq: From workflow automation to actionable insights
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
RNA-seq 数据的生物信息学分析自动化具有挑战性,因为每个项目都需要独特的分析步骤组合以及频繁的、针对具体情况的调整。这些项目特定的流程限制了工作流程在其他分析中的可复用性,并需要大量的手动编码。大语言模型(LLM)智能体非常适合应对这些挑战,因为它们能够解读自然语言指令、动态规划工作流程,并在无需手动编码的情况下适应特定研究的需求。我们开发了一个智能体 AI 平台,利用 LLM 通过自然语言指令来规划、执行和解读批量 RNA-Seq 分析。该平台包括两种实现方式:一个交互式的 Streamlit 应用,用户可上传数据并描述项目;以及一个非交互式 API,可集成到更大的智能体生态系统中,从而扩展到单细胞 RNA-Seq 和突变分析。我们的系统利用经过验证的最先进方法以确保可复现性,同时在 Python 中执行 PCA、差异表达和通路富集等分析。该平台生成可操作的报告,将表格和图表置于项目背景中进行解读。关键能力包括自动对比生成、协变量处理,以及对差异表达基因和富集通路的准确识别。验证研究表明,PCA 聚类、差异表达和通路评分与预期的生物学相符,并与手动流程的准确性相匹配。这种智能体方法减少了编码工作量,提高了可复现性,并使 RNA-Seq 分析的可及性普及化,同时具备扩展到多组学分析的可能性。
查看英文原文 English abstract
Automating bioinformatic analyses of RNA-seq data is challenging because each project requires unique combinations of analytical steps and frequent, case-specific adjustments. These project-specific processes limit the reusability of workflows to other analyses and require extensive manual coding. Large language model (LLM) agents are well-suited to address these challenges because they can interpret natural language instructions, dynamically plan workflows, and adapt to study-specific requirements without manual coding.We developed an agentic AI platform that uses LLMs to plan, execute, and interpret bulk RNA-Seq analyses via natural language instructions. The platform includes two implementations: an interactive Streamlit app where users can upload data and describe the project, and a non-interactive API for integration into larger agentic ecosystems to enable extension to single-cell RNA-Seq and mutation analyses. Our system leverages vetted, state-of-the-art methods to ensure reproducibility while performing analyses such as PCA, differential expression, and pathway enrichment in Python. The platform generates actionable reports that contextualize tables and figures in the project context. Key capabilities include automatic contrast generation, covariate handling, and accurate identification of differentially expressed genes and enriched pathways. Validation studies demonstrate that PCA clustering, differential expression and pathway scores align with expected biology and match manual pipeline accuracy.This agentic approach reduces coding effort, improves reproducibility, and democratizes the accessibility of RNA-Seq analysis, with the possibility of expanding multi-omics analyses.
利益披露 Disclosure
A. Liberzon,
AstraZeneca Employment.
P. Cingolani,
AstraZeneca Employment.
S. W. Criscione,
AstraZeneca Employment.
E. Jacob,
AstraZeneca Employment.