PO.BCS02.01 · 生物信息与计算
Charles:一个用于癌症的自我批判式智能体AI药物发现分析师
Charles: A self-critical agentic AI drug discovery analyst for cancer
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
背景:智能体AI将成为未来科学团队的成员。AI幻觉以及不可靠、不可追溯的公共数据给这一角色蒙上了阴影。我们开发了据我们所知首个自我批判、自我纠正的智能体AI药物发现分析师:Charles。它在保持科学稳健性的同时,为癌症药物发现提供信息综合和假设生成支持。
方法:Charles是一个基于GPT的智能体LLM。构建Charles包括由人类首席研究员(PI)主导的、基于事实约束和模仿的教学训练。人类监督者精心优化环境,将Charles约束于可靠来源,防止其检索外部知识。一个对抗性AI智能体充当自我批判的“良知”,对照真实数据检验Charles回答的真实性。我们进一步通过注入诱饵原始数据来评估这一点,从而对模型的真实性进行定量评价。经训练的模型被用于综合生成基于事实的靶点摘要。随后,我们开发了一个多智能体LLM,利用这些以数据为依据的摘要来回答真实世界的药物发现问题。一个规划者智能体消化用户的问题,并协调制定行动计划以指挥专业智能体(生物/化学/系统等)。最后,一个批判性AI智能体对专业智能体的响应进行事实核查,并将其报告给规划者智能体以进一步优化行动。这些模块共同构成了Charles。
结果:我们证明Charles的响应严格基于经审校的数据。初步评估显示Charles复现了所有对抗性诱饵,表明无外部泄露。为进一步评估Charles的推理和事实准确性,我们利用一个问答数据集,Charles通过解读超过1000份蛋白质摘要来回答问题。借助自我批判智能体,我们对摘要和问题答案均进行了评估。对多智能体框架的初步评估显示Charles响应的准确率达99%。规划者成功招募了适当的专业智能体,而批判性智能体在向规划者检测语义和事实不一致方面表现出良好行为。输出摘要已通过canSAR.ai公开提供。
结论:Charles代表了一个自我批判、经药物发现训练的AI分析师,能够综合多模态信息、评估自身输出并据此优化后续步骤。通过将基于LLM交互的推理能力与经验证的审校数据的保障相结合,Charles在最大限度减少幻觉的同时提供了对复杂肿瘤学知识的可靠访问,使研究人员能够在癌症研究中更有信心地做出数据驱动的决策。
查看英文原文 English abstract
Background: Agentic AI will be members of scientific teams of the future. AI hallucinations and unreliable and untraceable public data cast shadows on such a role. We developed what is to our knowledge the first self-critical, self-correcting agentic-AI drug discovery analyst: Charles. It supports information synthesis and hypothesis generation for cancer drug discovery while maintaining scientific robustness.
Methods: Charles is an agentic GPT-based LLM. Building Charles consisted of fact-constrained, imitation-based teaching exercise led by the human PI. A human supervisor carefully optimized the environment to constrain Charles to reliable sources and prevent retrieving external knowledge. An adversarial AI agent acts as the self-critical ‘conscience' to test the veracity of Charles's answers against the true data. We further evaluated this by injecting decoy raw data, allowing quantitative evaluation of the models' veracity. The trained model was used to synthesize factual target summaries. We then developed a multi-agent LLM that leverages these data-grounded summaries to answer real-word drug discovery questions. A planner agent digests the user's questions and orchestrates a coordinated plan of action to direct the specialist agents (bio-/chemo-/systems etc). Finally, a critical AI agent fact-checks the responses from specialist agents, reporting them to the planner agent for further refinement of actions. Together, these modules form Charles.
Results: We demonstrate that Charles' responses are strictly grounded in the curated data. Initial evaluations show Charles reproduces all adversarial decoys indicating no outside leakage. To further evaluate the reasoning and factual accuracy of Charles, we utilized a question-answer dataset that Charles answers through interpretation of the >1000 protein summaries. Using the self-critical agent, we evaluated both the summaries and the answers to questions. The initial evaluation of the multi-agent framework shows 99% accuracy in Charles's responses. The planner successfully recruited the appropriate specialized agents while the critical agent showed promising behavior in detecting semantic and factual inconsistencies to the planner. The output summaries are publicly available via canSAR.ai.
Conclusions: Charles represents a self-critical, drug discovery-trained, AI analyst capable of synthesizing multimodal information, evaluating its own outputs, and refining the next steps based on that. By combining the reasoning power of LLM-based interactions with the assurance of verified curated data, Charles provides reliable access to complex oncological knowledge while minimizing hallucinations and enabling researchers to make data-driven decisions with more confidence in cancer research.
利益披露 Disclosure
S. Orouji, None..
Y. Zhu, None..
D. Maxwell, None..
K. Russell, None..
B. Al-Lazikani, None.