PO.BCS01.05 · 生物信息与计算

Omicology:一个基于AI/LLM/NLP的综合性网络资源,用于组学文献的本体、系统发生学及实用导航

Omicology: A comprehensive AI/LLM/NLP-based web resource for the ontology, phylogeny, and practical navigation of omics literature

海报缩略图:Omicology:一个基于AI/LLM/NLP的综合性网络资源,用于组学文献的本体、系统发生学及实用导航
编号 5451 展板 18 时间 4/21 02:00–05:00 区域 Section 1 主讲 John Weinstein, MD;PhD
分会场 Application of Bioinformatics to Cancer Biology 5
查看 PDF 下载 PDF 🔒 查看 / 下载完整 PDF 需登录并开通下载套餐 · 查看套餐 / 开通 AACR 官方页面

作者与单位 Authors & Affiliations

James M. Melott, John N. Weinstein

UT MD Anderson Cancer Center, Houston, TX

摘要 Abstract

中文摘要
引言:我们在此介绍一个公共网站omicology.com,它包含关于组学研究领域的丰富信息,并具备专门的组学搜索引擎。组学(Omics)最初只是一个简单的后缀,用于基于基因、蛋白质及其他(生物)分子整体分子图谱分析的研究。其词源来自古希腊语或可能是梵语,尚存争议。自1990年代起,该后缀被应用于数十个研究领域(Weinstein 1998),常被讥讽为行话。"假设在哪里?"是一个常见的批评——也是拒绝稿件或课题申请的理由。但鉴于2001年人类基因组草图的发布,组学研究与假设驱动研究之间的协同作用逐渐得到认可。在2000、2010、2020和2025年,组学术语分别有6、61、267和超过700个。近年来,高通量单细胞、空间分辨和时间分辨的组学技术为我们追求癌症精准医学提供了新的维度。 方法:我们开发搜索引擎数据库的起点是6,980,243篇全文开放获取的PubMed Central文章及各类元数据,其中包含5,699个组学术语。通过我们的NLP、LLM资源的运用以及细致的人工审编,产出了702个组学术语。此类列表的精炼因语言的特殊性、拼写错误、异体拼写、诸如'nomics'(如economics中)而非'omics'的后缀,以及'chromosomics'等歧义术语而变得复杂。 AI工具(Elicit、Cursor、Google和ChatGPT)为一个静态html原型网站提供了有用的设计和内容思路。一个具备自主能力(agentic)的AI编码工具协助我们基于该原型开发了一个动态的全栈、人工审编版本(目前有34,124行代码和内容)(将于AACR 2026之前公开发布)。 代表性结果:最常见的组学术语是基因组学(Genomics,257,617篇文章)、蛋白质组学(Proteomics,123,697篇)、代谢组学(Metabolomics,80,023篇)和转录组学(Transcriptomics,78,279篇)。宏基因组学、影像组学、脂质组学、表观基因组学、药物基因组学和磷酸化蛋白质组学位列前十。 我们关于前702个术语的数据目前包括每个文档章节的使用数量、发表日期、审编状态、LLM生成的简要描述等。 结论:这个开源、可更新的Omicology.com网站将提供不断扩充的组学信息与视角,以及专门的组学搜索能力。生物医学视角将得益于资深作者在组学研究领域的长期经验,他发起并领导了首个组学/多组学NCI项目——NCI-60的分子图谱分析(如Weinstein等,Science,1997)。此处用于组学的方法可为其他难以分析的领域的研究提供模板,补充诸如PubMed Central等资源的能力。
查看英文原文 English abstract
Introduction: We present here a public website, omicology.com, which includes extensive information on omic research domains and features a specialized omics search engine. Omics started as a simple suffix for research based on molecular profiling of genes, proteins, and other (biological) molecules in aggregate. The etymology, from ancient Greek or possibly Sanskrit, is debated. Application of the suffix to dozens of research fields starting in the 1990s (Weinstein 1998) was often derided as jargon. “Where's the hypothesis?” was a common critique - and reason for rejecting manuscripts or proposals. But given the draft human genome in 2001, synergy between omic and hypothesis-driven research gradually gained acceptance. In 2000, 2010, 2020, and 2025, there were 6, 61, 267, and >700 omics terms, respectively. More recently, high-throughput single-cell, spatially-resolved, and temporally-resolved omic technologies have provided new dimensions to our pursuit of precision medicine for cancer. Methods: Our starting point for development of the search engine database was 6,980,243 full-text Open Access PubMed Central articles plus various metadata. It included 5,699 omics terms. Our NLP, use of LLM resources, and careful manual curation have produced 702 omics terms. Refinement of such lists is complicated by eccentricities of language, misspellings, alternative spellings, suffixes like ‘nomics' (e.g., in economics), not ‘omics,' and ambiguous terms like ‘chromosomics.' AI tools (Elicit, Cursor, Google, and ChatGPT) have provided useful design and content ideas for a static html prototype website. An agentic AI coding tool has assisted our development of a dynamic full-stack, human-curated version (currently 34,124 lines of code and content) based on the prototype (for publicly roll-out before AACR 2026). Representative Results: The most frequent omic terms are Genomics (257,617 articles), Proteomics (123,697), Metabolomics, (80,023), and Transcriptomics (78,279). Metagenomics, radiomics, lipidomics, epigenomics, pharmacogenomics, and phosphoproteomics complete the top 10. Our data on the top 702 terms currently include usage numbers for each document section, publication date, curation status, a brief LLM-generated description, and more. Conclusions: The open-source, updatable Omicology.com website will provide an expanding repertoire of information and perspectives on omics plus specialized omics search capabilities. Biomedical perspective will be aided by the senior author's long-term experience in omic research beginning with his initiation and leadership of the first omic/multi-omic NCI project, molecular profiling of the NCI-60 (e.g., Weinstein, et al., Science, 1997). Methods used here for omics can provide a template for research on other hard-to-analyze fields, complementing the capabilities of such resources as PubMed Central.
利益披露 Disclosure
J. M. Melott, None.. J. N. Weinstein, None.

← 返回 AACR 2026 检索