PO.BCS01.08 · 生物信息与计算

解读地图:一份对人类肿瘤图谱网络(Human Tumor Atlas Network)资源的邀约

Reading the map: An invitation to the resources of the Human Tumor Atlas Network

海报缩略图:解读地图:一份对人类肿瘤图谱网络(Human Tumor Atlas Network)资源的邀约
编号 4169 展板 19 时间 4/21 09:00–12:00 区域 Section 3 主讲 Adam Taylor, PhD
分会场 Digital Pathology 3
查看 PDF 下载 PDF 🔒 查看 / 下载完整 PDF 需登录并开通下载套餐 · 查看套餐 / 开通 AACR 官方页面

作者与单位 Authors & Affiliations

Adam Taylor1, Ashley Clayton1, Aditi Gopalan1, Milen Nikolov1, Thomas Yu1, David Gibbs2, Yamina Katariya2, Dar'ya Pozhidayeva1, Ino de Bruijn3, Seluck Onur Sumer3, Kristen Anton4, Jennifer Altreuter5, Alex Lash5, Ethan Cerami5, Nikolaus Schultz3, Vesteinn Thorsson2

1Sage Bionetworks, Seattle, WA,2Institute for Systems Biology, Seattle, WA,3Memorial Sloan Kettering Cancer Center, New York, NY,4Dartmouth Health, Lebanon, NH,5Dana-Farber Cancer Institute, Boston, MA

摘要 Abstract

中文摘要
目的:人类肿瘤图谱网络(HTAN)旨在跨越空间与时间绘制人类肿瘤的细胞和分子结构,以推进精准肿瘤学。HTAN数据协调中心(DCC)通过标准化、整合和分发多模态数据来支撑这一使命,确保其长期价值与影响力。 方法:DCC开发了一个可扩展的云生态系统(Synapse.org、Google BigQuery、定制数据门户),用于数据摄取、治理、验证和传播。一个符合NCI标准、由社区驱动的流程产生了一致的元数据模式(schema),确保了丰富的临床和生物标本信息以及所有数据模态之间的互操作性。HTAN数据门户为HTAN数据用户提供了统一的着陆页,具备基于筛选的搜索、影像和单细胞数据集的可视化以及详细文档。数据通过分层模型进行传播,影像和dbGaP受控访问测序数据可从NCI Cancer Research Data Commons General Commons获取,开放访问的已处理数据则可在Synapse.org获取。当前工作正过渡到模块化的LinkML数据模型,增加AI辅助的策展界面,并纳入BigQuery中精简的奖章架构(medallion architecture)以及门户增强功能。 结果:HTAN前五年(v7.0版本)产生了来自2,372个病例和11,378份生物标本的334 TB多模态数据(0.23M个文件),涵盖>60种疾病类型和>25种检测方法。HTAN每月支持>3,400名全球独立用户。最常访问的数据包括scRNA-seq(1075名下载用户)、多重成像(197名)和空间转录组学(419名)。对迄今为止dbGaP请求(N=127)的分析显示,学术界(75%)和产业界均有广泛使用,研究主题聚焦于基因组不稳定性(115项请求)、免疫逃逸(52项)和转移(49项),且常采用多组学和AI驱动的方法。 结论:HTAN已建立起一个被全球使用的、协调统一的基础,用于空间分辨的多模态癌症研究。这一持久的社区驱动发现平台建立在稳健的数据管理和透明的访问模型之上。下一阶段将深化空间和临床数据的整合,扩大AI就绪度,并强化对数据复用和可重复性的支持。我们诚邀研究者参与使用这一资源。 使用了ChatGPT 5、Gemini 2.5 Pro和Claude Sonnet 4.5来汇总使用指标、对dbGaP申请进行主题分析以及撰写摘要初稿。所有内容均经作者评估和批准。
查看英文原文 English abstract
Purpose: The Human Tumor Atlas Network (HTAN) sets out to map the cellular and molecular architecture of human tumors over space and time to advance precision oncology. The HTAN Data Coordinating Center (DCC) underpins this mission by standardizing, integrating, and distributing multimodal data; ensuring legacy and impact. Methods: The DCC developed a scalable cloud ecosystem (Synapse.org, Google BigQuery, custom data portal) for data ingestion, governance, validation, and dissemination. A community-driven process, aligned with NCI standards, produced a consistent metadata schema that ensures interoperability across rich clinical and biospecimen information and all data modalities. The HTAN data portal provides a single landing page for users of HTAN data featuring filter-based search, visualization of imaging and single-cell datasets and detailed documentation. Data are disseminated via a tiered model, with imaging and dbGaP-controlled access sequencing data available from NCI Cancer Research Data Commons General Commons and open-access processed data on Synapse.org. Current work transitions to a modular LinkML data model, adds an AI-assisted curation interface, and includes a streamlined medallion architecture in BigQuery and portal enhancements. Results: The first five years of HTAN (release v7.0) produced 334 TB of multimodal data (0.23M files) from 2,372 cases and 11,378 biospecimens, spanning >60 disease types and >25 assays. HTAN supports >3,400 unique global users a month. The most frequently accessed data includes scRNA-seq (1075 downloading users), multiplex imaging (197), and spatial transcriptomics (419). Analysis of dbGaP requests to-date (N=127) shows broad academic (75%) and industry use with research themes focused on genome instability (115 requests), immune evasion (52), and metastasis (49), often with multi-omic and AI-driven approaches. Conclusions: HTAN has established a globally utilized, harmonized foundation for spatially resolved, multimodal cancer research. This enduring platform for community-driven discovery is built on robust data management and transparent access models. The next phase deepens integration of spatial and clinical data, expanding AI-readiness, and strengthening support for reuse and reproducibility. We invite researchers to engage with this resource. ChatGPT 5, Gemini 2.5 Pro, & Claude Sonnet 4.5 were used to summarize usage metrics, conduct thematic analysis of dbGaP applications, and initial abstract drafting. All content was evaluated and approved by the authors.
利益披露 Disclosure
A. Taylor, None.. A. Clayton, None.. A. Gopalan, None.. M. Nikolov, None.. T. Yu, None.. D. Gibbs, None.. Y. Katariya, None.. D. Pozhidayeva, None.. I. de Bruijn, None.. S. Sumer, None.. K. Anton, None.. J. Altreuter, None.. A. Lash, None.. E. Cerami, None.. N. Schultz, None.. V. Thorsson, None.

← 返回 AACR 2026 检索