PO.BCS01.12 · 生物信息与计算
用于癌症药物筛选数据可重复分析的计算框架
A computational framework for reproducible analysis of cancer drug screening data
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
背景:在患者来源或模型来源的癌细胞中筛选治疗反应,为评估治疗疗效提供了一种直接的实验策略。这些实验产生复杂的数据集,需要可扩展、标准化的分析流程以确保可重复性和跨研究可比性。然而,目前的数据管理和分析工作流程仍然是碎片化的,依赖于人工整理和临时脚本,阻碍了可重复性。本研究引入了一个开源计算框架,用于标准化高通量筛选(HTS)数据的存储、分析和可视化,在临床和学术环境中提供对常用分析方法的便捷访问,同时保持可重复性。
方法:我们开发了一个Python框架,为存储、处理和分析高通量药物筛选数据提供了连贯的工作流程。它能够以跨实验和数据集一致的数据结构实现高效、统一的存储。该框架整合了自动药物名称标准化、标准化数据预处理,以及包括IC50、EC50和DSS计算在内的常用剂量-反应建模方法。此外,该框架支持队列水平的汇总和治疗优先级排序,允许用户跨患者比较药物反应。这些组件共同构建了一个可重复且适应性强的流程,将原始实验测量数据转化为可解释的生物学和转化医学见解。
结果:为展示其广泛的实用性,我们将该框架应用于三个使用案例。使用导入的GDSC2数据集,我们重新计算的IC50和AUC值与已发表数据高度一致,同时提供整合的可视化和更多情境性见解。我们重新分析了一个乳腺癌PDX来源的类器官数据集,能够评估批次效应并跨多个模型比较反应。我们还分析了一项正在进行的儿童脑肿瘤精准肿瘤学计划中的离体药物筛选数据。这些输出结果被用于在分子肿瘤委员会会议上协助为个体患者优先排序治疗候选方案,展示了其在转化医学情境中的适用性。在所有案例中,该框架确保了一致的预处理,最大程度减少了人工数据操作,并提供了全面的跨患者和跨药物比较,以识别个性化的候选药物。
结论:我们的框架为管理癌症药物筛选数据提供了全面、可互操作的解决方案。通过最大程度减少人工数据操作并实现可重复分析,它弥合了实验数据生成与可操作见解之间的鸿沟。除了个性化的离体实验外,该平台还支持跨公共和临床数据集的系统性分析,促进了转化研究和个性化治疗选择。
查看英文原文 English abstract
Background: Screening therapeutic responses in patient-derived or model-based cancer cells provides a direct experimental strategy to evaluate treatment efficacy. These assays generate complex datasets that require scalable, standardized analysis pipelines to ensure reproducibility and cross-study comparability. However, current data management and analysis workflows remain fragmented, relying on manual curation and ad hoc scripts that hinder reproducibility. This study introduces an open-source computational framework that standardizes storage, analysis, and visualization of high-throughput screening (HTS) data, providing easy access to common analysis methods while maintaining reproducibility in clinical and academic settings.
Methods: We developed a Python framework that provides a coherent workflow for storing, processing, and analyzing high-throughput drug screening data. It enables efficient and uniform storage with consistent data structure across experiments and datasets. The framework integrates automated drug name standardization, standardized data preprocessing, and common dose-response modeling methods including IC50, EC50, and DSS calculations. In addition, the framework supports cohort-level summarization and treatment prioritization, allowing users to compare drug responses across patients. Together, these components create a reproducible and adaptable pipeline that transfers raw experimental measurements to interpretable biological and translational insights.
Results: To demonstrate its broad utility, we applied the framework to three use cases. With the imported GDSC2 dataset, our recomputed IC50 and AUC values were highly consistent with published data, while providing integrated visualization and more contextual insights. We re-analyzed a breast cancer PDX-derived organoid dataset and were able to evaluate batch effects and compare responses across multiple models. We also analyzed an ex-vivo drug screening data from an ongoing pediatric brain tumor precision oncology initiative. The outputs were used to aid in prioritizing treatment candidates for individual patients at molecular tumor board meetings, demonstrating applicability in translational contexts. Across all cases, the framework ensured consistent preprocessing, minimized manual data manipulation, and provided comprehensive cross patient and cross drug comparison for identifying personalized drug candidates.
Conclusion: Our framework provides a comprehensive, interoperable solution for managing cancer drug screening data. By minimizing manual data manipulation and enabling reproducible analysis, it bridges the gap between experimental data generation and actionable insight. Beyond personalized ex-vivo assays, this platform empowers systematic analysis across public and clinical datasets, facilitating translational research and personalized treatment selection.
利益披露 Disclosure
H. Yang, None..
J. Lubkowitz, None..
X. Huang, None..
G. Marth, None..
S. Cheshier, None..
P. Moos, None..
Y. Qiao, None.