PO.BCS02.04 · 生物信息与计算

用于全面肿瘤负荷评估的自动病灶分类的多中心开发:解决RECIST在试验与实践之间的脱节

Multi-site development of automated lesion classification for comprehensive tumor burden assessment: Addressing the RECIST trial-practice disconnect

编号 2788 展板 19 时间 4/20 02:00–05:00 区域 Section 4 主讲 Sean Khozin, PhD
分会场 Radiomics and AI in Medical Imaging
该海报暂无可下载的资料 AACR 官方页面

作者与单位 Authors & Affiliations

Ella Pavlechko1, Xi Jiang1, Ravikumar Komandur Elayavilli2, Eleanor McCabe2, Jon McDunn2, Sean Khozin3

1SAS Institure, Cary, NC,2Project Data Sphere, Cary, NC,3CEO Roundtable on Cancer & Project Data Sphere, Cary, NC

摘要 Abstract

中文摘要
背景:实体瘤疗效评价标准(RECIST)要求人工选择2-5个靶病灶并进行单维测量,产生超过30%的阅片者间不一致,同时将三维肿瘤动态压缩为生物学可解释性有限的分类结果。RECIST方案在常规实践中鲜少使用,造成试验与实践之间的脱节,削弱了终点的有效性。总肿瘤负荷的自动体积定量是一种替代方案,其前提是模型能在多样的影像环境中实现可靠的自主性能。 方法:我们开发了一个模块化的双架构系统,将UNet分割与ResNet50分类相结合,在来自三大洲(北美/南美与亚洲)1,324例患者的2,464次CT扫描上进行训练,这些扫描采集自异构的扫描仪平台(GE、Siemens、Philips、Toshiba)。该11,705个病灶的数据集具有临床代表性的类别分布:2,125个恶性(18%)、193个良性(2%)和9,387个其他发现(80%)。预处理采用Hounsfield加窗(窗位:-600,窗宽:1500)和Lungmask肺分割。为提高分类性能并解决类别不平衡问题,我们使用翻转、旋转和锐化对少数类图像进行增强。 结果:分割训练在各训练轮次中损失降低了86%,训练集和验证集同步收敛。在留出的开发数据(n=1,332个病灶)上的分类性能,在0.5概率阈值下取得了86%的准确率、91%的敏感性、17%的特异性、94%的阳性预测值(PPV)和11%的阴性预测值。精确率-召回率曲线下面积为0.89。模型性能在各扫描仪制造商之间保持稳定,无需针对特定平台进行重新校准。 结论:高敏感性(91%)与受限的特异性(17%)体现了针对不平衡数据集中恶性检测的成功优化。94%的PPV证实了可靠的恶性识别,而0.89的精确率-召回率AUC支持在11:1不平衡下的有效少数类判别。稳定的跨平台性能使多中心训练成为可用于泛化部署的可行范式。这项初步工作为全自主系统内的病灶级分类奠定了基础,以实现体积肿瘤负荷定量。我们的开发路径包括分类精修、扩展至跨器官系统的全身CT,以及自主检测和分割模块的整合。要将TTB确立为解决RECIST局限性的监管级终点,后续步骤需要在独立队列中进行外部验证、前瞻性临床试验证据以及与临床结局的相关性分析。
查看英文原文 English abstract
Background: Response Evaluation Criteria in Solid Tumors (RECIST) mandate manual selection of 2-5 target lesions with unidimensional measurements, generating inter-reader discordance beyond 30% while folding 3D tumor dynamics into categorical outcomes with limited biological interpretability. RECIST protocols are rarely used in routine practice, creating a trial-practice disconnect that undermines endpoint validity. Automated volumetric quantification of total tumor burden represents an alternative contingent on reliable autonomous model performance across diverse imaging environments. Methods: We developed a modular dual-architecture system combining UNet segmentation with ResNet50 classification, trained on 2,464 CT scans from 1,324 patients across three continents (North/South America & Asia) acquired on heterogeneous scanner platforms (GE, Siemens, Philips, Toshiba). The 11,705-lesion dataset had a clinically representative class distribution: 2,125 malignant (18%), 193 benign (2%), and 9,387 other findings (80%). Preprocessing applied Hounsfield windowing (level: -600, width: 1500) and Lungmask segmentation. To improve classification performance and address class imbalance, we augmented minority-class images using flips, rotations, and sharpening. Results: Segmentation training demonstrated loss reduction of 86% across epochs with parallel convergence in training and validation sets. Classification performance on held-out development data (n=1,332 lesions) yielded 86% accuracy, 91% sensitivity, 17% specificity, 94% positive predictive value (PPV), and 11% negative predictive value at 0.5 probability threshold. Precision-recall area under curve was 0.89. Model performance remained stable across scanner manufacturers without platform-specific recalibration. Conclusions: High sensitivity (91%) with constrained specificity (17%) exhibits successful optimization for malignancy detection in imbalanced datasets. The 94% PPV confirms reliable malignancy identification, while precision-recall AUC of 0.89 supports effective minority class discrimination despite 11:1 imbalance. Stable cross-platform performance allows for multi-site training to be a viable paradigm for generalized deployment. This initial work provides the foundation for lesion-level classification within a fully autonomous system for volumetric tumor burden quantification. Our development path includes classification refinement, expansion to whole-body CT across organ systems, and integration of autonomous detection and segmentation modules. The next steps in establishing TTB as a regulatory-grade endpoint addressing RECIST's limitations require external confirmation in independent cohorts, prospective clinical-trial evidence, and correlation with clinical outcomes.
利益披露 Disclosure
E. Pavlechko, None.. X. Jiang, None.. R. Komandur Elayavilli, None.. E. McCabe, None.. J. McDunn, None.. S. Khozin, None.

← 返回 AACR 2026 检索