PO.BCS02.06 · 生物信息与计算

利用AI技术从临床和文本特征预测肝细胞癌患者的免疫治疗反应

Predicting immunotherapy response in patients with hepatocellular carcinoma from clinical and textual features using AI techniques

海报缩略图:利用AI技术从临床和文本特征预测肝细胞癌患者的免疫治疗反应
编号 4219 展板 15 时间 4/21 09:00–12:00 区域 Section 5 主讲 Sola Adeleke
分会场 Machine Learning Approaches for Cancer Prediction
查看 PDF 下载 PDF 🔒 查看 / 下载完整 PDF 需登录并开通下载套餐 · 查看套餐 / 开通 AACR 官方页面

作者与单位 Authors & Affiliations

Anwaar Saeed1, Meghana Singh1, Yuming Shi1, Alireza Tojjari1, Vaishnavi Balaji2, Lakshya Sharma3, Azhar Saeed4, Thant Hoe2, Yuxi Zhang2, Sola Adeleke2

1Department of Medicine, Division of Hematology & Oncology, University of Pittsburgh Medical Cancer & UPMC Hillman Cancer Center, Pittsburgh, PA,2Curenetics Ltd, London, United Kingdom,3School of Medicine, University of St Andrews, Scotland, United Kingdom,4University of Vermont Medical Center, Colchester, VT

摘要 Abstract

中文摘要
背景: 免疫治疗(IO)改善了晚期肝细胞癌(HCC)的生存,但不到30%的患者对治疗有反应。现有生物标志物的预测准确性有限。机器学习(ML)和自然语言处理(NLP)技术可用于开发支持个体化治疗的预测模型。我们旨在开发和评估预测HCC患者IO反应的机器学习模型。 方法: 我们回顾性分析了2014年12月至2023年12月期间在UPMC Hillman Cancer Centre接受免疫治疗的302例HCC患者的数据。开发了五种机器学习模型来预测免疫治疗反应,包括逻辑回归、随机森林、XGBoost、支持向量机和多层感知机。模型最初使用20个临床特征进行训练,随后扩展为纳入50个临床特征与文本嵌入特征相结合的特征。放射学和临床记录通过自然语言处理(NLP)模型进行处理以生成文本嵌入。数据分为训练集(80%)和测试集(20%)。采用Shapley加性解释(SHAP)来解读预测模型。 结果: 在302例患者中,215例(71%)为疾病稳定,87例(29%)为疾病进展。表现最佳的模型是纳入临床特征和NLP衍生特征的随机森林分类器(AUC 0.77,精确率:0.72,召回率0.72)。当限制为仅使用临床变量时,模型性能略有下降(AUC 0.71,精确率:0.70,召回率0.70)。免疫治疗反应的关键预测因子包括较低的甲胎蛋白(AFP)、肝功能检查在正常范围内(AST、ALT、ALP、白蛋白、胆红素)、较高的总蛋白以及较低的ECOG体能状态分级。142例患者接受了一线IO治疗阿替利珠单抗联合贝伐珠单抗(Atezo/Bev),57例患者接受了度伐利尤单抗联合替西木单抗(Durva/Treme)。亚组分析显示,接受Atezo/Bev患者的模型性能(测试AUC-ROC 0.97)优于接受Durva/Treme的患者(测试AUC-ROC 0.66)。然而,两种IO方案之间的预测平均反应无统计学显著差异(Atezo/Bev 0.55,Durva/Treme 0.64,T统计量:-1.54,p值0.13)。 结论: 本研究表明,整合临床特征和NLP衍生特征的机器学习模型能够准确预测HCC患者的IO反应。疾病进展的关键预测因子包括AFP、肝功能血液检查和ECOG体能状态。未来工作将在更大数据集上对这些结果进行外部验证,旨在开发具有泛化能力且临床实用的预测模型。
查看英文原文 English abstract
Background: Immunotherapy (IO) improves survival in advanced hepatocellular carcinoma (HCC), yet under 30% of patients respond to treatment. Existing biomarkers have shown limited predictive accuracy. Machine learning (ML) and natural language processing (NLP) techniques could be used to develop prediction models that support personalised treatment. We aimed to develop and evaluate machine learning models that predict response to IO in patients with HCC. Methods: We retrospectively analyzed data from 302 patients with HCC treated with immunotherapy at UPMC Hillman Cancer Centre between December 2014 and December 2023. Five machine learning models were developed to predict immunotherapy response including logistic regression, random forest, XGBoost, support vector machine and multi-layer perceptron. Models were initially trained using 20 clinical features, then models were expanded to include 50 combined clinical and text-embedded features. Radiological and clinical notes were processed using a natural language processing (NLP) model to generate text embeddings. Data was split into training (80%) and test (20%) sets. Shapley Additive Explanations (SHAP) was used to interpret the prediction models. Results: Of the 302 patients, 215 (71%) had stable disease and 87 (29%) had progression. The best-performing model was the Random Forest classifier incorporating both clinical and NLP-derived features (AUC 0.77, Precision: 0.72, Recall 0.72). Model performance marginally decreased when restricted to clinical variables alone (AUC 0.71, Precision: 0.70, Recall 0.70). Key predictors of response to immunotherapy included lower alpha-fetoprotein (AFP), liver function tests within normal range (AST, ALT, ALP, albumin, bilirubin), higher total protein and lower grade of ECOG performance status. 142 patients had first-line IO treatment with atezolizumab and bevacizumab (Atezo/Bev) and 57 patients had durvalumab and tremelimumab (Durva/Treme). A subgroup analysis showed that model performance for patients receiving Atezo/Bev (test AUC-ROC 0.97) was superior to those receiving Durva/Treme (test AUC-ROC 0.66). However, there was no statistically significant difference in predicted mean response between the two IO regimens (Atezo/Bev 0.55, Durva/Treme 0.64, T-statistic: -1.54, p-value 0.13). Conclusions:​​ This study demonstrates that ML models integrating both clinical and NLP-derived features can accurately predict IO response in patients with HCC. Key predictors of disease progression included AFP, liver function blood tests and ECOG performance status. Future work will externally validate these results on larger datasets, with the aim of developing generalizable and clinically useful predictive models.
利益披露 Disclosure
A. Saeed, None.. M. Singh, None.. Y. Shi, None.. A. Tojjari, None. V. Balaji, Curenetics Ltd Employment. L. Sharma, None. T. Hoe, Curenetics Ltd Employment. Y. Zhang, Curenetics Ltd Employment. S. Adeleke, Curenetics Ltd Employment, g., Board of Directors, non-salaried role), Stock.

← 返回 AACR 2026 检索