PO.BCS02.05 · 生物信息与计算
使用体模训练模型并应用于人体数据的Dyna-Q强化学习用于乳腺肿瘤恶性程度分类
Dyna-Q Reinforcement Learning for breast tumor malignancy classification using phantom-trained models applied to human data
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
背景:获取大规模、标注完善的人体数据集仍是生物医学AI进展的主要障碍之一。伦理限制、患者负担和机构审查流程限制了对临床样本的获取,使得训练具有泛化能力的机器学习模型变得困难。这一问题在组织病理学和肿瘤分类研究中尤为突出。基于Das等人(IEEE Sensors 2024)所展示的体模到人体迁移学习概念,我们提出了一种新的Dyna-Q强化学习框架,利用触觉感知来评估肿瘤恶性程度。该模型仅在体模数据上训练,并采用零样本知识迁移策略应用于未见过的人体数据集,使用可解释的力学特征作为合成域与生物域之间的桥梁。
方法:在本研究中,我们利用触觉感知系统生成与乳腺肿瘤力学特性(大小、深度和弹性)相关的图像。首先从乳腺体模采集触觉数据。使用这些数据,我们开发了一个用于肿瘤分类的深度Dyna-Q学习模型。一个深度神经网络作为决策智能体,以自监督方式训练以对肿瘤恶性程度进行分类。训练过程完全依赖于从乳腺肿瘤体模衍生的触觉特征,共使用9,000个平衡的体模数据样本。随后,我们从40名人体患者采集了触觉数据。经训练的深度Dyna-Q强化学习模型随后直接在未见过的人体数据上进行测试,以对乳腺肿瘤恶性程度进行分类,以病理结果作为真值。
结果:尽管体模模型与人体组织之间存在固有的域差异,我们的方法取得了有前景的性能:总体准确率为76.5%,敏感性为62.9%,特异性为73.5%。这些发现表明,即使没有直接的人体训练样本,也可以利用触觉特征跨域迁移具有临床意义的信号。通过消除对人体数据重新训练的需求,该方法大幅减少了对稀缺临床数据集的依赖,并为早期诊断建模提供了可扩展的解决方案。
结论:本研究凸显了在数据有限的生物医学环境中,基于触觉、强化学习的AI系统用于肿瘤分类的可行性。通过利用体模数据作为训练基质,该方法最大限度地减少了对稀缺临床样本的依赖,并为早期诊断建模提供了一个符合伦理、可扩展的框架。除乳腺癌外,这一范式还可扩展到其他有组织仿真体模可用的器官系统,为经济高效且可泛化的AI驱动诊断提供了一条切实可行的路径。
查看英文原文 English abstract
Background: Obtaining large, well-annotated human datasets remains one of the major barriers to progress in biomedical AI. Ethical restrictions, patient burden, and institutional review processes limit access to clinical samples, making it difficult to train generalizable machine learning models. This issue is particularly acute in histopathology and tumor classification studies. Building on the concept of phantom-to-human transfer learning demonstrated by Das et al. (IEEE Sensors 2024), we propose a new Dyna-Q reinforcement learning framework that leverages tactile sensing to assess tumor malignancy. The model is trained exclusively on phantom data and applies a zero-shot knowledge transfer strategy to unseen human datasets, using interpretable mechanical features as the bridge between synthetic and biological domains.
Methods: In this study, we utilized a Tactile Sensing System to generate images correlated with the mechanical properties (size, depth, and elasticity) of breast tumors. Tactile data were first collected from breast phantoms. Using these data, we developed a deep Dyna-Q learning model for tumor classification. A deep neural network served as the decision-making agent and was trained in a self-supervised manner to classify tumor malignancy. The training process relied exclusively on phantom-derived tactile features of breast tumors, using a total of 9,000 balanced phantom data samples. Subsequently, we collected tactile data from 40 human patients. The trained deep Dyna-Q reinforcement learning model was then tested directly on the unseen human data to classify breast tumor malignancy, with pathology results serving as the ground truth.
Results: Despite the inherent domain gap between phantom models and human tissue, our approach achieved promising performance: an overall accuracy of 76.5%, sensitivity of 62.9%, and specificity of 73.5%. These findings indicate that clinically meaningful signals can be transferred across domains using tactile features, even without direct human training samples. By eliminating the need for retraining human data, this method substantially reduces reliance on scarce clinical datasets and offers a scalable solution for early-stage diagnostic modeling.
Conclusions: This study highlights the feasibility of a tactile, reinforcement learning-based AI system for tumor classification in data-limited biomedical settings. By exploiting phantom data as a training substrate, the approach minimizes dependence on scarce clinical samples and provides an ethical, scalable framework for early diagnostic modeling. Beyond breast cancer, this paradigm could be extended to other organ systems where tissue-mimicking phantoms are available, offering a practical path toward cost-effective and generalizable AI-driven diagnostics.
利益披露 Disclosure
C. Won, None..
A. Das, None..
D. Caroline, None.