PO.BCS02.02 · 生物信息与计算
在癌症沟通中患者更青睐ChatGPT而非医生的回应
Patients prefer ChatGPT to physician responses in cancer communication
作者与单位 Authors & Affiliations
摘要 Abstract
中文摘要
引言
像ChatGPT(GPT)这样的大语言模型正越来越多地被患者和医生使用,然而尚不清楚这两个群体中是否有任何一方更青睐GPT的输出而非医生撰写的内容。本研究评估了患者和医生如何看待并比较ChatGPT生成的建议与医生撰写的建议。
方法
我们调查了新墨西哥大学综合癌症中心的51名成年女性乳腺癌患者和15名医生,比较他们对GPT与医生撰写的针对四种涉及治疗、家庭关系和就业的癌症情景回应的评价。每个情景包含针对同一问题的两个盲法回应——一个来自GPT,一个来自医生。参与者在7分李克特量表上对每个回应的有用性、共情性和信息量进行评分,并指出他们偏好的回应(1=非常同意,7=非常不同意)。主要结局是偏好GPT的比例(相对50%的二项检验);次要结局是采用Wilcoxon符号秩检验分析的受试者内李克特差异(Δ = GPT - 医生)。
结果
在66名参与者中(51名患者;15名医生),患者偏好GPT的比例为71/104(68.3%,p = 0.00025),而医生偏好GPT的比例为31/52(59.6%,p = 0.21),两组间无显著差异(p = 0.82)。偏好因情景而异:患者在S1或S2中无显著差异(偏好GPT的比例为52.6%和58.6%;均p > 0.45),在S3中强烈偏好GPT(85.7%,p = 0.00018),在S4中显著偏好GPT(71.4%,p = 0.036)。医生对GPT的偏好在所有情景中均接近均势(36.4-75.0%,均p ≥ 0.15)。患者评分在有用性(Δ = −0.36,p = 0.009)、共情性(Δ = −0.46,p = 0.009)和信息量(Δ = −0.42,p = 0.013)上偏向GPT。医生评分在有用性(Δ = −0.10,p = 0.59)或信息量(Δ = +0.18,p = 0.57)上无显著差异,在共情性上有偏向GPT的趋势(Δ = −0.44,p = 0.058)。
结论
患者对GPT生成的建议表现出强烈的总体偏好,并存在情景特异性差异。医生表现出较小的、不显著的对GPT的偏好。虽然各维度的评分在数值上接近,但患者一致认为GPT的回应更有用、更具共情性和信息量,而医生对两个来源的评价相似。这些发现提示,医生可以将患者所偏好的GPT生成回应与自己的建议相结合,以改善临床沟通。
查看英文原文 English abstract
Introduction
Large language models like ChatGPT (GPT) are increasingly utilized by patients and physicians yet it is unknown whether either group prefers the output of GPT compared to physician-authored content. This study evaluated how patients and physicians perceive and compare ChatGPT-generated versus physician-authored recommendations.
Methods
We surveyed 51 adult female breast cancer patients and 15 physicians at the University of New Mexico Comprehensive Cancer Center to compare their evaluations of GPT versus physician-authored responses to four cancer scenarios related to treatment, family dynamics and employment. Each scenario included two blinded, responses to the same question-one from GPT and one from a physician. Participants rated each on 7-point Likert scales for helpfulness, empathy, and informativeness and indicated their preferred response (1 = strongly agree, 7 = strongly disagree). The primary outcome was the proportion preferring GPT (binomial test vs 50%); secondary outcomes were within-subject Likert differences (Δ = GPT - physician) analyzed with Wilcoxon signed-rank tests.
Results
Among 66 participants (51 patients; 15 physicians), patients preferred GPT 71/104 (68.3%, p = 0.00025), while physicians preferred GPT 31/52 (59.6%, p = 0.21), with no significant difference between groups (p = 0.82). Preferences varied by scenario: patients showed no significant differences in S1 or S2 (52.6% and 58.6% preferring GPT; both p > 0.45), strongly favored GPT in S3 (85.7%, p = 0.00018), and significantly favored GPT in S4 (71.4%, p = 0.036). Physician preference for GPT was near equipoise for all scenarios (36.4-75.0%, all p ≥ 0.15). Patient ratings favored GPT for helpfulness (Δ = −0.36, p = 0.009), empathy (Δ = −0.46, p = 0.009), and informativeness (Δ = −0.42, p = 0.013). Physician ratings showed no significant differences for helpfulness (Δ = −0.10, p = 0.59) or informativeness (Δ = +0.18, p = 0.57), with a trend toward higher empathy for GPT (Δ = −0.44, p = 0.058).
Conclusion
Patients showed a strong overall preference for GPT-generated recommendations, with scenario-specific differences. Physicians demonstrated a smaller, nonsignificant preference for GPT. While ratings were numerically close for all domains, patients consistently judged GPT responses as more helpful, empathetic, and informative, whereas physicians rated the two sources similarly. These findings suggest that physicians could combine GPT-generated responses preferred by patients with their own recommendations to improve clinical communication.
利益披露 Disclosure
A. Segura, None..
B. Tawfik, None..
B. Liem, None..
Z. Dayao, None..
D. Y. Lee, None..
J. M. Nemunaitis, None..
D. Savage, None..
M. Harari-Turquie, None..
N. Hill, None..
C. Foucar, None..
A. Tarnower, None..
D. Quintana, None..
J. Khatib, None..
T. Schroeder, None..
M. Mapalo, None..
Y. Sanchez, None..
R. Gopal, None.