Performance of a large language model (ChatGPT-3.5) for Pooled Cohort Equation estimation of atherosclerotic cardiovascular disease risk — Ben J. Marafino (2023) | RDL Network
Abstract Despite demonstrated facility for arithmetic and other quantitative tasks, the performance of ChatGPT and other large language models for clinical risk calculation have yet to be assessed. Using synthetic patient data, this preliminary study aimed to assess the calibration, reproducibility, and potential for sociodemographic bias of ChatGPT-derived Pooled Cohort Equation (PCE) scores of atherosclerotic cardiovascular disease risk as compared to true scores. We found that ChatGPT-derived PCE scores, despite being moderately associated with the true PCE scores, displayed poor calibration with respect to true PCE scores, and exhibited instability between repeated rounds of prompting, suggesting lack of reproducibility. Moreover, ChatGPT-derived PCE scores also appeared inappropriately sensitive to contextual indicators of the sociodemographic status of the synthetic patients in this study. Further work is needed to confirm these results, and to assess performance on a wider variety of prompts as well as in other settings beyond cardiovascular disease prevention where accurate risk calculation is also vital to appropriate clinical decision-making. Abstract Figure Figure. Underestimation of true PCE risk estimates ( x -axis) by ChatGPT ( y -axis) on synthetic patient data.
L Fauchier, Arnaud Bisson, Alexandre Bodin, Julien Herbert, P Spiesser, Bertrand Pierre, Nicolas Clémenty, D Babuty, Anne Bernard, Professor Gregory Lip
Bernard Srour, Léopold Fezeu, Emmanuelle Kesse‐Guyot, Benjamin Allès, Eloi Chazelas, Mélanie Deschasaux, Serge Hercberg, Carlos Augusto Monteiro, C Julia, Mathilde Touvier
Discussion(0)
No comments yet. Be the first to comment.