What does ChatGPT-3.5* know about neonatology? A new paper in @JAMAPeds with @swanbeams, @bhaweshiitk, Cindy Wang, Dara Brodksy, and @crmartin90 takes a look at this question.
We asked all 900+ questions from the well known Brodsky and Martin neonatal board prep book and had Drs. Brodsky and Martin evaluate the answer justifications ChatGPT provided.
On average, it got about 50% of the questions correct, with a lot of variation in accuracy by topic. It seemed well-aligned to the scientific consensus, but often omitted key details or facts in its answer justifications.
ChatGPT-3.5 does seem to know quite a bit about neonatology, but still falls short in important ways.
Paper: https://t.co/rH12m7QSSK
*Since AI progress moves faster than scientific publishing, GPT-4 came out while this paper was under review. We are in the process of updating the results with the new model - stay tuned!