Assessing the Accuracy of Generative Conversational Artificial Intelligence in Debunking Sleep Health Myths: Mixed-Methods Comparative Study with Expert Analysis.

Nicola Luigi Bragazzi, Sergio Garbarino

JMIR Formative Research 2024 March 15

BACKGROUND: Adequate sleep is essential for maintaining both individual and public health, positively affecting cognition and well-being, and reducing chronic disease risks. It also plays a significant role in driving the economy, public safety, and managing healthcare costs. Digital tools, including websites, sleep trackers, and apps, are key in promoting sleep health education. Conversational Artificial Intelligence (AI) like ChatGPT offers accessible, personalized advice on sleep health but raises concerns about potential misinformation. This underscores the importance of ensuring AI-driven sleep health information is accurate, given its significant impact on individual and public health, and the spread of sleep-related myths.

OBJECTIVE: The study aims to examine ChatGPT's capability to debunk sleep-related disbeliefs.

METHODS: A mixed-methods design was leveraged. ChatGPT was asked to categorize twenty sleep-related false myths identified by ten sleep experts and to rate them in terms of falseness and public health significance, on a 5-point Likert scale. Sensitivity, positive predictive value, and inter-rater agreement were also calculated. A qualitative comparative analysis was also conducted.

RESULTS: ChatGPT labeled a significant portion (85%, n=17) of the statements as "false" (45%, n=9) or "generally false" (40%, n=8), with varying accuracy across different domains. For instance, it correctly identified most myths about "sleep timing", "sleep duration", and "behaviors during sleep", while it had varying degrees of success with other categories like "pre-sleep behaviors" and "brain function and sleep". ChatGPT's assessment of the degree of falseness and public health significance, on the 5-point Likert scale, showed an average score of 3.45 (SD=0.85) and 3.15 (SD=0.96), respectively, indicating a good level of accuracy in identifying the falseness of statements and a good understanding of their impact on public health. The AI-based tool showed a sensitivity of 85% and a perfect positive predictive value of 100%. Overall, this indicates that when ChatGPT labels a statement as false, it is highly reliable, but it may miss identifying some false statements. When comparing with expert ratings, high intra-class correlation coefficients (ICCs) between ChatGPT's appraisals and expert opinions could be found, suggesting that the AI's ratings were generally aligned with expert views on falseness (ICC=.83, P<.0001) and public health significance (ICC=.79, P=.001) of sleep-related myths. From a qualitative standpoint, both ChatGPT and sleep experts refuted sleep-related misconceptions. However, ChatGPT adopted a more accessible style and provided a more generalized, focusing on broad concepts, while experts sometimes used technical jargon, providing evidence-based explanations.

CONCLUSIONS: ChatGPT-4 can accurately address sleep-related queries and debunk sleep-related myths, with a performance comparable to sleep experts, even if, given its limitations, the AI cannot completely replace expert opinions, especially in nuanced and complex fields like sleep health, but can be a valuable complement in the dissemination of updated information and promotion of healthy behaviors.

Full text links

We have located links that may give you full text access.

Show additional links to paperHide additional links to paper

PubMed

Add to Saved Papers

Get 1-tap access

Related Resources

For the best experience, use the Read mobile app

Get seemless 1-tap access through your institution/university

For the best experience, use the Read mobile app

All material on this website is protected by copyright, Copyright © 1994-2024 by WebMD LLC.
This website also contains material copyrighted by 3rd parties.

By using this service, you agree to our terms of use and privacy policy.

Your Privacy Choices

You can now claim free CME credits for this literature searchClaim now

Get seemless 1-tap access through your institution/university

For the best experience, use the Read mobile app