JCCVS

Journal of Cardiology & Cardiovascular Surgery scientific, open-access, double-blind peer-reviewed journal covering a wide spectrum of topics in cardiology and cardiovascular surgery.

EndNote Style
Index
Original Article
Comprehensive comparison of Turkish hypertension patient education texts generated by large language models: a cross-sectional content analysis
Aims: This study aimed to compare Turkish patient education texts on hypertension generated by large language models in terms of quality, understandability, actionability, readability, guideline adherence, and patient safety.
Methods: In a cross-sectional descriptive content analysis design, standard questions covering eight core domains of hypertension patient education were posed to ChatGPT, Gemini, and Claude using the same Turkish prompt. For each model, only the first response from a separate new chat session was analyzed, yielding a total of 24 texts. The quality of responses containing treatment or self-management information was evaluated using DISCERN, while the understandability and actionability of all texts were assessed with PEMAT-P, and readability was evaluated using the Ateşman formula. In addition, adherence to the 2025 AHA/ACC hypertension guideline was scored, and patient safety was classified dichotomously. Friedman, Wilcoxon, and Cochran’s Q tests were used for statistical analysis.
Results: No significant difference was found among the models in total DISCERN scores (p=0.066). There was an overall difference in the DISCERN overall quality score (p=0.022), but no significant pairwise difference was identified in post hoc comparisons. PEMAT-P understandability scores differed significantly among the models (p=0.004); Gemini and Claude scored higher than ChatGPT, and the ChatGPT–Gemini difference remained significant in post hoc analysis (p=0.012). PEMAT-P actionability scores were similar across models (p=0.472). For Ateşman readability scores, Claude scored higher than both ChatGPT and Gemini (p=0.030; p=0.017 for both pairwise comparisons). No significant difference was found in guideline adherence (p=0.717) or patient safety (p=0.135).
Conclusion: Large language models may be useful for generating Turkish hypertension patient education materials; however, their performance varies according to the evaluation domain. While differences were observed among models in understandability and readability, no clear superiority was demonstrated in guideline adherence or patient safety. Therefore, LLM-generated outputs should be used under physician supervision, particularly with respect to clinical content and safety.


1. Mills KT, Stefanescu A, He J. The global epidemiology of hypertension. Nat Rev Nephrol. 2020;16(4):223-237. doi:10.1038/s41581-019-0244-2
2. WHO. Hypertension. 20.04.2026. https://www.who.int/news-room/fact-sheets/detail/hypertension
3. Jones D, Ferdinand K, Taler S, Johnson H, Shimbo D, Abdalla M. AHA/ACC/multi-society guideline for the prevention, detection, evaluation, and management of high blood pressure in adults: a report of the American College of Cardiology/American Heart Association Joint Committee on Clinical Practice Guidelines. J Am Coll Cardiol. 2025; 86(18):1567-1678. doi:10.1016/j.jacc.2025.05.007
4. McMullan M. Patients using the Internet to obtain health information: how this affects the patient-health professional relationship. Patient Educ Couns. 2006;63(1-2):24-28. doi:10.1016/j.pec.2005.10.006
5. Diviani N, Van Den Putte B, Giani S, van Weert JC. Low health literacy and evaluation of online health information: a systematic review of the literature. J Med Internet Res. 2015;17(5):e112. doi:10.2196/jmir.4018
6. Do Nascimento IJB, Pizarro AB, Almeida JM, et al. Infodemics and health misinformation: a systematic review of reviews. Bull World Health Organ. 2022;100(9):544-561. doi:10.2471/BLT.21.287654
7. Almagazzachi A, Mustafa A, Sedeh AE, et al. Generative Artificial Intelligence in patient education: ChatGPT takes on hypertension questions. Cureus. 2024;16(2):e53441. doi:10.7759/cureus.53441
8. OpenAI. Introducing ChatGPT. 20.04.2026. https://openai.com/tr-TR/index/chatgpt/
9. Walsh S. Timeline Of ChatGPT Updates & Key Events. 20.04.2026. https://www.searchenginejournal.com/history-of-chatgpt-timeline/488370/
10. Google. An important next step on our AI journey. 20.04.2026. https://blog.google/innovation-and-ai/technology/ai/bard-google-ai-search-updates/
11. Anthropic. Introducing Claude. 20.04.2026. https://www.anthropic.com/news/introducing-claude
12. Boztaş Demir G, Görgülü S. ChatGPT ve Gemini modellerinin ortodonti alanındaki Türkçe yeterliliği: hasta sorularına verilen yanıtların doğruluk, eksiksizlik ve okunabilirlik değerlendirilmesi. 7tepe Klinik Derg. 2025;21(3):151-158. doi:10.5505/yeditepe.2025.15010
13. Wójcik D, Adamiak O, Czerepak G, Tokarczuk O, Szalewski L. Comparing the performance of ChatGPT, Gemini, and Claude in English and Polish on medical examinations. Sci Rep. 2025;15(1):43849. doi:10.1038/s41598-025-31400-8
14. Liu L, Qu S, Zhao H, et al. Global trends and hotspots of ChatGPT in medical research: a bibliometric and visualized study. Front Med (Lausanne). 2024;11:1406842. doi:10.3389/fmed.2024.1406842
15. Peker RB. ChatGPT, ChatGPT Plus, Gemini ve Microsoft Copilot büyük dil modellerinin Türkiye’deki Diş Hekimliği Uzmanlık Eğitimi Giriş Sınavı’nda sorulan ağız, diş ve çene radyolojisi sorularını cevaplama performanslarının değerlendirilmesi. 7tepe Klinik Derg. 2025;21(3):130-135. doi:10.5505/yeditepe.2025.93265
16. Ruan T, Shao X, Sun Y, Ju X, Cui J. Evaluation of accuracy, quality, and readability of information on hypothyroidism provided by different artificial intelligence chatbot models. Front Public Health. 2025;13: 1698596. doi:10.3389/fpubh.2025.1698596
17. Naufal M, Kannan S, Senthilkumar S, et al. Patient education using AI: a cross-sectional study comparing ChatGPT and Google Gemini-generated patient education brochures on various surgical management of breast cancer. Cureus. 2025;17(11):e97525. doi:10.7759/cureus.97525
18. Liu Y, Li H, Ouyang J, et al. Evaluating large language models for preoperative patient education in superior capsular reconstruction: comparative study of Claude, GPT, and Gemini. JMIR Perioper Med. 2025;8:e70047. doi:10.2196/70047
19. Sivaramakrishnan G, Almuqahwi M, Ansari S, Lubbad M, Alagamawy E, Sridharan K. Assessing the power of AI: a comparative evaluation of large language models in generating patient education materials in dentistry. BDJ Open. 2025;11(1):59. doi:10.1038/s41405-025-00349-1
20. Will J, Gupta M, Zaretsky J, Dowlath A, Testa P, Feldman J. Enhancing the readability of online patient education materials using large language models: cross-sectional study. J Med Internet Res. 2025;27: e69955. doi:10.2196/69955
21. Shoemaker SJ, Wolf MS, Brach C. Development of the Patient Education Materials Assessment Tool (PEMAT): a new measure of understandability and actionability for print and audiovisual patient information. Patient Educ Couns. 2014;96(3):395-403. doi:10.1016/j.pec.2014.05.027
22. Zhan Y, Chen X, Ye F, et al. Evaluating AI chatbot responses to postkidney transplant inquiries. Elsevier; 2025:394-405.
23. Dihan QA, Brown AD, Zaldivar AT, et al. Advancing patient education in idiopathic intracranial hypertension: the promise of large language models. Neurol Clin Pract. 2025;15(1):e200366. doi:10.1212/CPJ.0000000000200366
24. Aydin S, Karabacak M, Vlachos V, Margetis K. Large language models in patient education: a scoping review of applications in medicine. Front Med (Lausanne). 2024;11:1477898. doi:10.3389/fmed.2024.1477898
25. AlSammarraie A, Househ M. The use of large language models in generating patient education materials: a scoping review. Acta Inform Med. 2025;33(1):4-10. doi:10.5455/aim.2024.33.4-10
26. Swisher AR, Wu AW, Liu GC, Lee MK, Carle TR, Tang DM. Enhancing health literacy: evaluating the readability of patient handouts revised by ChatGPT’s large language model. Otolaryngol Head Neck Surg. 2024; 171(6):1751-1757. doi:10.1002/ohn.927
27. Rosen KL, Sui M, Heydari K, Enichen EJ, Kvedar JC. The perils of politeness: how large language models may amplify medical misinformation. NPJ Digit Med. 2025;8(1):644. doi:10.1038/s41746-025-02135-7
28. Sharma S, Alaa AM, Daneshjou R. A longitudinal analysis of declining medical safety messaging in generative AI models. NPJ Digit Med. 2025;8(1):592. doi:10.1038/s41746-025-01943-1
29. Sismanoglu S, Capan BS. Performance of artificial intelligence on Turkish dental specialization exam: can ChatGPT-4.0 and gemini advanced achieve comparable results to humans? BMC Med Educ. 2025; 25(1):214. doi:10.1186/s12909-024-06389-9
30. Kinikoglu I. Evaluating ChatGPT and Google Gemini performance and implications in Turkish dental education. Cureus. 2025;17(1):e77292. doi:10.7759/cureus.77292
Volume 4, Issue 2, 2026
Page : 28-35
_Footer