Accuracy and Agreement of ChatGPT Based START Triage in the Emergency Department of Gunung Jati Hospital
DOI:
https://doi.org/10.36590/jika.v8i2.2042Keywords:
artificial intelligence, triage, patient treatmentAbstract
Triage is an essential process in emergency care services aimed at determining patient treatment priorities based on the severity of clinical conditions. This study aimed to analyze the accuracy and level of agreement of Artificial Intelligence (AI), particularly ChatGPT, in triage classification at the Emergency Department of RSUD Gunung Jati compared with medical personnel as the reference standard. This study employed a quantitative observational approach with a retrospective cross-sectional design. Data were obtained from 420 adult patient medical records collected between March and May 2025. Triage classification was conducted using the Simple Triage and Rapid Treatment (START) method by both medical personnel and AI ChatGPT. Data analysis was performed using a confusion matrix and diagnostic tests, including sensitivity, specificity, accuracy, positive predictive value (PPV), negative predictive value (NPV), as well as agreement analysis using Cohen’s kappa and weighted Cohen’s kappa. The results showed that AI ChatGPT achieved an overall accuracy of 72.1% (95% CI: 67.8–76.4). The Cohen’s kappa value was 0.313, while the weighted Cohen’s kappa values were 0.338 using linear weighting and 0.384 using quadratic weighting, indicating a fair level of agreement. These findings suggested that ChatGPT had the potential as a supportive system in triage classification; however, it could not yet fully replace the role of medical personnel and still required further development, particularly in cases involving higher levels of emergency severity.
Downloads
References
Ayoub, M., Ballout, A.A., Zayek, R.A., Ayoub, N.F., 2023. Mind + Machine: ChatGPT as a Basic Clinical Decisions Support Tool. Cureus Journal of Medical Science 15(8), 8–11. Https://Doi.Org/10.7759/Cureus.43690
Chen, Q., Qin, Y., Jin, Z., Zhao, X., He, J., Wu, C., 2024. Enhancing Performance of The National Field Triage Guidelines using Machine Learning: Development of A Prehospital Triage Model to Predict Severe Trauma. Journal of Medical Internet Research 26, 1–15. Https://Doi.Org/10.2196/58740
Dettori, J.R., Norvell, D.C., 2020. Kappa and Beyond: Is There Agreement? Global Spine Journal 10(4), 499–501. Https://Doi.Org/10.1177/2192568220911648
Hanegraaf, P., Wondimu, A., Mosselman, J.J., Jong, R.D., Abogunrin, S., Queiros, L., Lane, M., Postma, M.J., 2024. Reviewer Reliability of Human Literature Reviewing and Implications Assisted for The Introduction of Machine Systematic Reviews: A Mixed Methods Review. BMJ Open 14(3), 1–10. Https://Doi.Org/10.1136/Bmjopen-2023-076912
Harrigan, M.E., Boremski, P.A., Collier, B.R., Tegge, A.N., Gillen, J.R., 2023. Impact of Nonphysician, Technology-Guided Alert Level Selection on Rates of Appropriate Trauma Triage in The United States: A Before and After Study. Journal of Trauma and Injury 36(3), 231–241. https://doi.org/10.20408/jti.2023.0020
Hicks, S.A., Strümke, I., Thambawita, V., Hammou, M., Riegler, M.A., Halvorsen, P., Parasa, S., 2022. On Evaluation Metrics for Medical Applications of Artificial Intelligence. Scientific Reports 12(5979), 1–9. Https://Doi.Org/10.1038/S41598-022-09954-8
Hoyer, A., Zapf, A., 2021. Studies for The Evaluation of Diagnostic Tests. Deutsches Arzteblatt 118, 555-560. Https://Doi.Org/10.3238/Arztebl.M2021.0224
Huabbangyang, T., Rojsaengroeng, R., Tiyawat, G., Silakoon, A., Vanichkulbodee, A., On, J.S., Buathong, S., 2023. Associated Factors of Under and Over-Triage Based on The Emergency Severity Index; A Retrospective Cross- Sectional Study. Archives of Academic Emergency Medicine 11(1), 1–11. https://doi.org/10.22037/aaem.v11i1.2076
Kartika, A.P.T., Akbar, R., Atmaji, W.A., Dwina, E., Gennie, A., Krisna, A.A., 2025. Triase Instalasi Gawat Daruat Berbasis Artificial Intelligence Berdasarkan Emergency Severity Index. Jurnal Ners 10(1), 2522–2529. https://journal.universitaspahlawan.ac.id/index.php/ners/article/view/54523
Lee, S., Jung, S., Park, J.H., Cho, H., Moon, S., Ahn, S., 2025. Performance of ChatGPT, Gemini and Deepseek for Non-Critical Triage Support Using Real-World Conversations in Emergency Department. BMC Emergency Medicine 25(1), 1-20. https://doi.org/10.1186/s12873-025-01337-2
Li, M., Gao, Q., Yu, T., 2023. Kappa Statistic Considerations in Evaluating Inter-Rater Reliability Between Two Raters: Which, When, and Context Matters. BMC Cancer 23(1), 1–5. https://doi.org/10.1186/s12885-023-11325-z
Lim, C., 2021. Methods for Evaluating The Accuracy of Diagnostic Tests. Cardiovascular Prevention and Pharmacotherapy 3(1), 15–20. https://e-jcpp.org/journal/view.php?doi=10.36011/cpp.2021.3.e2
Margaret, E., 2023. Evaluation of The Version 4 of the Emergency Severity Index in US Emergency Departments for The Rate of Mistriage. JAMA Netw Open 6(3), 1-19. Https://Doi.Org/10.1001/Jamanetworkopen.2023.3404
[NCHS] National Center for Health Statistics., 2022. National Hospital Ambulatory Medical Care Survey: 2022 Emergency Department Summary Tables. National Center for Health Statistics, Washington.
Pranoto, Y.A., Wibowo, S.A., 2020. Aplikasi Desktop Sistem Triase untuk Pendukung Prioritas Tingkat Kegawatan. Jurnal Ilmu Komputer dan Teknok Informatika 3(1), 1–6. https://ejournal.itn.ac.id/index.php/mnemonic/article/view/2319
Rainio, O., Teuho, J., Klen, R., 2024. Evaluation Metrics and Statistical Tests for Machine Learning. Scientific Reports 14(6086), 1–14. Https://Doi.Org/10.1038/S41598-024-56706-X
Sandmann, S., Riepenhausen, S., Plagwitz, L., Varghese, J., 2024. Systematic Analysis of ChatGPT, Google Search and Llama 2 for Clinical Decision Support Tasks. Nature Communications 15, 1–8. Https://Doi.Org/10.1038/S41467-024-46411-8
Sari, S.R., Fajarini, M., 2022. The Emergency Severity Indeks (ESI) Usage: Triage Accuracy and Causes Of Mistriage. Jurnal Ilmu Kesehatan 7(S1), 243–248. Https://Doi.Org/10.30604/Jika.V7is1.1190
Sax, D.R., Warton, E.M., Mark, D.G., Reed, M.E., 2025. Emergency Department Triage Accuracy and Delays in Care for High-Risk Conditions. JAMA Network Open 8(5), 1–12. Https://Doi.Org/10.1001/Jamanetworkopen.2025.8498
Uly, R.G.Z., Sriyono, S., Qona'ah, A., 2025. Triage in Emergency Management in the Emergency Departement: A Systematic Review. Indonesian Journal of Global Health Research 7(2), 527–534. https://jurnal.globalhealthsciencegroup.com/index.php/IJGHR/article/view/5607
Widodo, S,P., Riani, D.A., 2026. Literatur Desain Instalasi Gawat Darurat (IGD) pada Rumah Sakit di Bekasi. Jurnal Riset Ilmiah 3(2), 513–522. https://manggalajournal.org/index.php/SINERGI/article/view/2328
Zuhairini, R., Natasyah, W., Nurfadillah, W., Putri, I.F., Sapikah, N., 2025. Pemanfaatan Artificial Intelligence dalam Pengambilan Keputusan Klinis dan Manajemen Pasien Gawat Darurat di Bidang Anestesiologi: Sebuah Scoping Review. Jurnal Siti Rufaidah 3(4), 310-326. https://journal.ppniunimman.org/index.php/JASIRA/article/view/271
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Efa Dwina, Agil Putra Tri Kartika, Rizaluddin Akbar, Widuri Azmi Atmaji

This work is licensed under a Creative Commons Attribution 4.0 International License.

