Accuracy and Agreement of ChatGPT Based START Triage in the Emergency Department of Gunung Jati Hospital

Authors

  • Efa Dwina Program Studi Ilmu Keperawatan, Universitas Muhammadiyah Cirebon, Cirebon, Indonesia
  • Agil Putra Tri Kartika Program Studi Ilmu Keperawatan, Universitas Muhammadiyah Cirebon, Cirebon, Indonesia
  • Rizaluddin Akbar Program Studi Ilmu Keperawatan, Universitas Muhammadiyah Cirebon, Cirebon, Indonesia
  • Widuri Azmi Atmaji Program Studi Ilmu Keperawatan, Universitas Muhammadiyah Cirebon, Cirebon, Indonesia

DOI:

https://doi.org/10.36590/jika.v8i2.2042

Keywords:

artificial intelligence, triage, patient treatment

Abstract

Triage is an essential process in emergency care services aimed at determining patient treatment priorities based on the severity of clinical conditions. This study aimed to analyze the accuracy and level of agreement of Artificial Intelligence (AI), particularly ChatGPT, in triage classification at the Emergency Department of RSUD Gunung Jati compared with medical personnel as the reference standard. This study employed a quantitative observational approach with a retrospective cross-sectional design. Data were obtained from 420 adult patient medical records collected between March and May 2025. Triage classification was conducted using the Simple Triage and Rapid Treatment (START) method by both medical personnel and AI ChatGPT. Data analysis was performed using a confusion matrix and diagnostic tests, including sensitivity, specificity, accuracy, positive predictive value (PPV), negative predictive value (NPV), as well as agreement analysis using Cohen’s kappa and weighted Cohen’s kappa. The results showed that AI ChatGPT achieved an overall accuracy of 72.1% (95% CI: 67.8–76.4). The Cohen’s kappa value was 0.313, while the weighted Cohen’s kappa values were 0.338 using linear weighting and 0.384 using quadratic weighting, indicating a fair level of agreement. These findings suggested that ChatGPT had the potential as a supportive system in triage classification; however, it could not yet fully replace the role of medical personnel and still required further development, particularly in cases involving higher levels of emergency severity.

Downloads

Download data is not yet available.

Author Biographies

  • Efa Dwina, Program Studi Ilmu Keperawatan, Universitas Muhammadiyah Cirebon, Cirebon, Indonesia

    Program Studi Ilmu Keperawatan, Universitas Muhammadiyah Cirebon, Cirebon, Indonesia

  • Agil Putra Tri Kartika, Program Studi Ilmu Keperawatan, Universitas Muhammadiyah Cirebon, Cirebon, Indonesia

    Program Studi Ilmu Keperawatan, Universitas Muhammadiyah Cirebon, Cirebon, Indonesia

    https://scholar.google.com

  • Rizaluddin Akbar, Program Studi Ilmu Keperawatan, Universitas Muhammadiyah Cirebon, Cirebon, Indonesia

    Program Studi Ilmu Keperawatan, Universitas Muhammadiyah Cirebon, Cirebon, Indonesia

  • Widuri Azmi Atmaji, Program Studi Ilmu Keperawatan, Universitas Muhammadiyah Cirebon, Cirebon, Indonesia

    Program Studi Ilmu Keperawatan, Universitas Muhammadiyah Cirebon, Cirebon, Indonesia

References

Ayoub, M., Ballout, A.A., Zayek, R.A., Ayoub, N.F., 2023. Mind + Machine: ChatGPT as a Basic Clinical Decisions Support Tool. Cureus Journal of Medical Science 15(8), 8–11. Https://Doi.Org/10.7759/Cureus.43690

Chen, Q., Qin, Y., Jin, Z., Zhao, X., He, J., Wu, C., 2024. Enhancing Performance of The National Field Triage Guidelines using Machine Learning: Development of A Prehospital Triage Model to Predict Severe Trauma. Journal of Medical Internet Research 26, 1–15. Https://Doi.Org/10.2196/58740

Dettori, J.R., Norvell, D.C., 2020. Kappa and Beyond: Is There Agreement? Global Spine Journal 10(4), 499–501. Https://Doi.Org/10.1177/2192568220911648

Hanegraaf, P., Wondimu, A., Mosselman, J.J., Jong, R.D., Abogunrin, S., Queiros, L., Lane, M., Postma, M.J., 2024. Reviewer Reliability of Human Literature Reviewing and Implications Assisted for The Introduction of Machine Systematic Reviews: A Mixed Methods Review. BMJ Open 14(3), 1–10. Https://Doi.Org/10.1136/Bmjopen-2023-076912

Harrigan, M.E., Boremski, P.A., Collier, B.R., Tegge, A.N., Gillen, J.R., 2023. Impact of Nonphysician, Technology-Guided Alert Level Selection on Rates of Appropriate Trauma Triage in The United States: A Before and After Study. Journal of Trauma and Injury 36(3), 231–241. https://doi.org/10.20408/jti.2023.0020

Hicks, S.A., Strümke, I., Thambawita, V., Hammou, M., Riegler, M.A., Halvorsen, P., Parasa, S., 2022. On Evaluation Metrics for Medical Applications of Artificial Intelligence. Scientific Reports 12(5979), 1–9. Https://Doi.Org/10.1038/S41598-022-09954-8

Hoyer, A., Zapf, A., 2021. Studies for The Evaluation of Diagnostic Tests. Deutsches Arzteblatt 118, 555-560. Https://Doi.Org/10.3238/Arztebl.M2021.0224

Huabbangyang, T., Rojsaengroeng, R., Tiyawat, G., Silakoon, A., Vanichkulbodee, A., On, J.S., Buathong, S., 2023. Associated Factors of Under and Over-Triage Based on The Emergency Severity Index; A Retrospective Cross- Sectional Study. Archives of Academic Emergency Medicine 11(1), 1–11. https://doi.org/10.22037/aaem.v11i1.2076

Kartika, A.P.T., Akbar, R., Atmaji, W.A., Dwina, E., Gennie, A., Krisna, A.A., 2025. Triase Instalasi Gawat Daruat Berbasis Artificial Intelligence Berdasarkan Emergency Severity Index. Jurnal Ners 10(1), 2522–2529. https://journal.universitaspahlawan.ac.id/index.php/ners/article/view/54523

Lee, S., Jung, S., Park, J.H., Cho, H., Moon, S., Ahn, S., 2025. Performance of ChatGPT, Gemini and Deepseek for Non-Critical Triage Support Using Real-World Conversations in Emergency Department. BMC Emergency Medicine 25(1), 1-20. https://doi.org/10.1186/s12873-025-01337-2

Li, M., Gao, Q., Yu, T., 2023. Kappa Statistic Considerations in Evaluating Inter-Rater Reliability Between Two Raters: Which, When, and Context Matters. BMC Cancer 23(1), 1–5. https://doi.org/10.1186/s12885-023-11325-z

Lim, C., 2021. Methods for Evaluating The Accuracy of Diagnostic Tests. Cardiovascular Prevention and Pharmacotherapy 3(1), 15–20. https://e-jcpp.org/journal/view.php?doi=10.36011/cpp.2021.3.e2

Margaret, E., 2023. Evaluation of The Version 4 of the Emergency Severity Index in US Emergency Departments for The Rate of Mistriage. JAMA Netw Open 6(3), 1-19. Https://Doi.Org/10.1001/Jamanetworkopen.2023.3404

[NCHS] National Center for Health Statistics., 2022. National Hospital Ambulatory Medical Care Survey: 2022 Emergency Department Summary Tables. National Center for Health Statistics, Washington.

Pranoto, Y.A., Wibowo, S.A., 2020. Aplikasi Desktop Sistem Triase untuk Pendukung Prioritas Tingkat Kegawatan. Jurnal Ilmu Komputer dan Teknok Informatika 3(1), 1–6. https://ejournal.itn.ac.id/index.php/mnemonic/article/view/2319

Rainio, O., Teuho, J., Klen, R., 2024. Evaluation Metrics and Statistical Tests for Machine Learning. Scientific Reports 14(6086), 1–14. Https://Doi.Org/10.1038/S41598-024-56706-X

Sandmann, S., Riepenhausen, S., Plagwitz, L., Varghese, J., 2024. Systematic Analysis of ChatGPT, Google Search and Llama 2 for Clinical Decision Support Tasks. Nature Communications 15, 1–8. Https://Doi.Org/10.1038/S41467-024-46411-8

Sari, S.R., Fajarini, M., 2022. The Emergency Severity Indeks (ESI) Usage: Triage Accuracy and Causes Of Mistriage. Jurnal Ilmu Kesehatan 7(S1), 243–248. Https://Doi.Org/10.30604/Jika.V7is1.1190

Sax, D.R., Warton, E.M., Mark, D.G., Reed, M.E., 2025. Emergency Department Triage Accuracy and Delays in Care for High-Risk Conditions. JAMA Network Open 8(5), 1–12. Https://Doi.Org/10.1001/Jamanetworkopen.2025.8498

Uly, R.G.Z., Sriyono, S., Qona'ah, A., 2025. Triage in Emergency Management in the Emergency Departement: A Systematic Review. Indonesian Journal of Global Health Research 7(2), 527–534. https://jurnal.globalhealthsciencegroup.com/index.php/IJGHR/article/view/5607

Widodo, S,P., Riani, D.A., 2026. Literatur Desain Instalasi Gawat Darurat (IGD) pada Rumah Sakit di Bekasi. Jurnal Riset Ilmiah 3(2), 513–522. https://manggalajournal.org/index.php/SINERGI/article/view/2328

Zuhairini, R., Natasyah, W., Nurfadillah, W., Putri, I.F., Sapikah, N., 2025. Pemanfaatan Artificial Intelligence dalam Pengambilan Keputusan Klinis dan Manajemen Pasien Gawat Darurat di Bidang Anestesiologi: Sebuah Scoping Review. Jurnal Siti Rufaidah 3(4), 310-326. https://journal.ppniunimman.org/index.php/JASIRA/article/view/271

Downloads

Published

2026-08-04

How to Cite

Dwina, E., Kartika, A. P. T., Akbar, R., & Atmaji, W. A. (2026). Accuracy and Agreement of ChatGPT Based START Triage in the Emergency Department of Gunung Jati Hospital. Jurnal Ilmiah Kesehatan (JIKA), 8(2), 333-345. https://doi.org/10.36590/jika.v8i2.2042

Similar Articles

1-10 of 87

You may also start an advanced similarity search for this article.