Performance of ChatGPT, Bard, Claude, and Bing on the Peruvian National Licensing Medical Examination: a cross-sectional study.
Descripción del Articulo
Purpose We aimed to describe the performance and evaluate the educational value of justifications provided by artificial intelligence chatbots, including GPT-3.5, GPT-4, Bard, Claude, and Bing, on the Peruvian National Medical Licensing Examination (P-NLME). Methods This was a cross-sectional analyt...
| Autores: | , , , , , , , , , , |
|---|---|
| Formato: | artículo |
| Fecha de Publicación: | 2023 |
| Institución: | Universidad Nacional de Cajamarca |
| Repositorio: | UNC-Institucional |
| Lenguaje: | inglés |
| OAI Identifier: | oai:repositorio.unc.edu.pe:20.500.14074/10222 |
| Enlace del recurso: | http://hdl.handle.net/20.500.14074/10222 https://doi.org/10.3352/jeehp.2023.20.30 |
| Nivel de acceso: | acceso abierto |
| Materia: | Medical education Educational measurement Artificial intelligence Peru https://purl.org/pe-repo/ocde/ford#5.03.01 |
| id |
RUNC_27be5e6d6eeacdf86beafd2314841d6e |
|---|---|
| oai_identifier_str |
oai:repositorio.unc.edu.pe:20.500.14074/10222 |
| network_acronym_str |
RUNC |
| network_name_str |
UNC-Institucional |
| repository_id_str |
4868 |
| spelling |
Torres-Zegarra, B.C.Rios-Garcia, W.Ñaña-Cordova, A.M.Arteaga-Cisneros, K.F.Benavente-Chalco, X.C.Bustamante-Ordoñez, M.A.Gutierrez-Rios, C.J.Ramos-Godoy, C.A.Teresa Panta Quezada, K.L.Gutiérrez-Arratia, J.D.Flores-Cohaila, J.A.2026-03-11T17:32:14Z2026-03-11T17:32:14Z2023http://hdl.handle.net/20.500.14074/10222https://doi.org/10.3352/jeehp.2023.20.30Purpose We aimed to describe the performance and evaluate the educational value of justifications provided by artificial intelligence chatbots, including GPT-3.5, GPT-4, Bard, Claude, and Bing, on the Peruvian National Medical Licensing Examination (P-NLME). Methods This was a cross-sectional analytical study. On July 25, 2023, each multiple-choice question (MCQ) from the P-NLME was entered into each chatbot (GPT-3, GPT-4, Bing, Bard, and Claude) 3 times. Then, 4 medical educators categorized the MCQs in terms of medical area, item type, and whether the MCQ required Peru-specific knowledge. They assessed the educational value of the justifications from the 2 top performers (GPT-4 and Bing). Results GPT-4 scored 86.7% and Bing scored 82.2%, followed by Bard and Claude, and the historical performance of Peruvian examinees was 55%. Among the factors associated with correct answers, only MCQs that required Peru-specific knowledge had lower odds (odds ratio, 0.23; 95% confidence interval, 0.09–0.61), whereas the remaining factors showed no associations. In assessing the educational value of justifications provided by GPT-4 and Bing, neither showed any significant differences in certainty, usefulness, or potential use in the classroom. Conclusion Among chatbots, GPT-4 and Bing were the top performers, with Bing performing better at Peru-specific MCQs. Moreover, the educational value of justifications provided by the GPT-4 and Bing could be deemed appropriate. However, it is essential to start addressing the educational value of these chatbots, rather than merely their performance on examinations.Este trabajo fue financiado por UK Research and Innovation, UKRI, (105173).application/pdfengKorea Health Personnel Licensing Examination Institute.https://www.scopus.com/pages/publications/85177454993urn:issn:19755937J. Educ. Eval. Health Prof. 2023; 20: 30info:eu-repo/semantics/openAccesshttp://creativecommons.org/licenses/by/4.0/Medical educationEducational measurementArtificial intelligencePeruhttps://purl.org/pe-repo/ocde/ford#5.03.01Performance of ChatGPT, Bard, Claude, and Bing on the Peruvian National Licensing Medical Examination: a cross-sectional study.info:eu-repo/semantics/articleinfo:eu-repo/semantics/publishedVersionreponame:UNC-Institucionalinstname:Universidad Nacional de Cajamarcainstacron:UNCORIGINALjeehp-20-30.pdfjeehp-20-30.pdfapplication/pdf821572http://repositorio.unc.edu.pe/bitstream/20.500.14074/10222/1/jeehp-20-30.pdf310e5371b2d6dd92c20926bad77fb4d8MD5120.500.14074/10222oai:repositorio.unc.edu.pe:20.500.14074/102222026-03-12 10:43:50.524Universidad Nacional de Cajamarcarepositorio@unc.edu.pe |
| dc.title.es_PE.fl_str_mv |
Performance of ChatGPT, Bard, Claude, and Bing on the Peruvian National Licensing Medical Examination: a cross-sectional study. |
| title |
Performance of ChatGPT, Bard, Claude, and Bing on the Peruvian National Licensing Medical Examination: a cross-sectional study. |
| spellingShingle |
Performance of ChatGPT, Bard, Claude, and Bing on the Peruvian National Licensing Medical Examination: a cross-sectional study. Torres-Zegarra, B.C. Medical education Educational measurement Artificial intelligence Peru https://purl.org/pe-repo/ocde/ford#5.03.01 |
| title_short |
Performance of ChatGPT, Bard, Claude, and Bing on the Peruvian National Licensing Medical Examination: a cross-sectional study. |
| title_full |
Performance of ChatGPT, Bard, Claude, and Bing on the Peruvian National Licensing Medical Examination: a cross-sectional study. |
| title_fullStr |
Performance of ChatGPT, Bard, Claude, and Bing on the Peruvian National Licensing Medical Examination: a cross-sectional study. |
| title_full_unstemmed |
Performance of ChatGPT, Bard, Claude, and Bing on the Peruvian National Licensing Medical Examination: a cross-sectional study. |
| title_sort |
Performance of ChatGPT, Bard, Claude, and Bing on the Peruvian National Licensing Medical Examination: a cross-sectional study. |
| author |
Torres-Zegarra, B.C. |
| author_facet |
Torres-Zegarra, B.C. Rios-Garcia, W. Ñaña-Cordova, A.M. Arteaga-Cisneros, K.F. Benavente-Chalco, X.C. Bustamante-Ordoñez, M.A. Gutierrez-Rios, C.J. Ramos-Godoy, C.A. Teresa Panta Quezada, K.L. Gutiérrez-Arratia, J.D. Flores-Cohaila, J.A. |
| author_role |
author |
| author2 |
Rios-Garcia, W. Ñaña-Cordova, A.M. Arteaga-Cisneros, K.F. Benavente-Chalco, X.C. Bustamante-Ordoñez, M.A. Gutierrez-Rios, C.J. Ramos-Godoy, C.A. Teresa Panta Quezada, K.L. Gutiérrez-Arratia, J.D. Flores-Cohaila, J.A. |
| author2_role |
author author author author author author author author author author |
| dc.contributor.author.fl_str_mv |
Torres-Zegarra, B.C. Rios-Garcia, W. Ñaña-Cordova, A.M. Arteaga-Cisneros, K.F. Benavente-Chalco, X.C. Bustamante-Ordoñez, M.A. Gutierrez-Rios, C.J. Ramos-Godoy, C.A. Teresa Panta Quezada, K.L. Gutiérrez-Arratia, J.D. Flores-Cohaila, J.A. |
| dc.subject.es_PE.fl_str_mv |
Medical education Educational measurement Artificial intelligence Peru |
| topic |
Medical education Educational measurement Artificial intelligence Peru https://purl.org/pe-repo/ocde/ford#5.03.01 |
| dc.subject.ocde.es_PE.fl_str_mv |
https://purl.org/pe-repo/ocde/ford#5.03.01 |
| description |
Purpose We aimed to describe the performance and evaluate the educational value of justifications provided by artificial intelligence chatbots, including GPT-3.5, GPT-4, Bard, Claude, and Bing, on the Peruvian National Medical Licensing Examination (P-NLME). Methods This was a cross-sectional analytical study. On July 25, 2023, each multiple-choice question (MCQ) from the P-NLME was entered into each chatbot (GPT-3, GPT-4, Bing, Bard, and Claude) 3 times. Then, 4 medical educators categorized the MCQs in terms of medical area, item type, and whether the MCQ required Peru-specific knowledge. They assessed the educational value of the justifications from the 2 top performers (GPT-4 and Bing). Results GPT-4 scored 86.7% and Bing scored 82.2%, followed by Bard and Claude, and the historical performance of Peruvian examinees was 55%. Among the factors associated with correct answers, only MCQs that required Peru-specific knowledge had lower odds (odds ratio, 0.23; 95% confidence interval, 0.09–0.61), whereas the remaining factors showed no associations. In assessing the educational value of justifications provided by GPT-4 and Bing, neither showed any significant differences in certainty, usefulness, or potential use in the classroom. Conclusion Among chatbots, GPT-4 and Bing were the top performers, with Bing performing better at Peru-specific MCQs. Moreover, the educational value of justifications provided by the GPT-4 and Bing could be deemed appropriate. However, it is essential to start addressing the educational value of these chatbots, rather than merely their performance on examinations. |
| publishDate |
2023 |
| dc.date.accessioned.none.fl_str_mv |
2026-03-11T17:32:14Z |
| dc.date.available.none.fl_str_mv |
2026-03-11T17:32:14Z |
| dc.date.issued.fl_str_mv |
2023 |
| dc.type.es_PE.fl_str_mv |
info:eu-repo/semantics/article |
| dc.type.version.es_PE.fl_str_mv |
info:eu-repo/semantics/publishedVersion |
| format |
article |
| status_str |
publishedVersion |
| dc.identifier.uri.none.fl_str_mv |
http://hdl.handle.net/20.500.14074/10222 |
| dc.identifier.doi.es_PE.fl_str_mv |
https://doi.org/10.3352/jeehp.2023.20.30 |
| url |
http://hdl.handle.net/20.500.14074/10222 https://doi.org/10.3352/jeehp.2023.20.30 |
| dc.language.iso.es_PE.fl_str_mv |
eng |
| language |
eng |
| dc.relation.ispartof.es_PE.fl_str_mv |
https://www.scopus.com/pages/publications/85177454993 urn:issn:19755937 J. Educ. Eval. Health Prof. 2023; 20: 30 |
| dc.rights.es_PE.fl_str_mv |
info:eu-repo/semantics/openAccess |
| dc.rights.uri.es_PE.fl_str_mv |
http://creativecommons.org/licenses/by/4.0/ |
| eu_rights_str_mv |
openAccess |
| rights_invalid_str_mv |
http://creativecommons.org/licenses/by/4.0/ |
| dc.format.es_PE.fl_str_mv |
application/pdf |
| dc.publisher.es_PE.fl_str_mv |
Korea Health Personnel Licensing Examination Institute. |
| dc.source.none.fl_str_mv |
reponame:UNC-Institucional instname:Universidad Nacional de Cajamarca instacron:UNC |
| instname_str |
Universidad Nacional de Cajamarca |
| instacron_str |
UNC |
| institution |
UNC |
| reponame_str |
UNC-Institucional |
| collection |
UNC-Institucional |
| bitstream.url.fl_str_mv |
http://repositorio.unc.edu.pe/bitstream/20.500.14074/10222/1/jeehp-20-30.pdf |
| bitstream.checksum.fl_str_mv |
310e5371b2d6dd92c20926bad77fb4d8 |
| bitstream.checksumAlgorithm.fl_str_mv |
MD5 |
| repository.name.fl_str_mv |
Universidad Nacional de Cajamarca |
| repository.mail.fl_str_mv |
repositorio@unc.edu.pe |
| _version_ |
1864825001218146304 |
| score |
13.411838 |
Nota importante:
La información contenida en este registro es de entera responsabilidad de la institución que gestiona el repositorio institucional donde esta contenido este documento o set de datos. El CONCYTEC no se hace responsable por los contenidos (publicaciones y/o datos) accesibles a través del Repositorio Nacional Digital de Ciencia, Tecnología e Innovación de Acceso Abierto (ALICIA).
La información contenida en este registro es de entera responsabilidad de la institución que gestiona el repositorio institucional donde esta contenido este documento o set de datos. El CONCYTEC no se hace responsable por los contenidos (publicaciones y/o datos) accesibles a través del Repositorio Nacional Digital de Ciencia, Tecnología e Innovación de Acceso Abierto (ALICIA).