Performance of ChatGPT, Bard, Claude, and Bing on the Peruvian National Licensing Medical Examination: a cross-sectional study.

Descripción del Articulo

Purpose We aimed to describe the performance and evaluate the educational value of justifications provided by artificial intelligence chatbots, including GPT-3.5, GPT-4, Bard, Claude, and Bing, on the Peruvian National Medical Licensing Examination (P-NLME). Methods This was a cross-sectional analyt...

Descripción completa

Detalles Bibliográficos
Autores: Torres-Zegarra, B.C., Rios-Garcia, W., Ñaña-Cordova, A.M., Arteaga-Cisneros, K.F., Benavente-Chalco, X.C., Bustamante-Ordoñez, M.A., Gutierrez-Rios, C.J., Ramos-Godoy, C.A., Teresa Panta Quezada, K.L., Gutiérrez-Arratia, J.D., Flores-Cohaila, J.A.
Formato: artículo
Fecha de Publicación:2023
Institución:Universidad Nacional de Cajamarca
Repositorio:UNC-Institucional
Lenguaje:inglés
OAI Identifier:oai:repositorio.unc.edu.pe:20.500.14074/10222
Enlace del recurso:http://hdl.handle.net/20.500.14074/10222
https://doi.org/10.3352/jeehp.2023.20.30
Nivel de acceso:acceso abierto
Materia:Medical education
Educational measurement
Artificial intelligence
Peru
https://purl.org/pe-repo/ocde/ford#5.03.01
id RUNC_27be5e6d6eeacdf86beafd2314841d6e
oai_identifier_str oai:repositorio.unc.edu.pe:20.500.14074/10222
network_acronym_str RUNC
network_name_str UNC-Institucional
repository_id_str 4868
spelling Torres-Zegarra, B.C.Rios-Garcia, W.Ñaña-Cordova, A.M.Arteaga-Cisneros, K.F.Benavente-Chalco, X.C.Bustamante-Ordoñez, M.A.Gutierrez-Rios, C.J.Ramos-Godoy, C.A.Teresa Panta Quezada, K.L.Gutiérrez-Arratia, J.D.Flores-Cohaila, J.A.2026-03-11T17:32:14Z2026-03-11T17:32:14Z2023http://hdl.handle.net/20.500.14074/10222https://doi.org/10.3352/jeehp.2023.20.30Purpose We aimed to describe the performance and evaluate the educational value of justifications provided by artificial intelligence chatbots, including GPT-3.5, GPT-4, Bard, Claude, and Bing, on the Peruvian National Medical Licensing Examination (P-NLME). Methods This was a cross-sectional analytical study. On July 25, 2023, each multiple-choice question (MCQ) from the P-NLME was entered into each chatbot (GPT-3, GPT-4, Bing, Bard, and Claude) 3 times. Then, 4 medical educators categorized the MCQs in terms of medical area, item type, and whether the MCQ required Peru-specific knowledge. They assessed the educational value of the justifications from the 2 top performers (GPT-4 and Bing). Results GPT-4 scored 86.7% and Bing scored 82.2%, followed by Bard and Claude, and the historical performance of Peruvian examinees was 55%. Among the factors associated with correct answers, only MCQs that required Peru-specific knowledge had lower odds (odds ratio, 0.23; 95% confidence interval, 0.09–0.61), whereas the remaining factors showed no associations. In assessing the educational value of justifications provided by GPT-4 and Bing, neither showed any significant differences in certainty, usefulness, or potential use in the classroom. Conclusion Among chatbots, GPT-4 and Bing were the top performers, with Bing performing better at Peru-specific MCQs. Moreover, the educational value of justifications provided by the GPT-4 and Bing could be deemed appropriate. However, it is essential to start addressing the educational value of these chatbots, rather than merely their performance on examinations.Este trabajo fue financiado por UK Research and Innovation, UKRI, (105173).application/pdfengKorea Health Personnel Licensing Examination Institute.https://www.scopus.com/pages/publications/85177454993urn:issn:19755937J. Educ. Eval. Health Prof. 2023; 20: 30info:eu-repo/semantics/openAccesshttp://creativecommons.org/licenses/by/4.0/Medical educationEducational measurementArtificial intelligencePeruhttps://purl.org/pe-repo/ocde/ford#5.03.01Performance of ChatGPT, Bard, Claude, and Bing on the Peruvian National Licensing Medical Examination: a cross-sectional study.info:eu-repo/semantics/articleinfo:eu-repo/semantics/publishedVersionreponame:UNC-Institucionalinstname:Universidad Nacional de Cajamarcainstacron:UNCORIGINALjeehp-20-30.pdfjeehp-20-30.pdfapplication/pdf821572http://repositorio.unc.edu.pe/bitstream/20.500.14074/10222/1/jeehp-20-30.pdf310e5371b2d6dd92c20926bad77fb4d8MD5120.500.14074/10222oai:repositorio.unc.edu.pe:20.500.14074/102222026-03-12 10:43:50.524Universidad Nacional de Cajamarcarepositorio@unc.edu.pe
dc.title.es_PE.fl_str_mv Performance of ChatGPT, Bard, Claude, and Bing on the Peruvian National Licensing Medical Examination: a cross-sectional study.
title Performance of ChatGPT, Bard, Claude, and Bing on the Peruvian National Licensing Medical Examination: a cross-sectional study.
spellingShingle Performance of ChatGPT, Bard, Claude, and Bing on the Peruvian National Licensing Medical Examination: a cross-sectional study.
Torres-Zegarra, B.C.
Medical education
Educational measurement
Artificial intelligence
Peru
https://purl.org/pe-repo/ocde/ford#5.03.01
title_short Performance of ChatGPT, Bard, Claude, and Bing on the Peruvian National Licensing Medical Examination: a cross-sectional study.
title_full Performance of ChatGPT, Bard, Claude, and Bing on the Peruvian National Licensing Medical Examination: a cross-sectional study.
title_fullStr Performance of ChatGPT, Bard, Claude, and Bing on the Peruvian National Licensing Medical Examination: a cross-sectional study.
title_full_unstemmed Performance of ChatGPT, Bard, Claude, and Bing on the Peruvian National Licensing Medical Examination: a cross-sectional study.
title_sort Performance of ChatGPT, Bard, Claude, and Bing on the Peruvian National Licensing Medical Examination: a cross-sectional study.
author Torres-Zegarra, B.C.
author_facet Torres-Zegarra, B.C.
Rios-Garcia, W.
Ñaña-Cordova, A.M.
Arteaga-Cisneros, K.F.
Benavente-Chalco, X.C.
Bustamante-Ordoñez, M.A.
Gutierrez-Rios, C.J.
Ramos-Godoy, C.A.
Teresa Panta Quezada, K.L.
Gutiérrez-Arratia, J.D.
Flores-Cohaila, J.A.
author_role author
author2 Rios-Garcia, W.
Ñaña-Cordova, A.M.
Arteaga-Cisneros, K.F.
Benavente-Chalco, X.C.
Bustamante-Ordoñez, M.A.
Gutierrez-Rios, C.J.
Ramos-Godoy, C.A.
Teresa Panta Quezada, K.L.
Gutiérrez-Arratia, J.D.
Flores-Cohaila, J.A.
author2_role author
author
author
author
author
author
author
author
author
author
dc.contributor.author.fl_str_mv Torres-Zegarra, B.C.
Rios-Garcia, W.
Ñaña-Cordova, A.M.
Arteaga-Cisneros, K.F.
Benavente-Chalco, X.C.
Bustamante-Ordoñez, M.A.
Gutierrez-Rios, C.J.
Ramos-Godoy, C.A.
Teresa Panta Quezada, K.L.
Gutiérrez-Arratia, J.D.
Flores-Cohaila, J.A.
dc.subject.es_PE.fl_str_mv Medical education
Educational measurement
Artificial intelligence
Peru
topic Medical education
Educational measurement
Artificial intelligence
Peru
https://purl.org/pe-repo/ocde/ford#5.03.01
dc.subject.ocde.es_PE.fl_str_mv https://purl.org/pe-repo/ocde/ford#5.03.01
description Purpose We aimed to describe the performance and evaluate the educational value of justifications provided by artificial intelligence chatbots, including GPT-3.5, GPT-4, Bard, Claude, and Bing, on the Peruvian National Medical Licensing Examination (P-NLME). Methods This was a cross-sectional analytical study. On July 25, 2023, each multiple-choice question (MCQ) from the P-NLME was entered into each chatbot (GPT-3, GPT-4, Bing, Bard, and Claude) 3 times. Then, 4 medical educators categorized the MCQs in terms of medical area, item type, and whether the MCQ required Peru-specific knowledge. They assessed the educational value of the justifications from the 2 top performers (GPT-4 and Bing). Results GPT-4 scored 86.7% and Bing scored 82.2%, followed by Bard and Claude, and the historical performance of Peruvian examinees was 55%. Among the factors associated with correct answers, only MCQs that required Peru-specific knowledge had lower odds (odds ratio, 0.23; 95% confidence interval, 0.09–0.61), whereas the remaining factors showed no associations. In assessing the educational value of justifications provided by GPT-4 and Bing, neither showed any significant differences in certainty, usefulness, or potential use in the classroom. Conclusion Among chatbots, GPT-4 and Bing were the top performers, with Bing performing better at Peru-specific MCQs. Moreover, the educational value of justifications provided by the GPT-4 and Bing could be deemed appropriate. However, it is essential to start addressing the educational value of these chatbots, rather than merely their performance on examinations.
publishDate 2023
dc.date.accessioned.none.fl_str_mv 2026-03-11T17:32:14Z
dc.date.available.none.fl_str_mv 2026-03-11T17:32:14Z
dc.date.issued.fl_str_mv 2023
dc.type.es_PE.fl_str_mv info:eu-repo/semantics/article
dc.type.version.es_PE.fl_str_mv info:eu-repo/semantics/publishedVersion
format article
status_str publishedVersion
dc.identifier.uri.none.fl_str_mv http://hdl.handle.net/20.500.14074/10222
dc.identifier.doi.es_PE.fl_str_mv https://doi.org/10.3352/jeehp.2023.20.30
url http://hdl.handle.net/20.500.14074/10222
https://doi.org/10.3352/jeehp.2023.20.30
dc.language.iso.es_PE.fl_str_mv eng
language eng
dc.relation.ispartof.es_PE.fl_str_mv https://www.scopus.com/pages/publications/85177454993
urn:issn:19755937
J. Educ. Eval. Health Prof. 2023; 20: 30
dc.rights.es_PE.fl_str_mv info:eu-repo/semantics/openAccess
dc.rights.uri.es_PE.fl_str_mv http://creativecommons.org/licenses/by/4.0/
eu_rights_str_mv openAccess
rights_invalid_str_mv http://creativecommons.org/licenses/by/4.0/
dc.format.es_PE.fl_str_mv application/pdf
dc.publisher.es_PE.fl_str_mv Korea Health Personnel Licensing Examination Institute.
dc.source.none.fl_str_mv reponame:UNC-Institucional
instname:Universidad Nacional de Cajamarca
instacron:UNC
instname_str Universidad Nacional de Cajamarca
instacron_str UNC
institution UNC
reponame_str UNC-Institucional
collection UNC-Institucional
bitstream.url.fl_str_mv http://repositorio.unc.edu.pe/bitstream/20.500.14074/10222/1/jeehp-20-30.pdf
bitstream.checksum.fl_str_mv 310e5371b2d6dd92c20926bad77fb4d8
bitstream.checksumAlgorithm.fl_str_mv MD5
repository.name.fl_str_mv Universidad Nacional de Cajamarca
repository.mail.fl_str_mv repositorio@unc.edu.pe
_version_ 1864825001218146304
score 13.411838
Nota importante:
La información contenida en este registro es de entera responsabilidad de la institución que gestiona el repositorio institucional donde esta contenido este documento o set de datos. El CONCYTEC no se hace responsable por los contenidos (publicaciones y/o datos) accesibles a través del Repositorio Nacional Digital de Ciencia, Tecnología e Innovación de Acceso Abierto (ALICIA).