Evaluating the Performance of Large Language Models for Spanish Language in Undergraduate Admissions Exams
Abstract: This study evaluates the performance of LLMs, specifically GPT-3.5 and BARD (supported by Gemini Pro model), in undergraduate admissions exams proposed by the National Polytechnic Institute in Mexico. The exams cover Engineering/Mathematical and Physical Sciences, Biological and Medical Sciences, and Social and Administrative Sciences. Both models demonstrated proficiency, exceeding the minimum acceptance scores for respective academic programs to up to 75% for some academic programs. GPT-3.5 outperformed BARD in Mathematics and Physics, while BARD performed better in History and questions related to factual information. Overall, GPT-3.5 marginally surpassed BARD with scores of 60.94% and 60.42%, respectively.
- Clinical knowledge and reasoning abilities of AI large language models in anesthesiology: A comparative study on the ABA exam. medRxiv, May 2023. DOI: 10.1101/2023.05.10.23289805.
- Language models are few-shot learners. arXiv, 2020. DOI: 10.48550/arXiv.2005.14165.
- Joost C. F. de Winter. Can chatgpt pass high school exams on english language comprehension? International Journal of Artificial Intelligence in Education, Sep 2023. DOI: 10.1007/s40593-023-00372-z.
- Peter A. Cotton Debby R. E. Cotton and J. Reuben Shipway. Chatting and cheating: Ensuring academic integrity in the era of chatgpt. Innovations in Education and Teaching International, 0(0):1–12, 2023. DOI: 10.1080/14703297.2023.2190148.
- The impact of chatgpt on higher education. Frontiers in Education, 8, 2023. DOI: 10.3389/feduc.2023.1206936.
- Google Gemini Team. Gemini: A Family of Highly Capable Multimodal Models. 2023. Available: https://storage.googleapis.com/deepmind-media/gemini/gemini_1_report.pdf [Accessed: 2023-12-06].
- Google Gemini Team. Introducing Gemini: our largest and most capable AI model. 2023. Available: https://blog.google/technology/ai/google-gemini-ai [Accessed: 2023-12-06].
- Google. Bard: Una herramienta de IA conversacional de Google. 2023. Available: https://bard.google.com [Accessed: 2023-12-06].
- Evaluating the efficacy of chatgpt in navigating the spanish medical residency entrance examination (mir): Promising horizons for ai in clinical medicine. Clinics and Practice, 13(6):1460–1487, 2023. DOI: 10.3390/clinpract13060130.
- Measuring massive multitask language understanding. arXiv, 2021. DOI: 10.48550/arXiv.2009.03300.
- IPN. IPN Programa institucional de mediano plazo . 2023. Available: https://www.ipn.mx/assets/files/coplaneval/docs/Planeacion/PIMP2123.pdf [Accessed: 2023-09-01].
- Instituto Kepler. EstadÃsticas del proceso de admisión IPN 2022. 2023. Available: https://institutokepler.com.mx/estadisticas-del-proceso-de-admision-ipn-nivel-superior [Accessed: 2023-10-23].
- Performance of chatgpt on usmle: Potential for ai-assisted medical education using large language models. PLOS Digital Health, 2(2):1–12, 02 2023. DOI: 10.1371/journal.pdig.0000198.
- Harnessing chatgpt and gpt-4 for evaluating the rheumatology questions of the spanish access exam to specialized medical training. Scientific Reports, 13(1):22129, Dec 2023. DOI: 10.1038/s41598-023-49483-6.
- OpenAI. ChatGPT de OpenAI. 2023. Available: https://chat.openai.com [Accessed: 2023-11-01].
- OpenAI. Gpt-4 Technical Report, 2023. Available: https://doi.org/10.48550/arXiv.2303.08774 [Accessed: 2023-12-01].
- Chatgpt and open-ai models: A preliminary review. Future Internet, 15(6), 2023. DOI: 10.3390/fi15060192.
- Llama 2: Open foundation and fine-tuned chat models. arXiv, 2023. DOI: 10.48550/arXiv.2307.09288.
- Wolfram. Mathematical notation characters. 2023. Available: https://reference.wolfram.com/language/guide/MathematicalNotationCharacters.html [Accessed: 2023-09-01].
Paper Prompts
Sign up for free to create and run prompts on this paper.