Keywords

Artificial Intelligence, Higher Education, Assessment, Educational Technology, Digital Competence, Language Models

Abstract

The integration of large language models into higher education has prompted growing interest in assessing the quality and reliability of their responses in academic tasks. However, there remains a lack of comparative studies combining algorithmic metrics with expert human perception while considering prompt design as a key performance variable. This study comparatively analyzes the performance of ChatGPT 4.5 and Gemini 2.5 across educational and research tasks using an explanatory-sequential mixed methods design. In the quantitative phase, responses generated through prompts structured with ASPECCT, CLEAR, and Chain of Thought techniques were assessed across seven quality parameters. In the qualitative phase, 78 evaluators with differentiated academic profiles rated responses using specialized rubrics. Results show that Gemini 2.5 achieved the highest overall mean score (92.68/100), with significant advantages in disciplinary contextualization and cognitive depth, while ChatGPT 4.5 demonstrated greater stability in quantitative data analysis. Qualitative analysis revealed a stance of cautious trust among evaluators, who acknowledged the models as powerful textual co-designers yet epistemologically fragile. The study concludes that generative AI effectiveness in academic settings critically depends on prompt quality, expert supervision, and disciplinary adaptation, positioning prompt engineering as an essential academic competency.

References

Al-Thani, S. N., Anjum, S., Bhutta, Z. A., Bashir, S., Majeed, M. A., Khan, A. S. y Bashir, K. (2025). Comparative performance of ChatGPT, Gemini, and final-year emergency medicine clerkship students in answering multiple-choice questions: implications for the use of AI in medical education. International Journal of Emergency Medicine, 18(1), 146. https://doi.org/10.1186/s12245-025-00949-6
Albadarin, Y., Saqr, M., Pope, N. y Tukiainen, M. (2024). A systematic literature review of empirical research on ChatGPT in education. Discover Education, 3(1), 60. https://doi.org/10.1007/s44217-024-00138-2
Ali, M., Harbieh, I. y Haider, K. H. (2025). Bytes versus brains: A comparative study of AI-generated feedback and human tutor feedback in medical education. Medical Teacher, 48(1), 131–141. https://doi.org/10.1080/0142159X.2025.2519639
Avello-Martínez, R., Gajderowicz, T. y Gómez-Rodríguez, V. G. (2024). ¿ChatGPT es útil para que los estudiantes de posgrado adquieran conocimientos sobre narración digital y reduzcan su carga cognitiva? Un experimento. Revista de Educación a Distancia (RED), 24(78), 8. https://doi.org/10.6018/red.604621
Bas Graells, G., Tinoco Devia, R., Salinas Leyva, L. C. y Sevilla Molina, J. (2024). Systematic review of taxonomies of risks associated with Artificial Intelligence. Analecta Política, 14(26), 01–25. https://doi.org/10.18566/apolit.v14n26.a08
Berghea, F., Berghea, E. C., Daia, C. O., Ciuc, D. y Dinca, G. V. (2025). In the search for the perfect prompt in medical AI queries. Frontiers in Artificial Intelligence, 8, 1689178. https://doi.org/10.3389/frai.2025.1689178
Cain, W. (2023). Prompting Change: Exploring Prompt Engineering in Large Language Model AI and Its Potential to Transform Education. TechTrends, 68(1), 47–57. https://doi.org/10.1007/s11528-023-00896-0
Escalante, J., Pack, A. y Barrett, A. (2023). AI-generated feedback on writing: insights into efficacy and ENL student preference. International Journal of Educational Technology in Higher Education, 20(1), 57. https://doi.org/10.1186/s41239-023-00425-2
García-Peñalvo, F. J. (2024). Inteligencia artificial generativa y educación: Un análisis desde múltiples perspectivas. Education in the Knowledge Society (EKS), 25, e31942. https://doi.org/10.14201/eks.31942
Google LLC. (2025). Google Cloud for faculty. https://cloud.google.com/edu/faculty
Guan, Q. y Han, Y. (2025). From AI to authorship: Exploring the use of LLM detection tools for calling on “originality” of students in academic environments. Innovations in Education and Teaching International, 62(5), 1514–1528. https://doi.org/10.1080/14703297.2025.2511062
Holmes, W., Bialik, M. y Fadel, C. (2019). Inteligencia artificial en la educación: Promesas e implicaciones para la enseñanza y el aprendizaje. Centro para el Rediseño Curricular. https://discovery.ucl.ac.uk/id/eprint/10139722
ISO. (2011). ISO/IEC 25010:2011: Systems and software engineering — Systems and software Quality Requirements and Evaluation (SQuaRE) — System and software quality models. https://www.iso.org/standard/35733.html
Jara, I. y Ochoa, J. M. (2020). Usos y efectos de la inteligencia artificial en educación. Banco Interamericano de Desarrollo. https://doi.org/10.18235/0002380
Kim, J., Yu, S., Detrick, R. y Li, N. (2024). Exploring students’ perspectives on Generative AI-assisted academic writing. Education and Information Technologies, 30(1), 1265–1300. https://doi.org/10.1007/s10639-024-12878-7
Lee, D. y Palmer, E. (2025). Prompt engineering in higher education: a systematic review to help inform curricula. International Journal of Educational Technology in Higher Education, 22(1), 7. https://doi.org/10.1186/s41239-025-00503-7
Liu, Y., Park, J. y McMinn, S. (2024). Using generative artificial intelligence/ChatGPT for academic communication: Students’ perspectives. International Journal of Applied Linguistics, 34(4), 1437–1461. https://doi.org/10.1111/ijal.12574
Lo, L. S. (2023). The CLEAR path: A framework for enhancing information literacy through prompt engineering. The Journal of Academic Librarianship, 49(4), 102720. https://doi.org/10.1016/j.acalib.2023.102720
López Trinidad, N., Lara Inocencio, G., Silverio Hurtado, M. y Pascual Peralta, J. (2026). La inteligencia artificial en la educación: vacíos, tendencias y desafíos éticos. Una revisión de literatura (2019–2024). Comunicar, 34(84), 42–56. https://doi.org/10.5281/zenodo.18113730
Meyer, J. G., Urbanowicz, R. J., Martin, P. C. N., O’Connor, K., Li, R., Peng, P.-C., Bright, T. J., et al. (2023). ChatGPT and large language models in academia: opportunities and challenges. BioData Mining, 16(1), 20. https://doi.org/10.1186/s13040-023-00339-9
Mohammadi, M., Tajik, E., Martinez-Maldonado, R., Sadiq, S., Tomaszewski, W. y Khosravi, H. (2025). Artificial intelligence in multimodal learning analytics: A systematic literature review. Computers and Education: Artificial Intelligence, 8, 100426. https://doi.org/10.1016/j.caeai.2025.100426
Niño-Carrasco, S. A., Castellanos-Ramírez, J. C., Perezchica Vega, J. E. y Sepúlveda Rodríguez, J. A. (2025). Percepciones de estudiantes universitarios sobre los usos de inteligencia artificial en educación. Revista Fuentes, 27(1), 94–106. https://doi.org/10.12795/revistafuentes.2025.26356
Ouyang, F., Wu, M., Zheng, L., Zhang, L. y Jiao, P. (2023). Integration of artificial intelligence performance prediction and learning analytics to improve student learning in online engineering course. International Journal of Educational Technology in Higher Education, 20(1), 4. https://doi.org/10.1186/s41239-022-00372-4
Shen, M., Shen, Y., Liu, F. y Jin, J. (2025). Prompts, privacy, and personalized learning: integrating AI into nursing education—a qualitative study. BMC Nursing, 24(1), 470. https://doi.org/10.1186/s12912-025-03115-8
Singh, S. y Strzelecki, A. (2025). Academics as adopters of generative AI: an application of diffusion of innovations theory. Education and Information Technologies, 31(2), 621–645. https://doi.org/10.1007/s10639-025-13835-8
Wang, S., Xu, T., Li, H., Zhang, C., Liang, J., Tang, J., Yu, P. S., et al. (2025). Large Language Models for Education: A survey and outlook. IEEE Signal Processing Magazine, 42(6), 51–63. https://doi.org/10.1109/MSP.2025.3594309
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., et al. (2022). Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. En Advances in Neural Information Processing Systems 35 (NeurIPS 2022) (pp. 24824–24837). proceedings.com. https://doi.org/10.52202/068431-1800
Yan, L., Sha, L., Zhao, L., Li, Y., Martinez-Maldonado, R., Chen, G., Li, X., et al. (2023). Practical and ethical challenges of large language models in education: A systematic scoping review. British Journal of Educational Technology, 55(1), 90–112. https://doi.org/10.1111/bjet.13370
Zawacki-Richter, O., Marín, V. I., Bond, M. y Gouverneur, F. (2019). Systematic review of research on artificial intelligence applications in higher education – where are the educators? International Journal of Educational Technology in Higher Education, 16(1), 39. https://doi.org/10.1186/s41239-019-0171-0

Fundref

Los autores expresan su agradecimiento a la Universidad del Istmo (UDelIstmo) por el financiamiento recibido para llevar a cabo el proyecto “Análisis de buenas prácticas en diseño de prompts y su impacto en el rendimiento de modelos LLM”, correspondiente al PdT N.º PRO-CII-D-03-25. También reconocen el apoyo institucional y académico del Centro de Investigación Educativa AIP (CIEDU AIP), así como la colaboración de la Universidad Monteávila (UMA), la Universidad Pedagógica Experimental Libertador (UPEL) y la Universidad Internacional de Ciencia y Tecnología (UNICyT), cuya participación fue fundamental para el desarrollo, el análisis y la difusión de los resultados de esta investigación.

Crossmark

Technical information

Received: 2026-03-25 | Reviewed: 2026-05-05 | Accepted: 2026-05-25 | Online First: 2026-10-01 | Published: 2026-10-05

Metrics

Metrics of this article

Views: 38099

Abstract readings: 36810

PDF downloads: 1289

Full metrics of Comunicar 77

Views: 459033

Abstract readings: 446071

PDF downloads: 12962

Cited by

Cites in Web of Science

Currently there are no citations to this document

Cites in Scopus

Currently there are no citations to this document

Cites in Google Scholar

Currently there are no citations to this document

Download

Alternative metrics

How to cite

Gustavo Quintero Barreto., Aura L. López de Ramos., Yuly Esteves González., Nelly Meléndez. (2026). AI Educational Responses: Quality and Reliability in ChatGPT and Gemini. Comunicar, 34(87). 10.5281/zenodo.23185063

Share

        

Oxbridge Publishing House

4 White House Way

B91 1SE Sollihul United Kingdom

Administration

Editorial office

Creative Commons

This website uses cookies to obtain statistical data on the navigation of its users. If you continue to browse we consider that you accept its use. +info X