AI Educational Responses: Quality and Reliability in ChatGPT and Gemini
DOI:
https://doi.org/10.5281/zenodo.23185063Keywords:
Artificial Intelligence, Higher Education, Assessment, Educational Technology, Digital Competence, Language ModelsAbstract
The integration of large language models into higher education has prompted growing interest in assessing the quality and reliability of their responses in academic tasks. However, there remains a lack of comparative studies combining algorithmic metrics with expert human perception while considering prompt design as a key performance variable. This study comparatively analyzes the performance of ChatGPT 4.5 and Gemini 2.5 across educational and research tasks using an explanatory-sequential mixed methods design. In the quantitative phase, responses generated through prompts structured with ASPECCT, CLEAR, and Chain of Thought techniques were assessed across seven quality parameters. In the qualitative phase, 78 evaluators with differentiated academic profiles rated responses using specialized rubrics. Results show that Gemini 2.5 achieved the highest overall mean score (92.68/100), with significant advantages in disciplinary contextualization and cognitive depth, while ChatGPT 4.5 demonstrated greater stability in quantitative data analysis. Qualitative analysis revealed a stance of cautious trust among evaluators, who acknowledged the models as powerful textual co-designers yet epistemologically fragile. The study concludes that generative AI effectiveness in academic settings critically depends on prompt quality, expert supervision, and disciplinary adaptation, positioning prompt engineering as an essential academic competency.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Comunicar

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.