Inicio
todoELE
  • Inicio
  • Materiales
    • πŸ“‹ Actividades
    • πŸ“ Conjugación
    • πŸ“Š Corpus
    • πŸ“” Diccionarios
    • βœ… Evaluación
    • βš™οΈ Gramática
    • πŸ“— Manuales
    • ✍️ Ortografía
    • πŸ“… Programación
    • πŸ—£οΈ Pronunciación
    • πŸ“ Recursos
    • πŸ”€ Vocabulario
    • πŸ’» Herramientas digitales
  • Formación
    • πŸ“š Bibliografía
    • πŸ‘₯ Congresos
    • πŸŽ“ Cursos
    • 🏫 Centros
    • 🏒 Organizaciones
    • πŸ“° Revistas
    • 🌍 Atlas de ELE
  • Trabajo
    • πŸ’Ό Ofertas de trabajo
    • ℹ️ Trabajo - Recursos
  • En la red
    • 🌐 Sitios ELE
    • πŸ“° Agregador
    • πŸ“§ Formespa
  • IA
    • ✨ Nuevos contenidos
    • πŸ“š Bibliografía IA
    • 🧰 Herramientas IA
    • πŸ’¬ Prompts
    • πŸ§ͺ Experiencias IA
    • 🌐 Sitios web IA
    • πŸ“° Actualidad IA
  • Comunidad
    • πŸ“° Actualidad ELE
    • 😊 Anécdotas ELE
    • πŸ“ Blog
    • πŸ“ŒTablón de anuncios
  • Buscar

Ruta de navegación

  • Inicio
  • Bibliografia
  • Towards responsible AI in education: A Delphi-AHP-based framework for evaluating educational large language models

Sección IA: Inteligencia artificial Bibliografía

Towards responsible AI in education: A Delphi-AHP-based framework for evaluating educational large language models

Pingrong Lin
Qin Deng
Yanbian Zhou
2026
Computers & Education: Artificial Intelligence
10
https://www.sciencedirect.com/science/a…
artículo
estudio empírico
inteligencia artificial
grandes modelos de lenguaje
ética
evaluación
modelos de lenguaje (LLM)
ética de la IA
estudio empírico

Texto completo

As large language models (LLMs) become deeply integrated into the educational landscape, evaluation criteria focusing solely on performance are insufficient to mitigate the risks of value misalignment and socioethical concerns. To steer educational LLMs towards responsible and beneficial development, this study aims to construct a multidimensional evaluation framework grounded in educational theory. Initially, a preliminary pool of evaluation indicators was established on the basis of a review of the literature and pedagogical theories. The Delphi method was subsequently employed to refine the indicator structure by integrating opinions from 21 cross-disciplinary experts. The analytic hierarchy process (AHP) was then applied to weigh these indicators and determine their priorities. The final framework comprises 5 first-level indicators and 21 s-level indicators. Learning effectiveness, knowledge construction capability, and social alignment are assigned critical weights, whereas intelligent interaction capability is less prioritized. Among the second-level indicators, information veracity was weighted the highest, while educational equity had the weakest influence. This study not only provides direction for the development and optimization of educational LLMs but also offers a reference for establishing a responsible artificial intelligence in education ecosystem.

Texto completo en abierto (CC BY 4.0).
  • Inicie sesión para enviar comentarios

Enviar publicación

Contenidos relacionados

  • Is GPT-4 fair? An empirical analysis in automatic short answer grading
  • Not for people like me: How frontier AI models redirect skeptical rural school staff
  • Fine-tuning ChatGPT for automatic scoring
  • Comparative analysis of NLP-driven MCQ generators from text sources
  • Beyond binary outcomes: Evaluating and mitigating bias in national standardized test score prediction
  • LLMs do not grade essays like humans
Sobre Todoele Índice Publica Contacto: todoele@gmail.com
Política de privacidad Créditos