Inicio
todoELE
  • Inicio
  • Materiales
    • πŸ“‹ Actividades
    • πŸ“ Conjugación
    • πŸ“Š Corpus
    • πŸ“” Diccionarios
    • βœ… Evaluación
    • βš™οΈ Gramática
    • πŸ“— Manuales
    • ✍️ Ortografía
    • πŸ“… Programación
    • πŸ—£οΈ Pronunciación
    • πŸ“ Recursos
    • πŸ”€ Vocabulario
    • πŸ’» Herramientas digitales
  • Formación
    • πŸ“š Bibliografía
    • πŸ‘₯ Congresos
    • πŸŽ“ Cursos
    • 🏫 Centros
    • 🏒 Organizaciones
    • πŸ“° Revistas
    • 🌍 Atlas de ELE
  • Trabajo
    • πŸ’Ό Ofertas de trabajo
    • ℹ️ Trabajo - Recursos
  • En la red
    • 🌐 Sitios ELE
    • πŸ“° Agregador
    • πŸ“§ Formespa
  • IA
    • ✨ Nuevos contenidos
    • πŸ“š Bibliografía IA
    • 🧰 Herramientas IA
    • πŸ’¬ Prompts
    • πŸ§ͺ Experiencias IA
    • 🌐 Sitios web IA
    • πŸ“° Actualidad IA
  • Comunidad
    • πŸ“° Actualidad ELE
    • 😊 Anécdotas ELE
    • πŸ“ Blog
    • πŸ“ŒTablón de anuncios
  • Buscar

Ruta de navegación

  • Inicio
  • Bibliografia
  • How reliable are large language models in analyzing the quality of written lesson plans? A mixed-methods study from a teacher internship program

Sección IA: Inteligencia artificial Bibliografía

How reliable are large language models in analyzing the quality of written lesson plans? A mixed-methods study from a teacher internship program

Dennis Hauk
Nina Soujon
2026
Computers & Education: Artificial Intelligence
10
https://www.sciencedirect.com/science/a…
artículo
estudio empírico
inteligencia artificial
grandes modelos de lenguaje
formación de profesores
diseño curricular
evaluación
modelos de lenguaje (LLM)
IA y evaluación
análisis de producción de IA
estudio empírico

Texto completo

This study investigates the reliability of Large Language Models (LLMs) in evaluating the quality of written lesson plans from pre-service teachers. A total of 32 lesson plans, each ranging from 60 to 100 pages, were collected during a teacher internship program for civic education pre-service teachers. Using the ChatGPT-o1 reasoning model, we compared a human expert standard with LLM coding outcomes in a two-phase explanatory sequential mixed-methods design that combined quantitative reliability testing with a qualitative follow-up analysis to interpret inter-dimensional patterns of agreement. Quantitatively, overall reliability across six qualitative components of written lessons plans (Content Transformation, Task Creation, Adaptation, Goal Clarification, Contextualization and Sequencing ) reached a moderate alignment in identifying explicit instructional features (α = .689; 73.8% exact agreement). Qualitative analyses further revealed that the LLM struggled with high-inferential criteria, such as the depth of pedagogical reasoning and the coherence of instructional decisions, as it often relied on surface-level textual cues rather than deeper contextual understanding. These findings indicate that LLMs can support teacher educators and educational researchers as a design-stage screening tool, but human judgment remains essential for interpreting complex pedagogical constructs in written lesson plans and for ensuring the ethical and pedagogical integrity of evaluation processes. We outline implications for integrating LLM-based analysis into teacher education and emphasize improved prompt design and systematic human oversight to ensure reliable qualitative use.

Highlights:

  • Focus on reliability of LLMs rating of written lesson plans (CODE-PLAN)
  • Explanatory sequential design: quantitative α, then follow-up qualitative analysis
  • LLMs align with experts on explicit, surface-identifiable plan features
  • Lower reliability on high-inference pedagogy and instructional coherence
  • Practice guidance: LLM-supported mentor–mentee feedback workflow
Texto completo en abierto (CC BY 4.0).
  • Inicie sesión para enviar comentarios

Enviar publicación

Contenidos relacionados

  • Automated reading passage generation with OpenAI's large language model
  • Assessing the proficiency of large language models in automatic feedback generation: An evaluation study
  • Assessing the quality of automatic-generated short answers using GPT-4
  • LLMs do not grade essays like humans
  • Applying large language models and chain-of-thought for automatic scoring
  • Can large language models meet the challenge of generating school-level questions?
Sobre Todoele Índice Publica Contacto: todoele@gmail.com
Política de privacidad Créditos