Inicio
todoELE
  • Inicio
  • Materiales
    • πŸ“‹ Actividades
    • πŸ“ Conjugación
    • πŸ“Š Corpus
    • πŸ“” Diccionarios
    • βœ… Evaluación
    • βš™οΈ Gramática
    • πŸ“— Manuales
    • ✍️ Ortografía
    • πŸ“… Programación
    • πŸ—£οΈ Pronunciación
    • πŸ“ Recursos
    • πŸ”€ Vocabulario
    • πŸ’» Herramientas digitales
  • Formación
    • πŸ“š Bibliografía
    • πŸ‘₯ Congresos
    • πŸŽ“ Cursos
    • 🏫 Centros
    • 🏒 Organizaciones
    • πŸ“° Revistas
    • 🌍 Atlas de ELE
  • Trabajo
    • πŸ’Ό Ofertas de trabajo
    • ℹ️ Trabajo - Recursos
  • En la red
    • 🌐 Sitios ELE
    • πŸ“° Agregador
    • πŸ“§ Formespa
  • IA
    • ✨ Nuevos contenidos
    • πŸ“š Bibliografía IA
    • 🧰 Herramientas IA
    • πŸ’¬ Prompts
    • πŸ§ͺ Experiencias IA
    • 🌐 Sitios web IA
    • πŸ“° Actualidad IA
  • Comunidad
    • πŸ“° Actualidad ELE
    • 😊 Anécdotas ELE
    • πŸ“ Blog
    • πŸ“ŒTablón de anuncios
  • Buscar

Ruta de navegación

  • Inicio
  • Bibliografia
  • Validating AI-generated classroom observations: Reliability, accuracy, and limits of LLM-based pedagogical judgment

Sección IA: Inteligencia artificial Bibliografía

Validating AI-generated classroom observations: Reliability, accuracy, and limits of LLM-based pedagogical judgment

Carolina Melo
Javiera de la Maza
Matías Recabarren
2026
Computers & Education: Artificial Intelligence
10
https://www.sciencedirect.com/science/a…
artículo
estudio empírico
inteligencia artificial
grandes modelos de lenguaje
práctica docente
formación de profesores
educación primaria y secundaria
modelos de lenguaje (LLM)
análisis de producción de IA
estudio empírico

Texto completo

This study examines the reliability and accuracy of large language models (LLMs) for automated classroom observation using the World Bank's TEACH Primary framework. As education systems increasingly explore AI-based tools to scale teacher feedback and professional development, empirical validation of these systems is critical. Using a corpus of 12 primary classroom videos, we compared 8618 AI-generated evaluations from eight LLM endpoints against consensus-based ratings from certified TEACH experts. To account for model stochasticity, each model produced 10 independent evaluations per video–element pair. Reliability was assessed using variability and inter-rater consistency indicators, while accuracy was evaluated using error-based and concordance-based agreement measures. Results show substantial stochastic variability across repeated evaluations, with no model achieving uniformly high reliability across instructional elements. Agreement with expert ratings remained moderate at best. Importantly, reliability and accuracy did not co-vary systematically: models producing more stable scores did not necessarily align better with expert judgments, and models with stronger expert agreement often exhibited higher internal variability. In an exploratory analysis of model justifications, patterns suggest that LLMs tend to prioritize explicit verbal cues over contextual or implicit pedagogical evidence when generating high-inference judgments. These findings highlight structural limitations of current text-based AI observation pipelines and demonstrate that automated classroom observation cannot be treated as a uniform capability. The study provides empirical evidence to inform the design and validation of AI-assisted observation systems that integrate pedagogical expertise, measurement constraints, and complementary human judgment.

Texto completo en abierto (CC BY-NC-ND 4.0).
  • Inicie sesión para enviar comentarios

Enviar publicación

Contenidos relacionados

  • Evaluating the performance of ChatGPT and GPT-4o in coding classroom discourse data: A study of synchronous online mathematics instruction
  • How reliable are large language models in analyzing the quality of written lesson plans? A mixed-methods study from a teacher internship program
  • El léxico ELE en los modelos de lenguaje
  • ¿Tienen GPT-3.5 y GPT-4 un estilo de escritura diferente del estilo humano?: un estudio exploratorio para el español
  • FermBench: A new benchmark for measuring the capabilities of LLMs on fermentation knowledge
  • Analysis of LLMs for educational question classification and generation
Sobre Todoele Índice Publica Contacto: todoele@gmail.com
Política de privacidad Créditos