Inicio
todoELE
  • Inicio
  • Materiales
    • πŸ“‹ Actividades
    • πŸ“ Conjugación
    • πŸ“Š Corpus
    • πŸ“” Diccionarios
    • βœ… Evaluación
    • βš™οΈ Gramática
    • πŸ“— Manuales
    • ✍️ Ortografía
    • πŸ“… Programación
    • πŸ—£οΈ Pronunciación
    • πŸ“ Recursos
    • πŸ”€ Vocabulario
    • πŸ’» Herramientas digitales
  • Formación
    • πŸ“š Bibliografía
    • πŸ‘₯ Congresos
    • πŸŽ“ Cursos
    • 🏫 Centros
    • 🏒 Organizaciones
    • πŸ“° Revistas
    • 🌍 Atlas de ELE
  • Trabajo
    • πŸ’Ό Ofertas de trabajo
    • ℹ️ Trabajo - Recursos
  • En la red
    • 🌐 Sitios ELE
    • πŸ“° Agregador
    • πŸ“§ Formespa
  • IA
    • ✨ Nuevos contenidos
    • πŸ“š Bibliografía IA
    • 🧰 Herramientas IA
    • πŸ’¬ Prompts
    • πŸ§ͺ Experiencias IA
    • 🌐 Sitios web IA
    • πŸ“° Actualidad IA
  • Comunidad
    • πŸ“° Actualidad ELE
    • 😊 Anécdotas ELE
    • πŸ“ Blog
    • πŸ“ŒTablón de anuncios
  • Buscar

Ruta de navegación

  • Inicio
  • Bibliografia
  • A hybrid reasoning framework for artificial intelligence assessment rubric generation in human and automated contexts: Evidence from an undergraduate programming course

Sección IA: Inteligencia artificial Bibliografía

A hybrid reasoning framework for artificial intelligence assessment rubric generation in human and automated contexts: Evidence from an undergraduate programming course

Pedro C. Mendonça
Filipe Quintal
Mário Figueiredo
Karolina Baras
Fábio Mendonça
2026
Computers & Education: Artificial Intelligence
11
https://www.sciencedirect.com/science/a…
artículo
estudio empírico
inteligencia artificial
evaluación
grandes modelos de lenguaje
pensamiento computacional
educación superior
IA y evaluación
modelos de lenguaje (LLM)
estudio empírico

Texto completo

Developing assessment rubrics is resource-intensive, limiting frequent formative assessment. This study introduces HARMOGEN-R (Hybrid Assessment Rubric Model Generation with Reasoning), a framework that uses reasoning-enhanced Large Language Models (LLMs) for initial rubric generation and standard models for synthesis. It supports two generation approaches, namely, structured, with predefined evaluation criteria, and free-form, with Artificial Intelligence (AI) defined criteria. Using a within-subjects design, four AI-generated rubrics and a human-created baseline were compared across 308 open-ended responses (text and code) from three formative programming assignments in an undergraduate computer science course. Four human evaluators and three LLM evaluators scored all responses under each rubric. When applied by human evaluators, all AI-generated rubrics met pooled equivalence within an operational ±5-point margin (±1 point on the Portuguese 0-20 scale, 5% of the maximum), with correlations from 0.948 to 0.973 and quadratic weighted kappa for grade agreement from 0.929 to 0.988. At the assignment and question-type level, five cells exceeded this margin, and OpenAI Structured was the only source meeting equivalence across all assignments. In automated evaluation, equivalence depended on the evaluator model. DeepSeek V3 met equivalence for all rubric sources, whereas GPT-4.1 and GPT-4o showed systematic negative deviations. Open-weight (DeepSeek) and proprietary (OpenAI) models produced equivalent rubrics in most conditions, with structured generation showing higher cross-model consistency. These findings indicate that AI-generated rubrics can achieve scoring comparability with human-created rubrics for technical content, although outcomes depend on rubric format, evaluator model, and assessment context; the evidence concerns scoring comparability rather than broader rubric validity.

Texto completo en abierto (CC BY 4.0).
  • Inicie sesión para enviar comentarios

Enviar publicación

Contenidos relacionados

  • Automatic question-answer pairs generation using pre-trained large language models in higher education
  • Assessing the quality of automatic-generated short answers using GPT-4
  • Reimagining feedback through generative AI in engineering education
  • A framework for evaluation of large language models in essay assessment: Reliability, alignment, and causal reasoning
  • Harnessing the power of AI-instructor collaborative grading approach: Topic-based effective grading for semi open-ended multipart questions
  • GPT-3.5 para la nivelación de ELE en un entorno universitario: Un estudio comparativo de zero-shot learning y fine-tuning
Sobre Todoele Índice Publica Contacto: todoele@gmail.com
Política de privacidad Créditos