Inicio
todoELE
  • Inicio
  • Materiales
    • πŸ“‹ Actividades
    • πŸ“ Conjugación
    • πŸ“Š Corpus
    • πŸ“” Diccionarios
    • βœ… Evaluación
    • βš™οΈ Gramática
    • πŸ“— Manuales
    • ✍️ Ortografía
    • πŸ“… Programación
    • πŸ—£οΈ Pronunciación
    • πŸ“ Recursos
    • πŸ”€ Vocabulario
    • πŸ’» Herramientas digitales
  • Formación
    • πŸ“š Bibliografía
    • πŸ‘₯ Congresos
    • πŸŽ“ Cursos
    • 🏫 Centros
    • 🏒 Organizaciones
    • πŸ“° Revistas
    • 🌍 Atlas de ELE
  • Trabajo
    • πŸ’Ό Ofertas de trabajo
    • ℹ️ Trabajo - Recursos
  • En la red
    • 🌐 Sitios ELE
    • πŸ“° Agregador
    • πŸ“§ Formespa
  • IA
    • ✨ Nuevos contenidos
    • πŸ“š Bibliografía IA
    • 🧰 Herramientas IA
    • πŸ’¬ Prompts
    • πŸ§ͺ Experiencias IA
    • 🌐 Sitios web IA
    • πŸ“° Actualidad IA
  • Comunidad
    • πŸ“° Actualidad ELE
    • 😊 Anécdotas ELE
    • πŸ“ Blog
    • πŸ“ŒTablón de anuncios
  • Buscar

Ruta de navegación

  • Inicio
  • Bibliografia
  • Harnessing the power of AI-instructor collaborative grading approach: Topic-based effective grading for semi open-ended multipart questions

Sección IA: Inteligencia artificial Bibliografía

Harnessing the power of AI-instructor collaborative grading approach: Topic-based effective grading for semi open-ended multipart questions

Phyo Yi Win Myint
Siaw Ling Lo
Yuhao Zhang
2024
Computers & Education: Artificial Intelligence
7
https://www.sciencedirect.com/science/a…
artículo
estudio empírico
inteligencia artificial
evaluación
grandes modelos de lenguaje
educación superior
IA y evaluación
modelos de lenguaje (LLM)
estudio empírico

Texto completo

Semi open-ended multipart questions consist of multiple sub questions within a single question, requiring students to provide certain factual information while allowing them to express their opinion within a defined context. Human grading of such questions can be tedious, constrained by the marking scheme and susceptible to the subjective judgement of instructors. The emergence of large language models (LLMs) such as ChatGPT has significantly advanced the prospect of automatic grading in educational settings. This paper introduces a topic-based grading approach that harnesses LLM capabilities alongside a refined marking scheme to ensure fair and explainable assessment processes. The proposed approach involves segmenting student responses according to sub questions, extracting topics utilizing LLM, and refining the marking scheme in consultation with instructors. The refined marking scheme is derived from LLM-extracted topics, validated by instructors to augment the original grading criteria. Leveraging LLM, we match student responses with refined marking scheme topics and employ a Python program to assign marks based on the matches. Various prompt versions are compared using relevant metrics to determine the most effective prompts. We evaluate LLM's grading proficiency through three approaches: zero-shot prompting, few-shot prompting, and our proposed method. Results indicate that while zero-shot and few-shot prompting methods fall short compared to human grading, the proposed approach achieves the best performance (highest percentage of exact match marks, lowest mean absolute error, highest Spearman correlation, highest Cohen's weighted kappa) and closely mirrors the distribution observed in human grading. Specifically, the collaborative approach enhances the grading process by refining the marking scheme to student responses, improving transparency and explainability through topic-based matching, and significantly increasing the effectiveness of LLMs when combined with instructor input, rather than as standalone automated grading systems.

Texto completo en abierto (CC BY-NC 4.0).
  • Inicie sesión para enviar comentarios

Enviar publicación

Contenidos relacionados

  • A hybrid reasoning framework for artificial intelligence assessment rubric generation in human and automated contexts: Evidence from an undergraduate programming course
  • Automatic question-answer pairs generation using pre-trained large language models in higher education
  • Assessing the quality of automatic-generated short answers using GPT-4
  • Reimagining feedback through generative AI in engineering education
  • A framework for evaluation of large language models in essay assessment: Reliability, alignment, and causal reasoning
  • GPT-3.5 para la nivelación de ELE en un entorno universitario: Un estudio comparativo de zero-shot learning y fine-tuning
Sobre Todoele Índice Publica Contacto: todoele@gmail.com
Política de privacidad Créditos