Inicio
todoELE
  • Inicio
  • Materiales
    • πŸ“‹ Actividades
    • πŸ“ Conjugación
    • πŸ“Š Corpus
    • πŸ“” Diccionarios
    • βœ… Evaluación
    • βš™οΈ Gramática
    • πŸ“— Manuales
    • ✍️ Ortografía
    • πŸ“… Programación
    • πŸ—£οΈ Pronunciación
    • πŸ“ Recursos
    • πŸ”€ Vocabulario
    • πŸ’» Herramientas digitales
  • Formación
    • πŸ“š Bibliografía
    • πŸ‘₯ Congresos
    • πŸŽ“ Cursos
    • 🏫 Centros
    • 🏒 Organizaciones
    • πŸ“° Revistas
    • 🌍 Atlas de ELE
  • Trabajo
    • πŸ’Ό Ofertas de trabajo
    • ℹ️ Trabajo - Recursos
  • En la red
    • 🌐 Sitios ELE
    • πŸ“° Agregador
    • πŸ“§ Formespa
  • IA
    • ✨ Nuevos contenidos
    • πŸ“š Bibliografía IA
    • 🧰 Herramientas IA
    • πŸ’¬ Prompts
    • πŸ§ͺ Experiencias IA
    • 🌐 Sitios web IA
    • πŸ“° Actualidad IA
  • Comunidad
    • πŸ“° Actualidad ELE
    • 😊 Anécdotas ELE
    • πŸ“ Blog
    • πŸ“ŒTablón de anuncios
  • Buscar

Ruta de navegación

  • Inicio
  • Bibliografia
  • LLM sentiment quantification reveals selective alignment with human course-evaluation raters

Sección IA: Inteligencia artificial Bibliografía

LLM sentiment quantification reveals selective alignment with human course-evaluation raters

Joyce W. Lacy
Chi Nnoka
Zachary Jock
Cathleen Morreale
2026
Computers & Education: Artificial Intelligence
10
https://www.sciencedirect.com/science/a…
artículo
estudio empírico
inteligencia artificial
grandes modelos de lenguaje
procesamiento del lenguaje natural
educación superior
modelos de lenguaje (LLM)
análisis de producción de IA
estudio empírico

Texto completo

Student course evaluations contain rich qualitative feedback in the form of comments written in response to open-ended questions. However, this qualitative data, which may be more nuanced and detailed than quantitative ratings, is often unexamined in both administrative and research settings due to the labor-intensive nature of manual analysis. We investigate whether large language models (LLMs), including BERT, RoBERTa, and OpenAI model variants, can accurately replicate human judgments of sentiment in these comments. We compare masked and generative language models, using both naïve and fine-tuned approaches, to analyze a curated dataset of 1000 de-identified course evaluation responses. Results show that some artificial intelligence (AI) models can approach inter-rater reliability with humans remarkably well and quickly with limited tuning or training data provided. However, performance varied and not all models were able to produce a reliable sentiment analysis, even after training. This has implications for future avenues of qualitative data analysis within course evaluations as well as the large repositories of course evaluations available at institutions of higher education. Importantly, consideration should be taken when selecting an AI model as this decision has ramifications for the reliability and validity of the generated output.

Texto completo en abierto (CC BY-NC-ND 4.0).
  • Inicie sesión para enviar comentarios

Enviar publicación

Contenidos relacionados

  • Can large language models write reflectively
  • How do LLMs perform in the context of MCQs across different levels of thinking skills in a business education course at higher education? A comparison of ChatGPT, Gemini, and Copilot
  • Assessing the quality of automatic-generated short answers using GPT-4
  • Leveraging generative AI for course learning outcome categorization using Bloom's taxonomy
  • El léxico ELE en los modelos de lenguaje
  • ¿Tienen GPT-3.5 y GPT-4 un estilo de escritura diferente del estilo humano?: un estudio exploratorio para el español
Sobre Todoele Índice Publica Contacto: todoele@gmail.com
Política de privacidad Créditos