Inicio
todoELE
  • Inicio
  • Materiales
    • πŸ“‹ Actividades
    • πŸ“ Conjugación
    • πŸ“Š Corpus
    • πŸ“” Diccionarios
    • βœ… Evaluación
    • βš™οΈ Gramática
    • πŸ“— Manuales
    • ✍️ Ortografía
    • πŸ“… Programación
    • πŸ—£οΈ Pronunciación
    • πŸ“ Recursos
    • πŸ”€ Vocabulario
    • πŸ’» Herramientas digitales
  • Formación
    • πŸ“š Bibliografía
    • πŸ‘₯ Congresos
    • πŸŽ“ Cursos
    • 🏫 Centros
    • 🏒 Organizaciones
    • πŸ“° Revistas
    • 🌍 Atlas de ELE
  • Trabajo
    • πŸ’Ό Ofertas de trabajo
    • ℹ️ Trabajo - Recursos
  • En la red
    • 🌐 Sitios ELE
    • πŸ“° Agregador
    • πŸ“§ Formespa
  • IA
    • ✨ Nuevos contenidos
    • πŸ“š Bibliografía IA
    • 🧰 Herramientas IA
    • πŸ’¬ Prompts
    • πŸ§ͺ Experiencias IA
    • 🌐 Sitios web IA
    • πŸ“° Actualidad IA
  • Comunidad
    • πŸ“° Actualidad ELE
    • 😊 Anécdotas ELE
    • πŸ“ Blog
    • πŸ“ŒTablón de anuncios
  • Buscar

Ruta de navegación

  • Inicio
  • Bibliografia
  • Assessing how accurately large language models encode and apply the common European framework of reference for languages

Sección IA: Inteligencia artificial Bibliografía

Assessing how accurately large language models encode and apply the common European framework of reference for languages

Luca Benedetto
Gabrielle Gaudeau
Andrew Caines
Paula Buttery
2025
Computers & Education: Artificial Intelligence
8
https://www.sciencedirect.com/science/a…
artículo
estudio empírico
inteligencia artificial
MCER
grandes modelos de lenguaje
enseñanza/aprendizaje de lenguas
evaluación
IA y enseñanza-aprendizaje de lenguas
modelos de lenguaje (LLM)
análisis de producción de IA
estudio empírico

Texto completo

Large Language Models (LLMs) can have a transformative effect on a variety of domains, including education, and it is therefore pressing to understand whether these models have knowledge of – or, in other words, how they have encoded – the specific pedagogical requirements of different educational domains, and whether they use this when performing educational tasks. In this work, we propose an approach to evaluate the knowledge – or encoding – that the LLMs have of the Common European Framework of Reference for Languages (CEFR), and use it to evaluate five modern LLMs. Our study shows that the suite of tasks we propose is quite challenging for all the LLMs, and they often provide results which are not satisfactory and would be unusable in educational applications, suggesting that – even if they encode some information about the CEFR – this knowledge is not really leveraged when performing downstream tasks.

Texto completo en abierto (CC BY-NC-ND 4.0).
  • Inicie sesión para enviar comentarios

Enviar publicación

Contenidos relacionados

  • EvalYaks: Instruction tuning datasets and LoRA fine-tuned models for automated scoring of CEFR B2 speaking assessment transcripts
  • Standardized assessment of LLM English proficiency
  • LLMs do not grade essays like humans
  • Large language models and automated essay scoring of English language learner writing: Insights into validity and reliability
  • A framework for evaluation of large language models in essay assessment: Reliability, alignment, and causal reasoning
  • Automated reading passage generation with OpenAI's large language model
Sobre Todoele Índice Publica Contacto: todoele@gmail.com
Política de privacidad Créditos