Inicio
todoELE
  • Inicio
  • Materiales
    • πŸ“‹ Actividades
    • πŸ“ Conjugación
    • πŸ“Š Corpus
    • πŸ“” Diccionarios
    • βœ… Evaluación
    • βš™οΈ Gramática
    • πŸ“— Manuales
    • ✍️ Ortografía
    • πŸ“… Programación
    • πŸ—£οΈ Pronunciación
    • πŸ“ Recursos
    • πŸ”€ Vocabulario
    • πŸ’» Herramientas digitales
  • Formación
    • πŸ“š Bibliografía
    • πŸ‘₯ Congresos
    • πŸŽ“ Cursos
    • 🏫 Centros
    • 🏒 Organizaciones
    • πŸ“° Revistas
    • 🌍 Atlas de ELE
  • Trabajo
    • πŸ’Ό Ofertas de trabajo
    • ℹ️ Trabajo - Recursos
  • En la red
    • 🌐 Sitios ELE
    • πŸ“° Agregador
    • πŸ“§ Formespa
  • IA
    • ✨ Nuevos contenidos
    • πŸ“š Bibliografía IA
    • 🧰 Herramientas IA
    • πŸ’¬ Prompts
    • πŸ§ͺ Experiencias IA
    • 🌐 Sitios web IA
    • πŸ“° Actualidad IA
  • Comunidad
    • πŸ“° Actualidad ELE
    • 😊 Anécdotas ELE
    • πŸ“ Blog
    • πŸ“ŒTablón de anuncios
  • Buscar

Ruta de navegación

  • Inicio
  • Bibliografia
  • EvalYaks: Instruction tuning datasets and LoRA fine-tuned models for automated scoring of CEFR B2 speaking assessment transcripts

Sección IA: Inteligencia artificial Bibliografía

EvalYaks: Instruction tuning datasets and LoRA fine-tuned models for automated scoring of CEFR B2 speaking assessment transcripts

Nicy Scaria
Silvester John Joseph Kennedy
Thomas Latinovich
Deepak Subramani
2026
Computers & Education: Artificial Intelligence
10
https://www.sciencedirect.com/science/a…
artículo
estudio empírico
inteligencia artificial
expresión oral
evaluación
MCER
grandes modelos de lenguaje
enseñanza/aprendizaje de lenguas
IA y enseñanza-aprendizaje de lenguas
IA y evaluación
modelos de lenguaje (LLM)
estudio empírico

Texto completo

Relying on human experts to evaluate the Common European Framework of Reference for Languages (CEFR) speaking assessments in an e-learning environment creates scalability challenges, as it limits how quickly and widely assessments can be conducted. We aim to automate the evaluation of CEFR B2 English speaking assessments in e-learning environments from conversation transcripts. First, we evaluate the capability of leading open source and commercial Large Language Models (LLMs) to score a candidate’s performance across various criteria in the CEFR B2 speaking exam in both global and India-specific contexts. Next, we create a new expert-validated, CEFR-aligned synthetic conversational dataset with transcripts that are rated at different assessment scores. In addition, new instruction-tuned datasets are developed from the English Vocabulary Profile (up to CEFR B2 level) and the CEFR-SP WikiAuto datasets. Finally, using these new datasets, we perform parameter efficient instruction tuning of Mistral Instruct 7B v0.2 to develop a family of models called EvalYaks . Four models in this family are for assessing the four sections of the CEFR B2 speaking exam, one for identifying the CEFR level of vocabulary and generating level-specific vocabulary, and another for detecting the CEFR level of text and generating level-specific text. EvalYaks achieved an average acceptable accuracy of 96 %, a degree of variation of 0.35 levels, achieving performance competitive with state-of-the-art frontier models like GPT-4o and Gemini Flash 2.5. Furthermore, a pilot validation on real-world learner transcripts verified the model’s transferability to real-world assessment contexts. This demonstrates that a 7B parameter LLM instruction tuned with high-quality CEFR-aligned assessment data can effectively evaluate and score CEFR B2 English speaking assessments, offering a promising solution for scalable, automated language proficiency evaluation. The methodology is adaptable to other regional contexts and CEFR levels through appropriate data generation and validation protocols.

Texto completo en abierto (CC BY-NC 4.0).
  • Inicie sesión para enviar comentarios

Enviar publicación

Contenidos relacionados

  • Opening the blackbox of LLM-based automated essay scoring: Insights into feature weighting patterns and score validity
  • Assessing how accurately large language models encode and apply the common European framework of reference for languages
  • Large language models and automated essay scoring of English language learner writing: Insights into validity and reliability
  • A framework for evaluation of large language models in essay assessment: Reliability, alignment, and causal reasoning
  • Automated reading passage generation with OpenAI's large language model
  • Evaluating large language models as raters in large-scale writing assessments: A psychometric framework for reliability and validity
Sobre Todoele Índice Publica Contacto: todoele@gmail.com
Política de privacidad Créditos