Inicio
todoELE
  • Inicio
  • Materiales
    • πŸ“‹ Actividades
    • πŸ“ Conjugación
    • πŸ“Š Corpus
    • πŸ“” Diccionarios
    • βœ… Evaluación
    • βš™οΈ Gramática
    • πŸ“— Manuales
    • ✍️ Ortografía
    • πŸ“… Programación
    • πŸ—£οΈ Pronunciación
    • πŸ“ Recursos
    • πŸ”€ Vocabulario
    • πŸ’» Herramientas digitales
  • Formación
    • πŸ“š Bibliografía
    • πŸ‘₯ Congresos
    • πŸŽ“ Cursos
    • 🏫 Centros
    • 🏒 Organizaciones
    • πŸ“° Revistas
    • 🌍 Atlas de ELE
  • Trabajo
    • πŸ’Ό Ofertas de trabajo
    • ℹ️ Trabajo - Recursos
  • En la red
    • 🌐 Sitios ELE
    • πŸ“° Agregador
    • πŸ“§ Formespa
  • IA
    • ✨ Nuevos contenidos
    • πŸ“š Bibliografía IA
    • 🧰 Herramientas IA
    • πŸ’¬ Prompts
    • πŸ§ͺ Experiencias IA
    • 🌐 Sitios web IA
    • πŸ“° Actualidad IA
  • Comunidad
    • πŸ“° Actualidad ELE
    • 😊 Anécdotas ELE
    • πŸ“ Blog
    • πŸ“ŒTablón de anuncios
  • Buscar

Ruta de navegación

  • Inicio
  • Bibliografia
  • Standardized assessment of LLM English proficiency

Sección IA: Inteligencia artificial Bibliografía

Standardized assessment of LLM English proficiency

Shaonan Wang
Shangchao Min
Hui Wang
Xinyu Gao
Nai Ding
2026
Computers & Education: Artificial Intelligence
11
https://www.sciencedirect.com/science/a…
artículo
estudio empírico
inteligencia artificial
grandes modelos de lenguaje
evaluación
enseñanza/aprendizaje de lenguas
IA y enseñanza-aprendizaje de lenguas
modelos de lenguaje (LLM)
análisis de producción de IA
estudio empírico

Texto completo

Large language models (LLMs) are increasingly used in language learning and assessment, yet their English proficiency is seldom reported against interpretable proficiency standards. We introduce the China’s Standards of English Language Ability (CSE) framework to assess the proficiency levels and subskills of LLMs. The test is referred to as the CSEBench and comprises 624 expert-annotated multiple-choice items across CSE Levels 2–7. Each item is accompanied by metadata, including difficulty level and subskill labels covering vocabulary, syntax, phonology, and cohesion/discourse. Critically, the dataset includes test responses from 2,050 middle school and and sophomore college students who are learning English as a second language. We evaluate closed-source models, open-source baselines, and enhanced open-source variants incorporating additional supervision and external knowledge. Results show a clear proficiency divide: after mapping model scores to CSE levels, closed-source models consistently reach CSE Level 6, whereas most open-source baselines cluster around CSE Levels 3–4. A follow-up cognitive diagnostic analysis reveals that while closed-source LLMs exhibit broad competence across subskills, open-source models display persistent deficits—most pronounced in phonology. Crucially, these weaknesses are shown to be substantially reducible through targeted enhancements. CSEBench thus offers a proficiency-interpretable testbed for reporting LLM English ability and diagnosing subskill gaps.

Texto completo en abierto (CC BY-NC 4.0).
  • Inicie sesión para enviar comentarios

Enviar publicación

Contenidos relacionados

  • Assessing how accurately large language models encode and apply the common European framework of reference for languages
  • Automated reading passage generation with OpenAI's large language model
  • EvalYaks: Instruction tuning datasets and LoRA fine-tuned models for automated scoring of CEFR B2 speaking assessment transcripts
  • A framework for evaluation of large language models in essay assessment: Reliability, alignment, and causal reasoning
  • Opening the blackbox of LLM-based automated essay scoring: Insights into feature weighting patterns and score validity
  • LLMs do not grade essays like humans
Sobre Todoele Índice Publica Contacto: todoele@gmail.com
Política de privacidad Créditos