Inicio
todoELE
  • Inicio
  • Materiales
    • πŸ“‹ Actividades
    • πŸ“ Conjugación
    • πŸ“Š Corpus
    • πŸ“” Diccionarios
    • βœ… Evaluación
    • βš™οΈ Gramática
    • πŸ“— Manuales
    • ✍️ Ortografía
    • πŸ“… Programación
    • πŸ—£οΈ Pronunciación
    • πŸ“ Recursos
    • πŸ”€ Vocabulario
    • πŸ’» Herramientas digitales
  • Formación
    • πŸ“š Bibliografía
    • πŸ‘₯ Congresos
    • πŸŽ“ Cursos
    • 🏫 Centros
    • 🏒 Organizaciones
    • πŸ“° Revistas
    • 🌍 Atlas de ELE
  • Trabajo
    • πŸ’Ό Ofertas de trabajo
    • ℹ️ Trabajo - Recursos
  • En la red
    • 🌐 Sitios ELE
    • πŸ“° Agregador
    • πŸ“§ Formespa
  • IA
    • ✨ Nuevos contenidos
    • πŸ“š Bibliografía IA
    • 🧰 Herramientas IA
    • πŸ’¬ Prompts
    • πŸ§ͺ Experiencias IA
    • 🌐 Sitios web IA
    • πŸ“° Actualidad IA
  • Comunidad
    • πŸ“° Actualidad ELE
    • 😊 Anécdotas ELE
    • πŸ“ Blog
    • πŸ“ŒTablón de anuncios
  • Buscar

Ruta de navegación

  • Inicio
  • Bibliografia
  • FermBench: A new benchmark for measuring the capabilities of LLMs on fermentation knowledge

Sección IA: Inteligencia artificial Bibliografía

FermBench: A new benchmark for measuring the capabilities of LLMs on fermentation knowledge

Fiammetta Caccavale
Adem R.N. Aouichaoui
Ulrich Krühne
Krist V. Gernaey
Carina L. Gargalo
2026
Computers & Education: Artificial Intelligence
10
https://www.sciencedirect.com/science/a…
artículo
estudio empírico
inteligencia artificial
grandes modelos de lenguaje
chatbots
análisis de producción de IA
modelos de lenguaje (LLM)
estudio empírico

Texto completo

Generative Artificial Intelligence (GenAI) chatbots continue to amaze users worldwide with their rapid improvements. These tools possess vast general knowledge and can thus be used in various fields, including education. However, before rolling out these models in pedagogical applications, it is fundamental to understand whether the information provided is reliable and if any of the currently available chatbots are best suited for domain-specific tasks. The objective of this study is to thoroughly investigate these aspects in a specific domain, fermentation, with the overarching goal of providing guidelines to students and teachers to select the best GenAI assistant. To achieve this goal, we introduce FermBench , a dataset specifically designed for fermentation processes. We use the collected data to benchmark five large language models (LLMs) powering commercially available GenAI chatbots, including ChatGPT, Gemini, DeepSeek, Claude and le Chat. To evaluate the responses of these models, we propose a robust experimental framework that includes automated metrics, human annotations, and the LLM-as-a-Judge approach. The obtained results suggest that, given the high baseline and the fact that the judges were unable to agree on an overall best model, the current knowledge embedded within these models is adequate and the standalone results cannot provide pedagogical guidelines regarding which chatbot should be used in education. These results suggest that the choice of which GenAI chatbot should be supported by institutional or government guidance, as well as individual preferences, perhaps informed by the parameters identified in our analysis. An interesting finding of the study is that curated answers are not necessarily better than generated ones.

Texto completo en abierto (CC BY 4.0).
  • Inicie sesión para enviar comentarios

Enviar publicación

Contenidos relacionados

  • ChatGMP: A case of AI chatbots in chemical engineering education towards the automation of repetitive tasks
  • El léxico ELE en los modelos de lenguaje
  • ¿Tienen GPT-3.5 y GPT-4 un estilo de escritura diferente del estilo humano?: un estudio exploratorio para el español
  • The promise and limits of LLMs in constructing proofs and hints for logic problems in intelligent tutoring systems
  • How reliable are large language models in analyzing the quality of written lesson plans? A mixed-methods study from a teacher internship program
  • Standardized assessment of LLM English proficiency
Sobre Todoele Índice Publica Contacto: todoele@gmail.com
Política de privacidad Créditos