Inicio
todoELE
  • Inicio
  • Materiales
    • πŸ“‹ Actividades
    • πŸ“ Conjugación
    • πŸ“Š Corpus
    • πŸ“” Diccionarios
    • βœ… Evaluación
    • βš™οΈ Gramática
    • πŸ“— Manuales
    • ✍️ Ortografía
    • πŸ“… Programación
    • πŸ—£οΈ Pronunciación
    • πŸ“ Recursos
    • πŸ”€ Vocabulario
    • πŸ’» Herramientas digitales
  • Formación
    • πŸ“š Bibliografía
    • πŸ‘₯ Congresos
    • πŸŽ“ Cursos
    • 🏫 Centros
    • 🏒 Organizaciones
    • πŸ“° Revistas
    • 🌍 Atlas de ELE
  • Trabajo
    • πŸ’Ό Ofertas de trabajo
    • ℹ️ Trabajo - Recursos
  • En la red
    • 🌐 Sitios ELE
    • πŸ“° Agregador
    • πŸ“§ Formespa
  • IA
    • ✨ Nuevos contenidos
    • πŸ“š Bibliografía IA
    • 🧰 Herramientas IA
    • πŸ’¬ Prompts
    • πŸ§ͺ Experiencias IA
    • 🌐 Sitios web IA
    • πŸ“° Actualidad IA
  • Comunidad
    • πŸ“° Actualidad ELE
    • 😊 Anécdotas ELE
    • πŸ“ Blog
    • πŸ“ŒTablón de anuncios
  • Buscar

Ruta de navegación

  • Inicio
  • Bibliografia
  • Fine-tuning AI models for enhanced consistency and precision in chemistry educational assessments

Sección IA: Inteligencia artificial Bibliografía

Fine-tuning AI models for enhanced consistency and precision in chemistry educational assessments

Sri Yamtinah
Antuni Wiyarsi
Hayuni Retno Widarti
Ari Syahidul Shidiq
Dimas Gilang Ramadhani
2025
Computers & Education: Artificial Intelligence
8
https://www.sciencedirect.com/science/a…
artículo
estudio empírico
inteligencia artificial
evaluación
grandes modelos de lenguaje
IA y evaluación
modelos de lenguaje (LLM)
estudio empírico

Texto completo

Integrating artificial intelligence (AI) into educational assessments represents a paradigm shift, especially in STEM subjects such as chemistry, which require complex problem-solving and written feedback. This study focuses on the effects of fine-tuning on each of the four AI models—Gemini 1.5, Gemini 1.0, BERT, and XLNet—and their performance on chemistry tasks. The models were implemented using three evaluation methods: Seq2Seq for Sodium Reaction Grading, a Regression Task for Overall Grading, and Attention-Based Grading for Key Steps Grading. The fine-tuning process significantly improved the models' accuracy, precision, and stability. Gemini 1.5 outperformed the other models across all three metrics, with accuracy increasing from 80 % to 89.5 % and TPR from 0.73 to 0.93, whereas Gemini 1.0 achieved TPR gains from 0.69 to 0.89. BERT and XLNet, which had substantially lower baselines, also showed significant improvements, particularly in identifying fundamental steps in the evaluation. These advances highlight the critical role of fine-tuning in refining AI model output to align with expert grading standards, ensuring accuracy and reliability in assessment. The results confirm that fine-tuning is essential in preparing AI models for teaching applications, particularly for complex tasks such as chemistry evaluation, thus enabling scalable solutions. The findings of this study provide further justification for the wider adoption of fine-tuned AI models to improve the reliability, scalability, and effectiveness of grading systems in STEM education.

Texto completo en abierto (CC BY-NC 4.0).
  • Inicie sesión para enviar comentarios

Enviar publicación

Contenidos relacionados

  • Assessing the proficiency of large language models in automatic feedback generation: An evaluation study
  • Harnessing the power of AI-instructor collaborative grading approach: Topic-based effective grading for semi open-ended multipart questions
  • How well can LLMs grade essays in Arabic?
  • A hybrid reasoning framework for artificial intelligence assessment rubric generation in human and automated contexts: Evidence from an undergraduate programming course
  • Automatic question-answer pairs generation using pre-trained large language models in higher education
  • Optimizing automated scoring in ILSAs with prompt compression
Sobre Todoele Índice Publica Contacto: todoele@gmail.com
Política de privacidad Créditos