Scene description tasks effectively enhance students' English writing skills in contextual settings, facilitating the establishment of authentic situational connections. However, evaluating descriptive quality and providing accurate, level-appropriate feedback present significant challenges. Although Multimodal Large Language Models (MLLMs) have demonstrated strong capabilities in vision-language tasks, their generated feedback for scene description tasks often remains generic. It fails to account for students' educational stages. To address this limitation, we construct a novel level-specific feedback dataset for scene description tasks. This dataset is constructed using GPT-4o with Retrieval-Augmented Generation (RAG), guided by the Hong Kong primary and secondary school English word lists, which categorize vocabulary into four educational stages (key stages 1–4). We fine-tuned a designed MLLM on this dataset and evaluated its performance against open-source and closed-source baselines. Experimental results demonstrate that the proposed fine-tuned MLLM significantly enhances educational stage relevance in feedback generation while reducing hallucinated content. These findings substantiate the efficacy of fine-tuned MLLM in providing level-specific feedback for scene description tasks, advancing the potential for more adaptive AI-assisted writing support in educational contexts.
- Inicie sesión para enviar comentarios