From Classification to Captioning: Adapting Vision-Language Models for Fruit Quality Description
Keywords: Language Models, Captioning; Fruit quality; Vision
Abstract
The increasing demand for efficient and automated quality assessment in agriculture underscores the need for intelligent systems that can accurately interpret the visual characteristics of produce. Traditional image-based classification models often fail to generate descriptive insights, thereby limiting their utility for professionals seeking richer semantic interpretations. To address this gap, we implement and evaluate state-of-the-art Vision-Language Models to generate improved textual descriptions of fruit quality from images. This paper presents a systematic comparison of vision models, applying them to a manually curated dataset of images with captions for three fruit classes. This dataset enables the use of zero-shot inference and fine-tuning during training. Our results demonstrate the effectiveness of the Vision-Models in providing an enhanced description in this evaluation setup, thereby validating the hypothesis that fine-tuning multimodal models enhances their descriptive precision in domain-specific contexts. These findings underscore the significance of rich textual datasets and the potential of fine-tuned Vision-Language models in enhancing post-harvest processing and supply chain decisions. Future work should focus on scaling to more fruit types, incorporating crowd-sourced descriptions, and integrating spatial or temporal cues to improve generalizability. © 2025 IEEE.
Más información
| Título de la Revista: | Proceedings - IEEE CHILEAN Conference on Electrical, Electronics Engineering, Information and Communication Technologies, ChileCon |
| Editorial: | Institute of Electrical and Electronics Engineers Inc. |
| Fecha de publicación: | 2025 |
| Idioma: | English |
| DOI: |
10.1109/CHILECON66915.2025.11476116 |