From Classification to Captioning: Adapting Vision-Language Models for Fruit Quality Description

Donoso; José (60585182300); Nicolis; Orietta (16305118900); Pieringer; Christian (36608838800); Peralta; Billy (54403707300)

Keywords: Language Models, Captioning; Fruit quality; Vision

Abstract

The increasing demand for efficient and automated quality assessment in agriculture underscores the need for intelligent systems that can accurately interpret the visual characteristics of produce. Traditional image-based classification models often fail to generate descriptive insights, thereby limiting their utility for professionals seeking richer semantic interpretations. To address this gap, we implement and evaluate state-of-the-art Vision-Language Models to generate improved textual descriptions of fruit quality from images. This paper presents a systematic comparison of vision models, applying them to a manually curated dataset of images with captions for three fruit classes. This dataset enables the use of zero-shot inference and fine-tuning during training. Our results demonstrate the effectiveness of the Vision-Models in providing an enhanced description in this evaluation setup, thereby validating the hypothesis that fine-tuning multimodal models enhances their descriptive precision in domain-specific contexts. These findings underscore the significance of rich textual datasets and the potential of fine-tuned Vision-Language models in enhancing post-harvest processing and supply chain decisions. Future work should focus on scaling to more fruit types, incorporating crowd-sourced descriptions, and integrating spatial or temporal cues to improve generalizability. © 2025 IEEE.

Más información

Título de la Revista: Proceedings - IEEE CHILEAN Conference on Electrical, Electronics Engineering, Information and Communication Technologies, ChileCon
Editorial: Institute of Electrical and Electronics Engineers Inc.
Fecha de publicación: 2025
Idioma: English
DOI:

10.1109/CHILECON66915.2025.11476116