Machine learning models combined with feature importance methods for honey yield classification: A replicable approach
Keywords: climatic variables, machine learning, supervised classification, apiculture, Model interpretability, Honey yield prediction
Abstract
Beekeepers require planning tools supported by modern technologies, such as machine learning and the Internet of Things, to address agricultural challenges such as the decrease and irregularity in honey production. To ensure replicability, this article presents a research workflow that begins with the creation of an open-access database, developed from annual production records and climatic variables (temperature and rainfall), integrating data construction, explainability analysis, and model evaluation. Then, feature importance methods and explainability techniques are applied, such as feature importance, the depth-wise frequency of each feature in random forest, and the Shapley Additive Explanations method. Finally, machine learning approaches are evaluated for honey yield prediction: logistic regression, k-nearest neighbors, support vector machine, decision tree, multilayer perceptron, random forest, linear discriminant analysis, gradient boosting, and Naive Bayes. These algorithms are compared considering: (1) a baseline corresponding to models without hyperparameter optimization, using leave-one-out cross-validation and stratified 10-fold cross-validation; (2) the baseline plus normalization/standardization (div-max, min-max, and z-score); (3) the configuration in point 2 plus bagging; (4) evaluation of a data-augmentation and class-balancing strategy using SMOTE, together with model combination via the Voting Classifier. The results suggest that rainfall is one of the most important variables for honey yield prediction. By selecting certain features, the models improve in some cases or do not significantly degrade their performance. The min-max and z-score methods led to improved predictions in some algorithms; for example, support vector machine achieved an accuracy of 0.80, compared with the 0.67 accuracy reported in the reference study based on random forest, representing an increase of 13 percentage points. Finally, bagging techniques, SMOTE oversampling, and the Voting Classifier, using algorithms such as KNN and SVM, can achieve an ACC of 0.82. Overall, this study proposes a replicable data mining-based workflow that integrates machine learning techniques for predicting honey yield from climatic variables, including the use of an open-access dataset, explainability analysis, and a comparative evaluation of machine learning models, contributing to the development of future analysis and planning tools in the beekeeping sector.
Más información
| Título de la Revista: | Smart Agricultural Technology |
| Volumen: | 15 |
| Editorial: | Sciencedirect |
| Fecha de publicación: | 2026 |
| Idioma: | Ingles |
| Notas: | WOS |