A comparative benchmark of vision transformer architectures for chili leaf disease classification

Authors

  • Acihmah Sidauruk Department of Computer Science, Universitas AMIKOM Yogyakarta, Indonesia
  • Danang Wijayanto Department of Computer Science, Universitas AMIKOM Yogyakarta, Indonesia
  • I Made Artha Agastya Department of Computer Science, Universitas AMIKOM Yogyakarta, Indonesia
  • Jumanto Unjung Department of Computer Science, Universitas Negeri Semarang, Indonesia
  • Mulia Sulistiyono Department of Computer Science, Universitas AMIKOM Yogyakarta, Indonesia

DOI:

https://doi.org/10.52465/joscex.v7i3.62

Keywords:

Image classification, Chili Plant Disease, Self-supervised learning, DINOv2, Vision transformer

Abstract

Chili plant disease detection represents a critical component for enhancing agricultural productivity. Although Convolutional Neural Networks (CNN) have demonstrated promising results, they encounter limitations in capturing global contextual relationships within images. However, existing Vision Transformer studies on plant disease commonly assess only a single architecture, leaving the relative performance of different Vision Transformer families on chili disease data largely unexamined. This research aims to conduct a comparative benchmark analysis of five Vision Transformer-based architectures ViT, Swin Transformer, MaxViT, DINOv2, and EVA-02 to identify the most optimal model for chili plant disease classification. The methodology begins with data preprocessing and augmentation on a chili leaf dataset comprising five classes: healthy, leaf curl, leaf spot, whitefly, and yellowish. Each model is then fine-tuned under consistent training configurations with early stopping to prevent overfitting, and evaluated using accuracy, precision, recall, F1-score, and AUC. The results indicate that DINOv2 achieves superior performance with 96% accuracy, 96% precision, 96% recall, 96% F1-score, and 99% AUC, along with the highest training efficiency through convergence at epoch 12, outperforming ViT (92%), Swin (88%), MaxViT (88%), EVA-02 (86%), and previous CNN-based approaches. These findings confirm the superior potential of Vision Transformers, particularly self-supervised models, as a promising alternative for agricultural disease detection applications. The main contribution of this study is the first unified, head-to-head benchmark of five distinct Vision Transformer families for chili leaf disease classification, providing practical guidance on model selection in terms of both accuracy and training efficiency.

Downloads

Published

11-08-2026

How to Cite

A comparative benchmark of vision transformer architectures for chili leaf disease classification. (2026). Journal of Soft Computing Exploration, 7(3), 549-560. https://doi.org/10.52465/joscex.v7i3.62

Similar Articles

11-20 of 42

You may also start an advanced similarity search for this article.