A Hybrid CNN-Transformer Architecture with Bidirectional Cross-Attention Fusion for Efficient Flower Recognition
DOI:
https://doi.org/10.33022/ijcs.v15i3.5181Abstract
Automated flower recognition is crucial to the fields of agriculture and biodiversity monitoring, but deep learning models are extremely high in computing demand for resource-constrained devices. In this paper, a compact and efficient hybrid model based on EfficientNet-B4 and ViT-Small/16 is introduced. The design uses a bidirectional cross-attention fusion mechanism that uses both local edge-level features and global context to learn highly discriminative representations for fine-grained classification. The model was tested using 3,500 images from 35 species of flowers from a curated, augmented database, using 5-fold cross-validation. It outperformed the state-of-the-art architectures such as Swin Transformer and ResNet50, with 98.71% accuracy and a 98.80% F1 score with high statistical stability (95% CI:98.34%–99.09%) and with faster convergence rate. Statistically, these improvements were confirmed by pair-wise and Wilcoxon signed-rank statistical tests, which indicate the model's potential in resource-efficient real-world automated plant identification.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Hersh Hama, Hersh M. Hama, Shayan I. Jalal, Saman M. Omer, Mohammed H. Ahmed

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.





