MEGA Hub

How Merge-Tolerant Are Vision Transformers for Wheat Phenotyping?

Authors

Do you know Simon Ravé?You can claim authorship or link another user.Do you know Pejman Rasti?You can claim authorship or link another user.Do you know David Rousseau?You can claim authorship or link another user.

Abstract

Vision-based wheat phenotyping requires repeated measurements under deployment constraints, from growth-stage recognition to wheat-head counting and organ segmentation. Plain Vision Transformers (ViTs) provide a common architecture for these tasks, but quadratic attention limits high-throughput and edge inference. Training-free token merging is attractive because it can be inserted into trained models without retraining. We provide a systematic benchmark of ToMe and Mutual Pair Merging across growth-stage classification, wheat-head detection, and wheat-organ segmentation, measuring task quality, throughput, token count, and peak GPU memory, with additional Raspberry Pi 5 measurements. The benchmark reveals a clear hierarchy: classification is highly merge-tolerant, while detection and segmentation are constrained by repeated instances, thin organs, dense boundaries, reconstruction, and runtime overhead. Optimized attention backends can erase apparent speedups, so deployment value must be profiled on the target runtime rather than inferred from token count.

Community

00

Publication notes

Author note
Accepted to the CVPPA workshop at ECCV 2026