MEGA Hub

WADE: A Reasoning-Annotated Benchmark for Multi-Instance Floating-Waste Grounding with Compact Vision-Language Models

Authors

Do you know Md. Asaduzzaman Shuvo?You can claim authorship or link another user.Do you know Ahsan Farabi?You can claim authorship or link another user.Do you know Md. Abdul Ahad Minhaz?You can claim authorship or link another user.Do you know Mahedi Hasan?You can claim authorship or link another user.Do you know Israt Khandaker?You can claim authorship or link another user.Do you know Ibrahim Khalil Shanto?You can claim authorship or link another user.Do you know Muhammad Nomani Kabir?You can claim authorship or link another user.

Abstract

Floating waste in inland waterways threatens aquatic ecosystems and requires timely monitoring under cluttered, multi-object conditions. Existing aquatic-waste datasets provide limited geographic coverage, sparse multi-instance annotations, and little supervision beyond boxes and labels. Compact vision-language models (VLMs) therefore remain insufficiently evaluated for jointly localizing, classifying, counting, and explaining floating waste. We introduce WADE, a reasoning-annotated benchmark containing 2,167 images from rural Bangladesh, 13,608 bounding boxes, and ten waste categories. Each annotation is associated with class-level recognition rules covering visual cues, likely confusions, and discriminative features. We evaluate six VLMs under zero-shot, two-shot, reasoning-guided, and fine-tuned settings using detection, counting, and hallucination metrics. For resource-efficient adaptation, we jointly fine-tune Qwen3-VL-2B on boxes, labels, and reasoning chains using QLoRA. Fine-tuning increases recall from 0.0248 to 0.2339 and F1 from 0.0257 to 0.2163, while reducing image-level hallucination from 0.6836 to 0.0883. However, over three-quarters of instances remain undetected, establishing WADE as a challenging benchmark for dense floating-waste grounding with compact VLMs.

Community

00