MEGA Hub

The Role of Natural Language Understanding in Multimodal Video-Based Dengue Diagnosis

Authors

Do you know Danial Sharifrazi?You can claim authorship or link another user.Do you know Saadat Behzadi?You can claim authorship or link another user.Do you know Julakha Jahan Jui?You can claim authorship or link another user.Do you know Mojtaba Mohammadi?You can claim authorship or link another user.Do you know Nouman Javed?You can claim authorship or link another user.Do you know Roohallah Alizadehsani?You can claim authorship or link another user.Do you know Prasad N. Paradkar?You can claim authorship or link another user.Do you know Asim Bhatti?You can claim authorship or link another user.

Abstract

Detecting infection-related behavioral changes in mosquitoes from video data is challenging because mosquitoes are small, move rapidly and irregularly, and are affected by environmental factors such as background, lighting, and shadows, which can make reliable feature extraction difficult. In this study, a YOLO- and Contrastive Language-Image Pre-training (CLIP)-based vision-language framework is proposed to classify mosquito flight frames of uninfected and Dengue virus serotype 2 (DENV2)-infected mosquitoes. First, YOLO is used to isolate mosquito regions from the background. Then, visual features extracted from video frames are aligned with biologically meaningful textual prompts in a shared embedding space. The multimodal model was fine-tuned using supervised bidirectional contrastive learning and evaluated through frame-level image-text similarity-based classification. The results show that the proposed method achieved 98.54% accuracy and 99.91% sensitivity at the frame level. After temporal aggregation of frame-level information, the model achieved complete video-level performance. The ablation results showed that fine-tuning and CLIP-based representations were essential for this domain, while the textual branch provided semantic image-text alignment rather than an accuracy advantage over the vision-only model. These findings suggest that vision-language models can provide a useful framework for analyzing infection-related biological behaviors from video data.

Community

00