MEGA Hub

Visual Semantic Decoding of Electrocorticography from Video Stimuli using End-to-End Deep Learning

Authors

Do you know Stella Ho?You can claim authorship or link another user.Do you know Joel Villalobos?You can claim authorship or link another user.Do you know Joseph West?You can claim authorship or link another user.Do you know Jingyang Liu?You can claim authorship or link another user.Do you know Weijie Qi?You can claim authorship or link another user.Do you know Haruhiko Kishima?You can claim authorship or link another user.Do you know Ryohei Fukuma?You can claim authorship or link another user.Do you know Takufumi Yanagisawa?You can claim authorship or link another user.Do you know Sam E. John?You can claim authorship or link another user.Do you know David B. Grayden?You can claim authorship or link another user.

Abstract

ECoG-based visual semantic decoding enables inference of semantic interpretation of visual perception from complex, noisy brain activity. This study examines the feasibility of visual semantic decoding using an end-to-end deep learning framework using electrocorticography (ECoG). Specifically, the decoding task is to predict visual categories from video stimuli using time-series neural inputs. A previously collected ECoG dataset from participants ($n=17$) with drug-resistant epilepsy is used for analysis. With fewer than 50 training samples per visual category, this study evaluates multiple deep learning approaches, artificial neural network architectures, and frequency-band filtered inputs. The best-performing approach is analyzed to shed light on the discriminative information it relies on across spectral, temporal, and cortical dimensions. The selected decoding system uses mixup augmentation, a Transformer-based encoder, and high-gamma (80-150 Hz) inputs with a 900 ms post-stimulus window. Further analysis shows that early visual cortex (V2-V4), ventral stream visual cortex, MT+ complex with neighbouring visual areas, and lateral temporal cortex contributed substantially to decoding performance. This study demonstrates that an end-to-end deep learning framework can yield promising decoding performance from dynamic visual stimuli without handcrafted features, while the model behavior remains interpretable through spectral, temporal, and cortical dimensions, which are broadly consistent with established neuroscience knowledge.

Community

00

Publication notes

Author note
This is a preprint and has not yet undergone peer review