MEGA Hub

Seen, Said, or Forgotten? A Causal Audit of Visual KV Memory Across Dialog Turns

Authors

Do you know Hong Chen?You can claim authorship or link another user.Do you know Kang Chen?You can claim authorship or link another user.Do you know Yuxuan Fan?You can claim authorship or link another user.Do you know Bo Wang?You can claim authorship or link another user.Do you know Yubo Gao?You can claim authorship or link another user.Do you know Yuanlin Chu?You can claim authorship or link another user.Do you know Xuming Hu?You can claim authorship or link another user.

Abstract

Stateful multimodal assistants encode an image once but may answer questions about it many turns later. Attention-guided visual-KV eviction assumes that evidence irrelevant now will remain dispensable, although future questions are unknown. We ask when a visual fact is actually safe to forget and introduce the Causal Visual Memory Audit (CVMA), a paired single-prefill framework that tests what later answers lose when a visual region, the whole image, or prior assistant text becomes unavailable. On VisDial and ConvBench, current attention can rank future-useful regions worse than random even though a diagnostic marginal-utility control shows substantial selection headroom. Aggregate scores hide this failure when later turns do not need vision; controlled and stock-generated histories reveal a second escape route, in which assistant-text KV replaces image KV for facts already stated but not reliably for unstated facts. In the tested stacks, safe forgetting is supported by low future visual dependence or fact-specific verbalization---not by low current attention.

Community

00