MEGA Hub

RUMBA: Russian User Memory Benchmark

Authors

Do you know Elizaveta Shevtsova?You can claim authorship or link another user.Do you know Inna Glebkina?You can claim authorship or link another user.Do you know Mark Baushenko?You can claim authorship or link another user.Do you know Pavel Gulyaev?You can claim authorship or link another user.Do you know Alena Fenogenova?You can claim authorship or link another user.

Abstract

The ability to handle long-term memory in LLMs is becoming increasingly critical, yet existing benchmarks remain English-centric and rely on aggregate retrieval metrics, failing to capture interactions between long-range context, temporal information, and reasoning. To address this, we introduce RUMBA (Russian User Memory BenchmArk) - a new benchmark for long-term conversational memory that provides a fine-grained taxonomy of memory-centric question types and a unified methodology accounting for semantic type, session scope, temporal reasoning, and the explicitness of temporal expressions. RUMBA consists of timestamped user-assistant dialogues with QA pairs requiring retrieval, combination, and reasoning across sessions. While designed for Russian, we also provide an aligned English subset under the same methodology. We evaluate contemporary memory systems and long-context models, and show how RUMBA serves as a diagnostic tool to analyze model behavior across benchmark slices and identify strengths and failure modes of different memory mechanisms.

Community

00