MEGA Hub

Explicit State Elicitation Is Not Enough: A Controlled Audit of Memory-Policy Classification

Authors

Do you know Yihang Chen?You can claim authorship or link another user.Do you know Pin Qian?You can claim authorship or link another user.Do you know Su Wang?You can claim authorship or link another user.Do you know Chong Peng?You can claim authorship or link another user.Do you know Huan Xu?You can claim authorship or link another user.Do you know Shuaiting Li?You can claim authorship or link another user.Do you know Yiqi Sun?You can claim authorship or link another user.

Abstract

Personalized agents must decide whether retrieved user memory should be used, ignored, updated, or queried before it affects a current task. We use this setting to develop an empirical audit protocol for structured intermediate outputs: first audit dataset shortcuts, then isolate bundled prompt changes, check whether intermediate labels are answer-associated, test decomposed semantic evidence, and audit provider-level execution failures. A 480-example synthetic development set initially suggested large gains from a state-structured prompt bundle, but TF-IDF diagnostics showed lexical separability and no positive standalone Ignore cases. We therefore construct a frozen 160-example controlled counterfactual set with 40 matched four-way families and rule-derived reference policies. On this set, exposing the four state definitions improves accuracy, but an isolated explicit state-output field does not significantly improve policy accuracy for Llama-3.3-70B and gives only a marginal, non-significant gain for GPT-OSS-120B. Supplying benchmark-associated state labels shifts policy predictions, but because those labels deterministically map to policies, this is a label-conditioning diagnostic rather than evidence of a faithful internal mechanism. Family-level and seed-stability analyses further show that example-level accuracy overstates counterfactual consistency: complete four-way family success is rare. An exploratory follow-up that elicits decomposed semantic evidence also fails to improve routing for the cleanly evaluated endpoint; the corresponding GPT-OSS condition was unavailable because of provider-side request validation. We evaluate policy classification only, not downstream responses, tool actions, or memory-store mutation.

Community

00

Publication notes

Author note
34 pages, 1 figure