MEGA Hub

Where Animacy Lives in Large Language Models: Tracing the Circuits of the Animacy Concept

Authors

Do you know Samuele Punzo?You can claim authorship or link another user.Do you know Giovanni Cinà?You can claim authorship or link another user.Do you know Sandro Pezzelle?You can claim authorship or link another user.

Abstract

Distinguishing animate from inanimate concepts in written language requires more than shallow text processing, as it involves recognizing complex selectional constraints and contextual cues, such as verb-argument interactions. Yet, current large language models (LLMs) appear to be capable of doing it. We investigate whether this animacy-sensitive behavior of LLMs can be traced to a localized set of causally relevant components and connections. To do so, we construct a controlled dataset of minimal pairs and perform circuit discovery on four open-weight models. Through in-depth experiments and ablations, we show that a causal mechanism responsible for handling animacy in these models does exist, thus discovering an animacy circuit. At the same time, this circuit appears to be less localized compared to other known ones and generalizes only partially across models and animacy tasks, confirming the distributed, context-dependent, and somewhat graded nature of the animacy concept.

Community

00