MEGA Hub

From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop

Authors

Do you know Rahul Gupta?You can claim authorship or link another user.Do you know Abhinav Mohanty?You can claim authorship or link another user.Do you know Anaelia Ovalle?You can claim authorship or link another user.Do you know Anil Ramakrishna?You can claim authorship or link another user.Do you know Anubrata Das?You can claim authorship or link another user.Do you know Apurv Verma?You can claim authorship or link another user.Do you know Jwala Dhamala?You can claim authorship or link another user.Do you know Ninareh Mehrabi?You can claim authorship or link another user.Do you know Tharindu Kumarage?You can claim authorship or link another user.Do you know Yada Pruksachatkun?You can claim authorship or link another user.Do you know Yang Trista Cao?You can claim authorship or link another user.Do you know Kai-Wei Chang?You can claim authorship or link another user.Do you know Aram Galstyan?You can claim authorship or link another user.

Abstract

The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located with major ACL conferences since 2021, has grown from 8 proceedings papers to 41 over six editions, documenting a field-wide transition from post-hoc interpretability of static models to mechanistic understanding and proactive control of generative systems. We synthesize insights from all 144 proceedings papers, classifying them along six trust dimensions grounded in established frameworks (TrustLLM, DecodingTrust). We observe co-occurrences with capability emergence. The release of the first high-impact chat models activated all trust dimensions simultaneously, while subsequent model generations shifted focus toward truthfulness and safety alignment. Analysis from the classification study reveals that truthfulness is the fastest-growing dimension (absent in 2021-2022, comprising 37% of papers by 2025-2026), fairness remains the most consistent theme, and explainability exhibits a U-shaped trajectory; declining as post-hoc methods lost relevance but resurging in 2026 through mechanistic interpretability. A cross-venue comparison with ACL, NAACL, EACL, and EMNLP (~2K papers) in the same period shows that TrustNLP's topical distribution closely follows the field average. We identify four structural insights and conclude with actionable directions for the research community.

Community

00

Publication notes

Author note
17 pages, 2 figures, 3 tables. Submitted to ACL ARR August 2026 cycle (EACL 2027)