MEGA Hub

Evidence-in-the-Loop: Trace-Driven Optimization for Customer-Service LLM Agents

Authors

Do you know Chunming Wu?You can claim authorship or link another user.Do you know Dafei Qiu?You can claim authorship or link another user.Do you know Congde Yuan?You can claim authorship or link another user.Do you know Charles Quan?You can claim authorship or link another user.Do you know Jun Wu?You can claim authorship or link another user.Do you know Suipeng Li?You can claim authorship or link another user.Do you know Mo Wu?You can claim authorship or link another user.Do you know Gavin Xie?You can claim authorship or link another user.Do you know Hope Chen?You can claim authorship or link another user.Do you know Max Yao?You can claim authorship or link another user.

Abstract

Production customer-service bots must improve answer quality across iterative releases, yet large language models must not bypass evidence boundaries, policy rules, or human-handoff safeguards. We present an \textbf{Evidence-Grounded Customer-Service Agent Workflow} deployed in a real-world customer-service setting. BM25 recall, issue-title-vector recall, issue-description-vector recall, weighted RRF fusion, and cross-encoder reranking construct grounded FAQ evidence for controlled LLM decisions. Policy-guided orchestration then combines this RAG evidence with scenario-specific rule evidence, conversation memory, and clarification state inside a fixed LangGraph DAG~\cite{langgraph2024}. The paper contributes three reusable deployment patterns: \textbf{hybrid RAG evidence construction}, where multi-channel retrieval and reranking produce auditable FAQ candidates; \textbf{evidence-grounded issue/action decision}, where an Evidence-Grounded Decision Module selects an issue/action from typed FAQ evidence and scenario-specific rule evidence; and \textbf{trace-driven RAG and reranker improvement}, where traces diagnose whether failures come from recall, ranking, final candidate selection, clarification, rule-derived evidence, or action policy, and where reranker fine-tuning is evaluated not only for in-domain gain but also for forgetting risk.

Community

00