MEGA Hub

TEAMMix: Taxonomy Enrichment Augmentation and Minority-augmented Mixing Strategy for LLM-enhanced Weak-Supervised Hierarchical Text Classification

Authors

Do you know Jian Zhang?You can claim authorship or link another user.Do you know Zhuohao Yang?You can claim authorship or link another user.Do you know Songlin Lei?You can claim authorship or link another user.Do you know Bangli Liu?You can claim authorship or link another user.Do you know Ziwei Wang?You can claim authorship or link another user.Do you know Xufeng Weng?You can claim authorship or link another user.Do you know Gehan Amaratunga?You can claim authorship or link another user.Do you know Yu Lin?You can claim authorship or link another user.Do you know Hongwei Wang?You can claim authorship or link another user.

Abstract

Hierarchical Text Classification (HTC), as a critical text mining task, faces challenges such as complex label hierarchies and class imbalance. Existing methods based on large language models (LLMs) struggle to be efficiently applied to this task due to issues like lengthy prompts and loss of label structural information. To address these limitations, this paper proposes a weakly supervised HTC framework enhanced by LLM-based data augmentation. The framework first enriches the label hierarchy semantically through keyword generation and corpus mining, thereby enhancing the model's understanding of labels. Subsequently, it guides the LLM to generate pseudo-samples to mitigate the long-tail problem, and employs a Gaussian mixture model for confidence-based resampling to optimize the quality of generated data. Experimental results demonstrate that the proposed method effectively improves the reliability of LLM-generated pseudo-labels and significantly enhances classification performance on fine-grained and imbalanced datasets.

Community

00

Publication notes

Author note
Accepted by IEEE CSCWD 2026
DOI
10.1109/CSCWD68734.2026.11582679