Detection and Safeguarding of Chinese Toxic Content from Users and LLMs: A Survey
Junyu Lu, Weiming Wang, Deyi Ji, Lanyun Zhu, Bo Xu, Liang Yang (+2 more)
Abstract
Abstract Toxic content detection and safeguarding have become important topics in natural language processing and affective computing, as they involve the analysis of offensive expressions and harmful attitudes in language. With the widespread adoption of social media, diverse forms of user-generated toxic content have imposed substantial harm on individuals and society. In parallel, as large language models (LLMs) rapidly advance, they are increasingly exploited for malicious purposes, which in turn amplifies the proliferation of LLM-generated toxic content. These risks are particularly salient in Chinese online ecosystems, where evolving diverse slang and implicit expressions are widespread. However, compared with research on Western languages, Chinese toxic content has received comparatively less systematic attention, and existing studies remain fragmented across tasks, datasets, and methodology. A comprehensive synthesis is therefore needed to consolidate progress and identify open challenges. To this end, this paper systematically reviews research on Chinese toxic content detection and safeguarding along two tracks. First, we survey toxic content detection on Chinese social media, covering diverse forms of user-generated toxicity. Second, we review efforts to safeguard Chinese LLMs against LLM-generated toxic content. By organizing representative methods and resources, this review maps the technical landscape and identifies key challenges for guiding future research.
Identifiers
Radar topics