Detect abusive content automatically and act gradually.
Traditional filters rely on word lists, so they are fooled by the first letter swap and delete innocent messages that happen to contain a word. This system reads the message and understands its meaning and context, telling an insult apart from the same word used in a discussion, and catching abuse written in a roundabout way.
A word filter is broken by design: it is fooled by the first swapped letter or added space, and it deletes an innocent message because it mentioned a word in ordinary context. This system reads the message and understands its meaning and context, telling a directed insult apart from the same word in a discussion, and catching abuse written in a roundabout way that matches no list.
Sensitivity comes in three levels: low for clear violations only, medium which is the recommendation, and high for a strict stance. Alongside it you pick the content categories detected, and you can ignore administrators and scan stickers.
Despite the smart detection, manual lists stay useful: blocked words and emojis you add yourself with how seriously a match is treated. Use them for words specific to your community that no general system knows, and leave general abuse to the smart detection.
The strike system is what makes it fair: a first violation is treated as a mistake with a deleted message and a warning, while repetition inside a window you set is treated as behavior and escalates to a timeout or jail. Alongside them sit ignored roles and channels, so serious discussion channels are not judged like chat channels.