Benchmarks and accuracy
Tuteliq is built to catch the harm that general-purpose moderation misses: coded slang, emoji, algospeak, and deliberate filter evasion, read in the context of the conversation rather than matched against a word list. This page summarises how it performs, how that is measured, and how it compares.Headline result
In an internal benchmark of coded-language and filter-evasion cases, Tuteliq detected roughly 1.7x more of them than leading general-purpose moderation APIs, on a 319-case evasion set. The gap is largest exactly where keyword filters and general-purpose classifiers are weakest: novel algospeak, emoji payloads, and context-dependent slang that only reads as harmful given the messages around it.What is measured
The evasion benchmark is a set of messages that carry a harmful payload through obfuscation rather than plain words: coded acronyms, emoji substitution, leetspeak, homoglyphs, deliberate misspellings, and algospeak (for example “unalive”, “seggs”, “camping”). It also includes benign look-alikes that use the same surface vocabulary, so a detector cannot score well simply by flagging everything.- Positive cases: messages where an obfuscated harmful meaning is present.
- Negative controls: benign messages that share vocabulary or symbols with the positives.
- Metric: recall on the positive cases (what fraction of real evasion is caught), reported alongside behaviour on the negative controls.
Methodology
- Comparison is against three leading general-purpose moderation APIs, averaged. Vendors are not named.
- Figures reflect this internal benchmark and are not a guarantee of performance on any specific dataset or deployment.
- The advantage is a platform-wide property of how Tuteliq reads content, so it applies across every detector (grooming, romance and fraud, radicalisation, self-harm, bullying), not only to toxicity.