Benchmarks and accuracy
Tuteliq is built to catch the harm that general-purpose moderation misses: coded slang, emoji, algospeak, and deliberate filter evasion, read in the context of the conversation rather than matched against a word list. This page describes what we measure and how.Where the numbers are
We previously published a comparative recall figure against general-purpose moderation APIs. That figure is withdrawn pending re-validation: later runs on the same case set materially changed our own recall, and the vendor baselines have not been re-run against them, so no honest ratio can be quoted right now. Rather than restate a number we cannot currently stand behind, this page describes what we measure and how. Updated figures will be published once the comparison has been re-run end to end. If you are evaluating Tuteliq and need evidence now, the useful path is a benchmark on your own content: we will help you assemble a labelled set from your platform and run it, which is a better predictor for your deployment than any figure of ours.What is measured
The evasion benchmark is a set of messages that carry a harmful payload through obfuscation rather than plain words: coded acronyms, emoji substitution, leetspeak, homoglyphs, deliberate misspellings, and algospeak (for example “unalive”, “seggs”, “camping”). It also includes benign look-alikes that use the same surface vocabulary, so a detector cannot score well simply by flagging everything.- Positive cases: messages where an obfuscated harmful meaning is present.
- Negative controls: benign messages that share vocabulary or symbols with the positives.
- Metric: recall on the positive cases (what fraction of real evasion is caught), reported alongside behaviour on the negative controls.
Methodology
- Comparison is against three leading general-purpose moderation APIs, averaged. Vendors are not named.
- Figures reflect this internal benchmark and are not a guarantee of performance on any specific dataset or deployment.
- The advantage is a platform-wide property of how Tuteliq reads content, so it applies across every detector (grooming, romance and fraud, radicalisation, self-harm, bullying), not only to toxicity.