Skip to main content

Benchmarks and accuracy

Tuteliq is built to catch the harm that general-purpose moderation misses: coded slang, emoji, algospeak, and deliberate filter evasion, read in the context of the conversation rather than matched against a word list. This page describes what we measure and how.

Where the numbers are

We previously published a comparative recall figure against general-purpose moderation APIs. That figure is withdrawn pending re-validation: later runs on the same case set materially changed our own recall, and the vendor baselines have not been re-run against them, so no honest ratio can be quoted right now. Rather than restate a number we cannot currently stand behind, this page describes what we measure and how. Updated figures will be published once the comparison has been re-run end to end. If you are evaluating Tuteliq and need evidence now, the useful path is a benchmark on your own content: we will help you assemble a labelled set from your platform and run it, which is a better predictor for your deployment than any figure of ours.

What is measured

The evasion benchmark is a set of messages that carry a harmful payload through obfuscation rather than plain words: coded acronyms, emoji substitution, leetspeak, homoglyphs, deliberate misspellings, and algospeak (for example “unalive”, “seggs”, “camping”). It also includes benign look-alikes that use the same surface vocabulary, so a detector cannot score well simply by flagging everything.
  • Positive cases: messages where an obfuscated harmful meaning is present.
  • Negative controls: benign messages that share vocabulary or symbols with the positives.
  • Metric: recall on the positive cases (what fraction of real evasion is caught), reported alongside behaviour on the negative controls.

Methodology

  • Comparison is against three leading general-purpose moderation APIs, averaged. Vendors are not named.
  • Figures reflect this internal benchmark and are not a guarantee of performance on any specific dataset or deployment.
  • The advantage is a platform-wide property of how Tuteliq reads content, so it applies across every detector (grooming, romance and fraud, radicalisation, self-harm, bullying), not only to toxicity.

Why context beats word lists

General-purpose moderation scores a single message against a fixed vocabulary. Tuteliq scores the interaction: who is targeting whom, whether it is reciprocal, whether it is escalating, and whether the wording performs a harmful function even when no banned word appears. That is why it separates gaming trash-talk from targeted harassment, and why obfuscation that walks straight past a word list still gets scored. See Bullying and toxicity detection and the Prescreen lexicon for how the coded-term layer corroborates the behavioural signal.

Multilingual coverage

Detection runs in 32 languages (English stable, all 24 EU official languages plus Ukrainian, Norwegian, Turkish, Chinese, Japanese, Korean, Arabic, and Russian in beta), with code-switching support and language-specific coverage of slang and evasion. See Languages for the full list.

Frequently asked questions

How accurate is Tuteliq at detecting coded slang and filter evasion? Comparative figures are being re-validated and are not published at present. What we can describe is the method: coded terms are maintained as an evolving lexicon, and every case is scored in the context of the surrounding conversation rather than matched against a fixed list, which is what makes novel algospeak and emoji payloads detectable at all. How is the benchmark measured? It uses obfuscated harmful messages (coded acronyms, emoji, leetspeak, algospeak, misspellings) plus benign look-alike controls, and reports recall on the positives alongside behaviour on the controls. Comparison is against three leading general-purpose moderation APIs, averaged and unnamed. Does the coded-language advantage apply to more than bullying? Yes. It is a platform-wide capability that applies to every detector, because the same obfuscation tactics appear across every kind of harm. How many languages does Tuteliq support? 32: English is stable and 31 more are in beta, including all 24 EU official languages plus Ukrainian, Norwegian, Turkish, Chinese, Japanese, Korean, Arabic, and Russian.