Skip to main content
Tuteliq’s own disclosure against our framework. We wrote the questions. Here are our answers, including where we fall short.This is a vendor reference disclosure under clause P0.2.5 of TUT-CSDT-001, the Child Safety Detection Transparency Framework. It confers no conformance level, because levels attach to platforms rather than to detection systems. It covers the vendor-side clauses only: the properties of the detection system, which only we can evidence. The platform-side clauses, covering how a given deployment is configured and what happens to a flagged user, belong to the platform using us and cannot be discharged by citing this document.Our grooming detection shows zero false positives across 450 benign hard-negative turns and per-tactic recall between 0.913 and 1.00 on a held-out expert corpus. These figures are self-measured; independent evaluation is in progress via COVA-X. Precision is the metric most sensitive to base rate, so compute your own figure from your own prevalence rather than citing ours: at 1% turn-level prevalence it is roughly 0.54.Corrected 18 August 2026. An earlier version of this page described this evaluation as conversation-level and reported a 1% prevalence precision of roughly 0.07. Both were wrong. The details are below, and the correction moves every figure in our favour.

The short version

The machine-readable form is at csdt-disclosure.json. Where this page and that file differ, the JSON is authoritative.

What we measured

Grooming, behavioural and multi-turn. Measured at turn level, not conversation level. The positive class is 1,082 expert-annotated anchor turns drawn from 350 harmful trajectories. The negative class is 450 adult turns drawn from 150 benign hard-negative trajectories written in adjacent registers such as sex education, so these are same-domain hard cases rather than unrelated benign text. Zero false positives across all seven tactics on those 450 turns, giving a 95% CI on the turn-level false-positive rate of [0, 0.0085]. Precision 1.00 (95% CI 0.996 to 1.0). Micro recall 0.968 (95% CI 0.955 to 0.977) across all 1,082 anchors. Macro F1 0.987 across the seven tactics, which is the figure quoted on our model card and in external materials. Per tactic: flattery 1.000; photo_request 1.000; gift_giving 1.000; secrecy_request 0.994; meeting_request 0.988; isolation 0.975; boundary_pushing 0.955. F1 sits above recall on this set because precision is 1.000, so both are given rather than only the lower. Per-tactic recall with 95% intervals: flattery 1.00 [0.875, 1.0]; secrecy_request 0.989 [0.972, 0.996]; isolation 0.951 [0.910, 0.974]; boundary_pushing 0.913 [0.871, 0.943]; photo_request 1.00 [0.980, 1.0]; meeting_request 0.976 [0.877, 0.996]; gift_giving 1.00 [0.912, 1.0]. The macro recall of 0.976 quoted elsewhere is the unweighted mean of those seven, so it weights flattery (27 anchors) equally with secrecy_request (362). Neither macro nor micro is a binomial proportion; the per-tactic intervals are the ones to use. Corrected 18 August 2026. An earlier version described this evaluation as conversation-level, with 500 positive conversations against 27 negative. The underlying measurement was always turn-level, as stated above. Every corrected figure moves in our favour, which is why the change is recorded rather than applied silently. A separate set of 27 benign control conversations exists and is not part of the figures above. Fraud and social engineering. Recall 0.98 (147 of 150) on the COVA-X scam corpus, used with the authorisation of its authors. Precision is not reported for this category: the evaluated set contains scam conversations only, with no benign class, so there is no false-positive denominator. Sextortion and coercive control are now separately measured. Both against the 150 benign hard-negative trajectories as a shared negative class, with zero false positives. Sextortion: 85 of 85 conversations detected, recall 1.000 (95% CI 0.957 to 1.000), precision 1.000 (95% CI 0.958 to 1.000). Coercive control: 50 of 50, recall 1.000 (95% CI 0.929 to 1.000), precision 1.000 (95% CI 0.930 to 1.000). Read the unit carefully. These are conversation-level figures, meaning the system flagged the conversation at all. That is a coarser task than the per-turn tactic labelling behind the grooming numbers above, and the two must not be compared or combined. Corrected 18 August 2026. Both were previously recorded as not measured, on the basis that no figure was attributable to either category alone. The held-out corpus carries a harm type per trajectory and the conversation-level verdicts were archived, so category-specific figures were always derivable and had simply never been aggregated.

The public benchmark

PAN-2012 Sexual Predator Identification (CLEF 2012, Inches and Crestani, DOI 10.5281/zenodo.3713280): precision 0.549 (95% CI 0.434 to 0.660), recall 0.780 (95% CI 0.648 to 0.872). We publish this because it is the most comparable public number we have. It should be read with the limitations below, which are substantial. It should not be read as a production false-positive rate. PAN-2012’s positive class comes from decoy operations in which adults posed as children, so it measures detection of an adult performing a child rather than of a child. Positive and negative classes were drawn from different sources, so a classifier can score well on artefacts of the source. It is English-only and of 2012 vintage. Most consequentially: the PAN-2012 “clean” set is raw adult IRC chat, and on our own inspection 82% of it (41 of 50) contains grooming-pattern tactics genuinely present between adults. The benign class is not benign. The overlap is checkable rather than assertable. We flagged 32 of the 50. All 32 fall inside those 41, being 78% of the tactic-bearing subset and none outside it. The two figures are dimensionally consistent, which they did not have to be. Held to the same standard as the adjudication below: that judgement was made by Tuteliq, after the run, with no blinding protocol applied or recorded. Whoever judged which conversations carried tactics could see which had been flagged, so perfect containment is what this method would tend to produce whether or not it were true. The residual subset adjudicated as genuinely clean is nine conversations, none flagged, which bounds the false-positive rate on truly benign adult text at 0 to 0.299. That interval, not the containment argument, is the honest measure. The caution that follows, and platforms should carry it into their own P6.10: if adult-to-adult conversation reliably carries these tactics, that is a production false-positive characteristic and not merely a benchmark artefact. A deployment analysing adult-to-adult traffic should expect flags of this kind. Our curated benign sets return zero false positives across 27 control conversations and across 450 benign turns, but those are child-context conversations, not adult IRC.

How the corpus was built

External consultants in criminology, child psychology and linguistics supplied the material and ran validation tests in their own domains against the taxonomy and the detection behaviour, with their findings acted on. Dr Nicola Harding, our Chief Scientific Officer, a criminologist and Specialist Advisor to the UK Home Office, led the programme and labelled and validated the corpus. It is grounded in published frameworks in the grooming and child-sexual-exploitation literature. The conversations are fabricated rather than collected from real people, and expansion was AI-assisted under expert direction with expert validation of the output. The experts declined to work from real case material so that no victim’s personal data could enter the corpus, be retained, or leak through the model. No row of training or evaluation data traces back to a real child. Labels on our own corpora were applied by paid independent consultants in criminology, child psychology and linguistics, each working within their own discipline. Dr Harding led the programme, validated the labelling and signed it off; she did not apply every label. The consultants’ identities are confidential under their consultancy agreements. No inter-annotator agreement figure exists for those corpora, and none is recoverable, but the reason is provenance recording rather than rater count. Annotator identity was never captured against individual labels, so although our labelling protocol specified a 15% double-coded overlap targeting Cohen’s kappa of 0.7 or better, the double-coded subset cannot be reconstructed. Labelling from here records a pseudonymous identifier per label, which supports agreement statistics without disclosing who the consultants are. None of this makes our labelling independent. Every contributor was paid by Tuteliq or advises it, and our own definition of independence excludes anyone with a financial, employment or contractual relationship. It does not apply to the public benchmarks below, whose labels are not ours. The expert-authored material is the English backbone. Multilingual coverage is derived from it by translation and in-language generation with native-speaker review, rather than authored independently in each language.

Where we fall short

Our labels are not independent, and the expert panel does not change that. Under clause P0.1, a false positive or negative determination requires a reviewer with no financial, employment or contractual relationship with the vendor. Our evaluation labels come from our Chief Scientific Officer, who is a co-founder and shareholder, and the consultants above were engaged and paid by us. A credentialled panel is a statement about expertise, not about independence, and the two are easily conflated. Every figure above rests on ground truth we produced ourselves. Inter-annotator agreement has not been measured on our own corpora. A single annotator labelled them, so no agreement statistic exists to compute. This is recorded as not measured rather than withheld: there is no figure being held back. The study is designed and the sample drawn; external labellers are being commissioned. This applies to the corpora we authored. It does not apply to our two public benchmark runs, where the labels are not ours: PAN-2012 is an externally constructed and externally labelled corpus from the CLEF 2012 shared task, and the COVA-X scam corpus is externally constructed and used with its authors’ authorisation. Under clause P0.1 a public benchmark run sits above self-conducted evaluation in evidential weight, because the vendor did not choose the cases or set the answers. It sits below independent evaluation, because we executed the run. One boundary within that: our re-adjudication of the PAN-2012 clean set is our own judgement, not the corpus authors’, and it is described as such below. Ten of twelve covered categories have no published figures. Detection runs for all of them. Performance has not been measured per category for most. Per-language recall is measured for 26 languages; per-language precision is not. The held-out multilingual evaluation covers 728 expert-pattern cases written in language rather than translated: 28 per language, seven tactics, four cases per tactic per language. Exact-tactic recall by language, ten at 1.000 and nineteen of twenty-six at 0.95 or above: Greek 1.000; Spanish 1.000; Finnish 1.000; French 1.000; Latvian 1.000; Dutch 1.000; Portuguese 1.000; Slovenian 1.000; Swedish 1.000; Ukrainian 1.000; Danish 0.964; German 0.964; Irish 0.964; Croatian 0.964; Italian 0.964; Norwegian 0.964; Polish 0.964; Romanian 0.964; Turkish 0.964; Estonian 0.929; Hungarian 0.929; Lithuanian 0.929; Maltese 0.929; Bulgarian 0.893; Czech 0.893; Slovak 0.893; By tactic across all 26 languages: secrecy_request 1.0; gift_giving 0.99; meeting_request 0.99; photo_request 0.981; isolation 0.971; boundary_pushing 0.933; flattery 0.885; Overall exact-tactic recall 0.964 across the 728 cases, and zero silent misses: the system never returned an absence of risk on a case carrying a tactic. Where it erred it named a different tactic, which a reviewer sees and can correct. No confidence intervals are computed for these, and the reason is arithmetic rather than reticence: four observations per tactic per language will not support one. Treat them as point estimates on 28 cases. What does not exist is per-language precision. Every one of the 728 cases is a positive, so the set holds no per-language negatives from which a false-positive rate could be computed. That is the substantive remaining gap, and it is what would move a language to stable. Corrected 18 August 2026. An earlier version said twenty-six languages including English were covered, leaving six unevaluated. English is not in that matrix; it is measured separately. The correct split is twenty-seven measured and five unmeasured: Chinese, Japanese, Korean, Arabic and Russian. Five of the thirty-two have no evaluation of their own. Chinese, Japanese, Korean, Arabic and Russian carry culture-specific guidelines and run the same multilingual model. Their beta tier rests on two in-language smoke tests run before promotion, not on evaluation: a single grooming payload run eight times in each of Chinese, Japanese, Korean and Arabic, and a two-case probe run six times in each of Arabic and Russian, all correct. Those are repetition counts on one to two cases, against the production service rather than the model artifact, with no independent ground truth. They show the system responds in these scripts; they measure nothing, and we cite no figure from them. Arabic script has been exercised but diglossia and dialect have not. Twenty-seven measured plus five unmeasured is thirty-two. Detection nevertheless runs everywhere (clause P4.11). There is no language in which detection is disabled, sampled, gated behind a flag, or run at reduced escalation. The performance tier describes what has been measured, not what is switched on, and we would not withdraw protection from a language in order to avoid declaring a tier for it. Subgroup equity is unmeasured. We have not measured differential performance across languages, dialects, demographics or content types. Differential false-positive rates across protected groups are a discrimination risk and a child-rights issue. Known-material hash matching reaches IWF members only. We hold IWF membership and hash matching runs in the request path, but the IWF Hash List cannot be sublicensed, so it is applied only for a customer that holds its own membership. We verify that membership out of band and resolve the entitlement server-side; a caller cannot assert it, because a caller that could would be able to obtain a licensed dataset it does not hold, or to switch matching off. Customers without their own membership receive compositional analysis only, and the response records that the check did not run rather than that the content was checked and cleared.

One correction we made to ourselves

An earlier figure of 0.316 for the boundary_pushing tactic has been withdrawn and should not be cited. On expert re-adjudication, all sixteen anchor turns behind that figure were found to be mislabelled: four were concealment framing belonging to secrecy_request, and twelve were supervision probes belonging to reconnaissance, an arc-level tactic that the seven-tactic per-message benchmark did not carry. No genuine boundary_pushing anchors remained, so the tactic was untested by that benchmark rather than failed. Corrected, from identical predictions with no retraining: per-message macro F1 0.886 at precision 1.00. On a purpose-built set of genuine anchors, boundary_pushing scores F1 0.750. How that re-adjudication was conducted, because it bears on the weight it carries: it was performed by the same annotator who produced the original labels, unblinded, on a corpus for which no inter-annotator agreement statistic exists. The per-item reasoning for all sixteen turns will be published so the adjudication can be assessed rather than taken on trust. We record this because a withdrawn number that stays in circulation is worse than a bad number that is explained.

Determinism of the detection path

Until 21 August 2026 our detection endpoints decoded with sampling enabled. The setting was inherited from a general-purpose client default written for prose generation. No detection endpoint ever overrode it, so the value was never a decision anyone made about classification, and it was unreachable through the intended parameter: a request for deterministic decoding was treated as absent and the default substituted. Detection now decodes deterministically. What the variance was. Ten identical calls per case, on the production service, across a fixed panel spanning the decision boundary: Variance concentrated on cases near the decision boundary; cases away from it were stable across every repeat. That is the profile sampling predicts, and it is why the effect was invisible to every evaluation we had run: all of them were single runs. What this means for a caller. A retried request could previously return a different routing decision for identical input. No platform reported this to us, and we did not find it ourselves until we tested for it. Any figure we published before this date is a single draw from that distribution rather than a stable point estimate. That applies to our favourable figures as much as our unfavourable ones. What determinism does not fix. It removes the sampling component. It does not make output byte-identical, because the model is served behind infrastructure whose batching we do not control. We keep repeated-measures testing as a standing check rather than treating the matter as closed, and we will not publish a single-run figure as a point estimate again. How the fix was validated. As a two-by-two factorial over decoding mode and detection configuration, seven cases by fifteen repeats in each of four cells. We ran a factorial because our first comparison changed two things at once and could not attribute the result to either. It found a significant interaction: the configuration change was harmful under sampling and beneficial under deterministic decoding, so the two had to ship together or not at all. Per-case differences carry Fisher exact tests. The panel is seven fixed cases, so those tests support statements about those cases and not a generalisation to production traffic, and we do not make one. A defect the same work exposed, now resolved. The panel included crisis conversations written to require escalation. One class was under-routed: where the behaviour indicating acute risk was present but expressed in understated language, it escalated in 2 of 15 repeats, while the same behaviour expressed explicitly escalated in 15 of 15. The detector was keying on explicit phrasing rather than on the underlying behaviour, which is the wrong way round, because understatement is common in the population this exists to protect. After the change the understated case escalates in 15 of 15. We report it because anyone assessing our recall should know it was measured on material that skews explicit. What we changed about how we evaluate. Comparisons between system versions previously reported routing decisions only. That is insufficient: a case can keep its routing while losing a harm category, which removes it from category-keyed customer policy rules and from escalation eligibility with nothing visible changing. We found one such case in this work. Version comparisons now diff categories as well as routing, and treat the loss of a safety-relevant category as a regression.

Check it yourself

The disclosure validates against the published schema with zero errors. You can reproduce that:
The validator computes the conformance level from the disclosure’s own contents rather than trusting what it claims, and reports every item recorded as withheld or not measured. One thing to read carefully before using any figure here: every number was measured against the model artifact directly, not through the production API. The API adds prompt construction, input normalisation and threshold handling on top of the model, none of which these figures exercise. Treat them as an upper bound on what the deployed path delivers.
This disclosure is reviewed at least annually and on material change to the detection system, the harm categories covered, the operating points used in production, or data-handling arrangements. A level claim is not made from a disclosure more than twelve months old.

Revision history