Skip to main content
The per-item reasoning for all sixteen turns, published so the adjudication can be checked rather than taken on trust.Every turn is quoted verbatim with its corpus identifier, so any individual call can be contested without accepting or rejecting the set.It corrects the counts published in clause P2.6.1 of our vendor reference disclosure, against us. See section 4.

Why this document exists

A label correction that raises our own score, made by the person who set the original labels, is weak evidence on its own. We said so in the disclosure itself. Publishing the per-item reasoning is what makes it checkable: a reader can disagree with any individual call rather than accept or reject the whole adjudication. The sixteenth turn was subsequently ruled on by Dr Nicola Harding, Chief Scientific Officer and the author of the taxonomy. Section 4 sets out that ruling and the correction it forces on the published counts.

What was corrected, and what it changed

TUT-100 v1 labelled sixteen adult turns boundary_pushing. On re-adjudication (13 August 2026, confirmed and extended 25 September 2026) those turns were found to belong to other tactics. The consequence for the published figure: boundary_pushing went from F1 0.316 to a per-message macro of 0.886 at precision 1.00, because the corpus had been scoring the tactic against turns that were never examples of it. The withdrawn 0.316 should not be cited. Nor should 0.886 be read as an improvement in the model: nothing about the detector changed. The labels did.

The cause, stated plainly

TUT-100 was labelled against a seven-tactic set. The Tuteliq taxonomy carries an eighth, reconnaissance, defined at conversation level. With no bucket available for adult-initiated information-gathering, boundary_pushing was reached for as the least-wrong option. Twelve of the sixteen turns are that error. This is a scope decision in the original benchmark, not taxonomy drift. The annotation guide has carried reconnaissance throughout.

1. Concealment framing, to secrecy_request (4 turns)

Adjudicated rationale. Channel concealment; an instruction to withhold the conversation from parents. No access-advance in the utterance itself, so boundary_pushing does not apply: concealment is the primary function, not a delivery mechanism for a push. All four contain the canonical channel-concealment formulation (“just between us”, “no need to make it a big family announcement”), which the annotation guide already carries as a fenced example under what boundary_pushing is NOT.

2. Supervision profiling, to reconnaissance (6 turns)

Adjudicated rationale. Establishing parental oversight of the child’s device, channels or person. Serves a later secrecy, isolation or meeting move, but the operative function of this utterance is information-gathering. At the time of the re-adjudication reconnaissance was scored only across a whole conversation; it became a per-message tactic on 25 September 2026, which does not change any label below.

3. Availability profiling, to reconnaissance (5 turns)

Adjudicated rationale. Establishing unsupervised windows. boundary_pushing does not apply because no stated limit is being tested or eroded in the utterance.

4. The sixteenth turn, and a correction to P2.6.1

TUT-POS-03 t9: “By the way, how old are you exactly? You sound really mature for your age.” This turn was left unadjudicated in the original pass. Working it through has produced a correction that runs against our own published figure, so we state it plainly rather than leave the item open. The definitional argument. reconnaissance is scoped to supervision, availability, and the child’s degree of unobserved access. An age question is none of the three. Whatever else a bare age probe is, it does not establish who is watching the child, when the child is alone, or how observed they are. On that reading T16 is not a probe, and the second clause, “mature for your age”, is the first set-on example under flattery in the detection prompt and also the relational sub-register of boundary_pushing. We checked the reading rather than asserting it. The scope question was put to our own tooling across six phrasings of the age probe, bare and wrapped in praise, formal and informal. All six return not-reconnaissance on the same ground: no supervision or availability content. Controls rule out the obvious confound: “you’re so switched on for your age, do your parents check your phone?” and “you’re really mature, what time are you usually on your own in the evenings?” both register as probes. Praise does not suppress the reading. The probe target does. The consequence, which does not favour us. P2.6.1 states four concealment turns and twelve probes, and that no genuine boundary_pushing anchors remained. The tables above give four and eleven. That arithmetic reconciles only if T16 is the twelfth probe, and on the definitional argument it is not. So one anchor did remain that is not a probe, and P2.6.1’s claim that none remained is wrong. We could have closed the gap by calling T16 a probe. The taxonomy does not support it, so the published disclosure takes the correction instead. Confirmed by the Chief Scientific Officer. The ruling that an age question is not reconnaissance is Dr Nicola Harding’s, the author of the taxonomy, not ours alone. The correction below therefore rests on the taxonomy author’s own reading of it rather than on our interpretation. What is still open. Whether T16 is flattery, relational boundary_pushing, or both firing together is a labelling decision that belongs to the annotator who owns the corpus, and it is not settled here. What is settled is the narrower point that it is not reconnaissance, and therefore that the published count of twelve is wrong.

How to disagree with this

Every turn above is quoted verbatim with its file and turn number, so any individual call can be contested without accepting or rejecting the set. If you believe a specific turn is wrongly re-adjudicated, the corpus identifiers are sufficient to locate it. The standing caveat in P2.6.1 remains: the original re-adjudication was performed by the same annotator who produced the original labels, unblinded, on a corpus for which no inter-annotator agreement statistic exists. Publishing the reasoning does not remove that limitation. It makes it checkable.