The per-item reasoning for all sixteen turns, published so the adjudication can be checked rather than taken on trust.Every turn is quoted verbatim with its corpus identifier, so any individual call can be contested without accepting or rejecting the set.It corrects the counts published in clause P2.6.1 of our vendor reference disclosure, against us. See section 4.
Why this document exists
A label correction that raises our own score, made by the person who set the original labels, is weak evidence on its own. We said so in the disclosure itself. Publishing the per-item reasoning is what makes it checkable: a reader can disagree with any individual call rather than accept or reject the whole adjudication. The sixteenth turn was subsequently ruled on by Dr Nicola Harding, Chief Scientific Officer and the author of the taxonomy. Section 4 sets out that ruling and the correction it forces on the published counts.What was corrected, and what it changed
TUT-100 v1 labelled sixteen adult turnsboundary_pushing. On re-adjudication
(13 August 2026, confirmed and extended 25 September 2026) those turns were found
to belong to other tactics. The consequence for the published figure:
boundary_pushing went from F1 0.316 to a per-message macro of 0.886 at
precision 1.00, because the corpus had been scoring the tactic against turns that
were never examples of it.
The withdrawn 0.316 should not be cited. Nor should 0.886 be read as an
improvement in the model: nothing about the detector changed. The labels did.
The cause, stated plainly
TUT-100 was labelled against a seven-tactic set. The Tuteliq taxonomy carries an eighth,reconnaissance, defined at conversation level. With no bucket available
for adult-initiated information-gathering, boundary_pushing was reached for as
the least-wrong option. Twelve of the sixteen turns are that error.
This is a scope decision in the original benchmark, not taxonomy drift. The
annotation guide has carried reconnaissance throughout.
1. Concealment framing, to secrecy_request (4 turns)
Adjudicated rationale. Channel concealment; an instruction to withhold the
conversation from parents. No access-advance in the utterance itself, so
boundary_pushing does not apply: concealment is the primary function, not a
delivery mechanism for a push.
All four contain the canonical channel-concealment formulation (“just between
us”, “no need to make it a big family announcement”), which the annotation guide
already carries as a fenced example under what
boundary_pushing is NOT.
2. Supervision profiling, to reconnaissance (6 turns)
Adjudicated rationale. Establishing parental oversight of the child’s device,
channels or person. Serves a later secrecy, isolation or meeting move, but the
operative function of this utterance is information-gathering. At the time of
the re-adjudication reconnaissance was scored only across a whole
conversation; it became a per-message tactic on 25 September 2026, which does
not change any label below.
3. Availability profiling, to reconnaissance (5 turns)
Adjudicated rationale. Establishing unsupervised windows. boundary_pushing
does not apply because no stated limit is being tested or eroded in the
utterance.
4. The sixteenth turn, and a correction to P2.6.1
TUT-POS-03 t9: “By the way, how old are you exactly? You sound really mature for your age.” This turn was left unadjudicated in the original pass. Working it through has produced a correction that runs against our own published figure, so we state it plainly rather than leave the item open. The definitional argument.reconnaissance is scoped to supervision,
availability, and the child’s degree of unobserved access. An age question is
none of the three. Whatever else a bare age probe is, it does not establish who
is watching the child, when the child is alone, or how observed they are.
On that reading T16 is not a probe, and the second clause, “mature for your age”,
is the first set-on example under flattery in the detection prompt and also
the relational sub-register of boundary_pushing.
We checked the reading rather than asserting it. The scope question was put
to our own tooling across six phrasings of the age probe, bare and wrapped in
praise, formal and informal. All six return not-reconnaissance on the same
ground: no supervision or availability content. Controls rule out the obvious
confound: “you’re so switched on for your age, do your parents check your phone?”
and “you’re really mature, what time are you usually on your own in the
evenings?” both register as probes. Praise does not suppress the reading. The
probe target does.
The consequence, which does not favour us. P2.6.1 states four concealment
turns and twelve probes, and that no genuine boundary_pushing anchors
remained. The tables above give four and eleven. That arithmetic reconciles
only if T16 is the twelfth probe, and on the definitional argument it is not. So
one anchor did remain that is not a probe, and P2.6.1’s claim that none remained
is wrong.
We could have closed the gap by calling T16 a probe. The taxonomy does not
support it, so the published disclosure takes the correction instead.
Confirmed by the Chief Scientific Officer. The ruling that an age question
is not reconnaissance is Dr Nicola Harding’s, the author of the taxonomy, not
ours alone. The correction below therefore rests on the taxonomy author’s own
reading of it rather than on our interpretation.
What is still open. Whether T16 is flattery, relational
boundary_pushing, or both firing together is a labelling decision that belongs
to the annotator who owns the corpus, and it is not settled here. What is settled
is the narrower point that it is not reconnaissance, and therefore that the
published count of twelve is wrong.