Redact or Keep?
A Fully Local AI Cascade for Educational Dialogue De-Identification
Haocheng Zhang, Zhuqian Zhou, Kirk Vanacore, Bakhtawar Ahtisham, René F. Kizilcec
KDD 2026 Workshop on AI for Education · August 2026 · Jeju, Korea
MOTIVATION
Educational Dialogue Is Valuable — and Sensitive
Transcripts that capture how students learn also capture what they disclose about themselves. De-identification is required before data can be shared for research.
KEEP
“Let’s practice Riemann sums today.”
Curricular context: “Jordan has 15 apples.” Over-redaction destroys research value.
REDACT
“Hi, I’m Maria Riemann!”
Identity-bearing context: “Thanks, Jordan!” Missing it is a privacy failure.
Same surface form, opposite decisions — context is what separates them.
2
WHY THIS MATTERS
Tutoring Data Captures Learning at Human Scale
A single session can combine conversation, equations, whiteboard work, audio, and video—rich evidence for research, and rich privacy risk.
Digital tutoring produces detailed, turn-by-turn learning traces.
The same infrastructure increasingly captures multimodal interaction.
Representative tutoring contexts · official National Tutoring Observatory photography
3
The National Tutoring Observatory (NTO) is a tutoring-focused AI research infrastructure initiative that aims to:
Together, these efforts support research and development aimed at identifying what makes tutoring effective —and improving human and AI tutoring at scale
MOTIVATION
Current Methods Trade Governance for Accuracy
Commercial LLM APIs
Model context across a dialogue
But require data egress: student transcripts leave institutional control — a governance and compliance risk
Local NER Systems
spaCy and Presidio are fast and fully local
But generic name detection over-redacts curricular content and offers no Redact/Keep decision
A practical system must balance four goals: detection accuracy · curricular preservation · privacy governance · deployment cost
5
RESEARCH QUESTIONS
1
Cascade vs. LLM-only
Holding Gemma 31B constant, does per-span cascade review outperform single-pass full-dialogue extraction?
2
Robustness on ambiguous names
When real student names overlap with curricular-content names, which reviewer remains robust under deliberate ambiguity shift?
3
Deployment footprint
Can training and inference stay on-device — and what accuracy, memory, and throughput tradeoffs result?
6
APPROACH
A Fully Local Two-Stage Cascade
INPUT
Tutoring dialogue
STAGE 1
PROPOSE SPANS
High recall
DeBERTa + ModernBERT
CANDIDATES
Sarah
NAME · EMAIL · …
STAGE 2
REVIEW IN CONTEXT
One decision per span
Gemma4 31B
REDACT
or
KEEP
OUTPUT
Apply local policy
DIRECT RULES email · phone · URL · identifying number
High-precision identifiers bypass review and redact automatically.
7
APPROACH
From Entity Recognition to Privacy Triage
“Jordan has 15 apples…”
KEEP
“Thanks, Taylor!”
REDACT
“Hi, I’m Morgan.”
REDACT
One constrained decision per candidate: is this real PII in context? When uncertain → Redact.
Three reviewers, same candidates
Candidate proposal is more tractable; contextual review determines privacy risk.
8
METHOD
Math Tutoring Transcripts from Two Platforms
Platform A
Short, question-focused K–12 tutoring dialogues
Platform B
Longer scheduled 1-to-1 lessons with extended multi-turn interaction
1,611
dialogues
manually annotated
by 3 annotators
4.36M
tokens across train / val / test splits
97.1%
of PII turns contain a NAME — the ambiguity that matters
130
challenge dialogues where real & curricular names co-occur
Evaluation: span-overlap precision / recall / F1 · 200 canonical test dialogues · 46 zero-PII dialogues
9
RESULTS · RQ1
The 31B Cascade Leads
SAME GEMMA 31B
.767
FULL-DIALOGUE EXTRACTION
↓ TASK FORMULATION ↓
.958
PER-SPAN CASCADE REVIEW
+0.191 macro F1
10
RESULTS · RQ2
The Challenge Set Reveals the Reviewer Gap
The canonical set makes reviewers look similar.
Challenge set macro F1:
local 31B cascade .932 vs. Gemini 3.1 Pro .873.
11
RESULTS · RQ3
One Laptop Is Enough for Local Batch Processing
18 GB
31B cascade inference memory
(.958 macro F1)
0.09/s
reviewer throughput
on the canonical workload
4.3 h
training time
per seed
12
CONCLUSIONS
Formulation Matters More Than Scale
01
Reframing beats scaling. Holding Gemma 31B constant, cascade review lifts macro F1 from .767 to .958.
02
De-identification is privacy triage, not NER. Recall-first proposal finds candidates; contextual review decides whether each span is real PII.
03
Governance without accuracy loss. The fully local cascade outperforms Gemini 3.1 Pro (.958 vs. .706) with zero transcript egress.
13
LIMITATIONS & FUTURE WORK
Scope and What Comes Next
Future work: generalize the cascade to multimodal models for de-identifying video tutoring sessions
14
Questions?