1 of 15

Redact or Keep?

A Fully Local AI Cascade for Educational Dialogue De-Identification

Haocheng Zhang, Zhuqian Zhou, Kirk Vanacore, Bakhtawar Ahtisham, René F. Kizilcec

KDD 2026 Workshop on AI for Education · August 2026 · Jeju, Korea

2 of 15

MOTIVATION

Educational Dialogue Is Valuable — and Sensitive

Transcripts that capture how students learn also capture what they disclose about themselves. De-identification is required before data can be shared for research.

KEEP

“Let’s practice Riemann sums today.”

Curricular context: “Jordan has 15 apples.” Over-redaction destroys research value.

REDACT

“Hi, I’m Maria Riemann!”

Identity-bearing context: “Thanks, Jordan!” Missing it is a privacy failure.

Same surface form, opposite decisions — context is what separates them.

2

3 of 15

WHY THIS MATTERS

Tutoring Data Captures Learning at Human Scale

A single session can combine conversation, equations, whiteboard work, audio, and video—rich evidence for research, and rich privacy risk.

Digital tutoring produces detailed, turn-by-turn learning traces.

The same infrastructure increasingly captures multimodal interaction.

Representative tutoring contexts · official National Tutoring Observatory photography

3

4 of 15

The National Tutoring Observatory (NTO) is a tutoring-focused AI research infrastructure initiative that aims to:

  • Safely build and share the largest repository of tutoring data
  • Create open-source tools for analyzing tutoring sessions at scale
  • Connect tutoring practices to student engagement and learning outcomes

Together, these efforts support research and development aimed at identifying what makes tutoring effective —and improving human and AI tutoring at scale

5 of 15

MOTIVATION

Current Methods Trade Governance for Accuracy

Commercial LLM APIs

Model context across a dialogue

But require data egress: student transcripts leave institutional control — a governance and compliance risk

Local NER Systems

spaCy and Presidio are fast and fully local

But generic name detection over-redacts curricular content and offers no Redact/Keep decision

A practical system must balance four goals: detection accuracy · curricular preservation · privacy governance · deployment cost

5

6 of 15

RESEARCH QUESTIONS

1

Cascade vs. LLM-only

Holding Gemma 31B constant, does per-span cascade review outperform single-pass full-dialogue extraction?

2

Robustness on ambiguous names

When real student names overlap with curricular-content names, which reviewer remains robust under deliberate ambiguity shift?

3

Deployment footprint

Can training and inference stay on-device — and what accuracy, memory, and throughput tradeoffs result?

6

7 of 15

APPROACH

A Fully Local Two-Stage Cascade

INPUT

Tutoring dialogue

STAGE 1

PROPOSE SPANS

High recall

DeBERTa + ModernBERT

CANDIDATES

Sarah

NAME · EMAIL · …

STAGE 2

REVIEW IN CONTEXT

One decision per span

Gemma4 31B

REDACT

or

KEEP

OUTPUT

Apply local policy

DIRECT RULES email · phone · URL · identifying number

High-precision identifiers bypass review and redact automatically.

7

8 of 15

APPROACH

From Entity Recognition to Privacy Triage

Jordan has 15 apples…”

KEEP

“Thanks, Taylor!”

REDACT

“Hi, I’m Morgan.”

REDACT

One constrained decision per candidate: is this real PII in context? When uncertain → Redact.

Three reviewers, same candidates

  • RoBERTa — 125M, full fine-tune
  • Gemma E4B — 4B, LoRA
  • Gemma 31B — LoRA

Candidate proposal is more tractable; contextual review determines privacy risk.

8

9 of 15

METHOD

Math Tutoring Transcripts from Two Platforms

Platform A

Short, question-focused K–12 tutoring dialogues

Platform B

Longer scheduled 1-to-1 lessons with extended multi-turn interaction

1,611

dialogues

manually annotated

by 3 annotators

4.36M

tokens across train / val / test splits

97.1%

of PII turns contain a NAME — the ambiguity that matters

130

challenge dialogues where real & curricular names co-occur

Evaluation: span-overlap precision / recall / F1 · 200 canonical test dialogues · 46 zero-PII dialogues

9

10 of 15

RESULTS · RQ1

The 31B Cascade Leads

SAME GEMMA 31B

.767

FULL-DIALOGUE EXTRACTION

↓ TASK FORMULATION ↓

.958

PER-SPAN CASCADE REVIEW

+0.191 macro F1

10

11 of 15

RESULTS · RQ2

The Challenge Set Reveals the Reviewer Gap

The canonical set makes reviewers look similar.

Challenge set macro F1:

local 31B cascade .932 vs. Gemini 3.1 Pro .873.

11

12 of 15

RESULTS · RQ3

One Laptop Is Enough for Local Batch Processing

18 GB

31B cascade inference memory

(.958 macro F1)

0.09/s

reviewer throughput

on the canonical workload

4.3 h

training time

per seed

12

13 of 15

CONCLUSIONS

Formulation Matters More Than Scale

01

Reframing beats scaling. Holding Gemma 31B constant, cascade review lifts macro F1 from .767 to .958.

02

De-identification is privacy triage, not NER. Recall-first proposal finds candidates; contextual review decides whether each span is real PII.

03

Governance without accuracy loss. The fully local cascade outperforms Gemini 3.1 Pro (.958 vs. .706) with zero transcript egress.

13

14 of 15

LIMITATIONS & FUTURE WORK

Scope and What Comes Next

  • Domain scope: English math tutoring from two platforms; other subjects, languages, and age groups remain untested
  • Latency: 31B review is heavy: ~1,000–1,500-session batches on one large AWS EC2 machine. Prefix caching, FlashAttention, continuous batching, and paged KV cache have already made processing 4× faster

Future work: generalize the cascade to multimodal models for de-identifying video tutoring sessions

14

15 of 15

Questions?