CFL-KG
A Knowledge Graph for Instructional Support�in Chinese as a Foreign Language
Han Wang, Qianyu Wang, Yunshi Lan, Ye Wang, Yuanyuan Liang, and Anqi Ding
East China Normal University
KDD 2026 Special Day of AI for Education
Presenter: Han Wang
Roadmap
The talk moves from the problem to the graph design,�then to evidence from evaluation.
1
Motivation
Why CFL tasks need explicit relations
2
Representation
How CFL-KG organizes educational evidence
3
Construction and system
How the graph is built and deployed
4
Evaluation
Intrinsic quality and teacher-facing utility
5
Takeaways
When graph structure adds value
2
Motivation: CFL Tasks Need Traceable Relations
3
CFL resources are fragmented, but teacher-facing tasks require connected evidence.
Fragmented resources
Proficiency standards and level descriptors
Characters, words, grammar, and cultural content
Assessment items, rubrics, corpora, and learner errors
Why keyword search is insufficient
Teachers must explain what a question assesses
They must trace errors to knowledge and causes
They need inspectable evidence chains, not isolated results
Core challenge: connect standards, knowledge points, questions, learner errors, and cultural evidence in one reusable structure.
Research Gap and Research Questions
Research gap
Many educational KGs focus on single-course or single-standard settings.
CFL resources coexist across multiple standards and heterogeneous sources.
High-value links such as item–knowledge and error–knowledge still require explicit construction and validation.
RQ1
How can a multi-standard graph unify language knowledge, assessment resources, learner-error evidence, and cultural content?
RQ2
Can such a graph support teacher-facing tasks more effectively and more traceably than document search and keyword-only retrieval?
Main idea: treat educational relations as first-class, inspectable objects.
4
CFL-KG Overview: Data, Construction, and Functions
5
CFL-KG connects heterogeneous CFL resources, a construction pipeline, and teacher-facing functions.
Three parts are connected: data sources → KG construction pipeline → teacher-facing functions.
The graph turns heterogeneous resources into inspectable evidence paths for teaching and assessment.
Ontology: Three-Layer Educational Schema
Key design
Language layer: characters, words, grammar, radicals, and dependencies
Assessment layer: questions, templates, papers, and scoring standards
Cultural layer: cultural points, idioms, and intercultural cases
Level and Error nodes act as cross-layer hubs
The schema makes task-relevant relations explicit: level alignment, item–knowledge links, and error–knowledge links.
6
Data Construction and Graph Scale
7
Construction combines controlled automation with expert review for high-impact educational relations.
Automated construction
Parsing and cleaning from standards, textbooks, dictionaries, corpora, and assessment resources
Stable identifiers, normalization, deduplication, and provenance preservation
Rule-based and statistical candidate generation for alignments and links
Expert review
Domain experts review item–knowledge, error–knowledge, and cross-standard links
Approved, revised, pending, and rejected candidates are recorded
This protects assessment tracing and intervention planning from unsupported links
Graph scale
20K+
characters
250K+
words
1K+
grammar�points
500+
cultural�nodes
10K+
questions
2K+
reading�materials
High-impact educational relations are reviewed before entering the deployed graph.
Deployed System: From Graph Data to Teacher Workflows
8
What users can do
Browse by standard and level
Search entities and inspect node details
Explore one-hop and multi-hop relations
Trace questions back to target knowledge points
Combine graph paths with corpus evidence
Intrinsic Evaluation: Is the Graph Reliable Enough?
Before testing teacher workflows, we first audit the graph itself.
Evaluation dimensions
Standards coverage
Schema and structural quality
High-value relation correctness
Short-hop reachability
Manual samples
220 item–knowledge links
210 error–knowledge links
200 level-alignment links
Human judgment
Two raters with CFL backgrounds
Endpoint type, relation direction, source evidence, and pedagogical justification
Third adjudicator for disagreements
Inter-rater agreement is substantial, with Cohen’s κ ranging from .81 to .90 for key relation groups.
9
Intrinsic Results: Coverage, Structure, and Relation Quality
97.8%
entities aligned to standards
0.6%
schema-violating edge rate
92.0%
item–knowledge precision
89.0%
error–knowledge precision
95.0%
level-alignment precision
91.2%
reachability within ≤3 hops
Interpretation: high-value educational links reach usable precision, and evidence chains remain short enough to inspect.
10
Extrinsic Evaluation: Does CFL-KG Improve Teacher Workflows?
We compare graph-based evidence tracing with document search and keyword-only retrieval.
Participants
12 teachers or researchers
International Chinese education backgrounds
Within-subject design
Conditions
Manual document search
Keyword-only retrieval
Full CFL-KG
Task families
T1: standards-aligned preparation
T2: assessment traceability
T3: error-informed planning
Measures: task quality, evidence completeness, and completion time. The setup separates graph relations from simple retrieval.
11
Extrinsic Results: Higher Quality, Stronger Evidence, Faster Work
Task quality
0.78 → 0.91
Evidence completeness
0.71 → 0.88
Median time
245s → 109s
Largest gains appear in relation-intensive tasks, especially assessment tracing and error-informed planning.
12
Discussion: When Does Graph Structure Matter?
Graph structure matters most when teachers need alignment, diagnosis, and justification.
What CFL-KG adds
Relational verification: inspect how items, standards, errors, and examples connect.
Evidence packaging: support intervention planning and explanation.
Efficiency for relation-intensive teacher workflows.
Limitations and next steps
Sample size is modest and tasks are controlled.
Some cultural and learner-error links still need expert judgment.
Future work: confidence metadata, review status, and authentic classroom deployment.
CFL-KG supports teachers and researchers rather than autonomous high-stakes decisions about learners.
13
Takeaways
1
CFL-KG unifies language knowledge, standards, assessment resources, learner errors, and cultural content.
2
Its value comes from inspectable educational relations, not only from data scale.
3
Evaluations show improved task quality, stronger evidence completeness, and faster teacher-facing workflows.
Thank you! Q&A
Han Wang · East China Normal University
14