1 of 14

CFL-KG

A Knowledge Graph for Instructional Support�in Chinese as a Foreign Language

Han Wang, Qianyu Wang, Yunshi Lan, Ye Wang, Yuanyuan Liang, and Anqi Ding

East China Normal University

KDD 2026 Special Day of AI for Education

Presenter: Han Wang

2 of 14

Roadmap

The talk moves from the problem to the graph design,�then to evidence from evaluation.

1

Motivation

Why CFL tasks need explicit relations

2

Representation

How CFL-KG organizes educational evidence

3

Construction and system

How the graph is built and deployed

4

Evaluation

Intrinsic quality and teacher-facing utility

5

Takeaways

When graph structure adds value

2

3 of 14

Motivation: CFL Tasks Need Traceable Relations

3

CFL resources are fragmented, but teacher-facing tasks require connected evidence.

Fragmented resources

Proficiency standards and level descriptors

Characters, words, grammar, and cultural content

Assessment items, rubrics, corpora, and learner errors

Why keyword search is insufficient

Teachers must explain what a question assesses

They must trace errors to knowledge and causes

They need inspectable evidence chains, not isolated results

Core challenge: connect standards, knowledge points, questions, learner errors, and cultural evidence in one reusable structure.

4 of 14

Research Gap and Research Questions

Research gap

Many educational KGs focus on single-course or single-standard settings.

CFL resources coexist across multiple standards and heterogeneous sources.

High-value links such as item–knowledge and error–knowledge still require explicit construction and validation.

RQ1

How can a multi-standard graph unify language knowledge, assessment resources, learner-error evidence, and cultural content?

RQ2

Can such a graph support teacher-facing tasks more effectively and more traceably than document search and keyword-only retrieval?

Main idea: treat educational relations as first-class, inspectable objects.

4

5 of 14

CFL-KG Overview: Data, Construction, and Functions

5

CFL-KG connects heterogeneous CFL resources, a construction pipeline, and teacher-facing functions.

Three parts are connected: data sources → KG construction pipeline → teacher-facing functions.

The graph turns heterogeneous resources into inspectable evidence paths for teaching and assessment.

6 of 14

Ontology: Three-Layer Educational Schema

Key design

Language layer: characters, words, grammar, radicals, and dependencies

Assessment layer: questions, templates, papers, and scoring standards

Cultural layer: cultural points, idioms, and intercultural cases

Level and Error nodes act as cross-layer hubs

The schema makes task-relevant relations explicit: level alignment, item–knowledge links, and error–knowledge links.

6

7 of 14

Data Construction and Graph Scale

7

Construction combines controlled automation with expert review for high-impact educational relations.

Automated construction

Parsing and cleaning from standards, textbooks, dictionaries, corpora, and assessment resources

Stable identifiers, normalization, deduplication, and provenance preservation

Rule-based and statistical candidate generation for alignments and links

Expert review

Domain experts review item–knowledge, error–knowledge, and cross-standard links

Approved, revised, pending, and rejected candidates are recorded

This protects assessment tracing and intervention planning from unsupported links

Graph scale

20K+

characters

250K+

words

1K+

grammar�points

500+

cultural�nodes

10K+

questions

2K+

reading�materials

High-impact educational relations are reviewed before entering the deployed graph.

8 of 14

Deployed System: From Graph Data to Teacher Workflows

8

What users can do

Browse by standard and level

Search entities and inspect node details

Explore one-hop and multi-hop relations

Trace questions back to target knowledge points

Combine graph paths with corpus evidence

9 of 14

Intrinsic Evaluation: Is the Graph Reliable Enough?

Before testing teacher workflows, we first audit the graph itself.

Evaluation dimensions

Standards coverage

Schema and structural quality

High-value relation correctness

Short-hop reachability

Manual samples

220 item–knowledge links

210 error–knowledge links

200 level-alignment links

Human judgment

Two raters with CFL backgrounds

Endpoint type, relation direction, source evidence, and pedagogical justification

Third adjudicator for disagreements

Inter-rater agreement is substantial, with Cohen’s κ ranging from .81 to .90 for key relation groups.

9

10 of 14

Intrinsic Results: Coverage, Structure, and Relation Quality

97.8%

entities aligned to standards

0.6%

schema-violating edge rate

92.0%

item–knowledge precision

89.0%

error–knowledge precision

95.0%

level-alignment precision

91.2%

reachability within ≤3 hops

Interpretation: high-value educational links reach usable precision, and evidence chains remain short enough to inspect.

10

11 of 14

Extrinsic Evaluation: Does CFL-KG Improve Teacher Workflows?

We compare graph-based evidence tracing with document search and keyword-only retrieval.

Participants

12 teachers or researchers

International Chinese education backgrounds

Within-subject design

Conditions

Manual document search

Keyword-only retrieval

Full CFL-KG

Task families

T1: standards-aligned preparation

T2: assessment traceability

T3: error-informed planning

Measures: task quality, evidence completeness, and completion time. The setup separates graph relations from simple retrieval.

11

12 of 14

Extrinsic Results: Higher Quality, Stronger Evidence, Faster Work

Task quality

0.78 → 0.91

Evidence completeness

0.71 → 0.88

Median time

245s → 109s

Largest gains appear in relation-intensive tasks, especially assessment tracing and error-informed planning.

12

13 of 14

Discussion: When Does Graph Structure Matter?

Graph structure matters most when teachers need alignment, diagnosis, and justification.

What CFL-KG adds

Relational verification: inspect how items, standards, errors, and examples connect.

Evidence packaging: support intervention planning and explanation.

Efficiency for relation-intensive teacher workflows.

Limitations and next steps

Sample size is modest and tasks are controlled.

Some cultural and learner-error links still need expert judgment.

Future work: confidence metadata, review status, and authentic classroom deployment.

CFL-KG supports teachers and researchers rather than autonomous high-stakes decisions about learners.

13

14 of 14

Takeaways

1

CFL-KG unifies language knowledge, standards, assessment resources, learner errors, and cultural content.

2

Its value comes from inspectable educational relations, not only from data scale.

3

Evaluations show improved task quality, stronger evidence completeness, and faster teacher-facing workflows.

Thank you! Q&A

Han Wang · East China Normal University

14