1 of 61

JASH PAREKH

SIEBEL SCHOOL OF COMPUTING AND DATA SCIENCE

UNIVERSITY OF ILLINOIS URBANA-CHAMPAIGN

AUGUST 10, 2026

1

Reasoning with Structures for Large Language Models

1

2 of 61

Outline

  • How LLMs Reason and Where Reasoning Breaks
  • Structure-Augmented Reasoning: Graphs That Guide the Chain
  • Where Structure Comes From: Inducing Relations and Schemas
  • When the Context Doesn't Fit: Structuring Long Documents
  • Where Structure Lives: In the Knowledge, and in the Environment
  • The Road Ahead: Structure for Long-Horizon Agents

2

3 of 61

Why do LLMs Need to Reason?

Wei, Wang, Schuurmans, Bosma, Xia, Chi, Le, Zhou, "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models", NeurIPS 2022

GSM8K solve rate, direct answering (black) vs. chain-of-thought (blue); dashed = prior supervised SOTA. The gap only opens at scale.

Real questions are multi-step: the answer depends on intermediate results the model must work out first!

Reasoning lets the model spend more computation per problem, and leaves a trace that can be checked

3

4 of 61

Why is it Important?

Accuracy: intermediate steps provide extra computation on hard problems

Fewer errors and hallucinations: intermediate conclusions can be used to ground future steps

Verifiability: intermediate steps can make the reasoning process easier to inspect and verify

Wei, Wang, Schuurmans, Bosma, Xia, Chi, Le, Zhou, "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models", NeurIPS 2022

(a) answer directly; (b) write the steps first.

4

5 of 61

Chain-of-Thought Reasoning

Wei, Wang, Schuurmans, Bosma, Xia, Chi, Le, Zhou, "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models", NeurIPS 2022

One pass answers 27 and sounds certain; the chain answers 9.

CoT prompts the model to write the steps before the conclusion

Key Issue: One sample, one path — a single wrong step poisons the rest

Takeaway: Chain is generated, not verified — nothing forces step k to actually follow from step k−1

5

6 of 61

Overcoming the Limitations of CoT

Wang et al., "Self-Consistency Improves Chain of Thought Reasoning in Language Models", ICLR 2023

Yao et al., "Tree of Thoughts: Deliberate Problem Solving with LLMs", NeurIPS 2023

Self-Consistency: sample many chains, keep the majority answer

Tree-of-Thoughts: branch, score, prune — reasoning as search

Similarities: Both show that more inference compute buys higher accuracy

What they do not fix: every path is still free-form text over the same flat context

6

7 of 61

Evolution of Reasoning - From Prompting for Reasoning to Training for It

DeepSeek proposed R1-Zero, a training paradigm that:

[1] Skips the intermediate supervised fine-tuning process, and directly apply RL to update the LLM

[2] Eliminates the need to train a reward model, but uses a rule-based reward (e.g., exact match) instead.

By 2026 this is table stakes: thinking is always on, with an effort dial

DeepSeek-AI (Guo et al.), "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning", 2025

DeepSeek-R1-Zero’s results on AIME (American Invitational Mathematics Examination)

The LLM expands its reasoning chain to get more accurate answer, like a human!

7

8 of 61

What Is Structure, and Why Does It Matter?

Text states facts; the relations between them stay implicit

Structure makes the relations explicit (e.g., graphs, trees, tables, schemas)

Different shapes answer different questions — hops, hierarchy, aggregation

Text tells you what is known. Structure tells you how it connects.

“Aspirin thins the blood.

Warfarin also thins the blood."

raw text

extract relations

Aspirin

Warfarin

Bleeding risk

Interaction stated nowhere

increases

increases

the same knowledge, with relations made explicit

8

9 of 61

Pitfalls of Current Reasoning Methods - 1: Reasoning Is Not Knowledge

Parametric knowledge: frozen at pretraining — stale, thin in the long tail

Longer chain over wrong premises → more confident wrong answer

Traditional RAG and even trained methods (e.g., Search-R1) fix what the model sees…

…not how it reasons over it

Jin, Zeng, Yue, Yoon, Arik, Wang, Zamani, Han, "Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning", COLM 2025

A real Search-R1 rollout: the model interleaves <think> and <search> until it can answer (github.com/PeterGriffinJin/Search-R1)

9

10 of 61

Pitfalls of Current Reasoning Methods - 2: Flat Context Defeats the Chain

Liu, Lin, Hewitt, Paranjape, Bevilacqua, Petroni, Liang, "Lost in the Middle: How Language Models Use Long Contexts", TACL 2024

Jin, Yoon, Han, Arik, "Long-Context LLMs Meet RAG", 2024

Retrieved chunks arrive flat — the model must

reconnect facts across documents

Accuracy depends on where the evidence sits, not

just whether it is present

Multi-hop is where it bites: the bridge fact is

there, and never connected

10

11 of 61

Pitfalls of Current Reasoning Methods - 3: The Chain Is Unverifiable

A chain-of-thought is a plausible narrative, not a

derivation

No provenance: no pointer from a claim back to its

evidence

In domains such as medicine and finance, the audit

trail is the deliverable!

Verification has to come from outside the chain

a fluent chain…

Step 1 — “Drug X was approved in 2019”

Step 2 — “Trial B reported no interactions”

Step 3 — “Therefore X is safe to combine with warfarin”

Doc A

Doc B

?

Steps 1–2 trace back to documents. The conclusion traces back to nothing — and nothing in the chain flags it.

11

12 of 61

Each Pitfall Has a Structural Remedy

Three pitfalls, one shape: relations left implicit

Not knowledge:

ground the chain in organized evidence

Flat context:

adjacency — bridge facts co-located

Unverifiable:

the traversal path is the explanation

Retrieve → STRUCTURE → reason

12

13 of 61

Outline

  • How LLMs Reason and Where Reasoning Breaks
  • Structure-Augmented Reasoning: Graphs That Guide the Chain
  • Where Structure Comes From: Inducing Relations and Schemas
  • When the Context Doesn't Fit: Structuring Long Documents
  • Where Structure Lives: In the Knowledge, and in the Environment
  • The Road Ahead: Structure for Long-Horizon Agents

13

14 of 61

GraphRAG: Graph Community Detection and Textualization

Edge, et al. "From local to global: A graph rag approach to query-focused summarization." arXiv preprint arXiv:2404.16130 

Overview of GraphRAG

Corpus

Triple

Extraction

Knowledge Graph

Community Detection

Graph Communities

Community Summarization

New Corpus

The core idea is to “reindex” the corpus into structurally meaningful text chunks

14

15 of 61

GraphRAG: Graph Community Detection and Textualization

Edge, et al. "From local to global: A graph rag approach to query-focused summarization." arXiv preprint arXiv:2404.16130 

Leiden algorithm

Knowledge Graph

Graph Communities

(Leiden is a community detection algorithm that identifies well-connected groups of nodes in networks while guaranteeing connected communities and optimal node assignments through iterative refinement.)

Community Detection

15

16 of 61

GraphRAG: Graph Community Detection and Textualization

Edge, et al. "From local to global: A graph rag approach to query-focused summarization." arXiv preprint arXiv:2404.16130 

Graph Communities at Different Hierarchical Level

Level 0

Level 1

Community Summaries

 

Summarize each community

New corpus for RAG

 

Community Textualization

16

17 of 61

GraphRAG: Graph Community Detection and Textualization

Edge, et al. "From local to global: A graph rag approach to query-focused summarization." arXiv preprint arXiv:2404.16130 

GraphRAG outperforms Naïve RAG on comprehensiveness, diversity, and empowerment.

This shows the effectiveness of re-indexing the corpus with structure information.

17

18 of 61

Bernal, et al. "Hipporag: Neurobiologically inspired long-term memory for large language models." NeurIPS 2024

HippoRAG

HIppoRAG

Constructs KG for all the passages in the corpus using LLM.

Retrieve relevant information by finding paths in the KG between the entities in the query, using PageRank.

HippoRAG – Knowledge Graph Construction and Retrieval

18

19 of 61

SARG: Structure Without the Index

Build the reasoning structure on demand, over only the documents the retriever returned!

Jash Parekh, Pengcheng Jiang, Jiawei Han, "SARG: Structure-Augmented Reasoning Generation", arXiv:2506.08364

19

20 of 61

SARG Results: Multi-Hop and Domain-Specific Benchmarks

Bold = best within each backbone group; † = best overall across all methods

SARG leads on every multi-hop metric!

20

21 of 61

LLM-Powered Output Generation with Justification

SARG reconstructs the full pathway the human-annotated KG misses

SARG's key advantages:�

[1] Plug-and-play: a post-retrieval layer for any RAG pipeline �[2] Interpretable: the exact traversal path is surfaced

[3] Domain-adaptive: few-shot extraction, no fine-tuning

21

22 of 61

CGR: Condition-Gated Reasoning for Context-Dependent Biomedical Question Answering

Motivation: ALL clinical decisions are conditional (patient-dependent) – age, comorbidities, allergies, contraindications.

Existing benchmarks & RAG studies ignore patient context!

Parekh et al."Condition-Gated Reasoning for Context-Dependent Biomedical Question Answering“, KDD 2026

22

23 of 61

CGR: Condition-Gated Reasoning for Context-Dependent Biomedical Question Answering

Condition-aware KG: extract n-tuples that carry condition tags with LLM

Gating: prune paths by patient conditions

Inference sees only condition-appropriate knowledge

Parekh et al."Condition-Gated Reasoning for Context-Dependent Biomedical Question Answering“, KDD 2026

23

24 of 61

CGR Performance Across QA Benchmarks

Parekh et al."Condition-Gated Reasoning for Context-Dependent Biomedical Question Answering“, KDD 2026

24

25 of 61

Query-Time Structuring, Learned: StructRAG and Structure-R1

StructRAG: router picks the structure — table, graph, algorithm, catalogue, chunk

Structure-R1: RL-learned structuring — reward: answer from the structure alone

Results show that 7B + learned structure ≈ much larger models

Li, Chen, Yu, Lin, Lu, Tang, Huang, Han, Sun, Li, "StructRAG", ICLR 2026

Wu, Zhong, Sun, Li, Jin, Han, Zeng, "Structure-R1", 2025, arXiv:2510.15191

25

26 of 61

Outline

  • How LLMs Reason and Where Reasoning Breaks
  • Structure-Augmented Reasoning: Graphs That Guide the Chain
  • Where Structure Comes From: Inducing Relations and Schemas
  • When the Context Doesn't Fit: Structuring Long Documents
  • Where Structure Lives: In the Knowledge, and in the Environment
  • The Road Ahead: Structure for Long-Horizon Agents

26

27 of 61

Where Does Structure Actually Come From?

Needed before reasoning: taxonomy, relations, roles, schema�

Old: expert annotation, fixed vocabularies�

New: LLMs propose, corpus evidence decides, humans constrain

Taxonomy

Relations

Roles

Schema

27

28 of 61

PriORE: Aligning Relation Extraction to a Specific Topic

Linyi Ding, Jinfeng Xiao, Sizhe Zhou, Chaoqi Yang, Jiawei Han, “Topic-Oriented Open Relation Extraction with A Priori Seed Generation”, EMNLP’24

Closed RE: needs a predefined list · open RE: drifts off-topic

PriORE: a-priori seed relations → classify mentions

Why it matters here: the relation vocabulary for SARG & CGR graphs

28

29 of 61

PriORE: Biggest Gains on Narrow, Domain-Specific Topics

Takeaways:�

Biggest wins: narrow specialist topics — no existing KG�

Largest gain: informativeness

Coarse-grained topics (DocRED, FewRel-Bio)

Fine-grained topics (EV batteries, redox reactions)

29

30 of 61

ObSchema: Inducing Attribute Schemas From Scientific Text

Attributes stated locally — values, units, conditions tangled�

Closed schemas: can't enumerate · open extraction: never consolidates�

Goal: induced, type-level attribute schema

Patrick Xu et al., ObSchema: From Local Observations to Global Entity Attribute Schemas in Scientific Literature , submitted

Local, condition-laden observations (left) and the type-level attribute schema they should roll up into (right).

30

31 of 61

Inside ObSchema: Extract, Organize, Refine

Extract n-ary observations (schema-guided)�

Promote only on corpus-level evidence�

Schema improves batch over batch

The schema S is carried forward as each batch arrives, so later batches are extracted against a better schema.

31

32 of 61

ObSchema: Best Schema Quality Under Both Base LLMs

Four metrics — Attribute Validity, Path Granularity, Sibling Coherence, Uniqueness — averaged over five independent LLM judges.

SOTA performance across the board!

Higher is better; bold marks the best method in each block.

32

33 of 61

ObSchema: Faster and Cheaper Than Iterative Induction

33

34 of 61

Takeaway: Structure Is Now Inducible

Taxonomies, relations, roles, and schemas are ALL now inducible without any predefined vocabulary!

Corpus

Candidates

Evidence gate

Structure

34

35 of 61

Outline

  • How LLMs Reason and Where Reasoning Breaks
  • Structure-Augmented Reasoning: Graphs That Guide the Chain
  • Where Structure Comes From: Inducing Relations and Schemas
  • When the Context Doesn't Fit: Structuring Long Documents
  • Where Structure Lives: In the Knowledge, and in the Environment
  • The Road Ahead: Structure for Long-Horizon Agents

35

36 of 61

Long Document QA: No Document Can Be Ignored

Wang et al. "Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA", ArXiv: 2405-17419v2

Every document required — top-k ≠ all

Tasks: locate · compare · cluster · chain

Loong’s four task types: locate the evidence; locate and compare; locate and cluster; locate and reason along a chain.

36

37 of 61

RAPTOR: A Tree of Summaries, Retrieval at Every Level

Embed → cluster → summarize, recursively�

Retrieve at any level: theme + detail in one pass

Sarthi, Abdullah, Tuli, Khanna, Goldie, Manning, "RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval", ICLR 2024

37

38 of 61

SLIDERS: Turn a Document Set Into a Database, Then Query It

Harshit Joshi, Priyank Shethia, Jadelynn Dao, and Monica S. Lam, "Contexts are Never Long Enough: Structured Reasoning for Scalable Question Answering over Long Document Sets", ArXiv:2604-22294

Chunk pipelines: the long-context problem regenerates

SLIDERS: DB rows + SQL answers + per-row provenance

Top: a text intermediate state hits the aggregation bottleneck. Bottom: the intermediate state is a database, and the answer is a query over it.

38

39 of 61

SLIDERS: From Raw Documents to a Queryable Database

Documents → contextualized chunks → induced schema → extracted rows → reconciled tables → SQL answer

1. Induce typed schema (query + metadata)

2. Extract rows w/ provenance · reconcile

3. Answer with SQL

39

40 of 61

SLIDERS: Scales Past 3.9M Tokens at ~$0.76 per Query

Only system past 3.9M tokens

~$0.76 / query · ~3 min end-to-end

Every answer auditable to the row

Benchmarks

FinanceBench — analyst questions on filings

Loong — every document required

Oolong — classify locally, aggregate globally

40

41 of 61

Outline

  • How LLMs Reason and Where Reasoning Breaks
  • Structure-Augmented Reasoning: Graphs That Guide the Chain
  • Where Structure Comes From: Inducing Relations and Schemas
  • When the Context Doesn't Fit: Structuring Long Documents
  • Where Structure Lives: In the Knowledge, and in the Environment
  • The Road Ahead: Structure for Long-Horizon Agents

41

42 of 61

Jiang et al., "Reasoning-Enhanced Healthcare Predictions with Knowledge Graph Community Retrieval", ICLR 2025

Large Language Models are not good at everything.

For example, they are bad at clinical prediction, if no external knowledge is provided:

Brown, et al. "Large language models are less effective at clinical prediction tasks than locally trained machine learning models." JAMIA 2025

Chen et al. "ClinicalBench: Can LLMs Beat Traditional ML Models in Clinical Prediction?." arXiv preprint arXiv:2411.06469 (2024).

This is because clinical predictions given patient context needs highly relevant medical knowledge and complex reasoning process.

KARE: Graph Community-Indexed Corpus and Personalized Retrieval

42

43 of 61

Jiang et al., "Reasoning-Enhanced Healthcare Predictions with Knowledge Graph Community Retrieval", ICLR 2025

Naïve RAG?

EHR example

Retrieval Result from PubMed

Unwanted Information

KARE: Graph Community-Indexed Corpus and Personalized Retrieval

43

44 of 61

Jiang et al., "Reasoning-Enhanced Healthcare Predictions with Knowledge Graph Community Retrieval", ICLR 2025

KARE, a framework for clinical predictions, introducing fine-grained medical knowledge indexing, personalized (theme-specific) retrieval, and LLM reasoning over patient context and retrieved knowledge.

KARE: Graph Community-Indexed Corpus and Personalized Retrieval

44

45 of 61

We construct theme-specific summarization of KG communities w/ LLMs.

Jiang et al., "Reasoning-Enhanced Healthcare Predictions with Knowledge Graph Community Retrieval", ICLR 2025

We retrieve the KG communities based on multiple metrics:

Node Hits: the overlapped node between patient graph and the KG

Coherence: semantic similarity between the patient context and the KG community

Recency: communities that are more relevant to recent visits should be weighted higher

Theme Relevance: communities more relevant to the targeted theme should be weighted higher

KARE: Graph Community-Indexed Corpus and Personalized Retrieval

45

46 of 61

Jiang et al., "Reasoning-Enhanced Healthcare Predictions with Knowledge Graph Community Retrieval", ICLR 2025

We then finetune the LLM in a multi-task setting, performing both reasoning generation (w/ data distilled from larger models) and label prediction

KARE: Graph Community-Indexed Corpus and Personalized Retrieval

46

47 of 61

Jiang et al., "Reasoning-Enhanced Healthcare Predictions with Knowledge Graph Community Retrieval", ICLR 2025

For mortality prediction where the data is extremely imbalanced (5.42% positive labels), most ML models perform poor

LM+ML based methods improved the performance by leveraing external knowledge

Naïve RAG and few-shot prompting-based LLM methods perform worse than traditional ML methods in most cases (see full results in paper).

KARE significantly outperforms all the previous methods

Key results:

(accuracy of correctly predicting patients who cannot survive)

(accuracy of correctly predicting patients who will be readmitted within 15 days)

Machine Learning Methods

ML+LM Methods

LLM Methods

KARE: Graph Community-Indexed Corpus and Personalized Retrieval

47

48 of 61

Jiang et al., "Reasoning-Enhanced Healthcare Predictions with Knowledge Graph Community Retrieval", ICLR 2025

KARE: Graph Community-Indexed Corpus and Personalized Retrieval

(Base: raw patient context, Aug: augmented patient context)

KARE shows that:

fine-grained knowledge (from the constructed medical KG communities) and

structured, graph-information-awareness

are crucial for truly useful information retrieval.

48

49 of 61

What Is a Harness?

The harness is the program wrapped around the model — it renders what the model sees, parses what the model emits, calls the tools, and keeps the state�

The model itself is stateless between turns; whatever is “remembered” lives in the harness

49

50 of 61

Harness-1: The Problem — a Policy Over a Growing Transcript

Standard training: the policy owns a growing transcript

Search decisions AND bookkeeping compete inside one policy

Harness-1: move the state into the environment; keep only the semantic decisions

Pengcheng Jiang, Zhiyi Shi, Kelly Hong, Xueqiang Xu, Jiashuo Sun, Jimeng Sun, Hammad Bashir, Jiawei Han, "Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnesses", 2026, arXiv:2606.02373

50

51 of 61

Harness-1: The State Lives in the Environment

The harness remembers: candidate pool, curated set, evidence graph, verification cache, full-text store

The policy decides: what to search, keep, verify — and when to stop

Pengcheng Jiang, Zhiyi Shi, Kelly Hong, Xueqiang Xu, Jiashuo Sun, Jimeng Sun, Hammad Bashir, Jiawei Han, "Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnesses", 2026, arXiv:2606.02373

The harness state: six externalized stores the environment maintains across turns.

51

52 of 61

Harness-1: The State Lives in the Environment

Working memory rendered at each turn:

Pengcheng Jiang, Zhiyi Shi, Kelly Hong, Xueqiang Xu, Jiashuo Sun, Jimeng Sun, Hammad Bashir, Jiawei Han, "Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnesses", 2026, arXiv:2606.02373

52

53 of 61

Harness-1: Rendered Working Memory, and a Reward That Sees It

Each turn: a budget-aware rendered working memory, not the raw transcript

Training: teacher rollouts → SFT → RL (CISPO) over full trajectories

SFT teaches the interface, RL the decisions.

Pengcheng Jiang, Zhiyi Shi, Kelly Hong, Xueqiang Xu, Jiashuo Sun, Jimeng Sun, Hammad Bashir, Jiawei Han, "Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnesses", 2026, arXiv:2606.02373

53

54 of 61

Harness-1: Externalized State Generalizes

Structure moved from the knowledge into the environment

🡪 SOTA: Harness-1 outperforms frontier models (e.g., GPT-5.4) on long-horizon search!

Pengcheng Jiang, Zhiyi Shi, Kelly Hong, Xueqiang Xu, Jiashuo Sun, Jimeng Sun, Hammad Bashir, Jiawei Han, "Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnesses", 2026, arXiv:2606.02373

54

55 of 61

Where Are We Heading?

build the index once, retrieve from it (e.g., GraphRAG, HippoRAG, RAPTOR)

Static Structure

build (or learn to build) the structure per question (e.g., SARG, StructRAG, Structure-R1)

Query-Time Structure

State and tasks live outside the model! (e.g., Harness-1, SLIDERS, Kimi K3)

Structure in the Environment

Flog-retrieval hacks → trained consolidation over explicit structure

Agent Memory, Trained Consolidation

55

56 of 61

This Loop Already Runs at Frontier Scale: Kimi K3

K3 RL tasks: self-evolving knowledge graph — sample concepts → retrieve materials → synthesize tasks

The knowledge graph is now training infrastructure

Kimi Team, “Kimi K3: Open Frontier Intelligence”, arXiv:2607.24653, 2026 — RL task-synthesis workflow over the concept graph

56

57 of 61

Conclusion

Reasoning with Structures: use it, induce it, scale it, live in it

Use it

Graphs guide the chain!

GraphRAG, SARG , CGR

Induce it

Structure from raw corpora!

PriORE , ObSchema

Scale it

Beyond the context window!

RAPTOR , SLIDERS

Live in it

Index and environment!

KARE , Harness-1

Text tells you what is known. Structure tells you how it connects. — Thank you!

57

58 of 61

References I

  • Xiang Chen, Ningyu Zhang, Xin Xie, Shumin Deng, Yunzhi Yao, Chuanqi Tan, Fei Huang, Luo Si, Huajun Chen, “KnowPrompt: Knowledge-aware Prompt-tuning with Synergistic Optimization for Relation Extraction”, WWW’22
  • Yew Ken Chia, Lidong Bing, Soujanya Poria, Luo Si, “RelationPrompt: Leveraging Prompts to Generate Synthetic Data for Zero-Shot Relation Triplet Extraction”, ACL’22 Findings
  • Linyi Ding, Jinfeng Xiao, Sizhe Zhou, Chaoqi Yang, Jiawei Han, “Topic-Oriented Open Relation Extraction with A Priori Seed Generation”, EMNLP’24
  • Gao, T., Fisch, A., & Chen, D. (2021). Making pre-trained language models better few-shot learners. ACL
  • Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., ... & Gelly, S. (2019). Parameter-efficient transfer learning for NLP. ICML
  • Pengcheng Jiang, Cao Xiao, Minhao Jiang, Parminder Bhatia, Taha Kass-Hout, Jimeng Sun, Jiawei Han, "Reasoning-Enhanced Healthcare Predictions with Knowledge Graph Community Retrieval", ICLR’2025
  • Z. Jiang, F. F. Xu, L. Gao, Z. Sun, Q. Liu, J. Dwivedi-Yu, Y. Yang, J. Callan, and G. Neubig, “Active retrieval augmented generation,” 2023.
  • Yizhu Jiao, Sha Li, Yiqing Xie, Ming Zhong, Heng Ji, Jiawei Han. "Open-vocabulary argument role prediction for event extraction," EMNLP’22 Findings
  • Yizhu Jiao, Ming Zhong, Sha Li, Ruining Zhao, Siru Ouyang, Heng Ji, and Jiawei Han. “Instruct and Extract: Instruction Tuning for On-Demand Information Extraction.” EMNLP 2023.
  • Yizhu Jiao, Sha Li, Sizhe Zhou, Heng Ji, and Jiawei Han. “ Text2DB: Integration-Aware Information Extraction with Large Language Model Agents .” ACL 2024 findings.
  • Priyanka Kargupta, Yunyi Zhang, Yizhu Jiao, Siru Ouyang, Jiawei Han, "Synergizing Unsupervised Episode Detection with LLMs for Large-Scale News Events", ACL 2025
  • Priyanka Kargupta, Runchu Tian, Jiawei Han, "Beyond True or False: Retrieval-Augmented Hierarchical Analysis of Nuanced Claims", ACL 2025
  • Priyanka Kargupta, Ishika Agarwal, Tal August, Jiawei Han, "Tree-of-Debate: Multi-Persona Debate Trees Elicit Critical Thinking for Scientific Comparative Analysis", ACL 2025
  • Seoyeon Kim, Kwangwook Seo, Hyungjoo Chae, Jinyoung Yeo, Dongha Lee, “VerifiNER: Verification-augmented NER via Knowledge-grounded Reasoning with Large Language Models”, ACL’24
  • Le Scao, T., & Rush, A. M. (2021). How many data points is a prompt worth? NAACL

58

59 of 61

References II

  • Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., ... & Zettlemoyer, L. (2020). BART: Denoising sequence-to-sequence Li, X. L., & Liang, P. (2021). Prefix-tuning: Optimizing continuous prompts for generation. ACL
  • P. Manakul, A. Liusie, and M. J. F. Gales, “SelfcheckGPT: Zero-resource black-box hallucination detection for generative large language models,” 2023
  • Yubo Ma, Yixin Cao, YongChing Hong, Aixin Sun, “Large Language Model Is Not a Good Few-shot Information Extractor, but a Good Reranker for Hard Samples!”, EMNLP’23 Findings
  • S. Min, K. Krishna, X. Lyu, M. Lewis, W. tau Yih, P. W. Koh, M. Iyyer, L. Zettlemoyer, and H. Hajishirzi, “Factscore: Fine-grained atomic evaluation of factual precision in long form text generation,” 2023
  • O. Ovadia, et al (2023), “Fine-tuning or retrieval? comparing knowledge injection in LLMs,” arXiv:2312.05934
  • B. Paranjape, S. Lundberg, S. Singh, H. Hajishirzi, L. Zettlemoyer, and M. T. Ribeiro, “Art: Automatic multi-step reasoning and tool-use for large language models,” 2023
  • Jash Parekh, Pengcheng Jiang, Jiawei Han, "SARG: Structure-Augmented Reasoning Generation", arXiv:2506.08364
  • X. Qiu, T. Sun, Y. Xu, Y. Shao, N. Dai and X. Huang, Pre-trained models for natural language processing: A survey, Science China-Technological Sciences, 2020
  • Radford, A., Narasimhan, K., Salimans, T., & Sutskever, I. (2018). Improving language understanding by generative pre-training. OpenAI blog.
  • Schick, T., & Schütze, H. (2021). Exploiting cloze questions for few shot text classification and natural language inference. EACL
  • Omar Shaikh, et al. (2022), "On Second Thought, Let’s Not Think Step by Step! Bias and Toxicity in Zero-Shot Reasoning" arXiv.2212.08061
  • Suzgun, M., Scales, N., Schärli, N., Gehrmann, S., Tay, Y., Chung, H. W., Chowdhery, A., Le, Q. V., Chi, E. H., Zhou, D., Wei, J. (2023). Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them. ACL Findings.
  • N. Shinn, F. Cassano, E. Berman, A. Gopinath, K. Narasimhan, and S. Yao, “Reflexion: Language agents with verbal reinforcement learning,” 2023.
  • Chengen Wang and Murat Kantarcioglu, "A Review of DeepSeek Models’ Key Innovative Techniques", arXiv:2503.11486
  • Wei, J., et al. (2022). Emergent Abilities of Large Language Models. TMLR.
  • T. Wu, E. Jiang, A. Donsbach, J. Gray, A. Molina, M. Terry, and C. J. Cai, “PromptChainer: Chaining large language model prompts through visual programming,” 2022

59

60 of 61

References III

  • Xueqiang Xu, Wonbin Kweon, Jimeng Shi, Runchu Tian, Yunfan Kang, Wei Hu, Shaowen Wang, Jiawei Han, "ObSchema: From Local Observations to Global Entity Attribute Schemas in Scientific Literature", arXiv: May 2026
  • Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., Narasimhan, K. (2023). Tree of Thoughts: Deliberate Problem Solving with Large Language Models. NeurIPS.
  • Changlong Yu, Weiqi Wang, Xin Liu, Jiaxin Bai, Yangqiu Song, Zheng Li, Yifan Gao, Tianyu Cao, and Bing Yin. “FolkScope: Intention Knowledge Graph Construction for E-commerce Commonsense Discovery”, ACL’23 Findings
  • S. J. Zhang, et al., “Exploring the MIT mathematics and EECS curriculum using large language models,” 2023
  • Zhang, Z., Zhang, A., Li, M., Smola, A. (2023). Automatic Chain of Thought Prompting in LLMs. ICLR
  • Kai Zhang, Bernal Jiménez Gutiérrez, Yu Su, “Aligning Instruction Tasks Unlocks Large Language Models as Zero-Shot Relation Extractors”, ACL'23 Findings
  • Yu Zhang, Yu Meng, Xuan Wang, Sheng Wang, Jiawei Han. "Seed-guided topic discovery with out-of-vocabulary seeds." NAACL’22.
  • Yu Zhang, Yunyi Zhang, Yucheng Jiang, Martin Michalski, Yu Deng, Lucian Popa, ChengXiang Zhai, Jiawei Han. "Entity Set Co-Expansion in StackOverflow." IEEE Big Data 2022
  • Ming Zhong, Siru Ouyang, Minhao Jiang, Vivian Hu, Yizhu Jiao, Xuan Wang, Jiawei Han, “ReactIE: Enhancing Chemical Reaction Extraction with Weak Supervision”, ACL’23 Findings
  • Ming Zhong, Siru Ouyang, Yizhu Jiao, Priyanka Kargupta, Leo Luo, Yanzhen Shen, Bobby Zhou, Xianrui Zhong, Xuan Liu, Hongxiang Li, Jinfeng Xiao, Minhao Jiang, Vivian Hu, Xuan Wang, Heng Ji, Martin Burke, Huimin Zhao and Jiawei Han, "Reaction Miner: An Integrated System for Chemical Reaction Extraction from Textual Data", EMNLP’23
  • Zhou, C., Liu, P., Xu, P., Iyer, S., Sun, J., Mao, Y., Ma, X., Efrat, A., Yu, P., Yu, L., Zhang, S., Ghosh, G., Lewis, M., Zettlemoyer, L., & Levy, O. (2023). LIMA: Less Is More for Alignment.
  • Sizhe Zhou, Suyu Ge, Jiawei Han, “Corpus-Based Relation Extraction by Identifying and Refining Relation Patterns”, ECMLPKDD’23
  • Sizhe Zhou, Yu Meng, Bowen Jin, Jiawei Han, “Grasping the Essentials: Tailoring Large Language Models for Zero-Shot Relation Extraction”, arXiv’24
  • Wenxuan Zhou, Muhao Chen, “An Improved Baseline for Sentence-level Relation Extraction”, AACL’22
  • Y. Zhou, A. I. Muresanu, Z. Han, K. Paster, S. Pitis, H. Chan, and J. Ba, “Large language models are human-level prompt engineers,” 2023

60

61 of 61

References IV

  • DeepSeek-AI, “DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning”, 2025
  • Bowen Jin, Hansi Zeng, Zhenrui Yue, Jinsung Yoon, Sercan Arik, Dong Wang, Hamed Zamani, Jiawei Han, “Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning”, COLM 2025
  • Zhuoqun Li, Xuanang Chen, Haiyang Yu, Hongyu Lin, Yaojie Lu, Qiaoyu Tang, Fei Huang, Xianpei Han, Le Sun, Yongbin Li, “StructRAG: Boosting Knowledge Intensive Reasoning of LLMs via Inference-time Hybrid Information Structurization”, ICLR 2025
  • OpenAI, “Learning to Reason with LLMs” (o1), 2024
  • Qwen Team, “Qwen3 Technical Report”, 2025
  • Junlin Wu, Xianrui Zhong, Jiashuo Sun, Bolian Li, Bowen Jin, Jiawei Han, Qingkai Zeng, “Structure-R1: Dynamically Leveraging Structural Knowledge in LLM Reasoning through Reinforcement Learning”, arXiv:2510.15191, 2025
  • Jash Rajesh Parekh, Wonbin Kweon, Joey Chan, Rezarta Islamaj, Robert Leaman, Pengcheng Jiang, Chih-Hsuan Wei, Zhizheng Wang, Zhiyong Lu, Jiawei Han, “Condition-Gated Reasoning for Context-Dependent Biomedical Question Answering”, KDD 2026.

61