1 of 95

Computing Syntax in�Large Language Models

Week 7 Update�Thursday | July 9, 2026

2 of 95

Roadmap

  1. Causal Abstraction Framework
    1. Counterfactual Prompts
  2. Literature Review
    • CausalGym
    • Distributed Alignment Search
    • Cross-linguistic machinery
    • Dyck-k,n parsing and Depth
  3. Motivating GPT-2 Architecture Experiments
  4. RASP Language
    • D-RASP Results
    • RASP-L length generalizing program for HI → HF

2

3 of 95

Roadmap

  • Causal Abstraction Framework
    • Counterfactual Prompts
  • Literature Review
    • CausalGym
    • Distributed Alignment Search
    • Cross-linguistic machinery
    • Dyck-k,n parsing and Depth
  • Motivating GPT-2 Architecture Experiments
  • RASP Language
    • D-RASP Results
    • RASP-L length generalizing program for HI → HF

3

4 of 95

Causal Abstraction

4

5 of 95

Causal Abstraction

5

6 of 95

Causal Abstraction

6

7 of 95

Causal Abstraction

7

8 of 95

Causal Abstraction

8

9 of 95

Causal Abstraction

9

10 of 95

High Level Causal Diagram (Addition)

10

11 of 95

High Level Causal Diagram (Addition)

11

12 of 95

High Level Causal Diagram (Addition)

12

13 of 95

High Level Causal Diagram (Addition)

13

14 of 95

High Level Causal Diagram (Addition)

14

15 of 95

Causal Mediation

  • Hypothesis of a high level model (computation graph)
  • Explore where in the low level model (Transformer)

15

16 of 95

Causal Mediation

  • Hypothesis of a high level model (computation graph)
  • Explore where in the low level model (Transformer)

16

17 of 95

Causal Mediation

  • Hypothesis of a high level model (computation graph)
  • Explore where in the low level model (Transformer)

17

18 of 95

Causal Mediation

  • Hypothesis of a high level model (computation graph)
  • Explore where in the low level model (Transformer)

18

19 of 95

Causal Mediation

  • Hypothesis of a high level model (computation graph)
  • Explore where in the low level model (Transformer)

19

20 of 95

Causal Mediation

  • Hypothesis of a high level model (computation graph)
  • Explore where in the low level model (Transformer)

20

21 of 95

Causal Mediation

  • Hypothesis of a high level model (computation graph)
  • Explore where in the low level model (Transformer)

21

22 of 95

Causal Mediation

  • Hypothesis of a high level model (computation graph)
  • Explore where in the low level model (Transformer)

22

23 of 95

Predicting the First Noun

23

24 of 95

Predicting the First Noun

24

25 of 95

Predicting the First Noun

25

26 of 95

Predicting the First Noun

26

27 of 95

Predicting the First Noun

27

28 of 95

Predicting the First Noun

28

29 of 95

Predicting the First Noun

29

30 of 95

Predicting the First Noun

  1. Change Syntactic structure

30

31 of 95

Roadmap

  • Causal Abstraction Framework
    • Counterfactual Prompts
  • Literature Review
    • CausalGym
    • Cross-linguistic machinery
    • Dyck-k,n parsing and Depth
  • Motivating GPT-2 Architecture Experiments
  • RASP Language
    • D-RASP Results
    • RASP-L length generalizing program for HI → HF

31

32 of 95

Lit Review:: CausalGym

  1. Two tasks
    1. Negative Polarity Item (NPI) Licensing
      1. “No athlete scored any…”
      2. “The athlete scored {*any, some}…”
    2. Filler Gap Subject
      • “Shri reported that Arora bought the cake with the help of him
      • “Shri reported who Arora bought the cake with the help of *him

32

33 of 95

Lit Review:: CausalGym

  • Subject Agreement
    • “the keys near the table are … “
    • “the key near the tables is … “
  • Metric: Log Odds-Ratio
    • 0 = no change
    • + = correct change
    • - = opposite change

33

34 of 95

Lit Review:: CausalGym

How detailed can patching be?

  1. Residual Stream Patching
  2. Rotation Matrix
  3. Distributed Alignment Search (DAS)
    1. Single Linear Direction
  4. Desiderata-based Component Masking

34

35 of 95

Roadmap

  • Causal Abstraction Framework
    • Counterfactual Prompts
  • Literature Review
    • CausalGym
    • Cross-linguistic machinery
    • Dyck-k,n parsing and Depth
  • Motivating GPT-2 Architecture Experiments
  • RASP Language
    • D-RASP Results
    • RASP-L length generalizing program for HI → HF

35

36 of 95

Lit Review:: Multilingual Representations

  1. Word-translation task
  2. Early layers encode output language
  3. Concept in language agnostic latent-space later

  • Cross-linguistic head-initial to head-final machinery?
    • Japanese, Korean, Turkish, Hindi, Farsi

36

37 of 95

Roadmap

  • Causal Abstraction Framework
    • Counterfactual Prompts
  • Literature Review
    • CausalGym
    • Cross-linguistic machinery
    • Dyck-k,n parsing and Depth
  • Motivating GPT-2 Architecture Experiments
  • RASP Language
    • D-RASP Results
    • RASP-L length generalizing program for HI → HF

37

38 of 95

Lit Review:: Dyck-k,d

  1. Dyck-k=1
    1. (())()()(())
  2. Dyck-k=2
    • ({}() {})(){}
  3. Dyck-k=1,d=2
    • (())()
    • ((()))

  • Encoder only: Hardmax attention
    • D+1 layers
    • O(log k) memory size (d_model)
  • Decoder Only Softmax
    • 2 layer with O(log k) memory

38

39 of 95

Roadmap

  • Causal Abstraction Framework
    • Counterfactual Prompts
  • Literature Review
    • CausalGym
    • Cross-linguistic machinery
    • Dyck-k,n parsing and Depth
  • Motivating GPT-2 Architecture Experiments
  • RASP Language
    • D-RASP Results
    • RASP-L length generalizing program for HI → HF

39

40 of 95

Motivating Causal Attention Architectures

  1. Only Causal Self attention (narrow hypothesis space)
  2. More literature on Causal Attention-based architectures
  3. Easier to translate into pretrained models
  4. May have correspondence to real-time psycholinguistic processing

40

41 of 95

Roadmap

  • Causal Abstraction Framework
    • Counterfactual Prompts
  • Literature Review
    • CausalGym
    • Cross-linguistic machinery
    • Dyck-k,n parsing and Depth
  • Motivating GPT-2 Architecture Experiments
  • RASP Language
    • D-RASP Results
    • RASP-L length generalizing program for HI → HF

41

42 of 95

Transformer Programming

  1. Algorithmic programs in transformer weights
  2. E.g. search algorithms

42

43 of 95

RASP Language

Restricted Access Sequence Processing Language (2021)

  1. “analyzing a RASP program implies a maximum number of heads and layers necessary to encode a task in a transformer”
  2. Constrained instruction set/assembly language for transformers

43

44 of 95

RASP Intuition

  1. Two types of operations
    1. Elementwise-ops = MLP
    2. Select (sequence-wise)-ops = Attention
  2. Every operation maps a sequence to a sequence

44

45 of 95

Element Wise Ops

45

46 of 95

Element Wise Ops

46

47 of 95

Element Wise Ops

47

48 of 95

Element Wise Ops

48

49 of 95

Element Wise Ops

49

50 of 95

Element Wise Ops

50

51 of 95

Selection Ops

51

52 of 95

Selection Ops

52

53 of 95

Selection Ops

53

54 of 95

Selection Ops

54

55 of 95

Selection Ops

55

56 of 95

Selection Ops

56

57 of 95

Selection Ops

57

58 of 95

Selection Ops

58

59 of 95

Selection Ops

59

60 of 95

Selection Ops

60

61 of 95

Selection Ops

61

62 of 95

Selection Ops

62

63 of 95

Selection Ops

63

64 of 95

Selection Ops

64

65 of 95

Selection Ops

65

66 of 95

Selection Ops

66

67 of 95

Selection Ops

67

68 of 95

Selection Ops

68

69 of 95

Selection Ops

69

70 of 95

RASP Language

D-RASP

  1. Decompile GPT2Model weights into D-RASP code
    1. Works well for algorithmic tasks/toy models
    2. Uses Optimal Ablation + Uniform Gradient Sampling to create a sparse computation graph
    3. Goal: Create a high level computational model for Causal Abstraction

70

71 of 95

71

72 of 95

D-RASP Program for HI -> HF Translation

1. s1 = select(q=token, k=token, op=\circled{a}) # layer 0 head 0

2. s2 = select(q=pos, k=pos, op=\circled{b}) # layer 0 head 0

3. s3 = select(k=token, op=\circled{c})

4. s4 = select(k=pos, op=\circled{d})

5. a1 = aggregate(s=s1+s2+s3+s4, v=token) # layer 0 head 0

6. s5 = select(q=token, k=token, op=\circled{e}) # layer 0 head 1

7. s6 = select(q=pos, k=pos, op=\circled{f}) # layer 0 head 1

8. s7 = select(k=token, op=\circled{g})

9. a2 = aggregate(s=s5+s6+s7, v=token) # layer 0 head 1

10. a3 = aggregate(s=s5+s6+s7, v=pos) # layer 0 head 1

11. new_a1 = element_wise_op(a1) # layer 0 mlp

12. new_a2 = element_wise_op(a2) # layer 0 mlp

13. s8 = select(q=a1, k=a1, op=\circled{h}) # layer 1 head 0

14. s9 = select(q=a2, k=a1, op=\circled{i}) # layer 1 head 0

15. s10 = select(q=a2, k=token, op=\circled{j}) # layer 1 head 0

16. s11 = select(q=a2, k=new_a1, op=(q==k)) # layer 1 head 0

17. s12 = select(q=a2, k=new_a2, op=(q==k)) # layer 1 head 0

18. s13 = select(q=token, k=a1, op=\circled{k}) # layer 1 head 0

19. s14 = select(q=token, k=token, op=\circled{l}) # layer 1 head 0

20. s15 = select(q=token, k=new_a1, op=(uniform selection),

special_op=(k==q)) # layer 1 head 0

21. s16 = select(q=token, k=new_a2, op=(q==k)) # layer 1 head 0

22. s17 = select(q=pos, k=pos, op=\circled{n}) # layer 1 head 0

23. s18 = select(q=new_a1, k=a1, op=(q==k)) # layer 1 head 0

24. s19 = select(q=new_a1, k=token, op=(q==k)) # layer 1 head 0

25. s20 = select(q=new_a1, k=new_a1, op=(q==k)) # layer 1 head 0

26. s21 = select(q=new_a1, k=new_a2, op=(q==k)) # layer 1 head 0

27. s22 = select(q=new_a1, k=a3, op=(q==k)) # layer 1 head 0

28. s23 = select(q=new_a2, k=token, op=(q==k)) # layer 1 head 0

29. s24 = select(q=new_a2, k=new_a2, op=(q==k)) # layer 1 head 0

30. a4 = aggregate(s=s8+s9+s10+s11+s12+s13+s14+s15+s16+s17+s18+s19+s20+s21+s22+s23+s24, v=a1) # layer 1 head 0

31. a5 = aggregate(s=s8+s9+s10+s11+s12+s13+s14+s15+s16+s17+s18+s19+s20+s21+s22+s23+s24, v=a2) # layer 1 head 0

32. a6 = aggregate(s=s8+s9+s10+s11+s12+s13+s14+s15+s16+s17+s18+s19+s20+s21+s22+s23+s24, v=token) # layer 1 head 0

33. s25 = select(q=a1, k=a1, op=\circled{o}) # layer 1 head 1

34. s26 = select(q=a1, k=token, op=\circled{p}) # layer 1 head 1

35. s27 = select(q=a1, k=new_a1, op=(q==k)) # layer 1 head 1

36. s28 = select(q=a2, k=a1, op=\circled{q}) # layer 1 head 1

37. s29 = select(q=a2, k=a2, op=\circled{r}) # layer 1 head 1

38. s30 = select(q=a2, k=new_a1, op=(q==k)) # layer 1 head 1

39. s31 = select(q=a2, k=new_a2, op=(q==k)) # layer 1 head 1

40. s32 = select(q=token, k=a2, op=\circled{s}) # layer 1 head 1

41. s33 = select(q=token, k=token, op=\circled{t}) # layer 1 head 1

42. s34 = select(q=token, k=new_a1, op=(q==k)) # layer 1 head 1

43. s35 = select(q=token, k=new_a2, op=(q==k)) # layer 1 head 1

44. s36 = select(q=new_a1, k=a1, op=(q==k)) # layer 1 head 1

45. s37 = select(q=new_a1, k=new_a1, op=(q==k)) # layer 1 head 1

46. s38 = select(q=new_a1, k=new_a2, op=(q==k)) # layer 1 head 1

47. s39 = select(q=new_a1, k=a3, op=(q==k)) # layer 1 head 1

48. s40 = select(q=new_a2, k=token, op=(q==k)) # layer 1 head 1

49. s41 = select(q=new_a2, k=new_a1, op=(q==k)) # layer 1 head 1

50. a7 = aggregate(s=s25+s26+s27+s28+s29+s30+s31+s32+s33+s34+s35+s36+s37+s38+s39+s40+s41, v=token) # layer 1 head 1

51. a8 = aggregate(s=s25+s26+s27+s28+s29+s30+s31+s32+s33+s34+s35+s36+s37+s38+s39+s40+s41, v=a2) # layer 1 head 1

52. new_new_a1 = element_wise_op(new_a1) # layer 1 mlp

53. new_new_a2 = element_wise_op(new_a2) # layer 1 mlp

54. logits1 = project(inp=a4, op=\circled{u})

55. logits2 = project(inp=a5, op=\circled{v})

56. logits3 = project(inp=a6, op=\circled{w})

57. logits4 = project(inp=a7, op=\circled{x})

58. logits5 = project(inp=a8, op=\circled{y})

59. logits6 = project(inp=a1, op=\circled{z})

60. logits7 = project(inp=a4, op=\circled{{})

61. logits8 = project(inp=a5, op=\circled{|})

62. logits9 = project(inp=a6, op=\circled{}})

63. logits10 = project(inp=a7, op=\circled{~})

64. logits11 = project(inp=token, op=\circled{�})

65. logits12 = project(inp=new_new_a1, op=(inp==out))

66. logits13 = project(inp=new_new_a2, op=(inp==out))

67. prediction = softmax(logits1+

logits2+

logits3+

logits4+

logits5+

logits6+

logits7+

logits8+

logits9+

logits10+

logits11+

logits12+

logits13)

72

73 of 95

D-RASP Program for HI -> HF Translation

73

74 of 95

D-RASP Program for HI -> HF Translation

74

75 of 95

D-RASP Program for HI -> HF Translation

75

76 of 95

D-RASP Program for HI -> HF Translation

76

77 of 95

RASP Language

RASP-L

  • L stands for “learnable”/length generalization
  • RASP Generalization Conjecture:
    • Transformers tend to length generalize on a task if the task can be solved by a short RASP program which works for all input lengths.

Results for HI -> HF Task

  1. AI generated RASP-L Program to solve HI -> HF task
    1. Length generalizable → O(D) layers to solve D layers of nesting (relative clauses)

77

78 of 95

RASP-L HI -> HF Translation

  1. Phase 1: POS Tagging
  2. Phase 2: Constituency Parsing
  3. Phase 3: Reordering

  1. Phase 1 and 3 verified by hand
  2. Phase 2 is more involved and has 3 independent implementations

78

79 of 95

Subtree Reversal as Addition (Displacement)

79

80 of 95

Subtree Reversal as Addition (Displacement)

80

81 of 95

Subtree Reversal as Addition (Displacement)

81

82 of 95

Subtree Reversal as Addition (Displacement)

82

83 of 95

Subtree Reversal as Addition (Displacement)

83

84 of 95

Subtree Reversal as Addition (Displacement)

84

85 of 95

Subtree Reversal as Addition (Displacement)

85

86 of 95

Subtree Reversal as Addition (Displacement)

86

87 of 95

Subtree Reversal as Addition (Displacement)

87

88 of 95

Subtree Reversal as Addition (Displacement)

88

89 of 95

Subtree Reversal as Addition (Displacement)

89

90 of 95

Subtree Reversal as Addition (Displacement)

90

91 of 95

Subtree Reversal as Addition (Displacement)

91

92 of 95

Subtree Reversal as Addition (Displacement)

92

93 of 95

Algorithm for Calculating Displacement

  1. displace(x) = for all flipping ancestors F:
    1. + |F.comp| if head(F)
    2. - |F.head| if comp(F)

displace(the) = 0 because no flipping ancestors

displace(boy) = + |NPsing.comp| = +2

displace(that) = +|CP_rel.comp| - |NPsing.head| = 1-1 = 0

displace(swims) = -|CP_rel.head| - |NPsing.head| = 1-2 = -2

displace(dances) = 0 because no flipping ancestors

93

94 of 95

Evaluating RASP-L

  1. RASP-L is a language for expressing transformer computations
    1. SGD not necessarily learns this algorithm proposed
  2. RASP-L is not a RASP dialect that can be directly compiled into transformer weights
  3. Can allow us to hypothesize high-level causal model for causal abstraction experiments

94

95 of 95

Next Steps

  1. Finish verifying the RASP-L algorithm to build intuitions on transformer mechanisms
  2. Start Causal Mediation experiment for first noun prediction
  3. Understand the Dyck-k,d literature
    1. Smaller models may be less interpretable because the algorithms are distributed across hidden dimensions rather than layers
    2. Upper bound limits on pretrained models for translation

95