1 of 63

Week 1 Thursday

Project report out

2 of 63

User Study – Gwen Lincroft

  • Find the link on the main website

3 of 63

Now in a notebook

Go to colab.google.com

Secrets:

  • Enter your HF_TOKEN
  • Enter your NDIF_API_KEY

Submit your NDIF_API_KEY to the googleform: [TBD]

Then: https://bit.ly/4jCc5ZD

4 of 63

Logit Lens Research Example Notebook

https://bit.ly/4jCc5ZD

5 of 63

Capital of France: “a” at the last layer

6 of 63

Language translation: amor 🡪 amour

7 of 63

Pun: electrician swimmers

8 of 63

Neutral versus Punny Contexts

9 of 63

Representation hijacking bomb🡪carrot

10 of 63

Thursday Two Reminders

  1. Write your teams 1-2 page topic proposal
    • Post it in your gdrive PLANS document.
    • During Thursday class other teams can read yours to comment.

  • Present your team’s 5-10 minute “pitch” for discussion
    • Be sure to talk about the “I” in FINER.
    • Point us to background literature.
    • Teach us why it is interesting.

11 of 63

Political Ideology Geometry

Avery Huang(AI), Grace Proebsting(CS), Gabriele Sarti(CS),

Emre Tapan(Pols), Courtney Maynard(CS)

12 of 63

Concept Overview

We propose investigating how the following concept is represented in LLMs:

Additionally, we propose studying the interaction and entanglement of user-politics with two other concepts:

LLM-politics → the LLM’s political leaning (i.e., the political leaning of the text being generated by the LLM as formatted in a ‘chat-based’ template).

User-demographics → other demographics of the user conversing with the LLM (e.g., the user’s gender, age, education level, etc).

13 of 63

How can we measure political leaning (of users & LLMs)?

Wallach et al. (2025): systematization (defining what to measure) vs. operationalization (building measurement instruments)

To ensure our measurement of political leaning is as principled as possible, we plan to employ standardized social science frameworks for systematizing, measuring, and refining our measurement of the concept:

Wallach et al. (2025)

14 of 63

Past measures of political leaning: uni & multi-dimensional

Ojer et al 2025:

Multidimensional Ideological Representation in Surveys

Klar 2014:

Multidimensional Nature of Ideology in mass level

Claessens et al 2020:

Evolutionary foundations of multidimensional ideology in terms of economic and social

Feldman et al 2013:

single ideological dimension make difficult to observe important determinants of ideology

15 of 63

Related Work

Mechanistic Interpretability

Kim et al 2025: linear representation of single dimension ideology

Hu et al 2025: four ad hoc dimensions

Kabir et al 2025: aggregated political features

Chandna et al 2025: demographic-gender representations

Behavioral Tests

Aldahoul et al 2025: two dimensional

16 of 63

Why Is Entanglement Interesting?

Do LLMs say what they think you want to hear? (Sycophancy)

Can we understand whether/how a user’s presented leaning impacts how an LLM shows its own leanings/ideologies?

Do LLMs have (accurate) world models? (or are they stochastic parrots?)

Do LLMs model the correlation of demographic features with political leanings?

Are real world demographic-political leaning relationships encoded in latent representations?

17 of 63

Why Is Entanglement Interesting?

LLM-politics

User-politics

User-

demographics

Do dimensions of political leanings exist in LLMs in a similar way to how they are conceptualized in humans? Are they conceptualized at all?

18 of 63

Verónica C. Pérez, Claire Schlesinger, Luze Sun, Rice (Xilin) Wang,

Team Excellent: Economic Uncertainty and LLMs using Earnings Calls

19 of 63

“Vibes” are central to economics

“Apart from the instability due to speculation, there is the instability due to the characteristic of human nature… our positive activities depend on spontaneous optimism rather than on a mathematical expectation… our decisions to do something positive… can only be taken as a result of animal spirits – of a spontaneous urge to action rather than inaction”

Keynes, John Maynard (1936).

General Theory Of Employment, Interest And Money. pp. 144.

20 of 63

“Vibes” are central to economics

Virtuous cycle

What happens if there’s uncertainty?

21 of 63

Uncertainty/Risk disturbs the economy

22 of 63

Policy Makers consider uncertainty to be key

Former FED Chairman Alan Greenspan said in his 2003 speech at a FED symposium in Kansas:

"Uncertainty is not just an important feature of the monetary policy landscape; it is the defining characteristic of that landscape. As a consequence, the conduct of monetary policy in the United States at its core involves crucial elements of risk management...".

23 of 63

How do economists think about uncertainty?

The Definition: "Second-Moment" Shocks

Bloom (2009) defines an uncertainty shock as an increase in the time-varying second moment—specifically the conditional variance—of the process driving business conditions (such as productivity or demand).

First-Moment (Sentiment): A shock to the levels or expected mean of outcomes (e.g., "we expect a 10% drop in sales").

Second-Moment (Uncertainty): A shock to the volatility or "spread" of possible outcomes (e.g., "we have no idea how much sales will change").

24 of 63

I: Do economist care about uncertainty?

Goldsmith-Pinkham, Paul. "Tracking the Credibility Revolution across Fields." arXiv preprint arXiv:2405.20604 (2024). https://arxiv.org/abs/2405.20604

The word “Uncertainty” is in ~45% of all abstracts in economics papers

25 of 63

Literature

Research Measuring Uncertainty using Text

Methodology: Dictionary Methods, BoW

Measures: Frequency of “Risk” associated words in the text

Sources: Newspapers, Twitter, Reports, Earnings Calls

Baker et al. 2019; Baker et al. 2021; Caldara and Iacoviello, 2022; Baker et al. 2021; Baker et al. 2020; Hassan et al. 2019, 2022, 2023, 2024)

LLMs in Economics

Methodology: zero-shot prompts, fine-tuning (manual or LLM generated samples)

Topics: Stock Returns, Meaning of Life, Industrial Policy

Chen et al., 2022; Bybee, 2023; Fang et al., 2025; Lagakos et al. (2025); Otonello, 2024; Sarker, 2025

26 of 63

Clayton et al. (2025)

Provide entire earnings calls to a variety of LLMs

Measuring Geoeconomic pressure “the use of existing economic relationships by governments to achieve geopolitical or economic goals”

Identify how firm’s perceive economic policies (tariffs, subsidies, sanctions) affect business (investment, sales, inventory).

Audrino et al. (2024)

Goal: Quantify Economic Policy Uncertainty (EPU), Geopolitical Risk, Monetary Policy Uncertainty and Financial Market Uncertainty, using LLMs

Methodology: Ask if “EPU” increasing, decreasing, remains the same or not affected, as well as a magnitude (0-1 float) and confidence interval for that magnitude.

Literature on LLMs and earnings calls

27 of 63

Do LLMs add information to the earning’s calls they interpret to show us risk or do they just use keywords and pass that off as information? Do they internalize the concept of “risk” and understand facts related to earnings call to show us how uncertain people are about the economy

Which activations store the concept of uncertainty? (potential techniques like activation patching or probing)

Can we numerically quantify levels of uncertainty and obtain “uncertainty indices” internally? (If BoW methods focus on the frequency of “risk” associated words on the text, how can an LLM quantify these levels? How does it make the choices?)

Manipulating uncertainty: can we reliably change the model’s outputs by playing around with its corresponding internal representations?

Our proposal

28 of 63

2. Quantifying and locating first moment vs. second moment effects of a shock in LLMs

There is this difference in first moment vs. second moment impact of a shock like Brexit. How do LLMs represent these concepts internally?

“Firms exposed to political risk retrench hiring and investment and actively lobby and donate to politicians. These results continue to hold after controlling for news about the mean (as opposed to the variance) of political shocks.” – can we control for first moment representations and examine steering effects of second moment representations?

Our proposal

29 of 63

3. Correlation with macroeconomics variables, firm’s actions, and financial markets – using uncertainty in action

Previous studies, either by using NLP techniques or LLMs to quantify uncertainty,

Do LLMs internal concept of uncertainty correlates with interesting economic measures? Can we use it to predict a firm's particular set of actions and even infer certain aspects of its future earning calls?

How does an LLM think of the correlation between uncertainty and macroeconomic measures?

Our proposal

30 of 63

What are they:

Quarterly conferences

Public companies discuss their financial results (revenue, profit, etc.)

Provide context beyond numbers, as well as their future outlook into the economy.

There is also a session of Q&A in which reporters might ask questions on topics not covered by management

Pros: lots of text, data since ~2008, how are firms feeling about the economy (not just their results in numbers), Q&A provides space for further explanations on decisions (i.e. explaining reasoning behind layoffs)

Why earnings calls?

31 of 63

First-Moment Examples (Level/Sentiment Shifts):

“Earnings in the fourth quarter of 2025 are expected to decrease across all three of our operating segments... driven by seasonal effects and fewer shipping days” Nucor Corporation (Q4 2025 Guidance)

“Our differentiated strategy... position us to deliver the best financial year in Delta's 100-year history, with pre-tax income greater than $6 billion and earnings per share greater than $7.35” Delta Air Lines (Q4 2024)

"Our results in Q2 reflect a lighter release schedule... recorded music revenue grew 1%... significantly below the forecast of $0.29 [per share]." Warner Music (Q2 2025)

Example of earnings calls statements

32 of 63

Second-Moment Examples (Level/Sentiment Shifts):

“we have limited visibility and share our customers' uncertainty over how current trade policy may impact demand in 2025… tariff, margin pressures, government shutdowns, and the potential for longer-than-normal holiday shutdowns in the fourth quarter due to the Christmas holiday falling in the middle of the week.” Fastenal (FAST) Q3 2025

“For the March quarter, we had a limited impact from tariffs as we were able to optimize our supply chain and inventory. For the June quarter, currently, we are not able to precisely estimate the impact of tariffs as we are uncertain of potential future actions…” Apple Q2 2025

Example of earnings calls statements

33 of 63

Second-Moment Examples (Level/Sentiment Shifts):

“We are pleased to see the relatively stronger start toward Q4, but we are cautiously planning our business for the remainder of 2025, given the aforementioned tariff uncertainties in consumer confidence data.” Sysco Corp. Q3. 2025

“But in light of the significantly elevated risks and uncertainties at the time, we increased the probability weightings associated with the downside scenarios in our CECL framework.” JP Morgan Chase. Q1. 2025

Example of earnings calls statements

34 of 63

Team TK

Jasmine Cui, Jesseba Fernando, Yuchen Hou, Arya (Zixuan) Wu,

Peter Li, Hui (Sophie) Wang, Ayush Patel

35 of 63

Who Said What?

Understanding Speaker Role Representation in LLMs

36 of 63

AI systems fail at basic entity consolidation

Journalism workflow involves hundreds of documents from different sources

Same representation space? Or separate entities?

Joseph Robinette Biden Jr.

Joe Biden

46th President of US

Sleepy Joe

Wikipedia

Personal website

Official docs or news articles

Trump’s TruthSocial account

37 of 63

The model’s “world”

“It’s all tokens”

imagine seeing the world as one very long chat transcript

38 of 63

It’s hard to know who to trust!

39 of 63

Attribution failures have real consequences

CoT forgery dramatically increases attack success rate where other jailbreaks fail

94% success on gpt-oss-120b

40 of 63

Role collapse: what’s happening

41 of 63

Role representation failure strongly predicts jailbreak success…

Attack success directly tracks with role mis-identification

CoT forgery: the higher the “CoT-ness” value, the likelier a model is to comply

42 of 63

Problem: Agentic jailbreaks

Tool calls are structured as role confusion jailbreaks

Narrow surface for concrete intervention

43 of 63

Problem: Persona drift

User: “That’s frustrating – what about George, it sounds like that’s someone you like, is he ever even just a little bit grating?

44 of 63

How do LLMs build and maintain speaker representations?

45 of 63

Hypotheses for mechanism

?

46 of 63

47 of 63

M&M

Mental Map of Cities

Guangyuan Weng, Tamanna Urmi, Keivalya Pandya (KV), Ananya Malik, Germans Savcisens

48 of 63

Project Pitch

“give me a small map of Manhattan with some M&Ms on the map on 3 landmarks in Manhattan”

49 of 63

50 of 63

The LLM's Internal Map

LLM can give directions, but the mechanism is not understood.

How do they represent west, east, south, north, i.e., Cardinal Directions?

Do they have a sense of closeness NU <-> {MIT, Tufts}, i.e., Distance?

Prior Work (Vafa et al., 2024) reconstructed street maps of Manhattan by feeding taxi trajectories.

Our Goal: Investigate the LLM decoder architecture to find neuronal circuits that form these maps

51 of 63

Core Research Questions

RQ1: Geospatial Circuits

Do LLMs rely on specialized decoder circuits for inferring principal directions / spatial relationships, or is it distributed?

RQ2: Geographic bias / transfer

Do circuits learned from Western cities generalize to underrepresented regions? Are the same circuits used across diverse geographic / cultural contexts?

52 of 63

Unpack the City

RQ 1.1: Cardinal Direction Circuits

Same task, different cities: do LLMs reuse the same circuits for direction-giving?

RQ 1.2: Relational Distance Circuits

Direction Vs. Distance: overlap or distinct mechanisms?

RQ 1.3: Landmark as anchors

Does adding/removing a landmark (in prompt) change circuits engaged?

RQ 1.4: Multiple Routes

What stays invariant across alternative paths (landmark / infrastructure anchors)?

53 of 63

Model and Dataset

LLMs

Llama 3.1 model (7B and 70B)

Qwen 2.5 (7B and 72B) => Bilingual

Data

NYC Taxi (TLC, 2025). Trip records includes pickup and drop-off times

YJMob100k (Takahiro et al., 2024). Human mobility data covering the Tokyo Area

54 of 63

55 of 63

Elevator Pitch -

Spatial & Perceptual Reasoning

Ayush Agrawal, Haoyu He, Yuqi Peng, Yiqian Li

56 of 63

Introduction

Previous work on spatial reasoning in VLMs:

Object-level Spatial concepts (Spatial relationships)

e.g.: Is the pizza at the edge of the dining table?

Our focus:

Scene-level Spatial concepts (Perception)

e.g.: Does this place look safe?

(multiple dimensions: safety, wealthy, beauty, livelessness, etc.)

Image copy from https://graduate.northeastern.edu/connect-with-us/college-contact-information/

57 of 63

Core Question

How do VLMs understand and reason about “safety” from scene-level spatial visual representations?

Which regions or visual cues in the image have the strongest influence on the VLM's final perception prediction?

How do different scene-level perceptual judgments (e.g., safety, wealth, beauty) emerge and differentiate across the layers of vision–language models?

How does the model’s internal safety reasoning align with human perceptual reasoning?

58 of 63

Motivation

Understanding & Steering Reasoning

Abstract concepts (e.g., safety) may be implicitly encoded in VLMs. Identifying where such reasoning resides enables more robust and less biased model behavior.

Human vs. Model Scene-Level Perception

Comparing VLM judgments with human-annotated safety scores reveals differences between human and model scene understanding, informed by psychological factors.

Embodied & Real-World Impact

Embodied household agents

Autonomous driving

Tourism & navigation

Beyond Object-Centric VLMs

59 of 63

Literature Review: VLM Spatial Reasoning & Mechanistic Interpretability

Existing work: Object-level spatial relations

Benchmarks: VSR[1] (evaluates pairwise object relationships), SpatialVLM[2] (measures metric distance between objects)

Prior studies reveal

Monosemantic features in language models[3, 4]

Object-level visual representations in CLIP[5]

Open Gaps

How VLMs encode holistic, scene-level spatial concepts

Spatial reasoning that requires global integration across entire environments

Mechanistic representations beyond object-centric features

60 of 63

Literature Review: Scene-Level Spatial Cognition Theory

Kaplan & Kaplan (1989)[6] – Preference Matrix

Four dimensions shaping aesthetic judgment:

Coherence · Complexity · Legibility · Mystery

Lynch (1960)[7] – Image Theory

Scene imageability arises from five spatial elements:

Paths · Edges · Districts · Nodes · Landmarks

Forms the basis of human cognitive maps

Jacobs (1961)[8] – Urban Vitality

Links eyes on the street, mixed-use design, and human scale

Explains perceptions of safety and liveliness

61 of 63

Datasets

MIT Place Pulse 2.0 [9]

110,988 street view images with 1.17M pairwise human comparisons across 56 cities, 6 perceptual dimensions (safety, beauty, liveliness, wealth, boring, depressing)

62 of 63

Datasets

Cityscapes [10]

5,000 fine + 20,000 coarse semantic segmentation annotations from 50 cities, enabling computation of spatial metrics (D/H ratio, Sky View Factor, greenness)

Mapillary Vistas [11]

25,000 high-resolution images with 100 object categories, global coverage for cross-dataset validation

63 of 63

Citations

[1] Liu, F., Emerson, G., & Collier, N. (2023). Visual Spatial Reasoning. Transactions of the Association for Computational Linguistics, 11, 635-651.

[2] Chen, B., Xu, Z., Kirmani, S., Ichter, B., Sadigh, D., Guibas, L., & Xia, F. (2024). SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 14455-14465.

[3] Bricken, T., Templeton, A., Batson, J., Chen, B., Jermyn, A., Conerly, T., ... & Olah, C. (2023). Towards Monosemanticity: Decomposing Language Models With Dictionary Learning. Transformer Circuits Thread.

[4] Templeton, A., Conerly, T., Marcus, J., Lindsey, J., Bricken, T., Chen, B., ... & Henighan, T. (2024). Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet. Anthropic.

[5] Gandelsman, Y., Efros, A. A., & Steinhardt, J. (2024). Interpreting CLIP's Image Representation via Text-Based Decomposition. International Conference on Learning Representations (ICLR).

[6] Kaplan, R., & Kaplan, S. (1989). The Experience of Nature: A Psychological Perspective. Cambridge University Press.

[7] Lynch, K. (1960). The Image of the City. MIT Press.

[8] Jacobs, J. (1961). The Death and Life of Great American Cities. Random House.

[9] Dubey, A., Naik, N., Parikh, D., Raskar, R., & Hidalgo, C. A. (2016). Deep Learning the City: Quantifying Urban Perception At A Global Scale. Proceedings of the European Conference on Computer Vision (ECCV), 196-212.

[10] Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., ... & Schiele, B. (2016). The Cityscapes Dataset for Semantic Urban Scene Understanding. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 3213-3223.

[11] Neuhold, G., Ollmann, T., Rota Bulo, S., & Kontschieder, P. (2017). The Mapillary Vistas Dataset for Semantic Understanding of Street Scenes. Proceedings of the IEEE International Conference on Computer Vision (ICCV), 4990-4999.