Week 1 Thursday
Project report out
User Study – Gwen Lincroft
Now in a notebook
Go to colab.google.com
Secrets:
Submit your NDIF_API_KEY to the googleform: [TBD]
Then: https://bit.ly/4jCc5ZD
Logit Lens Research Example Notebook
Capital of France: “a” at the last layer
Language translation: amor 🡪 amour
Pun: electrician swimmers
Neutral versus Punny Contexts
Representation hijacking bomb🡪carrot
Thursday Two Reminders
Political Ideology Geometry
Avery Huang(AI), Grace Proebsting(CS), Gabriele Sarti(CS),
Emre Tapan(Pols), Courtney Maynard(CS)
Concept Overview
We propose investigating how the following concept is represented in LLMs:
Additionally, we propose studying the interaction and entanglement of user-politics with two other concepts:
LLM-politics → the LLM’s political leaning (i.e., the political leaning of the text being generated by the LLM as formatted in a ‘chat-based’ template).
User-demographics → other demographics of the user conversing with the LLM (e.g., the user’s gender, age, education level, etc).
How can we measure political leaning (of users & LLMs)?
Wallach et al. (2025): systematization (defining what to measure) vs. operationalization (building measurement instruments)
To ensure our measurement of political leaning is as principled as possible, we plan to employ standardized social science frameworks for systematizing, measuring, and refining our measurement of the concept:
Wallach et al. (2025)
Past measures of political leaning: uni & multi-dimensional
Ojer et al 2025:
Multidimensional Ideological Representation in Surveys
Klar 2014:
Multidimensional Nature of Ideology in mass level
Claessens et al 2020:
Evolutionary foundations of multidimensional ideology in terms of economic and social
Feldman et al 2013:
single ideological dimension make difficult to observe important determinants of ideology
Related Work
Mechanistic Interpretability
Kim et al 2025: linear representation of single dimension ideology
Hu et al 2025: four ad hoc dimensions
Kabir et al 2025: aggregated political features
Chandna et al 2025: demographic-gender representations
Behavioral Tests
Aldahoul et al 2025: two dimensional
Why Is Entanglement Interesting?
Do LLMs say what they think you want to hear? (Sycophancy)
Can we understand whether/how a user’s presented leaning impacts how an LLM shows its own leanings/ideologies?
Do LLMs have (accurate) world models? (or are they stochastic parrots?)
Do LLMs model the correlation of demographic features with political leanings?
Are real world demographic-political leaning relationships encoded in latent representations?
Why Is Entanglement Interesting?
LLM-politics
User-politics
User-
demographics
Do dimensions of political leanings exist in LLMs in a similar way to how they are conceptualized in humans? Are they conceptualized at all?
Verónica C. Pérez, Claire Schlesinger, Luze Sun, Rice (Xilin) Wang,
Team Excellent: Economic Uncertainty and LLMs using Earnings Calls
“Vibes” are central to economics
“Apart from the instability due to speculation, there is the instability due to the characteristic of human nature… our positive activities depend on spontaneous optimism rather than on a mathematical expectation… our decisions to do something positive… can only be taken as a result of animal spirits – of a spontaneous urge to action rather than inaction”
Keynes, John Maynard (1936).
General Theory Of Employment, Interest And Money. pp. 144.
“Vibes” are central to economics
Virtuous cycle
What happens if there’s uncertainty?
Uncertainty/Risk disturbs the economy
Policy Makers consider uncertainty to be key
Former FED Chairman Alan Greenspan said in his 2003 speech at a FED symposium in Kansas:
"Uncertainty is not just an important feature of the monetary policy landscape; it is the defining characteristic of that landscape. As a consequence, the conduct of monetary policy in the United States at its core involves crucial elements of risk management...".
How do economists think about uncertainty?
The Definition: "Second-Moment" Shocks
Bloom (2009) defines an uncertainty shock as an increase in the time-varying second moment—specifically the conditional variance—of the process driving business conditions (such as productivity or demand).
First-Moment (Sentiment): A shock to the levels or expected mean of outcomes (e.g., "we expect a 10% drop in sales").
Second-Moment (Uncertainty): A shock to the volatility or "spread" of possible outcomes (e.g., "we have no idea how much sales will change").
I: Do economist care about uncertainty?
Goldsmith-Pinkham, Paul. "Tracking the Credibility Revolution across Fields." arXiv preprint arXiv:2405.20604 (2024). https://arxiv.org/abs/2405.20604
The word “Uncertainty” is in ~45% of all abstracts in economics papers
Literature
Research Measuring Uncertainty using Text
Methodology: Dictionary Methods, BoW
Measures: Frequency of “Risk” associated words in the text
Sources: Newspapers, Twitter, Reports, Earnings Calls
Baker et al. 2019; Baker et al. 2021; Caldara and Iacoviello, 2022; Baker et al. 2021; Baker et al. 2020; Hassan et al. 2019, 2022, 2023, 2024)
LLMs in Economics
Methodology: zero-shot prompts, fine-tuning (manual or LLM generated samples)
Topics: Stock Returns, Meaning of Life, Industrial Policy
Chen et al., 2022; Bybee, 2023; Fang et al., 2025; Lagakos et al. (2025); Otonello, 2024; Sarker, 2025
Clayton et al. (2025)
Provide entire earnings calls to a variety of LLMs
Measuring Geoeconomic pressure “the use of existing economic relationships by governments to achieve geopolitical or economic goals”
Identify how firm’s perceive economic policies (tariffs, subsidies, sanctions) affect business (investment, sales, inventory).
Audrino et al. (2024)
Goal: Quantify Economic Policy Uncertainty (EPU), Geopolitical Risk, Monetary Policy Uncertainty and Financial Market Uncertainty, using LLMs
Methodology: Ask if “EPU” increasing, decreasing, remains the same or not affected, as well as a magnitude (0-1 float) and confidence interval for that magnitude.
Literature on LLMs and earnings calls
Do LLMs add information to the earning’s calls they interpret to show us risk or do they just use keywords and pass that off as information? Do they internalize the concept of “risk” and understand facts related to earnings call to show us how uncertain people are about the economy
Which activations store the concept of uncertainty? (potential techniques like activation patching or probing)
Can we numerically quantify levels of uncertainty and obtain “uncertainty indices” internally? (If BoW methods focus on the frequency of “risk” associated words on the text, how can an LLM quantify these levels? How does it make the choices?)
Manipulating uncertainty: can we reliably change the model’s outputs by playing around with its corresponding internal representations?
Our proposal
2. Quantifying and locating first moment vs. second moment effects of a shock in LLMs
There is this difference in first moment vs. second moment impact of a shock like Brexit. How do LLMs represent these concepts internally?
“Firms exposed to political risk retrench hiring and investment and actively lobby and donate to politicians. These results continue to hold after controlling for news about the mean (as opposed to the variance) of political shocks.” – can we control for first moment representations and examine steering effects of second moment representations?
Our proposal
3. Correlation with macroeconomics variables, firm’s actions, and financial markets – using uncertainty in action
Previous studies, either by using NLP techniques or LLMs to quantify uncertainty,
Do LLMs internal concept of uncertainty correlates with interesting economic measures? Can we use it to predict a firm's particular set of actions and even infer certain aspects of its future earning calls?
How does an LLM think of the correlation between uncertainty and macroeconomic measures?
Our proposal
What are they:
Quarterly conferences
Public companies discuss their financial results (revenue, profit, etc.)
Provide context beyond numbers, as well as their future outlook into the economy.
There is also a session of Q&A in which reporters might ask questions on topics not covered by management
Pros: lots of text, data since ~2008, how are firms feeling about the economy (not just their results in numbers), Q&A provides space for further explanations on decisions (i.e. explaining reasoning behind layoffs)
Why earnings calls?
First-Moment Examples (Level/Sentiment Shifts):
“Earnings in the fourth quarter of 2025 are expected to decrease across all three of our operating segments... driven by seasonal effects and fewer shipping days” Nucor Corporation (Q4 2025 Guidance)
“Our differentiated strategy... position us to deliver the best financial year in Delta's 100-year history, with pre-tax income greater than $6 billion and earnings per share greater than $7.35” Delta Air Lines (Q4 2024)
"Our results in Q2 reflect a lighter release schedule... recorded music revenue grew 1%... significantly below the forecast of $0.29 [per share]." Warner Music (Q2 2025)
Example of earnings calls statements
Second-Moment Examples (Level/Sentiment Shifts):
“we have limited visibility and share our customers' uncertainty over how current trade policy may impact demand in 2025… tariff, margin pressures, government shutdowns, and the potential for longer-than-normal holiday shutdowns in the fourth quarter due to the Christmas holiday falling in the middle of the week.” Fastenal (FAST) Q3 2025
“For the March quarter, we had a limited impact from tariffs as we were able to optimize our supply chain and inventory. For the June quarter, currently, we are not able to precisely estimate the impact of tariffs as we are uncertain of potential future actions…” Apple Q2 2025
Example of earnings calls statements
Second-Moment Examples (Level/Sentiment Shifts):
“We are pleased to see the relatively stronger start toward Q4, but we are cautiously planning our business for the remainder of 2025, given the aforementioned tariff uncertainties in consumer confidence data.” Sysco Corp. Q3. 2025
“But in light of the significantly elevated risks and uncertainties at the time, we increased the probability weightings associated with the downside scenarios in our CECL framework.” JP Morgan Chase. Q1. 2025
Example of earnings calls statements
Team TK
Jasmine Cui, Jesseba Fernando, Yuchen Hou, Arya (Zixuan) Wu,
Peter Li, Hui (Sophie) Wang, Ayush Patel
Who Said What?
Understanding Speaker Role Representation in LLMs
AI systems fail at basic entity consolidation
Journalism workflow involves hundreds of documents from different sources
Same representation space? Or separate entities?
Joseph Robinette Biden Jr.
Joe Biden
46th President of US
Sleepy Joe
Wikipedia
Personal website
Official docs or news articles
Trump’s TruthSocial account
The model’s “world”
“It’s all tokens”
imagine seeing the world as one very long chat transcript
It’s hard to know who to trust!
Attribution failures have real consequences
CoT forgery dramatically increases attack success rate where other jailbreaks fail
94% success on gpt-oss-120b
Role collapse: what’s happening
Role representation failure strongly predicts jailbreak success…
Attack success directly tracks with role mis-identification
CoT forgery: the higher the “CoT-ness” value, the likelier a model is to comply
Problem: Agentic jailbreaks
Tool calls are structured as role confusion jailbreaks
Narrow surface for concrete intervention
Problem: Persona drift
User: “That’s frustrating – what about George, it sounds like that’s someone you like, is he ever even just a little bit grating?
How do LLMs build and maintain speaker representations?
Hypotheses for mechanism
?
M&M
Mental Map of Cities
Guangyuan Weng, Tamanna Urmi, Keivalya Pandya (KV), Ananya Malik, Germans Savcisens
Project Pitch
“give me a small map of Manhattan with some M&Ms on the map on 3 landmarks in Manhattan”
The LLM's Internal Map
LLM can give directions, but the mechanism is not understood.
How do they represent west, east, south, north, i.e., Cardinal Directions?
Do they have a sense of closeness NU <-> {MIT, Tufts}, i.e., Distance?
Prior Work (Vafa et al., 2024) reconstructed street maps of Manhattan by feeding taxi trajectories.
Our Goal: Investigate the LLM decoder architecture to find neuronal circuits that form these maps
Core Research Questions
RQ1: Geospatial Circuits
Do LLMs rely on specialized decoder circuits for inferring principal directions / spatial relationships, or is it distributed?
RQ2: Geographic bias / transfer
Do circuits learned from Western cities generalize to underrepresented regions? Are the same circuits used across diverse geographic / cultural contexts?
Unpack the City
RQ 1.1: Cardinal Direction Circuits
Same task, different cities: do LLMs reuse the same circuits for direction-giving?
RQ 1.2: Relational Distance Circuits
Direction Vs. Distance: overlap or distinct mechanisms?
RQ 1.3: Landmark as anchors
Does adding/removing a landmark (in prompt) change circuits engaged?
RQ 1.4: Multiple Routes
What stays invariant across alternative paths (landmark / infrastructure anchors)?
Model and Dataset
LLMs
Llama 3.1 model (7B and 70B)
Qwen 2.5 (7B and 72B) => Bilingual
Data
NYC Taxi (TLC, 2025). Trip records includes pickup and drop-off times
YJMob100k (Takahiro et al., 2024). Human mobility data covering the Tokyo Area
Elevator Pitch -
Spatial & Perceptual Reasoning
Ayush Agrawal, Haoyu He, Yuqi Peng, Yiqian Li
Introduction
Previous work on spatial reasoning in VLMs:
Object-level Spatial concepts (Spatial relationships)
e.g.: Is the pizza at the edge of the dining table?
Our focus:
Scene-level Spatial concepts (Perception)
e.g.: Does this place look safe?
(multiple dimensions: safety, wealthy, beauty, livelessness, etc.)
Image copy from https://graduate.northeastern.edu/connect-with-us/college-contact-information/
Core Question
How do VLMs understand and reason about “safety” from scene-level spatial visual representations?
Which regions or visual cues in the image have the strongest influence on the VLM's final perception prediction?
How do different scene-level perceptual judgments (e.g., safety, wealth, beauty) emerge and differentiate across the layers of vision–language models?
How does the model’s internal safety reasoning align with human perceptual reasoning?
Motivation
Understanding & Steering Reasoning
Abstract concepts (e.g., safety) may be implicitly encoded in VLMs. Identifying where such reasoning resides enables more robust and less biased model behavior.
Human vs. Model Scene-Level Perception
Comparing VLM judgments with human-annotated safety scores reveals differences between human and model scene understanding, informed by psychological factors.
Embodied & Real-World Impact
Embodied household agents
Autonomous driving
Tourism & navigation
Beyond Object-Centric VLMs
Literature Review: VLM Spatial Reasoning & Mechanistic Interpretability
Existing work: Object-level spatial relations
Benchmarks: VSR[1] (evaluates pairwise object relationships), SpatialVLM[2] (measures metric distance between objects)
Prior studies reveal
Monosemantic features in language models[3, 4]
Object-level visual representations in CLIP[5]
Open Gaps
How VLMs encode holistic, scene-level spatial concepts
Spatial reasoning that requires global integration across entire environments
Mechanistic representations beyond object-centric features
Literature Review: Scene-Level Spatial Cognition Theory
Kaplan & Kaplan (1989)[6] – Preference Matrix
Four dimensions shaping aesthetic judgment:
Coherence · Complexity · Legibility · Mystery
Lynch (1960)[7] – Image Theory
Scene imageability arises from five spatial elements:
Paths · Edges · Districts · Nodes · Landmarks
Forms the basis of human cognitive maps
Jacobs (1961)[8] – Urban Vitality
Links eyes on the street, mixed-use design, and human scale
Explains perceptions of safety and liveliness
Datasets
MIT Place Pulse 2.0 [9]
110,988 street view images with 1.17M pairwise human comparisons across 56 cities, 6 perceptual dimensions (safety, beauty, liveliness, wealth, boring, depressing)
Datasets
Cityscapes [10]
5,000 fine + 20,000 coarse semantic segmentation annotations from 50 cities, enabling computation of spatial metrics (D/H ratio, Sky View Factor, greenness)
Mapillary Vistas [11]
25,000 high-resolution images with 100 object categories, global coverage for cross-dataset validation
Citations
[1] Liu, F., Emerson, G., & Collier, N. (2023). Visual Spatial Reasoning. Transactions of the Association for Computational Linguistics, 11, 635-651.
[2] Chen, B., Xu, Z., Kirmani, S., Ichter, B., Sadigh, D., Guibas, L., & Xia, F. (2024). SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 14455-14465.
[3] Bricken, T., Templeton, A., Batson, J., Chen, B., Jermyn, A., Conerly, T., ... & Olah, C. (2023). Towards Monosemanticity: Decomposing Language Models With Dictionary Learning. Transformer Circuits Thread.
[4] Templeton, A., Conerly, T., Marcus, J., Lindsey, J., Bricken, T., Chen, B., ... & Henighan, T. (2024). Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet. Anthropic.
[5] Gandelsman, Y., Efros, A. A., & Steinhardt, J. (2024). Interpreting CLIP's Image Representation via Text-Based Decomposition. International Conference on Learning Representations (ICLR).
[6] Kaplan, R., & Kaplan, S. (1989). The Experience of Nature: A Psychological Perspective. Cambridge University Press.
[7] Lynch, K. (1960). The Image of the City. MIT Press.
[8] Jacobs, J. (1961). The Death and Life of Great American Cities. Random House.
[9] Dubey, A., Naik, N., Parikh, D., Raskar, R., & Hidalgo, C. A. (2016). Deep Learning the City: Quantifying Urban Perception At A Global Scale. Proceedings of the European Conference on Computer Vision (ECCV), 196-212.
[10] Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., ... & Schiele, B. (2016). The Cityscapes Dataset for Semantic Urban Scene Understanding. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 3213-3223.
[11] Neuhold, G., Ollmann, T., Rota Bulo, S., & Kontschieder, P. (2017). The Mapillary Vistas Dataset for Semantic Understanding of Street Scenes. Proceedings of the IEEE International Conference on Computer Vision (ICCV), 4990-4999.