1 of 50

Brain-Inspired AI 2.0: Aligning Language Models Across Languages and Modalities

Subba Reddy Oota1, Tanmoy Chakraborty2, Manish Gupta3,4, Raju S. Bapi3

1TU Berlin, Germany; 2IIT Delhi, India; 3IIIT Hyderabad, India; 4Microsoft, India

subba.reddy.oota@tu-berlin.de, tanchak@iitd.ac.in, gmanish@microsoft.com, raju.bapi@iiit.ac.in

AAAI 2026: Brain-Inspired AI 2.0

2 of 50

Agenda

  • Introduction to the tutorial [10 min]
  • Introduction to Brain Encoding and Decoding [50 min]
  • Brain Encoding: Scaling Laws, Multilinguality, Multimodal and Instruction-tuned Models [60 min]
  • Coffee Break & Networking [30 min]
  • Brain-informed Fine-tuning of Language Models [30 min]
  • Brain-based Interpretability and Causal Testing of AI Models [20 min]
  • Brain Decoding [30 min]
  • Summary and Future Trends [10 min]

AAAI 2026: Brain-Inspired AI 2.0

3 of 50

Agenda

  • Introduction to the tutorial [10 min]
  • Introduction to Brain Encoding and Decoding [50 min]
  • Brain Encoding: Scaling Laws, Multilinguality, Multimodal and Instruction-tuned Models [60 min]
  • Coffee Break & Networking [30 min]
  • Brain-informed Fine-tuning of Language Models [30 min]
  • Brain-based Interpretability and Causal Testing of AI Models [20 min]
  • Brain Decoding [30 min]
  • Summary and Future Trends [10 min]

AAAI 2026: Brain-Inspired AI 2.0

4 of 50

Agenda

  • Introduction to the tutorial [10 min]
  • Introduction to Brain Encoding and Decoding [50 min]
    • Brain Encoding/Decoding and applications
    • Introduction to popular datasets
    • Text Stimulus Representations
    • Alignment Between AI Models and Human Brain Language Comprehension
  • Brain Encoding: Scaling Laws, Multilinguality, Multimodal and Instruction-tuned Models [60 min]
  • Coffee Break & Networking [30 min]
  • Brain-informed Fine-tuning of Language Models [30 min]
  • Brain-based Interpretability and Causal Testing of AI Models [20 min]
  • Brain Decoding [30 min]
  • Summary and Future Trends [10 min]

AAAI 2026: Brain-Inspired AI 2.0

5 of 50

Agenda

  • Introduction to the tutorial [10 min]
  • Introduction to Brain Encoding and Decoding [50 min]
    • Brain Encoding/Decoding and applications
    • Introduction to popular datasets
    • Text Stimulus Representations
    • Alignment Between AI Models and Human Brain Language Comprehension
  • Brain Encoding: Scaling Laws, Multilinguality, Multimodal and Instruction-tuned Models [60 min]
  • Coffee Break & Networking [30 min]
  • Brain-informed Fine-tuning of Language Models [30 min]
  • Brain-based Interpretability and Causal Testing of AI Models [20 min]
  • Brain Decoding [30 min]
  • Summary and Future Trends [10 min]

AAAI 2026: Brain-Inspired AI 2.0

6 of 50

Neuroscience

  • Field of science that studies the structure and function of the nervous system of different species.
  • How is information represented and processed in the brain?
    • How does the brain integrate text, vision, sound, and touch?
    • How do abstract concepts (like numbers or language) get encoded?
  • How does the brain learn and change?
    • What are the neural mechanisms that allow learning, memory, and adaptation?
    • How are memories formed but not overwritten?
    • How does the brain learn from very few examples?
    • How does RL work biologically?
  • How do large-scale brain circuits coordinate behavior?
    • How do distributed brain regions work together to produce coherent behavior?
    • How does the brain decide when to act or inhibit action?
    • How do sensory inputs get transformed into motor outputs?

AAAI 2026: Brain-Inspired AI 2.0

7 of 50

Neuroscience

  • What goes wrong in neurological and psychiatric disorders, and how can we fix it?
    • How do disruptions in brain structure or function cause diseases (Alzheimer’s, depression, Parkinson’s, autism, and schizophrenia), and how can we treat them?
    • Precision psychiatry: Can we predict which treatment will work for whom, when, and why?
  • How does the brain give rise to the mind and consciousness?
    • How do electrical and chemical signals in neurons produce subjective experiences like thoughts, emotions, awareness, and the sense of self?
    • Is consciousness localized or distributed?
    • Can machines ever be conscious?
    • What changes in the brain when consciousness is lost?
  • Brain-computer interfaces

AAAI 2026: Brain-Inspired AI 2.0

8 of 50

Brain encoding and decoding in cognitive neuroscience

8

20-Jan-26

 

AAAI 2026: Brain-Inspired AI 2.0

9 of 50

Brain encoding and decoding

  •  

AAAI 2026: Brain-Inspired AI 2.0

10 of 50

Techniques for studying the brain function

  • fMRI: high spatial but low time resolution.
    • Good to study a specific location in the brain
    • Unsuitable for sentence-level analysis. fMRI takes about two seconds to complete a scan. This is far lower than the speed at which humans can process language.
    • Cannot capture syntactic information (Gauthier and Levy, 2019)
  • EEG: high time but low spatial resolution.
    • Can preserve rich syntactic information (Hale et al., 2018)
    • But cannot use for source analysis.
  • fNIRS: compromise option
    • Time resolution better than fMRI
    • Spatial resolution better than EEG
    • Balance of spatial and temporal resolution may not be enough to compensate for the loss in both.

10

20-Jan-26

Single Micro-Electrode (ME), Micro-Electrode array (MEA), Electro-Cortico Graphy (ECoG), Positron emission tomography (PET), functional MRI (fMRI), Magneto-encephalography (MEG), Electro-encephalography (EEG), Near-Infrared Spectroscopy (NIRS)

AAAI 2026: Brain-Inspired AI 2.0

11 of 50

Computational Cognitive Science Research goals

  • Predictive Accuracy
    • Compare feature sets: Which feature set provides the most faithful reflection of the neural representational space?
    • Test feature decodability: “Does neural data Y contain information about features X?”
    • Build accurate models of brain data: Aim is to enable simulations of neuroscience experiments.
  • Interpretability
    • Examine individual features: Which features contribute the most to neural activity?
    • Test correspondences between representational spaces
      • “CNNs vs ventral visual stream” or “Two text representations”
    • Interpret feature sets
      • Do features X, generated by a known process, accurately describe the space of neural responses Y?
      • Do voxels respond to a single feature or exhibit mixed selectivity?
    • How does the mapping relate to other models or theories of brain function?

AAAI 2026: Brain-Inspired AI 2.0

12 of 50

Agenda

  • Introduction to the tutorial [10 min]
  • Introduction to Brain Encoding and Decoding [50 min]
    • Brain Encoding/Decoding and applications
    • Introduction to popular datasets
    • Text Stimulus Representations
    • Alignment Between AI Models and Human Brain Language Comprehension
  • Brain Encoding: Scaling Laws, Multilinguality, Multimodal and Instruction-tuned Models [60 min]
  • Coffee Break & Networking [30 min]
  • Brain-informed Fine-tuning of Language Models [30 min]
  • Brain-based Interpretability and Causal Testing of AI Models [20 min]
  • Brain Decoding [30 min]
  • Summary and Future Trends [10 min]

AAAI 2026: Brain-Inspired AI 2.0

13 of 50

Types of stimuli and popular datasets

  • Text (Words, Sentences, Paragraphs): Harry Potter Story, ZUCO EEG, Question-Answering MEG.
  • Visual: Binary visual patterns, Natural Images (Vim-1), BOLD5000, Algonauts and SS-fMRI.
  • Audio: Alice’s Adventures in Wonderland, Narratives, The Moth Radio Hour, Audio stories.
  • Videos: BBC’s Doctor Who, Japanese Ads, Pippi Langkous, Algonauts.
  • Other Multimodal Stimuli: Words + line drawing of concept named by each word, Pereira.

AAAI 2026: Brain-Inspired AI 2.0

14 of 50

Forms of stimulus presentation and data collection

  • Type: fMRI, EEG, MEG, …
  • TR: Sampling time.
  • Fixation points: location, color, shape.
  • Form of stimuli presentation: text, video, audio, images.
  • Task: question answering, property generation, understanding, …
  • Time given to participants: 1 minute to list properties, …
  • Type of participants: males/females, sighted/blind, …
  • Number of times the response to stimuli was recorded.
  • Stimulus Language

AAAI 2026: Brain-Inspired AI 2.0

15 of 50

Text Stimulus Datasets

Dataset

Type

Language

Stimulus

#Subjects

Paradigm

Size

Task

Wehbe et al., 2014

fMRI

English

Chapter 9 of Harry Potter and the Sorcerer's Stone

9

Reading stories

5000 word chapter was presented in 45 minutes.

Story understanding

Handjaras et al., 2016

fMRI

Italian

Verbal, pictorial or auditory presentation of 40 concrete nouns

20

Reading, viewing or listening

40 nouns * 4 times.

Property Generation

Anderson et al., 2017

fMRI

Italian

70 concrete and abstract nouns from law/music.

7

Reading

70 nouns * 5 times.

Imagine a situation that they personally associate with the noun

Zurich Cognitive Language Processing Corpus (ZuCo): Hollenstein et al., 2018

EEG and eye-tracking

English

Sentences from movie reviews or Wikipedia

12

Reading natural sentences

21,629 words in 1107 sentences and 154,173 fixations

Rate movie quality, answer control questions, check for existence of a relation

Anderson et al., 2019

fMRI

English

240 active voice sentences describing everyday situations

14

Reading

240 sentences seen 12 times (by 10 subjects) and 6 times (by 4 subjects)

Passive reading

BCCWJ-EEG: Oseki and Asahara, 2020

EEG

Japanese

20 newspaper articles

40

Reading

1 time reading for ~30-40 minutes

Passive reading

Deniz et al., 2019

fMRI

English

Subset of Moth Radio Hour. 11 stories

9

Reading

11 10- to 15 min stories presented twice word by word

Passive reading and Listening

AAAI 2026: Brain-Inspired AI 2.0

16 of 50

Visual Stimulus Datasets

Dataset

Type

Stimulus

#S

Paradigm

Size

Task

Thirion et al., 2006

fMRI

Rotating wedges, expanding/contracting rings, rotating Gabor filters, grid

9

Viewing visual patterns

Wedges/rings for 8 times, 36 Gabor filters for 4 times, grid 36 times

Passive viewing, imagine one of the 6 domino stimuli when prompted to.

Vim-1: Kay et al., 2008

fMRI

Sequences of natural photos

2

Viewing natural images

Each subject viewed 1750 (Stage 1)+ 120 (Stage 2) novel natural images

Passive viewing

Horikawa et al., 2017

fMRI

Object images

5

Viewing and Reading

Each subject: (1) Image presentation: 1,200 images from 150 object categories and 50 images from 50 object categories; (2) Imagery: 10 times.

One-back repetition detection task, imagine object images pertaining to the category

BOLD5000: Chang et al., 2019

fMRI

5254 images depicting real-world scenes

4

Viewing natural images

∼20 hours of MRI scans per each of four participants

Passive viewing

Algonauts: Cichy et al., 2019

fMRI (EVC and IT)/MEG (early and late in time)

Object images

15

Viewing object images

92 silhouette object images and 118 images of objects on natural background

Passive viewing

Natural Scenes Dataset: Allen et al., 2022

fMRI

73000 natural scenes

8

Viewing natural scenes

~73000 distinct natural scene images from MSCOCO.

Passive viewing

THINGS: Hebart et al., 2023

fMRI/EEG

31188 natural images across 1,854 object concepts.

8

Viewing natural images

fMRI: 3 Participants. 8,740 unique images. 720 objects. MEG: 4 Participants. 22,448 unique images. 1,854 objects

oddball detection task (synthetic image).

AAAI 2026: Brain-Inspired AI 2.0

17 of 50

Audio Stimulus Datasets

Dataset

Type

Language

Stimulus

#S

Paradigm

Size

Task

Handjaras et al., 2016

fMRI

Italian

Verbal, pictorial or auditory presentation of 40 concrete nouns

20

Reading, viewing or listening

40 nouns * 4 times.

Property Generation

Huth et al., 2016

fMRI

English

Eleven 10-minute stories

7

Listening

2 hours of stories from The Moth Radio Hour

Passive Listening

Brennan and Hale, 2019

EEG

English

Chapter one of Alice’s Adventures in Wonderland as read by Kristen McQuillan

33

Listening

2,129 words in 84 sentences. The entire experimental session lasted 1–1.5 h (including QA).

8 MCQ Question answering concerning the contents of the story

Anderson et al., 2020

fMRI

English

One of 20 scenario names

26

Listening scenario name

20 scenario prompts displayed 5 times.

Imagine themselves personally experiencing common scenarios

Narratives: Nastase et al., 2021

fMRI

English

27 diverse naturalistic spoken stories

345

Listening

891 functional scans, totaling ~4.6 hours of unique stimuli (~43,000 words)

Passive Listening

Natural Stories: Zhang et al., 2020

fMRI

English

Moth-Radio-Hour naturalistic spoken stories

19

Listening

5 h 33 m (repeated twice). Each story is 6 m 48 s avg or 2492 words.

Passive Listening

The Little Prince: Li et al., 2021

fMRI

English, Chinese, French

Audiobook

112

Listening

English audiobook is 94 minutes long. Chinese: 99min. French: 97 min.

Passive Listening. 4 quiz questions.

MEG-MASC: Gwilliams et al., 2022

MEG

English

4 English fictional stories: Cable spool boy, LW1, Black willow, Easy money.

27

Listening

Two hours of naturalistic stories. 208 MEG sensors.

Passive Listening

AAAI 2026: Brain-Inspired AI 2.0

18 of 50

Video Stimulus Datasets

Dataset

Type

Language

Stimulus

#Subjects

Paradigm

Size

Task

BBC’s Doctor Who: Seeliger et al., 2019

fMRI

English

Spatiotemporal visual and auditory naturalistic stimuli (30 episodes of BBC’s Doctor Who)

1

Viewing episode videos

120.830 whole-brain volumes (approx. 23 h) of single-presentation data, and 1.178 volumes (11 min) of repeated narrative short episodes (22 repetitions)

Passive viewing

Japanese Ads: Nishida et al., 2020

fMRI

Japanese

368 web and 2452 TV Japanese ad movies (15-30s)

40 and 28 for web and TV ads. 16 were overlapped

Viewing Ads

7200 train and 1200 test fMRIs for web; fMRIs from 420 ads.

Passive viewing

Pippi Langkous: Berezutskaya et al., 2020

ECoG

The movie was originally in Swedish but dubbed in Dutch

30 s excerpts of a feature film (in total, 6.5 min long), edited together for a coherent story

37 patients

Viewing

6.5 min movie.

Passive viewing

Algonauts: Cichy et al., 2021

fMRI

English

1000 short video clips

10

Viewing video clips

1000 short video clips (3 sec each)

Passive viewing

Natural Short Clips: Huth et al., 2022

fMRI

English

Natural short movie clips

5

Watching natural short movie clips

3870 responses per subject.

Passive viewing

AAAI 2026: Brain-Inspired AI 2.0

19 of 50

Other Multimodal Stimulus Datasets

Dataset

Type

Language

Stimulus

#Subjects

Paradigm

Size

Task

Mitchell et al., 2008

fMRI

English

60 different word-picture pairs from 12 categories.

9

Viewing word-picture pairs

60 different word-picture pairs presented six times each

Passive viewing

Sudre et al., 2012

MEG

English

60 concrete nouns along with line drawings

9

Reading

60 stimuli × 20 questions = 1200 examples

Question answering

Zinszer et al., 2017

fNIRS

English

8 concrete nouns (audiovisual word and picture stimuli): bunny, bear, kitty, dog, mouth, foot, hand, and nose

24

Viewing and listening

12 blocks with the 8 stimuli per subject.

Passive viewing and listening

Pereira et al., 2018

fMRI

English

180 Words with Picture, Sentences, word clouds; 96 text passages; 72 passages

16

Viewing WP, sentences or word clouds

180 WP, S and WC per subject; 96+72 passages shown 3 times

Passive viewing

Cao et al., 2021

fNIRS

Chinese

50 concrete nouns from 10 semantic categories

7

Viewing and listening

Each stimulus is presented 7 times.

Passive viewing and listening

Courtois Neuromod

fMRI

English

full-length movies and TV show

6

Viewing and Listening

~100 hours of data per participant

Passive viewing

AAAI 2026: Brain-Inspired AI 2.0

20 of 50

Agenda

  • Introduction to the tutorial [10 min]
  • Introduction to Brain Encoding and Decoding [50 min]
    • Brain Encoding/Decoding and applications
    • Introduction to popular datasets
    • Text Stimulus Representations
    • Alignment Between AI Models and Human Brain Language Comprehension
  • Brain Encoding: Scaling Laws, Multilinguality, Multimodal and Instruction-tuned Models [60 min]
  • Coffee Break & Networking [30 min]
  • Brain-informed Fine-tuning of Language Models [30 min]
  • Brain-based Interpretability and Causal Testing of AI Models [20 min]
  • Brain Decoding [30 min]
  • Summary and Future Trends [10 min]

AAAI 2026: Brain-Inspired AI 2.0

21 of 50

Text Stimulus Representations

  • Basic NLP Representations
    • Corpus co-occurrence counts
    • Topic models
    • Linguistic: POS, dependencies, roles.
  • Discourse
    • Characters, motion, speech, emotions, non-motion verbs
  • Deep Learning based Representations
    • Embeddings
    • Longer context using LSTMs
    • Transformers
  • Experiential attributes
    • Rated on 0-6 scale
    • Binary

AAAI 2026: Brain-Inspired AI 2.0

22 of 50

Basic NLP Representations for Word Stimuli

  • Corpus co-occurrence counts
    • 25 verbs (Mitchell et al., 2008; Pereira et al., 2013)
      • Basic sensory and motor activities, actions performed on objects, and actions involving changes to spatial relationships.
      • Verbs: see, hear, listen, taste, smell, eat, touch, nib, lift, manipulate, run, push, fill, move, ride, say, fear, open, approach, near, enter, drive, wear, break, and clean.
      • For each (verb, stimulus word w), feature value = normalized co-occurrence count of w with any of three forms of the verb (e.g., taste, tastes, or tasted) over the text corpus.
    • 985 common English words (such as above, worry, and mother) in (Huth et al., 2016).
  • Topic models (Pereira et al., 2013)
    • Get relevant Wiki pages (e.g., “airplane” is “Fixed-Wing Aircraft”) and other linked pages (e.g. “Aircraft cabin”)
    • LDA topic modelling on 3500 pages with #topics from 10 to 100, in increments of 5, setting the α parameter to 25/#topics.
    • LSA topic modelling (Wang et al., 2017)

22

20-Jan-26

AAAI 2026: Brain-Inspired AI 2.0

23 of 50

Basic NLP Representations for Word Stimuli

  • Word length
  • Is the word related to one of the 28 unique parts of speech and 17 unique dependency relationships?
  • Position of word in the sentence
  • Roles
    • Main verb
    • Agent or experiencer
    • Patient or recipient
    • Predicate of a sentence (The window was dusty)
    • Modifier (The angry activist broke the chair)
    • Complement in adjunct and propositional phrase, including direction, location, and time (The restaurant was loud at night).

AAAI 2026: Brain-Inspired AI 2.0

24 of 50

Discourse features (for Harry Potter dataset)

  • Characters: Binary features to signal which of the 10 characters are mentioned.
  • Motions: Identify a set of motions that occurred frequently in the chapter (e.g. fly, manipulate, collide physically, etc.).
  • Speech: Indicate the parts of the story that correspond to direct speech between the characters. Used the presence of dialog as a feature. 
  • Emotions: Identified a set of emotions that were felt by the characters in the chapter (e.g. annoyance, nervousness, pride, etc.).
  • Verbs: Identified a set of actions that occurred frequently in the chapter that were distinct from motion (e.g. hear, know, see, etc.).

AAAI 2026: Brain-Inspired AI 2.0

25 of 50

DL Representations: Using embeddings for word stimuli

  • GloVe 300D vectors (Pereira et al., 2016; Wang et al., 2017; Pereira et al., 2018; Anderson et al., 2019)
  • 300D embeddings by training a skip-gram model using negative sampling (SGNS) on Italian and English Wikipedia dumps using Gensim. (Anderson et al., 2017a)
  • FastText (Berezutskaya et al., 2020)
  • Comparison across multiple embedding methods
    • GloVe, word2vec, WordNet2Vec, FastText, ELMo (Hollenstein et al., 2019)
    • word2Vec, fastText, GloVe, Dependency-based word2vec, RWSGwn, ConceptNet, ELMo, averaged and concatenated combinations (Wang et al., 2020)

AAAI 2026: Brain-Inspired AI 2.0

26 of 50

DL Representations: Using longer context for word stimuli

  • Multi-task LSTMs
    • Predict next word and POS of next word.
  • ELMo embeddings: LSTM based pretrained language model

26

20-Jan-26

AAAI 2026: Brain-Inspired AI 2.0

27 of 50

DL Representations: Using sentence embeddings

  • Unstructured Models: Ignore sentence structure
    • Simple Pooling Methods
      • Average/max/concat(max, avg) pooling over word embeddings.
    • Advanced Pooling Methods
      • FastSent (Hill, Cho, and Korhonen 2016)
      • SIF (Arora, Liang, and Ma 2016) adapts the naïve averaging of word embeddings to weighted averaging.
  • Structured Models
    • Unsupervised Methods: Skip-thought, QuickThought.
    • Supervised Methods: InferSent, GenSen (Subramanian et al. 2018), Universal Sentence Encoder

27

20-Jan-26

AAAI 2026: Brain-Inspired AI 2.0

28 of 50

DL Representations: Transformer-based methods for text stimuli (Layer #, context length, architecture)

28

20-Jan-26

Transformer-XL is the only model that continues to increase performance as the context length is increased. In all networks, the middle layers perform the best for contexts longer than 15 words. The deepest layers across all networks show a sharp increase in performance at short-range context (fewer than 10 words), followed by a decrease in performance. [Toneva and Wehbe, 2019]

AAAI 2026: Brain-Inspired AI 2.0

29 of 50

DL Representations: Transformer-based methods for text stimuli (NLP task finetuning)

29

20-Jan-26

Tasks

Paraphrase, Summarization, Question Answering, Sentiment Analysis, NER, Word Sense Disambiguation, Natural Language Inference, Semantic Role Labeling, Coreference Resolution, Shallow Syntax Parsing

Pereira dataset: CR, NER, and SS perform the best.

Dendrogram constructed using similarity on representations from task-specific Transformer encoder models with stimuli from the dataset passed as input.

AAAI 2026: Brain-Inspired AI 2.0

30 of 50

DL Representations: Transformer-based methods for text stimuli (Multi-task setup)

  • Settings
    • Finetune BERT vs not
    • Finetune BERT using one representative subject and train dense layer for each subject, vs finetune BERT for each subject.
    • Finetune BERT on MEG for all subjects, then finetune BERT on fMRI.
    • Multi-task finetune BERT for fMRI+MEG prediction task
  • Results
    • Fine-tuned models predict fMRI data better than vanilla BERT
    • Using MEG data can improve fMRI predictions.
    • A single model can be used to predict fMRI activity across multiple experiment participants.

30

20-Jan-26

AAAI 2026: Brain-Inspired AI 2.0

31 of 50

Experiential attributes model for text stimuli

  • Represents words in terms of human ratings of their degree of association with different attributes of experience
    • “On a scale of 0 to 6, to what degree do you think of a banana as having a characteristic or defining color?”
    • Anderson et al., 2019: 65 attributes spanning sensory, motor, affective, spatial, temporal, causal, social, and abstract cognitive experiences.
  • Value-add on top of text models: a lot of experiential information goes unstated in natural verbal communication.
    • E.g., it is rarely useful to communicate the color of bananas because it is obvious to all those with experience of bananas.
    • E.g., it would be unusual to specify that dropping things involves movement.
  • Nishida et al., 2020 use a subset of 20 attributes.

31

20-Jan-26

AAAI 2026: Brain-Inspired AI 2.0

32 of 50

Binary attribute representations

  • Each stimulus is represented using a binary vector capturing membership to one of the eight semantic categories.

  • 42 neurally plausible semantic features (NPSFs)
    • 10 Perceptual and affective characteristics of an entity. E.g, man-made, size, color, temperature, positive affective valence, high affective arousal
    • Animate beings (person, human-group, animal)
    • Time and space properties (e.g. unenclosed setting, change of location)

32

20-Jan-26

AAAI 2026: Brain-Inspired AI 2.0

33 of 50

Agenda

  • Introduction to the tutorial [10 min]
  • Introduction to Brain Encoding and Decoding [50 min]
    • Brain Encoding/Decoding and applications
    • Introduction to popular datasets
    • Text Stimulus Representations
    • Alignment Between AI Models and Human Brain Language Comprehension
  • Brain Encoding: Scaling Laws, Multilinguality, Multimodal and Instruction-tuned Models [60 min]
  • Coffee Break & Networking [30 min]
  • Brain-informed Fine-tuning of Language Models [30 min]
  • Brain-based Interpretability and Causal Testing of AI Models [20 min]
  • Brain Decoding [30 min]
  • Summary and Future Trends [10 min]

AAAI 2026: Brain-Inspired AI 2.0

34 of 50

Does context size impact alignment?

  • ELMo, BERT, USE (Universal Sentence Encoder) and T-XL.
  • Transformer-XL continues to increase perf as the context length is increased.
  • Middle layers perform the best for contexts of 15+ words.
  • Deepest layers show a sharp increase in perf at short-range context (<10 words).

Toneva, Mariya, and Leila Wehbe. "Interpreting and improving natural-language processing (in machines) with natural language-processing (in the brain)." Advances in neural information processing systems 32 (2019).

34

20-Jan-26

AAAI 2026: Brain-Inspired AI 2.0

35 of 50

What if we use uniform attention in first few layers?

  • Shallow layers benefit from the uniform attention for context lengths up to 25 words.

Toneva, Mariya, and Leila Wehbe. "Interpreting and improving natural-language processing (in machines) with natural language-processing (in the brain)." Advances in neural information processing systems 32 (2019).

  •  

35

20-Jan-26

AAAI 2026: Brain-Inspired AI 2.0

36 of 50

Is the representation of semantic information different when reading vs listening?

Deniz, Fatma, Anwar O. Nunez-Elizalde, Alexander G. Huth, and Jack L. Gallant. "The representation of semantic information across human cerebral cortex during listening versus reading is invariant to stimulus modality." Journal of Neuroscience 39, no. 39 (2019): 7722-7736.

  • Train on fMRI for 10 narrative stories (from The Moth Radio Hour) while participants listened to or read several hours
  • Voxel-wise modeling with banded ridge regression
  • During reading, each word was presented for a duration=duration of that word in the spoken story.
  • Features
    • 39 Motion-energy features using spatiotemporal Gabor pyramid
    • 80 spectral audio features based on cochleogram.
    • Word rate, phoneme rate, letter rate, word length variation per TR
    • 39 phoneme frequency
    • 26 letter frequency
    • 56 syntactic binary features: 12 POS and 44 dependency
    • 985 co-occurrence semantics

36

20-Jan-26

AAAI 2026: Brain-Inspired AI 2.0

37 of 50

Is the representation of semantic information different when reading vs listening?

Deniz, Fatma, Anwar O. Nunez-Elizalde, Alexander G. Huth, and Jack L. Gallant. "The representation of semantic information across human cerebral cortex during listening versus reading is invariant to stimulus modality." Journal of Neuroscience 39, no. 39 (2019): 7722-7736.

  • Semantic tuning during listening and reading are highly correlated in most semantically selective regions of cortex
  • Models estimated using one modality accurately predict voxel responses in the other modality.

37

20-Jan-26

AAAI 2026: Brain-Inspired AI 2.0

38 of 50

Do larger Transformer models lead to better brain‑encoding accuracy?

Schrimpf, Martin, Idan Asher Blank, Greta Tuckute, Carina Kauf, Eghbal A. Hosseini, Nancy Kanwisher, Joshua B. Tenenbaum, and Evelina Fedorenko. "The neural architecture of language: Integrative modeling converges on predictive processing." Proceedings of the National Academy of Sciences 118, no. 45 (2021): e2105646118.

  • Transformer models predict nearly 100% of explainable variance in neural responses to sentences.
    • Larger models are better. GPT2-XL is the best.

38

20-Jan-26

AAAI 2026: Brain-Inspired AI 2.0

39 of 50

Does improved perf of Transformer models on NLP benchmarks translate to better brain‑encoding accuracy?

Schrimpf, Martin, Idan Asher Blank, Greta Tuckute, Carina Kauf, Eghbal A. Hosseini, Nancy Kanwisher, Joshua B. Tenenbaum, and Evelina Fedorenko. "The neural architecture of language: Integrative modeling converges on predictive processing." Proceedings of the National Academy of Sciences 118, no. 45 (2021): e2105646118.

  • Performance on GLUE tasks does not predict brain scores.

  • Next-Word-Prediction Task Performance Selectively Predicts Brain Scores.

39

20-Jan-26

AAAI 2026: Brain-Inspired AI 2.0

40 of 50

Does improved perf of Transformer models on NLP benchmarks translate to better brain‑encoding accuracy?

Schrimpf, Martin, Idan Asher Blank, Greta Tuckute, Carina Kauf, Eghbal A. Hosseini, Nancy Kanwisher, Joshua B. Tenenbaum, and Evelina Fedorenko. "The neural architecture of language: Integrative modeling converges on predictive processing." Proceedings of the National Academy of Sciences 118, no. 45 (2021): e2105646118.

  • Perf on other GLUE tasks does not correlate with behavioral scores
  • Behavioral scores (self-paced reading times), brain scores, and next-word-prediction task perf are pairwise correlated.

40

20-Jan-26

AAAI 2026: Brain-Inspired AI 2.0

41 of 50

Which NLP Tasks are the most Predictive of fMRI Brain Activity?

  • Coreference resolution, NER, and shallow syntax parsing explain greater variance for the reading activity.
  • For the listening activity, paraphrase generation, summarization, and natural language inference show better encoding performance.
  • Tree derived from predicted brain responses vs tree based on Transformer encoder embeddings.

Oota, Subba Reddy, Jashn Arora, Veeral Agarwal, Mounika Marreddy, Manish Gupta, and Bapi Surampudi. "Neural language taskonomy: Which NLP tasks are the most predictive of fMRI brain activity?." In NAACL-HLT, pp. 3220-3237. 2022.

41

20-Jan-26

AAAI 2026: Brain-Inspired AI 2.0

42 of 50

Does the brain also perform next word prediction and is surprised when next word does not match the prediction?

  • ECoG while listening to a 30-min podcast.
  • Human brains and DLMs
    • Both are engaged in continuous next-word prediction before word onset
    • Both match their pre-onset predictions to the incoming word to calculate post-onset surprise
    • Both rely on contextual embeddings to represent words in natural contexts.

Goldstein, Ariel, Zaid Zada, Eliav Buchnik, Mariano Schain, Amy Price, Bobbi Aubrey, Samuel A. Nastase et al. "Shared computational principles for language processing in humans and deep language models." Nature neuroscience 25, no. 3 (2022): 369-380.

42

20-Jan-26

AAAI 2026: Brain-Inspired AI 2.0

43 of 50

Does the brain also perform next word prediction and is surprised when next word does not match the prediction?

  • GloVe yielded significant correlations with predicted neural responses to upcoming words up to 800 ms before word onset.
  • Pre-onset activity associated with next-word prediction matches prediction content even when the prediction was incorrect.

Goldstein, Ariel, Zaid Zada, Eliav Buchnik, Mariano Schain, Amy Price, Bobbi Aubrey, Samuel A. Nastase et al. "Shared computational principles for language processing in humans and deep language models." Nature neuroscience 25, no. 3 (2022): 369-380.

  • Post-onset activity matches content of the incoming word, even if it was unpredicted.
  • Increase in encoding perf for surprising words compared to predicted words 400 ms after word onset.

43

20-Jan-26

AAAI 2026: Brain-Inspired AI 2.0

44 of 50

Does the brain depend on context for processing text?

  • Match between confidence level and the accuracy level of GPT-2 and human predictions
    • Under-confidence in their predictions and were above 95% correct when the probabilities were higher than 40%.
  • Correlation between human and GPT2 word predictions improved as the contextual window increased.

Goldstein, Ariel, Zaid Zada, Eliav Buchnik, Mariano Schain, Amy Price, Bobbi Aubrey, Samuel A. Nastase et al. "Shared computational principles for language processing in humans and deep language models." Nature neuroscience 25, no. 3 (2022): 369-380.

  • Contextual embeddings (GPT-2) > static embeddings (GloVe)
  • Contextual embeddings (GPT-2) > concatenated GloVe embeddings for 10 prev words
  • Avg context: Averaging all embeddings for each unique word (all occurrences of ‘monkey’) into 1 vector
  • Scramble the embeddings across different occurrences of the same word in the story (switch embedding of ‘monkey’ in sentence 5 with that in sentence 50).

44

20-Jan-26

AAAI 2026: Brain-Inspired AI 2.0

45 of 50

Do speech models align with language and speech ROIs during speech production and comprehension?

  • Extracted low-level acoustic, mid-level speech and contextual word embeddings from a multimodal speech-to-text model (Whisper).
  • Sensory and motor regions better align with the model’s speech embeddings
  • Higher-level language areas better align with the model’s language embeddings.

Goldstein, Ariel, Haocheng Wang, Leonard Niekerken, Mariano Schain, Zaid Zada, Bobbi Aubrey, Tom Sheffer et al. "A unified acoustic-to-speech-to-language embedding space captures the neural basis of natural language processing in everyday conversations." Nature human behaviour (2025): 1-15.

45

20-Jan-26

AAAI 2026: Brain-Inspired AI 2.0

46 of 50

Are both text and audio needed for speech alignment?

  • Only text input (conversation transcripts) vs language receiving audio and text inputs (speech recordings and conversation transcripts)
  • Language embeddings fused with auditory features outperform text-only language embeddings in predicting neural activity across multiple electrodes.
  • Even though IFG is associated with linguistic processing, across multiple lags, the audio-fused language embeddings yield higher encoding performance during both production and comprehension.

Goldstein, Ariel, Haocheng Wang, Leonard Niekerken, Mariano Schain, Zaid Zada, Bobbi Aubrey, Tom Sheffer et al. "A unified acoustic-to-speech-to-language embedding space captures the neural basis of natural language processing in everyday conversations." Nature human behaviour (2025): 1-15.

46

20-Jan-26

AAAI 2026: Brain-Inspired AI 2.0

47 of 50

Which speech model aligns the best?

  • Wav2Vec2.0, HuBERT, Data2Vec align well.
  • Data2Vec aligns best with both language and auditory brain regions

Oota, Subba Reddy, Khushbu Pahwa, Mounika Marreddy, Manish Gupta, and Bapi S. Raju. "Neural architecture of speech." In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 1-5. IEEE, 2023.

47

20-Jan-26

AAAI 2026: Brain-Inspired AI 2.0

48 of 50

Which Speech Tasks are the most Predictive of fMRI Brain Activity?

  • 8 tasks from Speech processing Universal PERformance Benchmark (SUPERB): Phoneme Recognition (PR), Automatic Speech Recognition (ASR), Keyword Spotting (KS), Intent Classification (IC), Speaker Diarization (SD), Speaker Verification (SV), Speaker Identification (SID), and Emotion Recognition (ER)

Oota, Subba Reddy, Veeral Agarwal, Mounika Marreddy, Manish Gupta, and Raju Surampudi Bapi. "Speech taskonomy: Which speech tasks are the most predictive of fmri brain activity?." In INTERSPEECH 2023-24th INTERSPEECH Conference, pp. 5167-5171. 2023.

  • ASR finetuning (middle layers) yields the best encoding performance for the whole brain, language and auditory regions.
  • Finetuning on ER, SID and IC leads to the best alignment for the early auditory cortex.

48

20-Jan-26

AAAI 2026: Brain-Inspired AI 2.0

49 of 50

A big thank you!

49

20-Jan-26

Tutorial, Code and Material:

Material from AAAI 2026 Tutorial would be uploaded soon!

(Past): Deep Learning for Brain Encoding and Decoding, Cogsci-2022

https://tinyurl.com/DL4Brain

(Past): Language and the Brain: Deep Learning for Brain Encoding and Decoding, IJCNN 2023

https://tinyurl.com/DLBrainIJCNN2023

(Past): Deep Neural Networks and Brain Alignment: Brain Encoding and Decoding, IJCAI 2023

https://tinyurl.com/DLBrainIJCAI2023

AAAI 2026: Brain-Inspired AI 2.0

50 of 50

Thanks!

AAAI 2026: Brain-Inspired AI 2.0