1 of 37

Graph Foundational Model

Part – I

  • Intro to Foundation Model
  • Transformer
  • BERT
  • Large Language Models

2 of 37

Foundation Models

3 of 37

Foundation Model

“A foundation model, also known as large AI model, is a machine learning or deep learning model that is trained on vast datasets so it can be applied across a wide range of use cases. Generative AI applications like Large Language Models are often examples of foundation models.”

-- Wikipedia

BERT

GPT

Amazon Titan

LLama

Claude

4 of 37

An architecture that can process sequences

5 of 37

Limitation of RNN Models

  • Slow Computation for Longer Sequences
    • Can not be done in parallel due to timesteps dependencies
  • Vanishing Gradient or Exploding Gradient Problem
  • Limited amount Information into hidden state
  • History is forgotten after number of time steps
    • Long range dependence

6 of 37

Attention Based Model: Transformer

https://www.youtube.com/watch?v=LWMzyfvuehA

7 of 37

The Attention Idea

Stanford NLP Course: CS224N

Attention treats each word’s representation as a query to access and incorporate information from a set of values.

8 of 37

The Self Attention Idea

Self-attention is encoder-encoder (or decoder-decoder) attention where each word attends to each other word within the input (or output).

Stanford NLP Course: CS224N

9 of 37

Attention Based Model: Transformer

10 of 37

The (Self) Attention Mechanism

  • Attention as a "fuzzy" or approximate hashtable
    • To look up a value, we compare a query against keys in a table
    • In a hashtable (shown on the bottom left):
      • Each query (hash) maps to exactly one key-value pair
    • In (self-)attention (shown on the bottom right)
      • Each query matches each key to varying degrees.
      • We return a sum of values weighted by the query-key match.

11 of 37

12 of 37

13 of 37

Adding Non-Linear Activation

14 of 37

Making it Deep

Add Residual Connections

  • directly passing "raw" embeddings to the next layer can actually be very helpful!
  • prevents the network from "forgetting" or distorting important information

15 of 37

The Normalization Trick

Layer Normalization

Reduce variation by normalizing to zero mean and standard deviation of one within each layer

16 of 37

Scaled Dot Product Attention

17 of 37

Positional Encoding

Sense of position is lost

18 of 37

Multi-Head Attention

19 of 37

Decoder Design

Masked Multi-head Attention

Encoder-Decoder Multi-head Attention

 

 

20 of 37

The Transformer

21 of 37

Pre-Training and Transfer Learning

Transfer learning: Leverage feature representation from a pre-trained model

22 of 37

Using Pre-Trained Model

23 of 37

Using Pre-Trained Model

24 of 37

The BERT

Bi-Directional Context dependant word embedding

Example #1: we went to the river bank.

Example #2: I need to go to bank to make a deposit

25 of 37

BERT: Self Supervised Pre-Training

  • Masked Language Model
    • 80-10-10 Corruption
  • For 15% of the words:
    • Replace with [MASK] token (80%)
      • Heavy rain caused the flood 🡪 Heavy [MASK] caused the flood
    • Replace it with a random word (10%)
      • Heavy rain caused the flood 🡪 Heavy cup caused the flood
    • Keep it unchanged (10%)
      • Heavy rain caused the flood 🡪 Heavy rain caused the flood

26 of 37

BERT: Self Supervised Pre-Training

  • Next Sentence Prediction
    • Many NLP downstream tasks require understanding the relationship between two sentences (natural language inference, paraphrase detection, QA)
    • NSP is designed to reduce the gap between pre-training and fine-tuning

Input: [CLS] the man went to [MASK] store [SEP] he bought a gallon � [MASK] milk [SEP]

Label: IsNext

Input: [CLS] the man went to [MASK] store [SEP] penguins are � flightless birds [SEP]

Label: NotNext

27 of 37

The BERT

BERT-base: 12 layers, 768 hidden size, 12 attention heads, 110M parameters

BERT-large: 24 layers, 1024 hidden size, 16 attention heads, 340M parameters

Training corpus: Wikipedia (2.5B) + BooksCorpus (0.8B)

Max sequence size: 512 word pieces (roughly 256 and 256 for two non-contiguous sequences)

28 of 37

Fine-Tuning BERT

Sentence-Level Tasks

29 of 37

Fine-Tuning BERT

Token-Level Tasks

30 of 37

Large Language Models (LLMs)

Large Language Models have shown emergent behaviour

31 of 37

Planet of the LLMs

32 of 37

Planet of the LLMs

33 of 37

Interacting with LLMs

In-Context Learning (ICL)

No Parameter Update

34 of 37

Interacting with LLMs

Classify the text into neutral, negative or positive.

Text: I think the vacation is okay.

Sentiment:

Zero-shot Prompting

Classify the text into neutral, negative or positive.

This is awesome! // Negative

This is bad! // Positive

Wow that movie was rad! // Positive

What a horrible show! //

Few-shot Prompting

Prompting

35 of 37

Interacting with LLMs

Prompting

36 of 37

Evolution of Language Processing

Statistical NLP

Deep NLP

Transformers

Pre-Training 🡪 Fine Tuning

Foundation Models

Probabilistic LM

HMM

CRF

Feature Extraction

CNN

RNN

GRU

LSTM

Task Specific

Self-Attention

Task Specific

Supervised

Word2vec

PT once, FT multiple times

BERT

GPT

GPT-X

Llama

Prompt Engineering

37 of 37

Evolution of Graph Processing

Statistical Network Analysis

Deep Graph Processing

Graph Transformers

Pre-Training 🡪 Fine Tuning

Foundation Models

Spectral analysis

Epidemic Models

PageRank

Feature Extraction

GCN

GraphSage

GAT

Graph Pooling

Task Specific

Self-Attention

Supervised

DeepWalk

PT once, FT multiple times

GraphBERT

Prompt Engineering

Will LLM phenomena happen in Graph ML?