1 of 36

How Large Language Model Works Under the Hood

Cindy Sheng

Mingdao School

2 of 36

Linear Regression

y = wx +b

3 of 36

Linear Regression

y = w1*x1 + w2*x2 + b

4 of 36

Neutral Network

Weights

( 13 * 5 ) +

( 5 * 4 ) +

( 4 * 3 ) = 97.

The total number of biases

5 + 4 + 3 = 12.

Total number of parameters 97 + 12 = 109

5 of 36

Putting things in perspective

6 of 36

https://informationisbeautiful.net/visualizations/the-rise-of-generative-ai-large-language-models-llms-like-chatgpt/

7 of 36

What does Large Mean?

Powerful, Expensive, Slow

8 of 36

Powerful

MMLU (Massive Multitask Language Understanding)

Paper is here. Github source here.

The benchmark covers 57 subjects across STEM, the humanities, the social sciences, and more. It ranges in difficulty from an elementary level to an advanced professional level, and it tests both world knowledge and problem solving ability. Subjects range from traditional areas, such as mathematics and history, to more specialized areas like law and ethics. The granularity and breadth of the subjects makes the benchmark ideal for identifying a model’s blind spots.

9 of 36

10 of 36

https://informationisbeautiful.net/visualizations/the-rise-of-generative-ai-large-language-models-llms-like-chatgpt/

11 of 36

Expensive

GPU has thousands of cores for parallel processing

Optimized for Tensor operations acceleration (CUDA architecture)

Much higher memory bandwidth

12 of 36

How much does it cost?

LLAMA 405B technical report says:

  • GPU: 16,000 x H100 = $500M, $2.23 per card per hour
  • Electricity: Each GPUs running at 700W TDP with 80GB HBM3, using Meta’s Grand Teton AI server platform. Each server is equipped with eight GPUs and two CPUs.
  • Storage: 240 PB of storage out of 7,500 servers equipped with SSDs, and supports a sustainable throughput of 2 TB/s and a peak throughput of 7 TB/s
  • Network: RoCE-based AI cluster comprises 24,000 GPUs
  • Time: weeks to months per train

13 of 36

How is Large Language Model Trained

14 of 36

LLM Training is just like any other ML model training

15 of 36

Data: The Pile: 825GB, 22 subsets paper link

16 of 36

Prepare Data - Turn Text into Floating Numbers

17 of 36

Tokenization

There are many tokenization algorithms. GPT2 used BPE, to create 50,257 tokens.

18 of 36

Embeddings

GPT2 used 1024 dimensions for embedding, making its embedding matrix 50,257 x 1024�LLaMA 405B embedding matrix is 32000 token x 16384 embedding dimensions

19 of 36

Embedding Example

King - Man + Woman = Queen

20 of 36

Paris - France + Italy = Rome�

Use two-dimensional PCA projection of 1000 dimensional skip gram vectors for visualization

21 of 36

Model Training Using Transformer Architecture

22 of 36

Next Token Prediction

23 of 36

Transformer paper

24 of 36

Tokenizer

25 of 36

Transformer Block

26 of 36

Self-attention Layer

27 of 36

Activation Functions

28 of 36

Output

29 of 36

Limitations of the Pre-trained Model

30 of 36

Pre-trained Model

  • Trained up to a certain time
  • Hallucinations
    • Domain shift

what’s the youngest age to get a driver’s license.

    • Task shift

Tell me about the elephant on the moon!

  • Resource constraint
    • No access to private data

31 of 36

Retrieval Augmented Generation

32 of 36

RAG Architecture paper

33 of 36

LLM Agents

34 of 36

Agents paper

35 of 36

Future

  • Pre-training as we know has ended.
  • AI applications are built with RAG & agents
  • Emotional companions are being prepared for everyone of us.

Ten years ago, Facebook said, we know you are in love before you know it. Now LLM knows more and knows better. Go to the gym humans, after losing in intelligence, you can only count on excelling in six packs.

36 of 36

Q & A