1 of 37

AI Foundations

NLP - Natural Language Processing

PRESENTED BY:- SOMIL AGRAWAL

2 of 37

What is NLP ?

NLP stands for "Natural Language Processing," which is a subfield of Artificial Intelligence (AI) that focuses on enabling computers to understand and generate human language, utilizing techniques from both Machine Learning (ML) and Deep Learning (DL) to achieve this goal

3 of 37

What are the prerequisites ?

  • Python
  • Mathematics
  • Basic Machine Learning: K-means Clustering
  • Basic Deep Learning

4 of 37

NLP Pipeline

  1. Data Acquisition - We need data for our model to learn
  2. Text Preparation - We need to clean the data for use
  3. Feature Engineering - We need to prepare usable features from data
  4. Modelling - We need to create a model around the features and then evaluate - Different method for ML and DL
  5. Deployment - We need to deploy 😮‍💨 - Monitoring Performance - Continuous Learning, etc..

5 of 37

Data Acquisition

6 of 37

Text Processing

  • Cleaning
  • HTML Tag Removal
  • Emoji Removal - Replace or Remove
  • Spell Check

I am - i m

You - u

laughing out loud - lol 😂

7 of 37

Text Processing

Basic Processing

  • Tokenization [Important] - Into Words, Sentences
  • Optional
      • Stop Word removal
      • Stemming / lemmatization
      • Removing Digits
      • Lowercasing
      • Language Translation

8 of 37

Text Processing

Advanced Processing

  • POS (Parts of Speech)Tagging
  • Syntactic Parsing
  • Conference Resolution

[ I ] wrote scripts for most of [ my ] plays.

9 of 37

Stemming vs Lemmatization

10 of 37

POS Tagging (part-of-speech)

11 of 37

Synatctic Analysis

12 of 37

Feature Engineering

How to engineer features ?

Examples..

  • Length of Document - longer means autogenerated
  • Average word size within a document
  • use of punctuation in text
  • capitalization of words - Maybe use to infer emphasis on word
  • Word Embedding [imp]

13 of 37

Word Embedding

14 of 37

Text Processing

15 of 37

Common Terms

Corpus -> Whole dataset

Vocabulary -> All the unique words in the corpus

Document -> Piece of text in the corpus, maybe a sentence. Usually a single training example

Word -> each group of letters in document, i.e a normal word.

16 of 37

One Hot Vector

17 of 37

One Hot Vector

18 of 37

Bag Of Words

19 of 37

N-Grams

Instead of treating each word separate, what if we treat them in a group of 2, 3 or N

20 of 37

TF-IDF

21 of 37

Word2Vec

22 of 37

CBOW (Continuous Bag of Words)

  • Predicts word based on the context

google dream company software engineer

-> google dream company

-> dream company software

-> company software engineer

23 of 37

Skip-gram

  • Predicts context based on the word

google dream company software engineer

-> google dream company

-> dream company software

-> company software engineer

24 of 37

Where are the embeddings ?

25 of 37

Other Embeddings

    • GloVe (Global Vectors for Word Representation): Combines global word co-occurrence statistics with local context-based learning.

    • FastText: Extends Word2Vec by representing words as n-grams of characters, improving representations for rare words.

    • ELMo (Embeddings from Language Models): Uses deep, contextualized word representations.

    • BERT (Bidirectional Encoder Representations from Transformers): Uses transformers for contextualized word embeddings, considering both left and right context.

26 of 37

Playing With Embedding

27 of 37

Modelling

28 of 37

Naive Bayes Classifier

Bayes Theorem (Mathematics)

29 of 37

Naive Bayes Classifier

Email ID

Free

Win

Money

BUY

Spam ?

1

Yes

Yes

Yes

No

Spam

2

No

Yes

No

Yes

Not-Spam

3

Yes

No

Yes

No

Spam

4

No

No

No

Yes

Not-Spam

5

No

Yes

No

No

Spam

30 of 37

Naive Bayes Classifier

Email ID

Free

Win

Money

BUY

Spam ?

1

Yes

Yes

Yes

No

Spam

2

No

Yes

No

Yes

Not-Spam

3

Yes

No

Yes

No

Spam

4

No

No

No

Yes

Not-Spam

5

No

Yes

No

No

Spam

31 of 37

Naive Bayes Classifier

New Sentence. “Win Free Money”

32 of 37

Naive Bayes Classifier

New Sentence. “Win Free Money”

33 of 37

Naive Bayes Classifier

34 of 37

RNN (Recurrent Neural Network)

RNN (Recurrent Neural Network)

Meant to be worked with sequential data, for example textual data, time-series data, or speech, etc.

Why ?

-> We lose sequence information in ANN

-> Unnecessary padding is done to maintain the input size to the ANN model - Thus computation wasted

-> Bad at prediction for sequential information.

Input to RNN :-

Input to RNN is of the form, (timestamp, features)

e.g. Text - My name is Somil

Is of size (4,d) where d is the no. of dimension in the embedding vector.

35 of 37

RNN (Recurrent Neural Network)

36 of 37

RNN (Recurrent Neural Network)

New Sentence. “Everybody(x1) Loves(x2) Chocolate(x3)”

37 of 37

RNN (Recurrent Neural Network)

Problems:-

-> Forgets long context

-> Vanishing Gradients during

backpropagation

Variations:-