1 of 28

UNIT – 3

​

​

  • Generation of Images: Introduction to Generative Adversarial Networks,
  • Adversarial Training Process,
  • Nash Equilibrium,
  • Variational Autoencoders,
  • Encoder-Decoder Architectures,
  • Stable Diffusion Models,
  • Introduction to Transformer-based Image Generation,
  • CLIP,
  • Visual Transformers ViT- Dall-E2 and DallE3,
  • GPT-4V,
  • Issues of Image Generation models like Mode Collapse and Stability.

2 of 28

  • Generative Adversarial Networks (GANs) revolutionized AI image generation by creating realistic and high-quality images from random noise.
  • In this article, we will train a GAN model on the MNIST dataset to generate handwritten digit images.

​

A GAN consists of two neural networks that compete with each other:

  1. Generator (G): Creates fake images from random noise.
  2. Discriminator (D): Determines whether an image is real (from the dataset) or fake (created by the generator).

​

Generation of Images: Introduction to Generative Adversarial Networks

3 of 28

4 of 28

5 of 28

6 of 28

2. Working of GAN

Random Noise → Generator → Fake Image → Discriminator → Real/Fake

​

​

The process is:

  • Random noise is given to the Generator.
  • The Generator produces a new image.
  • The Discriminator receives both real images and generated images.
  • It predicts whether each image is real or fake.
  • The Generator improves its ability to create realistic images.
  • The Discriminator improves its ability to detect fake images.
  • Through repeated training, the Generator learns to produce highly realistic images.

7 of 28

ADVERSARIAL TRAINING PROCESS IN GENERATIVE AI

Adversarial training is a learning process in which two neural networks compete with each other to improve the quality of generated data. It is most commonly used in Generative Adversarial Networks (GANs).

Main Components

Generator (G)

    • Creates synthetic data such as images, text, or audio.
    • Takes random noise as input.
    • Tries to generate samples that look like real data.

Discriminator (D)

    • Acts as a classifier.
    • Receives both real samples from the dataset and fake samples produced by the Generator.
    • Predicts whether each sample is Real or Fake.

​

8 of 28

Training Process

​

Step 1: Generate Fake Data�Random noise is given to the Generator.

Step 2: Discriminator Evaluation�The Discriminator receives:

  • Real data → from the training dataset
  • Fake data → from the Generator

Step 3: Calculate Errors�The Discriminator learns to correctly distinguish real samples from generated samples.

Step 4: Update the Generator�The Generator is trained to produce increasingly realistic samples that can fool the Discriminator.

Step 5: Repeat the Competition�Generator and Discriminator are trained repeatedly in an adversarial manner.

Step 6: Convergence�Ideally, the Generator becomes capable of producing highly realistic data, while the Discriminator can no longer reliably distinguish fake data from real data.

9 of 28

NASH EQUILIBRIUM

  • Nash equilibrium is an important concept in game theory that provides the optimal outcome in case the player doesn’t deviate from their initial strategy.
  • This is done in response to no incentive provided to the players for such deviation.
  • This was named after the Mathematician, John Nash, defining the solution of a non-cooperative game involving two or more players.

​

10 of 28

Key Points:

​

  • Nash equilibrium provides an optimal solution to encounter the required outcome by not deviating from their initial strategy.
  • Since individuals are already aware of each other’s strategy, both the players win as everyone gets the outcome that they thought for.
  • One such example is the Prisoner's dilemma.
  • Since the players quest to win, they would be performing that strategy that would lead them to such a state. Thus the chosen strategy is the best and the optimal solution that they can use. This is also in conjunction with the dominant strategy.
  • Further, as stated there cannot be any Nash equilibrium in a game. So, it is not always true that the strategy chosen is the optimal one.

11 of 28

Nash equilibrium vs Dominant strategy

  • Nash equilibrium reveals the optimal solution to be one when the player doesn’t deviate from their initial strategy and when he is aware of other players' strategies as well.
  • None of the players will be changing their strategy and will be keeping it the same.

Dominant strategy

  • on the other hand, makes sure that the chosen strategy will lead to better results out of all the possible strategies available.
  • It doesn’t conform to the strategy the opponent is supposed to use.
  • Though both seem quite similar, there is a grave difference between them.
  • Nash equilibrium connotes that none of the players will win if one of the users deviates from its strategy, keeping the other players constant, though deviating from its initial strategy, as per dominant strategy, will lead to win of the player, as this being his best solution.

12 of 28

Example of Nash equilibrium

In the given game, we have two players with strategies S1 and S2.

S1 being the optimal strategy considering both the players know each other’s approach or they are made aware of it and both will be moving their initial strategy. They will opt for it because they know deviating from it will not be in the interest of the match.

Hence in case they opt for S1, both will win. In other cases, deviating to S2 will result in defeat of one of them.

13 of 28

Variational Autoencoders

A Variational Autoencoder (VAE) is a Generative AI model that learns the hidden/latent representation of data and uses it to generate new, realistic samples similar to the training data.

  • Learns a continuous latent representation
  • Enables controlled and meaningful data generation
  • Widely used in image synthesis, anomaly detection, and representation learning

14 of 28

15 of 28

1. Encoder (Understanding the Input)

The encoder takes input data like images or text and learns its key features. Instead of outputting one fixed value, it produces two vectors for each feature:

  • Mean (μ): A central value representing the data.
  • Standard Deviation (σ): It is a measure of how much the values can vary.

These two values define a range of possibilities instead of a single number.

2. Latent Space (Adding Some Randomness)

  • Fixed point: Input is not encoded as one fixed value.
  • Random point: A random value is selected based on Mean and Standard Deviation.
  • Introduces randomness: This creates small variations in the data.
  • Generates new data: The model can produce different versions of the original data.
  • Realistic samples: Helps generate new and realistic data samples.

16 of 28

3. Decoder (Reconstructing or Creating New Data)

The decoder takes the random sample from the latent space and tries to reconstruct the original input. Since the encoder gives a range, the decoder can produce new data that is similar but not identical to what it has seen.

17 of 28

Encoder-Decoder Architectures

The encoder-decoder model is a neural network used for tasks where both input and output are sequences, often of different lengths. It is commonly applied in areas like translation, summarization and speech processing.

  • The encoder processes the input sequence and converts it into a fixed representation (context vector)
  • The decoder uses this representation to generate the output sequence step by step
  • Works well for tasks where input and output lengths are different

18 of 28

19 of 28

Encoder

The encoder processes the input sequence and converts it into a fixed representation (context vector) using an RNN or LSTM.

  • Processes input tokens sequentially and updates hidden states
  • Captures relationships between words in the sequence
  • Produces final hidden and cell states forming the context vector

20 of 28

21 of 28

Decoder

The decoder uses the context vector from the encoder to generate the output sequence step by step.

  • Takes previous output and context to predict the next token
  • Generates output sequentially until an end token is reached
  • Initializes its states using the encoder’s final states

22 of 28

Working of Encoder Decoder Model

​

The actual working of the encoder decoder model is shown in below diagram. Now we will understand it stepwise

23 of 28

Step 1: Tokenizing the Input Sentence

  • The sentence "I am learning AI" is first broken into tokens: ["I", "am", "learning", "AI"].
  • Each word (token) is converted into a vector that a machine can understand. This process is called embedding.

Step 2: Encoding the Input

  • The encoder processes these embeddings sequentially using an LSTM network.
  • At each step, it updates its hidden state based on the current word and previous context. This helps the model understand the sequence order and relationships between words.
  • After processing the full sentence, the encoder generates a context vector (final hidden and cell states), which represents the meaning of the entire input sentence.

24 of 28

Step 3: Passing the Context to the Decoder

  • The Context Vector is passed to the Decoder as shown in image.
  • It acts like a summary of the full input sentence.

Step 4: Decoder Generates Output Step-by-Step

  • The Decoder uses the context and starts creating the output one word at a time.
  • First it predicts the first word then uses that to predict the second word and so on.

25 of 28

Step 5: Attention Mechanism

  • Basic encoder-decoder uses a single context vector, which can limit performance for long sequences.
  • Attention mechanism helps the decoder focus on different parts of the input at each step.
  • Improves accuracy by not relying only on one fixed representation.

Step 6: Producing the Final Output

  • The decoder continues generating until the full translated sentence is produced.
  • Each output token depends on the previous ones and the input context. You finally see the output tokens generated on the right side of the diagram completing the translation.

26 of 28

Stable Diffusion is a latent diffusion-based Generative AI model used to generate images from text prompts. It creates an image by starting with random noise and gradually removing the noise until a meaningful image is produced.

Stable Diffusion Models

27 of 28

28 of 28