1 of 37

AI in Photography

Technology Snapshot 2024

2 of 37

The Human Neuron

The average human brain contains 86,000,000,000 neurons

3 of 37

Modelling the Neuron in Maths

4 of 37

Neural Networks

5 of 37

6 of 37

Oops!

AI was supposed to take the boring jobs and allow people to have time to be creative but the opposite has happened.

Boston Dynamics Spot

7 of 37

Such Recent History – This is Now!

From Wikipedia…

​

Starting in 2018, the OpenAI GPT series of decoder-only Transformers became state of the art in natural language generation.

​

In 2022, a chatbot based on GPT-3, ChatGPT, became unexpectedly popular, triggering a boom around large language models.

8 of 37

A.I. Concepts 1

CPU

Central Processing Unit

​

GPU

Graphics Processing Unit

​

NPU

Neural Processing Unit

​

TPU

Tensor Processing Unit (Google)

​

​

​

​

​

​

TOPS

Trillion operations per second; a metric used to measure and compare the computational power of AI chips. The minimum requirement for Microsoft Copilot+ PCs in 2024 is 40 TOPS.

9 of 37

A.I. Concepts 2

AI – Artificial Intelligence

The science of making machines that can think like humans.

AGI – Artificial General Intelligence

The stage at which humans can be replaced by machines.

ASI – Artificial Superintelligence

The stage at which machines dominate humans and humans may struggle to survive.

​

ML – Machine Learning

Machine learning is a subset of artificial intelligence that performs data analysis tasks without explicit instructions. Machine learning technology can process large quantities of historical data, identify patterns, and predict new relationships between previously unknown data.

​

​

DL – Deep Learning

Deep learning is a subset of machine learning that uses multilayered neural networks, called deep neural networks, to simulate the complex decision-making power of the human brain. It is a method in artificial intelligence that teaches computers to process data in a way that is inspired by the human brain. Deep learning models can recognize complex patterns in pictures, text, sounds, and other data to produce accurate insights and predictions.

​

​

NLP – Natural Language Processing

A machine learning technology that gives computers the ability to interpret, manipulate, and comprehend human language.

​

​

10 of 37

A.I. Concepts 3

Generative AI

Generative AI models use neural networks to identify the patterns and structures within existing data to generate new and original content.

​

​

LLM – Large Language Model

Large language models, also known as LLMs, are very large deep learning models that are pre-trained on vast amounts of data. The underlying transformer is a set of neural networks that consist of an encoder and a decoder with self-attention capabilities. The encoder and decoder extract meanings from a sequence of text and understand the relationships between words and phrases in it.

​

​

Transformer

Transformers are a type of neural network architecture that transforms or changes an input sequence into an output sequence. They do this by learning context and tracking relationships between sequence components.

​

​

GPT – Generative Pre-trained Transformer

GPT models are neural network-based language prediction models built on the Transformer architecture. They analyse natural language queries, known as prompts, and predict the best possible response based on their understanding of language.

​

11 of 37

The Transformer Paper 2017

12 of 37

The Transformer Model

MLP = Multi-Layer Perceptron

13 of 37

A.I. Concepts 4

RNN – Recurrent Neural Network

Unlike feedforward neural networks, which process data in a single pass, RNNs process data across multiple time steps. The network maintains a hidden state as a form of memory.

​

GAN – General Adversarial Network

These pit two neural networks against each other: a generator that generates new examples and a discriminator that learns to distinguish the generated content as either real (from the samples) or fake (generated).

​

CNN – Convolutional Neural Network

A neural network with a hidden layer that performs a dot product of the convolution kernel with the layer's input matrix to make image processing vastly more efficient.

​

VAE – Variational Autoencoders

Two neural networks comprising an encoder and decoder. The encoder and decoder work together to learn an efficient and simple latent data representation; allowing the user to easily sample new latent representations that can be mapped through the decoder to generate novel data.

​

Stable Diffusion

A text to image model; it works by teaching a neural network model to predict the noise added to an image in gradually increasing steps and then reversing the process so it can convert pure noise to a recognisable image. It uses a variational autoencoder to speed-up the processing and links to text prompts with a cross-attention mechanism.

​

14 of 37

A.I. Concepts 5

LoRA – Low-Rank Adaptation

A technique designed to refine and optimise large language models. Unlike traditional fine-tuning methods that require extensive retraining of the entire model, LoRA focuses on adapting only specific parts of the neural network. This approach allows for targeted improvements without the need for comprehensive retraining, which can be time-consuming and resource-intensive.

​

Flux

A new AI image generation model developed by Black Forest Labs. It represents a significant advancement in AI-generated art, utilizing a “hybrid architecture” that combines transformer and diffusion techniques, scaled up to 12 billion parameters.

​

15 of 37

Stable Diffusion – Forwards Direction

16 of 37

Stable Diffusion – Reverse Direction

17 of 37

Stable Diffusion and Text to Image

Google DeepMind

18 of 37

AI in Photography 1

Text to Image Generation

Many examples like Microsoft Image Creator, Open AI DALL.E and Midjourney

​

​

​

​

​

​

​

19 of 37

AI in Photography 2

Scene Recognition

Various camera manufacturers have an intelligent-Auto mode that can recognise landscapes, portraits, food, pets, etc.

20 of 37

AI in Photography 3

(Sony) Eye AF

Uses machine learning to recognise the eyes of humans and animals and maintain focus on them in real-time.

21 of 37

AI in Photography 4

Background Blur

Many examples using AI to recognise the background in an image and apply a Gaussian blur. This is equivalent to what a wide aperture or long focal length can do in conventional lens-based photography.

​

22 of 37

AI in Photography 5

(Google) Magic Eraser

Click on people/objects to remove them (with shadows) and fill-in the background with A.I generated scenery and textures.

​

23 of 37

AI in Photography 6

Generative Fill & Crop

Photoshop 2024 and ON1 Photo Raw 2025

​

24 of 37

AI in Photography 7

(Google) Magic Editor

Lets you move, remove, and resize people or objects in an image and change the background

​

25 of 37

AI in Photography 8

(Google) Best Take

Combines multiple similar photos into a single image where everyone looks their best i.e. not blinking, smiling, looking into camera

​

26 of 37

AI in Photography 9

(Google) Portrait Light

Uses AI to add and re-balance the light falling on a portrait

​

27 of 37

AI in Photography 10

(Google) Add Me

Allows the photographer of a group shot to be merged in from a 2nd shot

​

28 of 37

AI in Photography 11

(Microsoft) Restyle Image

Use a text prompt to create a new style for your selected photo

​

29 of 37

AI in Photography 12

(Google) Reimagine

Lets you select an object from an image ( including sky, water etc. ) and change it in any way using text prompts

​

30 of 37

AI in Photography 13

(Microsoft) Cocreator

Sketch and Prompt to Image Generation

​

31 of 37

AI in Photography 14

Photo Unblur

Use AI to regenerate lost detail e.g. https://www.artguru.ai/unblur-image/

​

32 of 37

AI in Photography 15

Automatic Image Captioning

​

Image Alt Text originally for the visually impaired : Image 🡪 Text 🡪 Voice

​

33 of 37

AI in Photography 16

Image to Video Generation

State of the art. Provide a static image and directorial text prompts and receive a short video. See Minimax, Runway and Kling.

​

deepai.org/video

34 of 37

AI Image Analysis in Medicine

The Future of A.I.

35 of 37

Tesla Cybercab (Production 2026, Cost £23K)

Showcased on 10th October 2024

The Future of A.I.

36 of 37

Tesla Bot – Optimus (Production unknown)

Showcased on 10th October 2024

The Future of A.I.

37 of 37

The End