AI in Photography
Technology Snapshot 2024
The Human Neuron
The average human brain contains 86,000,000,000 neurons
Modelling the Neuron in Maths
Neural Networks
Oops!
AI was supposed to take the boring jobs and allow people to have time to be creative but the opposite has happened.
Boston Dynamics Spot
Such Recent History – This is Now!
From Wikipedia…
Starting in 2018, the OpenAI GPT series of decoder-only Transformers became state of the art in natural language generation.
In 2022, a chatbot based on GPT-3, ChatGPT, became unexpectedly popular, triggering a boom around large language models.
A.I. Concepts 1
CPU
Central Processing Unit
GPU
Graphics Processing Unit
NPU
Neural Processing Unit
TPU
Tensor Processing Unit (Google)
TOPS
Trillion operations per second; a metric used to measure and compare the computational power of AI chips. The minimum requirement for Microsoft Copilot+ PCs in 2024 is 40 TOPS.
A.I. Concepts 2
AI – Artificial Intelligence
The science of making machines that can think like humans.
AGI – Artificial General Intelligence
The stage at which humans can be replaced by machines.
ASI – Artificial Superintelligence
The stage at which machines dominate humans and humans may struggle to survive.
ML – Machine Learning
Machine learning is a subset of artificial intelligence that performs data analysis tasks without explicit instructions. Machine learning technology can process large quantities of historical data, identify patterns, and predict new relationships between previously unknown data.
DL – Deep Learning
Deep learning is a subset of machine learning that uses multilayered neural networks, called deep neural networks, to simulate the complex decision-making power of the human brain. It is a method in artificial intelligence that teaches computers to process data in a way that is inspired by the human brain. Deep learning models can recognize complex patterns in pictures, text, sounds, and other data to produce accurate insights and predictions.
NLP – Natural Language Processing
A machine learning technology that gives computers the ability to interpret, manipulate, and comprehend human language.
A.I. Concepts 3
Generative AI
Generative AI models use neural networks to identify the patterns and structures within existing data to generate new and original content.
LLM – Large Language Model
Large language models, also known as LLMs, are very large deep learning models that are pre-trained on vast amounts of data. The underlying transformer is a set of neural networks that consist of an encoder and a decoder with self-attention capabilities. The encoder and decoder extract meanings from a sequence of text and understand the relationships between words and phrases in it.
Transformer
Transformers are a type of neural network architecture that transforms or changes an input sequence into an output sequence. They do this by learning context and tracking relationships between sequence components.
GPT – Generative Pre-trained Transformer
GPT models are neural network-based language prediction models built on the Transformer architecture. They analyse natural language queries, known as prompts, and predict the best possible response based on their understanding of language.
The Transformer Paper 2017
The Transformer Model
MLP = Multi-Layer Perceptron
A.I. Concepts 4
RNN – Recurrent Neural Network
Unlike feedforward neural networks, which process data in a single pass, RNNs process data across multiple time steps. The network maintains a hidden state as a form of memory.
GAN – General Adversarial Network
These pit two neural networks against each other: a generator that generates new examples and a discriminator that learns to distinguish the generated content as either real (from the samples) or fake (generated).
CNN – Convolutional Neural Network
A neural network with a hidden layer that performs a dot product of the convolution kernel with the layer's input matrix to make image processing vastly more efficient.
VAE – Variational Autoencoders
Two neural networks comprising an encoder and decoder. The encoder and decoder work together to learn an efficient and simple latent data representation; allowing the user to easily sample new latent representations that can be mapped through the decoder to generate novel data.
Stable Diffusion
A text to image model; it works by teaching a neural network model to predict the noise added to an image in gradually increasing steps and then reversing the process so it can convert pure noise to a recognisable image. It uses a variational autoencoder to speed-up the processing and links to text prompts with a cross-attention mechanism.
A.I. Concepts 5
LoRA – Low-Rank Adaptation
A technique designed to refine and optimise large language models. Unlike traditional fine-tuning methods that require extensive retraining of the entire model, LoRA focuses on adapting only specific parts of the neural network. This approach allows for targeted improvements without the need for comprehensive retraining, which can be time-consuming and resource-intensive.
Flux
A new AI image generation model developed by Black Forest Labs. It represents a significant advancement in AI-generated art, utilizing a “hybrid architecture” that combines transformer and diffusion techniques, scaled up to 12 billion parameters.
Stable Diffusion – Forwards Direction
Stable Diffusion – Reverse Direction
Stable Diffusion and Text to Image
Google DeepMind
AI in Photography 1
Text to Image Generation
Many examples like Microsoft Image Creator, Open AI DALL.E and Midjourney
AI in Photography 2
Scene Recognition
Various camera manufacturers have an intelligent-Auto mode that can recognise landscapes, portraits, food, pets, etc.
AI in Photography 3
(Sony) Eye AF
Uses machine learning to recognise the eyes of humans and animals and maintain focus on them in real-time.
AI in Photography 4
Background Blur
Many examples using AI to recognise the background in an image and apply a Gaussian blur. This is equivalent to what a wide aperture or long focal length can do in conventional lens-based photography.
AI in Photography 5
(Google) Magic Eraser
Click on people/objects to remove them (with shadows) and fill-in the background with A.I generated scenery and textures.
AI in Photography 6
Generative Fill & Crop
Photoshop 2024 and ON1 Photo Raw 2025
AI in Photography 7
(Google) Magic Editor
Lets you move, remove, and resize people or objects in an image and change the background
AI in Photography 8
(Google) Best Take
Combines multiple similar photos into a single image where everyone looks their best i.e. not blinking, smiling, looking into camera
AI in Photography 9
(Google) Portrait Light
Uses AI to add and re-balance the light falling on a portrait
AI in Photography 10
(Google) Add Me
Allows the photographer of a group shot to be merged in from a 2nd shot
AI in Photography 11
(Microsoft) Restyle Image
Use a text prompt to create a new style for your selected photo
AI in Photography 12
(Google) Reimagine
Lets you select an object from an image ( including sky, water etc. ) and change it in any way using text prompts
AI in Photography 13
(Microsoft) Cocreator
Sketch and Prompt to Image Generation
AI in Photography 14
AI in Photography 15
Automatic Image Captioning
Image Alt Text originally for the visually impaired : Image 🡪 Text 🡪 Voice
AI in Photography 16
Image to Video Generation
State of the art. Provide a static image and directorial text prompts and receive a short video. See Minimax, Runway and Kling.
deepai.org/video
AI Image Analysis in Medicine
The Future of A.I.
Tesla Cybercab (Production 2026, Cost £23K)
Showcased on 10th October 2024
The Future of A.I.
Tesla Bot – Optimus (Production unknown)
Showcased on 10th October 2024
The Future of A.I.
The End