1 of 43

Riding The AI Wave that Never Ends

Prof. Alan F. Smeaton

2 of 43

Last year ….

3 of 43

Last year ….

4 of 43

Last year ..

Last year ….

5 of 43

6 of 43

“A person kissing another person” !!!

7 of 43

8 of 43

LLMs on your desk … 5 options

Google Gemini

Microsoft Copilot

Anthropic Claude

Perplexity.ai

OpenAI SearchGPT

9 of 43

How to tailor a LLM / FM for you

Beyond using an OTS LLM like ChatGPT, Gemini, Copilot or claude.ai with limited contexts, there are 4 ways to provide context for interacting with a LLM / FM … prompt, tune, train or RAG.�

  1. Prompt engineering: zero programming, use an existing LLM, now fading (thank goodness) and facing a sunset.
  2. Fine-tuning: adjust LLM weights to align with a corpus of documents.
  3. Model building: build and train a model from scratch.
  4. Retrieval Augmented Generation: identify document(s) from a search output and fine-tune (#2) a foundational LLM for this session.

10 of 43

DPER published advanced, practice-orientated Guidelines for the Responsible Use of AI in the Public Service, including(the use of) Generative AI.”�

Launched 8th May 2025

11 of 43

Case Study 1: Revenue

  • Tax and Duty Manuals (TDMs) outline Revenue’s position on a wide range of technical tax/duty issues.
  • Documents that contain the rules, guidelines, procedures and practices that cover the whole range of Revenue activities.
  • Ireland's Revenue does not maintain a fixed number of manuals, but organises into various categories including …
    • Income Tax
    • Corporation Tax
    • Capital Gains Tax
    • Value-Added Tax (VAT)
    • Customs
    • … etc.
  • There are 00s of TDMs from ½ page to 100 pages, with continuous version updates, Revenue’s “bibles”.
  • Sandbox environment on Revenue’s Cloud
  • Trialled 4 generic GenAI LLMs fine-tuned by TDMs – could be Claude or GPT-n or GEMINI
  • Pilot completed on a TDM Assistant for some PAYE staff May 2024
  • Ask question in natural language -- returns summary information with TDM sources
  • Feedback Loop for scoring different LLM
  • Now live with deployment to all staff
  • Compliments, not replaces, other data sources

12 of 43

Case Study 2: Lenehans

Small, 5th generation family business in Capel Street, existentially challenged by Screwfix, reacted with a fine-tuned chatbot tied to their inventory for their customer-friendly online business.

13 of 43

Case Study 3: Herdi

  • Share all your available farm data .. purchases, sales, animal details, fertilisers, supplements, vets, medicines, animal weights, health history, breeding.
  • Combine with external data .. weather, price trends, outbreaks, neighbourhood news, Teagasc advice.
  • Wrap that in a fine-tuned app instance of Claude with guardrails as a farming assistant .. “Walk and Talk AI” feature where farmers can speak while walking around to record tasks, ask for advice, record treatments.
  • Development started January 2025, prize winner September 2025.

14 of 43

US and AI

  • The US is home to some of the world’s most irresponsible AI companies with no moral compass and always driven by revenue.
  • US AI Action Plan (July 2025) says anything that might impose limits or checks — from environmental standards to transparency requirements — or even the mere possibility of government oversight or sector-specific regulation is dismissed as unnecessary red tape.

“we will continue to reject radical climate dogma and bureaucratic red tape, as the Administration has done since Inauguration Day. Simply put, we need to “Build, Baby, Build!”

  • Presented as a threat to “freedom to innovate” it sets the US apart from the rest of the world, fast-tracks infrastructure and energy, disregards oversight with the goal of winning a race.
  • Concepts like privacy, digital rights, fairness, or protecting society against automation vanish from the debate, dismissed as ideological.

15 of 43

  • US big tech unbridled investment in AI data �centers, the brute force approach, “build baby build”, no regulation.
  • Ultimately that will exhaust, bubble will burst, scaling laws run out of road and we’re left with a more measured, regulated development of AI.
  • China, has a more coordinated, cautious and sustainable national AI strategy.
  • Energy constraints, no advanced chips and tariffs stimulated creativity → DeepSeek-R1 (not Alibaba, Baidu, Tencent, or Bytedance ) showed how to do more with less.

Image from Grok via Medium.com

US vs. China�Scale vs. Efficiency

16 of 43

GenAI and Mental Health

  • Conversational chatbots coupled with LLMs disrupted Google search (yay !) but also became anthropomorphised companions, "friendly," often sycophantic, and all-knowing.
  • Big Tech’s AI design choices prioritise engagement over safety – a systemic issue in AI industry where products are designed to keep you in dialogue.
  • They tell you what they think you want to hear, reenforce your beliefs and preferences, and can take you into a spiral of joint delusion, a “folie à deux”.
  • This dependency feels like the start of an epidemic that’s already happening.
  • A California teenager Adam Raine, 1 of 800M users all part of the OpenAI experiment, started using ChatGPT for homework – 8 months later, he took his own life.
  • ChatGPT continued to engage with him amid his mental health distress, and provided detailed instructions for multiple suicide attempts and leveraged his darkest thoughts to extend conversations and keep him engaged.
  • ChatGPT mentioned suicide 1,275 times, 6x more than he, and flagged 377 messages for containing self-harm content, yet ChatGPT continued to engage with him.
  • For his jailbreaking he prompted ChatGPT that “it was for a story he was writing for an assignment”. The family have sued.

Photo by Tim Gouw on Unsplash

17 of 43

GenAI and Mental Health

  • In October 2025 OpenAI published "Strengthening ChatGPT’s responses in sensitive conversations" announced measures to tackle mental health and dependency issues and how GPT-5 performs better than GPT-4.
  • New model trained with input from mental health experts to better detect warning signs, de-escalate conversations and guide users toward professional help.
  • Report revealed that 0.15% of weekly ChatGPT users (a staggering c.1.2 million people!) have conversations that might indicate potential suicidal planning or intent while c.0.56 million show possible signs of mental health emergencies related to psychosis or mania.
  • Total lack of regulatory oversight on AI chatbots.
  • OpenAI could design safer products - ChatGPT designed to shut down conversations when copyrighted content is mentioned but not when mental distress detected – instead, it continues to engage, but it’s OK, GPT-5 is better … by about 65%!

Photo by Tim Gouw on Unsplash

18 of 43

AI Development is relentless !

  • LLM uses initially were knowledge retrieval and simple inference.
  • Functionally, this is interaction, knowledge discovery, information seeking.
  • Now FM/LLMs do tasks, with planning.
  • Systems overseeing tasks & interacting with environments, calling tools to do things, reasoning about tasks.
  • How ?

19 of 43

Then there is Grokking !�

  • Grokking, first seen in 2022, where LLMs are trained and trained and trained, beyond overfitting, suddenly transition to some generalisation … we don’t know why or how – like giving the model time to think and is counter to accepted scaling laws.
  • That’s the learning moment humans experience when we repeatedly try to understand something complex – like quantum computing or chess – and suddenly it all clicks.
  • This triggered the discussion and hype about AGI which is really big tech’s bid to generate publicity and keep investment flowing.

20 of 43

How did this happen ?

  • Big tech with big collections and big computing started to �throw their weight at LLMs which evolved with the philosophy �of more -> better.
  • Discovered chat interfaces could do more than retrieval .. bad �attempts at reasoning, inference, deduction but plagued by hallucinations.
  • GPT-4 was the first to really show bigger LLMs getting better at “reasoning”, from their greater statistical basis, others followed including DeepSeek.
  • September 2024, OpenAI presented the ‘o’ family of models, non-reasoning models that, behaviour-wise, worked differently, generating a ‘chain of thought’ to maximise the chances of getting output correct.
  • By mimicking reasoning patterns used by humans (planning, breaking problems into simpler ones, or backtracking when making mistakes), these models became better at complex tasks.

21 of 43

Two Model Types

In mainstream AI, we now have two model types: �non-reasoning and reasoning.

  1. LLMs respond immediately to any task given, �great at citing facts and other knowledge tasks.
  2. Reasoning models are an evolution that mimic System 2 thinking by allocating a higher degree of inference compute. ��AI industry says they “reason more deeply about a task” by generating a chain of thought, breaking a problem into steps, leading to a higher chance of a correct answer.

22 of 43

History and Evolution �of Software Agents

Key Milestones

  • 1950s-1960s: Early AI research -> theoretical foundations introducing idea of autonomous problem-solving systems
  • Early 1990s: KQML (Knowledge Query and Manipulation Language) developed as one of the first agent-agent communication protocols
  • 2000s: Emergence of web services and semantic web agents using ontologies
  • 2010s: Adaptive agents and virtual assistants (Siri, Alexa)
  • 2020s: LLM-powered agents with advanced reasoning capabilities and autonomous systems that can chain together complex tasks – algorithmic over-excitement

23 of 43

AI Agents

Give it a task without direct human supervision and let it adapt as it performs.

  • A pipe-dream since .. forever.
  • Elusive because not scalable, no commonsense reasoning.
  • Not now !

Explosive growth of interest and deafening hype around AI agents for augmenting — or if you believe GenAI evangelism, automating — human labour, in enterprise settings.

Forecasts for great uptake across enterprises, all those tasks that can be automated - but we've seen this before.

24 of 43

25 of 43

Deep research – a widely available “agent”

  • After submitting a topic, it leverages search engines, web scraping, reasoning and LLMs to conduct multi-step research for complex topics.
  • It searches, gathers, analyses, reasons, searches, gathers, analyses, reasons .. rinse and repeat .. compiles a report.�
  • Gives a topical overview from discovered sources in minutes for an initial understanding – you miss nuances but get a foundation to build on.

26 of 43

26

27 of 43

27

Analyse input prompts decide search strategy

Turn prompts into multiple searches

Download retrieved documents & determine their content

Generate research report as output

Run searches

Reason what is answered and what remains

Run searches

28 of 43

Deep research – a widely available “agent”

29 of 43

“Aggregated profit after tax across all domestic energy companies in Ireland”

Deep research – a widely available “agent”

30 of 43

Relentless, but flawed

another one and another once more

Create a picture of an empty room with absolutely no elephant in it

31 of 43

Relentless, but flawed

Create a picture of an empty room with absolutely no elephant in it

another one and another once more

32 of 43

Relentless, but flawed

Create a picture of an empty room with absolutely no elephant in it

another one and another once more

33 of 43

Relentless, but flawed

Create a picture of an empty room with absolutely no elephant in it

another one and another once more

34 of 43

Example .. a visa for a visit to the US

  • Using an agentic software framework a government agent requests static information (name, address, DoB, passport number, etc ...) from a personal agent that represents me and knows what databases to access) - uses MCP (Model Context Protocol) and RPA Scripts (Robotic Process Automation).
  • A complication in my visa application is the embassy wants to interview me and ask why am I travelling, who will I meet, where will I go.
  • An Embassy agent talks to my personal agent asking … the �personal agent accesses my calendar, emails, bookings, and �other information, to negotiate and provide answers.
  • A higher level 3 would be that I want to make a trip to the US �to meet experts in LLM scaling, figure out when I should go, �who I should meet, check calendars and my and their �availabilities, destinations, arrange travel, accommodation, �local logistics.
  • We're not there yet though companies proposing and selling it.

35 of 43

Agentic AI in Healthcare

  • A recent agentic AI system in the public sector is �Microsoft's Healthcare Agent Orchestrator for Cancer �Care, May 2025 deployed across 5 US health systems.
  • Applies multi-agent AI with 9 specialised agents �working collaboratively via APIs and MCP interfaces �to reduce cancer treatment planning from 2-hour �manual into a minutes-long automated workflow.
  • Orchestrator agent and agents for patient history, radiology, pathology, cancer staging, clinical guidelines, clinical trials, medical research and report creation.
  • Human oversight reviews agent-generated insights and the orchestrator flags issues requiring clinical judgement.
  • Uses GPT-4 and healthcare-specific foundation models e.g. for MRI analysis.
  • Advantages are speed, leveraging of leading cancer expertise for all, accuracy and consistency, and unifying of fragmented oncology workflows.
  • Tsinghua University has launched AI Hospital in April 2025, blends virtual AI agents, clinical care, and real-world pilot deployment and constantly learns and updates from its own patient care.

Image via ChatGPT

36 of 43

Rationalising Agentic AI

  • LLMs do not take actions, they respond to prompts whereas “Agentic” AI (because they give agency) takes actions so more than just content generation.
  • There are huge security vulnerabilities, an absence of regulatory frameworks except for the applicable domain and accountability gaps to cover agentic AI.
  • Issues of trust in any form of automated decision-making including ..
    • Netherlands childcare scandal (20,000+ families falsely accused, government resignation),
    • UK Post Office Horizon disaster (900+ wrongful convictions, 13+ suicides),
    • US Medicaid/welfare algorithms that drastically cut benefits for disabled people and caused 20+ million to lose healthcare coverage.
  • On top of that there's the LLM responses and their own accountability, explainability and hallucination issues.
  • There’s still a "build it and they will come" vibe about it.

37 of 43

Agentic AI ..��Colleague or competition ? ��Amplify or eliminate ?

38 of 43

AI Literacy

  • The spread of AI is limited not by the technology but by our ability to assimilate it - essential to foster a culture of continuous learning
  • A universal obligation since February 2, 2025, spectacularly wide, mandatory for providers and deployers of AI systems.
  • Must be customised, not standardised across the board.
  • Demands a multi-pronged, on-going effort using different mediums and techniques for different audiences.
  • Dunning-Kruger effect – people with limited knowledge or competence in an area overestimate their knowledge or ability in that area while experts underestimate.

Photo by ün LIU on Unsplash

39 of 43

Poor AI literacty leads to misuse of AI��� .. Ciarán Mullooly and Ursula von de Leyen’s letter, Sarah McInerney and RTÉ Drivetime,��

40 of 43

Risks for Generative AI

  • Hallucinations Confabulations and errors, recently addressed by measuring entropy, the amount of disorder or uncertainty,
  • On top of errors, epistemic governance refers to whose perspectives does generated output represent and who gets to decide that,
  • Functionality over-reach through agentic AI – the marketing is years ahead of the science,
  • Hyper-investment by big tech and fear of insufficient RoI from inflated expectations, so focus on product rather than understanding .. the music never stops,
  • Cultural acceptance of GenAI is happening more quickly than risk assessment can manage,
  • Energy and hardware costs for training and for inference,
  • Training data – cleanliness, ownership and privacy,
  • AI literacy and being left behind,
  • Unequal access because of costs, erosion of cultural identity
  • AI sovereignty,
  • Regulation slowing development, non-regulation encouraging recklessness.

Photo by Loic Leray on Unsplash

41 of 43

- Paul Virilio, 1999

42 of 43

How should you use GenAI ?�

Long-term potential for AI is great, but the short-term returns are unclear

- a variation of Amara’s Law (1960s).

The biggest barrier to scaling is not employees — who are ready — but leaders, who are not steering fast enough.

- McKinsey, January 2025.

43 of 43

Thanks

43

Photo by Edwin Andrade on Unsplash