1 of 29

Everything I Need to Know About AI/LLM, I Learned From a Pasadena Search Company 24 Years Ago

By Gene Chuang �Southern California Linux Expo

03/09/2025

2 of 29

Pasadena Sighting

3 of 29

Francis Arnold and John Hopfield (Geoff Hinton)

4 of 29

Gene’s Background

5 of 29

Search Background

  • GoTo.com was an Idealab company, foundering engineers are from Caltech
  • Joined GoTo.com in 2001 - $500M in revenue
  • We became Overture in 2002, became pure Sponsored Search
  • Acquired by Yahoo! In 2004 for $1.7B, was doing $2.0B rev at end of 2004, 2 Billion searches a day
  • Google AdWord is a copy of Overture launched in 2002 - doing $2B in revenue by 2004
  • Project Panama 2005-2007 - rewrite of Overture (mod perl, J2EE, Oracle on RHEL to C++, Tomcat/jBoss, MySQL on FreeBSD). 30% revenue uplift to $3B. Google in that time 10x to $20B
  • Yahoo! Gave up on search, outsourced to Microsoft�

6 of 29

This is my visit of AT&T Labs in 2010. Information Theory by Claude Shannon is the mathematical study of the quantification, storage, and communication of information. It is at the intersection of electronic engineering, mathematics, statistics, computer science, neurobiology, physics, and electrical engineering. ��I also sat in the office of Dennis Ritchie(!)

7 of 29

Information retrieval (IR) is the process of finding and presenting relevant information from a collection of data. IR systems are the interface between users and large data repositories, such as databases, the internet, and digital libraries. They use a variety of techniques and algorithms to help users find information that matches their search query.

�IR systems have advanced over time, from manual cataloging and archiving to AI-powered search technologies. They use sophisticated algorithms to filter out noise and present users with the most relevant information.

8 of 29

Search vs LLM - UX

Guess user intent from search box of keywords, retrieve links to documents websites and list in order of relevance

Guess user intent from chat dialog, retrieve documents, merge and rewrite in plain English

9 of 29

Search vs LLM - Algorithms

  • Naive Bayes
  • k-Nearest Neighbor
  • Clustering
  • Cosine Similarity
  • SVM
  • Markov Chain
  • NLP for Exact, Phrase, Orthogonal Match
  • CPU
  • Deep Neural Network
  • Transformer
  • Retrieval Augmented Generation
  • Vectors and embedding
  • GPU

10 of 29

Search vs LLM - Technology

  • Apache Nutch
  • Lucene
  • Hadoop
  • Apache Mahout
  • Handspun algo in R
  • Oracle DB
  • BerkeleyDB
  • Mod Perl
  • RegEx

  • PyTorch/Pandas
  • Tensorflow
  • Nvidia TensorRT
  • Amazon Rekognize
  • Snowflake
  • Spark
  • DataStax
  • Pinecone
  • Llama 3.1

11 of 29

Search vs LLM - Ingestion and Transformation

Crawling the entire public website and generating index for information retrieval

Crawling the entire public website and generating Large Language Model multi-level graph weighed by probability for answering.

12 of 29

Search vs LLM - Size Matters(?)

Size of Index or Number of pages crawled:��Google 20B Pages Crawled!�Bing 18B Pages Crawled!

Large Language Model Size

13 of 29

Search vs LLM - Talent War

At peak of Search War in 2008, Google, MSN and Yahoo! Were poaching each others’ talent, $500K+ comp packages

At current AI/LLM War in 2024, Google, Meta, Microsoft and well funded AI startups are offering $1M+ comp packages

14 of 29

Search vs LLM - Hit Missed!

Precision vs Recall

Hallucination

15 of 29

Search vs LLM - Agentic AI

Cron Job

/etc/crontab

3 * * * * root python /user/me/ai.py

Agentic AI

16 of 29

Can LLM Really Reason?

  • Yann LeCun Chief Scientist at Meta thinks AI is Dumber Than a Cat

17 of 29

“I am not interested anymore in LLMs. They are just token generators and those are limited because tokens are in discrete space. I am more interested in next-gen model architectures, that should be able to do 4 things: understand physical world, have persistent memory and ultimately be more capable to plan and reason.”

Yann LeCun at Nvidia GTC 2025

18 of 29

Can LLM Really Reason?

  • Apple Paper: Understanding the Limitations of Mathematical Reasoning in Large Language Models
  • Terence Tao UCLA Math Professor:
    • OpenAI o1 is a mediocre or incompetent research assistant.
    • It does not “reason”, is not a source of knowledge or ideas, but is a really useful glue
    • We will always need humans and AI. They have complementary strengths. AI is very good at converting billions of pieces of data into one good answer. Humans are good at taking 10 observations and making really inspired guesses.�

19 of 29

20 of 29

21 of 29

22 of 29

LLM is a Powerful Tool and Can Disrupt

23 of 29

Economics of LLM

  • $100B to build the next AI Factory Data center - Microsoft and OpenAI planning to build one
  • Wall Street frenzy creates $11bn debt market for AI groups buying Nvidia chips
  • OpenAI, Google and Anthropic are seeing diminishing returns of new models
  • ChatGPT Agent $20K/mo
  • Search Monetization is arbitrage. LLM Monetization will be arbitrage
  • Consolidation of AI companies is inevitable. Like Search

24 of 29

25 of 29

26 of 29

Open Source Model is changing the game

  • DeepSeek Benchmark https://arxiv.org/html/2412.19437v1
  • Open Source LLM will eliminate all close weight LLMs that cannot outperform
  • Yann Lecun: “open source models are surpassing proprietary ones”

27 of 29

Lessons Learned in Search That Apply to LLM

  • LLM is a power tool that interfaces with systems we are used to. It is not AGI
  • Data Cleanliness/Quality is King: Garbage In Garbage Out
  • Search is good at joining indexed data. LLM is good at joining unstructured data.
  • It’s just pattern matching and data interpolation. LLM cannot extrapolate
  • Agentic AI = cron + awk|sed|regex + powerful GPU cluster

28 of 29

What’s Next for AI?

  • Multi Modal LLM is the next Killer App
  • Edge Inference and Feedback loop
  • Spatial Intelligence -> Robotics
  • Fei-Fei Li’s World Labs - PhD Caltech, ImageNet and Caltech 101
  • in vitro -> in vivo -> in silico - Simulation
  • Computing Causality - Caltech paper
  • Satya Nadella: “Us self-claiming some AGI milestone, that’s just nonsensical benchmark hacking to me.” Instead “The real benchmark is the world growing at 10%”
  • When is Blackwell H200 coming out Jensen?

29 of 29

Q&A

genechuang@gmail.com

LinkedIn: Gene Chuang�Substack: genechuang.substack.com

Coda: My SCaLE 22x talk 3/9/25 recorded: https://www.youtube.com/live/TP_TUAFu7JU?t=6228s