1 of 9

1

AI-POWERED

Document Insights

and Data Extraction

Pharmaceutical document intelligence with OCR, retrieval and grounded answers

John Raymond TYABA

Email: u2678314@uel.ac.uk

1

OCR + CLEAN

2

CHUNK + TAG

3

RETRIEVE

4

GENERATE

5

CITE

AI-Powered Document Insights and Data Extraction

2 of 9

2

Problem Statement | Pharmaceutical Document Intelligence

Extracting trustworthy answers from varied, unstructured pharmaceutical PDFs

CHALLENGE

Digital PDFs, scans, certificates and packaging specifications vary in structure and quality. Users must manually search pages, interpret document types and verify where each answer came from.

SOLUTION

The pipeline extracts or OCRs text, creates metadata-tagged chunks, retrieves relevant evidence and returns a source-grounded answer through Gradio.

PAIN POINT

HOW THE PIPELINE SOLVES IT

Low-text scans

Tesseract OCR fallback

Mixed document types

Metadata classification and filtering

Long unstructured content

Overlapping page-level chunks

Low trust in answers

Sources, pages, scores and confidence

AI-Powered Document Insights and Data Extraction

3 of 9

3

System Architecture | Resilient Open-Source RAG Pipeline

Document Input → OCR → Chunking → Embeddings → Vector Search → Retrieval → LLM / Extractive Answer

INPUT

PDF

OCR

Tesseract

CHUNK

900 chars

150 overlap

EMBED

MiniLM

INDEX

FAISS

RETRIEVE

Top-k 4

ANSWER

FLAN-T5

Component Specifications

Component

Technology choice

Configuration

OCR

Tesseract + PDF2Image

220 DPI; sparse-page trigger

Chunking

Fixed page chunks

900 characters; 150 overlap

Embeddings

all-MiniLM-L6-v2

Normalised dense vectors

Vector index

FAISS IndexFlatIP

Inner-product similarity; in memory

Retriever

Metadata-filtered dense search

Top-k = 4; manual filter + auto-route

LLM

FLAN-T5-small

max_new_tokens = 180; deterministic

Fallbacks

TF-IDF + extractive answer

Keeps app usable if optional models fail

AI-Powered Document Insights and Data Extraction

4 of 9

4

Pipeline Performance Metrics | Four-query functional test

Measured on the supplied two-page pharmaceutical test PDF using the reproducible evaluation cell

100%

Recall@2

1.00

MRR

50%

Precision@2

100%

Hit rate

END-TO-END QUALITY

100%

Answer accuracy

100%

Citation accuracy

SYSTEM PERFORMANCE

4.0 ms

two-page PDF extraction

0.8 ms

average fallback retrieval

Note: this small functional test demonstrates behaviour, not production-scale performance.

AI-Powered Document Insights and Data Extraction

5 of 9

5

Pipeline in Action | Factual Quality Lookup

Query 1 validates quality-document routing, numeric extraction and source attribution

QUESTION

“What was the aspirin assay result?”

RETRIEVED CHUNK

Certificate of Quality — page 1

“The aspirin assay result was 99.4 percent. The approved assay range was 98.0 to 102.0 percent.”

TOP MATCH

AI-Powered Pharmaceutical Document Intelligence

Upload digital or scanned PDFs

Drop PDFs here

Process and Index

Indexed 1 file • 2 pages • 2 chunks

Document type: All

☑ Auto-route

Top-k: 4

Grounded answers

What was the aspirin assay result?

The aspirin assay result was 99.4%, within the approved range.

Confidence: 78.6% | Chunks used: 2

Source: rag_test_document_final.pdf, page 1

Certificate of Quality | score 0.812

Ask a question...

Ask

FINAL ANSWER

The aspirin assay result was 99.4%, within the approved range.

AI-Powered Document Insights and Data Extraction

6 of 9

6

Pipeline in Action | Packaging Requirements

Query 2 demonstrates metadata routing between two document types

QUESTION

“What information must appear on the product label?”

RETRIEVED CONTEXT

Packaging Specification — page 2

“The product label must show the batch number, expiry date, storage instructions, and regulatory identifier.”

RESPONSE QUALITY

Correct type

Packaging Specification

Completeness

Four required label fields

Source

Page 2 citation

Limitation

No structured table extraction

FINAL ANSWER

Batch number, expiry date, storage instructions and regulatory identifier.

AI-Powered Document Insights and Data Extraction

7 of 9

7

Design Decision Analysis | Reliable on Limited Hardware

The design favours transparency, graceful degradation and maintainable components

Decision

Rationale

Trade-off

MiniLM embeddings

Fast and lightweight for Colab

Less domain-specific than larger encoders

900/150 chunking

Simple; preserves page evidence

May split clauses across boundaries

FLAN-T5-small

Open-source and CPU-friendly

Lower reasoning quality than larger LLMs

FAISS in memory

Simple, fast demo indexing

Not persistent or distributed

TF-IDF fallback

Prevents indexing failure

Weaker semantic matching

Extractive fallback

Always grounded in evidence

Less fluent than generated answers

KEY TRADE-OFFS

Speed vs. accuracy

Smaller models improve accessibility but reduce generation depth.

Complexity vs. maintainability

Fallbacks add code paths, but prevent the entire demo from failing when optional components are unavailable.

AI-Powered Document Insights and Data Extraction

8 of 9

8

Current Limitations & Next Steps | Improvement Roadmap

LOW-QUALITY INPUT

Handwriting, skew and low-DPI scans can weaken OCR.

RETRIEVAL GAPS

Broad questions may match more than one document type.

SCALABILITY

In-memory indexes and synchronous processing suit a demo, not 1,000+ PDFs.

PROPOSED ENHANCEMENTS

SHORT TERM

Add image preprocessing, confidence thresholds and more annotated test questions.

MEDIUM TERM

Add hybrid BM25 + vector retrieval, a cross-encoder reranker and structured table parsing.

LONG TERM

Persist embeddings in a production vector database; add asynchronous ingestion, access control, monitoring and audit logs.

AI-Powered Document Insights and Data Extraction

9 of 9

9

Project Impact & Learning Outcomes

A practical demonstration of technical delivery, transparent reporting and product thinking

TECHNICAL LEARNING

Python pipeline design

OCR and PDF extraction

RAG retrieval and prompting

FAISS and TF-IDF

Gradio event wiring

Evaluation metrics

POTENTIAL IMPACT

Faster document lookup

Traceable evidence for review

Reduced manual page searching

Reusable architecture for new document types

Clearer escalation when context is insufficient

PROFESSIONAL SKILLS

Problem solving

Debugging and resilience

Documentation

Technical storytelling

Balancing speed, accuracy and maintainability

Outcome: a working, transparent pipeline that can be explained to both technical and non-technical audiences.

AI-Powered Document Insights and Data Extraction