1
AI-POWERED
Document Insights
and Data Extraction
Pharmaceutical document intelligence with OCR, retrieval and grounded answers
John Raymond TYABA
Email: u2678314@uel.ac.uk
1
OCR + CLEAN
2
CHUNK + TAG
3
RETRIEVE
4
GENERATE
5
CITE
AI-Powered Document Insights and Data Extraction
2
Problem Statement | Pharmaceutical Document Intelligence
Extracting trustworthy answers from varied, unstructured pharmaceutical PDFs
CHALLENGE
Digital PDFs, scans, certificates and packaging specifications vary in structure and quality. Users must manually search pages, interpret document types and verify where each answer came from.
SOLUTION
The pipeline extracts or OCRs text, creates metadata-tagged chunks, retrieves relevant evidence and returns a source-grounded answer through Gradio.
PAIN POINT
HOW THE PIPELINE SOLVES IT
Low-text scans
Tesseract OCR fallback
Mixed document types
Metadata classification and filtering
Long unstructured content
Overlapping page-level chunks
Low trust in answers
Sources, pages, scores and confidence
AI-Powered Document Insights and Data Extraction
3
System Architecture | Resilient Open-Source RAG Pipeline
Document Input → OCR → Chunking → Embeddings → Vector Search → Retrieval → LLM / Extractive Answer
INPUT
OCR
Tesseract
CHUNK
900 chars
150 overlap
EMBED
MiniLM
INDEX
FAISS
RETRIEVE
Top-k 4
ANSWER
FLAN-T5
Component Specifications
Component | Technology choice | Configuration |
OCR | Tesseract + PDF2Image | 220 DPI; sparse-page trigger |
Chunking | Fixed page chunks | 900 characters; 150 overlap |
Embeddings | all-MiniLM-L6-v2 | Normalised dense vectors |
Vector index | FAISS IndexFlatIP | Inner-product similarity; in memory |
Retriever | Metadata-filtered dense search | Top-k = 4; manual filter + auto-route |
LLM | FLAN-T5-small | max_new_tokens = 180; deterministic |
Fallbacks | TF-IDF + extractive answer | Keeps app usable if optional models fail |
AI-Powered Document Insights and Data Extraction
4
Pipeline Performance Metrics | Four-query functional test
Measured on the supplied two-page pharmaceutical test PDF using the reproducible evaluation cell
100%
Recall@2
1.00
MRR
50%
Precision@2
100%
Hit rate
END-TO-END QUALITY
100%
Answer accuracy
100%
Citation accuracy
SYSTEM PERFORMANCE
4.0 ms
two-page PDF extraction
0.8 ms
average fallback retrieval
Note: this small functional test demonstrates behaviour, not production-scale performance.
AI-Powered Document Insights and Data Extraction
5
Pipeline in Action | Factual Quality Lookup
Query 1 validates quality-document routing, numeric extraction and source attribution
QUESTION
“What was the aspirin assay result?”
RETRIEVED CHUNK
Certificate of Quality — page 1
“The aspirin assay result was 99.4 percent. The approved assay range was 98.0 to 102.0 percent.”
TOP MATCH
AI-Powered Pharmaceutical Document Intelligence
Upload digital or scanned PDFs
Drop PDFs here
Process and Index
Indexed 1 file • 2 pages • 2 chunks
Document type: All
☑ Auto-route
Top-k: 4
Grounded answers
What was the aspirin assay result?
The aspirin assay result was 99.4%, within the approved range.
Confidence: 78.6% | Chunks used: 2
Source: rag_test_document_final.pdf, page 1
Certificate of Quality | score 0.812
Ask a question...
Ask
FINAL ANSWER
The aspirin assay result was 99.4%, within the approved range.
AI-Powered Document Insights and Data Extraction
6
Pipeline in Action | Packaging Requirements
Query 2 demonstrates metadata routing between two document types
QUESTION
“What information must appear on the product label?”
RETRIEVED CONTEXT
Packaging Specification — page 2
“The product label must show the batch number, expiry date, storage instructions, and regulatory identifier.”
RESPONSE QUALITY
Correct type
Packaging Specification
Completeness
Four required label fields
Source
Page 2 citation
Limitation
No structured table extraction
FINAL ANSWER
Batch number, expiry date, storage instructions and regulatory identifier.
AI-Powered Document Insights and Data Extraction
7
Design Decision Analysis | Reliable on Limited Hardware
The design favours transparency, graceful degradation and maintainable components
Decision | Rationale | Trade-off |
MiniLM embeddings | Fast and lightweight for Colab | Less domain-specific than larger encoders |
900/150 chunking | Simple; preserves page evidence | May split clauses across boundaries |
FLAN-T5-small | Open-source and CPU-friendly | Lower reasoning quality than larger LLMs |
FAISS in memory | Simple, fast demo indexing | Not persistent or distributed |
TF-IDF fallback | Prevents indexing failure | Weaker semantic matching |
Extractive fallback | Always grounded in evidence | Less fluent than generated answers |
KEY TRADE-OFFS
Speed vs. accuracy
Smaller models improve accessibility but reduce generation depth.
Complexity vs. maintainability
Fallbacks add code paths, but prevent the entire demo from failing when optional components are unavailable.
AI-Powered Document Insights and Data Extraction
8
Current Limitations & Next Steps | Improvement Roadmap
LOW-QUALITY INPUT
Handwriting, skew and low-DPI scans can weaken OCR.
RETRIEVAL GAPS
Broad questions may match more than one document type.
SCALABILITY
In-memory indexes and synchronous processing suit a demo, not 1,000+ PDFs.
PROPOSED ENHANCEMENTS
SHORT TERM
Add image preprocessing, confidence thresholds and more annotated test questions.
MEDIUM TERM
Add hybrid BM25 + vector retrieval, a cross-encoder reranker and structured table parsing.
LONG TERM
Persist embeddings in a production vector database; add asynchronous ingestion, access control, monitoring and audit logs.
AI-Powered Document Insights and Data Extraction
9
Project Impact & Learning Outcomes
A practical demonstration of technical delivery, transparent reporting and product thinking
TECHNICAL LEARNING
Python pipeline design
OCR and PDF extraction
RAG retrieval and prompting
FAISS and TF-IDF
Gradio event wiring
Evaluation metrics
POTENTIAL IMPACT
Faster document lookup
Traceable evidence for review
Reduced manual page searching
Reusable architecture for new document types
Clearer escalation when context is insufficient
PROFESSIONAL SKILLS
Problem solving
Debugging and resilience
Documentation
Technical storytelling
Balancing speed, accuracy and maintainability
Outcome: a working, transparent pipeline that can be explained to both technical and non-technical audiences.
AI-Powered Document Insights and Data Extraction