1 of 10

2 of 10

Create a Contextual Chatbot with LLM and VectorDb in 10 Minutes

Raahul Dutta

MLOps Lead

@Elsevier

3 of 10

USE CASE

  • Scientific research should be leveraged to address the pressing global challenges our society faces.

  • However, due to the growing number of publications and conferences, the pressure on policy makers to stay up to date with the latest research has grown significantly in recent years, making evidence-based policy decisions hard and resource intensive.

  • Using generative AI to analyse data and generate comprehensive briefs thus allows.

Distinguish Engineer�Elsevier

Intern Research Engineer�Elsevier

MLOps Lead�Elsevier

4 of 10

Implementation

5 of 10

Embedding

  • TensorRT Pytorch Pipeline (KServe bato produce the embeddings.�
  • We used “"intfloat/simlm-base-msmarco-finetuned".�
  • We stored the embeddings in numpy vectors.

6 of 10

Vector Database (Qdrant)

  • We used Qdrant - quadrant.tech as Vector Database
    • Documentation : 😎
    • Multi Language : 👌�
  • We used Rust Code to upload the embedding Vectors.�
  • We stored the embeddings in numpy vectors.

  • We tuned
    • Indexing Optimizer
    • Memmap_threshold

  • Deployed the docker in K8 cluster.

7 of 10

LLM - Fine Tuning

  • We are trying -
    • Vicuna 13b - 👍
    • facebook/opt-6.7b - 👍
    • Bloom 7b - 🥱
    • RedPajama-INCITE-Instruct-3B-v1 - 🛠️
    • Fine tuned Falcon 7b Instruct - 👍�
  • Fine Tuning - PeFT and QLoRA
  • OnnxOptimizer and Triton Inference Server to host.

8 of 10

Web App - Gradio

9 of 10

Demo

10 of 10

Thank

YOU!