1 of 11

AI QA COPILOT

Kiểm thử RAG chatbot

không thể chỉ bằng “expected output”

Từ requirement đến bằng chứng chất lượng có thể truy vết.

REQ

TEST

RUN

EVAL

REPORT

AI Agent hỗ trợ QA tự động sinh, thực thi và đánh giá test case cho RAG Chatbot

01

2 of 11

PAIN

QA truyền thống không được thiết kế để kiểm thử câu trả lời xác suất

Phần mềm truyền thống

Input

Logic

Output

PASS / FAIL rõ ràng

Expected output cố định

RAG chatbot

Natural language

Xác suất

Context-dependent

QA phải đánh giá thêm

Relevancy

Groundedness

Hallucination

PII / Safety

Prompt injection

02

3 of 11

PAIN

Một test case đang kéo QA qua 7 bước thủ công

1

Đọc

requirement

2

Xác định

scenario

3

Chuẩn bị

data

4

Gọi

API

5

Đọc

response

6

Đánh giá

7

Lưu

kết quả

Tốn thời gian

Nhiều thao tác lặp

Coverage khó

Dễ bỏ sót failure mode

Khó tái hiện

Input/version rời rạc

Khó so sánh

Regression thiếu baseline

03

4 of 11

PRODUCT

Từ requirement đến test report trong một workflow

01

Requirement

02

Clarify

03

Generate

04

Execute

05

Evaluate

06

Report

AI hỗ trợ

Human-in-the-loop

Traceability xuyên suốt

Evidence thay vì “cảm giác đúng”

AI copilot cho QA — không tự quyết định release thay con người.

04

5 of 11

LIVE DEMO

Thay vì giải thích thêm,

hãy xem một requirement

trở thành test report như thế nào.

05

6 of 11

DIFFERENTIATION

6 AI roles biến requirement thành test case có kiểm soát

1

Requirement

Analyst

2

Risk & Failure

Analyst

3

Coverage

Designer

4

Scenario

Generator

5

Scenario Critic

/ Curator

6

Testcase

Designer

critique → gap → regenerate · tối đa 5 vòng

Điều kiện dừng

Coverage ≥ 92% • Critical requirements = 100% • Không còn critical gap

Không phải một prompt “hãy sinh test case”.

06

7 of 11

EVALUATION

Không chỉ biết test fail — QA biết fail vì lý do gì

RAG RESPONSE

“Câu trả lời

chatbot”

LAYER 1

Deterministic checks

HTTP / Schema • Regex

PII • Forbidden content

LAYER 2

LLM / DeepEval

Faithfulness • Relevancy

Safety • RAG metrics

TRACEABLE REPORT

PASS / FAIL / ERROR

Rule + metric + evidence

Input + version liên quan

Mandatory rule fail → skip LLM cost

07

8 of 11

ARCHITECTURE

Modular monolith giúp MVP nhanh — nhưng vẫn tách rõ execution và evaluation

QA

User

Web UI

Next.js

API

FastAPI

Queue

Redis + RQ

Worker

Isolated

RAG

Target API

Eval

DeepEval

Report

PostgreSQL

Realtime progress: SSE

Deployment: Docker Compose

SECURITY CALLOUT

Token không vào PostgreSQL/log/report • Redis encrypted + TTL • Isolated runner + egress policy

08

9 of 11

FEASIBILITY

MVP tập trung vào một golden path có thể demo end-to-end

Configure

Clarify

Generate

Execute

Evaluate

Report

Kiểm soát

HITL trong mỗi bước dùng LLM

→ Đảm bảo chất lượng đầu ra.

Chi phí

Ưu tiên deterministic checks

→ chỉ gọi LLM khi cần

Độ ổn định

1 LLM adapter + giới hạn concurrency

→ giảm biến số trong MVP

09

10 of 11

FUTURE

Từ công cụ demo đến quality platform cho AI application

NOW

RAG QA MVP

• Golden path

• Traceable evidence

NEXT

Regression at scale

• CI/CD Quality Gate

• Quality Trend Dashboard

• Team Collaboration

LATER

AI App Quality Platform

• Multi-turn / tools

• Domain-specific QA

• Red Team & Safety

Quality evidence trở thành một phần của software delivery pipeline.

10

11 of 11

AI đang thay đổi cách chúng ta xây phần mềm.

QA cũng cần thay đổi cách chúng ta kiểm thử.

Từ RAG requirements

đến test có cấu trúc

đến bằng chứng chất lượng

From RAG requirements to traceable quality evidence.

AI QA Copilot for RAG

11