1 of 20

Semi-automatic LoRA Model Evaluator

CSCI 2952Y - Ziheng Huang, Yuki Zang, Yang Xiang

2 of 20

All activities 110/150

  • Core Generation Pipeline MVP - 30/30
  • Interactive Scoring Interface + Ablation - 30/40
  • Automated Report Generation - 30/30
  • AI-Assisted Evaluation - 0/20
  • LLM-Assisted Metric Generation - 10/10
  • System Self-Evaluation (User Study) - 10/20

3 of 20

Core Generation Pipeline MVP - 30/30

  • Accept LoRA, Metrics, base model, lists of prompts, and lists of seeds. ✅
  • Our system will iterate through all combinations of (Weight x Prompt x Seed) and call the inference engine. ✅
  • Save images with metadata and display them in a web-based grid✅
  • Our system will have a simple scoring ui for each image, based on a single criterion defined by user or by default. ✅
  • Generate the basic report including image grids. ✅

4 of 20

5 of 20

6 of 20

7 of 20

8 of 20

Interactive Scoring Interface - 30/40

  • Custom evaluation metrics. ✅
  • Image pair rating. ✅
  • Scores saving. ✅
  • Extra points for optimized ui when there are multiple metrics. ❌

9 of 20

10 of 20

11 of 20

Automated Report Generation - 30/30

  • Score analysis and optimal weight calculation. ✅
  • Representative "best" and "worst" prompts. ✅
  • Representative "best" and "worst" images. ✅
  • Final report. ✅

12 of 20

13 of 20

14 of 20

15 of 20

16 of 20

17 of 20

18 of 20

System Self-Evaluation (User Study) - 10/20

Metric

Baseline (Manual Testing)

Our Method (Semi-auto Evaluator)

Task Completion Time (minutes, Mean ± Std)

17.0 (± 4.6)

18.0 (± 4.8)

Perceived Usability (1-5)

2.6

3.4

Quality of Generated Report (1-5)

3.0

3.9

Preferred Method for Future Use (1-5)

3.1

4.2

N = 100 images

19 of 20

User comments

Pros:

  • Concise weight-quality plot
  • Representative pairs are straight and cool
  • Analysis drive unexpected conclusions

Cons:

  • A bit slow
  • Report loses prompt-specific information
  • The lora model sometimes does not work (due to system setting and different base model)

20 of 20

All activities 110/150

  • Core Generation Pipeline MVP - 30/30
  • Interactive Scoring Interface + Ablation - 30/40
  • Automated Report Generation - 30/30
  • AI-Assisted Evaluation - 0/20
  • LLM-Assisted Metric Generation - 10/10
  • System Self-Evaluation (User Study) - 10/20