ABCDEFGHIJKLMNOPQRSTUVWXYZ
1
GenAI Evaluation Across the World — Workshop Templates
2
Simplified workbook for a small-scale G11n GenAI evaluation exercise.
3
4
Workshop flowWhat participants doWorkbook tabsPurpose
5
1. Select CategoryChoose one of the four evaluation categories.SC & Prompt DesignSelect SC, create prompts, localize to es-ES, define gold truth.
6
2. Select 5 Success CriteriaPick 5 SC from the selected category.Testing TrackingRecord run results for ChatGPT and Gemini.
7
3. Create 5 Base PromptsCreate one en-EN prompt per SC.Issue LogDocument issues, severity, evidence, and notes.
8
4. Localize PromptsLocalize each prompt to es-ES.Model ScoreAutomatically calculates score per SC and final score per model.
9
5. Define Gold TruthWrite the expected result for each prompt.Reference_SC_ListSimplified list of real model success criteria for selection.
10
6. Run EvaluationExecute each prompt twice in two AI models.ListsDropdown values used by the workbook.
11
7. Register ResultsRecord Pass/Fail results and evidence.
12
8. Log IssuesClassify issues by severity and type.
13
9. Calculate ScoreUse the Model Score sheet to compare final results.
14
15
16
Scoring logic
17
Pass + Pass = 100%
Pass + Fail = 50%
Fail + Fail = 0%
Final AI Score = average of the 5 SC scores
18
Use Evidence Link / Notes to reference screenshots, transcripts, or issue details.
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100