1 of 10

XI INTERNATIONAL CONFERENCE

“INFORMATION TECHNOLOGY AND IMPLEMENTATION” (IT&I-2024)

Features of the practical use of LLM for generating quiz

Oleh Ilarionov , Hanna Krasovska , �Iryna Domanetska , Olena Fedusenko

2 of 10

The potential of LLM in education

Переваги LLM в освіті:

  • Автоматизація рутинних завдань.
  • Персоналізація навчальних матеріалів.
  • Адаптація завдань до рівня знань студентів.

​

Моделі, використані у дослідженні:

  • GPT-4 (OpenAI), Claude (Anthropic), Copilot (Microsoft), Gemini (Google DeepMind).

​

Приклади можливостей:

  • Генерація тестів різного рівня складності.
  • Підтримка різних форматів завдань: тести з вибором, запитання з кількома правильними відповідями, відкриті питання тощо.

3 of 10

A query optimisation

Query optimisation for generating high-quality test tasks :

  • Identification of the subject matter of the note.
  • Formulation of the role for LLM (e.g., "virtual teacher").
  • Checking for key details in the request.
  • Unmasking of ambiguities.

Algorithm for forming an optimal query

4 of 10

Limitations of models in quiz generation

Оптимізація запиту для генерації якісних завдань:

  • Визначення мети запиту.
  • Формулювання ролі для LLM (наприклад, "віртуальний викладач").
  • Перевірка наявності ключових деталей у запиті.
  • Усунення неоднозначностей.

5 of 10

Research models and tools

Models : GPT, Claude, Copilot та Gemini.

​

A fragment of lecture material :

  • Format : PDF, size 361 КБ.
  • Characters without spaces : 9710;
  • Number of words : 1523;

​

Main evaluation criteria :

  • Compliance with Bloom's Taxonomy.
  • Structure and clarity.
  • Variety of task types.
  • Validity and discriminative power.

6 of 10

Comparative analysis of models

7 of 10

Comparative analysis of models

Parameter

GPT

Claude

Δ

Average question length (characters)

85

110

-25

Average length of justification (characters)

120

150

-30

Median number of words per question

15

18

-3

8 of 10

Quality of created tasks

Criterion

GPT

Claude

Δ

Content validity

4.5

4.5

0

Construct validity

4.0

4.5

+0.5

Clarity of wording

4.5

4.5

0

No ambiguity

4.0

4.5

+0.5

Relevance of answer options

4.5

4.5

0

Cognitive level (according to Bloom)

3.5

4.0

+0.5

Average score

4.17

4.42

+0.25

9 of 10

Conclusion

  • GPT has demonstrated the greatest flexibility in creating tasks of different formats and at different cognitive levels, according to Bloom's Taxonomy.
  • Claude, on the other hand, received higher scores for construct validity and clarity of task wording.
  • Copilot and Gemini proved to be less versatile than GPT and Claude, in particular due to the limited number of available task formats in the free mode.
  • The formulation of queries is an important factor in obtaining relevant answers from the models. Teachers need to take into account both the limitations of the models (e.g., the amount of text processed) and the peculiarities of generating questions at different cognitive levels.

10 of 10

THANK YOU FOR ATTENTION