1 of 12

2024

Soumen Paul, IIT Kharagpur, India

Srijoni Majumdar, University of Leeds, UK

Raj Shah, IIT Goa, India

Susmita Das, University of Glasgow, UK

Madhusudan Ghosh, IACS, India

Debasis Ganguly, University of Glasgow, UK

Gul Calikli, University of Glasgow, UK

Debarshi Sanyal, IACS, India

Partha Pratim Das, IIT Kharagpur, India

Paul D Clough, University of Sheffield, UK

2 of 12

3 of 12

4 of 12

5 of 12

Generative AI based Software Metadata Classification

ST-1: Comment Usefulness Prediction

a) 9048 pairs of code and comments written in C, labeled as either Useful or Not Useful.

b) Code and Comment Pairs, written in C with generated labels of useful / not useful using any Large Language Model Architecture

Classification model with and without the new set of code comment pairs and generated labels

ST-2: Code Quality Estimation

Estimate the functional correctness of code snippets generated via LLMs in response to a prompt specifying a programming task.

https://huggingface.co/datasets/openai/openai_humaneval

6 of 12

Features

Vector Space Representation

7 of 12

Inconsistencies and Redundancies in Code and Comment

Similarities between Vectors

8 of 12

Participation

14 Submissions

Industry and Academia

Multiple runs

Total 56 Runs

9 of 12

Binary Classification

10 of 12

Pre-trained Embeddings - Vectors

11 of 12

In a nutshell

IRSE 2024 track empirically investigates approaches to evaluate the LLM capabilities in terms of code metadata quality and code generation.

The LLM-generated labels reduce the overfitting of the overall classification model for comment quality

Improve the F1 score when the combined data from all participants

12 of 12

IRSE 2022 and IRSE 2023

11 Teams, 34 runs -> 17 Teams, 56 runs

Increased participation from industry developers - Microsoft, American Express, Amazon, Bosch Research

Dataset Curated: 2 sets (9048, 11467 rows)

New Datasets Generated: 9500 code and comment evaluated pairs