2024
Soumen Paul, IIT Kharagpur, India
Srijoni Majumdar, University of Leeds, UK
Raj Shah, IIT Goa, India
Susmita Das, University of Glasgow, UK
Madhusudan Ghosh, IACS, India
Debasis Ganguly, University of Glasgow, UK
Gul Calikli, University of Glasgow, UK
Debarshi Sanyal, IACS, India
Partha Pratim Das, IIT Kharagpur, India
Paul D Clough, University of Sheffield, UK
Generative AI based Software Metadata Classification
ST-1: Comment Usefulness Prediction
a) 9048 pairs of code and comments written in C, labeled as either Useful or Not Useful.
b) Code and Comment Pairs, written in C with generated labels of useful / not useful using any Large Language Model Architecture
Classification model with and without the new set of code comment pairs and generated labels
ST-2: Code Quality Estimation
Estimate the functional correctness of code snippets generated via LLMs in response to a prompt specifying a programming task.
https://huggingface.co/datasets/openai/openai_humaneval
Features
Vector Space Representation
Inconsistencies and Redundancies in Code and Comment
Similarities between Vectors
Participation
14 Submissions
Industry and Academia
Multiple runs
Total 56 Runs
Binary Classification
Pre-trained Embeddings - Vectors
In a nutshell
IRSE 2024 track empirically investigates approaches to evaluate the LLM capabilities in terms of code metadata quality and code generation.
The LLM-generated labels reduce the overfitting of the overall classification model for comment quality
Improve the F1 score when the combined data from all participants
IRSE 2022 and IRSE 2023
11 Teams, 34 runs -> 17 Teams, 56 runs
Increased participation from industry developers - Microsoft, American Express, Amazon, Bosch Research
Dataset Curated: 2 sets (9048, 11467 rows)
New Datasets Generated: 9500 code and comment evaluated pairs