PAT: Parameter-Free Audio-Text Aligner to Boost Zero-Shot Audio
Classification
Ashish Seth, Ramaneswaran Selvakumar, Sonal Kumar, Sreyan Ghosh,
Dinesh Manocha
(University Of Maryland, College Park)
Motivation
Problem Statement
“How can we adapt Audio-Language Models to out-of-distribution audio classification tasks in parameter/training free fashion”
Key contributions
Methodology (Date Store)
Guidelines for prompt generation
Methodology (Weighted Prompt Ensemble)
Methodology (Weighted Prompt Ensemble)
Methodology (Cross Modal Alignment)
Results (Main Results)
Performance comparison between PAT and vanilla zero-shot classification (ZS) across six ALEs and 16 diverse audio classification tasks, including 10 sound and 8 music datasets. The best scores for each ALE are bolded. Overall, PAT outperforms vanilla ZS, achieving improvements ranging from 0.42% to 27%.
Results (Ablations)
Zero-shot performance evaluation of MSCLAP-23 using PAT across 16 audio classification tasks under noisy conditions. Each audio sample undergoes five different types of audio augmentations. PAT outperforms vanilla zero-shot (ZS) classification, achieving an absolute improvement of 0.10%–11.15%.
Thank You:)