1 of 7

Abstractive Summarization of large articles in Hindi�Language using IndicBART�

Authors - Chaitanya Subhedar, Sheetal Sonawane

2 of 7

Experimentation challenges

IndicBERT –

  • Tokenizer often stripped vowel markers and diacritics, which are critical for meaning in Hindi.

mT5-small –

  • mT5’s multilingual nature spreads focus across 100+ languages, reducing representation for Hindi.
  • Used mT5-small variant due to limited GPU resources; larger models could not be tested.

Finally, since the news articles were large, the number of tokens generated far exceeded the input sequence limit for both models, which may have impacted performance.

3 of 7

Number of tokens

4 of 7

Results obtained for mT5

5 of 7

IndicBART

  • Pre-trained specifically on a diverse corpus of Indian languages, providing better linguistic coverage for Hindi

6 of 7

Methodology

Combining Headline and Article:

  • A "Combined" column was created by merging headlines with their corresponding articles to provide

better context for each data sample.

Chunking Strategy:

  • Texts in the "Combined" column were split into sentences and tokenized using the IndicBART tokenizer.
  • Sentences were grouped into chunks, ensuring each chunk stayed within the model’s token limit (1024 tokens).
  • The ID of the initial data sample was spread to all chunks from the same data sample.

Training Process:

  • The model was trained using Hugging Face’s Trainer API for 10 epochs, with a batch size of 4 and a weight decay of

0.01.

  • Each chunk was summarized individually during training.

Recombination and Evaluation:

  • Sub-summaries from all chunks of the same article were recombined using their common IDs to form the final

summary, which was then evaluated against target summaries using ROUGE scores.

7 of 7

Future Scope

  • Currently, semantic relations between chunked sentences are not considered.
  • Ensemble of extractive and abstractive approaches

Chunk 1

Chunk 2

Chunk 3

Chunk 4

Extractive Summarization

Small sized Chunk 1

Small sized Chunk 2

Small sized Chunk 3

Small sized Chunk 4

Abstractive Summarization

Final Summary