1 of 7

Abstractive Summarization of large articles in Hindi�Language using IndicBART�

Authors - Chaitanya Subhedar, Sheetal Sonawane

​

2 of 7

Experimentation challenges

IndicBERT –

  • Tokenizer often stripped vowel markers and diacritics, which are critical for meaning in Hindi.

mT5-small –

  • mT5’s multilingual nature spreads focus across 100+ languages, reducing representation for Hindi.
  • Used mT5-small variant due to limited GPU resources; larger models could not be tested.

​

Finally, since the news articles were large, the number of tokens generated far exceeded the input sequence limit for both models, which may have impacted performance.

3 of 7

Number of tokens

4 of 7

Results obtained for mT5

5 of 7

IndicBART

  • Pre-trained specifically on a diverse corpus of Indian languages, providing better linguistic coverage for Hindi

​

​

​

​

​

6 of 7

Methodology

Combining Headline and Article:

  • A "Combined" column was created by merging headlines with their corresponding articles to provide

better context for each data sample.

Chunking Strategy:

  • Texts in the "Combined" column were split into sentences and tokenized using the IndicBART tokenizer.
  • Sentences were grouped into chunks, ensuring each chunk stayed within the model’s token limit (1024 tokens).
  • The ID of the initial data sample was spread to all chunks from the same data sample.

Training Process:

  • The model was trained using Hugging Face’s Trainer API for 10 epochs, with a batch size of 4 and a weight decay of

0.01.

  • Each chunk was summarized individually during training.

Recombination and Evaluation:

  • Sub-summaries from all chunks of the same article were recombined using their common IDs to form the final

summary, which was then evaluated against target summaries using ROUGE scores.

7 of 7

Future Scope

  • Currently, semantic relations between chunked sentences are not considered.
  • Ensemble of extractive and abstractive approaches

​

​

Chunk 1

Chunk 2

Chunk 3

Chunk 4

Extractive Summarization

Small sized Chunk 1

Small sized Chunk 2

Small sized Chunk 3

Small sized Chunk 4

Abstractive Summarization

Final Summary