1 of 27

XAlign: Cross-lingual Fact-to-Text Alignment and Generation for Low-Resource Languages

Tushar Abhishek, Shivprasad Sagare, Bhavyajeet Singh, Anubhav Sharma, Manish Gupta and Vasudeva Varma 

�����Information Retrieval and Extraction Lab, IIIT Hyderabad, India 

1

25-29 April 2022 I Lyon, France

2 of 27

Agenda

  • Motivation
  • Related Work
  • Data Collection and Pre-processing
  • F2T Alignment
    • Stage-1 (Candidate Generation)
    • Stage-2 (Candidate Selection)
  • Manual Annotation and Ground-truth data
  • XF2T Generation
  • Conclusion

2

25-29 April 2022 I Lyon, France

3 of 27

XF2T and XAlign

  • Cross-Lingual Fact to Text XF2T
    • English fact triples 🡪 descriptive text in low resource (LR) languages
    • Needs aligned (English structured facts, LR sentences) data.
  • We contribute XAlign
    • An XF2T dataset with 0.45M pairs
    • 8 languages
    • 5402 pairs have been manually annotated

3

25-29 April 2022 I Lyon, France

4 of 27

XF2T Example

4

25-29 April 2022 I Lyon, France

5 of 27

Agenda

  • Motivation
  • Related Work
  • Data Collection and Pre-processing
  • F2T Alignment
    • Stage-1 (Candidate Generation)
    • Stage-2 (Candidate Selection)
  • Manual Annotation and Ground-truth data
  • XF2T Generation
  • Conclusion

5

25-29 April 2022 I Lyon, France

6 of 27

Related Work

  • Related problems: Entity linking, fact linking
  • Previous work is mainly on en-F2T only.
  • F2T methods: template-based methods [8], Seq-2-seq attention networks [17, 18, 25], hierarchical attention networks [6, 19], pretrained Transformers [6, 11, 27].

6

25-29 April 2022 I Lyon, France

7 of 27

Agenda

  • Motivation
  • Related Work
  • Data Collection and Pre-processing
  • F2T Alignment
    • Stage-1 (Candidate Generation)
    • Stage-2 (Candidate Selection)
  • Manual Annotation and Ground-truth data
  • XF2T Generation
  • Conclusion

7

25-29 April 2022 I Lyon, France

8 of 27

Gathering Sentences

  • For each of the 8 languages
    • Extract text from 2021-05-20 Wikipedia xml dump.
    • Split text into sentences using IndicNLP.​
    • Prune out
      • Other language sentences using Polyglot language detector
      • Sentences with <5 words or >100 words
      • Sentences with no factual information (i.e., no noun or verb).

8

25-29 April 2022 I Lyon, France

9 of 27

Gathering Facts

  • Extract English facts from 2020-12-21 WikiData dump.​
  • Retain useful factual information for person entities
    • WikibaseItem
    • Time
    • Quantity
    • Monolingualtext
  • ∼0.91M facts extracted for ∼85K entities across all the 8 languages

9

25-29 April 2022 I Lyon, France

10 of 27

Agenda

  • Motivation
  • Related Work
  • Data Collection and Pre-processing
  • F2T Alignment
    • Stage-1 (Candidate Generation)
    • Stage-2 (Candidate Selection)
  • Manual Annotation and Ground-truth data
  • XF2T Generation
  • Conclusion

10

25-29 April 2022 I Lyon, France

11 of 27

Two Stage Approach

  • First stage (Candidate Generation)
    • Provides high recall
    • Generates (facts, sentence) candidates based on automated translation and syntactic & semantic match. ​
  • Second stage (Candidate Selection)
    • Provides high precision
    • Retains only those candidates which are strongly aligned using transfer learning and distant supervision. ​

11

25-29 April 2022 I Lyon, France

12 of 27

XF2T System Architecture

12

25-29 April 2022 I Lyon, France

13 of 27

Agenda

  • Motivation
  • Related Work
  • Data Collection and Pre-processing
  • F2T Alignment
    • Stage-1 (Candidate Generation)
    • Stage-2 (Candidate Selection)
  • Manual Annotation and Ground-truth data
  • XF2T Generation
  • Conclusion

13

25-29 April 2022 I Lyon, France

14 of 27

Candidate Generation

  • Compute a similarity score that captures syntactic as well as semantic similarity between a (fact, sentence) pair. ​
  • Syntactic similarity: TF-IDF by translating either the fact to native language or the sentence to English. ​
  • Semantic similarity: Cosine similarity between MuRIL representations of the fact and the sentence​

14

25-29 April 2022 I Lyon, France

15 of 27

Agenda

  • Motivation
  • Related Work
  • Data Collection and Pre-processing
  • F2T Alignment
    • Stage-1 (Candidate Generation)
    • Stage-2 (Candidate Selection)
  • Manual Annotation and Ground-truth data
  • XF2T Generation
  • Conclusion

15

25-29 April 2022 I Lyon, France

16 of 27

Candidate Selection

Goal: Retain only strongly aligned (fact, sentence) pairs.

    • Distant supervision from another English-only F2T task
      • Train a binary classifier on KELM dataset
      • Task: predict whether the fact is associated with the LR language sentence or not.
    • Transfer learning from NLI (Natural language Inference) task
      • F2T~NLI
      • We experimented with: XLM-R, mT5, MuRIL.

16

25-29 April 2022 I Lyon, France

17 of 27

(Fact, Sentence) Alignment Accuracy

17

mT5 with transfer learning performs the best.

Stage-2 (Fact, Sentence) Candidate Selection F1

25-29 April 2022 I Lyon, France

18 of 27

Agenda

  • Motivation
  • Related Work
  • Data Collection and Pre-processing
  • F2T Alignment
    • Stage-1 (Candidate Generation)
    • Stage-2 (Candidate Selection)
  • Manual Annotation and Ground-truth data
  • XF2T Generation
  • Conclusion

18

25-29 April 2022 I Lyon, France

19 of 27

XAlign Dataset Statistics

  • mT5 Stage-2 aligner run on Stage-1 output gives Train+Validation part of XAlign.
  • Manual annotations to get the test set.

19

 

25-29 April 2022 I Lyon, France

20 of 27

Fact Distribution across Languages

20

25-29 April 2022 I Lyon, France

21 of 27

Agenda

  • Motivation
  • Related Work
  • Data Collection and Pre-processing
  • F2T Alignment
    • Stage-1 (Candidate Generation)
    • Stage-2 (Candidate Selection)
  • Manual Annotation and Ground-truth data
  • XF2T Generation
  • Conclusion

21

25-29 April 2022 I Lyon, France

22 of 27

XF2T BLEU on XAlign Test Set

  • mT5 performs best.
  • BLEU performance is good for en, hi and bn.
  • Further work needed for other LR languages.​

22

25-29 April 2022 I Lyon, France

23 of 27

Test Dataset Examples

23

25-29 April 2022 I Lyon, France

24 of 27

Agenda

  • Motivation
  • Related Work
  • Data Collection and Pre-processing
  • F2T Alignment
    • Stage-1 (Candidate Generation)
    • Stage-2 (Candidate Selection)
  • Manual Annotation and Ground-truth data
  • XF2T Generation
  • Conclusion

24

25-29 April 2022 I Lyon, France

25 of 27

Conclusion

  • Proposed a novel XF2T problem
  • Contributed a new XAlign dataset for 8 languages
  • Proposed two effective F2T alignment methods
  • Reported BLEU results on strong baseline multi-lingual models

Acknowledgements: This research was partially funded by Ministry of Electronics and Information Technology (MeitY), Government of India.

Code: https://github.com/tushar117/XAlign

Paper: https://arxiv.org/abs/2202.00291

Email: tushar.abhishek@research.iiit.ac.in

25

25-29 April 2022 I Lyon, France

26 of 27

References

[1] O Agarwal, H Ge, S Shakeri, and R Al-Rfou 2021 Knowledge Graph Based Synthetic Corpus Generation for Knowledge-Enhanced Language Model Pretraining In NAACL-HLT 3554–3565

[2] G Attardi 2015 WikiExtractor https://githubcom/attardi/wikiextractor

[3] K Bali, M Choudhury, and P Biswas 2010 Indian Language POS Tagset: Bengali Linguistic Data Consortium, LDC2010T16 (2010)

[4] J A Botha, Z Shan, and D Gillick 2020 Entity Linking in 100 Languages In EMNLP 7833–7845

[5] M Chen, S Wiseman, and K Gimpel 2021 WIKITABLET: A Large-Scale Data-toText Dataset for Generating Wikipedia Article Sections In ACL-IJCNLP Findings 193–209

[6] W Chen, Y Su, X Yan, and W Y Wang 2020 KGPT: Knowledge-Grounded Pre-Training for Data-to-Text Generation arXiv:201002307 (2020)

[7] A Conneau, K Khandelwal, N Goyal, V Chaudhary, G Wenzek, F Guzmán, É Grave, M Ott, L Zettlemoyer, and V Stoyanov 2020 Unsupervised Cross-lingual Representation Learning at Scale In ACL 8440–8451

[8] D Duma and E Klein 2013 Generating natural language from linked data: Unsupervised template extraction In IWCS 83–94

[9] H Elsahar, P Vougiouklis, A Remaci, C Gravier, J Hare, F Laforest, and E Simperl 2018 T-rex: A large scale alignment of natural language with knowledge base triples In LREC

[10] T Ferreira, C Gardent, N Ilinykh, C van der Lee, S Mille, D Moussallem, and A Shimorina 2020 The 2020 Bilingual, Bi-Directional WebNLG+ Shared Task: Overview and Evaluation Results In WebNLG+ 55–76

[11] Z Fu, B Shi, W Lam, L Bing, and Z Liu 2020 Partially-aligned data-to-text generation with distant supervision arXiv:201001268 (2020)

[12] C Gardent, A Shimorina, S Narayan, and L Perez-Beltrachini 2017 The WebNLG challenge: Generating text from RDF data In INLG 124–133

[13] Z Jin, Q Guo, X Qiu, and Z Zhang 2020 Genwiki: A dataset of 13 million content-sharing text and graphs for unsupervised graph-to-text generation In COLING 2398–2409

[14] S Khanuja, D Bansal, S Mehtani, S Khosla, A Dey, B Gopalan, D K Margam, P Aggarwal, R T Nagipogu, S Dave, et al 2021 Muril: Multilingual representations for indian languages arXiv:210310730 (2021)

26

25-29 April 2022 I Lyon, France

27 of 27

References

[15] K Kolluru, M Rezk, P Verga, W W Cohen, and P Talukdar 2021 Multilingual Fact Linking In AKBC

[16] A Kunchukuttan 2020 The IndicNLP Library https://githubcom/ anoopkunchukuttan/indic_nlp_library/blob/master/docs/indicnlppdf

[17] R Lebret, D Grangier, and M Auli 2016 Neural Text Generation from Structured Data with Application to the Biography Domain In EMNLP 1203–1213

[18] H Mei, M Bansal, and M R Walter 2016 What to talk about and how? Selective Gen using LSTMs with Coarse-to-Fine Alignment In NAACL-HLT 720–730

[19] P Nema, S Shetty, P Jain, A Laha, K Sankaranarayanan, and M M Khapra 2018 Generating Descriptions from Structured Data Using a Bifocal Attention Mechanism and Gated Orthogonalization In NAACL-HLT 1539–1550

[20] J Novikova, O Dušek, and V Rieser 2017 The E2E dataset: New challenges for end-to-end generation arXiv:170609254 (2017)

[21] C Patel and K Gali 2008 Part-of-speech tagging for Gujarati using conditional random fields In IJCNLP Workshop on NLP for Less Privileged Languages

[22] P Qi, Y Zhang, Y Zhang, J Bolton, and C D Manning 2020 Stanza: A Python Natural Language Processing Toolkit for Many Human Languages In ACL Demos https://nlpstanfordedu/pubs/qi2020stanzapdf

[23] G Ramesh, S Doddapaneni, A Bheemaraj, M Jobanputra, Raghavan AK, A Sharma, S Sahoo, H Diddee, D Kakwani, N Kumar, et al 2021 Samanantar: The Largest Publicly Available Parallel Corpora Collection for 11 Indic Languages arXiv:210405596 (2021)

[24] E Reiter and R Dale 1997 Building applied natural language generation systems NL Engineering 3, 1 (1997), 57–87

[25] H Shahidi, M Li, and J Lin 2020 Two Birds, One Stone: A Simple, Unified Model for Text Generation from Structured and Unstructured Data In ACL 3864–3870

[26] L Xue, N Constant, A Roberts, M Kale, R Al-Rfou, A Siddhant, A Barua, and C Raffel 2021 mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer In NAACL-HLT 483–498

[27] C Zhao, M Walker, and S Chaturvedi 2020 Bridging the structural gap between encoding and decoding for data-to-text generation In ACL 2481–2491

27

25-29 April 2022 I Lyon, France