XAlign: Cross-lingual Fact-to-Text Alignment and Generation for Low-Resource Languages
�Tushar Abhishek, Shivprasad Sagare, Bhavyajeet Singh, Anubhav Sharma, Manish Gupta and Vasudeva Varma
�����Information Retrieval and Extraction Lab, IIIT Hyderabad, India
1
25-29 April 2022 I Lyon, France
Agenda
2
25-29 April 2022 I Lyon, France
XF2T and XAlign
3
25-29 April 2022 I Lyon, France
XF2T Example
4
25-29 April 2022 I Lyon, France
Agenda
5
25-29 April 2022 I Lyon, France
Related Work
6
25-29 April 2022 I Lyon, France
Agenda
7
25-29 April 2022 I Lyon, France
Gathering Sentences
8
25-29 April 2022 I Lyon, France
Gathering Facts
9
25-29 April 2022 I Lyon, France
Agenda
10
25-29 April 2022 I Lyon, France
Two Stage Approach
11
25-29 April 2022 I Lyon, France
XF2T System Architecture
12
25-29 April 2022 I Lyon, France
Agenda
13
25-29 April 2022 I Lyon, France
Candidate Generation
14
25-29 April 2022 I Lyon, France
Agenda
15
25-29 April 2022 I Lyon, France
Candidate Selection
Goal: Retain only strongly aligned (fact, sentence) pairs.
16
25-29 April 2022 I Lyon, France
(Fact, Sentence) Alignment Accuracy
17
mT5 with transfer learning performs the best.
Stage-2 (Fact, Sentence) Candidate Selection F1
25-29 April 2022 I Lyon, France
Agenda
18
25-29 April 2022 I Lyon, France
XAlign Dataset Statistics
19
25-29 April 2022 I Lyon, France
Fact Distribution across Languages
20
25-29 April 2022 I Lyon, France
Agenda
21
25-29 April 2022 I Lyon, France
XF2T BLEU on XAlign Test Set
22
25-29 April 2022 I Lyon, France
Test Dataset Examples
23
25-29 April 2022 I Lyon, France
Agenda
24
25-29 April 2022 I Lyon, France
Conclusion
Acknowledgements: This research was partially funded by Ministry of Electronics and Information Technology (MeitY), Government of India.
Code: https://github.com/tushar117/XAlign
25
25-29 April 2022 I Lyon, France
References
[1] O Agarwal, H Ge, S Shakeri, and R Al-Rfou 2021 Knowledge Graph Based Synthetic Corpus Generation for Knowledge-Enhanced Language Model Pretraining In NAACL-HLT 3554–3565
[2] G Attardi 2015 WikiExtractor https://githubcom/attardi/wikiextractor
[3] K Bali, M Choudhury, and P Biswas 2010 Indian Language POS Tagset: Bengali Linguistic Data Consortium, LDC2010T16 (2010)
[4] J A Botha, Z Shan, and D Gillick 2020 Entity Linking in 100 Languages In EMNLP 7833–7845
[5] M Chen, S Wiseman, and K Gimpel 2021 WIKITABLET: A Large-Scale Data-toText Dataset for Generating Wikipedia Article Sections In ACL-IJCNLP Findings 193–209
[6] W Chen, Y Su, X Yan, and W Y Wang 2020 KGPT: Knowledge-Grounded Pre-Training for Data-to-Text Generation arXiv:201002307 (2020)
[7] A Conneau, K Khandelwal, N Goyal, V Chaudhary, G Wenzek, F Guzmán, É Grave, M Ott, L Zettlemoyer, and V Stoyanov 2020 Unsupervised Cross-lingual Representation Learning at Scale In ACL 8440–8451
[8] D Duma and E Klein 2013 Generating natural language from linked data: Unsupervised template extraction In IWCS 83–94
[9] H Elsahar, P Vougiouklis, A Remaci, C Gravier, J Hare, F Laforest, and E Simperl 2018 T-rex: A large scale alignment of natural language with knowledge base triples In LREC
[10] T Ferreira, C Gardent, N Ilinykh, C van der Lee, S Mille, D Moussallem, and A Shimorina 2020 The 2020 Bilingual, Bi-Directional WebNLG+ Shared Task: Overview and Evaluation Results In WebNLG+ 55–76
[11] Z Fu, B Shi, W Lam, L Bing, and Z Liu 2020 Partially-aligned data-to-text generation with distant supervision arXiv:201001268 (2020)
[12] C Gardent, A Shimorina, S Narayan, and L Perez-Beltrachini 2017 The WebNLG challenge: Generating text from RDF data In INLG 124–133
[13] Z Jin, Q Guo, X Qiu, and Z Zhang 2020 Genwiki: A dataset of 13 million content-sharing text and graphs for unsupervised graph-to-text generation In COLING 2398–2409
[14] S Khanuja, D Bansal, S Mehtani, S Khosla, A Dey, B Gopalan, D K Margam, P Aggarwal, R T Nagipogu, S Dave, et al 2021 Muril: Multilingual representations for indian languages arXiv:210310730 (2021)
26
25-29 April 2022 I Lyon, France
References
[15] K Kolluru, M Rezk, P Verga, W W Cohen, and P Talukdar 2021 Multilingual Fact Linking In AKBC
[16] A Kunchukuttan 2020 The IndicNLP Library https://githubcom/ anoopkunchukuttan/indic_nlp_library/blob/master/docs/indicnlppdf
[17] R Lebret, D Grangier, and M Auli 2016 Neural Text Generation from Structured Data with Application to the Biography Domain In EMNLP 1203–1213
[18] H Mei, M Bansal, and M R Walter 2016 What to talk about and how? Selective Gen using LSTMs with Coarse-to-Fine Alignment In NAACL-HLT 720–730
[19] P Nema, S Shetty, P Jain, A Laha, K Sankaranarayanan, and M M Khapra 2018 Generating Descriptions from Structured Data Using a Bifocal Attention Mechanism and Gated Orthogonalization In NAACL-HLT 1539–1550
[20] J Novikova, O Dušek, and V Rieser 2017 The E2E dataset: New challenges for end-to-end generation arXiv:170609254 (2017)
[21] C Patel and K Gali 2008 Part-of-speech tagging for Gujarati using conditional random fields In IJCNLP Workshop on NLP for Less Privileged Languages
[22] P Qi, Y Zhang, Y Zhang, J Bolton, and C D Manning 2020 Stanza: A Python Natural Language Processing Toolkit for Many Human Languages In ACL Demos https://nlpstanfordedu/pubs/qi2020stanzapdf
[23] G Ramesh, S Doddapaneni, A Bheemaraj, M Jobanputra, Raghavan AK, A Sharma, S Sahoo, H Diddee, D Kakwani, N Kumar, et al 2021 Samanantar: The Largest Publicly Available Parallel Corpora Collection for 11 Indic Languages arXiv:210405596 (2021)
[24] E Reiter and R Dale 1997 Building applied natural language generation systems NL Engineering 3, 1 (1997), 57–87
[25] H Shahidi, M Li, and J Lin 2020 Two Birds, One Stone: A Simple, Unified Model for Text Generation from Structured and Unstructured Data In ACL 3864–3870
[26] L Xue, N Constant, A Roberts, M Kale, R Al-Rfou, A Siddhant, A Barua, and C Raffel 2021 mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer In NAACL-HLT 483–498
[27] C Zhao, M Walker, and S Chaturvedi 2020 Bridging the structural gap between encoding and decoding for data-to-text generation In ACL 2481–2491
27
25-29 April 2022 I Lyon, France