1 of 23

NatGen: Generative Pre-Training by “Naturalizing” Source Code

Saikat Chakraborty, Toufique Ahmed, Yangruibo Ding,

Prem Devanbu, Baishakhi Ray

Columbia University and UC Davis

1

2 of 23

Machine Learning for Code

2

Coding

Test

Bug Fix

Bug

Finding

Document

3 of 23

Machine Learning for Code

3

Language Model

Task Model

4 of 23

Machine Learning for Code

4

Language Model

Task Model

Pretraining

Fine-Tuning

5 of 23

Hurdles of Using Pretrained Language Model

5

  • Resource Constraint
    • Not all large model is feasible for everyone to use.

  • Misalignment of Pretraining and Finetuning tasks
    • e.g., CodeBERT and GraphCodeBERT is designed to understand code. Not very suitable for generative tasks without substantial fine tuning.

  • Lack of sufficient Finetuning Data
    • Often requires large quantity of labelled data.
      • Lower quantity of data than that of training from scratch

6 of 23

PLBART – Token Based Denoising

6

Correct Code

Noisy Code

Encoder

Decoder

Noise Injector

PLBART

Token Masking

Token Deletion

Token Infilling

7 of 23

Dual Information Channel of Code

7

[1] Casalanuovo et. al. 2020

[2] Karampatsis et. al. 2020

Natural Channel

Formal Channel

8 of 23

Pre-Training through Formal Channel Mutation

8

Encoder

Decoder

Noise Injector

NatGen

Semantic Preserving Formal Channel Mutations

Correct Code

Noisy Code

9 of 23

Pre-Training through Formal Channel Mutation

9

De-Naturalizing Transformation

10 of 23

De-Naturalizing Transformations

10

11 of 23

De-Naturalizing Transformations

11

12 of 23

Pre-Training through Formal Channel Mutation

12

Encoder

Decoder

Noise Injector

NatGen

Semantic Preserving Formal Channel Mutations

Correct Code

Noisy Code

13 of 23

Pre-training NatGen

13

  • Start from salesforce/codet5-base checkpoint.
  • CodeSearchNet dataset.
  • 25000 steps after CodeT5

Language

BlockSwap

OperandSwap

Confusion

DeadCode

VarRenamer

LoopTransformer

Total

%

Go

13299

116756

0

250506

322325

0

702886

8.65

Java

41286

182666

43925

557741

628122

55546

1509286

18.57

JavaScript

14170

251474

0

576047

765796

90389

1697876

20.89

Php

8411

95506

0

379711

451493

11353

946474

11.64

Python

27395

79854

0

480264

516266

3194

1106973

13.62

Ruby

3375

10990

0

74086

74482

0

162933

2.00

C

19618

121522

16850

380702

411401

49534

999627

12.30

C#

6563

63714

5703

395560

515125

12084

998749

12.29

Total

134117

922482

66478

3094617

3685010

222100

8124804

100

%

1.65

11.35

0.82

38.09

45.35

2.73

100

14 of 23

Q: Fine-tuning on downstream tasks

14

15 of 23

15

NL to Code Generation

Code Translation

Bug Fix

16 of 23

Q: NatGen �on Resource Constraint Environment�(Zero-shot)

16

17 of 23

Q: NatGen �on Resource Constraint Environment�(200 �Training Examples)

17

18 of 23

18

19 of 23

Other Comparison

19

Code Summarization

20 of 23

Lessons Learned

20

  • Semantic Preserving Rewriting can teach models write robust code.
  • Teaching rewriting rules can embed PL-knowledge into model.
  • [Future idea] Can be used for semantic style transfer learning.
  • [Future idea] Can be used for code refactoring.

21 of 23

Code and Model

21

  • Code for Semantic Preserving Transformation and Pretraining

  • Pretrained Model

from transformers import AutoTokenizer, AutoModelForSeq2SeqLM

tokenizer = AutoTokenizer.from_pretrained("saikatc/NatGen")

model = AutoModelForSeq2SeqLM.from_pretrained("saikatc/NatGen")

22 of 23

22

Prem Devanbu

UC Davis

Yangruibo Ding

Columbia

Toufique Ahmed Parag

UC Davis

Baishakhi Ray

Columbia

Thanks (for feedback)

  • Wasi Uddin Ahmad – Amazon
  • Vikram Nitin - Columbia

ARiSE lab, Columbia & DECAL lab, UC Davis.

23 of 23

23

Thanks!