1 of 29

Ph.D. Yearly Review (2020-21)

Yaman Kumar Singla

PhD19024

SUNY-Buffalo, IIITD

2 of 29

Background

  • May 2020-Present: PhD student at IIITD and SUNY-Buffalo
    • Advisors: Dr Rajiv Ratn Shah, Dr Changyou Chen
    • Awarded Google PhD fellowship (along with 39 other candidates all over the world)
  • Jun 2018-Present: Researcher @ Adobe MDSR
    • Before: SDE in Adobe
  • May 2018: B.Engg (Honors) in Computer Engineering from NSIT (Delhi University)
  • Nov 2017-Present: Graduate RA @ MIDAS Lab.

3 of 29

Course Work

Number

Course Name

Credits

Grade Point

1

Information Retrieval

4

10

2

Trustworthy AI Systems

4

10

3

Independent Study - Language Distance and Minimum Distance b/w two Languages

4

10

4

Independent Project - Interpretability of Transformers

4

10

CGPA: 10, Course Credits Remaining: 16, Thesis Credits Accumulated: 20

4 of 29

My Research Questions

RQ1: How can we make ML systems understand creativity?

  1. Can we integrate creativity with data?

RQ2: How do the deep learning systems work?

    • When do they fail (pass)? Why?
    • How and Why are the different ML models, “different”?

5 of 29

Projects Overview (May’20 Onwards)

  1. (RQ1) Automatic Scoring
    1. (RQ1) Speaker-Conditioned Hierarchical Modeling for Automated Speech Scoring
      1. Accepted at CIKM 2021
    2. LTRC 2021 Presentation
    3. System trial with Second Language Testing Inc in 8 countries including Japan, Philippines, USA, Iran, etc
    4. (RQ1, RQ2) AES Interpretability
  2. (RQ2) Interpretability of DL models
    • Audio Transformers Interpretability - Ongoing
    • Visual Speech
      • ICASSP 2021 Paper Published
    • Patent Granted by USPTO (US Patent 10,937,428) (Only Inventor)
  3. (RQ1, RQ2) Data Driven Content Workflows
    • 1 Patents Filed at USPTO (1st among 20 inventors)
    • Best Adobe Sneaks Award
    • Covered by Forrester, Fast Company and other media houses
    • Will be submitting paper to upcoming conferences

6 of 29

RQ1: Integrating Creativity with Data

  1. Automatic Essay Scoring
    1. Speaker Conditioned Hierarchical Modeling for Automated Speech Scoring (Next Slide)
    2. Accepted at CIKM-2021
  2. Project Hemingway: Predictive, Prescriptive, and Descriptive modeling of content analytics
    • Worked with Telegraph Media Group (UK), Datacom (Australia), Adobe Blogs team
    • Led to ~10% increase in customer engagement
    • Patent filed at USPTO (1st amongst 20 inventors)
  3. Project LearnAd: Information retrieval of creative multimodal advertisements
    • Sentiment, What (Action), Why (Reason), and Topic Prediction
    • Worked with Adobe Spark Team and Adobe Universal Search Team

7 of 29

8 of 29

9 of 29

10 of 29

11 of 29

12 of 29

13 of 29

14 of 29

RQ2: How Do Deep Learning Systems Work?

The Directions I am Exploring:

  1. (RQ2) Visual Speech
    1. LIFI: Towards Linguistically Informed Frame Interpolation
    2. Published at ICASSP
  2. (RQ2) Text and Speech Transformers
    • Powerhouse for the full ML stack: NLP, Speech, and Vision
    • Submitting to Computational Linguistics Journal
  3. (RQ1, RQ2) Automatic Essay Scoring -
      • Overstability and Oversensitivity of AES systems (Next Slide)
      • How to measure them?
      • Why does it arise?
      • Can we do something about it? How to cure it?
      • Submitting to Linguistics Journal

15 of 29

16 of 29

17 of 29

18 of 29

19 of 29

20 of 29

21 of 29

22 of 29

RQ1: How Do Deep Learning Systems Work?

The Directions I am Exploring:

  • (RQ1) Visual Speech
    • LIFI: Towards Linguistically Informed Frame Interpolation (Next Slide)
    • Published at ICASSP
  • (RQ1) Text and Speech Transformers
    • Powerhouse for the full ML stack: NLP, Speech, and Vision
    • Submitting to Computational Linguistics Journal
  • (RQ1, RQ2) Automatic Essay Scoring -
      • Overstability and Oversensitivity of AES systems
      • How to measure them?
      • Why does it arise?
      • Can we do something about it? How to cure it?
      • Submitting to Linguistics Journal

23 of 29

24 of 29

25 of 29

26 of 29

Visemic Corruption

(visemes of a particular type being corrupted and requiring regeneration)

Intra Word Corruption

(Corruption of frames within the occurrence of a large word)

Inter Word Corruption

(Corruption of frames across word boundaries)

27 of 29

Extra Curriculars

  1. Got selected for and attended 8th Heidelberg Laureate Forum
  2. Reviewed papers for Neurips 2020, IJCAI 2020, AIRE Journal, BigMM 2020
  3. Delivered tutorial on the topic “Marriage of Computer Vision, Speech, and Natural Language” at ICVGIP 2020 with Dr Rajiv Ratn Shah
  4. Collaborating with and helping NGOs such as Nyaaya through Artificial Intelligence and Language Processing
  5. Mentored 2 high school students on how to conduct research at Summer STEM Institute 2021

28 of 29

Overall Learnings

  1. In first year: Lot of emphasis on learning new tools. Second year should be utilized to transition from this phase to a more focused research mindset.
  2. Estimate your capabilities appropriately. I had to drop 2 courses because of mismatch between time commitment and course requirements
  3. Compartmentalizing projects and reducing uncertainty around objectives.
  4. Theory vs Useful Theory
  5. The secret art of designing good experiments
  6. Maintaining a scientific notion about Machine Learning.

29 of 29

Future Plans

  1. Courses in upcoming semesters: Advanced Deep Learning, Meta Learning, ASP, Theory of DL
  2. Prepare myself mathematically for more rigorous exploration of my focus area.
  3. Continuing and completing the projects listed in the initial slides, especially, my focus areas would be:
    1. Integrating creativity with data
    2. Understanding black box nature of deep learning algorithms