1 of 37

Seminar/Proseminar “Novel and non-mainstream advances in Data Science”Winter 2021��Online Kick-Off Meeting - 21.10.2021 10:00�Online Room: tba

Topic specific meetings if possible offline / up to the Supervisor

KIT – Die Forschungsuniversität in der Helmholtz-Gemeinschaft

IPD Böhm - Lehrstuhl für Systeme der Informationsverwaltung

www.kit.edu

IPD Böhm

1

2 of 37

Agenda

Motivation

Requirements for passing the seminar

“How To” Guide

    • Finding papers
    • Reading papers
    • Writing report
    • Formatting report

Topic presentation

Topic assignment

IPD Böhm

2

3 of 37

MOTIVATION

IPD Böhm

3

4 of 37

Motivation

  • Imagine you have just started working at �R&D department of a big company using �Information Systems on a daily basis.

  • Your boss asked you to prepare a report �that will be used to evaluate possible improvements of the core technology.

  • We provide you with example topics and promising papers to get you started.

Effective communication in research and business is very similar!

IPD Böhm

4

5 of 37

REQUIREMENTS

IPD Böhm

5

6 of 37

Requirements

  • Postpone everything to the last moment
  • Be silent in presentation round
  • Ignore supervisor’s comments

  • Meet with advisor
  • Actively participate in all presentations
  • Meet all deadlines

Time

Choose �a topic (Today)

Propose�report structure

Present�your work

Prepare�final report

IPD Böhm

6

7 of 37

Requirements

Proseminar

(Bachelor)

Seminar

(Master)

Report length

min 10–12 pages

min 12–15 pages

Extra references

min 2

min 4

Language

(report & presentation)

English

Presentation

20–25 minutes (~17 slides)

IPD Böhm

7

8 of 37

Requirements

Not a copy-paste of papers’ abstracts but a consistent story:

  • Summarizes
  • Evaluates
  • Compares
  • Puts in the context

Your report

Papers you read

You

IPD Böhm

8

9 of 37

Important Dates

28.10.21

Final registration deadline

  • Prüfungsnummer Seminar: 7500133
  • Prüfungsnummer Proseminar: 7500134

28.11.21

Submit report structure and literature list

28.11–17.12.22

Discuss report structure and literature list with the supervisor

21.01.22

Submit presentation draft (1 week before presentation)

27.01 and,

04.02.22

Presentation round

03.03.22

Submit final report

IPD Böhm

9

10 of 37

“How to” Guide

IPD Böhm

10

11 of 37

Finding Relevant Papers

1. Search the paper by its title or keywords

2. Look if the papers citing your or related papers are relevant

3. Look at previous work featured in your paper

IPD Böhm

11

12 of 37

Reading and Summarizing (by Andrew Ng)

Read in multiple passes:

  • Title + abstract + figures
    • worth to continue?
  • Intro + conclusions + figures + quickly glance rest
  • Read but skip math
  • Read whole but skip parts that don’t make sense to you

Summarize by answering questions:

  • What do authors try to accomplish?
  • What are the key elements of the approach?
  • What can you use?
  • Which related work is relevant for you?

twitter.com/andrewyng

After reading 5–20 papers:

basic area understanding, able to apply algorithms

IPD Böhm

12

13 of 37

Report Writing

Writing process

  • 70% Prewriting (organize information)
  • 10% Writing the first draft
  • 20% Revision (check and improve)

Use a good style:

Learn style:

“This paper reviews quantum physics.”

The aim of this paper is to provide a review of the basic principles of quantum physics.

IPD Böhm

13

14 of 37

Formatting Report: LaTeX

Takes longer to write than with Word,�but looks more professional and clean.

Take a look at examples here:�http://liinwww.ira.uka.de/~thw/vl-latex-co/

Offers efficient referencing of the literature, plots, and equations.

IPD Böhm

14

15 of 37

TOPIC PRESENTATION

IPD Böhm

15

16 of 37

(T1: Review of dependency data sets and data generation (BA))

T2: Review of dependency data generation from Graph Models (BA or MA)

(T3: Visualization of complex data Dependencies (BA or MA))

T4_1: (Active) Learning of complex data Dependencies (BA or MA)

T4_2: (Active) Learning of Causation (BA or MA)

T6: A review of regression models with uncertainty estimates

T7_1: Review of Surrogate Model based optimization with active search strategies Optimal_vs_robust_process_parameters�T7_2: Review of Surrogate Model based optimization with active search strategies scenario_discovery�(T7_3: Discovering governing equations from data (BA or MA))

(T7_4: Sample efficient reinforcement learning (MA))

Bela Böhnke

IPD Böhm

16

17 of 37

T8: Annotation Inconsistencies in Image Datasets (BA or MA)

T9: Learning Taxonomies From Data (BA or MA)

T10: Using Taxonomies to improve Machine Learning Tasks (BA or MA)

T11: Detecting Taxonomy Inconsistencies (BA or MA)

T12: Bandit Algorithms with Domain Knowledge (BA or MA)

T13: ML Methods for Solving Differential and Difference Equations (BA or MA)�T14: Supervised Uncoupled Feature Extraction (BA or MA)

T15: Unsupervised Uncoupled Feature Extraction (BA or MA)

Pawel Bielski

Vadim Arzamasov

IPD Böhm

17

18 of 37

Overview / Motivation:

[R. Manhaeve et al., ‘DeepProbLog: Neural Probabilistic Logic Programming’, 2018]�[S. Badreddine et al. ‘Logic Tensor Networks’, 2021]

Task: 1,...,7

IPD Böhm

18

19 of 37

T1, 2: Review of dependency data sets and data generation (with Graph Models)

Mooij, Joris M., ... “Distinguishing Cause from Effect Using Observational Data: Methods and Benchmarks.”

Arzamasov, Vadim, and Klemens Böhm. “REDS: Rule Extraction for Discovering Scenarios,”

Fouché, Edouard, and Klemens Böhm. “Monte Carlo Dependency Estimation,” 2019, 12. https://doi.org/10/gjr349.

  • For testing the mentioned knowledge discovery tasks, test data with ground truth is required
  • Review the data used in these tasks:
    • What data is useful for testing the tasks?
    • What data sets do exist?
    • What data generation methods do exist?
    • What are the pros/cons of different data sources?
    • Which dataset is good/bad for testing certain aspects?
    • Can we categorize accordingly?�
  • Implement an Oracle that unifies different data acquisition strategies (so a scientist can use it for testing algorithms)

IPD Böhm

19

20 of 37

T3: Visualisation of complex data Dependencies

[m.A., Dependency Graph - an overview“, o. J. https://www.sciencedirect.com/topics/computer-science/dependency-graph]

Krzywinski, M., I. Birol, S. J. Jones, and M. A. Marra. “Hive Plots--Rational Approach to Visualizing Networks.”

How to visualize and explore Dependencies?

  • Relational Dependency Networks
  • Dependency graphs
  • Bayesian network

A goal of knowledge discovery is to present the knowledge in a way a human can understand it.

IPD Böhm

20

21 of 37

T4: Learning Dependencies and Causation

Mooij, Joris M., ... “Distinguishing Cause from Effect Using Observational Data: Methods and Benchmarks.”

Fouché, Edouard, and Klemens Böhm. “Monte Carlo Dependency Estimation,” 2019, 12. https://doi.org/10/gjr349

Yuan, Changhe. “Optimal Algorithms for Learning Bayesian Network Structures:,”

Koller, Daphne, and Nir Friedman. Probabilistic Graphical Models: Principles and Techniques

  • Review Algorithms
    • Learning complex dependencies from data
    • Find causal direction from data

IPD Böhm

21

22 of 37

T6: A review of regression models with uncertainty estimates

Mullachery, Vikram, Aniruddh Khera, und Amir Husain. „Bayesian Neural Networks“, ArXiv: 1801.07710

Müller Peter et. al.. Nonparametric Bayesian Inference, 2013. https://doi.org/10.1214/cbms/1362163742.

Wilson Andrew Gordon et. al., Bayesian Deep Learning and a Probabilistic Perspective of Generalization, ArXiv:2002.08791

In current ML one tries to optimise for one goal directly by taking one “path” (frequentistic).

The goal of this seminar is to explore different methods for estimating such promising paths for learning with uncertainty

It can be beneficial to start with multiple paths, and only pursue the ones that are most promising, i.e., which provide most information.

IPD Böhm

22

23 of 37

T7: Review of Surrogate Model based optimization with active search strategies

Arzamasov, Vadim, and Klemens Böhm. “REDS: Rule Extraction for Discovering Scenarios,”

Zimmerling, Clemens, Patrick Schindler, Julian Seuffert, and Luise Kärger. “Deep Neural Networks as Surrogate Models for Time-Efficient Manufacturing Process Optimisation.”

  • Provided is one or multiple starting papers
  • Find other approaches that solve similar problems
  • Explain the basic intuition behind each approach
  • Compare the different approaches theoretically:
    • What are categories of approaches?
    • What do they differently?
    • Why are they good/bad?

IPD Böhm

23

24 of 37

T7_3: Discovering governing equations from data

[Sparse identification of nonlinear dynamics, Steven L. Brunton et. al., 2016, DOI: 10.1073/pnas.1517384113]

[Data-driven discovery of coordinates and governing equations, Kathleen Champion et. al., 2019, DOI: 10.1073/pnas.1906995116]

One method to model dependencies are physical equations.�They describe the interaction between influencing variables and influenced variables.

The goal of this seminar is to explore different methods for finding such physical equations from data, and find good criteria for evaluation of equation quality.

IPD Böhm

24

25 of 37

T7_4: Sample efficient reinforcement learning

Buckman Jacob et. al., Sample-Efficient Reinforcement Learning with Stochastic Ensemble Value Expansion

Deisenroth Marc Peter et. al, PILCO: A Model-Based and Data-Efficient Approach to Policy Search

In RL one has to explore the environment, but such exploration can be expensive.

So a goal in RL is to explore only as much as necessary, but also not to little. This can be done by bayesian learning.

The goal of this seminar is to explore different applications of bayesian learning in RL.

IPD Böhm

25

26 of 37

T8: Annotation Inconsistencies in Image Datasets

26

21.09.2020

[Andrew Ng, “A Chat with Andrew on MLOps: From Model-centric to Data-centric AI”, 2021]

Labeling instruction: use bounding boxes to label the position of iguanas

In order to increase accuracy from 50% to 60% we can clean the data or increase �the data size 3 times (from 500 to 1500).

www.aimino.de

MLOps: systematically improve data quality

What are annotation inconsistencies?��How do they arise?��What types exist?

�Create a taxonomy based on the literature, blogs or tutorials.

IPD Böhm

26

27 of 37

T9: Learning Causal Knowledge �Graphs from Text Log Data

27

21.09.2020

[ P, Aggarwal et. al., Localization of Operational Faults in Cloud Applications by� Mining Causal Dependencies in Logs using Golden Signals, 2020]

Dependency graphs captures the patterns from log data.

Example approach: model the log data as time series �and apply causal inference.

Pre-process

Log data

Infer taxonomy

AIOps: AI for IT (Cloud) Operations

Perform literature search to find methods for inferring causal knowledge graphs from temporal (text log) data.�

Create a taxonomy of methods

IPD Böhm

27

28 of 37

T10: Using Taxonomies to improve �Machine Learning Tasks

28

21.09.2020

[M. Elhamod et. al., “Hierarchy-guided Neural Networks for Species Classification”, 2021]

Features sharing

Biology taxonomy groups fish species into families.

Fish class prediction

Image

Fish family prediction

Yclass

Yfamily

IPD Böhm

28

29 of 37

T11: Detecting Taxonomy Inconsistencies

29

21.09.2020

[C. Yin et. al., "Domain Knowledge Guided Deep Learning with Electronic Health Records,", 2019]�[W. Ceusters et. al., “Mistakes in Medical Ontologies: Where Do They Come From and How Can They Be Detected?”, 2004]

Clinical Risk

Example: Medical Diagnosis Prediction Models -

Incorporating Expert Knowledge

RNN

Symptom Embeddings

Symptoms �per Visit

Knowledge Graph (DAG)

ICD9-785 Symptoms involving cardiovascular system

ICD9-785.5 Shock without mention of trauma

ICD9-785.52 Septic shock

attention weights

IPD Böhm

29

30 of 37

T12: Bandit Algorithms with Domain Knowledge

30

21.09.2020

[R. Singh et. al., “Multi-Armed Bandits with Dependent Arms”, 2020]

[S. Pandey et. al. “Multi-armed bandit problems with dependent arms”, 2007]

Bandit algorithms for sequential decisions.

Can be used for monitoring of complex systems, eg. Cloud centres.

Some approaches integrate domain knowledge.

What approaches are there?

IPD Böhm

30

31 of 37

T13 ML Methods for Solving Differential

and Difference Equations

x1

x2

xm

y

Solver

IPD Böhm

31

32 of 37

T13 ML Methods for Solving Differential

and Difference Equations

Solves / helps to solve differential equations

Predicts y given {x1,...,xm} values or vice versa

Solver

x1

x2

xm

y

xn1

...

x11

xnm

...

x1m

...

...

...

yn

...

y1

Solver

ML model

ML model

IPD Böhm

32

33 of 37

T13 ML Methods for Solving Differential

and Difference Equations

Literature:

  • Weinan, E., Han, J. and Jentzen, A., 2017. Deep learning-based numerical methods for high-dimensional parabolic partial differential equations and backward stochastic differential equations. Communications in Mathematics and Statistics, 5(4), pp.349-380.
  • Sirignano, J. and Spiliopoulos, K., 2018. DGM: A deep learning algorithm for solving partial differential equations. Journal of computational physics, 375, pp.1339-1364.
  • Bar-Sinai, Y., Hoyer, S., Hickey, J. and Brenner, M.P., 2019. Learning data-driven discretizations for partial differential equations. Proceedings of the National Academy of Sciences, 116(31), pp.15344-15349.
  • Cockayne, J., Oates, C., Sullivan, T. and Girolami, M., 2016. Probabilistic numerical methods for partial differential equations and Bayesian inverse problems. arXiv preprint arXiv:1605.07811.

IPD Böhm

33

34 of 37

T14–T15 Supervised / Unsupervised

Uncoupled Feature Extraction

f1=(x1-1)^2

f2=(x2-2)^2

  • min 4 methods
  • avoid
    • simple methods (PCA)
    • autoencoders
  • start from

Storcheus, D., Rostamizadeh, A. and Kumar, S. A survey of modern questions and challenges in feature extraction. 2015

...

1

x1

...

x2

1

...

0

0

y

...

0

f1

...

f2

0

...

1

1

y

IPD Böhm

34

35 of 37

TOPIC ASSIGNMENT

IPD Böhm

35

36 of 37

(T1: Review of dependency data sets and data generation (BA))

T2: Review of dependency data generation from Graph Models (BA or MA) Johannes N.

T3: Visualization of complex data Dependencies (BA or MA) Samuel B.

T4_1: (Active) Learning of complex data Dependencies (BA or MA) David N.

T4_2: (Active) Learning of Causation (BA or MA) Brandon S.

T6: A review of regression models with uncertainty estimates Jia D.

T7_1: Review of Surrogate Model based optimization with active search strategies Optimal_vs_robust_process_parameters Isabel A.�T7_2: Review of Surrogate Model based optimization with active search strategies scenario_discovery Jonas H.�(T7_3: Discovering governing equations from data (BA or MA))

(T7_4: Sample efficient reinforcement learning (MA))

Bela Böhnke

IPD Böhm

36

37 of 37

T8: Annotation Inconsistencies in Image Datasets (BA or MA) Niklas K.

T9: Learning Taxonomies From Data (BA or MA) Aleksandr E.

T10: Using Taxonomies to improve Machine Learning Tasks (BA or MA) Elena S.

T11: Detecting Taxonomy Inconsistencies (BA or MA) Sönke J.

T12: Bandit Algorithms with Domain Knowledge (BA or MA) Ola S.

T13: ML Methods for Solving Differential and Difference Equations Dmitrii S.�T14: Supervised Uncoupled Feature Extraction (BA or MA) Tilo S.

T15: Unsupervised Uncoupled Feature Extraction (BA or MA) Kevin H.

T16: Data Augmentation for Tabular Data Ufuk G.

Pawel Bielski

Vadim Arzamasov

IPD Böhm

37