1 of 31

Day 5: Uncertainty and chaos are helpful

What does uncertainty buy you?

Choice.

How do you choose?

Data!

2 of 31

What is biomed?

Multi-scale,

Cryptic,

Multi-source

Personal

3 of 31

Fields of study: component view

Any list of “parts” will likely be incomplete (many parts to count, but also a classification problem!)

A list of parts ignores interactions and dynamics

Integumentary

(connective tissues)

Skeletal

(bones)

Excretory

(digestion)

Circulatory

(blood)

Nervous

(nerves, brain)

Respiratory

(breathing)

Reproductive

(sex)

Cells

genes

4 of 31

Fields of study: Pathology view

(is death health?)

5 of 31

Fields of study: systems view

6 of 31

How do we study biology

Topic

Scale

Method

Personal

Niche

3. Cancer

2. nm-mm

1. Microscopy

Cancer cell

migration

1. Cancer

2. SNPs

3. genetics

Cancer

mutations

Data

What you care most about

constrains your options

7 of 31

Algorithms

(from Al-Kwarizmi, in honor of “the man of Khiva” a 9th century writer on algebra (al jabara – the uniting of broken pieces))

  • Known and Reproducible: Inputs, Processes, Outputs

  • Explainability in modern algorithms: like randomness or chaos, a “thing” can be deterministic but not known (or knowable)

  • High dimensionality / complex data doesn’t have to 🡺 black box

Take thing(s)

Do thing(s)

Get thing(s)

Real world constraints!

Compliance, cost, reliability, etc…

Real world metrics!

Cost, usability, information, etc…

8 of 31

What’s a model?

Combined factors that fit data and predict future observations.

Mathematical model – hypothesis driven (by necessity)

Machine learning model – computation driven

(hypothesis driven if you feel like it)

Dependent on two sets of assumptions:

1) what features carry information (variables)

2) what are relationships (equations)

9 of 31

Unsupervised learning

  • Let the data drive – not training towards a label
  • Clustering on data-defined distances
  • Dimension reduction
  • Hypothesis generation

Supervised learning

  • Let the label drive –train for a specific target
  • Iterating to identify manifold of separation
  • Ideally only independent dimensions
  • Hypothesis testing / deployment (mostly)

10 of 31

40-50 yo

Men

Women

-/+ Hypertension

How can we explore data landscapes?

Add up features

Add in time

Include broadly

Guided approach:

11 of 31

What is bias?

Prediction error in a given test compared to training

Note: this is not = variance or effect size

What is systematic bias?

Bias that is predictable due to condition

Note: this can arise from variance or effect size

Why is this especially relevant

in biomedicine?

Complexity 🡺 unpredictable clustering

Note: cluster identities act like other identities

12 of 31

  • Story game

13 of 31

  • break

14 of 31

Types of error

  • Random error, systematic error, transient error

Imprecision

White noise

Uniform, unaccounted variance

(remember, stochastic ≠ random)

Uncalibrated

Non-white noise

Unaccounted variance

Interference

Spikiness

Unaccounted events

15 of 31

Types of error

  • Can you have systematic random error?

  • What about transient systematic error?

  • What about transient random error?

16 of 31

Types of Shift

Expected Data X with Labels Y, and actual (current) Data X and Labels Y

  • Label shift p(Y) vs p(Y)

  • Covariate/Feature shift p(X) vs p(X)

  • Concept shift p(Y|X) vs p(Y|X).

˜

˜

˜

˜

˜

˜

17 of 31

Types of variance

  • We want to accurately measure inaccuracy

  • When is variance error, and when is variance not error
    • Explainable

    • Unexplainable

  • How can you tell?

18 of 31

What is chaos? Complexity

Why a butterfly?

Small things add up.

“That thou canst not stir a flower

Without troubling a star”

19 of 31

Chaos theory

  • Chaos theory: study/modeling dynamical systems with deterministic components.
  • NOT: randomness in step functions
  • BUT: hyper dependence on initial conditions
  • Can be patterned (deterministic), but not precisely predictable

  • Turbulent flow
  • Heat convection
  • Jointed pendula
  • All of Biology!

20 of 31

Recall: you can’t compute true randomness…

Nature is complex (and stochastic), but…

We should be able to predict deterministic systems, dang it! Deterministic: no “random” input (no dice involved in the game) Enter: “the three body problem!”

Lyapunov time

MY

21 of 31

Lyapunov time

How do we quantify the extent to which a system is chaotic?

Calculate divergence across iterations in a system with 2 starting conditions: X vs. X+e

D amplitude ∝ time and frequency

What’s big for your system?

22 of 31

23 of 31

Lorenz Attractors as example chaotic systems

  • Kinda oscillating?
  • No fixed period
  • Non-deterministic results
  • YET: bounded

What’s a bioengineer to do?

24 of 31

When chaos / entropy is a feature

Burks et al., PLoS Digital Health, 2025

25 of 31

Featurizing complexity – signal processing but not frequency

Lower overall complexity

Higher overall complexity

26 of 31

Entropy at different timescales changes in age & condition

27 of 31

In complex systems we can still measure stuff.

  • Can we tell who might have diabetes before they come to the clinic?
  • Coupled oscillators suggest we might measure changes coupled to metabolic health.
  • So we hypothesize: wearables let us predict diabetes without using glucose.
  • We can’t know 100%, can’t understand all possible causes,
  • …but we can find non-random predictions.
  • If predictions work, we support our hypothesis (and help people). But we’d still like to know how!

28 of 31

Argue from numbers.

Leave things out, one at a time, with the full model as a positive control.

Show the cost.

Whatever costs you the most is the most valuable.

29 of 31

Consider independence too

If feature overlap carry the same information, they are redundant

Including redundant features costs compute, adds noise

independence, “orthogonality”, correlation, co-information, etc.,

Past glucose data

Can we predict

future glucose?

30 of 31

LLMs: biology is more complex than language

LLMs on language prediction:

Sentences are complicated but not complex.

We know the full state of every sentence.

Sentence accuracy is 100% verifiable upon creation.

LLMs on biological prediction:

Organisms are complex but bounded by inputs.

We have very few stable ground truths.

Prediction accuracy can never be fully evaluated (in many cases)

“you have COVID” can

“you are aging well” cannot

Because of bounding, predictions on biological systems are not random or deterministic, but must be probabilistic

Bounded chaotic systems are still chaotic

31 of 31

HW 5. combine projects and compare

You all have features related to estrus detection.

Merge with another group.

Tally how many right and wrong detections you each make.

Are your features redundant or additive?

Make a rule for voting on or combining your features to help improve on your individual detectors.