1 of 23

48th Transport Modellers Forum Meeting

18/04/2024

Tills, Tensors & Trajectories

A Location Planning Perspective on Mobility Data and  Machine Learning

2 of 23

AGENDA

  • Location Planning

  • Mobility Data

  • Machine Learning

  • Location Planning & Mobility Data & Machine Learning

2

3 of 23

ABOUT GEOLYTIX

4 of 23

WHAT WE DO

£££

£££

4

5 of 23

WHAT WE DO

NEW STORE FORECASTS

CANNIBALISATION

COMPETITOR IMPACTS

SCENARIO PLANNING

NETWORK STRATEGY

EXTENSION FORECASTS

5

6 of 23

WHO WE DO IT FOR

6

7 of 23

Mobility Data

8 of 23

MOBILITY DATA

who? 3fe33ed0-610a-11ed-9b6a-0242ac120002

when? 2022-11-10T15:14:34+00:00

where? 51.36855491432777, 0.3822254854173581

--------------------------------------------

source: 610a3ed0-610a-11ed-9b6a-11edac120002

horizontal deviation: 24m

...

  • snapshots, not trajectories!
  • use appropriate reference data to adjust long-term trends
  • most characteristics of a mobility feed change over time
  • but some do not:

8

9 of 23

Machine Learning

10 of 23

MODEL TYPES (SELECTION)

Gravity

Increasing Complexity (Generally)

Scorecard

Linear

Regression

Waterfall

NNLS*

S Cubed ‘Spatial’

Analogue

* Non-negative Least Squares

Machine

Learning

10

11 of 23

Risk mitigation in the context of small sample sizes:

  • Oversampling: When dealing with imbalanced data (more data for some store formats than others), we use oversampling to create a more balanced training set.

  • Stratified Splitting: We avoid random splits for training and testing. Instead, we use stratified sampling to ensure a good spread of different store formats and location types in both sets.

  • Hyperparameter Tuning: We tune hyperparameters using techniques like random search with cross-validation to optimize model performance.

ML CHALLENGES AND SOLUTIONS

SAMPLE SIZE

RISK

Healthcare

Financial Services

Location Planning

CRM

A/B Testing

Scientific Research

11

12 of 23

INTERACTION MODELS VS “FLAT” MODELS

  • interaction models: stores (own and competition) interact on a demand level

  • flat models: store x feature matrix

    • manual scorecards

    • (constrained) regression models

    • gradient boosting algorithms, e.g. XGBoost, LightGBM

  • decision usually driven by whether external/transient factors deemed more salient than interaction between stores/competitors

vs

12

13 of 23

  1. Gravity Models
  2. Market Share ML Models

  • we model the flows of consumer demand from the smallest level of administrative geography to Retail Places or individual stores

  • Summing up these flows gives a robust and objective means of forecasting potential sales

  • The key underlying principle is consumers tend to behave rationally and select destinations that are more convenient and more attractive (greater choice)

SPATIAL INTERACTION MODELS

13

14 of 23

Where Machine Learning & Mobility Data (can) meet

15 of 23

  • Making meaningful input variables, rather than letting an ML model figuring out biases and coverage issues, usually works out better for us (GI/GO)

  • Consistency between what clients see and what goes into their models

INPUT VARIABLES

Footfall

Traffic

15

16 of 23

  • Calculating catchments helps us understand the role of a store/town and its attractiveness, how it overlaps with nearby stores/towns

  • Using customer data and/or modelled catchment data we map the flow of customer spend allowing us to generate ‘natural’ primary and secondary catchments

  • Catchment creation and analysis plays a crucial role in optimising a store network

  • The Retail Place type, catchment extents and coverage (household, population, demand, target customers) can all be used to help determine the importance of having a store in a given area

CATCHMENTS CREATION

16

17 of 23

  • Distance decay reflects how far people are willing to travel to a Retail Place, recognising that people in rural areas are often more willing to travel further than those people living in a city centre, whose provisions are typically met within a small distance from home

  • Our distance decay values are defined at OA level using the urbanity score of the OA

DISTANCE DECAY MODELLING

Penetration Level

17

18 of 23

  • Example: catchment overlap measure

  • As a standalone model, or

  • inbuilt in spatial components of feature matrix

  • leave transient model components untouched

RESIDENTIAL IMPACTS

NEW STORE

Nearest Catchments

Nearest Catchments

18

19 of 23

  • Interaction Surfaces Product

  • Probability of a person showing up in other locations on the same day

  • Measure of shared transient population potential

TRANSIENT IMPACTS

19

20 of 23

NETWORK BLUEPRINTS

20

21 of 23

The Future

22 of 23

site vs rest of the world (flat model)

site vs anything else (interaction model)

anything vs anything

WHAT’S NEXT

Top-down

(e.g. Graph Neural Nets)

Bottom-up

(e.g. Causally informed synthetic populations)

22

23 of 23

48th Transport Modellers Forum Meeting

18/04/2024

THANK YOU