Robin Gower

Models put data to work

Your data has potential. An evidence-based decision-making process leads to better outcomes because you’re acting on verifiable facts not fallible opinions or risky guesswork.

In order to realise this potential you need to put your data to work by creating a model. I’m not talking about database schemas, I mean a digital twin that simulates your domain.

A data model provides a simplified view of the complicated reality of your services and products. You can use this abstract representation to breathe life into your data.

At it’s heart this is a set of mathematical functions describing the processes that generate your data. Exploring the parameters and feeding-in alternate inputs provides insights about alternate scenarios. We can use models to help us optimise our choices.

Models turn data into decisions

Data is valuable because it helps us make better-informed decisions.

A model let’s us leverage data for decision-making potential. With it, we progress along a scale and are able to answer increasingly complex questions:

Descriptive questions - what happened?

Basic data analysis provides a description of the state of the world e.g. “are sales rising or falling”?

To figure this out we might chart sales over time:

We can see variation but the overall trend is downward trend with a loss of 130 units per year.

Diagnostic questions - why did it happen?

If we add a causal model we can begin to ask diagnostic questions e.g. “why did sales fall”? The causal model describes how we expect one variable to affect or interact with another. You can think of it as a network graph where each node is a variable and each edge a causal influence.

flowchart LR
    Price --> Sales

This leads us to equations that relate the quantities in each variable conceptually. In order to describe the relationships concretely we need to quantify them. These values are generally unknown and not explicit in the data - we call these parameters.

In order to find appropriate values for our parameters we fit our model to the data, using it to figure out the values that best match what we observe. We can fit an estimate of sales at a given price.

The parameter values themselves provide insight. We can use them to quantify effects - for example we sold 156 fewer units for every additional Euro charged.

Predictive questions - what is likely to happen?

Once we’ve described the data generating process like this we can also start to make predictions about what might happen in alternate scenarios e.g. “how much will sales fall if we drop the price from €9 to €6”?

The model predicts an increase in sales from 1135 to 1603 units (41%).

Prescriptive questions - what should we do?

We can then use these predictions to guide decisions we face which prescribe a course of action. In this case we know decreasing price will increase sales but it will also reduce the amount we earn per sale. What we really want to know is “which price will maximise revenue”?

Combining the predictive model of sales for a given price with the knowledge that revenue is the product of units sold times the price per unit we can model revenue in terms of price.

Looking at the revenue curve we can find the maximum revenue is €10,329 and the price point we should choose to achieve that is €8.

Models put data to work

The revenue-maximising price point we’ve found wasn’t even in our data in the first place. Notice on the second chart above where we fit the price model, there’s no dot at €9.

We were able to extract this insight from the data with the help of a model that let us interpolate between the observations. We didn’t ask the data, we asked the digital twin.

Models let us turn historical data into predictions about future decisions.

All models are wrong, some are useful

All models are wrong, some are useful - George Box

This isn’t a particularly good model. A straight line like this isn’t a very realistic model for a demand curve as it can predict negative sales at very high prices! In reality we would explore other functional forms like a log-log curve.

Nevertheless the model is useful. It confirmed and quantified a causal relationship and made predictions about the future we could use to prescribe a course of action.

Start building a model

Although this model has flaws its crucial to start with something simple.

Build models up gradually ensures two things. First that we understand it - each of the pieces are easier to reason about in isolation. Second that we create a parsimonious model - i.e. one that is only as complicated as we really need.

So how do you build a model?

You need to start by defining the decisions that you or your users/ customers are facing. From there you can begin prototyping a model can prescribe a course of action. At that point you’ll know what data you need to build it.

If you’d like to explore how a model could help exploit the potential of your own data then please feel free to schedule a call.

If you’d like to learn more then you might like to read about my process for making data-driven decisions.