I have led the building of a cross brand behavioural data foundation, and the most important thing I learned is that it is very easy to build one that produces excellent reporting and changes nothing.
Reporting is a by product. The point of a data foundation is activation, which means being able to recognise the same customer across channels and do something differently as a result, in the moment, without a person in the middle.
The test that separates the two
Can you take a behaviour observed in one place and act on it in another, this week, without an export?
If a customer browses a category on a phone, does the email they get tomorrow know that? If they contact support with a problem, does the marketing automation pause? If they are already a subscriber, does the acquisition campaign stop paying to reach them?
Those three questions have caught out every data project I have reviewed, including some very expensive ones.
If a data foundation only improves reporting, it was a reporting project with a larger budget.
What makes it hard, and it is not the technology
- Identity. Recognising one person across devices, channels and a physical store is where most of the difficulty lives.
- Consent. What you may use, for what purpose, and being able to prove it. This is design work, not a checkbox.
- Ownership. When the data spans brands or business units, someone has to be allowed to decide.
- Definitions. Two teams counting a visit differently will produce two truths and one long argument.
- Governance of tagging, which is dull, ongoing, and the reason foundations decay within a year.
Notice that four of those five are organisational. The technical build is the part with a plan and a vendor. The rest is why these projects run long.
Privacy as a design input
Building consent and purpose limitation in from the start is cheaper than retrofitting it, and it produces a better product. A team that knows what it may use stops arguing about what it might use and starts building.
It also survives regulatory change, which on a five year horizon is not optional.
And then AI has something to work with
Every conversation about personalisation and machine learning eventually arrives here. Models are not the constraint. The constraint is whether there is a trustworthy, consented, unified view of behaviour for a model to act on, and whether any channel is wired to receive the output.
Build that and the AI conversation becomes straightforward. Skip it and no model will save the project.