Customer Data Platforms (CDPs)
A CDP collects customer data from every touchpoint, resolves it into unified profiles, and makes those profiles available to other systems for activation.
CDP, CRM and Data Warehouse
These overlap enough to be confused constantly.
A CRM holds known-customer records, primarily for sales and service, entered largely by people.
A data warehouse stores large volumes of structured data for analysis. It is a storage and query layer, not an activation layer.
A CDP ingests behavioural and transactional data automatically, resolves identities across devices and channels into a single profile, and pushes segments out to marketing tools in near real time.
The distinguishing capabilities are identity resolution and activation. A warehouse can store the same data; it cannot natively push a segment into your ad platform.
Identity Resolution
This is the core function and the hardest part. The same person appears as an anonymous web visitor, an email subscriber, an app user and a purchase record, with no obvious link between them.
Identity resolution stitches these using deterministic matching where a shared identifier exists — an email address at login — and probabilistic matching where it does not, based on device, IP and behavioural signals.
Deterministic matching is accurate and limited. Probabilistic matching extends coverage and introduces error. Understand which your CDP uses and where the boundary sits, because merged profiles that are actually two people cause real problems.
Composable CDPs
Traditional CDPs store a copy of your data in the vendor's system. Composable CDPs sit on top of your existing warehouse, running identity resolution and activation without duplicating storage.
For organisations that already have a warehouse, composable approaches avoid maintaining two copies of the same data and keep governance in one place. This has become the dominant direction of travel.
Do You Actually Need One
The honest answer for many organisations is no.
A CDP earns its cost when you have genuinely fragmented data across many systems, enough customer volume for segmentation to matter, and marketing sophisticated enough to act on unified profiles.
It does not help when the underlying data is poor, when the organisation has no capacity to act on segments, or when the real problem is that two teams will not agree on definitions. A CDP implemented onto those problems produces an expensive version of the same problems.
Fix data collection quality first. Most of the value attributed to CDPs comes from the data hygiene work done during implementation rather than from the platform.
The Question That Decides Whether You Need One
CDP projects fail more often than they succeed, and the failures are visible in advance from one question: what will you do with a unified profile that you cannot do now?
If the answer is a specific activation — suppress existing customers from acquisition campaigns, trigger a message when behaviour across two systems combines, give support the purchase history at the moment of a call — the project has a destination. If the answer is *"have a single view of the customer"*, it does not, and the platform will be bought, populated, and left.
Three conditions that make one genuinely worth it: customer data genuinely scattered across systems that cannot talk to each other; a marketing team that needs to build audiences without engineering help, which is the honest core of the value proposition; and identity resolution that actually matters to your business, meaning the same person appears under different identifiers often enough to distort decisions.
Without at least two of those, a data warehouse and a scheduled export will do the same job for a fraction of the cost and effort.
Identity Resolution, and the Damage It Can Do
Identity resolution is what a CDP is really selling, and it is where the risk sits.
Deterministic matching — joining records that share a verified identifier such as an email or account ID — is reliable and limited to the customers who have given you one.
Probabilistic matching — inferring that two devices are the same person from behaviour, network and timing — extends coverage and is wrong some of the time in both directions. It merges two people into one profile and splits one person into two.
The consequences of a wrong merge are worse than most teams anticipate: one household member sees another's recommendations, a support agent reads the wrong purchase history, and in a regulated context you have combined two people's personal data. A subject access request against a probabilistically merged profile is a genuinely awkward thing to answer.
The defensible position: deterministic matching for anything a customer sees or that affects a decision about them; probabilistic only for aggregate analysis, clearly labelled as an estimate. And under a consent-based regime, be able to say which basis covers the joining itself — not just the collection of each part.