What data readiness means
Most Salesforce orgs are not short of data. They are short of data anyone is willing to make a decision on. Those are very different problems. Collecting more data solves the first and worsens the second.
Data is pillar two because everything after it inherits it. Automate a process on unreliable records and you automate the unreliability. Integrate two systems that disagree and you propagate the disagreement at speed. Ground an AI agent in it and you get answers delivered with confidence and no basis.
Book a Data Preparedness ReviewWhat we look at
Where data sits in the sequence
Data is pillar two because the three pillars after it all inherit whatever state it is in.
The cost of data nobody trusts
Poor data rarely announces itself. It shows up as behaviour, and by the time you can see the behaviour the cause is several steps back.
- People build their own version. The spreadsheet on someone’s desktop reflects a system that has been unreliable before. It is a rational response to a system they cannot rely on, and it takes the knowledge out of the business with them when they leave.
- Reporting becomes a negotiation. Two teams arrive with different numbers and the meeting is spent reconciling rather than deciding.
- Automation gets switched off. A workflow that fires on bad data does visible damage quickly, and the response is almost always to disable the automation rather than fix the input.
- Adoption quietly falls. Nobody announces they have stopped trusting the CRM. They just start checking elsewhere first.
- AI amplifies all of it. That amplification is why this pillar now carries a budget.
Everything above this inherits the state of your records
Data is the second foundation, and it is the one that determines whether anything built on top can be believed. Process, integration and AI do not improve data. They act on it, faster and at greater scale than a person would.
A process automated on unreliable records produces unreliable outcomes on a schedule. An agent grounded on them answers with total confidence and no basis. Neither is a failure of the layer above.
Structured and unstructured data
Every business holds two kinds of data, and until very recently only one of them was worth anything operationally.
Structured data is the record: fields, values, picklists, relationships. An account name, a close date, an amount. It is queryable, reportable and countable, and it is what a CRM was built to hold.
Unstructured data is everything else your business produces: the call recording, the email thread, the meeting transcript, the signed contract, the support conversation, the note someone typed at 6pm. It carries most of the actual meaning and almost none of the structure.
For thirty years unstructured data was effectively dead weight. You stored it for compliance, you searched it when there was a dispute, and no system could read it at scale. That is no longer true, and it is the single most important shift for any business planning AI work. Language models read unstructured content natively. The call transcript that was archive material is now a queryable asset.
Two consequences follow, and most businesses have addressed neither. First, the value of what you are already sitting on has gone up sharply, and nobody has re-audited it. Second, so has the liability. Content nobody ever classified, because nobody could read it, is now readable, which makes it discoverable, exportable and in scope for every question your Security policy has to answer.
Six dimensions of data quality
“Data quality” is too vague to act on. These six are the working definition, and each fails in a way you will recognise.
| Dimension | The question | How it fails |
|---|---|---|
| Completeness | Is the field populated where it matters? | Industry blank on 40% of accounts, so segmentation covers less than it appears to |
| Accuracy | Does the value match reality? | Job titles four years stale, addressed to people who left |
| Consistency | Does it agree across systems? | Three spellings of the same company, counted as three customers |
| Timeliness | Is it current enough to act on? | A last-contact date that predates the current account manager |
| Uniqueness | One record per real thing? | Duplicate contacts inflating reach and triggering the same email twice |
| Validity | Right format, right range? | Free text where a picklist belongs, so nothing can be grouped |
These are not equally expensive to fix. Validity and uniqueness are largely technical. Accuracy and timeliness are behavioural, which makes them process and ownership questions as much as data ones.
What machine learning does to bad data
This is the part that changes the business case, and it is not intuitive.
Traditional reporting is forgiving of bad data. A few wrong records barely move a total, and an average absorbs noise. You can run a business on a report that is 95% right and never notice the missing five.
Machine learning does the opposite. It does not average errors out, it learns them and then applies what it learned, at scale, with no indication of doubt. Train a model on a CRM where reps closed deals as “Other” because the picklist was awkward, and the model learns that Other predicts success. It will then tell you so, confidently, in a dashboard, for as long as you let it.
The same applies to grounding an agent. An assistant answering from partial records does not say “I only have half of this”. It answers as though it has all of it. Confidence is a property of the interface, not of the data.
Data work therefore comes before AI work rather than alongside it. Not as a matter of tidiness, but because the failure mode changes from a report that is slightly off to a system that is wrong at speed and believed.
How data is leveraged now
The tooling has moved a long way in two years, and most of the movement has been about using data where it lives rather than copying it somewhere central first.
- Harmonisation and identity resolution. Deciding that the record in the ERP, the one in the CRM and the one in the marketing tool are the same customer, and doing it as a rule rather than a clean-up project.
- Zero copy. Data Cloud, now positioned as Data 360, can connect to Snowflake, Databricks, BigQuery and Redshift without physically moving the data. You get harmonisation, identity and activation without a second copy to secure, reconcile and pay for.
- Grounding for AI. Unified, permissioned data is what an agent reasons over. This is the direct line between this pillar and the last one, and it is why the two are usually funded together.
- Analytics where the decision happens. Tableau surfacing the answer inside the workflow rather than in a report somebody opens on a Monday.
- Governance as a live control. Classification, retention and access reviewed on a cadence, not written once and filed.
The scale is not theoretical: Salesforce reported 32 trillion records ingested into Data Cloud in a single quarter of FY2026, of which 15 trillion arrived through zero-copy connectors, up 341% year on year. Whatever else that indicates, it says the industry has decided that copying data is the thing to avoid.
Do you know what state your data is in?
Most teams have a feeling about it. Very few have a measurement. The Data Preparedness Review turns the feeling into a number, so the conversation about what to fix can start from something specific.
It is a fixed-price engagement, priced on the size and complexity of your estate and agreed before anything begins. Five stages, running left to right:
You get a document you can act on and a conversation about what it means. Some of what comes back is usually process rather than technology, and we will point that out, because cleaning records without adjusting what produces them tends to hold for about six months.
If knowing would be useful, book one. It takes a short conversation to scope, you will have a price before any work starts, and you are under no obligation to do anything with what we find. Plenty of reviews conclude that the data is in better shape than the team expected, which is worth knowing too.
Data is the pillar that pays for the others
On its own, good data is an argument. Connected to what comes next, it is a return.
- Into Process. A defined process with reliable data is automatable. Either half missing and it is not.
- Into Integration. Agreeing which system owns which field is a data-governance question that happens to be delivered as an integration.
- Into AI. Grounding is what the whole thing rests on. An agent is only as good as what it can see and trust.
- Back to Security. Classification decides what may be exported, surfaced or reached by an agent, so data work is security work.
How to begin with data
Depending on how much you already know about the state of things.
Find out what state your data is in
Most businesses are somewhere between “better than we feared” and “that explains a lot”. Either way it can be measured, which costs less than estimating.
Salesforce data: frequently asked questions
The degree to which the records in your org can be relied on for a decision. In practice it breaks into six measurable dimensions: completeness, accuracy, consistency, timeliness, uniqueness and validity. Each fails differently and each has a different fix, which is why a single “clean the data” project rarely works.
Because machine learning does not tolerate bad data, it learns it. Traditional reporting averages errors out, so a report that is 95% right is usually fine. A model trained on flawed records learns the flaw and applies it at scale, presented with no indication of doubt. The failure mode changes from slightly wrong to confidently wrong.
Structured data is the record: fields, values, picklists, relationships. Unstructured data is everything else, such as call transcripts, emails, contracts, notes and support conversations. Until recently unstructured data could not be read at scale, so it was stored rather than used. Language models changed that, which raised both its value and its governance risk.
Zero copy lets Data Cloud, now positioned as Data 360, connect to a warehouse such as Snowflake, Databricks, BigQuery or Redshift and use the data where it sits, rather than duplicating it. You get harmonisation, identity resolution and activation without a second copy to secure and reconcile. Whether you need it depends on whether your data already lives in a warehouse.
We profile the org object by object for completeness, duplication, staleness and validity, assess what could safely ground an AI agent, look at what sits in your unstructured estate, trace where poor data is entering, and produce a prioritised plan. It is fixed price, scoped to the size and complexity of your estate, and agreed before we start. Remediation is quoted separately.
Both. The review tells you what is wrong and what it is costing you; the remediation is a separately agreed work package. We would rather you saw the findings and chose what is worth fixing than commit to a clean-up before anyone knows the scale of it.