Salesforce Data

Structured and unstructured data, the six dimensions of quality, and why machine learning does not average bad records out, it learns them as the truth.

Talk to an Expert

Pillar two of five

What data readiness means

Most Salesforce orgs are not short of data. They are short of data anyone is willing to make a decision on. Those are very different problems. Collecting more data solves the first and worsens the second.

Data is pillar two because everything after it inherits it. Automate a process on unreliable records and you automate the unreliability. Integrate two systems that disagree and you propagate the disagreement at speed. Ground an AI agent in it and you get answers delivered with confidence and no basis.

Book a Data Preparedness ReviewWhat we look at

Where data sits in the sequence

Data is pillar two because the three pillars after it all inherit whatever state it is in.

Why it matters

The cost of data nobody trusts

Poor data rarely announces itself. It shows up as behaviour, and by the time you can see the behaviour the cause is several steps back.

  • People build their own version. The spreadsheet on someone’s desktop reflects a system that has been unreliable before. It is a rational response to a system they cannot rely on, and it takes the knowledge out of the business with them when they leave.
  • Reporting becomes a negotiation. Two teams arrive with different numbers and the meeting is spent reconciling rather than deciding.
  • Automation gets switched off. A workflow that fires on bad data does visible damage quickly, and the response is almost always to disable the automation rather than fix the input.
  • Adoption quietly falls. Nobody announces they have stopped trusting the CRM. They just start checking elsewhere first.
  • AI amplifies all of it. That amplification is why this pillar now carries a budget.
Where it sits

Everything above this inherits the state of your records

Data is the second foundation, and it is the one that determines whether anything built on top can be believed. Process, integration and AI do not improve data. They act on it, faster and at greater scale than a person would.

Augment
05AI
Expand
04Integration
Establish
01Security02Data03Process

A process automated on unreliable records produces unreliable outcomes on a schedule. An agent grounded on them answers with total confidence and no basis. Neither is a failure of the layer above.

Understanding it

Structured and unstructured data

Every business holds two kinds of data, and until very recently only one of them was worth anything operationally.

Structured data is the record: fields, values, picklists, relationships. An account name, a close date, an amount. It is queryable, reportable and countable, and it is what a CRM was built to hold.

Unstructured data is everything else your business produces: the call recording, the email thread, the meeting transcript, the signed contract, the support conversation, the note someone typed at 6pm. It carries most of the actual meaning and almost none of the structure.

For thirty years unstructured data was effectively dead weight. You stored it for compliance, you searched it when there was a dispute, and no system could read it at scale. That is no longer true, and it is the single most important shift for any business planning AI work. Language models read unstructured content natively. The call transcript that was archive material is now a queryable asset.

Two consequences follow, and most businesses have addressed neither. First, the value of what you are already sitting on has gone up sharply, and nobody has re-audited it. Second, so has the liability. Content nobody ever classified, because nobody could read it, is now readable, which makes it discoverable, exportable and in scope for every question your Security policy has to answer.

The principles

Six dimensions of data quality

“Data quality” is too vague to act on. These six are the working definition, and each fails in a way you will recognise.

DimensionThe questionHow it fails
CompletenessIs the field populated where it matters?Industry blank on 40% of accounts, so segmentation covers less than it appears to
AccuracyDoes the value match reality?Job titles four years stale, addressed to people who left
ConsistencyDoes it agree across systems?Three spellings of the same company, counted as three customers
TimelinessIs it current enough to act on?A last-contact date that predates the current account manager
UniquenessOne record per real thing?Duplicate contacts inflating reach and triggering the same email twice
ValidityRight format, right range?Free text where a picklist belongs, so nothing can be grouped

These are not equally expensive to fix. Validity and uniqueness are largely technical. Accuracy and timeliness are behavioural, which makes them process and ownership questions as much as data ones.

The multiplier

What machine learning does to bad data

This is the part that changes the business case, and it is not intuitive.

Traditional reporting is forgiving of bad data. A few wrong records barely move a total, and an average absorbs noise. You can run a business on a report that is 95% right and never notice the missing five.

Machine learning does the opposite. It does not average errors out, it learns them and then applies what it learned, at scale, with no indication of doubt. Train a model on a CRM where reps closed deals as “Other” because the picklist was awkward, and the model learns that Other predicts success. It will then tell you so, confidently, in a dashboard, for as long as you let it.

The same applies to grounding an agent. An assistant answering from partial records does not say “I only have half of this”. It answers as though it has all of it. Confidence is a property of the interface, not of the data.

Data work therefore comes before AI work rather than alongside it. Not as a matter of tidiness, but because the failure mode changes from a report that is slightly off to a system that is wrong at speed and believed.

Modern practice

How data is leveraged now

The tooling has moved a long way in two years, and most of the movement has been about using data where it lives rather than copying it somewhere central first.

  • Harmonisation and identity resolution. Deciding that the record in the ERP, the one in the CRM and the one in the marketing tool are the same customer, and doing it as a rule rather than a clean-up project.
  • Zero copy. Data Cloud, now positioned as Data 360, can connect to Snowflake, Databricks, BigQuery and Redshift without physically moving the data. You get harmonisation, identity and activation without a second copy to secure, reconcile and pay for.
  • Grounding for AI. Unified, permissioned data is what an agent reasons over. This is the direct line between this pillar and the last one, and it is why the two are usually funded together.
  • Analytics where the decision happens. Tableau surfacing the answer inside the workflow rather than in a report somebody opens on a Monday.
  • Governance as a live control. Classification, retention and access reviewed on a cadence, not written once and filed.

The scale is not theoretical: Salesforce reported 32 trillion records ingested into Data Cloud in a single quarter of FY2026, of which 15 trillion arrived through zero-copy connectors, up 341% year on year. Whatever else that indicates, it says the industry has decided that copying data is the thing to avoid.

The offer

Do you know what state your data is in?

Most teams have a feeling about it. Very few have a measurement. The Data Preparedness Review turns the feeling into a number, so the conversation about what to fix can start from something specific.

It is a fixed-price engagement, priced on the size and complexity of your estate and agreed before anything begins. Five stages, running left to right:

01ProfileCompleteness, duplication, staleness and validity, measured object by object against your real records.
02AI readinessWhat could safely ground an agent today, and what would need work first.
03UnstructuredWhat sits in notes, files, transcripts and attachments, and whether it is classified.
04CausesWhere the gaps are entering. Usually a form, a permission, or an integration without validation.
05PlanWhat to address, in what order, with the remediation quoted separately.

You get a document you can act on and a conversation about what it means. Some of what comes back is usually process rather than technology, and we will point that out, because cleaning records without adjusting what produces them tends to hold for about six months.

If knowing would be useful, book one. It takes a short conversation to scope, you will have a price before any work starts, and you are under no obligation to do anything with what we find. Plenty of reviews conclude that the data is in better shape than the team expected, which is worth knowing too.

Book a Data Preparedness Review

Where this leads

Data is the pillar that pays for the others

On its own, good data is an argument. Connected to what comes next, it is a return.

  • Into Process. A defined process with reliable data is automatable. Either half missing and it is not.
  • Into Integration. Agreeing which system owns which field is a data-governance question that happens to be delivered as an integration.
  • Into AI. Grounding is what the whole thing rests on. An agent is only as good as what it can see and trust.
  • Back to Security. Classification decides what may be exported, surfaced or reached by an agent, so data work is security work.
Getting started

How to begin with data

Depending on how much you already know about the state of things.

If you suspect a problemData Preparedness ReviewA fixed-price diagnostic that measures it rather than debating it.
If AI is the driverReadiness assessmentSpecifically what would ground an agent safely, and what would not.
If you are not sure where to startLandscape ReviewBroad across all five pillars, and it will tell you if data is not your biggest problem.
Ready to talk

Find out what state your data is in

Most businesses are somewhere between “better than we feared” and “that explains a lot”. Either way it can be measured, which costs less than estimating.

Book a Data Preparedness Review

Salesforce data: frequently asked questions

What is Salesforce data quality?

The degree to which the records in your org can be relied on for a decision. In practice it breaks into six measurable dimensions: completeness, accuracy, consistency, timeliness, uniqueness and validity. Each fails differently and each has a different fix, which is why a single “clean the data” project rarely works.

Why does data quality matter more now than it did?

Because machine learning does not tolerate bad data, it learns it. Traditional reporting averages errors out, so a report that is 95% right is usually fine. A model trained on flawed records learns the flaw and applies it at scale, presented with no indication of doubt. The failure mode changes from slightly wrong to confidently wrong.

What is the difference between structured and unstructured data?

Structured data is the record: fields, values, picklists, relationships. Unstructured data is everything else, such as call transcripts, emails, contracts, notes and support conversations. Until recently unstructured data could not be read at scale, so it was stored rather than used. Language models changed that, which raised both its value and its governance risk.

What is zero copy, and do we need it?

Zero copy lets Data Cloud, now positioned as Data 360, connect to a warehouse such as Snowflake, Databricks, BigQuery or Redshift and use the data where it sits, rather than duplicating it. You get harmonisation, identity resolution and activation without a second copy to secure and reconcile. Whether you need it depends on whether your data already lives in a warehouse.

What does a Data Preparedness Review involve?

We profile the org object by object for completeness, duplication, staleness and validity, assess what could safely ground an AI agent, look at what sits in your unstructured estate, trace where poor data is entering, and produce a prioritised plan. It is fixed price, scoped to the size and complexity of your estate, and agreed before we start. Remediation is quoted separately.

Can you fix the data, or only report on it?

Both. The review tells you what is wrong and what it is costing you; the remediation is a separately agreed work package. We would rather you saw the findings and chose what is worth fixing than commit to a clean-up before anyone knows the scale of it.

Get in Touch