Coat of Arms

Coat of Arms

Enterprise data readiness is not a one-time cleanup project. It is the ability to provide data that is accurate, timely, understandable, permitted, secure, and reliable enough for a defined business use.

A large enterprise may need 18 to 36 months or longer to build reusable capabilities across multiple domains. A bounded AI use case can often reach sufficient readiness within 3 to 9 months.

The enterprise does not become AI-ready all at once. Specific data, decisions, and workflows become ready to a defined level.

Indicative timeline

These activities overlap. They should not be treated as sequential project phases.

ActivityIndicative timingExpected outcome
Use-case framingMonths 0 to 2Defined decision, baseline, required data, sponsor, and expected outcome
Data viability assessmentMonths 1 to 4Confirmed access, quality profile, ownership, constraints, and material gaps
Minimum governed data productMonths 2 to 8Authoritative sources, definitions, pipeline, quality rules, access, lineage, and monitoring
Pilot and workflow evaluationMonths 4 to 10Measured technical and business performance using representative data
Production hardeningMonths 7 to 15Security, privacy, resilience, support, cost control, and evidence
Domain scalingMonths 12 to 24Reusable data products and additional production use cases
Enterprise capabilityMonths 18 to 36+Cross-domain governance, master data, legacy remediation, and operating model

Start with a business decision

Do not begin with a general objective to become AI-ready.

Define:

“Improve customer service with AI” is too broad. “Recommend relevant knowledge articles to support agents and reduce handling time without lowering resolution quality” is testable.

Assess data viability early

Before committing to the use case, confirm:

The first output should be a viability assessment, not an enterprise-wide catalog.

Not every defect must be fixed. Readiness should be evaluated against the intended use. A missing field that does not affect the decision may be acceptable. A small error rate in a safety or financial attribute may not be.

Build a minimum governed data product

Create a supported data product for the bounded use case. It should include:

Do not confuse centralization with a single source of truth. CRM, ERP, identity, and master-data systems may each be authoritative for different attributes.

Centralizing conflicting data does not resolve disagreement. It creates one location containing several versions of the truth.

Correct the source of defects

Historical data may need correction, deduplication, mapping, enrichment, or migration. This is only part of the work.

Identify why the defects occur. Typical causes include weak validation, inconsistent business processes, conflicting identifiers, unclear ownership, legacy integration, and acquisition history.

Combine historical remediation with preventive controls such as:

Cleaning downstream data without correcting the source creates permanent remediation work.

Develop the pilot in parallel

Do not wait until every data problem is resolved before testing the use case.

Run data engineering, model evaluation, workflow design, security, and governance in parallel. A narrow pilot often reveals problems that cataloging cannot find, including inconsistent fields, unrepresentative history, late-arriving data, and workflows unable to act on recommendations.

Data preparation is often the largest and least predictable part of an AI initiative. The repeated observation that every project spends around 80 percent of its time cleaning data should not be used as only planning assumption.

Measure how much effort is spent rebuilding the same mappings, extracts, identity matching, and corrections. Repeated manual preparation shows that reusable data capabilities have not been established.

Evaluate the whole workflow

Model accuracy alone does not demonstrate business readiness.

Test whether:

Include data, infrastructure, model, support, oversight, and change-management costs when measuring value.

A pilot should also be allowed to show that the use case is not viable. Avoiding an expensive production failure is a useful result.

Harden the complete process

Production readiness requires more than a successful demonstration.

Add:

A pipeline can complete successfully while producing incomplete, duplicated, stale, or incorrectly transformed data. Monitor the business result, not only job completion.

Scale through reuse

Scale into adjacent use cases that can reuse the same data products, definitions, quality controls, and access patterns.

Do not call ten unrelated pilots enterprise scaling. If each pilot creates separate pipelines and definitions, the organization is scaling technical debt rather than capability.

Over time, data remediation should become part of normal operations. Business owners and source-system teams must share responsibility with data and platform teams for preventing deterioration.

Measure readiness with evidence

For each use case, demonstrate that:

Moving data to cloud storage does not make it AI-ready. Readiness depends on fitness for use, not platform location.

Planning ranges

Use these ranges as starting assumptions, not commitments:


Start with a valuable decision, not an enterprise-wide cleanup. Establish a minimum governed data product, test it through a real workflow, correct defects at their source, and scale through reuse. The objective is not to make all enterprise data AI-ready, this is usually to costly, and unrealistic. Objective is to make priority data reliably fit for specific decisions while building capabilities that can be reused.