Data Readiness for Generative AI Consulting
Fix the data layer GenAI actually needs — canonical sources, ACLs, PII, and eval sets — before you scale ChatGPT Enterprise, Copilot, or agents.
- Service
- Data
- Industry
- Enterprise
- Updated
- 2026-08-25
- Engagement
- 4 wks
Data readiness for generative AI consulting inventories the corpora a use case needs, repairs permissions and PII handling, names canonical vs stale sources, and builds the first eval questions — before ChatGPT Enterprise, Microsoft Copilot, or agents are scaled. Custom GPTs are prototypes and will happily answer from junk. A four-week standard produces a go/no-go per use case in your stack. You own catalogs and evals; you pay providers with no markup; SSO/IdP stays yours.
Why teams pick this engagement
Data × EnterprisePermissions before indexing
If SharePoint is overshared, Copilot and RAG will be overshared. We treat ACL repair as data readiness, not a later security ticket.
Ready vs not-ready, by use case
Each candidate GenAI use case gets a go/no-go on sources, freshness, PII, and a golden set. “We have a data lake” is not readiness.
Context engineering, not a dump
Chunking, metadata, and canonical-vs-archive rules are designed so agents get a budgeted context window — not every PDF the crawler found.
Stewards in the business
Every corpus needs an owner who will say which document is current. Data readiness without stewards becomes an IT index of contradictions.
Four-week standard
Week 1 inventory, week 2 ACL/PII/canonical design, week 3 sample index and eval questions, week 4 readiness report and backlog.
Artifacts you keep
Source catalog, PII handling, eval questions, and index jobs live in your cloud and IdP. We do not warehouse your documents.
Key takeaways
- 01
GenAI fails on contradictory, overshared, or unowned documents more often than on the choice of GPT vs Claude.
- 02
Copilot is often the wrong tool when write-actions and evals are required — and it is the wrong first buy if Graph ACLs are a mess.
- 03
Custom GPTs indexed on a file dump are prototypes of a hallucination engine. Readiness is canonical sources plus a golden set.
- 04
Context engineering and LLMOps sit on top of this work; they cannot invent a steward or an ACL.
- 05
Four weeks produces a scorecard and a sample index you own, not a multi-year data-platform rewrite.
What the engagement covers
How we work
- 01
Discover
Week 1: use cases, corpora, IdP groups, PII classes, current Copilot/ChatGPT file uploads, evals (usually none).
- 02
Design
Scorecard, gates, canonical rules, chunking approach, golden-question plan, what stays out of Custom GPTs.
- 03
Build
Week 2–3: ACL fixes or exclusions, sample index in your cloud, first golden set, steward assignments.
- 04
Validate
Permission tests, PII spot checks, retrieval on golden questions, go/no-go workshop with security and the business.
- 05
Enable
Week 4: catalog, backlog, pipeline docs, 30 days on-call. You own the data path; we do not become your data platform.
Take the playbook with you
The working documents from real engagements — free, in exchange for an email. They’re useful whether or not we ever talk.
GenAI Data Readiness Scorecard
Per-corpus scoring: canonical, ACL, PII, freshness, format, steward, eval coverage — the rubric we use in week 1.
Get the scorecard ·PII and Permission Gate for RAG
What must be true before a file is embeddable, and how live IdP entitlements wrap retrieval.
Get the gate ·