RAG Evaluation Consulting
Measure retrieval recall, answer faithfulness, and citations on your questions — then gate every index and prompt change in CI.
- Service
- Evaluation
- Industry
- Enterprise
- Updated
- 2026-08-25
- Engagement
- 4 wks
RAG evaluation consulting builds a golden set of real questions with verified answers, then measures retrieval hit@k, answer faithfulness, and citation quality, with refuse-when-empty and per-user ACL tests, wired as CI gates in the client cloud — typically a 4-week standard implementation, with the client owning the eval suite and IP.
Why teams pick this engagement
Evaluation × EnterpriseACL-aware evals
Restricted questions are run as the restricted user. A suite that searches as admin will not catch leaks.
Retrieval and generation split
Hit@k is scored before anyone debates the prose. Faithfulness is scored against retrieved context, not against the internet.
Harness in your CI
The suite runs in the client cloud on pull requests. A drop versus the last pin fails the build.
Labeling your people can sustain
Gold passages and answers are written by domain owners with a review queue. We do not leave you a one-off spreadsheet.
Four-week harness
Question sampling, gold labels, metrics, CI wiring, handover. Thirty days on-call for triage.
Citation as a metric
Every answer must point at a span. Missing, wrong, or decorative citations fail the check.
Key takeaways
- 01
If you cannot name hit@k, faithfulness, and citation coverage on a golden set, you are not evaluating RAG. You are demoing it.
- 02
Split retrieval from generation. Fixing prose will not help if the gold chunk is not in the top 10.
- 03
Golden sets plus CI evals are the release control. Offline notebooks that nobody runs on pull requests do not count.
- 04
Eval as the asking user. Admin-only suites miss ACL leaks. Refuse-when-empty must have dedicated cases.
- 05
Four weeks in your cloud. You own the suite. No shared training. You pay inference for judges and generators. Thirty days on-call.
What the engagement covers
How we work
- 01
Discover
Week 1: current evals, question sources, identity for ACL tests, metric targets, written scope.
- 02
Design
Strata, label schema, retrieval vs generation split, judge protocol, CI placement.
- 03
Build
Week 2: golden set, harness, traces, and pipeline jobs in the client cloud.
- 04
Validate
Week 3: shadow scores on the live retriever, inter-labeler checks, gate thresholds.
- 05
Enable
Week 4: CI live, runbooks, IP handover, 30 days on-call for eval triage.
Take the playbook with you
The working documents from real engagements — free, in exchange for an email. They’re useful whether or not we ever talk.
RAG Golden Set Labeling Guide
How to write questions, gold passages, gold answers, and identifier strata so labels stay stable across reviewers.
Get the guide ·RAG Eval Metric Sheet
Hit@k, faithfulness, citation coverage, refusal-on-empty, and ACL leak tests — definitions we wire into CI.
Get the sheet ·