Four-week pilot
Four weeks. Acceptance criteria written down first.
The decision at the end of the pilot does not rest on whether the interface impressed anyone. It rests on evidence, auditability, cost and real usage signals — all measured from the product’s own records.
Week by week
What happens, in order
Week 1 · Connectivity
Warehouse connection, user roles and row-level security, Semantic Contract import, and the first question-and-answer screen with its evidence panel.
Week 2 · First autonomous analysis
Scheduled scanning in one business domain you choose. The first discovery agent — job description written together — and the weekly rhythm. Your first brief arrives this week.
Weeks 3–4 · Second domain and evaluation
A second business domain, executive readers on the system, the feedback loop running, and the pilot assessed against the criteria below.
Acceptance criteria
Four sentences. Each one testable, none impressionistic.
- Every number shows its source Traceable to the query and its referenced evidence, in one step, on every surface.
- Every assumption is explicit An interpretation is visibly an interpretation and never carries a measurement’s stamp.
- The spend limit is known before the analysis starts Step, time and cost limits locked before a run begins; a run that hits one stops and says so.
- Every action is in the audit log Model, user, time, cost and outcome, append-only.
What gets measured
And where the number comes from
| Measure | Where the number comes from |
|---|---|
| Cost per analysis | Frozen per run against the price row in effect |
| Rate at which the provenance check catches figures | How often a number was stopped or flagged before delivery |
| Brief opens and feedback | First opening stamped once, never rewritten; votes with reasons |
| Refusal rate and its reasons | Held, duplicate and overflowed candidates, each with a recorded cause |
Request