Trust model
Every number has an address.
A number without one cannot reach the screen. The trust model does not rely on prompting, and it does not assume errors will fall as models improve. It relies on a deterministic control chain between the number and the sentence — and on recording what that chain refused.
Traceability
When a query runs, its result is sealed into an envelope. Every value inside it gets an address.
Which envelope, which row, which cell — and a check digit at the end, so the wrong door is not opened by accident. When the agent uses a number, it declares that address. The engine looks it up and compares. Even a percentage is never left to the model: it cannot say “divide”, it asks the engine for the ratio.
Four steps, in this order, every time
- The declared address is parsed — envelope, row, cell.
- The check digit is verified. If it fails, resolution is never attempted.
- The envelope is opened and the actual value in the cell is read.
- The number the agent wrote and the value read are compared.
The address reaches the query step — which query, which period, which filters — not the individual warehouse row. We would rather write that here than have you find it in week three of a pilot.
Trust tiers
A approved official metric · B derived · C exploratory or candidate. A tier C value cannot enter a brief raw, and promotion to A is a human decision.
There is deliberately no single “confidence score” badge.
FinalGuard
The last gate before delivery. Not an AI. Cannot be switched off.
Zero tolerance: one cent off is rejected, and an answer written before a query ran is rejected even if the number happens to be right.
The rejection ladder
- First rejection → the answer is held; the agent is told to use the envelope values verbatim.
- Second rejection → if there is exactly one clear envelope, the engine writes the answer itself.
- Third rejection → honest cut-off.
Withheld values
A number without evidence is removed from the body and logged in a box under the answer with its reason. It is structurally blocked from export, distribution and memory. If a correction was made, a note appears under the answer — never hidden.
Measured. Across our own operating record the provenance check stopped 282 draft answers across 275 runs; the agent fixed 87.2% on the next attempt. The review page that keeps this count refuses to draw a trend line through it: a rising rejection rate does not mean the agents got worse.
A deliberate absence. There is no “Verified ✓” badge. Only the unverifiable is flagged, so a coincidental match never earns a false stamp.
What the gates refused
Mid-run, the engine sent an entire draft answer back.
Any vendor can show you a good briefing. Ask to see what the system refused to say.
This is one scheduled run, opened in full. It began at 05:00, finished 285 seconds later, and cost 78 cents of model spend, frozen at the price in effect that day. Thirteen candidate records reached the database. Two did not survive the gates, and their reasons are still on file.
Its figures had no address the engine could resolve. This is what it wrote.
[engine] Structural rule (numeric provenance): every figure in a final answer is exactly one of three things. ADDRESSED (declared in the claims block with the ref you read it from), COMPUTED (declared with op and operands that are addresses), or STRUCTURE the engine or the user put there. Nothing else is proven. A value that merely EXISTS somewhere in the evidence is not proven until you name its address. Your answer was NOT delivered.

Across the whole operating record, eleven candidate objects were held. Nine were the same family: fastest-growing, weakest, highest. The class of sentence the gate catches most often is exactly the class an executive is most likely to act on.
In the product
The evidence panel, as it ships
Every figure carries the address the agent declared for it, and what the engine found when it resolved that address.

Four questions
Enterprise trust rests on four questions. None of them is left to the AI — and each answer states its limit.
Traceability — “Where did this number come from?”
Every number carries a three-part address: envelope, row, cell. The engine assigns it; the agent only declares it. Every record is linked to the run, the agent, the frozen model and the user authority that produced it. Values are frozen onto the record at creation, so the evidence survives even if the run log is cleaned.
The address reaches the query step — which query, which period, which filters — not the individual warehouse row.
Auditability — “Who, when, with what, at what cost?”
The audit log is append-only. No record is deleted, only archived; votes are stamped over. The publication chain is deterministic: same input, same result. Every dropped, held or duplicate candidate sits in the Guard Ledger with its reason. Step, time and cost limits are frozen before each run starts.
The audit log lives inside the product. Export to a SIEM or external audit system is scoped separately in the pilot.
Transparency — “What did the system do, and not do?”
Nothing silently disappears. A recurring topic is stamped ongoing / repeat and counted, not hidden. An empty period falls into a separate data state class. Causal or superlative language is badged interpretive: measured is separated from interpreted. Assumptions made in chat are written out.
Notification texts deliberately carry no values. Transparency is inside the record, not the notification.
Explainability — “Why did it reach this conclusion?”
In every finding the metric, period and filters are written by the engine. Every recommendation opens the finding it rests on in one click. “Drill down in chat” carries the finding into a conversation as quoted data, so “why?” is asked over the same evidence.
The product explains where the numbers in the sentence came from — not why the model phrased the sentence that way. Severity is an agent claim. Evidence is explained; judgement stays human.
Absence and assumption
The second thing as dangerous as a wrong number: a right number misunderstood.
- Assumptions are stated up front. “You did not state a period; full-year 2024 was assumed.” A fixed template sentence, identical every time.
- Freshness is measured, never assumed. “This model holds data up to 31 Dec 2025 — measured, not estimated.”
- A grain the model does not carry is refused, not approximated. “A quarterly view is not available in this model.” The engine will not build a quarter out of months to satisfy the question.
- A restricted user sees their own restriction. “Scope: Store — Maltepe Park.” The rule lives in the engine; with no resolvable identity it stops rather than falling back to “show everything”.
- Personal data is masked mechanically. A breakdown label reduces to initials by rule, not by prompt.
- Absence is published, not swallowed. A period with no data becomes a data state whose basis is
empty_result_set: “not a measured zero; there are no recorded transactions in this window.” No severity badge, its own region, and no recommendation may rest on it.
Self-audit
“Every number is checked” is a claim. So it is tested every week.
Golden set
A sealed question set. If two runs return the same fingerprint, behaviour has not changed. When the model or the code changes, the difference shows up immediately.
Zero-tolerance line
Unsourced numbers, refusing to estimate, personal-data masking. Hitting the threshold by deleting numbers is also caught.
Cadence
The tests run automatically every week and at every release close. If a claim contradicts the measurement, the system stops itself.
Live services screen
Measured on every load; it never shows a stored status.
The trust model does not assume errors will fall as models improve. Even if the model changes, addressing, the audit log, the publication chain and authority checks remain deterministic.