D-CAT Group ↗ Contact ENTRDEESAZ
Start a 4-week pilot

Trust model

Every number has an address.

A number without one cannot reach the screen. The trust model does not rely on prompting, and it does not assume errors will fall as models improve. It relies on a deterministic control chain between the number and the sentence — and on recording what that chain refused.

Traceability

When a query runs, its result is sealed into an envelope. Every value inside it gets an address.

Which envelope, which row, which cell — and a check digit at the end, so the wrong door is not opened by accident. When the agent uses a number, it declares that address. The engine looks it up and compares. Even a percentage is never left to the model: it cannot say “divide”, it asks the engine for the ratio.

The structure of an evidence address The address e1.r1.v.1u breaks into four parts: e1 identifies the sealed envelope of one query result, r1 the row within it, v.1 the cell in that row, and the final u is a check digit verified before the address is resolved. e1.r1.v.1u e1 Envelope The sealed result of one query. r1 Row Which row of that result. v.1 Cell Which value in the row. u Check digit Verified before the address is resolved.
The agent declares an address. The engine resolves it. Nothing on this diagram is a stage the model is in — which is the point.

Four steps, in this order, every time

  1. The declared address is parsed — envelope, row, cell.
  2. The check digit is verified. If it fails, resolution is never attempted.
  3. The envelope is opened and the actual value in the cell is read.
  4. The number the agent wrote and the value read are compared.

Measured catch rate against accidental resolution to a neighbouring cell: 97.63% with one check digit, 100% with two · internal measurement. FinalGuard comparison ≈ 2 ms · demo measurement.

Limit

The address reaches the query step — which query, which period, which filters — not the individual warehouse row. We would rather write that here than have you find it in week three of a pilot.

Trust tiers

A approved official metric · B derived · C exploratory or candidate. A tier C value cannot enter a brief raw, and promotion to A is a human decision.

There is deliberately no single “confidence score” badge.

FinalGuard

The last gate before delivery. Not an AI. Cannot be switched off.

Zero tolerance: one cent off is rejected, and an answer written before a query ran is rejected even if the number happens to be right.

The agent wrote…WhyVerdict
the exact value in the envelopeidenticalpasses
a standard rounding of itrounding rules are explicitpasses
a difference the engine had computedit is in the envelopepasses
a value one unit offneither identical nor a valid roundingrejected
a total the agent summed itselfthe agent does not do arithmeticrejected
“across nine regions” with no evidencesmall or not, unproven is unprovenrejected

The rejection ladder

  1. First rejection → the answer is held; the agent is told to use the envelope values verbatim.
  2. Second rejection → if there is exactly one clear envelope, the engine writes the answer itself.
  3. Third rejection → honest cut-off.

Withheld values

A number without evidence is removed from the body and logged in a box under the answer with its reason. It is structurally blocked from export, distribution and memory. If a correction was made, a note appears under the answer — never hidden.

Measured. Across our own operating record the provenance check stopped 282 draft answers across 275 runs; the agent fixed 87.2% on the next attempt. The review page that keeps this count refuses to draw a trend line through it: a rising rejection rate does not mean the agents got worse.

A deliberate absence. There is no “Verified ✓” badge. Only the unverifiable is flagged, so a coincidental match never earns a false stamp.

What the gates refused

Mid-run, the engine sent an entire draft answer back.

Any vendor can show you a good briefing. Ask to see what the system refused to say.

This is one scheduled run, opened in full. It began at 05:00, finished 285 seconds later, and cost 78 cents of model spend, frozen at the price in effect that day. Thirteen candidate records reached the database. Two did not survive the gates, and their reasons are still on file.

05:00
scheduled start
285 s
total duration
$0.78
model spend, frozen
13
candidates produced
11
published
2
held, with reasons

Scope: one real scheduled run from our own operating record, names masked, read from the product’s database. Not a customer benchmark.

Its figures had no address the engine could resolve. This is what it wrote.

engine · structural rule · numeric provenanceNOT DELIVERED
[engine] Structural rule (numeric provenance): every figure in a final answer
is exactly one of three things. ADDRESSED (declared in the claims block with
the ref you read it from), COMPUTED (declared with op and operands that are
addresses), or STRUCTURE the engine or the user put there. Nothing else is
proven. A value that merely EXISTS somewhere in the evidence is not proven
until you name its address.
Your answer was NOT delivered.
The closing summary of an ordinary run in the console, headed Operation trace, what this run did. It reads: the agent worked 10 turns and sent 50 queries to the warehouse. One answer attempt was sent back because it carried figures it could not source; the agent was told, and its next attempt was accepted. Six outputs were published: three findings, three recommendations. Cost 0.79 US dollars, 179 thousand input and 15 thousand output tokens, 100 thousand of them from cache. Below it, an account of the agent's work written out in prose.
demo dataThe same accounting at the end of an ordinary run, on the run itself rather than in a report about it: how many turns, how many queries, what was sent back and why, what was published, and what it cost down to the tokens it read from cache.

Across the whole operating record, eleven candidate objects were held. Nine were the same family: fastest-growing, weakest, highest. The class of sentence the gate catches most often is exactly the class an executive is most likely to act on.

Scope: whole operating record, 3 agents, 3.5 weeks, read from the product’s own database. Not a customer benchmark.

A scheduled run and the conversation that follows it, end to end. Demo data, one minute forty-five, no narration: the run stream, the settings frozen when the run started, what it published and what it held, then a finding carried into conversation and questioned there.

In the product

The evidence panel, as it ships

Every figure carries the address the agent declared for it, and what the engine found when it resolved that address.

A published finding headed: non-promoted Fresh Food gross sales grew only plus 57.01 percent, well below the network average of 96.97 percent. Two figures inside its narrative carry a dagger, and under the narrative a line reads: model statement, not verified by the engine. Below that an Evidence panel says every figure below carries the address the agent declared for it, the engine resolved that address and this is what it found. Each row names its metric, its period and its promotion status, and each is labelled either the engine computed it or measured at the source.
demo dataTwo values in the narrative above carry a dagger and the line model statement — not verified by the engine. Everything else is listed underneath with the address it was read from, and each row says whether the engine measured it at the source or computed it — a distinction most tools never make at all.
The evidence panel of a finding about promoted basket depth, with one derivation opened. The figure 1.5149 for units sold per transaction count is shown as: this figure came from 677.526 divided by 447.235, and each operand is listed underneath with its own metric, period and promotion status.
demo dataA derived figure, opened. It is not asserted — it is shown as the division it came from, and each operand carries its own address, exactly as the measured figures above it do. Nothing here was written by the model.

Four questions

Enterprise trust rests on four questions. None of them is left to the AI — and each answer states its limit.

Traceability — “Where did this number come from?”

Every number carries a three-part address: envelope, row, cell. The engine assigns it; the agent only declares it. Every record is linked to the run, the agent, the frozen model and the user authority that produced it. Values are frozen onto the record at creation, so the evidence survives even if the run log is cleaned.

Limit

The address reaches the query step — which query, which period, which filters — not the individual warehouse row.

Auditability — “Who, when, with what, at what cost?”

The audit log is append-only. No record is deleted, only archived; votes are stamped over. The publication chain is deterministic: same input, same result. Every dropped, held or duplicate candidate sits in the Guard Ledger with its reason. Step, time and cost limits are frozen before each run starts.

Limit

The audit log lives inside the product. Export to a SIEM or external audit system is scoped separately in the pilot.

Transparency — “What did the system do, and not do?”

Nothing silently disappears. A recurring topic is stamped ongoing / repeat and counted, not hidden. An empty period falls into a separate data state class. Causal or superlative language is badged interpretive: measured is separated from interpreted. Assumptions made in chat are written out.

Limit

Notification texts deliberately carry no values. Transparency is inside the record, not the notification.

Explainability — “Why did it reach this conclusion?”

In every finding the metric, period and filters are written by the engine. Every recommendation opens the finding it rests on in one click. “Drill down in chat” carries the finding into a conversation as quoted data, so “why?” is asked over the same evidence.

Limit

The product explains where the numbers in the sentence came from — not why the model phrased the sentence that way. Severity is an agent claim. Evidence is explained; judgement stays human.

The audit log in the console, subtitled: signin events plus every mutation, an append-only ledger. Columns read time, event, actor, summary and detail. Visible rows include a narrative_generated event recording 1001 tokens in, 291 out and 0.00736 US dollars; a scheduler_fire event whose actor is System, summarised as scheduled fire enqueued for agent 18, run 1325; create events naming the routes they came from; and sign-in events.
demo dataThe ledger behind the paragraph above. A scheduled firing is recorded against System rather than a person, and a generated narrative carries its own token count and cost. Two things it does not do yet: the detail column is empty on every row, and changing a setting through the interface does not produce a row of its own. Both are on the product’s list, and neither is hidden here.

Absence and assumption

The second thing as dangerous as a wrong number: a right number misunderstood.

  • Assumptions are stated up front. “You did not state a period; full-year 2024 was assumed.” A fixed template sentence, identical every time.
  • Freshness is measured, never assumed. “This model holds data up to 31 Dec 2025 — measured, not estimated.”
  • A grain the model does not carry is refused, not approximated. “A quarterly view is not available in this model.” The engine will not build a quarter out of months to satisfy the question.
  • A restricted user sees their own restriction. “Scope: Store — Maltepe Park.” The rule lives in the engine; with no resolvable identity it stops rather than falling back to “show everything”.
  • Personal data is masked mechanically. A breakdown label reduces to initials by rule, not by prompt.
  • Absence is published, not swallowed. A period with no data becomes a data state whose basis is empty_result_set: “not a measured zero; there are no recorded transactions in this window.” No severity badge, its own region, and no recommendation may rest on it.

Self-audit

“Every number is checked” is a claim. So it is tested every week.

Golden set

A sealed question set. If two runs return the same fingerprint, behaviour has not changed. When the model or the code changes, the difference shows up immediately.

Zero-tolerance line

Unsourced numbers, refusing to estimate, personal-data masking. Hitting the threshold by deleting numbers is also caught.

Cadence

The tests run automatically every week and at every release close. If a claim contradicts the measurement, the system stops itself.

Live services screen

Measured on every load; it never shows a stored status.

The trust model does not assume errors will fall as models improve. Even if the model changes, addressing, the audit log, the publication chain and authority checks remain deterministic.