AGENT RECORD AUDIT · FORENSIC RECORD & PROVENANCE VERIFICATION FOR AUTONOMOUS SYSTEMS

If one of your AI agents harmed a third party tomorrow, could you prove what it did — and that the record wasn't altered afterward?

Under whose authority it acted, what information it had, which model produced the action, what actually executed. Most organizations find out the answer during the incident. That is the wrong time to find out.

BOOK A 90-MINUTE RECONSTRUCTABILITY REVIEW — $1,500 →RUN THE 8-QUESTION SELF-CHECK
The self-check runs in your browser — nothing leaves it. Small door: open the live gate, the running instrument behind all of this.
WHAT $1,500 BUYS — EXACTLY
A 90-minute structured session against the Eight Questions with whoever owns your agent's logs, plus a 2–3 page written findings memo you can forward internally. It is not the full Baseline Review — the findings report, remediation sequence, and reproduction method are scoped and quoted after the session, in writing, before anything begins.

What a reconstructable record looks like

ONE RECONSTRUCTED INCIDENTSYNTHETIC EXAMPLE — INVENTED FOR ILLUSTRATION, NOT A CLIENT RECORD
A camera-analytics model flags a person on a perimeter feed at 02:14 and security dispatches a responder. The next morning, someone has to answer for it.
What most operators can prove today: a dashboard screenshot.
What they cannot prove: which model version scored the frame, which frames it actually used, what the alert threshold was that night, and who approved the dispatch.
AFTER A SEALED RECORD, THOSE FOUR FIELDS EXIST — HASHED, DOWNLOADABLE, RECOMPUTABLE:
MODELpercept-cam v3.2.1 (build 8841, weights sha256 ce41…a90f)
FRAMEScam-07 02:14:02–02:14:09 · 43 frames · sha256 7b2d…4c11
THRESHOLDperson-confidence ≥ 0.83 (night profile, set 2026-03-02 by ops-lead)
AUTHORIZATIONdispatch approved: badge S-114 at 02:15:37 · sealed leaf 9f0e…22ba
Every field above is a Merkle leaf under one signed root — alter any one afterward and recomputation fails. The mechanism is real and running: see it live or verify a sealed file yourself. This incident, again, is synthetic.

The findings memo you receive follows the same discipline — read a full sample report built from this synthetic incident, or download it as a PDF.

Start here — check your own records

Before any engagement, run the self-check. It is the same eight questions every audit answers, and it will show you where your evidentiary chain likely breaks — for free, and without sending us anything.

THE EIGHT QUESTIONS · SELF-CHECK

Answer these about your own agent deployment. This runs entirely in your browser — nothing is uploaded, and this is a self-assessment, not the audit. The audit reconciles your answers against independent evidence.

1.Can you reconstruct the exact state your system was in when the agent acted — not inferred from a later transcript?
2.Can you establish what information (sources, context, memory, tools) the agent actually had at that moment?
3.Can you prove which model and version produced the proposed action — not just an application label?
4.Can you reconstruct which policy, gate, or rule evaluated that action?
5.Can you distinguish a proposed action from an authorized one, and an attempted one from an executed one?
6.Can you establish which credentials and tools the agent could actually reach at that moment?
7.Can you reconcile the claimed execution against independent upstream or downstream evidence?
8.Can you prove who authored each field in the record — the model, the provider, the operator, or the harness?
Answer all eight to see where your evidentiary chain stands.

What it is — and is not

The method is reconciliation: we take what your system claims happened and reconcile it against what independently verifiable sources say happened — provider billing, upstream and downstream logs, execution records, cryptographic recomputation. The goal is to find where the chain of events cannot be independently reconstructed.

It is not a penetration test, not a model evaluation, and not a certificate that your system is “safe.” It does not manufacture a defense. It tells you what your system can actually prove — and where it cannot.

What we will not claim

This is not a legal opinion — your counsel determines the legal significance of any finding. No architecture creates a safe harbor; a tamper-evident record does not make unauthorized conduct lawful. We do not certify claims we cannot substantiate, and our bias-screening research measures toxicity and framing signals — it is not a protected-class discrimination test and will not be represented as one.

Beyond the first session

Baseline Review
One agent system, one environment. The Eight Questions assessment, a findings report, an evidentiary assessment, a remediation sequence, and a reproduction method your own team can re-run.
Extended Review
Multiple systems or environments, with deeper reconciliation against provider billing, upstream logs, and execution records where access is available.
Retained Review
Periodic re-verification for deployments that change frequently, or organizations under ongoing operational, regulatory, insurance, or counterparty scrutiny.

Scope and pricing for these are established after the initial session. No engagement begins before scope is agreed in writing. Every finding ships with the method used to reach it, so your own team can reproduce the test after remediation — the method transfers; you don't stay dependent on us.

To start

Bring one agent deployment and whoever owns its logs. We examine the Eight Questions. If your system already has good answers, we'll tell you. If it doesn't, we'll show you where the evidentiary chain breaks and what it would take to close the gap.

BOOK THE 90-MINUTE REVIEW →