Audit trails that survive an audit
Append-only design, hash chaining, what must never be logged, and how to prove to an assessor that the record in front of them is complete.
Most applications have logging. Very few have an audit trail. The difference is the question each answers: logging answers "what was the system doing", an audit trail answers "who did this, when, on whose authority, and can you prove the record has not been edited since". The second question is asked under adversarial conditions, sometimes years later, by someone who assumes you are wrong.
Append-only is a schema decision, not a policy
The most common failure we find in systems we inherit is an audit table that the application can UPDATE. Once that is possible, the trail proves nothing, because the honest explanation and the dishonest one produce identical rows. Make it structurally impossible instead of forbidden.
CREATE TABLE audit_events (
id bigserial PRIMARY KEY,
occurred_at timestamptz NOT NULL DEFAULT now(),
actor_id uuid NOT NULL,
actor_role text NOT NULL,
action text NOT NULL,
subject text NOT NULL, -- what was acted upon
reason text, -- required for privileged actions
payload jsonb NOT NULL, -- the change, never the secret
prev_hash bytea NOT NULL,
hash bytea NOT NULL
);
REVOKE UPDATE, DELETE ON audit_events FROM app_user;
GRANT INSERT, SELECT ON audit_events TO app_user;
-- Truncation is the quiet one people forget.
REVOKE TRUNCATE ON audit_events FROM app_user;Chaining: cheap, and it answers the only hard question
An assessor's difficult question is not "can I see the log" — it is "how do you know nothing was removed from it". Row counts do not answer that. A hash chain does: each row commits to its predecessor, so any deletion or edit breaks verification from that point forward, and the break is detectable without trusting the database.
def append(conn, event: dict) -> None:
with conn.transaction():
prev = conn.execute(
"SELECT hash FROM audit_events ORDER BY id DESC LIMIT 1 FOR UPDATE"
).scalar() or b"\x00" * 32
body = canonical_json(event) # stable key order, UTC, no floats
digest = hashlib.sha256(prev + body).digest()
conn.execute(
"INSERT INTO audit_events (..., prev_hash, hash) VALUES (..., %s, %s)",
(prev, digest),
)- Canonical serialisation matters more than the hash function. Two services writing the same event must produce identical bytes, or verification fails for reasons that have nothing to do with tampering.
- FOR UPDATE on the previous row serialises appends. On a pass system that is a few hundred events a minute at peak and the lock is uncontended; on a high-volume ledger, chain per-partition instead and verify partitions independently.
- Publish the head hash somewhere outside the database — a daily line in an operations channel is enough. It converts "the database says it is intact" into "the database and an external record agree", which is a materially stronger claim.
What must never enter the trail
An audit trail is the most widely-read table in the system and the least likely to be deleted. That makes it the worst possible place for anything sensitive, and the temptation to log the whole request body is exactly how personal data ends up retained for seven years by accident.
| Never | Instead |
|---|---|
| Full request or response bodies | The changed fields, by name, with before and after for non-sensitive values |
| Identity document numbers, full dates of birth | A stable pseudonymous subject id that resolves through a separate, access-controlled table |
| Photographs or biometric templates | A hash of the artefact, so the trail can prove which one was used without holding it |
| Tokens, session ids, OTPs, API keys | The credential type and its last four characters, if you need it at all |
| Free-text operator notes without review | A reason drawn from a controlled list, with free text permitted only in a field the retention policy covers |
Proving completeness in the room
When an assessor is in front of you, the demonstration that lands is not a document. It is running the verifier on their laptop against an export they choose, and showing them what a tampered row looks like. We keep a deliberately corrupted fixture for this — being able to show the failure mode is what makes the success case credible.