For the auditor. There is a companion document, audit-for-helpers.md, for
whoever from the auditee sits with you and operates the tool. You should not need
to read theirs; they should not need to read yours. If you find yourself needing
theirs, that is a defect in this one — say so.
You have two surfaces and they do different jobs:
Everything below was run against the Aureon Smart Systems pack
(testdata/iso9001.db). Ids, statuses and counts are quoted as the surfaces
actually returned them.
None of these can be fixed on the day, and each of them changes what your findings are worth. The one-page version to attach to the audit plan is
audit-engagement-checklist.md.
1. Read credentials, issued to you, before day one. Both surfaces: the web app and the chat. You are expected to use both. 2. The session runs under the audit's own account, even when the helper is at the keyboard. The audit trail (§6) records an account, not a person; under a shared account it records nothing useful. 3. The register is materialised and exported before the audit period opens, and the export is signed by both parties. Why this matters is §5. 4. Transcript retention agreed in writing: who keeps the chat transcript and the session trail, for how long, and in what form the certification body can have them. 5. The requirement library version is fixed and recorded. Any edit to it during the audit period is itself a finding. 6. AI hosting and retention terms agreed. See §7. 7. Out-of-KF evidence identified up front — the sibling legal entity, shared drives, email — with a conventional evidence route agreed for each. Half the evidence is usually outside KF, and that half needs planning, not reconciliation afterwards.
Record the helper's role for what it is: a competent intermediary, present at your discretion. Note where you take it back — the criteria judgement, the cross-reading (§4), and any question whose framing is itself in issue.
/audit?id=One page, and it is the shape of the audit. It carries:
If you read nothing else in this document, open that page and follow it.
/list/list enumerates. Filter by class, tag, code, subtree or modification date;
sort by any column; page through it; export it.
It always states the total. That is the point of it. A list of nine and the first nine of three hundred look identical, and without a denominator "is that all of them?" is answerable only by asking the organization you are auditing — which is exactly the question their answer should not settle.
Two flags to notice:
The moves worth making on day one:
| Question | URL |
|---|---|
| What moved in the fortnight before the audit? | /list?under= |
| Everything graded major | /list?under= |
| Every document in scope, newest first | /list?under= |
| Everything mentioning one asset id | /list?q=FTS-01 |
A punctuated term such as FTS-01 is searched as a phrase, so it finds the
records that name that instrument and not every record containing "01".
Open an item. Its tabs are components, tasks, events, types, people; each states how many rows it is showing and when each was last written, and by whom.
A blank author is a revision written before KF recorded authorship. It is left blank rather than filled in with a guess. Absence of a user on an old revision is not evidence of concealment.
/item/ is the record of when something was written and by whom. A date
typed into a document body is the organization's claim; this is the record.
The version message column is KF's answer to "what changed here?" — written by the author, per revision. For control of documented information that is the right record: the organization's own stated reason for a change.
Each row has a diff link, which is what makes that record checkable. "Version message says 'typo fix', the diff says the due date moved and the status went to done" is a finding, and it is the same claim-versus-record logic one level down.
Append ?asof=2026-07-26 to any item URL. The page shows the revision current at
the close of that day, editing is off, and a banner says so. Use it when you need
to establish what somebody could have known at the time.
/item//risk A risk level is an assertion until you can re-derive it. This page gives you the matrix that scored this item — a database may hold two — its rules in the order they are evaluated, the scales, and what every value means.
Re-derive one rating in every audit. It takes two minutes and it tells you whether the methodology is sound.
Look for values with a position and no words. The page marks them. In the Aureon pack severity 2–6 and 8 have a number on the scale and no description, so "severity 8" is ordinally meaningful and verbally undefined — two assessors cannot be expected to agree on it. That is a finding about the risk methodology, not about any one rating.
Everything you will want to attach to the file:
| What | Where |
|---|---|
| The register | /item/ and .xlsx |
| A listing | the CSV button on /list |
| Revisions | /item/ |
| The FMEA | /item/, /item/ |
| Your own session trail | /api/audit-trail?format=csv |
Every one of them carries a provenance header: which database, which build, which filter, what time. A register exported without that is a table of numbers whose meaning cannot be re-established later.
Joins, cross-cutting questions, re-derivations, and anything where you would otherwise open twenty pages. Not evidence.
Every id it gives you arrives as a link. Click it. Verification that costs one click gets done; verification that costs a search and three navigations gets skipped when the day runs late — and what gets skipped is the middle of the sample, the claims that never become findings and so are never checked by anyone.
Every answer carries a count and a complete flag. Read them. complete:
false with a note means you are looking at a subset.
Every answer carries the provenance block. It should match what the app's footer says. If it does not, you are looking at two databases.
The questions the chat is actually better at:
List every requirement underwhose status is missing, with its code. Which records mention FTS-01? Give me the id and title of each. Show me the risk model that applies to , then re-derive its level from its inputs and tell me whether the reported level is correct. Show me the revision history of , with the version message of each revision. For each claim you just made, give me the item id it rests on.
Ask for ids. An id is checkable; a paraphrase is not. Get into the habit on the first question and never drop it.
The two major nonconformities in the Aureon audit came from putting two records side by side. J03 says instrument FTS-01 went out of calibration on 2026-06-30. J11, the release record, says batch B-2026-17 — 1 450 Hub units — was verified on FTS-01 and released 2026-07-17. Seventeen days after the calibration lapsed.
No query connects those two records. KF has no edge between them. You find it by reading both and noticing.
This used to be the whole of the advice. Here is a method.
Three pairings account for most of what is found this way:
| Read this | Against this | What you are looking for |
|---|---|---|
| Calibration / equipment records | Release and test records | Product released on an instrument that was out of calibration, unverified, or not in the register |
| Supplier approval and evaluation | Incoming inspection, and the product it went into | Material accepted from a provider whose approval had lapsed, or whose evaluation never happened |
| Nonconformities and corrective actions | The product, process or clause they concern | A corrective action closed without effectiveness evidence; the same cause recurring under a different number |
Not titles. Titles are neither unique nor stable. Index on the things that appear verbatim in more than one record:
FTS-01, AS-114)B-2026-17)1. Export the documents in scope: /list?under= → CSV.
2. For each identifier you have collected, run /list?q= and note
every item it returns. Two or more hits on one identifier is a candidate.
3. Open both records in the app — not in the chat — and read them.
4. Where they conflict, check /item/ on both to establish which
statement was written first, and what was known when.
Step 2 is the one that used to have no query behind it at all.
Half a day for a small QMS, and do not compress it. This is the step that produces the findings that matter; everything before it produces the list of places to look.
MCP cannot write. Findings go back into KF over its REST API or web UI, entered by the auditee, typically as events linked to the requirement they fail and tagged with their grade.
There is a structural trap here, and §0(3) is the answer to it.
A finding can only attach to a requirement the organization has taken ownership
of. The requirements most in need of findings are exactly the inherited ones
still reading missing — and taking one on to hang a finding from it moves it
from missing to pending. So the act of recording a nonconformity makes the
register appear to advance.
The fix is sequencing. Before the opening meeting, the whole register is materialised in one operation and exported:
POST /api/items//checklist/materialise-all?dry=true # read this out first POST /api/items/ /checklist/materialise-all # then do it
or workflows/audit-open.sh snapshot , which does that, exports the
register, the population and the reconciliation, and hashes the lot.
Afterwards no row's status changes because a finding was raised. On the
Aureon pack this moved 38 rows from missing to pending once, before the
audit, on the record — instead of moving them one at a time during it.
The same snapshot gives you the drift baseline: audit-open.sh drift
lists everything modified since it was taken. Run it at the close.
The instance keeps a trail of every page you opened and every tool call you made: who, when, which route or tool, the arguments, and how much came back. Export it at the close:
/api/audit-trail?format=csv&user=&since=2026-07-26
It deliberately does not keep the payloads. The evidence is in the database; a second copy of the organization's records in a log file is a retention and confidentiality problem, not a benefit. If your certification body needs more than coverage, settle that in the audit plan (§0, §7) rather than discovering it afterwards.
This is what answers "how did you satisfy yourself on clause 8.4?" at surveillance.
State the position in the audit plan; do not leave it implicit.
Several behaviours changed with the build, and they change numbers, not just wording. Establish which apply from the provenance block — on the page footer, in the chat's envelope, and in the header of every export — rather than by asking the auditee.
| Behaviour | Applies to | How to tell |
|---|---|---|
Heading rows reported pending, inflating every status tally by the number of headings | Builds before 2026/07 | Compare rows_total, rows_requirements and rows_headings in the register summary: on a current build a heading carries no status at all |
search by=code returned exactly one holder, silently — and it was the library definition, not the organization's own | Older builds | On a current build a code lookup returns every holder, oldest first |
Materialised instances were given numeric codes such as 100282, so an instance may not answer to its requirement number | Packs built before mid-2026 | Look at the code column on the register: numeric codes on instance rows |
| The register's reconciliation had to be computed by hand | Builds before this one | reconcile=true on the register, or the summary block at the head of the register page |
| No listing, no last-modified outside an item, no result counts | Builds before this one | /list exists |
The counts quoted in this document are from a current build against the Aureon pack: 79 rows, 69 requirements, 10 headings, reconciled 98 = 98.
The full list of KF's own quirks is in the helper's document. These six are the ones that reach your report if nobody stops them.
1. done is not conformity. It is a tag somebody applied. In the Aureon
audit R30 was done with an overdue instrument in its evidence.
2. missing is not "no evidence exists". It means nobody drew the edge.
16 of 36 minor nonconformities in the reference audit were content that was
present in the QMS and simply not linked. Before writing up a gap, search
for the evidence by subject.
3. pending is not progress. Somebody took ownership; nothing was done.
4. Not every row is a requirement. Ten of the Aureon pack's 79 rows are
headings and carry no status. Count rows_requirements, never rows_total.
5. A done row can hide a subtree. The register shows what lies beneath each
row. A done cell with four missing obligations under it is not
four-fifths done.
6. A risk level without its matrix is an assertion. Open /item/.
Every status word in the app carries its meaning as a tooltip, and the register page opens with the full legend. Use those words as the app defines them, or translate them in your report — but do not carry them across untranslated. "Thirty-eight missing" written down at 4pm becomes thirty-eight gaps by the closing meeting.
Re-reading the record in the app is mandatory, not left to conscience, for:
done row in your sample — the tag says somebody marked it, not that
the requirement is met;/item//revs is the record;Everything else may be read in the chat.
Coverage is not "every row examined", because done is not conformity and
missing is not a gap. Adequate coverage is:
1. Every missing row chased by subject search before it is written up. This
is not optional: nearly half of them turn out to be unlinked evidence. One
/list?q= per gap; the gaps that survive it are real.
2. A full evidence walk on every requirement that is release-critical,
customer-facing, or the subject of a previous finding — regardless of
status.
3. A sample of the remainder, drawn from the listing so the population is
defensible: state the filter, the total it returned, and the row numbers you
drew. /list gives you all three.
4. One re-derived risk rating, one revision diff, and one cross-read (§4) in
every audit, as method checks rather than as sampling.
You are done when 1 and 2 are complete, 3 is drawn and worked, and 4 is on the record. Not when the register looks finished — it never will.
State these in your report rather than papering over them.
audit-criteria-validation.md — and it is not an hour inside a
client audit.The honest summary: conformity as a snapshot is determinable from KF; conformity as a history is determinable for anything KF was told about; and sufficiency is always yours to judge.
Companion documents: audit-for-helpers.md (KF mechanics, for the auditee's
expert), audit-criteria-validation.md (validating the requirement library
against the standard), audit-engagement-checklist.md (the one-pager for the
audit plan). The audit this method was drawn from: iso9001-audit.md.