Auditing an ISO 9001 QMS held in KF

For the auditor. There is a companion document, audit-for-helpers.md, for whoever from the auditee sits with you and operates the tool. You should not need to read theirs; they should not need to read yours. If you find yourself needing theirs, that is a defect in this one — say so.

You have two surfaces and they do different jobs:

Everything below was run against the Aureon Smart Systems pack (testdata/iso9001.db). Ids, statuses and counts are quoted as the surfaces actually returned them.

0. Settle these before the opening meeting

None of these can be fixed on the day, and each of them changes what your findings are worth. The one-page version to attach to the audit plan is

audit-engagement-checklist.md.

1. Read credentials, issued to you, before day one. Both surfaces: the web app and the chat. You are expected to use both. 2. The session runs under the audit's own account, even when the helper is at the keyboard. The audit trail (§6) records an account, not a person; under a shared account it records nothing useful. 3. The register is materialised and exported before the audit period opens, and the export is signed by both parties. Why this matters is §5. 4. Transcript retention agreed in writing: who keeps the chat transcript and the session trail, for how long, and in what form the certification body can have them. 5. The requirement library version is fixed and recorded. Any edit to it during the audit period is itself a finding. 6. AI hosting and retention terms agreed. See §7. 7. Out-of-KF evidence identified up front — the sibling legal entity, shared drives, email — with a conventional evidence route agreed for each. Half the evidence is usually outside KF, and that half needs planning, not reconciliation afterwards.

Record the helper's role for what it is: a competent intermediary, present at your discretion. Note where you take it back — the criteria judgement, the cross-reading (§4), and any question whose framing is itself in issue.

1. Start here: /audit?id=

One page, and it is the shape of the audit. It carries:

If you read nothing else in this document, open that page and follow it.

2. What the app is for: reading the record

The population, and therefore the sample: /list

/list enumerates. Filter by class, tag, code, subtree or modification date; sort by any column; page through it; export it.

It always states the total. That is the point of it. A list of nine and the first nine of three hundred look identical, and without a denominator "is that all of them?" is answerable only by asking the organization you are auditing — which is exactly the question their answer should not settle.

Two flags to notice:

The moves worth making on day one:

QuestionURL
What moved in the fortnight before the audit?/list?under=&since=2026-07-12&sort=time&desc=1
Everything graded major/list?under=&y=e&tag=major
Every document in scope, newest first/list?under=&y=i&sort=time&desc=1
Everything mentioning one asset id/list?q=FTS-01

A punctuated term such as FTS-01 is searched as a phrase, so it finds the records that name that instrument and not every record containing "01".

The record itself

Open an item. Its tabs are components, tasks, events, types, people; each states how many rows it is showing and when each was last written, and by whom.

A blank author is a revision written before KF recorded authorship. It is left blank rather than filled in with a guess. Absence of a user on an old revision is not evidence of concealment.

Revisions, and what the author said changed

/item//revs is the record of when something was written and by whom. A date typed into a document body is the organization's claim; this is the record.

The version message column is KF's answer to "what changed here?" — written by the author, per revision. For control of documented information that is the right record: the organization's own stated reason for a change.

Each row has a diff link, which is what makes that record checkable. "Version message says 'typo fix', the diff says the due date moved and the status went to done" is a finding, and it is the same claim-versus-record logic one level down.

Reading the record as it stood

Append ?asof=2026-07-26 to any item URL. The page shows the revision current at the close of that day, editing is off, and a banner says so. Use it when you need to establish what somebody could have known at the time.

The risk model: /item//risk

A risk level is an assertion until you can re-derive it. This page gives you the matrix that scored this item — a database may hold two — its rules in the order they are evaluated, the scales, and what every value means.

Re-derive one rating in every audit. It takes two minutes and it tells you whether the methodology is sound.

Look for values with a position and no words. The page marks them. In the Aureon pack severity 2–6 and 8 have a number on the scale and no description, so "severity 8" is ordinally meaningful and verbally undefined — two assessors cannot be expected to agree on it. That is a finding about the risk methodology, not about any one rating.

Exports

Everything you will want to attach to the file:

WhatWhere
The register/item//check/register.csv and .xlsx
A listingthe CSV button on /list
Revisions/item//revs.txt
The FMEA/item//fa-csv.csv, /item//fax.xlsx
Your own session trail/api/audit-trail?format=csv

Every one of them carries a provenance header: which database, which build, which filter, what time. A register exported without that is a table of numbers whose meaning cannot be re-established later.

3. What the chat is for

Joins, cross-cutting questions, re-derivations, and anything where you would otherwise open twenty pages. Not evidence.

Every id it gives you arrives as a link. Click it. Verification that costs one click gets done; verification that costs a search and three navigations gets skipped when the day runs late — and what gets skipped is the middle of the sample, the claims that never become findings and so are never checked by anyone.

Every answer carries a count and a complete flag. Read them. complete: false with a note means you are looking at a subset.

Every answer carries the provenance block. It should match what the app's footer says. If it does not, you are looking at two databases.

The questions the chat is actually better at:

List every requirement under  whose status is missing, with its code.

Which records mention FTS-01? Give me the id and title of each.

Show me the risk model that applies to , then re-derive its level from its
inputs and tell me whether the reported level is correct.

Show me the revision history of , with the version message of each revision.

For each claim you just made, give me the item id it rests on.

Ask for ids. An id is checkable; a paraphrase is not. Get into the habit on the first question and never drop it.

4. The step nobody can do for you: cross-reading

The two major nonconformities in the Aureon audit came from putting two records side by side. J03 says instrument FTS-01 went out of calibration on 2026-06-30. J11, the release record, says batch B-2026-17 — 1 450 Hub units — was verified on FTS-01 and released 2026-07-17. Seventeen days after the calibration lapsed.

No query connects those two records. KF has no edge between them. You find it by reading both and noticing.

This used to be the whole of the advice. Here is a method.

Which classes to cross-read

Three pairings account for most of what is found this way:

Read thisAgainst thisWhat you are looking for
Calibration / equipment recordsRelease and test recordsProduct released on an instrument that was out of calibration, unverified, or not in the register
Supplier approval and evaluationIncoming inspection, and the product it went intoMaterial accepted from a provider whose approval had lapsed, or whose evaluation never happened
Nonconformities and corrective actionsThe product, process or clause they concernA corrective action closed without effectiveness evidence; the same cause recurring under a different number

What to index on

Not titles. Titles are neither unique nor stable. Index on the things that appear verbatim in more than one record:

How to get them in front of you

1. Export the documents in scope: /list?under=&y=i → CSV. 2. For each identifier you have collected, run /list?q= and note every item it returns. Two or more hits on one identifier is a candidate. 3. Open both records in the app — not in the chat — and read them. 4. Where they conflict, check /item//revs on both to establish which statement was written first, and what was known when.

Step 2 is the one that used to have no query behind it at all.

Budget

Half a day for a small QMS, and do not compress it. This is the step that produces the findings that matter; everything before it produces the list of places to look.

5. Recording findings — and why the register moves

MCP cannot write. Findings go back into KF over its REST API or web UI, entered by the auditee, typically as events linked to the requirement they fail and tagged with their grade.

There is a structural trap here, and §0(3) is the answer to it.

A finding can only attach to a requirement the organization has taken ownership of. The requirements most in need of findings are exactly the inherited ones still reading missing — and taking one on to hang a finding from it moves it from missing to pending. So the act of recording a nonconformity makes the register appear to advance.

The fix is sequencing. Before the opening meeting, the whole register is materialised in one operation and exported:

POST /api/items//checklist/materialise-all?dry=true   # read this out first
POST /api/items//checklist/materialise-all            # then do it

or workflows/audit-open.sh snapshot , which does that, exports the register, the population and the reconciliation, and hashes the lot.

Afterwards no row's status changes because a finding was raised. On the Aureon pack this moved 38 rows from missing to pending once, before the audit, on the record — instead of moving them one at a time during it.

The same snapshot gives you the drift baseline: audit-open.sh drift lists everything modified since it was taken. Run it at the close.

6. Showing afterwards what you examined

The instance keeps a trail of every page you opened and every tool call you made: who, when, which route or tool, the arguments, and how much came back. Export it at the close:

/api/audit-trail?format=csv&user=&since=2026-07-26

It deliberately does not keep the payloads. The evidence is in the database; a second copy of the organization's records in a log file is a retention and confidentiality problem, not a benefit. If your certification body needs more than coverage, settle that in the audit plan (§0, §7) rather than discovering it afterwards.

This is what answers "how did you satisfy yourself on clause 8.4?" at surveillance.

7. Confidentiality

State the position in the audit plan; do not leave it implicit.

8. Which behaviours apply to your counts

Several behaviours changed with the build, and they change numbers, not just wording. Establish which apply from the provenance block — on the page footer, in the chat's envelope, and in the header of every export — rather than by asking the auditee.

BehaviourApplies toHow to tell
Heading rows reported pending, inflating every status tally by the number of headingsBuilds before 2026/07Compare rows_total, rows_requirements and rows_headings in the register summary: on a current build a heading carries no status at all
search by=code returned exactly one holder, silently — and it was the library definition, not the organization's ownOlder buildsOn a current build a code lookup returns every holder, oldest first
Materialised instances were given numeric codes such as 100282, so an instance may not answer to its requirement numberPacks built before mid-2026Look at the code column on the register: numeric codes on instance rows
The register's reconciliation had to be computed by handBuilds before this onereconcile=true on the register, or the summary block at the head of the register page
No listing, no last-modified outside an item, no result countsBuilds before this one/list exists

The counts quoted in this document are from a current build against the Aureon pack: 79 rows, 69 requirements, 10 headings, reconciled 98 = 98.

9. Six traps that survive into a report

The full list of KF's own quirks is in the helper's document. These six are the ones that reach your report if nobody stops them.

1. done is not conformity. It is a tag somebody applied. In the Aureon audit R30 was done with an overdue instrument in its evidence. 2. missing is not "no evidence exists". It means nobody drew the edge. 16 of 36 minor nonconformities in the reference audit were content that was present in the QMS and simply not linked. Before writing up a gap, search for the evidence by subject. 3. pending is not progress. Somebody took ownership; nothing was done. 4. Not every row is a requirement. Ten of the Aureon pack's 79 rows are headings and carry no status. Count rows_requirements, never rows_total. 5. A done row can hide a subtree. The register shows what lies beneath each row. A done cell with four missing obligations under it is not four-fifths done. 6. A risk level without its matrix is an assertion. Open /item//risk.

Every status word in the app carries its meaning as a tooltip, and the register page opens with the full legend. Use those words as the app defines them, or translate them in your report — but do not carry them across untranslated. "Thirty-eight missing" written down at 4pm becomes thirty-eight gaps by the closing meeting.

10. When to stop, and when to go and look

The verification rule

Re-reading the record in the app is mandatory, not left to conscience, for:

Everything else may be read in the chat.

The stopping rule

Coverage is not "every row examined", because done is not conformity and

missing is not a gap. Adequate coverage is:

1. Every missing row chased by subject search before it is written up. This is not optional: nearly half of them turn out to be unlinked evidence. One /list?q= per gap; the gaps that survive it are real. 2. A full evidence walk on every requirement that is release-critical, customer-facing, or the subject of a previous finding — regardless of status. 3. A sample of the remainder, drawn from the listing so the population is defensible: state the filter, the total it returned, and the row numbers you drew. /list gives you all three. 4. One re-derived risk rating, one revision diff, and one cross-read (§4) in every audit, as method checks rather than as sampling.

You are done when 1 and 2 are complete, 3 is drawn and worked, and 4 is on the record. Not when the register looks finished — it never will.

11. What KF cannot tell you

State these in your report rather than papering over them.

The honest summary: conformity as a snapshot is determinable from KF; conformity as a history is determinable for anything KF was told about; and sufficiency is always yours to judge.

Companion documents: audit-for-helpers.md (KF mechanics, for the auditee's expert), audit-criteria-validation.md (validating the requirement library against the standard), audit-engagement-checklist.md (the one-pager for the audit plan). The audit this method was drawn from: iso9001-audit.md.