The Auditman Cometh
Every child's boogeyman lives under the bed. Biopharma's lives in the gap between doing the work and proving you did it, and he arrives on a schedule. Three-quarters of FDA rejection letters cite quality and manufacturing problems, the same Form 483 findings recur year after year, and food now has twenty-four hours to produce lot-level traceability. Weekly Papers #5 on why audits are almost never lost on the science, and what it would mean for compliance to be a property of the system, not a project.
TL;DR: Companies rarely fail audits because the work was bad. They fail because the record of the work was assembled after the fact, in a different system from the one where the work happened. FDA's newly published Complete Response Letters show manufacturing and quality issues in roughly 74% of rejections, the same Form 483 citations have topped the list for years, and inspectors now scrutinise whether audit trails were reviewed rather than merely generated. The revised EU Annex 11 and the new Annex 22 on AI raise the bar again in 2026. Our position is that an audit trail which has to be reconstructed isn't an audit trail. Dalea records every state-changing action in the same transaction as the change itself, keeps a provenance graph of which run read and wrote which version of which artefact, marks each link by whether a server observed it or someone merely claimed it, and exports the result in the forms a regulator already reads.
The landscape
The inspector is not the problem
FDA carried out 1,248 drug quality assurance inspections in FY2025, and issued 303 drug and biologics warning letters, a 59% jump on the prior year [1]. That's the visible edge of enforcement. The more instructive number arrived with FDA's decision to publish its Complete Response Letters for the first time. Of 202 decision letters issued between 2020 and 2024, roughly 150, about 74%, involved quality or manufacturing issues, spanning process validation gaps, GMP non-compliance, and facility inspection findings [2][3]. More than half of the facility-related deficiencies came down to FDA being unable to complete the required pre-approval inspection at all [4].
Read that again, because it reframes the whole problem. The dominant reason a drug does not get approved on the first cycle is not that the molecule failed. It's that the evidence around the molecule did not survive contact with an inspector.
The same findings, every year
If failures were idiosyncratic, you'd expect the citation list to churn. It doesn't. Quality unit responsibilities and procedures under 21 CFR 211.22(d) has been the single most cited Form 483 observation for at least four consecutive years, joined at the top, year after year, by failure to thoroughly investigate discrepancies under 211.192, absent or unfollowed written procedures under 211.100(a), and laboratory controls under 211.160(b) [1]. Data integrity remains embedded in that most common citation, and appeared in around 15% of FY2025 warning letters overall [1].
A finding that recurs for years across thousands of facilities is not a training problem. It's a description of how the underlying systems work.
The problem
This is not a story of negligence
To be fair to the industry, nobody covered by these rules is ignoring them. Most organisations run validated systems, staff a quality unit, and pay for compliance software whose entire pitch is audit readiness. The eQMS market exists because everyone already knows the trail matters. The problem is where those systems sit: beside the work, not underneath it. A quality system can hold a perfectly controlled copy of a record that was authored somewhere else, after the fact. That seam, between where work happens and where it is documented, is exactly what inspectors keep finding.
Audits are lost in the gap, not in the lab
The boogeyman is only frightening because of what's under the bed. For most organisations, what's under the bed is a set of records that live somewhere other than where the work happened. A chromatography station with its own local storage. A spreadsheet on a shared drive. A paper logbook. An ELN that captured the conclusion but not the seventeen decisions behind it. The science can be excellent and the record still fail, because the record was never a by-product of the work. It was a separate act of authorship, performed later, by someone reconstructing what they believe took place.
Inspectors are extremely good at spotting the seam. In an April 2026 warning letter, FDA noted that the audit trail on a firm's FTIR instrument showed no activity across a week in September 2025, while the firm's own daily usage log recorded drug testing performed on that instrument during exactly that period. The same inspection found a common username and password on the HPLC used for impurity testing, and analysts holding administrator privileges that let them modify and delete data [5]. Nobody needed to prove intent. The two records simply disagreed, and only one of them was tamper-evident.
The subtler version is a system that produces a perfect audit trail nobody ever looks at. FDA cited a manufacturer in early 2025 precisely on this point, finding that production and the quality unit should have been reviewing the electronic raw data and audit trails, and were not [6]. Generating the trail is now table stakes. The emerging inspection focus is on documented evidence that someone competent reviewed it. The April 2026 letter demands, among other remediations, the firm's procedures for audit trail review [5][6].
The same failure, in clinical and in food
This isn't a manufacturing quirk. In FDA's Bioresearch Monitoring programme, the recurring clinical-investigator findings year after year are inadequate or inaccurate case histories and study records, and failure to maintain and retain records under 21 CFR 312.57 [7][8]. Different regulation, identical failure mode. The study was run, and the trail of how it was run did not hold together.
Food is arriving at the same place from the other direction. FSMA 204 requires covered entities handling Food Traceability List items to maintain Key Data Elements at every Critical Tracking Event and hand them to FDA within twenty-four hours of a request [9]. The compliance date has moved to 20 July 2028, but the obligation to produce records during an outbreak has not moved at all. FDA's own readiness tabletop exercises, run with industry between 9 March and 1 April 2026, found that most firms could respond inside the window, while flagging persistent difficulty with traceability lot codes, data standardisation, and retrieving records across multiple systems and supply chain partners [10].
Those three problems are one problem. And the cost of not solving it is unfolding in real time. As this paper goes out, FDA and CDC are tracing a multistate Cyclospora outbreak to iceberg lettuce from a single central-Mexico supplier, an outbreak that has grown to 9,481 reported illnesses across 17 states, with at least 398 hospitalisations and two deaths. The link was established through epidemiologic and traceback data: distribution records and patient interviews, working backwards through the supply chain [11].
The reconstruction tax
What companies do instead is pay a tax. Weeks or months of pre-audit remediation. Screenshots. Back-filled justifications. Someone senior reading two years of instrument logs to work out what a departed analyst meant. This is expensive, it scales with the size of the organisation, and it produces exactly the artefact that regulators are trained to distrust, which is a narrative written after the outcome was known.
And the bar is about to move. The draft revision of EU GMP Annex 11, published in July 2025, expands the guideline from five pages to nineteen across seventeen chapters, with final publication expected in 2026. It arrived alongside a revised Chapter 4 on documentation and an entirely new Annex 22 covering AI-based systems, which drew roughly 1,300 comments in consultation [12][13]. Meanwhile FDA issued its first warning letter citing AI misuse in CGMP documentation in April 2026, on the principle that AI output requires authorised human review before it can become a controlled record [14]. As we argued in Biology has no compiler, agents multiply the number of untraceable steps unless the substrate underneath them is deterministic. The regulators have now put that in writing.
The solution
First, a structured language for science
The reason records have to be reconstructed is that the work was never written down in a form anything could keep. Prose is not auditable. A protocol described in a paragraph, a result pasted into a spreadsheet, a decision explained in a comment thread, none of these can be traced by a machine, so tracing them falls to a person, later, under time pressure.
Dalea starts by giving science a structured language. Protocols, data models, results, and inventory are first-class objects with defined shapes rather than free text. That change is not primarily about tidiness. It is what makes provenance possible at all, because once work is expressed in structured terms, every operation on it can be recorded without anyone deciding to record it. This is step one, and it is the step most compliance tooling skips.
Then, recording in flight
Step two is centralisation. When the notebook, the data, the inventory, and the AI assistant all sit on the same substrate, the record is written by the act itself, at the moment the act happens. Who did it, when, what changed, and what it was before.
The word worth defending there is same. Capture is not a listener bolted onto the side of the system: each recorded run is written in the same transaction as the change it describes, so a mutation without its provenance cannot exist. When code in an analysis session pulls workspace data, the pull is recorded with its query plan and a hash of the full result before any truncation, and if the pull cannot be recorded, the call fails. No evidence, no data.
That is the difference, but it is worth being precise about what disappears and what doesn't. Most organisations treat documentation as either an operational plan, meaning a set of intentions about how work will be written up, or an aftermath, meaning an exercise in remembering. Dalea replaces both with something stricter. The record is written by the act itself, every mutation and every entity captured as it happens, and the shape of that record is decided in advance by the organisation. An environment's schema fixes which fields a result has to carry. A locked template fixes which steps of a procedure can be reordered or skipped and which fields stay open for the operator. And the server re-validates every save against that configuration rather than trusting the browser to have done it.
So the planning doesn't disappear; it moves, and it happens once. Instead of planning how each piece of work will be written up afterwards, the organisation decides up front what an auditable record of its work looks like, and every workflow inherits that decision. What disappears is the reconstruction: there is no separate write-up to forget or perform later, because the work and the record of the work are the same event, landing inside a structure the organisation chose.
Two questions, not one
An inspector asks two different kinds of question, and most systems can only answer one of them.
The first is the event question. Who did what, when, from which device, under which authentication. That is an org-level audit log, and Dalea keeps one: sign-ins, role changes, schema changes with the reason the person typed at the time, result batches closed with their e-signatures, and every export of the log itself.
The second question is harder, and it is the one that decides investigations. Where did this number come from, and what else is wrong if it turns out to be wrong. That is not an event stream, it is a graph. Dalea records runs — an import, a data pull, a code execution, a document edit, an export — together with the immutable versions they read and wrote, pinned by content hash. The graph never points at "the file"; it points at the exact bytes that run consumed. Walking backwards from an artefact gives its origin. Walking forwards from a suspect input gives its impact: the list of everything that now has to be re-examined.
Read 21 CFR 211.192 with that distinction in mind. It does not merely require that an unexplained discrepancy be investigated. It requires the investigation to extend to the other batches of the same product, and to other products, that the same failure may have touched [18]. That is a forward walk over a lineage graph. In most organisations it is performed by a person with a spreadsheet and a recollection of what was running that month, which is a fair description of why it remains a top-cited observation. The food version of the same walk is the one FDA now wants completed inside twenty-four hours.
There is a side effect here worth naming, because it speaks to the trail nobody reads [6]. A lineage graph that answers where did this come from gets used on ordinary days, by the people doing the work, to settle ordinary arguments. Review stops being a ritual performed on the audit trail once a quarter and becomes the thing the trail is for.
Not all evidence is equal
A record assembled after the outcome is known is a narrative, and regulators are trained to distrust narratives. The defence is not to insist that everything in the record is equally solid. It is to state, inside the record, how solid each part of it is.
So every link in the graph carries two marks. Quality: did a server observe this directly, or is it recorded from a producer's own statement, or was it inferred by the platform from surrounding evidence. Precision: was this exact input observed feeding this output, or is it a candidate, a file that was merely open in the session, recorded as a possible source rather than a confirmed one. Coarse links over-approximate deliberately, because the safe direction for a recall is to flag one artefact too many rather than to miss one.
The same rule governs machine work, and this is where the AI question gets an answer rather than a policy. Every run records two things about its actor: what kind of thing performed it, human or agent or service, and which named person is accountable for it. The two are never collapsed, so an agent's work reads as the agent acting on behalf of a person. And what an assistant says its lineage was is stored separately from what the server observed it do, so the two can be compared instead of merged. A hallucinated citation cannot quietly become evidence.
FDA's April 2026 letter turned on the principle that AI output needs authorised human review before it becomes a controlled record [14]. A workspace can require exactly that, and when it does, the requirement is not a procedure someone might skip under deadline: an export whose trail ends in unreviewed AI outputs refuses to render. Clearing it means a person approving the work and re-authenticating with a second factor. An assistant can pre-screen the weakest links first, and its verdicts are labelled as assistive and as not counting.
The operations report
Which makes the moment the Auditman arrives fairly boring. Because the work and the record of the work are the same event, producing evidence is a traversal and an export rather than a project. Narrow to a person, an action type, a study, a date range, or the closure behind one figure, and render it (how both layers work is documented in audit logging and working with provenance on the Dalea wiki).
The forms it renders into are deliberately not ours. A human-readable trail report, W3C PROV for machine interchange [15], IEEE 2791 BioCompute, the standard FDA has supported for sequencing-analysis submissions since 2020 and lists in its Data Standards Catalog [16], and Define-XML at dataset grain. One traversal feeds all of them, so the copy a reviewer reads and the copy a machine parses cannot disagree.
Three things travel with every export, and they exist because the first question a good auditor asks about an evidence package is what it isn't. Its scope: which direction the walk ran, how deep, and whether it was truncated, so a complete closure is distinguishable from a slice. Its AI involvement: how many runs were machine-performed, on which models, and how much of that a human has reviewed. And a verification verdict: the integrity checks re-run over the exported segment, plus how far the tamper evidence actually reaches.
That last one deserves the honesty it costs us. Every workspace's history is committed into an append-only log of the kind certificate transparency uses [17], with signed checkpoints anchored outside our own control, because a log its operator can silently rewrite proves nothing. But structural consistency is not the same as truth: a record could have been rewritten consistently before it was ever anchored. So the verdict carries a flag for whether external anchoring covers that segment yet, and when it doesn't, it says so instead of rounding up. The same applies to silence. Capture begins when the first run touches an entity, so an empty trail is reported as an empty trail, never as proof that nothing happened.
And an auditor with no Dalea access at all can still take a file we released, fetch the public keys, and check the signature on the receipt that bound those bytes to that version, years later, without asking us.
That bundle is the operations report, and it is the same artefact whichever regulator is asking. An FDA investigator on a pre-approval inspection, an EU qualified person preparing for an Annex 11 assessment, a notified body, or a customer's own auditor all want the same thing, which is a complete and honest account of who did what, when. You are handing over the record, not building one.
Nothing under the bed
The boogeyman is a story about the dark. Turn the light on and there's a pile of laundry.
Audits work the same way. The firms that dread them are the ones who genuinely don't know what an inspector will find, because their own record of the last two years is scattered across instruments, drives, and memory. The firms that don't dread them aren't better at science. They simply have nothing to reconstruct, because the reconstruction was never necessary.
The Auditman cometh, on a schedule, to everyone. He only scares the ones with something under the bed.
References
- U.S. Food and Drug Administration, drug inspection and warning-letter data, FY2025. FDA's published inspection observation (Form 483) data and warning-letter records; FY2025 volumes and recurring top citations as compiled in industry analyses of FDA data (RAPS; Pharmaceutical Online). Supports the FY2025 inspection count, the 303 warning letters (+59% vs. FY2024's 190), the recurring top 483 citations, and the ~15% data-integrity share.
- Pharma Manufacturing. FDA's CRLs reveal 74% of applications rejected for quality, manufacturing issues. July 2025. Analysis of the 202 Complete Response Letters issued 2020-2024: 150 (74%) involved quality/manufacturing issues. pharmamanufacturing.com
- U.S. Food and Drug Administration. openFDA Complete Response Letters database. The primary dataset underlying [2]. open.fda.gov
- RSM US. FDA's complete response letters underscore outsourcing and quality challenges, February 2026. "More than half of facility-related deficiencies occurred because the FDA could not complete the required preapproval inspections." rsmus.com
- U.S. Food and Drug Administration. Warning Letter, Ava Inc., MARCS-CMS 721180, 14 April 2026. Cites 21 CFR 211.68(b) and 211.194(a): FTIR audit trail showing no activity 23-30 September 2025 while the daily usage log recorded testing; a common username and password on the HPLC used for impurity testing; analysts with administrator privileges to modify and delete data; and a demand for audit-trail-review procedures. fda.gov
- ECA Academy. FDA Warning Letter on missing Audit Trails and Raw Data Review, February 2025. On the expectation that production and the quality unit review electronic raw data and audit trails. gmp-compliance.org
- Goodwin. Common FDA Bioresearch Monitoring Violations, Updates from FY 2023 to Now. On recurring findings including failure to maintain and retain records under 21 CFR 312.57.
- FDA Compliance Program Guidance 7348.811 and BIMO annual metrics. On inadequate or inaccurate case histories and study records as a leading clinical-investigator citation.
- U.S. Food and Drug Administration. FSMA Final Rule on Requirements for Additional Traceability Records for Certain Foods. Key Data Elements, Critical Tracking Events, and the 24-hour records request. fda.gov
- U.S. Food and Drug Administration. Traceability Readiness Tabletop Exercises Final Report, June 2026. Exercises run 9 March to 1 April 2026; confirms the 24-hour electronic sortable spreadsheet requirement and the compliance-date extension from 20 January 2026 to 20 July 2028; remaining challenges around lot codes, data standardisation, and cross-system retrieval. fda.gov/media/192993
- Centers for Disease Control and Prevention. Cyclospora outbreak investigation, July-August 2026. Newsroom release (single supplier identified via FDA traceback; initially 1,644+ cases in five states) and investigation updates (9,481 illnesses across 17 states, at least 398 hospitalisations, two deaths as of mid-August 2026). cdc.gov newsroom | cdc.gov investigation
- European Commission. EudraLex Volume 4: draft revised Annex 11 (Computerised Systems), new Annex 22 (Artificial Intelligence), and revised Chapter 4 (Documentation), published for consultation 7 July 2025; consultation closed 7 October 2025. health.ec.europa.eu
- European Medicines Agency. Summary notes, HMA/EMA group focused on AI - industry stakeholders meeting, February 2026. Annex 22 received ~1,300 public comments and is under revision; final document expected by end of 2026. ema.europa.eu
- U.S. Food and Drug Administration. Warning Letter, Purolea Cosmetics Lab, 320-26-58, 2 April 2026. First FDA warning letter citing misuse of AI in CGMP documentation: AI agents used to create specifications, procedures, and master production/control records without authorised human review, cited under 21 CFR 211.22(c). fda.gov See also DLA Piper, FDA Warning Letter highlights risks of using AI in drug manufacturing, April 2026.
- Lebo T, Sahoo S, McGuinness D, et al. PROV-O: The PROV Ontology. W3C Recommendation, 30 April 2013. The interchange vocabulary for activities, entities, and the agents responsible for them. w3.org/TR/prov-o
- U.S. Food and Drug Administration. Electronic Submissions; Data Standards; Support for the IEEE Bioinformatics Computations and Analyses Standard for Bioinformatic Workflows. Federal Register notice, 22 July 2020, docket FDA-2020-N-1450. Announces FDA support for the BioCompute standard (IEEE 2791-2020) in regulatory submissions and its addition to the FDA Data Standards Catalog for HTS data in NDAs, ANDAs, BLAs, and INDs to CBER, CDER, and CFSAN. federalregister.gov
- Laurie B, et al. Certificate Transparency. IETF RFC 6962, June 2013; superseded by RFC 9162, Certificate Transparency Version 2.0. The append-only Merkle-tree log with signed tree heads, from which the tamper-evidence construction is borrowed. rfc-editor.org
- 21 CFR 211.192, Production record review. Requires that any unexplained discrepancy be thoroughly investigated whether or not the batch has been distributed, and that the investigation extend to other batches of the same drug product and other drug products that may have been associated with the failure. ecfr.gov