06
Who Signs
What happens when a human has to sign for what a machine proposed?
The applied case. Where the first five modules meet a domain that already has a written specification for what a valid signature requires.
6 lessons · 9 checks · 58 min
Why it sits here
Last because it needs all of it. The harness proposes, the workflow makes the wait durable, the compiled rule is what the agent invokes, the audit trail is what the signature emits. And it is the only module where the evidence will actively argue with your design instincts.
Lessons
- The contradiction at the centreThe industry has an answer to the signing problem. The best evidence says the answer makes reviewers worse.10m
- You are probably building the wrong layerThe only RCT in this space decomposes where the benefit comes from, and the split is not close.10m
- Your users are not the users in the literatureOne finding inverts a design assumption everything else rests on.8m
- What makes a review a controlAccounting already wrote the specification. Apply it and every shipped surface fails.10m
- The baseline runs backwardsWhat actually happens today when a human posts a wrong entry. The comparison everyone skips.8m
- The build orderEvidence-led rather than taste-led. Two of the five obvious design laws are contradicted.12m
Pre-reading
14 sources · ~7h if you read all of itRanked, not exhaustive. The lessons stand on their own, so treat the first tier as the genuinely load-bearing sources and the rest as depth when a lesson makes you want it.
Read properly
- How we contain Claude across productsAnthropic EngineeringThe 93% figure is both the most dangerous fact for the argument and the most useful. Empirical proof that approval surfaces decay.20 min
- When combinations of humans and AI are usefulVaccaro, Almaatouq & Malone · Nature Human Behaviour 8(12)Meta-analysis of 106 studies. Human-AI combinations performed worse than the better of either alone, with losses concentrated in decision tasks.35 min
- In Search of VerifiabilityFok & Weld · arXiv:2305.07722Turns "should we show an explanation" into "we should show the invoice". Makes your domain the exception rather than another instance of the problem.30 min
- The flaws of policies requiring human oversight of government algorithmsBen Green · Computer Law & Security Review 45The rigorous version of "your screen is theatre", with the demand that oversight be empirically justified rather than assumed. Sections 3 and 4 carry it.45 min
- Achieving Effective Internal Control Over Generative AICOSO, February 2026The reliance / non-reliance distinction that decides whether your model version lands in audit scope. Read that section, the reconciliation example, and Case 2.60 min
- Staff Audit Practice Alert No. 11PCAOB, 2013Pages 19 to 27 only. "Verifying that a review was signed off provides little or no evidence by itself" is still the most quotable line in the domain.25 min
- To Trust or to Think: cognitive forcing functionsBuçinca, Malaya & Gajos, CSCW 2021 · arXiv:2102.09692Three shippable interventions with measured effects, and the honest finding that the most effective ones test worst with users.45 min
- Moral Crumple ZonesMadeleine Clare Elish, Engaging STSThe aviation and nuclear cases are the substance, not the coinage. Concrete failure shapes to design against.40 min
- The Missing Layer Between AI Agents and the People Who Manage ThemChristopher Kobar, April 2026Names the gap for a design audience and stops before the domain. No finance, no regulatory frame, no built artifact, no engagement with the theatre critique.15 min
Skim one section
- Quality and Intent FiltersMediaWiki / Wikipedia ORESThe best production answer to the confidence question anywhere, and it is on a wiki. Bands named in the reviewer’s language with their precision published.10 min
- Measurement of clinical diagnostic accuracy with and without AI assistanceJabbour et al · JAMA 330(23)+4.4 points from a good model against −11.3 from a biased one, with explanations recovering almost none. The case for designing around failure, in one line.20 min
- Management review controlsGrant Thornton, November 2025The readable modern restatement of the precision requirement, and the source of the authorisation-versus-MRC boundary your screen sits on.30 min
- Who Signs the Opinion? Agentic AI, Accountabilitypreprints.org, July 2026The same framing applied to statutory actuarial opinions. Corroboration from an adjacent regulated profession, stopping at the doctrinal question.Blocks automated fetching; open in a browser.30 min