Book 45 · Patriola’s Guide to Claude
Annotation Scheme Design
Operational protocols for reliable annotation. What deployment reveals that a pilot calibration never does — codebook, monitoring, checklist, and a versioning system that survives schema evolution.
The scheme that ran clean for six months, then didn't
A 27-label annotation scheme ran for six months without incident. The pilot calibration had passed — kappa above threshold on every primary layer, four hundred items annotated and stored. Then a phenomenon appeared in a later batch that didn't fit any existing label cleanly. Three choices presented themselves: add a new label and accept that prior items don't carry it, reclassify all 400 prior items at significant cost, or collapse the distinction into an existing label and lose the signal. No protocol covered any of them. The researcher picked one in the moment and moved on.
Six months after that, the consequences arrived. The gap in prior items meant three different things depending on when in the study the item had been collected, and those four hundred gaps were uninterpretable. Remediation took two weeks and still couldn't fully resolve the ambiguity. The original design wasn't the problem — the label resolution was sound, the multi-layer architecture was appropriate, and the calibration data was real. Missing was a versioning protocol that should have been written before the first evolution happened.
Scheme design gets treated as the hard part, and deployment gets treated as the payoff. This book covers what happens after calibration passes: the codebook that operationalizes abstract labels for annotators working under time pressure, the inter-annotator agreement monitoring that catches drift before it compounds, the deployment checklist that covers what "calibration passed" doesn't, the versioning system that makes schema evolution safe rather than destructive, and an honest account of what Claude can and cannot contribute to annotation work.
What you’ll learnThe operational layer a pilot never surfaces
- Architecture before granularity — the decision that shapes every subsequent choice, building on the label-resolution foundation from Book 20.
- Calibrating the granularity decision — how fine a label distinction should be before it stops being reliably annotatable.
- The codebook as a production document — per-label definitions, positive examples, negative examples, and boundary rules that resolve disputes without anyone's memory.
- IRR beyond the pilot — monitoring protocols that catch inter-annotator agreement drift mid-project, before it compounds into unusable data.
- Deploying for production — the checklist covering calibration-set freeze, annotator readiness, and halt triggers that "calibration passed" doesn't.
- Versioning without invalidating, and Claude in the annotation loop — a schema evolution system that survives change, plus where Claude genuinely helps and where it introduces the errors it was meant to prevent.
A preview
Documentation written after the fact is reconstruction. It reads like a record but behaves like a guess.
Practitioners whose scheme survived the pilot and now has to survive deployment
This book is for anyone running an annotation project past the pilot stage — multiple annotators, a domain that keeps evolving, and months rather than weeks of runtime. Book 20's label-architecture chapter is the assumed foundation; this book starts where calibration passes and covers everything deployment reveals that a two-week pilot never does. It does assume you already have a working schema and are past the point of wondering whether documentation discipline is worth the ten minutes it costs in the moment.
A longer excerpt is available to newsletter subscribers.
More from Patriola
New books in this series
One short email per book launch.