// book 41
patriola.com

Book 41 · Patriola’s Guide to Claude

Behavioral Annotation Pipelines


Reproducible multi-tier annotation workflows. A formal ethogram and decision codebook, built before the data accumulates, with Claude enforcing the contract session after session.

Buy Ebook on Amazon

Patriola's Guide to Claude — Behavioral Annotation Pipelines: Reproducible Multi-Tier Annotation Workflows
What this book is

The coding scheme has to live outside the researcher's head

Six months of observation sessions, four behavioral categories, timestamps logged on a tablet — and nothing reproducible, because the coding scheme existed entirely in the researcher's head. Which vocalizations counted, what duration threshold separated a sustained gaze from a glance, how to handle overlapping events. Two sessions coded on different weeks produced contradicting counts, and there was no way to determine which one was correct, because correct had never been defined outside working memory. That drift is invisible until data forces a confrontation with it.

A reproducible annotation pipeline is not overhead. It is a technical precondition for trustworthy longitudinal data. This book documents building the formalization before the data accumulates: an ethogram that defines the behavioral categories, and a decision codebook that resolves the edge cases — what a vocalization includes, what threshold separates tiers, how to handle simultaneous activations across streams. ELAN handles time-coded tier annotation at the video level. BORIS handles ethogram coding with defined categories and exportable data. Claude's role is contract enforcement: given the codebook and the annotated data, checking whether a session's coding actually matches the criteria defined before the first observation was made.

The worked example is a non-speaking adult supported by AAC technology, whose communicative signal is distributed across four simultaneous behavioral streams — gaze direction, vocalizations, device activations, and physical indicators — where an activation during a sustained gaze at a communication partner carries different weight than one during gaze aversion. Annotating one stream at a time loses exactly the interaction the study needs to capture, which is why the pipeline has to hold all four tiers in simultaneous alignment, consistently, across sessions separated by weeks or months.

What you’ll learn

Building the ethogram before the data forces the question

  • Why formalization comes first — the compounding cost of deferring the ethogram and codebook until "the categories stabilize," and why that's usually too late.
  • ELAN's multi-tier structure — time-coded annotation across simultaneous behavioral streams, and why single-stream coding loses the interactions that matter.
  • BORIS ethogram coding — defined categories, frequency counts, and TSV export into downstream analysis.
  • The decision codebook — resolving edge cases explicitly, once, rather than re-deciding them session by session from memory.
  • Claude as contract enforcement — checking session coding against pre-defined criteria, making drift detectable instead of invisible.
  • Auditable judgment — why the pipeline doesn't replace researcher judgment, it makes that judgment checkable against a fixed standard.
From the book

A preview

Nothing about the sessions themselves had changed. Implicit criteria had drifted — and there was no external record of the decision rules to check them against.
You Given this codebook and this session's coded tiers, does anything here contradict a rule we defined before observation started?
Claude Checking each coded event against the codebook's explicit criteria rather than against what feels consistent with prior sessions is the whole point — memory drifts quietly, a written rule doesn't. Where a coded event falls into a genuinely ambiguous case the codebook doesn't cover yet, that's a gap in the codebook to fix, not a judgment call to make silently in the moment and hope it matches last week's.
Who it’s for

Researchers coding behavioral data across sessions

This book is for researchers running observational studies with multi-tier behavioral data, especially longitudinal work where sessions are separated by weeks or months and informal consistency isn't good enough. It assumes no prior ELAN or BORIS experience, but does assume you're past the point of wondering whether formalizing a coding scheme is worth the time — this book is about how to do it before the data accumulates, not why.

A longer excerpt is available to newsletter subscribers.

Buy Ebook on Amazon

Stay current

New books in this series

One short email per book launch.