Reliability check on my own dataset's annotation layer: five machine raters, one definition, answers from 0 to 78

I publish a small Turkish narrative corpus that ships a machine-generated annotation layer — six binary craft features per scene. Those flags were never validated against human judgement, so I ran three studies to check them, and the result is unflattering enough that I want it on the record rather than buried.

Setup: 120 scenes scored against my own blind labels (Study 1); then 100 disjoint scenes scored against an independent volunteer’s labels, locked before any machine ran (Study 2: rule-based detector, Gemini 2.5 Flash, Grok; Study 2b: Claude Fable 5 High, ChatGPT 5.5, identical protocol).

The finding, on the one rule that requires real inference — whether an abstract state has been rendered as a concrete physical detail:

Grok Gemini 2.5 ChatGPT 5.5 Detector Claude Fable 5 Human
Positives / 100 0 1 40 72 78 9

Cohen’s κ at or near chance for five of six labellers (0.004, 0.015, 0.000, 0.019, 0.027). Meanwhile ChatGPT posted the highest raw agreement in any study, 84.5% — because five of six rules have lopsided human distributions (positives: 0/1/9/96/99/44), so raw agreement mostly measures willingness to say “absent.” Four of my six rules were effectively not tested at all. That is a defect of my evaluation set, not of the models.

Two readings survive and I cannot separate them with one human rater per study: either the feature genuinely requires inference beyond current automatic raters, or the definition is not operational enough for anyone — including my human — to apply consistently. My own criterion drifted mid-pass in Study 1, which is evidence for the second.

What I would actually like from this forum is a second independent human rater. Everything needed is open: 100 scene texts, six definitions, locked human labels, all model label files, the ten prompt blocks, ID mapping and scoring scripts, in evaluation/.

Paper: How Reliable Are LLM Annotations? A Three-Study Benchmark Dataset: leventbulut/objective-projection · Datasets at Hugging Face

Disclosure: I wrote the rules and was the rater in Study 1. Claude is both one of the scored systems and was used to prepare the analysis scripts and the paper — the arithmetic is reproducible from published files, the framing is not neutral.

The disagreement between 0, 1, 40, 72, and 78 positives looks like a strong signal that the definition should be tested as an annotation instrument before comparing more automatic raters.

I would run a small calibration stage before asking a second human to label all 100 scenes:

  1. Select 15–20 scenes covering clear positives, clear negatives, and ambiguous cases.
  2. Ask both human raters to label them independently.
  3. Require each rater to record a short reason or the exact evidence span.
  4. Compare disagreements and rewrite the definition into a short decision tree.
  5. Repeat the calibration on a new small batch.
  6. Only then label the full evaluation set independently.

It may also help to add an uncertain or insufficient evidence option during calibration. Forcing an ambiguous example into a binary label can hide a definition problem as apparent model error.

Because several rules have extremely unbalanced positive rates, I would report per-class precision/recall and prevalence alongside raw agreement and Cohen’s kappa. A deliberately balanced diagnostic subset could also test whether each rule is operational, while the naturally distributed set estimates real-world behavior.

The evidence-span requirement may be especially useful here. If two raters choose the same label for different reasons, agreement on the binary flag can still overstate agreement on the actual concept.

This is the most concrete annotation-methodology advice the thread has had, and most of it is going in. One piece I want to argue with, because the order you propose would cost me the thing the second rater is for.

The ordering problem. Calibrating the definition with both raters, rewriting it into a decision tree, and only then labelling the full set is standard practice and it produces a better instrument. But it also changes the instrument before the diagnostic runs. The open question is whether the definition as published is operational — that’s what separates “the feature needs inference beyond current raters” from “my definition isn’t tight enough for anyone.” The machines were run against the published definition; my Study 1 labels were made against it. If the second human labels against a calibrated v2, their labels aren’t comparable to either, and I’ve improved the rule at the cost of never learning whether the original was the problem.

So I’d split it into two passes, in this order:

Pass 1, diagnostic. Second independent rater, same 100 scenes, definitions unchanged, no calibration, no contact with me beyond the written rules. This is the comparison that adjudicates the two readings.

Pass 2, instrument development. Everything you describe — calibration batch spanning clear positives, clear negatives and ambiguous cases, recorded reasons, disagreement review, decision tree, second calibration batch. That produces the v2 rule, and v2 requires re-running the machine raters too, since a definition change invalidates comparison with the existing model labels.

Pass 2 is the better use of effort for anyone who wants the corpus to be usable. Pass 1 is the one that settles what the published paper couldn’t, and it stops being possible the moment the definition moves.

Evidence spans: taken, and into Pass 1 as well. Your point that two raters can pick the same label for different reasons is the sharpest thing in your comment. It’s cheap to add without touching comparability: the rater records the span alongside the binary, the binary stays the unit of comparison with the machines, and span overlap gets reported separately as agreement on the concept rather than on the flag. If the binary agrees and the spans don’t, that’s a finding in itself and the current design would have missed it entirely.

The uncertain option: yes, but split by pass. In Pass 2, a full three-way option is right. In Pass 1 I’d keep the forced binary, because the machines answered binary and a third category breaks the comparison — but record confidence as a separate field. Then kappa gets reported twice, over everything and over the confident subset. That catches the failure mode you’re describing, forced ambiguity masquerading as model error, without losing comparability.

Per-class precision, recall and prevalence: partly done, and I’ll fix the presentation. Per-feature precision and recall are in the reliability paper; the prevalence figures are in the corpus paper, which means nobody sees them side by side. They’ll be in one table in the next version.

The balanced diagnostic subset: agreed, and there’s now a power argument for it. With both raters at nine positives in 100, exact Fisher needs at least three joint hits before it clears p < .05 — two joint hits gives p = .19, three gives p = .033 at kappa about 0.27. So the naturally distributed set cannot distinguish weak-but-real agreement from chance at this prevalence, whatever the second rater finds. A balanced subset raises that ceiling for the same n. The cost is that it needs its own machine runs, so it’s a second study rather than a re-analysis.

This matters beyond your comment: a separate critique elsewhere pointed out that none of my registered protocols carried a power calculation at all. They will now, including this one, before the second rater starts rather than after.

The offer stands and everything is still open in evaluation/. If you’d like to be credited for the evidence-span requirement and the calibration design, tell me how you’d prefer to be listed, and whether anything about your account or affiliation should be stated alongside it — I’m recording provenance for external contributions rather than just names now.