# LabelOwl context diagnostic — 20 September 2026

29/30 means exact agreement with predeclared expected label sets on 30 constructed boundary cases, with question and study context supplied. It is NOT an estimate of general accuracy on real respondents.

The same eight category questions, jev-1.13.0 and threshold 0.5 were used in all three arms. Codebook only: 25/30 exact sets; question + scope: 29/30; additional explicit boundaries/hierarchy/examples: 29/30. Expectations and fixtures were frozen before calls, with no post-result retuning. The minimal arm deliberately omits the actual question and study scope. One vague answer remains unresolved against the expected policy. This is an information-ablation diagnostic. More text is not always better.

A separate public-survey richer-context experiment used 30 responses and all 35 published categories: 21/30 exact sets, 47 matched labels, 9 extra labels, 4 missing labels, micro precision 83.9%, recall 92.2%, F1 87.9%. Reference coding is the dataset author's published coding, not an independently adjudicated human gold standard. Source: Kene David Nwosu, R Basics and Beyond dataset, https://zenodo.org/records/19857767, CC BY 4.0. Context derived from the study title and questionnaire headers; the exact original open-ended question wording was unavailable. These different experiments should not be pooled or presented as one product accuracy score.

The original frozen protocols, inputs and outputs remain in the local research workspace; this website distributes this aggregate methods note only, not respondent data or private benchmark files.
