LabelOwlAI RESEARCH TOOLS
Create account
LABELOWL / RESEARCH RESOURCES

Coding evidence, with its limits

LabelOwl is an AI-assisted survey coding tool operated by Shengxing Yang. The following experiments examine specific coding conditions; they are not a general product accuracy score.

Experiments dated 20 September 2026. Page reviewed 22 September 2026.

Constructed boundary cases: 29 of 30 expected label sets

Thirty constructed cases used eight category questions, jev-1.13.0 and a 0.5 threshold across three arms. Codebook only matched 25/30 expected sets; adding question and scope matched 29/30; adding explicit boundaries, hierarchy and examples also matched 29/30. Expected sets and fixtures were frozen before calls, with no post-result retuning.

This is an information-ablation diagnostic, not representative respondent sampling. One vague answer remained unresolved against the expected policy. More context is not always better, and 29/30 must not be advertised as general accuracy.

Separate public-survey richer-context experiment

Thirty responses were evaluated against all 35 published categories: 21/30 exact label sets, 47 matched labels, 9 extra labels and 4 missing labels. Micro precision was 83.9%, recall 92.2% and F1 87.9%, relative to the dataset author's published reference coding.

The reference was not independently adjudicated as a human gold standard. Context came from the study title and questionnaire headers; exact original open-question wording was unavailable. These experiments should not be pooled.

Source: Kene David Nwosu, R Basics and Beyond dataset, CC BY 4.0. Zenodo ↗

What is publicly available?

This site publishes aggregate results and a methods note. Frozen protocols, inputs and outputs remain in the local research workspace; respondent-level data and private benchmark files are not distributed here.

Read the original methods note

Apply a codebook-based workflow

Preview access and benefits