1 article with this tag.
The AI judge scored 60% of responses at Level 4. Manual review put the real number at 15%. We published the gap.