Fighting AI Slop in Production Codebases: How Huntress Improved Fable 5.1's API Recall
Huntress, Tuesday, September 22nd, 2026
Huntress raised Claude Fable 5.1's API-recall accuracy from about 48% to 100% on its evals using three small changes to its coding agent harness.
Huntress tested Anthropic's Claude Fable 5.1 against its coding evaluations and started with roughly 48% accuracy on API recall, whether an agent uses a language's built-in features instead of reinventing them, a known 'AI slop' problem.
Adding one short rule to the model's CLAUDE.md instructions alone raised accuracy to 86%. Adding a lookup tool for installed Ruby gems and having a second model review the diff before the agent finished its work closed the remaining gap, reaching 100% across seven eval tasks.
The post cites the Rails team's 'Agents on Rails' benchmark, which had found Claude Fable 5.1's API recall at 41%, the best score among frontier models.