AI Evals

Turn a fuzzy LLM feature into a golden set, rubric, judge plan and ship/no-ship threshold

Tested · Works

Test report

Verdict
Tested · Works
Score
9.2/10
Tested
Jul 18, 2026
Environment
Claude Code 2.x (agent harness)
Upstream re-checked
Jul 30, 2026 · aff68c9

This skill is no longer available upstream. Our re-check on Aug 10, 2026 couldn't find it any more (repo unreachable/deleted). The test below is what we measured on Jul 18, 2026 and we're leaving it up as a record — but there is nothing left to install, so we've removed the command.

Asked both baseline and skill to design evals for a support-reply assistant (no PII leakage, must cite KB, must refuse unsafe requests). Baseline produced ad-hoc 'try 10 tickets and eyeball it'; the skill delivered a decision-anchored pack with a ship/no-ship threshold, >=2 cases per target behavior including adversarial/safety tags, a severity-weighted error taxonomy, a rubric with concrete anchors and tie-breakers, and judge calibration plus a regression loop. Its references (RUBRIC, TEMPLATES, CHECKLISTS) are real substance, not stubs, and the repo even ships its own with-skill/without-skill showcase.

Scored on four weighted criteria — install, triggering, output vs. baseline, docs. How scoring works

  • Installs cleanly 5/5
  • Triggers reliably 5/5
  • Output vs. baseline 8/10
  • Docs & honesty 5/5

What AI Evals does

A 7-step workflow that produces an AI Evals Pack for an LLM feature: an eval PRD with acceptance thresholds, a tagged golden test set, an error taxonomy from open coding, a behaviorally-anchored rubric, a judge/harness plan with calibration, and a reporting + iteration loop where every new failure becomes a new test. Triggers on 'design evals for our AI feature', 'build a golden set and rubric', 'make our flaky AI quality a repeatable eval'.

How to install AI Evals

Nothing to install: the source repository no longer has this skill. If the author brings it back, our daily re-check will pick it up and the command will reappear here.

Commands — how to trigger AI Evals

  • /ai-evals Turn a fuzzy LLM feature into a golden set, rubric, judge plan and ship/no-ship threshold

It also activates on plain-language prompts like these:

  • Design evals for our support-reply assistant that must never leak PII.
  • Help me build a golden test set and rubric for this LLM feature.
  • Our AI quality is flaky, make it into a repeatable evaluation loop.

Frequently asked questions

Is the AI Evals skill free?
Yes. The skill itself is free from liqiongyu/lenny_skills_plus. SkillProof publishes the install command and an independent test verdict at no cost.
Does AI Evals work with Claude Code?
We tested it with Claude Code 2.x (agent harness) on Jul 18, 2026. Verdict: Tested · Works. Asked both baseline and skill to design evals for a support-reply assistant (no PII leakage, must cite KB, must refuse unsafe requests). Baseline produced ad-hoc 'try 10 tickets and eyeball it'; the skill delivered a decision-anchored pack with a ship/no-ship threshold, >=2 cases per target behavior including adversarial/safety tags, a severity-weighted error taxonomy, a rubric with concrete anchors and tie-breakers, and judge calibration plus a regression loop. Its references (RUBRIC, TEMPLATES, CHECKLISTS) are real substance, not stubs, and the repo even ships its own with-skill/without-skill showcase.
What is the AI Evals SkillProof Score?
9.2/10 — installs cleanly 5/5, triggers reliably 5/5, output vs. baseline 8/10, docs & honesty 5/5.
How do I install AI Evals?
Copy the install command from this page, run it in your terminal, and restart Claude Code. Skills live in ~/.claude/skills/ (global) or .claude/skills/ inside a project.
Can I use AI Evals with Cursor, Copilot, Gemini CLI, Codex or other AI tools?
The SKILL.md format is native to Claude (Claude Code, Desktop, claude.ai). The instructions inside adapt to other assistants: Cursor rules, GitHub Copilot instructions, Windsurf rules, Custom GPTs, AGENTS.md for OpenAI Codex, and GEMINI.md for Google Gemini CLI — our conversion guides cover each, and the free converter on the tools page does the wrapping for you.