AI Evals
Turn a fuzzy LLM feature into a golden set, rubric, judge plan and ship/no-ship threshold
Test report
- Verdict
- Tested · Works
- Score
- Tested
- Jul 18, 2026
- Environment
- Claude Code 2.x (agent harness)
- Upstream re-checked
- Jul 30, 2026 · aff68c9
⚠ This skill is no longer available upstream. Our re-check on Aug 10, 2026 couldn't find it any more (repo unreachable/deleted). The test below is what we measured on Jul 18, 2026 and we're leaving it up as a record — but there is nothing left to install, so we've removed the command.
Asked both baseline and skill to design evals for a support-reply assistant (no PII leakage, must cite KB, must refuse unsafe requests). Baseline produced ad-hoc 'try 10 tickets and eyeball it'; the skill delivered a decision-anchored pack with a ship/no-ship threshold, >=2 cases per target behavior including adversarial/safety tags, a severity-weighted error taxonomy, a rubric with concrete anchors and tie-breakers, and judge calibration plus a regression loop. Its references (RUBRIC, TEMPLATES, CHECKLISTS) are real substance, not stubs, and the repo even ships its own with-skill/without-skill showcase.
Scored on four weighted criteria — install, triggering, output vs. baseline, docs. How scoring works
- Installs cleanly 5/5
- Triggers reliably 5/5
- Output vs. baseline 8/10
- Docs & honesty 5/5
What AI Evals does
A 7-step workflow that produces an AI Evals Pack for an LLM feature: an eval PRD with acceptance thresholds, a tagged golden test set, an error taxonomy from open coding, a behaviorally-anchored rubric, a judge/harness plan with calibration, and a reporting + iteration loop where every new failure becomes a new test. Triggers on 'design evals for our AI feature', 'build a golden set and rubric', 'make our flaky AI quality a repeatable eval'.
How to install AI Evals
Nothing to install: the source repository no longer has this skill. If the author brings it back, our daily re-check will pick it up and the command will reappear here.
Commands — how to trigger AI Evals
-
/ai-evalsTurn a fuzzy LLM feature into a golden set, rubric, judge plan and ship/no-ship threshold
It also activates on plain-language prompts like these:
-
Design evals for our support-reply assistant that must never leak PII. -
Help me build a golden test set and rubric for this LLM feature. -
Our AI quality is flaky, make it into a repeatable evaluation loop.
Frequently asked questions
- Is the AI Evals skill free?
- Yes. The skill itself is free from liqiongyu/lenny_skills_plus. SkillProof publishes the install command and an independent test verdict at no cost.
- Does AI Evals work with Claude Code?
- We tested it with Claude Code 2.x (agent harness) on Jul 18, 2026. Verdict: Tested · Works. Asked both baseline and skill to design evals for a support-reply assistant (no PII leakage, must cite KB, must refuse unsafe requests). Baseline produced ad-hoc 'try 10 tickets and eyeball it'; the skill delivered a decision-anchored pack with a ship/no-ship threshold, >=2 cases per target behavior including adversarial/safety tags, a severity-weighted error taxonomy, a rubric with concrete anchors and tie-breakers, and judge calibration plus a regression loop. Its references (RUBRIC, TEMPLATES, CHECKLISTS) are real substance, not stubs, and the repo even ships its own with-skill/without-skill showcase.
- What is the AI Evals SkillProof Score?
- 9.2/10 — installs cleanly 5/5, triggers reliably 5/5, output vs. baseline 8/10, docs & honesty 5/5.
- How do I install AI Evals?
- Copy the install command from this page, run it in your terminal, and restart Claude Code. Skills live in ~/.claude/skills/ (global) or .claude/skills/ inside a project.
- Can I use AI Evals with Cursor, Copilot, Gemini CLI, Codex or other AI tools?
- The SKILL.md format is native to Claude (Claude Code, Desktop, claude.ai). The instructions inside adapt to other assistants: Cursor rules, GitHub Copilot instructions, Windsurf rules, Custom GPTs, AGENTS.md for OpenAI Codex, and GEMINI.md for Google Gemini CLI — our conversion guides cover each, and the free converter on the tools page does the wrapping for you.