Testing & QA
Testing skills are judged by what they catch, so we plant defects and count. A QA skill that misses planted regressions has no business holding a verdict.
143 skills listed · 65 fully verified · top 10 ranked →
143 skills
- 001 Tested · Works
Red-team LLM agents: prompt injection, MCP poisoning, memory poisoning, multi-turn jailbreaks
- 002 Tested · Works
Claude drives your web app in a real browser and reports what breaks.
- 003 Tested · Works
k6 load-testing reference and scenario/threshold patterns — k6 only, not Locust as the old card claimed.
- 004 Works with setup
Generates and runs OpenAPI-driven contract tests via SmartBear's Drift CLI — but `drift verify` hard-requires a PactFlow account login, undisclosed in the skill.
- 005 Works with setup
Python CLI that measures a skill's causal lift: paired with/without runs, leakage lint, ablations
- 006 Tested · Works
Structured 5-stage external recon methodology with confidence levels, severity rubric, and OpSec rules.
- 007 Tested · Works
Screen reader testing playbook for VoiceOver, NVDA, and JAWS with ARIA fixes.
- 008 Tested · Works
OWASP-mapped code/secrets/config audit with working bundled scan scripts
- 009 Tested · Works
Lighthouse-style audit across Performance, Accessibility, SEO, and Best Practices with severity tiers.
- 010 Tested · Works
Static-audits GitHub Actions workflows for prompt-injection paths into Claude/Gemini/Codex CI agents
- 011 Tested · Works
End-to-end Android APK acquisition, decompile, and secret-extraction pipeline for authorized red-team work.
- 012 Tested · Works
AD red-team methodology: Kerberoasting to ADCS ESC1-15 to DCSync, for authorized engagements
- 013 Tested · Works
Custom Playwright browser automation with dev-server auto-detection and safe /tmp scripts
- 014 Works with setup
CTF reference for ML attacks: model weight tricks, adversarial examples, LLM jailbreaks.
- 015 Tested · Works
CLI cheat-sheet for scripted Playwright browser automation: navigate, fill, screenshot, extract
- 016 Tested · Works
WCAG 2.1 AA audit across all 4 principles with P0-P3 severity and a remediation plan
- 017 Tested · Works
Offensive testing of LLM/ML systems: prompt injection, RAG poisoning, pickle-model scanning
- 018 Tested · Works
Turns bug reports into clean, path-free GitHub issues that survive refactors.
- 019 Tested · Works
Pre-launch QA of paid-ad tracking: UTM hygiene, conversion events, cross-platform dedup.
- 020 Tested · Works
Designing Distributed System Tests
Claim-driven test plans for distributed and stateful systems, tied to fault-injection scenarios
- 021 Tested · Works
Patterns for agent-relay workflows that validate features E2E before committing
- 022 Tested · Works
Reviews AI-written tests against nine rules: mock boundaries, no duplicate or empty tests
- 023 Tested · Works
Writes layered tests with boundary/edge coverage, mocking rules, and flaky-test triage
- 024 Tested · Works
Guidance for writing Swift Testing tests and migrating XCTest suites
- 025 Tested · Works
Applies Bugcrowd-specific VRT categorization and severity-override reporting tactics.
- 026 Works with setup
SecLists web-shell samples (PHP/ASP/JSP) as reference files for detection and IDS testing
- 027 Tested · Works
Scores a SKILL.md against Anthropic best practices with a validator script and weighted rubric
- 028 Tested · Works
Terraform/OpenTofu failure-mode workflow that catches identity churn and unsafe applies
- 029 Works with setup
Runbook to bring up an isolated local UAT compose stack and smoke-test it before hand-off
- 030 Works with setup
Reads Azure blob-stored integration test results to explain why a skill's tests are failing
- 031 Works with setup
Drives real-browser tests through Chrome DevTools MCP to catch DOM and console issues.
- 032 Tested · Works
Scaffolds eval.yaml test files with scenarios and rubrics for agent skills.
- 033 Works with setup
Automate Electron desktop apps via agent-browser over the Chrome DevTools Protocol
- 034 Works with setup
Executing Distributed System Tests
Runs a designed distributed-systems test plan, scoring each run on a 10-state verdict taxonomy
- 035 Works with setup
Reference playbook for attacking Windows AD: Kerberoasting, ADCS, coercion/relay, DCSync
- 036 Works with setup
Runbook for signing, notarizing, packaging and smoke-testing Apple release builds
- 037 Works with setup
Static UEFI/BIOS firmware analysis via Intel chipsec: scan dumps for known rootkits
- 038 Works with setup
Drives the reaper MITM proxy CLI to capture, search, and replay HTTP/HTTPS traffic
- 039 Works with setup
Playwright-automates AI image generation on higgsfield.ai across Soul 2.0 / Nano Banana models
- 040 Works with setup
Sets up or modifies CI/CD build and deployment pipelines with quality gates.
- 041 Works with setup
Runs WCAG 2.2 accessibility audits with automated checks and remediation guidance.
- 042 Works with setup
Evaluates A/B test results for significance, sample size, and ship/stop recommendations.
- 043 Works with setup
Fuzzes REST and GraphQL APIs to find IDOR and other bug-bounty-worthy flaws.
- 044 Works with setup
Methodology for scoring agent output quality and catching regressions against baselines
- 045 Works with setup
Generates and A/B tests Google Ads headlines and descriptions to improve CTR.
- 046 In test queue
Designs A/B and multivariate tests, including sample size and hypothesis setup.
- 047 In test queue
Analyzes ad spend across channels and recommends reallocation to improve ROAS and CAC.
- 048 In test queue
Adds or edits QA skills in seed-skills and publishes them to the qaskills.sh catalog.
- 049 In test queue
Report on design system adoption across teams, separating coverage from actual usage.
- 050 In test queue
Guides authoring and refining a SKILL.md through an empirical, test-first process.
- 051 In test queue
Design an AI Six Sigma model for property maintenance dispatch and quality dashboards.
- 052 In test queue
Pull analyst revenue and EPS estimates with ranges and coverage for a stock.
- 053 Tested · Didn't pass
Iteratively improves a real artifact against an evaluator using hypothesis tree refinement.
- 054 In test queue
Assess a module's coupling, data flow, and technical debt against SOLID principles.
- 055 In test queue
Automatically opens mergeable PRs against a repo by sourcing and fixing work items.
- 056 In test queue
Runs a structured bug-fix workflow producing a fix, regression test, and review gate.
- 057 In test queue
Builds React Native Nitro Modules with Nitrogen codegen and native bindings.
- 058 In test queue
Runs a pre-semester quality check on Canvas course structure, dates, and rubrics.
- 059 In test queue
Writes and reviews end-of-article calls-to-action for blog posts and newsletters.
- 060 In test queue
Guides writing subagent descriptions, tool selection, and prompts for auto-delegation.
- 061 In test queue
Create sectionized, TDD-oriented implementation plans via research and multi-LLM review.
- 062 In test queue
Runs autonomous multi-batch development overnight with PR review and testing.
- 063 In test queue
Implement, migrate, and test Drift/SQLite persistence in Flutter apps.
- 064 In test queue
Explains the GRACE methodology, its semantic markup, and knowledge graph conventions.
- 065 In test queue
Executes a batch of TDD-sized tasks using red-green-refactor discipline via subagents.
- 066 In test queue
Set up and harden a Linux VPS as an SSH host for remote Codex or Claude Code work.
- 067 In test queue
Audits a code module's health and flags whether it should be refactored.
- 068 In test queue
Guides migrating embedding models in Qdrant without downtime.
- 069 In test queue
Audits a Rails codebase for testing, security, and code design conventions.
- 070 In test queue
Manages requirement sets and traceability links in MATLAB's Requirements Toolbox.
- 071 In test queue
Runs a persona-based behavioral audit of a Lattice skill to surface usage gaps.
- 072 Works with setup
Audits ML experiments, blocking unverified numbers before a paper draft cites them.
- 073 Tested · Works
Playbook for wiring reference-based + LLM-as-judge evals into CI, with judge-calibration steps and named anti-patterns.
- 074 Tested · Works
CLI-driven Arize dataset CRUD, versioning, and export for LLM eval sets.
- 075 Tested · Works
Runs a disciplined baseline-vs-arms experiment to a sealed-test-set ship-or-kill verdict
- 076 Tested · Works
Diagnoses your weakest sub-concept, drills it live via predict/break/confirm, leaves a tracker.
- 077 Tested · Works
Writes, updates, and fixes Cypress E2E and component tests following house style rules.
- 078 Tested · Works
Flutter test patterns: layer isolation, Given-When-Then, Riverpod/Mockito/GetIt setup.
- 079 Works with setup
APK decompile + a systematic grep checklist for secrets, weak crypto, SQLi, and WebView holes.
- 080 Tested · Works
SWD/JTAG debug-port pentest probe via a physical SEGGER J-Link, classifies OPEN/LOCKED/DEAD.
- 081 Tested · Works
Drives the agent-browser CLI for real Chrome automation with accessibility-tree refs and an auth vault.
- 082 Tested · Works
Playwright E2E test architecture, POM, flaky-test debugging, and REST API testing.
- 083 Tested · Works
Offline security scanner for repos, AI skills, plugins, and MCP servers
- 084 Tested · Works
Turns tool exploration into hard, machine-gradable benchmark Q/A instead of trivia lookups.
- 085 Tested · Works
Swift Testing patterns, test-double taxonomy, and migration guidance for @Test-based suites.
- 086 Tested · Works
Audits time-series backtests for leakage before you trust the Sharpe
- 087 Tested · Works
Refute your own work before presenting it: attack inputs, assumptions, evidence
- 088 Works with setup
Orchestrated security/pattern/quality/language audit for AI agent code, with pattern-vs-heuristic tagging.
- 089 Tested · Works
Turn a fuzzy LLM feature into a golden set, rubric, judge plan and ship/no-ship threshold
- 090 Tested · Works
Generates API tests from real OpenAPI + route code, never guesses status codes
- 091 Tested · Works
Trace a runtime bug back to the spec gap, close it, and generate a fail-first regression test
- 092 Tested · Works
Proactive bug hunt that reports only findings backed by a failing test
- 093 Tested · Works
Enforces strict Red-Green-Refactor TDD with DUnitX, fakes-via-interface, and naming conventions in Delphi.
- 094 Tested · Works
Generates synthetic patients, claims, and pharmacy data in FHIR/HL7v2/X12/NCPDP for EMR test systems.
- 095 Tested · Works
Pentest IoT UART consoles via picocom: enumerate, hit bootloaders, gain root shells.
- 096 Tested · Works
Turns multi-parameter requirements into pairwise PICT models and test tables
- 097 Works with setup
Real Python validator for Claude Code plugin-marketplace repos — catches broken links and missing metadata.
- 098 Tested · Works
Forces research and benchmark claims through a frozen-verifier proof ladder instead of confident prose.
- 099 Tested · Works
AWS red-team cheat sheet: IAM privesc paths, metadata SSRF (incl. IMDSv2), S3/Lambda exploitation, persistence.
- 100 Tested · Works
Hurl + websocat + oauth2c workflow for validating OIDC-authenticated backend APIs and WebSockets.
- 101 Works with setup
Risk-scores a file, then writes or fills in the unit/integration/E2E tests it's missing.
- 102 Tested · Works
Chaos experiment design with runnable Litmus/Chaos Monkey manifests and a rollback safety checklist.
- 103 Works with setup
Runs the project's configured test command and writes a SHA-pinned pass/fail record other Hermit commands gate on.
- 104 Tested · Works
ffuf pentest guidance plus a working results-analysis and req.txt helper
- 105 Tested · Works
Test, build, and ship Godot 4.x games with GdUnit4 and PlayGodot
- 106 Tested · Works
Extract iOS IPA/Mach-O, map APIs, and scan for secrets/vulns with FP filtering
- 107 Works with setup
Lints SKILL.md files against 8 rules and auto-fixes the safe ones, with backup and undo.
- 108 Tested · Works
Reconnaissance & OSINT Automation
Working DNS/subdomain/tech-fingerprint recon scripts for authorized assessments
- 109 Tested · Works
Three-agent, spec-first E2E web testing with methodology-graded test cases
- 110 Tested · Works
WCAG-style checklist for headings, forms, contrast, focus, ARIA, and keyboard nav.
- 111 Tested · Works
Grep-driven checklist for ML deserialization, prompt injection, and untrusted model loading
- 112 Tested · Works
Forces a red-test-first bug fix plus a mandatory audit of why the test suite missed it.
- 113 Tested · Works
Multi-backend DB security auditor for Supabase and MongoDB (RLS, exposed keys, CVEs)
- 114 Tested · Works
Rubric-based 100-point quality scorer for Claude/Cursor/OpenClaw SKILL.md files, with anti-fabrication rules.
- 115 Works with setup
Test-Driven Development (Addy Osmani)
Enforces RED-GREEN-REFACTOR and a test-first 'Prove-It Pattern' for every bug fix or behavior change.
- 116 Works with setup
Annotated-snapshot UI reviews with severity-tagged findings via agent-browser.
- 117 Works with setup
Active Directory attack reference: BloodHound, Kerberos, ACL abuse, ADCS ESC1-8
- 118 Works with setup
Turns recon data into MITRE ATT&CK-mapped kill chains, scored and ranked.
- 119 Works with setup
CLI that measures whether an agent skill actually beats a no-skill baseline
- 120 Works with setup
Smoke-tests the harness-evolver pipeline offline (syntax/argparse/cross-ref) or online with a mock agent.
- 121 Works with setup
Scores a persona's roleplay replies on 5 dimensions and reports a 0-100 number
- 122 Tested · Works
Root-cause-first debugging discipline: reproduce, trace to source, fix with a test
- 123 Works with setup
Derives integration/E2E tests from your Design Doc instead of from the code
- 124 Works with setup
Scientific Hypothesis Generation
Turns an observation into 3-5 literature-grounded, falsifiable hypotheses with experiment designs and predictions.
- 125 Tested · Works
Meta-skill that scans a SKILL.md's wording for words with over-wide semantic boundaries.
- 126 Tested · Works
Merges Semgrep, gitleaks and npm audit into one scored security report
- 127 Works with setup
Drives real WordPress admin actions via Chrome DevTools MCP with safety rails.
- 128 Tested · Works
Schema-first validation layer for AI outputs before they hit downstream systems
- 129 Works with setup
Decodes UART/SPI/I2C/1-Wire from Saleae Logic MSO binary exports for CTF and hardware RE.
- 130 Works with setup
Runs a local zero-LLM QA pass over a PR diff and picks the validation command
- 131 Works with setup
Integration tests in .NET against real Docker containers, on current 4.x APIs
- 132 Works with setup
CI/CD design, caching, DevSecOps scanning and pipeline debugging
- 133 Works with setup
PCAP/live-capture wrapper that fingerprints MQTT, CoAP, Zigbee, Modbus and flags unencrypted or weakly-authed IoT traffic.
- 134 Works with setup
Weak-Agent Test (docx-cli Harness)
Adversarial Haiku harness that stress-tests docx-cli on 6 real document tasks
- 135 Works with setup
Android/AOSP kernel CVE lookups by version, branch, or date via the remote Dr. Binary vulnerability database.
- 136 Works with setup
Static scanner that audits a SKILL.md for injection, exfiltration and escalation
- 137 Works with setup
CLI command reference for uploading firmware, scanning, and pulling BA2 archives on Binarly's BTP.
- 138 Tested · Didn't pass
Sample skill that pings a URL with HEAD and prints only the status code
- 139 Tested · Didn't pass
Fake-Claude API detector whose 4.6-era answer keys misjudge newer models
- 140 Tested · Didn't pass
Orchestrator that fans out XSS, CSRF, injection and prototype-pollution testing subagents
- 141 Tested · Didn't pass
Hands your repo to Codex for review, with sandbox bypassed by default
- 142 Tested · Didn't pass
Sends your .ipa/.apk to ipaship.com's hosted AI to flag App Store/Play policy violations.
- 143 Tested · Didn't pass
Deterministic PreToolUse hooks blocking destructive commands — needs manual repair
Can't find what you need?
Request a skill — we'll test or build it
Tell us the job. Within a few days you get back a link to a tested skill that already does it, a SKILL.md you can build from, or a ready-to-use prompt. Popular asks get built for the catalog.
Request received — check your inbox for confirmation.
Describe the need in a sentence or two, and use a real email.