Testing & QA

Testing skills are judged by what they catch, so we plant defects and count. A QA skill that misses planted regressions has no business holding a verdict.

143 skills listed · 65 fully verified · top 10 ranked →

  1. 001

    AI Agent Redteam

    Red-team LLM agents: prompt injection, MCP poisoning, memory poisoning, multi-turn jailbreaks

    by hypnguyen1209 tested Jul 31, 2026

    9.2/10 Tested · Works
  2. 002

    Webapp Testing

    Claude drives your web app in a real browser and reports what breaks.

    by Anthropic tested Jun 19, 2026

    8.8/10 Tested · Works
  3. 003

    K6

    k6 load-testing reference and scenario/threshold patterns — k6 only, not Locust as the old card claimed.

    by grafana tested Jul 14, 2026

    8.0/10 Tested · Works
  4. 004

    Drift Testing

    Generates and runs OpenAPI-driven contract tests via SmartBear's Drift CLI — but `drift verify` hard-requires a PactFlow account login, undisclosed in the skill.

    by pactflow tested Jul 14, 2026

    7.6/10 Works with setup
  5. 005

    Skill Eval Harness

    Python CLI that measures a skill's causal lift: paired with/without runs, leakage lint, ablations

    by adewale tested Jul 20, 2026

    6.4/10 Works with setup
  6. 006

    OSINT Methodology

    Structured 5-stage external recon methodology with confidence levels, severity rubric, and OpSec rules.

    by elementalsouls tested Jul 10, 2026

    9.6/10 Tested · Works
  7. 007

    Screen Reader Testing

    Screen reader testing playbook for VoiceOver, NVDA, and JAWS with ARIA fixes.

    by wshobson tested Jul 10, 2026

    9.6/10 Tested · Works
  8. 008

    Skill Security Auditor

    OWASP-mapped code/secrets/config audit with working bundled scan scripts

    by eigent-ai tested Jul 10, 2026

    9.6/10 Tested · Works
  9. 009

    Web Quality Audit

    Lighthouse-style audit across Performance, Accessibility, SEO, and Best Practices with severity tiers.

    by addyosmani tested Jul 10, 2026

    9.6/10 Tested · Works
  10. 010

    Agentic Actions Auditor

    Static-audits GitHub Actions workflows for prompt-injection paths into Claude/Gemini/Codex CI agents

    by trailofbits tested Jul 10, 2026

    9.2/10 Tested · Works
  11. 011

    APK Red Team Pipeline

    End-to-end Android APK acquisition, decompile, and secret-extraction pipeline for authorized red-team work.

    by elementalsouls tested Jul 10, 2026

    9.2/10 Tested · Works
  12. 012

    Offensive Active Directory

    AD red-team methodology: Kerberoasting to ADCS ESC1-15 to DCSync, for authorized engagements

    by SnailSploit tested Jul 10, 2026

    9.2/10 Tested · Works
  13. 013

    Playwright Skill

    Custom Playwright browser automation with dev-server auto-detection and safe /tmp scripts

    by lackeyjb tested Jul 10, 2026

    9.2/10 Tested · Works
  14. 014

    CTF AI/ML

    CTF reference for ML attacks: model weight tricks, adversarial examples, LLM jailbreaks.

    by ljagiello tested Jul 10, 2026

    8.8/10 Works with setup
  15. 015

    Agent Browser

    CLI cheat-sheet for scripted Playwright browser automation: navigate, fill, screenshot, extract

    by code-yeongyu tested Jul 10, 2026

    8.4/10 Tested · Works
  16. 016

    Accessibility Audit

    WCAG 2.1 AA audit across all 4 principles with P0-P3 severity and a remediation plan

    by rampstackco tested Jul 21, 2026

    9.2/10 Tested · Works
  17. 017

    AI Security

    Offensive testing of LLM/ML systems: prompt injection, RAG poisoning, pickle-model scanning

    by hypnguyen1209 tested Jul 31, 2026

    9.2/10 Tested · Works
  18. 018

    Bug Capture

    Turns bug reports into clean, path-free GitHub issues that survive refactors.

    by rohitg00 tested Jul 14, 2026

    9.2/10 Tested · Works
  19. 019

    Conversion Signal QA

    Pre-launch QA of paid-ad tracking: UTM hygiene, conversion events, cross-platform dedup.

    by aaron-he-zhu tested Jul 20, 2026

    9.2/10 Tested · Works
  20. 020

    Designing Distributed System Tests

    Claim-driven test plans for distributed and stateful systems, tied to fault-injection scenarios

    by shenli tested Jul 31, 2026

    9.2/10 Tested · Works
  21. 021

    Relay 80 100 Workflow

    Patterns for agent-relay workflows that validate features E2E before committing

    by AgentWorkforce tested Jul 21, 2026

    9.2/10 Tested · Works
  22. 022

    Test Guard

    Reviews AI-written tests against nine rules: mock boundaries, no duplicate or empty tests

    by amElnagdy tested Jul 21, 2026

    9.2/10 Tested · Works
  23. 023

    Testing

    Writes layered tests with boundary/edge coverage, mocking rules, and flaky-test triage

    by rsmdt tested Jul 31, 2026

    9.2/10 Tested · Works
  24. 024

    Swift Testing Expert

    Guidance for writing Swift Testing tests and migrating XCTest suites

    by AvdLee tested Jul 21, 2026

    8.8/10 Tested · Works
  25. 025

    Bugcrowd Reporting

    Applies Bugcrowd-specific VRT categorization and severity-override reporting tactics.

    by elementalsouls tested Jul 13, 2026

    8.4/10 Tested · Works
  26. 026

    Security Webshells

    SecLists web-shell samples (PHP/ASP/JSP) as reference files for detection and IDS testing

    by Eyadkelleh tested Jul 21, 2026

    8.4/10 Works with setup
  27. 027

    Skill Evaluator

    Scores a SKILL.md against Anthropic best practices with a validator script and weighted rubric

    by gotalab tested Jul 21, 2026

    8.4/10 Tested · Works
  28. 028

    Terrashark

    Terraform/OpenTofu failure-mode workflow that catches identity churn and unsafe applies

    by LukasNiessen tested Jul 31, 2026

    8.4/10 Tested · Works
  29. 029

    Agf Deploying Uat

    Runbook to bring up an isolated local UAT compose stack and smoke-test it before hand-off

    by pcliangx tested Jul 21, 2026

    8.0/10 Works with setup
  30. 030

    Analyze Skill Issues

    Reads Azure blob-stored integration test results to explain why a skill's tests are failing

    by microsoft tested Jul 31, 2026

    8.0/10 Works with setup
  31. 031

    Browser Testing with DevTools

    Drives real-browser tests through Chrome DevTools MCP to catch DOM and console issues.

    by addyosmani tested Jul 11, 2026

    8.0/10 Works with setup
  32. 032

    Create Skill Test

    Scaffolds eval.yaml test files with scenarios and rubrics for agent skills.

    by dotnet tested Jul 12, 2026

    8.0/10 Tested · Works
  33. 033

    Electron

    Automate Electron desktop apps via agent-browser over the Chrome DevTools Protocol

    by fcakyon tested Jul 21, 2026

    8.0/10 Works with setup
  34. 034

    Executing Distributed System Tests

    Runs a designed distributed-systems test plan, scoring each run on a 10-state verdict taxonomy

    by shenli tested Jul 31, 2026

    8.0/10 Works with setup
  35. 035

    Active Directory Attack

    Reference playbook for attacking Windows AD: Kerberoasting, ADCS, coercion/relay, DCSync

    by hypnguyen1209 tested Jul 31, 2026

    7.6/10 Works with setup
  36. 036

    Agf Releasing Apple

    Runbook for signing, notarizing, packaging and smoke-testing Apple release builds

    by pcliangx tested Jul 21, 2026

    7.6/10 Works with setup
  37. 037

    Chipsec

    Static UEFI/BIOS firmware analysis via Intel chipsec: scan dumps for known rootkits

    by BrownFineSecurity tested Jul 21, 2026

    7.6/10 Works with setup
  38. 038

    Ghost Proxy

    Drives the reaper MITM proxy CLI to capture, search, and replay HTTP/HTTPS traffic

    by ghostsecurity tested Jul 21, 2026

    7.6/10 Works with setup
  39. 039

    Higgsfield Image Auto

    Playwright-automates AI image generation on higgsfield.ai across Soul 2.0 / Nano Banana models

    by AKCodez tested Jul 31, 2026

    7.6/10 Works with setup
  40. 040

    CI/CD and Automation

    Sets up or modifies CI/CD build and deployment pipelines with quality gates.

    by addyosmani tested Jul 11, 2026

    7.2/10 Works with setup
  41. 041

    WCAG Audit Patterns

    Runs WCAG 2.2 accessibility audits with automated checks and remediation guidance.

    by wshobson tested Jul 11, 2026

    7.2/10 Works with setup
  42. 042

    A/B Test Analysis

    Evaluates A/B test results for significance, sample size, and ship/stop recommendations.

    by phuryn tested Jul 11, 2026

    6.8/10 Works with setup
  43. 043

    API Fuzzing (Bug Bounty)

    Fuzzes REST and GraphQL APIs to find IDOR and other bug-bounty-worthy flaws.

    by zebbern tested Jul 12, 2026

    6.8/10 Works with setup
  44. 044

    Agent Benchmark

    Methodology for scoring agent output quality and catching regressions against baselines

    by vibeeval tested Jul 21, 2026

    5.6/10 Works with setup
  45. 045

    Google Ads Copy

    Generates and A/B tests Google Ads headlines and descriptions to improve CTR.

    by nowork-studio tested Jul 13, 2026

    5.2/10 Works with setup
  46. 046

    Ab Test Plan

    Designs A/B and multivariate tests, including sample size and hypothesis setup.

    by indranilbanerjee

    In test queue
  47. 047

    Ad Spend Optimizer

    Analyzes ad spend across channels and recommends reallocation to improve ROAS and CAC.

    by guia-matthieu

    In test queue
  48. 048

    Add Seed Skills

    Adds or edits QA skills in seed-skills and publishes them to the qaskills.sh catalog.

    by PramodDutta

    In test queue
  49. 049

    Adoption Report

    Report on design system adoption across teams, separating coverage from actual usage.

    by murphytrueman

    In test queue
  50. 050

    Agent Skill Author

    Guides authoring and refining a SKILL.md through an empirical, test-first process.

    by matlab

    In test queue
  51. 051

    AI Six Sigma Property Os

    Design an AI Six Sigma model for property maintenance dispatch and quality dashboards.

    by Mark393295827

    In test queue
  52. 052

    Analyst Estimates

    Pull analyst revenue and EPS estimates with ranges and coverage for a stock.

    by OctagonAI

    In test queue
  53. 053

    Arbor

    Iteratively improves a real artifact against an evaluator using hypothesis tree refinement.

    by K-Dense-AI tested Jul 11, 2026

    Tested · Didn't pass
  54. 054

    Architectural Analysis

    Assess a module's coupling, data flow, and technical debt against SOLID principles.

    by testdouble

    In test queue
  55. 055

    Auto Pr

    Automatically opens mergeable PRs against a repo by sourcing and fixing work items.

    by vouchdev

    In test queue
  56. 056

    Bug Fix

    Runs a structured bug-fix workflow producing a fix, regression test, and review gate.

    by sd0xdev

    In test queue
  57. 057

    Build Nitro Modules

    Builds React Native Nitro Modules with Nitrogen codegen and native bindings.

    by margelo

    In test queue
  58. 058

    Canvas Course Qc

    Runs a pre-semester quality check on Canvas course structure, dates, and rubrics.

    by vishalsachdev

    In test queue
  59. 059

    Copywriting Cta

    Writes and reviews end-of-article calls-to-action for blog posts and newsletters.

    by samber

    In test queue
  60. 060

    Creating An Agent

    Guides writing subagent descriptions, tool selection, and prompts for auto-delegation.

    by ed3dai

    In test queue
  61. 061

    Deep Plan

    Create sectionized, TDD-oriented implementation plans via research and multi-LLM review.

    by piercelamb

    In test queue
  62. 062

    Elves

    Runs autonomous multi-batch development overnight with PR review and testing.

    by aigorahub

    In test queue
  63. 063

    Flutter Drift

    Implement, migrate, and test Drift/SQLite persistence in Flutter apps.

    by MADTeacher

    In test queue
  64. 064

    Grace Explainer

    Explains the GRACE methodology, its semantic markup, and knowledge graph conventions.

    by osovv

    In test queue
  65. 065

    Implementing Tasks

    Executes a batch of TDD-sized tasks using red-green-refactor discipline via subagents.

    by prime-radiant-inc

    In test queue
  66. 066

    Krypton Vps Codex App

    Set up and harden a Linux VPS as an SSH host for remote Codex or Claude Code work.

    by jturntdev

    In test queue
  67. 067

    Module Audit Agent

    Audits a code module's health and flags whether it should be refactored.

    by wednesday-solutions

    In test queue
  68. 068

    Qdrant Model Migration

    Guides migrating embedding models in Qdrant without downtime.

    by qdrant

    In test queue
  69. 069

    Rails Audit Thoughtbot

    Audits a Rails codebase for testing, security, and code design conventions.

    by thoughtbot

    In test queue
  70. 070

    Simulink Requirements

    Manages requirement sets and traceability links in MATLAB's Requirements Toolbox.

    by matlab

    In test queue
  71. 071

    Skill Review

    Runs a persona-based behavioral audit of a Lattice skill to surface usage gaps.

    by techygarg

    In test queue
  72. 072

    Academic Experiments

    Audits ML experiments, blocking unverified numbers before a paper draft cites them.

    by joshua-zyy tested Jul 13, 2026

    8.8/10 Works with setup
  73. 073

    Add LLM Evals

    Playbook for wiring reference-based + LLM-as-judge evals into CI, with judge-calibration steps and named anti-patterns.

    by ContextJet-ai tested Jul 15, 2026

    9.6/10 Tested · Works
  74. 074

    Arize Dataset

    CLI-driven Arize dataset CRUD, versioning, and export for LLM eval sets.

    by github tested Jul 14, 2026

    9.6/10 Tested · Works
  75. 075

    Auto Itera

    Runs a disciplined baseline-vs-arms experiment to a sealed-test-set ship-or-kill verdict

    by clfhaha1234 tested Jul 16, 2026

    9.6/10 Tested · Works
  76. 076

    Concept Drill

    Diagnoses your weakest sub-concept, drills it live via predict/break/confirm, leaves a tracker.

    by Wamikmk tested Jul 15, 2026

    9.6/10 Tested · Works
  77. 077

    Cypress Author

    Writes, updates, and fixes Cypress E2E and component tests following house style rules.

    by cypress-io tested Jul 14, 2026

    9.6/10 Tested · Works
  78. 078

    Flutter Tester

    Flutter test patterns: layer isolation, Given-When-Then, Riverpod/Mockito/GetIt setup.

    by Harishwarrior tested Jul 14, 2026

    9.6/10 Tested · Works
  79. 079

    Jadx

    APK decompile + a systematic grep checklist for secrets, weak crypto, SQLi, and WebView holes.

    by BrownFineSecurity tested Jul 13, 2026

    9.6/10 Works with setup
  80. 080

    Jtagprobe

    SWD/JTAG debug-port pentest probe via a physical SEGGER J-Link, classifies OPEN/LOCKED/DEAD.

    by BrownFineSecurity tested Jul 13, 2026

    9.6/10 Tested · Works
  81. 081

    Mk Agent Browser

    Drives the agent-browser CLI for real Chrome automation with accessibility-tree refs and an auth vault.

    by ngocsangyem tested Jul 15, 2026

    9.6/10 Tested · Works
  82. 082

    Playwright Automation Expert

    Playwright E2E test architecture, POM, flaky-test debugging, and REST API testing.

    by jmr85 tested Jul 15, 2026

    9.6/10 Tested · Works
  83. 083

    Repo Forensics

    Offline security scanner for repos, AI skills, plugins, and MCP servers

    by alexgreensh tested Jul 17, 2026

    9.6/10 Tested · Works
  84. 084

    SkillHone Synthesis

    Turns tool exploration into hard, machine-gradable benchmark Q/A instead of trivia lookups.

    by Tencent tested Jul 14, 2026

    9.6/10 Tested · Works
  85. 085

    Swift Testing

    Swift Testing patterns, test-double taxonomy, and migration guidance for @Test-based suites.

    by bocato tested Jul 13, 2026

    9.6/10 Tested · Works
  86. 086

    Walk-Forward Leakage Auditor

    Audits time-series backtests for leakage before you trust the Sharpe

    by mospira tested Jul 30, 2026

    9.6/10 Tested · Works
  87. 087

    Adversarial Verify

    Refute your own work before presenting it: attack inputs, assumptions, evidence

    by NUX-Design tested Jul 19, 2026

    9.2/10 Tested · Works
  88. 088

    Agent Verifier (Verification)

    Orchestrated security/pattern/quality/language audit for AI agent code, with pattern-vs-heuristic tagging.

    by Aurite-ai tested Jul 14, 2026

    9.2/10 Works with setup
  89. 089

    AI Evals

    Turn a fuzzy LLM feature into a golden set, rubric, judge plan and ship/no-ship threshold

    by liqiongyu tested Jul 18, 2026

    9.2/10 Tested · Works
  90. 090

    API Tester

    Generates API tests from real OpenAPI + route code, never guesses status codes

    by laolaoshiren tested Jul 16, 2026

    9.2/10 Tested · Works
  91. 091

    Backpropagation

    Trace a runtime bug back to the spec gap, close it, and generate a fail-first regression test

    by LucasDuys tested Jul 18, 2026

    9.2/10 Tested · Works
  92. 092

    Bug Hunt

    Proactive bug hunt that reports only findings backed by a failing test

    by chrisallenlane tested Aug 7, 2026

    9.2/10 Tested · Works
  93. 093

    Delphi TDD and DUnitX

    Enforces strict Red-Green-Refactor TDD with DUnitX, fakes-via-interface, and naming conventions in Delphi.

    by delphicleancode tested Jul 14, 2026

    9.2/10 Tested · Works
  94. 094

    HealthSim

    Generates synthetic patients, claims, and pharmacy data in FHIR/HL7v2/X12/NCPDP for EMR test systems.

    by mark64oswald tested Jul 16, 2026

    9.2/10 Tested · Works
  95. 095

    IoT UART Console (Picocom)

    Pentest IoT UART consoles via picocom: enumerate, hit bootloaders, gain root shells.

    by BrownFineSecurity tested Jul 14, 2026

    9.2/10 Tested · Works
  96. 096

    PICT Test Designer

    Turns multi-parameter requirements into pairwise PICT models and test tables

    by omkamal tested Jul 17, 2026

    9.2/10 Tested · Works
  97. 097

    Plugin Release Checker

    Real Python validator for Claude Code plugin-marketplace repos — catches broken links and missing metadata.

    by joaquimscosta tested Jul 15, 2026

    9.2/10 Works with setup
  98. 098

    Research Proof

    Forces research and benchmark claims through a frozen-verifier proof ladder instead of confident prose.

    by tonyblu331 tested Jul 14, 2026

    9.2/10 Tested · Works
  99. 099

    AWS Penetration Testing

    AWS red-team cheat sheet: IAM privesc paths, metadata SSRF (incl. IMDSv2), S3/Lambda exploitation, persistence.

    by zebbern tested Jul 14, 2026

    8.8/10 Tested · Works
  100. 100

    Backend Validation

    Hurl + websocat + oauth2c workflow for validating OIDC-authenticated backend APIs and WebSockets.

    by johnkozaris tested Jul 15, 2026

    8.8/10 Tested · Works
  101. 101

    Build/Update Tests

    Risk-scores a file, then writes or fills in the unit/integration/E2E tests it's missing.

    by enbyaugust tested Jul 16, 2026

    8.8/10 Works with setup
  102. 102

    Chaos Engineer

    Chaos experiment design with runnable Litmus/Chaos Monkey manifests and a rollback safety checklist.

    by Jeffallan tested Jul 14, 2026

    8.8/10 Tested · Works
  103. 103

    Dev Test

    Runs the project's configured test command and writes a SHA-pinned pass/fail record other Hermit commands gate on.

    by gtapps tested Jul 14, 2026

    8.8/10 Works with setup
  104. 104

    FFUF Web Fuzzing

    ffuf pentest guidance plus a working results-analysis and req.txt helper

    by jthack tested Jul 17, 2026

    8.8/10 Tested · Works
  105. 105

    Godot

    Test, build, and ship Godot 4.x games with GdUnit4 and PlayGodot

    by Randroids-Dojo tested Jul 18, 2026

    8.8/10 Tested · Works
  106. 106

    iOS Reverse Engineering

    Extract iOS IPA/Mach-O, map APIs, and scan for secrets/vulns with FP filtering

    by incogbyte tested Jul 17, 2026

    8.8/10 Tested · Works
  107. 107

    Pulser

    Lints SKILL.md files against 8 rules and auto-fixes the safe ones, with backup and undo.

    by TheStack-ai tested Jul 15, 2026

    8.8/10 Works with setup
  108. 108

    Reconnaissance & OSINT Automation

    Working DNS/subdomain/tech-fingerprint recon scripts for authorized assessments

    by Masriyan tested Jul 17, 2026

    8.8/10 Tested · Works
  109. 109

    Spec-Driven Test

    Three-agent, spec-first E2E web testing with methodology-graded test cases

    by Loveforwa tested Jul 17, 2026

    8.8/10 Tested · Works
  110. 110

    Accessibility Audit Websites

    WCAG-style checklist for headings, forms, contrast, focus, ARIA, and keyboard nav.

    by TheGoat395 tested Jul 13, 2026

    8.4/10 Tested · Works
  111. 111

    AI/ML Attack Surface

    Grep-driven checklist for ML deserialization, prompt injection, and untrusted model loading

    by allsmog tested Jul 14, 2026

    8.4/10 Tested · Works
  112. 112

    Bug Fix Protocol

    Forces a red-test-first bug fix plus a mandatory audit of why the test suite missed it.

    by CodeAlive-AI tested Jul 13, 2026

    8.4/10 Tested · Works
  113. 113

    Database Sentinel

    Multi-backend DB security auditor for Supabase and MongoDB (RLS, exposed keys, CVEs)

    by Farenhytee tested Jul 17, 2026

    8.4/10 Tested · Works
  114. 114

    Skill Scorer

    Rubric-based 100-point quality scorer for Claude/Cursor/OpenClaw SKILL.md files, with anti-fabrication rules.

    by AndrewNgGirl tested Jul 14, 2026

    8.4/10 Tested · Works
  115. 115

    Test-Driven Development (Addy Osmani)

    Enforces RED-GREEN-REFACTOR and a test-first 'Prove-It Pattern' for every bug fix or behavior change.

    by kevinnft tested Jul 16, 2026

    8.4/10 Works with setup
  116. 116

    Visual Tester

    Annotated-snapshot UI reviews with severity-tagged findings via agent-browser.

    by incubrain tested Jul 14, 2026

    8.4/10 Works with setup
  117. 117

    AD Attacks Reference

    Active Directory attack reference: BloodHound, Kerberos, ACL abuse, ADCS ESC1-8

    by mukul975 tested Jul 17, 2026

    8.0/10 Works with setup
  118. 118

    Attack Path Architect

    Turns recon data into MITRE ATT&CK-mapped kill chains, scored and ranked.

    by Orizon-eu tested Jul 15, 2026

    8.0/10 Works with setup
  119. 119

    Cultivar

    CLI that measures whether an agent skill actually beats a no-skill baseline

    by pinecone-io tested Jul 18, 2026

    8.0/10 Works with setup
  120. 120

    Dev Dry Run

    Smoke-tests the harness-evolver pipeline offline (syntax/argparse/cross-ref) or online with a mock agent.

    by raphaelchristi tested Jul 14, 2026

    8.0/10 Works with setup
  121. 121

    Eval Consistency

    Scores a persona's roleplay replies on 5 dimensions and reports a 0-100 number

    by YIKUAIBANZI tested Aug 5, 2026

    8.0/10 Works with setup
  122. 122

    Foreman Debug

    Root-cause-first debugging discipline: reproduce, trace to source, fix with a test

    by VisionForge-OU tested Jul 17, 2026

    8.0/10 Tested · Works
  123. 123

    Recipe: Add Integration Tests

    Derives integration/E2E tests from your Design Doc instead of from the code

    by shinpr tested Jul 28, 2026

    8.0/10 Works with setup
  124. 124

    Scientific Hypothesis Generation

    Turns an observation into 3-5 literature-grounded, falsifiable hypotheses with experiment designs and predictions.

    by TobiasBlask tested Jul 15, 2026

    8.0/10 Works with setup
  125. 125

    Semantic Trap Detector

    Meta-skill that scans a SKILL.md's wording for words with over-wide semantic boundaries.

    by Jumbo-WJB tested Jul 14, 2026

    8.0/10 Tested · Works
  126. 126

    Shield Security Orchestrator

    Merges Semgrep, gitleaks and npm audit into one scored security report

    by alissonlinneker tested Jul 30, 2026

    8.0/10 Tested · Works
  127. 127

    WP Admin Browser

    Drives real WordPress admin actions via Chrome DevTools MCP with safety rails.

    by mralaminahamed tested Jul 14, 2026

    8.0/10 Works with setup
  128. 128

    AI Output Validation

    Schema-first validation layer for AI outputs before they hit downstream systems

    by DevelopersGlobal tested Jul 19, 2026

    7.6/10 Tested · Works
  129. 129

    Logic MSO Analysis

    Decodes UART/SPI/I2C/1-Wire from Saleae Logic MSO binary exports for CTF and hardware RE.

    by BrownFineSecurity tested Jul 13, 2026

    7.6/10 Works with setup
  130. 130

    QAMap PR QA

    Runs a local zero-LLM QA pass over a PR diff and picks the validation command

    by IvoryCanvas tested Jul 30, 2026

    7.6/10 Works with setup
  131. 131

    Testcontainers for .NET

    Integration tests in .NET against real Docker containers, on current 4.x APIs

    by testcontainers tested Jul 30, 2026

    7.6/10 Works with setup
  132. 132

    CI/CD Pipelines

    CI/CD design, caching, DevSecOps scanning and pipeline debugging

    by ahmedasmar tested Jul 16, 2026

    7.2/10 Works with setup
  133. 133

    IoTNet

    PCAP/live-capture wrapper that fingerprints MQTT, CoAP, Zigbee, Modbus and flags unencrypted or weakly-authed IoT traffic.

    by BrownFineSecurity tested Jul 13, 2026

    7.2/10 Works with setup
  134. 134

    Weak-Agent Test (docx-cli Harness)

    Adversarial Haiku harness that stress-tests docx-cli on 6 real document tasks

    by kklimuk tested Jul 17, 2026

    7.2/10 Works with setup
  135. 135

    Kernel CVE Analysis

    Android/AOSP kernel CVE lookups by version, branch, or date via the remote Dr. Binary vulnerability database.

    by DeepBitsTechnology tested Jul 14, 2026

    6.8/10 Works with setup
  136. 136

    SkillsGuard

    Static scanner that audits a SKILL.md for injection, exfiltration and escalation

    by Teycir tested Aug 7, 2026

    6.8/10 Works with setup
  137. 137

    BTP and BA2 CLI

    CLI command reference for uploading firmware, scanning, and pulling BA2 archives on Binarly's BTP.

    by vulhunt-re tested Jul 16, 2026

    6.4/10 Works with setup
  138. 138

    API Auditor

    Sample skill that pings a URL with HEAD and prints only the status code

    by saeed-vayghan tested Aug 5, 2026

    Tested · Didn't pass
  139. 139

    Claude Model Fingerprint

    Fake-Claude API detector whose 4.6-era answer keys misjudge newer models

    by rt22766 tested Jul 17, 2026

    Tested · Didn't pass
  140. 140

    Common AppSec Patterns

    Orchestrator that fans out XSS, CSRF, injection and prototype-pollution testing subagents

    by Stickman230 tested Jul 17, 2026

    Tested · Didn't pass
  141. 141

    Consensus Loop Audit

    Hands your repo to Codex for review, with sandbox bypassed by default

    by berrzebb tested Aug 7, 2026

    Tested · Didn't pass
  142. 142

    Ipaship Audit

    Sends your .ipa/.apk to ipaship.com's hosted AI to flag App Store/Play policy violations.

    by atharvnaik1 tested Jul 14, 2026

    Tested · Didn't pass
  143. 143

    Reins

    Deterministic PreToolUse hooks blocking destructive commands — needs manual repair

    by pegasi-ai tested Jul 16, 2026

    Tested · Didn't pass

Can't find what you need?

Request a skill — we'll test or build it

Tell us the job. Within a few days you get back a link to a tested skill that already does it, a SKILL.md you can build from, or a ready-to-use prompt. Popular asks get built for the catalog.