Screened · automated checks passed

autoresearch

from github/awesome-copilot · ★ 36,648 on GitHub · found by our crawler 2026-07-07

What the author says it does

Autonomous iterative experimentation loop for any programming task. Guides the user through defining goals, measurable metrics, and scope constraints, then runs an autonomous loop of code changes, testing, measuring, and keeping/discarding results. Inspired by Karpathy''s autoresearch. USE FOR: autonomous improvement, iterative optimization, experiment loop, auto research, performance tuning, automated experimentation, hill climbing, try things automatically, optimize code, run experiments, autonomous coding loop. DO NOT USE FOR: one-shot tasks, simple bug fixes, code review, or tasks without a measurable metric.

Quoted from the skill's own SKILL.md trigger description — this is what tells Claude when to activate it. Not yet verified by us.

Automated screening

100/100 validator score

Scored by the same rules as our free SKILL.md validator: trigger description quality, body substance, structure. Automated — a human bench test is the next step in the pipeline.

Install (unverified — review first)

git clone https://github.com/github/awesome-copilot
# skill lives at: skills/autoresearch/SKILL.md

SkillProof status

This skill is in our test queue. We install every skill in a clean environment, run a trigger battery and score output against a baseline before it earns a catalog page — the full protocol is public. Until then, treat it like any unreviewed dependency: read the SKILL.md and any scripts before installing.

Already tested in Testing & QA

  • Screen Reader Testing ★ 9.6/10 Screen reader testing playbook for VoiceOver, NVDA, and JAWS with ARIA fixes.
  • Skill Security Auditor ★ 9.6/10 OWASP-mapped code/secrets/config audit with working bundled scan scripts
  • Web Quality Audit ★ 9.6/10 Lighthouse-style audit across Performance, Accessibility, SEO, and Best Practices with severity tiers.
  • OSINT Methodology ★ 9.6/10 Structured 5-stage external recon methodology with confidence levels, severity rubric, and OpSec rules.

Top 10 Testing skills, ranked →