Screened · automated checks passed

webcrawler-deep-crawl

from browser-act/skills · ★ 4,444 on GitHub · found by our crawler 2026-07-07

What the author says it does

Deep-crawl any website from start URLs, return per-page LLM-ready text/markdown/HTML plus metadata (title, description, author, language, canonical URL, OG) and in-scope outbound links. Use when user mentions deep crawl website, recursive crawl, crawl a whole site, scrape entire website, scrape docs site, scrape documentation, scrape knowledge base, scrape blog, build RAG corpus, build vector database from website, knowledge base for chatbot, GPT knowledge files, llms.txt, sitemap crawl, BFS crawl, scrape with depth or page limit, include exclude URL globs, remove boilerplate, strip navigation header footer, website to markdown, website to text, multi-page extraction, bulk page scraping, clean markdown from URL, docs site to markdown corpus, site to clean corpus. Also applies to building RAG pipelines, indexing a customer site, syncing docs into a vector store, generating training corpora from any docs hub, or expanding a single start URL into a clean corpus of every reachable in-scope page.

Quoted from the skill's own SKILL.md trigger description — this is what tells Claude when to activate it. Not yet verified by us.

Automated screening

100/100 validator score

Scored by the same rules as our free SKILL.md validator: trigger description quality, body substance, structure. Automated — a human bench test is the next step in the pipeline.

Install (unverified — review first)

git clone https://github.com/browser-act/skills
# skill lives at: solutions/search-research/webcrawler-deep-crawl/SKILL.md

SkillProof status

This skill is in our test queue. We install every skill in a clean environment, run a trigger battery and score output against a baseline before it earns a catalog page — the full protocol is public. Until then, treat it like any unreviewed dependency: read the SKILL.md and any scripts before installing.

Already tested in SEO & Search

  • SEO Audit ★ 9.6/10 Structured technical/on-page/content SEO audit that explicitly flags a real trap: static fetch tools can't see JS-injected schema markup.
  • Programmatic SEO ★ 9.6/10 12-playbook framework for templated pages at scale (locations, comparisons, integrations, glossary...) with an explicit thin-content guardrail checklist.
  • Schema Markup ★ 9.6/10 JSON-LD generator/validator for schema.org rich-results markup.
  • AI SEO ★ 9.6/10 GEO: AI-crawler access checks, extractable content structure, citation tracking.

Top 10 SEO skills, ranked →