Screened · automated checks passed

awq-quantization

from Orchestra-Research/AI-Research-SKILLs · ★ 10,762 on GitHub · found by our crawler 2026-07-07

What the author says it does

Activation-aware weight quantization for 4-bit LLM compression with 3x speedup and minimal accuracy loss. Use when deploying large models (7B-70B) on limited GPU memory, when you need faster inference than GPTQ with better accuracy preservation, or for instruction-tuned and multimodal models. MLSys 2024 Best Paper Award winner.

Quoted from the skill's own SKILL.md trigger description — this is what tells Claude when to activate it. Not yet verified by us.

Automated screening

100/100 validator score

Scored by the same rules as our free SKILL.md validator: trigger description quality, body substance, structure. Automated — a human bench test is the next step in the pipeline.

Install (unverified — review first)

git clone https://github.com/Orchestra-Research/AI-Research-SKILLs
# skill lives at: 10-optimization/awq/SKILL.md

SkillProof status

This skill is in our test queue. We install every skill in a clean environment, run a trigger battery and score output against a baseline before it earns a catalog page — the full protocol is public. Until then, treat it like any unreviewed dependency: read the SKILL.md and any scripts before installing.

Already tested in Token Efficiency

  • Kill Dev Process ★ 9.6/10 Classifies and safely kills orphaned dev servers/browsers with a hard exclusion list for databases and the IDE.
  • Clean Cache ★ 9.6/10 Tiered, confirmation-gated cache cleanup for Flutter/Android/iOS/Node workspaces — verified not to touch lockfiles or source.
  • Agent Prompt Engineering ★ 9.6/10 A five-part checklist (role, paths, deliverables, constraints, comms) for writing sub-agent prompts that need zero follow-up.
  • Tidy Skill ★ 9.6/10 Read-only hygiene audits that stop agents littering repos with plan.md/todo.md junk.

Top 10 Efficiency skills, ranked →