We tested 47 skills
so you don't have to.
12 had broken setups. 6 crashed on launch. Only 3 actually saved us time. We read the docs, smoke-test what we can, and tell you which agent skills are worth your time. No sponsors. No affiliate links.
Recently tested
Real agent skills. Real smoke tests. Real verdicts.
Matt Pocock Skills
Matt Pocock's opinionated skill set brings grilling, deep-module design, and a tracker-backed ticket flow to coding a...
googleapis/mcp-toolbox
The biggest official-vendor agent-infrastructure blind-spot GearScope had not reviewed. Google ships a single Go bina...
Graphify
The 89,000-star graph engine that maps your codebase as a real graph (no embeddings) and lets agents query, path, and...
See all reviews →
๐ THE STORY: HEADROOM ENDS GRAPHIFY'S EIGHT-PEAT (+521 โ 60,483โ โ leads the gain board for the FIRST TIME, snapping Graphify's 8-scan streak) + ponytail's 66-window gap-widening streak BROKEN (headroom outgained ponytail, gap narrowed 116) + Graphify still pulls away from caveman (gap 663 โ 864) + NVIDIA/SkillSpector 93rd consecutive + Smithery directory scrape, NVIDIA/SkillSpector 93rd consecutive (+29 โ 13,447โ โ RECORD extended a 43RD TIME past 50; microsoft/SkillOpt narrowed the gap a 4TH consecutive scan โ 274 โ 257), NEW ESTABLISHED MCP TRACKS (10): borski/travel-hacking-toolkit + classfang/ssh-mcp-server + wshobson/maverick-mcp + RyanAlberts/best-of-Agent-Harnesses + Lyellr88/marm-memory + DAWNCR0W/affine-mcp-server + dmmulroy/overseer + finite-sample/rmcp + Qoyyuum/mcp-metatrader5-server + zilliztech/mcp-server-milvus [OFFICIAL], NEW ESTABLISHED SKILL TRACKS (10): JackyST0/awesome-agent-skills + JimLiu/baocut + Peiiii/nextclaw + contentful/skill-kit [OFFICIAL] + fallow-rs/fallow-skills + ollygarden/opentelemetry-agent-skills + simota/agent-skills + mblode/agent-skills + HLND2T/CS2_VibeSignatures + PixVerseAI/skills [OFFICIAL]...
How we test
Every skill goes through our testing pipeline. We're transparent about what we could and couldn't test. Each review is labeled with its test depth.
Read the docs, check the code
We load the full SKILL.md, every script, every reference file. We check dependencies, look for footguns, and map out what the skill actually does vs what it claims.
Test in an isolated sandbox
Every install runs inside a clean, throwaway Linux sandbox. Nothing touches our machine, and the full session (test script, commands, raw log) is published with the review so any reader can re-run the exact test that produced the verdict. When a skill can't be sandboxed (desktop apps, GPU-heavy workloads), we test live and explain why on the page. Every review labels its tier: sandboxed, hands-on, smoke test, or desk review.
Rate across four dimensions
1 to 5 gears each for docs, setup, value, reliability. No inflated scores. A 3 is average. Most skills land there.
Ship the verdict
KEEP IT, TRY IT, SKIP IT, or BROKEN. Popular skills get SKIP IT when they deserve it. Honest beats nice.
The skill install graveyard
Every hour spent untangling a broken skill is an hour not building. Here's what we catch before you do.
Why this exists
One weekend, one too many broken installs.
2
3
4
5
6
7
8
9
10
11
const broken = 12; // setup steps were wrong
const crashes = 6; // agent died on launch
const useful = 3; // actually saved time
if (broken + crashes > useful * 5) {
buildReviewSite("GearScope");
promise("no sponsors");
promise("real errors");
promise("honest ratings");
}
Questions you'll probably ask
Get the skill digest
Five skills that mattered. Honest ratings. Every Friday.
Keep the reviews honest
No sponsors means we answer only to you. Chip in if a review saved you an hour.