an independent review publication ยท est. 2026

We tested 90+ skills
so you don't have to.

2 arrived broken. 18+ were just okay. 85+ earned a KEEP IT. We read the docs, smoke-test what we can, and tell you which agent skills are worth your time. No sponsors. No affiliate links.

~/gearscope/reviews.log
$ tail -f ~/gearscope/reviews.log

[2026-05-15]  comfyui  .................  4/5  KEEP IT
[2026-05-17]  claude-code-design-ai  ...   โ€”   testing
[2026-05-19]  native-feel-skill  .......   โ€”   testing
[2026-05-21]  neuralinverse  ...........   โ€”   queued

# we read every doc. we run what we can. we're honest when we can't.
# no sponsors. no affiliate links.

$
// latest reviews

Recently tested

Real agent skills. Real smoke tests. Real verdicts.

TRY IT HANDS-ON

Unity Skills

Unity's first-party skill pack is written like documentation its own CLI team would sign, then drops 4 of 31 skills o...

3 / 5
KEEP IT HANDS-ON

Nuwa Skill (ๅฅณๅจฒ)

The 32,749-star thinking-distillation factory has the best eval methodology in its tier, and a quality gate that fail...

4 / 5
KEEP IT HANDS-ON

anything2explainer

A whole film studio as one skill: research, narration, storyboard, parallel build, measured QC. Toolchain verified ha...

4 / 5
IN QUEUE + 10 more

See all reviews →

๐Ÿ† PULSE (6h, Thursday 08:00): **ARCHIFY MILESTONE MORNING โ€” 65Kโ˜… CROSSED ~03:15 UTC (64,927 @02:00 โ†’ 65,284 @08:04, +357 โ‰ˆ 1,430/day re-accel; landed ~45min BEFORE the band's 04:00 early edge = 3rd straight band-edge landing) AND 90.0K installs crossed in the SAME window (88.8K โ†’ 90.0K โ‰ˆ 4.8K/day re-accel) AND archify-review 2nd-face install DEBUT 1.9K all-time (weekly bars 0ร—8 = roll pending) โ€” star-scale, install-scale and face-expansion axes LANDED TOGETHER (milestone-PAIRING is the new watch shape) ยท SOCKET AUDIT FAIL ร—3 = STRUCTURAL (Trust Hub PASS + Snyk PASS hold; first confirmed persistent audit regression on a charting mega-skill โ€” review-relevant at 65K)** ยท **K-DENSE/CMM widening #47 (lead 1,709โ†’1,710 ร—3 stable) โ€” the #253 โˆ’8 narrowing did NOT repeat = noise, race-PAUSED verdict vindicated; dormant KD +62 OUT-GAINED active CMM +33 (3rd dormant-win window)** ยท **GLAMA ACCELERATION ARTIFACT RESOLVED (270/day, below thaw band)** ยท **SMITHERY STALL #34 (flush #7 gap >139h NEW LONGEST)** ยท **GOOGLE/agents-cli Trending #21โ†’#44 (11.3Kโ†’9.5K) = cooling ร—2 CONFIRMED** ยท orca READ-MINUTE push again (08:03:51 vs ~08:05 read, ~42nd consecutive window-with-push).**, ๐Ÿ†• DISCOVERY (thin window โ€” 1 boarded + fork-less watch + CN cluster; line-level dedupe: atomica11y/navigate-business-situations/logo-concept-skill = prior log-only clusters; hyper3d-documentary/hehe-industry/simmsxd = prior logs re-seen; Paper-Reading dropped; Dungeon-Settlers twins stay excluded; yjn66627 POC-validation = competition agent SYSTEM not a skill, excluded):, ๐Ÿ“ก SIGNALS:, Signals & flags summary...

// the workflow

How we test

Every skill goes through our testing pipeline. We're transparent about what we could and couldn't test. Each review is labeled with its test depth.

01

Read the docs, check the code

We load the full SKILL.md, every script, every reference file. We check dependencies, look for footguns, and map out what the skill actually does vs what it claims.

02

Test in an isolated sandbox

Every install runs inside a clean, throwaway Linux sandbox. Nothing touches our machine, and the full session (test script, commands, raw log) is published with the review so any reader can re-run the exact test that produced the verdict. When a skill can't be sandboxed (desktop apps, GPU-heavy workloads), we test live and explain why on the page. Every review labels its tier: sandboxed, hands-on, smoke test, or desk review.

03

Rate across four dimensions

1 to 5 gears each for docs, setup, value, reliability. No inflated scores. A 3 is average. Most skills land there.

04

Ship the verdict

KEEP IT, TRY IT, SKIP IT, or BROKEN. Popular skills get SKIP IT when they deserve it. Honest beats nice.

// sound familiar?

The skill install graveyard

Every hour spent untangling a broken skill is an hour not building. Here's what we catch before you do.

The problem
How we catch it
โœ— Setup guide skipped three steps and your terminal is still throwing errors.
โ†’ We install on a clean machine and list every missing step in the review.
โœ— Two skills conflict and nobody warned you until your agent crashed mid-task.
โ†’ We test in isolation and alongside common skills, then flag the conflicts.
โœ— Every "top skills" list is just the README rephrased with extra adjectives.
โ†’ We paste the real errors, real tracebacks, and the config that broke.
โœ— Two hours configuring something that turns out not to work on macOS.
โ†’ We test on macOS, Linux, and WSL. If it's broken on yours, we say so.
// origin story

Why this exists

One weekend, one too many broken installs.

origin.js โ€” gearscope
1
2
3
4
5
6
7
8
9
10
11
// weekend project. 47 skills installed.
const broken = 12; // setup steps were wrong
const crashes = 6; // agent died on launch
const useful = 3; // actually saved time

if (broken + crashes > useful * 5) {
  buildReviewSite("GearScope");
  promise("no sponsors");
  promise("real errors");
  promise("honest ratings");
}
// frequently asked

Questions you'll probably ask

We track new releases, GitHub stars, and what's being talked about in the agent communities. If a skill is getting attention or a reader requests one, it goes on the bench. We prioritize skills people are trying to install today, not last year's leaderboard.
The moment a skill author pays us, every SKIP IT becomes a conversation we don't want to have. Reader tips and a future paid tier are the only revenue model.
Depends on the review. Our gold standard is sandboxed: we run the install inside a clean, isolated Linux sandbox and publish the test script and raw log so any reader can re-run it. When a skill can't be sandboxed (desktop apps, GPU work), the review is labeled hands-on and explains why. Smoke tests are quick verification of the basics. Desk reviews mean we read but didn't install. Every review's depth badge tells you which tier applies. Full methodology →
A 3 is average. Most skills land there. 4 means you should install it. 5 is reserved for the rare skill that's well-built and immediately useful. 1 and 2 mean the skill is broken or wasted your time. We publish the rubric next to each review.
Yes. Email reviews@gearscope.xyz or DM us on Twitter/X with the skill name and link. We can't promise we'll cover every request, but reader-requested skills jump the queue.
// weekly digest

Get the skill digest

The best and the broken from 100+ hands-on tests. One email a week.

one-click unsubscribe.
// support the mission

Keep the reviews honest

No sponsors, no affiliate links โ€” 90+ hands-on reviews and counting. If one saved you an hour, kick back a fraction of it.