Skill Test Lab -- Break Your Skill Before Buyers Do
SkillSkill
Test an agent skill against real tasks, edge cases, and permission traps before you publish it.
About
A valid SKILL.md can still produce terrible behavior.
Skill Test Lab evaluates what an agent actually does when your instructions meet realistic requests. It tests the happy path, near-misses, missing inputs, unavailable dependencies, and the permission traps that turn a useful assistant into a liability.
What you get
- A compact test matrix tied to the skill's promised outcome and boundaries
- Normal, near-match, underspecified, failure, and authorization-trap cases
- Isolated execution that stubs external mutations unless a real test is explicitly approved
- Evidence for actions attempted, artifacts produced, and observable results
- Pass, partial, fail, and not-run classifications with causes separated
- A
ready,ready_with_limits, ornot_readyverdict - The smallest revision supported by failed behavior, plus focused regression cases
- A record of what remains untested
This product tests behavior, not whether the response repeated expected words. It never publishes the skill, installs it globally, sends messages, spends money, or touches production as part of evaluation.
Use it before a marketplace launch, after a substantial skill revision, or when customers report that a skill follows its prose but misses its job. It pairs naturally with Release Candidate Audit when the package around the skill also needs inspection.
Core Capabilities
- Build realistic behavioral test matrices
- Probe routing and permission boundaries
- Separate skill failures from tool and environment failures
- Return an evidence-backed release verdict
Customer ratings
0 reviews
No ratings yet
- 5 star0
- 4 star0
- 3 star0
- 2 star0
- 1 star0
No reviews yet. Be the first buyer to share feedback.
Version History
This skill is actively maintained.
September 26, 2026
One-time purchase
$9
By continuing, you agree to the Buyer Terms of Service.
Creator
Brian Gorzelic — SpookyJuice.AI
Evidence-led tools for AI agents, software releases, and technical operations
Built from real operating workflows, with provenance, safety boundaries, and human approval designed in.
View creator profile →Details
- Type
- Skill
- Category
- Engineering
- Price
- $9
- Version
- 1
- License
- One-time purchase
Works With
Works with OpenClaw, Claude Projects, Custom GPTs, Cursor and other instruction-friendly AI tools.
Works great with
Personas that pair well with this skill.
Vibe Code Chaperone -- Because the Demo Works Is Not a Security Model
Persona
Supervise AI-built software with ownership, failure tests, and evidence before it meets production.
$9
Suite Smith Persona
Persona
The Suite Smith persona, standalone
$9
Code Explainer Persona
Persona
The Code Explainer persona, standalone
$9