wickedagilewicked-garden

the capability plane · the catalog · MIT

The tools your coding agentcan’t build alone.

Your harness already plans, swarms, and ships. wicked-garden is the capability plane it acts through — the catalog that fills only the gaps a planner-executor can’t fill alone, starting with the one that matters most: done is re-derived from evidence, never asserted.

load out garden and its peers — one command

interactive · picks your products across CLIs · direct install below

  • 143 skills
  • 14 domains
  • 40 QE specialists
  • 10 work-shapes
  • MIT · local-first
agent · buildarmed
all acceptance tests pass
the agent asserted this. the gate doesn’t trust the stamp.
re-deriving from evidence…
auto-playing · hover to hold · tap a dot to pin

01 / the toolbox

Six gaps your agent can’t close alone.

Your harness plans, swarms, and ships. These are the six things a planner-executor genuinely can’t do on its own — and they’re a sample: the full catalog runs to 142 skills across 14 domains. It plays itself; click any tool to pin it.

auto-playing the demos — click any tool to pin it and take control
“all tests pass”
REJECTED
the verifier never ran
proveevidence gate

Re-runs the proof behind the claim. A false pass is REJECTED; a missing backend FAILS CLOSED. Never a vacuous green.

garden skill

02 / the gate

Play the lying agent. Watch the gate refuse.

Garden’s one non-negotiable, made drivable. It plays itself — breaking the evidence and re-stamping the verdict. Flip a switch or pull the lever to take control: the gate re-derives the claim instead of taking your word for it.

evidence conditionsgate: armed
pull the lever — the gate stamps a verdict
auto-playing the gate — flip a switch or pull the lever to take control
claim card · archetype: build
build: all acceptance tests pass

The agent stamped this “done.” The gate re-runs the verifier, re-hashes the recording, and checks the vault is even there. Break a condition and prove it can’t be fooled.

03 / the whole catalog

Six tools were the sample. Here’s the catalog.

The toolbox shows the signature gap-fillers. Underneath sits the full surface — 143 skills across 14 domains, including the 40-specialist QE fleet absorbed from the retired wicked-testing plugin — all reading the same evidence-first discipline. Everything below is real and in the repo today.

quality engineering42 skills

The absorbed QE fleet — 40 specialists on five surfaces, the 3-agent acceptance pipeline, and the prove gate. See the wall below.

acceptplanauthorexecutereviewinsightprove
product & UX26 skills

Vague ask → SMART criteria that land in estate’s requirements graph — the forward source of truth; plus UX & a11y review, mockups, visual direction, user-signal synthesis, and the draft floor for document deliverables (contrast, page budget, cited claims).

requirements-analysisacceptance-criteriaux-reviewaccessibilitymockupstrategydraft
platform & ops16 skills

The rubric on demand — security, compliance, incident, infra, distributed traces, CI.

auditcomplianceincidentinfraobservabilitygithub-actions
engineering15 skills

Graph-driven change — rename & patch across files on estate's graph, debug from a trace, architecture, docs.

architecturedebuggingpatchsystem-designintegrationlarge-scale-migration
agentic review9 skills

Review the agentic codebase itself — topology, framework detection, trust & safety.

agentic-patternscontext-engineeringframeworksreview-methodologytrust-and-safety
orchestration9 skills

Orchestration — the governed-worker discipline every governed unit follows (creator / evaluator / neutral); swarm, worktrees, workflow runners.

governed-workerswarmworktreesworkflow
context7 skills

Pull-model context assembly on the estate-backed stack — briefing, intent, propose-skills, grounded in the repo.

discoveryintentpropose-skillsclassifyground
domain modeling4 skills

Extract the domain from the codebase — testable business rules with confidence + provenance, projected onto estate's graph as the forward source of truth.

domaindomain-extractordomain-coveragedomain-modeler
memory & knowledge4 skills

Estate-backed memory — store / recall / cited answers; ingest and session capture; plus repo-learn, which derives memories and policies as inert estate proposals. Agents propose; humans promote.

storerecallansweringestcapturerepo-learn
multi-model3 skills

The multi-model council IS the router — an independent external-model panel that classifies the work and returns a second opinion that isn’t self-grading. It absorbed the retired wicked-signals.

councilbrainstormmulti-model
data & ML2 skills

ETL, data-quality, ontology, and ML workflows under the same evidence discipline.

datadata-engineer
personas1 skills

Run any task under a named behavioral profile — a reusable review cast on demand.

persona
code intelligence2 skills

Impact & lineage over estate's real graph — the injected edges grep and a static call-graph can’t see.

blast-radiuslineagecodebase-narrator
work-shapes2 skills

Ten shapes read off each prompt — a typo to a cutover; deliberate challenges the ask first.

triagebuildmigratereviewdeliberate

One install bundles the family. Every peer is an opt-in layer you adopt when you want it — the kit works without any of them. The mem / search / patch / domain stack runs on wicked-estate, the system of record; the evidence backend the gate re-derives against (wicked-vault) installs directly; the QE pipeline ships in-catalog as the qe domain. Garden is the catalog agents act through: work rides wicked-crew’s control plane and lands in the record — where captured memories and derived policies become citable Steering doctrine only after a human promotes them.

wicked-estatethe foundation planewicked-vaultevidence backendwicked-busopt-in layerwicked-interactivedocument engine

04 / the qe domain · absorbed fleet

No agent grades its own homework.

The retired wicked-testing plugin lives on here as the qe domain — and it brought its wall. The 3-agent acceptance pipeline gives each role its own tool boundary; the reviewer reads cold evidence files only — never the executor’s context, reasoning, or stdout — so the self-grading that inflates AI PASS rates has nowhere to hide.

report card · same featureself-grade vs the wall
author grades itself
PASS
100% · every time
self-reported
independent reviewer
PASS
awaiting review…
contradicted
80%+ of self-graded PASSes don’t survive an independent read.

Illustrative of the measured self-grading gap.

The agent that wrote the code runs the tests — and calls them green.

Scripted frameworks — Playwright, pytest, k6, axe-core — only run what you already thought to test. They don’t tell you what to test, whether the tests are any good, or whether the results mean anything.

The qe domain separates authorship from judgment — and enforces it.

Crosses the wall — the reviewer reads
manifest.jsonevidence.jsonstep-N.jsoncontext.md
Bounced at the wall — never seen
executor contextchain-of-thoughtraw stdoutprior verdicts

Evidence lands in .wicked-qe/evidence/<run-id>/. Isolation is enforced by the reviewer’s context: fork boundary and an allowed-tools: Read frontmatter — the verdict is re-derived from the artifacts, never asserted by the agent that produced them.

garden skill · accept action

05 / the qe domain · the fleet

Forty specialists. Five surfaces. One contract.

The absorbed fleet: 40 qe-* specialist skills route beneath five orchestrator surfaces of the wicked-garden-qe router. Every specialist runs in an isolated forked context. Pick a surface — see who reports to it.

auto-cycling the surfaces — click one to take control
surface
qe-test-strategist
qe-risk-assessor
qe-testability-reviewer
qe-requirements-quality-analyst
qe-code-analyzer
qe-coverage-archaeologist
qe-test-impact-analyzer
qe-acceptance-test-writer
qe-test-designer
qe-test-automation-engineer
qe-contract-testing-engineer
qe-ui-component-test-engineer
qe-integration-test-engineer
qe-e2e-orchestrator
qe-fuzz-property-engineer
qe-test-data-manager
qe-incident-to-scenario-synthesizer
qe-acceptance-test-executor
qe-scenario-executor
qe-load-performance-engineer
qe-chaos-test-engineer
qe-security-test-engineer
qe-a11y-test-engineer
qe-mutation-test-engineer
qe-ai-feature-test-engineer
qe-iac-test-engineer
qe-localization-test-engineer
qe-observability-test-engineer
qe-data-quality-tester
qe-visual-regression-engineer
qe-exploratory-tester
qe-acceptance-test-reviewer
qe-semantic-reviewer
qe-test-code-quality-auditor
qe-snapshot-hygiene-auditor
qe-flaky-test-hunter
qe-release-readiness-engineer
qe-test-oracle
qe-production-quality-engineer
qe-compliance-test-engineer
40 specialists · real skills/qe-* dirsfive surfaces of one router — wicked-garden-qeevery specialist ships with context: fork

06 / build on it

Your domain. Your pack. Same gates.

The catalog is open — the built-in domains follow a naming contract anyone can follow. Ship a pack with your unique take — a fintech-risk domain, a games-QA fleet, a compliance cast — and it sits beside the built-ins, under the same evidence discipline.

skills/
acme-fintech/your domain router · user-invocable
acme-fintech-risk-modeler/fork worker · context: fork
acme-fintech-auditor/fork worker · evaluator ≠ creator
kebab-case · ≤64 chars · SKILL.md frontmatter — that’s the whole contract
one router per domain

A single user-invocable skill fronts the domain and routes its actions — the same shape as qe, mem, and product.

{vendor}-{domain}-{role} workers

Specialists are fork skills (context: fork) — isolated subagent contexts with their own tool boundaries.

evidence via vault

Your pack’s “done” goes through the same gate: claims are re-derived against wicked-vault, never asserted.

stamp the evidence gate into any repo — no garden required

07 / the bench

One command. The whole loadout.

The family installer is the fastest way in — one interactive command that picks your products across every CLI and installs the shared wicked CLI. Prefer just this plugin? The direct path is right below.

the family installer · recommended

interactive · picks your products across CLIs · ships the wicked CLI

or install just wicked-garden directly

add from marketplace
install the plugin

MIT · open-source · v12.38.1 · local-first — nothing leaves your machine.

The wicked platform

One surface. One control plane.
One catalog. One record.

Four planes, four contracts. Hover or tab through a plane to see its role and what crosses its seams — every cross-plane interaction goes through the contract, never around it.

ExperienceWhere product work happensRendering and editing surfaces — nothing semantic lives here.
wicked-studiothe surface

Brainstorm it, build it under a check nothing self-approves, then produce the doc, deck or demo.

submit intent · watch the run · answer gates — one crew API, the surface is a pure client
ControlIntent in, verified work outOrchestration, governance, gates — your coding agents as governed workers.
wicked-crewthe control plane

Evaluator ≠ creator — no agent grades its own homework. “Done” is re-derived from evidence, never asserted.

invokes skills as governed workers — deny dominates, every verdict lands in the record
CapabilityWhat agents can doSkills, tools, playbooks, councils, the QE fleet — how agents touch the record.
wicked-gardenthe catalog

Multi-model review councils, graph-aware refactors, repo playbooks — plus an open naming contract to ship your own pack.

you are here
reads & writes the record through its contract — never around it
FoundationThe system of recordCode graph · memory · knowledge · evidence · events. Zero-infra, local-first.
wicked-estatethe record

A 102-language code graph, memory, and knowledge in one binary (MCP) — including the injected edges grep never sees.

wicked-interactivethe document engine

Doc storage and lineage, HTML/PDF/PPTX rendering, demo recording. crew proxies it — you depend on it, you don’t visit it.