the capability plane · the catalog · MIT
The tools your coding agentcan’t build alone.
Your harness already plans, swarms, and ships. wicked-garden is the capability plane it acts through — the catalog that fills only the gaps a planner-executor can’t fill alone, starting with the one that matters most: done is re-derived from evidence, never asserted.
load out garden and its peers — one command
interactive · picks your products across CLIs · direct install below
- 143 skills
- 14 domains
- 40 QE specialists
- 10 work-shapes
- MIT · local-first
01 / the toolbox
Six gaps your agent can’t close alone.
Your harness plans, swarms, and ships. These are the six things a planner-executor genuinely can’t do on its own — and they’re a sample: the full catalog runs to 142 skills across 14 domains. It plays itself; click any tool to pin it.
Re-runs the proof behind the claim. A false pass is REJECTED; a missing backend FAILS CLOSED. Never a vacuous green.
02 / the gate
Play the lying agent. Watch the gate refuse.
Garden’s one non-negotiable, made drivable. It plays itself — breaking the evidence and re-stamping the verdict. Flip a switch or pull the lever to take control: the gate re-derives the claim instead of taking your word for it.
The agent stamped this “done.” The gate re-runs the verifier, re-hashes the recording, and checks the vault is even there. Break a condition and prove it can’t be fooled.
03 / the whole catalog
Six tools were the sample. Here’s the catalog.
The toolbox shows the signature gap-fillers. Underneath sits the full surface — 143 skills across 14 domains, including the 40-specialist QE fleet absorbed from the retired wicked-testing plugin — all reading the same evidence-first discipline. Everything below is real and in the repo today.
The absorbed QE fleet — 40 specialists on five surfaces, the 3-agent acceptance pipeline, and the prove gate. See the wall below.
Vague ask → SMART criteria that land in estate’s requirements graph — the forward source of truth; plus UX & a11y review, mockups, visual direction, user-signal synthesis, and the draft floor for document deliverables (contrast, page budget, cited claims).
The rubric on demand — security, compliance, incident, infra, distributed traces, CI.
Graph-driven change — rename & patch across files on estate's graph, debug from a trace, architecture, docs.
Review the agentic codebase itself — topology, framework detection, trust & safety.
Orchestration — the governed-worker discipline every governed unit follows (creator / evaluator / neutral); swarm, worktrees, workflow runners.
Pull-model context assembly on the estate-backed stack — briefing, intent, propose-skills, grounded in the repo.
Extract the domain from the codebase — testable business rules with confidence + provenance, projected onto estate's graph as the forward source of truth.
Estate-backed memory — store / recall / cited answers; ingest and session capture; plus repo-learn, which derives memories and policies as inert estate proposals. Agents propose; humans promote.
The multi-model council IS the router — an independent external-model panel that classifies the work and returns a second opinion that isn’t self-grading. It absorbed the retired wicked-signals.
ETL, data-quality, ontology, and ML workflows under the same evidence discipline.
Run any task under a named behavioral profile — a reusable review cast on demand.
Impact & lineage over estate's real graph — the injected edges grep and a static call-graph can’t see.
Ten shapes read off each prompt — a typo to a cutover; deliberate challenges the ask first.
One install bundles the family. Every peer is an opt-in layer you adopt when you want it — the kit works without any of them. The mem / search / patch / domain stack runs on wicked-estate, the system of record; the evidence backend the gate re-derives against (wicked-vault) installs directly; the QE pipeline ships in-catalog as the qe domain. Garden is the catalog agents act through: work rides wicked-crew’s control plane and lands in the record — where captured memories and derived policies become citable Steering doctrine only after a human promotes them.
04 / the qe domain · absorbed fleet
No agent grades its own homework.
The retired wicked-testing plugin lives on here as the qe domain — and it brought its wall. The 3-agent acceptance pipeline gives each role its own tool boundary; the reviewer reads cold evidence files only — never the executor’s context, reasoning, or stdout — so the self-grading that inflates AI PASS rates has nowhere to hide.
Illustrative of the measured self-grading gap.
The agent that wrote the code runs the tests — and calls them green.
Scripted frameworks — Playwright, pytest, k6, axe-core — only run what you already thought to test. They don’t tell you what to test, whether the tests are any good, or whether the results mean anything.
The qe domain separates authorship from judgment — and enforces it.
manifest.jsonevidence.jsonstep-N.jsoncontext.mdEvidence lands in .wicked-qe/evidence/<run-id>/. Isolation is enforced by the reviewer’s context: fork boundary and an allowed-tools: Read frontmatter — the verdict is re-derived from the artifacts, never asserted by the agent that produced them.
05 / the qe domain · the fleet
Forty specialists. Five surfaces. One contract.
The absorbed fleet: 40 qe-* specialist skills route beneath five orchestrator surfaces of the wicked-garden-qe router. Every specialist runs in an isolated forked context. Pick a surface — see who reports to it.
context: fork06 / build on it
Your domain. Your pack. Same gates.
The catalog is open — the built-in domains follow a naming contract anyone can follow. Ship a pack with your unique take — a fintech-risk domain, a games-QA fleet, a compliance cast — and it sits beside the built-ins, under the same evidence discipline.
A single user-invocable skill fronts the domain and routes its actions — the same shape as qe, mem, and product.
Specialists are fork skills (context: fork) — isolated subagent contexts with their own tool boundaries.
Your pack’s “done” goes through the same gate: claims are re-derived against wicked-vault, never asserted.
07 / the bench
One command. The whole loadout.
The family installer is the fastest way in — one interactive command that picks your products across every CLI and installs the shared wicked CLI. Prefer just this plugin? The direct path is right below.
interactive · picks your products across CLIs · ships the wicked CLI
or install just wicked-garden directly
MIT · open-source · v12.38.1 · local-first — nothing leaves your machine.
One surface. One control plane.
One catalog. One record.
Four planes, four contracts. Hover or tab through a plane to see its role and what crosses its seams — every cross-plane interaction goes through the contract, never around it.
Brainstorm it, build it under a check nothing self-approves, then produce the doc, deck or demo.
Evaluator ≠ creator — no agent grades its own homework. “Done” is re-derived from evidence, never asserted.
Multi-model review councils, graph-aware refactors, repo playbooks — plus an open naming contract to ship your own pack.
A 102-language code graph, memory, and knowledge in one binary (MCP) — including the injected edges grep never sees.
Doc storage and lineage, HTML/PDF/PPTX rendering, demo recording. crew proxies it — you depend on it, you don’t visit it.