AI-workflow generator · v0.21.0 · MIT
An AI‑workflow generator that publishes its own report card.
It reads your codebase and forges a project-specific workflow — rules, skills, commands, agents, MCP config, memory — then measures the pieces it ships against a pre-registered harness and commits the raw scores, weak grades included.
claude plugin marketplace add directiveforge/directiveforge claude plugin install directiveforge The re-measure · v0.20.0
v0.21.0 · released 2026-09-11 · The proportionality release
A hook set that reacts to a whole-suite test run.
The generated workflow wires a PreToolUse gate on Bash. When a whole-suite test run it recognises is about to start while only unmapped files have been touched, the gate refuses that one call, names a scoped command that runs instead, and writes the refusal to a log kept outside the worktree. A runner it does not recognise is allowed with a logged warning; an explicit override opens the gate and records that it was opened. The hooks README states the boundary in its own first line — these hooks add friction and a record, and they are not a lock.
The most it may claim, quoted from the hooks README §1:
On the pinned CLI, the generated workflow refuses a whole-suite test run it recognises — pytest, python -m pytest, a venv path to either, behind the wrappers it strips — while only unmapped files are touched, and names a scoped command that runs; a runner it does not recognise is allowed with a logged warning. Deliberate evasion is open.
The most it may claim — no larger sentence. hooks/README.md §1
What was measured — each figure with its n
- Step 0 — 8 of 16
- Replayed against 16 recorded scope-failure transcripts under the real emitted risk map, the shipped gate refused 8 of 16 (n = 16, one replay each, single-run directional). The other 8 were unlocked by a mapped edit — a touched router, model or service file inside a mapped risk area, which the gate reads as licence for the whole suite. CHANGELOG.md § v0.21.0
- K1″ — 0 refusals on the PASS tasks
- The first sweeps of the 10 run-1 PASS tasks, replayed through the same gate and the same map, were refused 0 times — 0 / 8 of the tasks that reached a whole-suite argv (single-run directional). CHANGELOG.md § v0.21.0
- The first pilot — 5 S + 5 C tasks per arm
- In a pilot of 5 standard + 5 consequential tasks per arm the gate refused 4 sweeps in 2 tasks, and 0 of 5 consequential tasks (Wilson 95% [0.00, 0.43]). The S leg moved 1/5 → 3/5, Wilson 95% [0.04, 0.62] and [0.23, 0.88] — the two intervals overlap, so the move is directional, not established. Two S tasks passed through a mapped router edit. The pilot is too small for a delta-table row. CHANGELOG.md § v0.21.0
- After a refusal — 4 of 4, and 0 of 4
- By the reader's labels all 4 post-refusal events were COMPLY:sentence — the agent followed the refusal's own sentence — and none was RE-SPELL, RATIONALISE or FORGE. That sentence named a root form the gate then refused, so on the CHANGELOG's other reading the compliance cleared the gate 0 of 4 times. The fallback was fixed after the pilot — it now names the mirrored test subdirectory or the file to add, never a testpaths root — and no live run of the fixed path exists yet. The key read was not blind, and the C key was unmeasurable under the pinned tools. CHANGELOG.md § v0.21.0
- What a whole suite costs over a scoped run
- 2.12× wall-clock and 32.9× output bytes on the Python code fixture; 0.99× and 29.1× on the docs fixture (n = 1 fixture each, median of three repeats, single-run directional). CHANGELOG.md § v0.21.0
Not shipped, not verified
The Stop instrument gate — NOT WIRED, NOT VERIFIED
A second hook was designed to hold the turn open while a touched risk area’s instrument fails. Its own reproduction read a crashing instrument as a pass, so the file ships unwired and the kit makes no claim for it. Quoted from the hooks README §9b:
Everything in §9 describes the gate as designed. It is not what ships. The file is in this directory, the kit's settings.json template carries no Stop entry for it, and a generated project that adds one fails the kit's own post-generation checklist (§24(n)).hooks/README.md §9b
- The CI templates
- Deferred to v0.22 — not in this release. CHANGELOG.md § v0.21.0
- What it does not do
- The residual list is published, not summarised away: a same-OS-user agent can rewrite the hooks, the state directory and the logs; a whitespace edit in a mapped path unlocks the whole suite for the session; a runner behind a wrapper the gate does not know is allowed and logged. Read §10 before relying on any of it. hooks/README.md §10
The v0.20.0 delta · measured 2026-07-03
Read the report card before you read the pitch.
The v0.20.0 re-measure, dated 2026-07-03. It stands unchanged — v0.21.0 did not re-run it. Every row below traces to a committed artifact that ships in the repo. Open it and check. These are the metrics that carried an F at baseline, re-measured after the fix.
Scores — higher is better · 0 to 1.0
v1.0 single-call channel, same instrument as baseline
Wilson 95% CIs do not overlap
v1.0 channel
CIs do not overlap
v1.1 batched channel + equivalence block
10/10; Heroku trap now a correctly-flagged negative
Defects — removed
all 4 planted traps now cited only with a drift flag
backup + audit-trail manifest; the hard F-cap lifts
same 2 real dead links kept (exit 1); 13 FP → 0
On the measured slice: decision pack F B band brownfield-api F A band brownfield-docs F-cap lifted
derivation & scope — DELTA.mdRegressions ship too · v0.20.0
The row that costs us, in the same table as the wins.
the boundary rewrite that killed a poem false-positive also lost a tone-check positive — a real trade, filed
We publish the row that costs us, because a report card that only shows the good numbers is marketing.
Measured, not claimed
The harness is pre-registered.
The measurement spec is committed to git before any run — a release verifier checks in history that the spec’s first commit predates the first results commit, so metrics can never be redefined after seeing the numbers. Every figure carries its n, its method, and a 95% CI — or an explicit “single-run, directional” caveat. A number that cannot be recomputed from a committed raw artifact does not ship.
A Fable-5 blind re-judge agreed with the rubric judge on 18/20 scores exactly (0.90) and 20/20 within ±1 (1.00).
What it is
A generator with a proof harness.
- A generator, not a catalog
- It reads your codebase and synthesizes a workflow for your stack and conventions — not a pile of prebuilt components to pick from.
- 22 trigger-scored skills
- Across three packs — decision (12), naming (6), design-elevation (4). Every skill carries a published trigger score; activation-repeatability is measured for the 12 decision skills (naming and design deferred per a spend circuit-breaker), and static and rubric scores cover all 22. The numbers live in the harness results, not just prose.
- A proof harness
- It scores the generator on itself: artifact-level statistical scoring (trigger F1, activation repeatability with CIs, anchored rubrics) plus an end-to-end benchmark on synthetic fixture repos with answer keys.
- A vigilance loop
- It scans the AI ecosystem daily, synthesizes weekly, and integrates monthly — so the kit’s knowledge does not silently rot as models and frameworks move.
What it is not
And what it refuses to be.
- Not a CLI or a SaaS
- There is no binary and no server. The plugin installs skills; the generator is a prompt you paste; upgrades run through a manifest.
- Not a 200-skill catalog
- 22 skills, each with a published trigger score. Breadth is not the pitch; measurement is.
- Not infrastructure benchmarks
- No latency, RSS, or throughput numbers. The harness measures workflow-artifact quality and generator behavior on fixtures — a different axis entirely.
- Not autonomous
- The generator asks before it writes. Upgrade mode dry-runs by default. /directiveforge:report-friction never submits anything without your explicit review. Nothing edits your repo behind your back.
- Not self-congratulatory
- The baseline published F grades. The numbers are recorded before they are fixed, and the regressions ship in the same table as the wins.
60-second quickstart
Three steps to a measured workflow.
-
Install the skills and commands as a plugin
Two commands. This installs the 22 skills (decision / naming / design packs) and the workflow commands, including /directiveforge:report-friction.
claude plugin marketplace add directiveforge/directiveforgeclaude plugin install directiveforgePinned form of line 1
claude plugin marketplace add directiveforge/directiveforge@v0.21.0-publicPinning to a release tag fixes the catalog and the plugin together — they are one repository.
-
Generate a project-specific workflow
Open a session at the root of your target project and paste generator/PROJECT_SETUP_PROMPT.md. It reads your codebase, asks a short profile, and writes your CLAUDE.md, .cursor/rules/*.mdc, .claude/commands/, agents, and MCP config.
generator/PROJECT_SETUP_PROMPT.md -
Cursor consumes skills as files
Copy templates/skills/<pack>/ into your project per the Cursor workflow.
workflows/WORKFLOW-CURSOR.md
Full walkthrough: QUICK_START.md
Honest limitations
Where the numbers stop.
None are softened; each links to where it lives in the repo.
- Not everything was re-measured
- The delta re-ran only the metrics that carried an F. Static gates (L1.1) and rubric scores (L1.4) were not re-measured post-fix; the greenfield fixture was deferred to public round 2. Absent, not zero. DELTA.md
- Single-run directionality
- Layer 2 and some claim rows are single re-measure runs. Where a metric is not CI’d, it is labelled directional, not proven.
- The self-checklist is our own instrument
- The L2.1 checklist measures self-consistency, not external quality — circular by construction. The answer-key-scored metrics (L2.2 / L2.4 / L2.6) are the external check.
- Judge-model dependence
- Trigger routing is simulated with Haiku and rubric scores are judged by Opus — both are LLM judges. A Fable-5 blind re-judge calibrates the rubric judge, but a different judge could move borderline rows. calibration-fable.md
- Synthetic-fixture boundary
- Layer 2 runs against synthetic reference repos with answer keys, not real projects. We make no claim that the kit improves real-project outcomes. Nobody in the ecosystem measures that yet — including us.
- A regression shipped in the v0.20.0 delta
- disconfirming-evidence-first F1 0.9091 → 0.8889, disclosed above and filed for the next patch. Six residual defects are open and dispositioned. feedback/DISPOSITIONS-v0.19.0.md
Freshness
Knowledge rots silently. The cadence is public.
- Scan / synthesize / integrate
- A daily scan of a severity-tagged watchlist, a weekly synthesis that separates signal from noise, and a monthly integration through a reviewed session — never an auto-merge. workflows/KIT-VIGILANCE.md
- A weekly public digest
- What moved upstream, what it means for kit users, what shipped or is queued, and the open friction-report count with disposition status — on GitHub Releases. vigilance/PUBLIC-DIGEST-FORMAT.md
Feedback
Every report answered. Zero silent drops.
- /directiveforge:report-friction
- Harvests the defect (file:line, expected vs actual, severity), sanitizes paths and names, and stops at a review gate. Nothing is submitted until you approve it.
- The disposition guarantee
- Every report is answered — on the issue when it is filed as one, and in a public disposition record (a per-release DISPOSITIONS file or a dated field-trial log) with an explicit disposition: fixed, deferred-with-reason, or rejected-with-reason. Zero silent drops. feedback/DISPOSITIONS-v0.21.0.md
For the skeptic
The questions a Hacker News reader asks first.
Is this a SaaS or a CLI I have to run?
Neither. There is no binary and no server. The plugin installs skills into Claude Code; the generator is a prompt you paste at your project root; upgrades run through a manifest. Nothing phones home.
Does it actually improve my real project?
We make no such claim. The harness measures generator behavior on synthetic fixture repos with answer keys — not real-project outcomes. Nobody in the ecosystem measures that yet, including us. What we publish is what we can recompute from a committed artifact.
What is a pre-registered harness?
The measurement spec is committed to git before any run. A release verifier checks in git history that the spec’s first commit predates the first results commit, so metrics cannot be redefined after seeing the numbers. Tuning a rubric to pass is a spec change, a new dated version, and a re-run — not a quiet edit.
Why publish your own F grades and a regression?
Because a report card that only shows the good numbers is marketing. The baseline F grades and the shipped regression sit in the same table as the wins — that is the credibility mechanism, not a footnote.
The F grades ship in the same table as the wins.
claude plugin marketplace add directiveforge/directiveforge claude plugin install directiveforge