The Product Development Council

Six product-delivery review lenses that write their verdicts before seeing each other's, argue to a synthesis, and hand it back with the dissent still in it.

Merged · PR #50 commit f5ea2a0 github.com/microsoft/amplifier-bundle-skills
“The council could not, and should not, average that away.” From a live product-council verdict, unprompted

Ask an AI whether your plan is good,
and it will usually find a way to say yes.

🤖

One reviewer, one temperament

A single assistant reading your plan has no structural reason to object. It reads what you wrote, in the frame you wrote it in, and helps. Agreement is the default failure mode.

🕔

The objection still arrives

Just later. The unvalidated problem, the missing metric, the alternative the customer will actually pick — those show up three weeks in, when reversing costs a sprint.

💬

The meeting you can't schedule

Someone who only cares whether the problem is real. Someone who only cares whether anyone wants it. Someone who only cares what it costs to land. In one room, on demand.

Ask six reviewers who each own exactly one question — and who write their verdicts before seeing each other's — and you get the objection you would otherwise hit three weeks in. Use it before you build, when reversing costs a conversation instead of a sprint.

Two commands. The difference is what the panel can see.

Forked · Isolated
/product-council <target>

Runs in its own session. It cannot see your chat — so you hand it something concrete: a file, a design doc, a repo path.

Clean-room review. Nothing from the conversation leaks in to soften it. It reviews a snapshot — see Limits.

Inline · In-session
/product-council-here

Runs inside the session you're already in and reviews the work in front of you — the plan, roadmap or scope decision you've been building this whole time.

Use this when the thing being reviewed is the conversation itself, or when the work moved in the last hour.

Three councils, complementary — not competing. /council reviews code and systems. /design-council reviews visual and UX work. /product-council reviews product plans, roadmaps and scope decisions.

Six lenses. Each owns exactly one load-bearing question.

outcomist Mandatory front gate

“Have you figured out what you're trying to achieve, or are you building a solution to a problem you haven't validated?”

intent-keeper Reused by reference

Goal drift: has the build wandered from the brief?

user-advocate Reused by reference

Desirability: does the person who has to live with this actually want it?

outcome-cartographer

“What measurable outcome defines success, and how will we know we moved it?”

positioning-critic

“Why would the customer choose this over the alternative, including doing nothing?”

bet-sizer

“What's most likely to make this slip or fail to land, and is the investment sized to our confidence?”

intent-keeper and user-advocate are reused by reference from the existing engineering /council — not duplicated. One definition, two benches.

Cold first, then the argument.

01 · Fan out

Cold and independent

Each lens writes its verdict before seeing any of the others'. No anchoring, no first-mover framing.

02 · Debate

To consensus — or not

The lenses see each other and argue. Positions move only when a lens's own test is met.

03 · Synthesize

One verdict, attributed

Every claim is traced to a named lens with a verbatim quote. You can audit who said what.

04 · Record

Dissent + roster manifest

Unresolved disagreement is written down as a finding. The manifest says exactly who sat on the panel.

A FAIL is never quietly downgraded

It cannot be dropped, softened in passing, or averaged into a milder aggregate. If it moves, the movement is on the record.

A lens that can't run is marked UNAVAILABLE

Never silently replaced with a substitute. A five-lens run tells you it was a five-lens run.

Every claim carries its author

Verbatim quote, named lens. No anonymous “the review found” smoothing over which lens actually held the position.

Dissent is preserved, not resolved.

Most review tools optimize toward a single confident answer. This one treats leftover disagreement as the product. When six lenses still disagree at the end, that tension is the finding — and it is written into the verdict rather than smoothed out of it.

Held its ground
“if I let five other lenses' severity move my needle instead of my own test, I've substituted 'consensus' for 'intent.'”
One lens holding CONCERN against five FAILs. It crossed to FAIL only when its own pre-stated trigger fired. The synthesis labeled the behavior: “a lens moving only when its own explicit, self-set criteria were met, not on pressure.” That verdict's header stated outright: “This is not an average.”
Changed hands
“The panel's FAIL did not resolve — it changed hands.”
One lens escalated CONCERN→FAIL while another softened FAIL→CONCERN, and a third refused to let the softened FAIL be absorbed.
Refused to round up
“Converting to FAIL now would be me borrowing their verdict instead of answering my own question.”
Severity is not contagious. A lens answers its own question or it abstains.
Reversed itself
“I withdraw 'real betting-table shape' as a description of the whole plan. It is half a betting table.”
“yes — my BYOK answer was a v2 positioning decision being forced into a v1 validation test. I argued myself out of my own segment.”
The lenses lose arguments on the record too. Reversals are logged, not quietly overwritten.

A human never wrote a single lens.

The roster was derived by councilify — a meta-skill for building councils, already public in the same bundle. It derived the axes blind, shipped a subset, and documented the reasoning for everything it left out.

12candidate axes derived blind by councilify
6shipped — with a written reason for each of the other six being dropped
4already-built lenses that were free to include and were deliberately excluded
The interesting exclusion

Free lenses, turned down

Three of the four excluded lenses were known “always-find-a-gap” reviewers. Adding them would have compounded over-criticism until every verdict read the same. Excluding them was calibration-positive — a bench that always fails everything carries no information.

Why a separate bench at all
“councils should generally be unique and complementary vs duplicative and overlapping.” The author, on scoping the roster

Roster size was not a guess.

The author rejected an initial eight-lens version and demanded data. Ten variants — five sizes × two derivation flavors — were run against three scenarios in isolated Digital Twin Universe containers, and scored across five dimensions. Six won.

10roster variants (5 sizes × 2 derivation flavors)
3scenarios, run in isolated DTU containers
5scored dimensions: quality, goal achievement, cost, wall time, turn count
6lenses — the size that won on the data
“let's really get some data-driven insights to leverage 'for real' here instead of us shooting from the hip here.” The author, rejecting the 8-lens version
“Remember, we don't want consensus, we DO want varied/tension perspectives as part of this — does that change anything?” The author, mid-build — a correction that reshaped the evaluation
It ships with a documented weakness. No lens asks “is this bold enough, should we do the bigger version?” The bench is deliberately caution-skewed — and says so in its own derivation notes.

It told us our best tool was wearing the wrong claim.

On 23 July 2026 the council reviewed our own simulated-user-research tool. positioning-critic found the product fighting an unwinnable fight: the winnable category was “automated pre-flight product audit,” not “user research” — and it noted that the genuine moat “appears in zero sentences of positioning anywhere in the repo.”

“A genuinely excellent engine wearing the wrong claim, with no dial on it.” The synthesized verdict
The panel's only FAIL

outcome-cartographer

“Nowhere does any artifact say what number a successful product would move.”

Closing line: “Name the number first.”

What shipped as a result
  • Reframed from “user research” to an audit filter
  • Gained a three-tier evidence-honesty system
  • Shipped a triage step producing a measured precision number
Traceable end-to-end: council verdict → backlog doc → commit → shipped README. This is the only case traced fully end-to-end. Every other outcome in this deck is traced to the assistant's report of a commit — a weaker chain, and named as such.

It overruled the person who commissioned it. It was right.

Overruled the commissioner

Self-hosting architecture: 6–0 against

The author instructed a self-hosting architecture. All six lenses came back against it. bet-sizer priced it: roughly 10% confidence a self-host v1 produces a usable signal, versus ~85% team-hosted.

The author accepted. The brief was amended.

Caught the AI misinforming the human

A factual error, found by the panel

“I told you 'all four automations are live.' True of the wiring. Not true of execution… The council found this before I did.”
Named its own operator's process failure
outcomist: “Question 4 — 'what is the measurable outcome' — is being asked AFTER two full investigations were already commissioned and produced complete competing designs. That's activity before outcome.”
The assistant conceded: “I ran two investigations before defining what 'fixed' means. That's on me.”
Killed before built

A 28-subcommand rewrite killed — and a live config-corruption path exposed on the way.

Killed before built

An expensive bezel-rail design killed after the panel proved both premises it rested on were false. A zero-cost alternative shipped instead.

Built, committed, live

In another project, all seven ranked fixes were built and committed in one session — including a metric-integrity fix from outcome-cartographer: “A metric that improves when the system breaks is worse than no metric.”

Built in five days. Used privately for eighteen. Rising, not decaying.

27real invocations in 18 days
10distinct projects reviewed
1person driving all of it
12of the 27 landed in the last five days

The timeline

  • 10–14 July 2026 — built
  • 18 days — used privately, in real work, before publishing
  • 31 July 2026 — published as PR #50, commit f5ea2a0
The adoption signal is the back half. Nearly half of all invocations came in the final five days of the private period. A tool that gets used more as it ages is one that survived contact with real decisions — but note the base: one person, ten of their own projects.

A council announcement that oversells itself undercuts the exact property being announced.

The honest ledger — 27 invocationsCountWhat that means
Substantive~20produced a substantive finding
Partial or ignored5verdict produced, action incomplete or absent
No verdict at all2provider overload / interrupt

One verdict was produced and flatly dropped — no action, no human response. There is no clean hit rate here and we are not claiming one.

Its advice is sometimes ignored — repeatedly

One recommendation — “get one real second human” — was raised four separate times and remains unactioned. The assistant itself logged “Third council ask, still unactioned.” A council whose advice is sometimes ignored is the true story.

The fork reviews a snapshot

Because /product-council forks, it can miss work shipped minutes earlier. In one real run it flagged something already shipped 40 minutes prior. That limitation is exactly why /product-council-here exists.

No lens asks “should we do the bigger version?”

The bench is deliberately caution-skewed and documents this in its own derivation notes. It will help you not overbuild. It will not push you to be bolder.

One author, self-applied

All 27 invocations are one person, across ten projects. One outcome is traced end-to-end to a shipped artifact; the others rest on the assistant's report of a commit. Read this as an early internal signal, not independent validation.

Point it at the plan before the plan becomes a sprint.

Review something concrete — forks, runs isolated, cannot see your chat:

/product-council docs/roadmap-q3.md /product-council ./my-repo # hand it a file, a doc, or a repo path

Review what's in front of you — runs inline, in this session:

/product-council-here # reviews the plan / scope decision # you've been building in this chat
Good moment 1

Before you build

When the cost of reversing is a conversation, not a sprint. This is the moment the bench is calibrated for.

Good moment 2

When scope is drifting

intent-keeper exists for exactly the build that has quietly wandered from its brief.

Good moment 3

Before a roadmap lands

Six named objections, on the record, are a defensible input to the prioritization conversation.

Where every number and quote came from.

Data

  • Data as of: 31 July 2026
  • Status: Active. Merged as PR #50, commit f5ea2a0, to microsoft/amplifier-bundle-skills
  • Timeline: built 10–14 July 2026; used privately 18 days; published 31 July 2026
  • Usage figures: 27 invocations across 10 distinct projects, 12 of them in the final five days — counted from session records
  • Quotes: verbatim from live council verdicts, synthesis headers, and the author's own messages during the build. No quote in this deck is paraphrased or reconstructed.
  • Roster derivation: councilify derivation notes — 12 candidate axes, 6 shipped, 6 documented drops, 4 available lenses deliberately excluded
  • Sizing experiment: 10 variants (5 sizes × 2 derivation flavors) × 3 scenarios, run in isolated Digital Twin Universe containers, scored on quality, goal achievement, cost, wall time and turn count

Gaps and disclosures

  • No clean hit rate. ~20 substantive, 5 partial or ignored, 2 returned no verdict at all. One verdict was dropped with no action and no human response.
  • One end-to-end trace. Only the simulated-user-research reframe is traced council → backlog doc → commit → shipped README. All other outcomes rest on the assistant's report of a commit.
  • n = 1 operator. Every invocation is one person, on their own projects. No external users, no independent replication.
  • Self-applied. The most-cited outcome is the council reviewing a tool built by the same author who commissioned the council.
  • Known blind spot, by design: no lens argues for ambition. Documented in the bench's derivation notes, repeated on the Limits slide.
  • Snapshot review: the forked command reviews state at fork time; one recorded run flagged work shipped 40 minutes earlier.
product-council · amplifier-bundle-skills
More Amplifier Stories