delphi-ground / references/why-briefings-fail.md

Why briefings fail

A reference shipped with the skill · 669 words

Why briefings fail

Three failure modes. Only the first rests on evidence — a measured run, described below. The second and third are arguments from how the contract behaves, and they are labelled as such: worth acting on because the countermeasures cost nothing, not because anyone has watched them happen.

1. A shared briefing manufactures the agreement you then cite

This is the one that looks most like success.

What happened. Twenty-four agents reviewed one design spec. They were blind to each other and all received the same verified-facts briefing, which included four real incidents. Twelve of them independently produced the same central criticism, and that convergence was reported as the strongest available evidence — agreement across parties with opposed interests.

What was actually true. The spec those agents reviewed contained zero occurrences of the terms at the centre of that criticism. The frame came from the briefing — the one input all twenty-four shared. Blinding removes cross-talk; it does not remove a shared prior.

The convergence measured briefing salience, not discovery.

The rule. Split the briefing. Environment facts are shareable; findings and prior conclusions are withheld from anyone whose independence you intend to rely on. And when reading a fan-out, check whether the converged frame appears in the briefing before crediting it as independent.

What survives. The finding itself was correct and was verified directly against the source. What did not survive was the evidential weight placed on agreement. Those are different things, and conflating them is how a real finding gets defended with a bad argument.

2. An empty briefing produces disciplined fiction

Reasoned, not observed. No run has been conducted on an artefact with an insufficient briefing — the refusal rule exists precisely to prevent one.

The mechanism. A review contract typically enforces form: cite sources, label grounding, name a specific failure, state a cost. On an artefact with nothing verifiable behind it — a greenfield proposal, a first-of-its-kind integration, a system with no history — every one of those constraints still passes. On invention.

The output is indistinguishable from a grounded run: labels, citations, specific-sounding scenarios, disciplined register. The structure is what makes it persuasive.

Why this is worse than an ungrounded review. Unlabelled generic output gets discounted. Output carrying citations and stated limitations reads as evidence, and nobody argues with it.

The rule. Rate strength honestly, and refuse at insufficient. The refusal is the highest-value thing this skill produces and the easiest to skip, because refusing feels like failing to deliver.

3. A stale briefing produces confident anachronism

Reasoned, not observed. No briefing here has yet been reused across a gap long enough to test this.

The mechanism. Briefings are expensive to build and cheap to reuse. A reused briefing carries its facts forward without carrying forward whether they still hold.

Reviewers then reason correctly about a system that has moved, and their output looks exactly like reasoning about the current one — same anchors, same confidence, same format.

The rule. Date every briefing. On reuse, spot-check the anchors and re-date it, or rebuild. A briefing with no date cannot be assessed for staleness at all, which is worse than one that is visibly old.

What the evidence does not show

Worth stating, because overclaiming here would repeat failure mode 1.

The twenty-four-agent run had no condition without a briefing, and no condition without personas. So it cannot distinguish "the briefing did the work" from "the personas did the work" from "both". The briefing is the component with a demonstrated effect on output quality — every checkable structural finding traced back to something in it — but that is an observation about one run, not a controlled result.

Three arms would settle it: briefing with no personas, personas with no briefing, and a plain direct read of the same artefact. None has run.

Until they do, treat this skill as unfalsified rather than validated, and be suspicious of anyone — including this document — who describes it more strongly.