career bridge

What good means

Define done before anything is built, with a numeric shipping bar written down while it is still uncomfortable.

You have a scope and a never-do list. The obvious next question, asked by everyone who funds this work, is how anybody will know whether it works. Most teams answer that after building, which means the answer turns out to be whatever the build happened to produce.

What you do

Define done before anything gets built. Take the workflow from the last stage and write the acceptance set: forty to sixty real cases someone would actually bring, drawn from real history if you can get it and invented carefully if you cannot. For each, write what a correct outcome is, and separately, what an acceptable but not ideal outcome is. That second column is what makes the set usable, because agents rarely fail cleanly. Group the cases into routine, ambiguous, out of scope, and adversarial. Then set the bar in public. What share of routine cases must be handled without a human before this ships? What is the maximum tolerable rate of wrong actions, not wrong answers, wrong actions? Write those numbers down before anyone has an incentive to move them. Take the set to an engineer and ask what would be expensive to measure, then decide out loud what you are willing to drop. Finish by writing how the set gets re-run and who owns it after launch, because an acceptance set nobody re-runs is a document rather than a control.

Done when

  • Forty or more real cases exist, grouped routine, ambiguous, out of scope, adversarial.
  • Each case has a correct outcome and an acceptable outcome, written separately.
  • A numeric shipping bar is written down and dated before any build starts.
  • An engineer has reviewed the set and told you which parts are expensive to measure.
  • A named owner and a re-run cadence exist for the set after launch.

What you end up with

An acceptance set of forty or more cases, with a dated numeric shipping bar written before the build and a named owner for re-running it.

If you get stuck

The most common failure is setting the bar after seeing the results, which is not a bar, it is a description. Write the number first and let it be uncomfortable. The second failure is forty happy-path cases and three hard ones, which produces a system that scores well and then fails in production. If your ambiguous and adversarial groups are smaller than your routine group, the set is flattering you.