Output vs outcome: shipping more and proving less
Your team ships visibly more than it did a year ago, and nobody can say which of it mattered. This page has the six-question outcome audit and an outcome-first definition of done, free to copy. It ends with how Aurora Coach turns the gap into changes a team commits to, shown on a worked example.
Output became cheap, so it stopped being evidence
For most of software's history building was the expensive part, so shipping a lot was itself evidence that someone had chosen carefully. Agents removed most of that cost. The scarce resource moved from building to deciding, and most teams have not moved their attention with it.
The test is not new. Josh Seiden's Outcomes Over Output defines an outcome as a change in what someone does: a self-serve export is not an outcome, support no longer running exports by hand is. What changed is that the cost of building used to enforce that test for you.
Throughput was a proxy for value because it was expensive. It is not expensive any more.
The outcome audit, free
Take the last ten things your team shipped and run these six questions over each one. It takes about an hour and it is uncomfortable the first time, which is the point.
- Was an intended outcome written down before the work started? Not a ticket description. A sentence saying what should be different for a user or the business once this exists.
- Who decided it was worth building, and on what evidence? A customer conversation, a support pattern, a metric, a competitor move, or someone senior asking. All are legitimate; only one of them is a hunch wearing a suit. Finding out before you build is dual-track discovery.
- Did anyone look after it shipped? Name the person and the date. If the honest answer is nobody, the item was output only, whatever it cost to build.
- What did the number do? Up, down, flat, or never instrumented. "Never instrumented" is the most common answer and the most useful one.
- Would you build it again knowing what you know now? The cheapest question on the list and the one teams skip. Ask it out loud, per item, with the people who built it.
- What happened to the items you would not rebuild? They are still shipped, still in the codebase, and still being maintained, tested and supported every sprint. Decide per item: keep it, or remove it and get the capacity back.
An outcome-first definition of done
The audit tells you where you stand. This keeps the gap from reopening: five fields to add to whatever your team already uses, filled in before the work starts.
- Intended outcome One sentence, in terms of someone outside the team. "Support stops fielding password resets", not "ship SSO".
- How it will be visible The signal you will look at afterwards. An existing metric, a query someone runs, or five customer conversations. It does not need a dashboard.
- When you will look A date far enough out that the signal exists, close enough that people still remember the decision. Two to six weeks for most things.
- What would make this a mistake Stated before you build. A team that cannot say what failure looks like has not made a bet, it has made a plan.
- What happens if it misses Iterate, remove, or accept and move on. Deciding this in advance is what stops every miss becoming permanent surface area.
Do not apply it to everything at once. Pick the next three items that cost more than a week and start there; a definition of done the team stops filling in after three weeks is worse than the one you had.
What is the difference between output and outcome in engineering?
Output is what the team produced: features shipped, pull requests merged, story points closed. An outcome is what changed for someone else as a result. Josh Seiden, in Outcomes Over Output, defines it as a change in human behaviour that drives a business result, which is a stricter test than most teams apply: a task got done faster, a support category disappeared, a customer renewed. Output sits fully inside the team’s control and is easy to count, which is exactly why it gets measured and outcomes do not.
How do you measure engineering outcomes without a full analytics stack?
Start by stating the intended outcome before the work begins and naming who will look afterwards and when. That single change surfaces most of the problem, because the items nobody can write an outcome for are usually the ones nobody checks. The signal itself can be an existing metric, a query someone runs by hand, or a handful of customer conversations.
Is velocity a bad metric?
It is a fine capacity signal and a poor value signal, and agents widened the gap between the two. Velocity tells you roughly how much a team can take on next period. It cannot tell you whether last period’s work was worth doing, and treating it as though it can is how throughput gets counted as impact.
From feature count to outcomes: a worked example
Simulated team · Real product output Lingon & Co, a fictional consumer e-commerce company. A seven-person storefront squad, top of the company's feature-count dashboard, and no idea which of last quarter's 23 shipped features changed anything. We scripted the inputs and ran them through Aurora Coach in production. Everything below is the product's real output.
1Sense and analyze
Every team member answers structured questions in their own words, and can take any thread further in a check-in conversation. The answers roll up into a team analysis that scores six domains and says what is holding each one back.

High output, low discovery
Two dots filled of five for Product Discovery and Development, while this squad's delivery domains stayed healthy and it kept releasing several times a week. The first line names the cause: customer access rated one of five for a consumer storefront squad. The second is the consequence, no post-release outcome measurement, which the analysis calls a risk of “efficiently building the wrong things.”

The smallest version of looking
A senior engineer took his own question to the coach: nobody could say whether the wishlist redesign shipped in April had been used. He pushes for something a seven-person squad could realistically do. The answer is one explicit step in the sprint review they already hold: five minutes on what happened with the last feature before demoing the new one, basic usage numbers pulled beforehand, and “did the last thing work?” becomes a required question before moving to “what's next?”
2Recommend, refine, commit
The low scores come with recommendations, and the recommendations become concrete suggestions the team votes on and commits to. The AI informs the decision, it does not make it.

From gap to commitment
One of twelve suggestions generated for this team: define and track one or two outcome metrics across the next three releases. Read the Context field. It argues from this squad's own situation, that it “ships with elite speed but has no systematic way to know if releases solve real customer problems”. That is what the product adds over this page. The definition of done above is general knowledge; this is what it becomes for one specific team, down to the sentence it asks them to write before each release: “We believe this feature will improve [metric] by [amount] within [timeframe]”. Commit turns it into a tracked improvement.
3Execute and re-evaluate
The squad does the work inside the roadmap it already has. The next analysis shows whether outcomes now get written down before work starts and whether anyone looked afterwards. Lingon & Co has run one period. The trend view starts when the second one lands.
This is one use case. How the full product works is on the product overview.
Not ready to change anything today? You already have the audit and the definition of done above, copy buttons and all. If you want one improvement loop like these in your inbox each month, leave your email.
What this page cannot tell you is which of it applies to your team, this quarter. Aurora Coach works that out from your team's own words, recommends next steps with the reasoning, and the next period shows whether it held.
Both are free. The ROI mapper needs no signup and takes about two minutes. What team members write stays private to them: see AI governance.