Output vs outcome: shipping more and proving less
Your team ships visibly more than it did a year ago. The common problem: output is measured continuously and outcomes are barely measured at all, so nobody can say which of it mattered. This page covers how Aurora Coach closes that gap, and ends with the audit and the outcome-first definition of done, free.
Output became cheap, so it stopped being evidence
For most of software's history, building was the expensive part. That made throughput a reasonable stand-in for value: if a team shipped a lot, it had spent its scarce resource on things someone thought were worth it. The proxy worked because it was costly to be wrong.
Agents removed most of that cost. A team can now produce more in a quarter than it could in a year, which means shipping a lot no longer implies anyone chose carefully. The scarce resource moved from building to deciding, and most teams have not moved their attention with it.
The symptom is recognisable: a roadmap full of delivered items, a leadership team asking what changed, and an engineering org that can answer the first question in detail and the second not at all. It is not a reporting problem. Nobody wrote down what was supposed to change, so there is nothing to report.
The vocabulary for this is not new. Josh Seiden's Outcomes Over Output defines an outcome as a change in someone's behaviour that produces a business result, which gives you a strict and useful test. Shipping SSO is not an outcome. Support no longer fielding password resets is. Marty Cagan has made the same argument about product teams for years. What changed is not the theory. It is that the cost of being wrong used to enforce the discipline by itself, and now nothing does.
Throughput was a proxy for value because it was expensive. It is not expensive any more.
How Aurora Coach addresses it
The target is a team that can say what its last quarter was for. This work lives in Aurora Coach's Product domain, discovery and development: building the right thing rather than only building things right. It is one of six domains of team effectiveness alongside Foundation, Engineering, Operations, Workflow, and Alignment, and the loop that moves it has seven stages.
- Sense + Analyze The whole team contributes context: the AI asks structured questions and the team answers in their own words. Engineers usually know which of the last quarter’s work mattered and are rarely asked. Aurora Coach for GitHub adds delivery signal alongside. The AI synthesizes it into a strengths-and-gaps read grounded in the team’s actual situation.
- Recommend + Refine + Commit The AI recommends concrete next steps with rationale, implementation steps, and success criteria. Team members vote, the team lead refines to fit reality. The AI’s job is to inform the decision, not to make it. Closing the outcome gap becomes a commitment owned by the team and tracked through periods, rather than a resolution after a bad quarter.
- Execute + Re-evaluate The team changes how it works alongside delivery. Next period’s analysis sees what changed: whether outcomes are being stated before work starts, whether anyone is looking afterwards, and whether the answers are changing what gets built. Context compounds.
Free-text check-ins keep the sensing going between periods. If the deeper issue is that nobody on the team has spoken to a user, see dual-track discovery. If your delivery numbers look fine and the value question is still unanswered, improving DORA metrics covers the other half.
What is the difference between output and outcome in engineering?
Output is what the team produced: features shipped, pull requests merged, story points closed. An outcome is what changed for someone else as a result. Josh Seiden, in Outcomes Over Output, defines it as a change in human behaviour that drives a business result, which is a stricter test than most teams apply: a task got done faster, a support category disappeared, a customer renewed. Output sits fully inside the team’s control and is easy to count, which is exactly why it gets measured and outcomes do not.
How do you measure engineering outcomes without a full analytics stack?
Start by stating the intended outcome before the work begins and naming who will look afterwards and when. That single change surfaces most of the problem, because the items nobody can write an outcome for are usually the ones nobody checks. The signal itself can be an existing metric, a query someone runs by hand, or a handful of customer conversations.
Is velocity a bad metric?
It is a fine capacity signal and a poor value signal, and it got worse at the second job when agents entered the picture. Velocity tells you roughly how much a team can take on next period. It cannot tell you whether last period’s work was worth doing, and treating it as though it can is how throughput becomes activity rather than impact.
The outcome audit, free
Take the last ten things your team shipped and run these six questions over each one. It takes about an hour and it is uncomfortable the first time, which is the point.
- Was an intended outcome written down before the work started? Not a ticket description. A sentence saying what should be different for a user or the business once this exists.
- Who decided it was worth building, and on what evidence? A customer conversation, a support pattern, a metric, a competitor move, or someone senior asking. All are legitimate; only one of them is a hunch wearing a suit.
- Did anyone look after it shipped? Name the person and the date. If the honest answer is nobody, the item was output only, whatever it cost to build.
- What did the number do? Up, down, flat, or never instrumented. "Never instrumented" is the most common answer and the most useful one.
- Would you build it again knowing what you know now? The cheapest question on the list and the one teams skip. Ask it out loud, per item, with the people who built it.
- What happened to the items you would not rebuild? Still shipped, still maintained, still in the codebase. Unbuilt work has a running cost that nobody bills for.
If you want an industry number to argue with, the most quoted one is the Standish Group's finding that roughly two thirds of software features are rarely or never used. Handle it carefully. The CHAOS reports have been criticised in peer-reviewed work, notably by Eveleens and Verhoef in IEEE Software, over how success is defined and how projects are selected. It is a good reason to measure your own ten items. It is not a fact about your product.
An outcome-first definition of done
The audit tells you where you stand. This keeps the gap from reopening: five fields to add to whatever your team already uses, filled in before the work starts.
- Intended outcome One sentence, in terms of someone outside the team. "Support stops fielding password resets", not "ship SSO".
- How it will be visible The signal you will look at afterwards. An existing metric, a query someone runs, or five customer conversations. It does not need a dashboard.
- When you will look A date far enough out that the signal exists, close enough that people still remember the decision. Two to six weeks for most things.
- What would make this a mistake Stated before you build. A team that cannot say what failure looks like has not made a bet, it has made a plan.
- What happens if it misses Iterate, remove, or accept and move on. Deciding this in advance is what stops every miss becoming permanent surface area.
Do not apply it to everything at once. Pick the next three items that cost more than a week and start there; a definition of done the team stops filling in after three weeks is worse than the one you had.
Not ready to change anything today? You already have the audit and the definition of done above, copy buttons and all. If you want one improvement loop like these in your inbox each month, leave your email.
You have the audit and the definition of done above, yours to keep and to run without buying anything. It is also generic, as any page has to be. What it cannot tell you is which of it applies to your team, in your situation, this quarter. That judgement is the actual work.
That judgement is what Aurora Coach does. Your team supplies the context through Coaching Sessions and check-ins, in its own words, every period. The AI works from that context rather than from a template, and recommends specific next steps with the reasoning and how you will know whether it worked. The team decides and commits. The next period shows whether it held. This page is one problem in the Product domain. The loop runs across all six, with every team, every period.
What it costs to run: one Coaching Session per person per period, fifteen to twenty minutes, and a period defaults to four weeks with the team setting its own. What the team writes stays private to that person. Only the team-level picture rolls up to leadership.
Both are free. No signup, about two minutes.