The forecast broke before the team did

Sprint planning rests on one assumption: that what the team got through recently predicts what it will get through next. Velocity is that assumption written down, and it holds as long as the spread around the average stays moderate.

Agent-assisted work widened the spread. The same nominal task lands in an hour when the agent gets it right, or takes three days when someone has to unpick what it produced, and the average of the two describes neither. That is one cause of six, and the other five never went away. They have different fixes.

Telling the team to plan harder is a diagnosis, not a fix. It decides the problem is the people, before anyone has asked which of the six causes is actually theirs.

Six causes of chronic slip, free

Each has a diagnostic question you can answer from the last two sprints without instrumenting anything. Most teams find two causes, not one.

  1. Unplanned work arrives every sprint and is never planned for Diagnostic: what share of last sprint was not in the plan on day one? Teams that answer "about a third, same as always" have a capacity model that ignores a third of reality.
  2. Work gets started faster than it gets finished Diagnostic: how many items were in progress at once, per engineer? Count the peak during the sprint, not the average. Everything moves and nothing lands, and this is the cause most often mistaken for the team being slow. It shows up in delivery data as lengthening lead time, one of the DORA metrics.
  3. The estimate assumes a middle case that no longer exists Diagnostic: for the last ten agent-assisted tasks, what was the fastest and what was the slowest? When the spread is tenfold, an average is not a forecast.
  4. Dependencies outside the team Diagnostic: how many committed items needed someone who does not attend your planning? Each one is a promise made on another team’s behalf without asking them.
  5. The commitment was negotiated, not forecast Diagnostic: did the number the team said out loud differ from the number they believed? A commitment produced by pressure is a prediction of the pressure, not of the work.
  6. Done keeps moving Diagnostic: how many items were finished by the team’s definition but not by the reviewer’s or the product owner’s? Slip is often the last mile being counted differently by different people.

The planning retro, fifteen minutes

Separate from the sprint retro and deliberately narrow. This one is only about the forecast: what you predicted, what happened, and which cause explains the difference.

  1. Committed versus landed, as a plain count No story points. Items promised, items delivered. Points make the conversation about the unit of measurement rather than about the plan. Whether the items that landed were worth landing at all is output vs outcome.
  2. For each item that missed, which of the six causes was it? One cause per item, chosen fast. The value is the tally across sprints, not accuracy on any single item.
  3. What arrived that nobody planned? List it, and note whether it was genuinely unforeseeable or the same recurring category as last time. Most of it is the same category as last time.
  4. Which estimate was furthest out, and in which direction? Both directions matter. Consistently finishing early is also a broken forecast, and it usually means the team is protecting itself from the last conversation about slipping.
  5. What will we change about the next plan? One change. Reducing the commitment counts as a change. So does reserving explicit capacity for the unplanned category that keeps recurring.
  6. What did we say last time, and did it help? The field that makes this a loop instead of a ritual. Without it you will diagnose the same cause for a year. Changes that get decided and then evaporate are retrospective action items.

Run it three or four times before drawing conclusions. One sprint is an anecdote.

Common questions about sprint planning

Why does our team never finish the sprint?

Usually not discipline. Six causes account for most chronic slip: unplanned work that is never planned for, too many items in progress at once, estimates built on a middle case that no longer exists, dependencies outside the team, commitments produced by pressure rather than forecast, and a definition of done that different people apply differently. They need different fixes, which is why planning harder does not work.

Should we still estimate now that agents write much of the code?

Estimate, but stop treating history as a forecast. Velocity is extrapolation from past throughput, and it works when variance is moderate. Agent-assisted work widened the spread: the same nominal task can land in an hour or take days when the output needs unpicking. Flow-based approaches that model a range, in the tradition of Kanban and the flow metrics literature, describe that world better than a single averaged number.

What is a healthy sprint commitment reliability?

There is no benchmark worth quoting, and anyone offering one is selling something. What matters is the trend and the spread. A team that lands a consistent share every sprint can plan around it, even if the share is not high. A team whose result swings widely cannot plan at all, and that instability is the thing to work on first.

From slipping sprints to changes: a worked example

Simulated team · Real product output Naming the cause only counts if the next plan changes. Here is what that looks like on one team: the Storefront Squad at Lingon & Co, a fictional consumer e-commerce company of about 200 people. A seven-person squad, top of the company's delivery dashboard, and a commitment that comes up short every sprint. We scripted the inputs and ran them through Aurora Coach in production. Everything below is the product's real output.

1Sense and analyze

Every team member answers structured questions in their own words, and can take any thread further with the coach in a check-in. The sessions roll up into one team analysis, with a score and recommendations per category.

Aurora Coach conversation with an engineer on the simulated Storefront Squad: 30 points planned and 19 landed for the third sprint running, and the coach naming unplanned work overwhelming planned capacity and proposing a 30 to 40 percent reserve for campaigns and partner requests

The plan that describes nothing

An engineer brings the numbers: 30 points planned, 19 landed, third sprint running. The coach names the cause as unplanned work overwhelming planned capacity, and puts a size on the fix: a 30 to 40 percent reserve for campaigns and partner requests. Her follow-up asks who has to agree to that, and the answer is specific. Cutting the commitment from 30 to 18-20 points escalates to whoever owns the roadmap promise, and it travels better as a delivery predictability argument than as a capacity one.

Aurora Coach conversation about planning poker taking 90 minutes every two weeks, with the coach proposing right-sizing work until each item fits the sprint instead of pointing it

Ninety minutes of planning poker

The same engineer on estimation: 90 minutes every two weeks producing a number that has never matched what happened and that nobody reads afterwards. The coach does not defend the ceremony. It offers right-sizing instead: ask whether each item fits the sprint, break it down until the answer is yes, and spend 20 to 30 minutes on scope and dependencies rather than 90 on points.

Aurora Coach growth opportunities detail for Adaptive Planning and Execution on the simulated Storefront Squad, scored two of five: velocity data exists but does not inform sprint commitments, flow metrics limited, retrospective insights not tracked to completion

What the analysis says about planning

Adaptive Planning and Execution is the category that scores how a team plans, and this team came back at two of five. The first line names the reason: a “significant gap between having planning practices and executing them with discipline”, velocity data that exists but never informs the commitment. That is one of the six causes above, in the product's own words.

2Recommend, refine, commit

The recommendations become concrete suggestions the team votes on and commits to. The AI informs the decision, it does not make it.

Aurora Coach improvement suggestion for the simulated Storefront Squad, fully expanded: start each sprint planning with a ten-minute velocity review using the last six sprints of data as the explicit basis for commitments, with expected outcome, context, implementation approaches, action steps, success metrics, timeframe, vote buttons, and Commit and Revise actions

From read to commitment

The planning read becomes one concrete move: a ten-minute velocity review at the start of each planning, with the last six sprints as the stated basis for the number the team commits to. It arrives with expected outcome, action steps, success metrics, and a timeframe. The Relevance field argues from this team's own numbers: definition of done already strong at four of five, so the missing piece is closing the loop between measurement and decision. The six causes above belong to every team that slips; the Relevance field is the part only this team's data could have written.

3Execute and re-evaluate

The team plans differently for a period, and the next analysis scores the same categories again: whether the unplanned share moved, whether committed and landed converged. The Storefront Squad has run one period, so the trend view starts when the second one lands.

This is one use case. How the full product works is on the product overview.

Not ready to change anything today? You already have the six causes and the planning retro above, copy buttons and all. If you want one improvement loop like these in your inbox each month, leave your email.

What this page cannot tell you is which of it applies to your team, this quarter. Aurora Coach works that out from your team's own words, recommends next steps with the reasoning, and the next period shows whether it held.

Both are free. The ROI mapper needs no signup and takes about two minutes. What team members write stays private to them: see AI governance.