Sprints that keep slipping: planning when estimates stopped working
Your team commits to a sprint and lands about two thirds of it, sprint after sprint. The common problem: this gets treated as a discipline failure, so everyone plans harder, when the inputs the plan is built from are what stopped being reliable. This page covers how Aurora Coach closes that gap, and ends with the six causes and the planning retro, free.
The forecast broke before the team did
Sprint planning rests on one assumption: that what the team got through recently predicts what it will get through next. Velocity is that assumption written down. It is a reasonable way to forecast, and it works as long as the spread around the average stays moderate.
Agent-assisted work widened the spread. The same nominal task now lands in an hour when the agent gets it right, or takes three days when someone has to understand and unpick what it produced. The average of those two numbers describes neither. History still exists, but it stopped describing the distribution, and a forecast built from it inherits the problem without anyone noticing the moment it happened.
Meanwhile the older causes never went away. Unplanned work still arrives every sprint and still gets planned for at zero. Too many things run at once, so everything moves and little finishes. Dependencies sit outside the team. Sometimes the commitment was never a forecast at all, just the number that ended an uncomfortable conversation.
These have different fixes, which is why the usual response fails. Planning harder treats six problems as one problem, and the one it assumes is the team.
Velocity is a forecast built from history. Agents widened the spread that history described.
How Aurora Coach addresses it
The target is a plan the team believes before the sprint starts. This lives in Aurora Coach's Workflow domain, adaptive planning and execution, which covers flow over velocity: planning effectiveness, work in progress, cycle time, and the feedback loops that keep planning honest. It is one of six domains of team effectiveness alongside Foundation, Product, Engineering, Operations, and Alignment.
- Sense + Analyze The whole team contributes context: the AI asks structured questions and each person answers in their own words. Planning problems are usually known in detail by the people doing the work and summarised out of existence by the time they reach a status report. Aurora Coach for GitHub adds delivery signal alongside. The AI synthesizes it into a strengths-and-gaps read grounded in the team’s actual situation.
- Recommend + Refine + Commit The AI recommends concrete next steps with rationale, implementation steps, and success criteria. Team members vote and the team lead refines to fit reality. The AI’s job is to inform the decision, not to make it. Changing how the team plans becomes a commitment it owns, rather than a resolution that lasts until the next busy sprint.
- Execute + Re-evaluate The team works differently alongside delivery. The next period’s analysis sees whether the change held: whether the unplanned share moved, whether items in flight came down, whether the gap between committed and landed narrowed. Context compounds, so the diagnosis gets sharper each period rather than starting over.
Free-text check-ins keep the sensing going between sprints, so a bad sprint is visible as a pattern rather than as an argument in the retro. If the plan is fine and the items never get done afterwards, see retrospective action items. If the question is whether the delivered work mattered at all, that is output vs outcome. For the delivery numbers themselves, improving DORA metrics.
Why does our team never finish the sprint?
Usually not discipline. Six causes account for most chronic slip: unplanned work that is never planned for, too many items in progress at once, estimates built on a middle case that no longer exists, dependencies outside the team, commitments produced by pressure rather than forecast, and a definition of done that different people apply differently. They need different fixes, which is why planning harder does not work.
Should we still estimate now that agents write much of the code?
Estimate, but stop treating history as a forecast. Velocity is extrapolation from past throughput, and it works when variance is moderate. Agent-assisted work widened the spread: the same nominal task can land in an hour or take days when the output needs unpicking. Flow-based approaches that model a range, in the tradition of Kanban and the flow metrics literature, describe that world better than a single averaged number.
What is a healthy sprint commitment reliability?
There is no benchmark worth quoting, and anyone offering one is selling something. What matters is the trend and the spread. A team that lands a consistent share every sprint can plan around it, even if the share is not high. A team whose result swings widely cannot plan at all, and that instability is the thing to work on first.
Six causes of chronic slip, free
Each has a diagnostic question you can answer from the last two sprints without instrumenting anything. Most teams find two causes, not one.
- Unplanned work arrives every sprint and is never planned for Diagnostic: what share of last sprint was not in the plan on day one? Teams that answer "about a third, same as always" have a capacity model that ignores a third of reality.
- Work gets started faster than it gets finished Diagnostic: how many items were in progress at once, per engineer? Everything moves and nothing lands. This is the cause most often mistaken for the team being slow.
- The estimate assumes a middle case that no longer exists Diagnostic: for the last ten agent-assisted tasks, what was the fastest and what was the slowest? When the spread is tenfold, an average is not a forecast.
- Dependencies outside the team Diagnostic: how many committed items needed someone who does not attend your planning? Each one is a promise made on another team’s behalf without asking them.
- The commitment was negotiated, not forecast Diagnostic: did the number the team said out loud differ from the number they believed? A commitment produced by pressure is a prediction of the pressure, not of the work.
- Done keeps moving Diagnostic: how many items were finished by the team’s definition but not by the reviewer’s or the product owner’s? Slip is often the last mile being counted differently by different people.
The planning retro, fifteen minutes
Separate from the sprint retro and deliberately narrow. This one is only about the forecast: what you predicted, what happened, and which cause explains the difference.
- Committed versus landed, as a plain count No story points. Items promised, items delivered. Points make the conversation about the unit of measurement rather than about the plan.
- For each item that missed, which of the six causes was it? One cause per item, chosen fast. The value is the tally across sprints, not accuracy on any single item.
- What arrived that nobody planned? List it, and note whether it was genuinely unforeseeable or the same recurring category as last time. Most of it is the same category as last time.
- Which estimate was furthest out, and in which direction? Both directions matter. Consistently finishing early is also a broken forecast, and it usually means the team is protecting itself from the last conversation about slipping.
- What will we change about the next plan? One change. Reducing the commitment counts as a change. So does reserving explicit capacity for the unplanned category that keeps recurring.
- What did we say last time, and did it help? The field that makes this a loop instead of a ritual. Without it you will diagnose the same cause for a year.
Run it three or four times before drawing conclusions. One sprint is an anecdote, and the tally across a handful of sprints is what names the cause. If the answer turns out to be the same category of unplanned work every time, that is not a planning problem to solve. It is capacity to reserve.
Not ready to change anything today? You already have the six causes and the planning retro above, copy buttons and all. If you want one improvement loop like these in your inbox each month, leave your email.
You have the six causes and the planning retro above, yours to keep and to run without buying anything. It is also generic, as any page has to be. What it cannot tell you is which of it applies to your team, in your situation, this quarter. That judgement is the actual work.
That judgement is what Aurora Coach does. Your team supplies the context through Coaching Sessions and check-ins, in its own words, every period. The AI works from that context rather than from a template, and recommends specific next steps with the reasoning and how you will know whether it worked. The team decides and commits. The next period shows whether it held. This page is one problem in the Workflow domain. The loop runs across all six, with every team, every period.
What it costs to run: one Coaching Session per person per period, fifteen to twenty minutes, and a period defaults to four weeks with the team setting its own. What the team writes stays private to that person. Only the team-level picture rolls up to leadership.
Both are free. No signup, about two minutes.