On-call: the load nobody reports until they leave
On-call is scheduled, so it looks like something the organization measures. The common problem: what gets measured is the schedule, and the load sits almost entirely outside it. This page covers how Aurora Coach closes that gap, and ends with the health pulse and the noise audit, free.
The rota looks fine on paper
A rotation is visible, fairly distributed, and reviewed when someone complains. It records who was available. It records nothing about what actually happened to them, and that is where the entire cost lives.
Three things go unrecorded. The pages that turned out to be noise, which cost the same sleep as the real ones. What a broken night takes out of the following day, which no incident record has ever captured. And the recurrence: work done at three in the morning restores service and rarely gets revisited in daylight, so the same page arrives on somebody else's week and counts as a new event.
Then there is the concentration. Most rotas have one person who gets called whether or not it is their week, because they are the one who knows. Their load does not appear in the schedule at all, and they are usually the engineer you can least afford to lose. That cost surfaces in a resignation conversation, months later, filed under something else.
The Google SRE book gave this vocabulary years ago with toil and error budgets, and plenty of teams can quote it. Quoting it is not the hard part. The hard part is that reducing on-call load competes with feature work every single period, and loses by default unless something keeps putting it back on the table.
The rota measures who was available. It does not measure who was woken.
How Aurora Coach addresses it
The target is a rotation that does not depend on particular people tolerating it. This lives in Aurora Coach's Operations domain, delivery and operational excellence: how quickly value reaches users, and how reliably it keeps working once it has. That is a wide area of practice. It takes in incident response and the on-call experience itself, the automation that removes repeat work, monitoring that catches problems before customers do, and the reliability decisions that determine whether a Tuesday night is quiet. DORA gives you four well-known measures of part of it, and they are worth having, but the domain is the practice underneath rather than the dashboard on top. It is one of six domains of team effectiveness alongside Foundation, Product, Engineering, Workflow, and Alignment.
- Sense + Analyze The whole team contributes context: the AI asks structured questions and each person answers in their own words. On-call load lives almost entirely in what people experienced and almost not at all in the incident record, so the team’s own words are the only place some of it exists. Aurora Coach for GitHub adds delivery signal alongside. The AI synthesizes it into a strengths-and-gaps read grounded in the team’s actual situation.
- Recommend + Refine + Commit The AI recommends concrete next steps with rationale, implementation steps, and success criteria. Team members vote and the team lead refines to fit reality. The AI’s job is to inform the decision, not to make it. Reducing on-call load becomes a commitment tracked through periods, which matters here more than most places, because this work competes with features and loses by default.
- Execute + Re-evaluate The team does the work alongside delivery. The next period’s analysis sees whether the noise share moved, whether the same person is still being escalated to off-rota, and whether the recurring page came back. Context compounds, so a quiet quarter is visible as improvement rather than as luck.
Free-text check-ins catch a bad week while people still remember it. If the fixes get agreed and never shipped, that is postmortem follow-through. If the concern is the person absorbing the load rather than the rota itself, see senior engineer attrition. For time to restore and the rest of the delivery picture, improving DORA metrics.
What is alert fatigue?
What happens when a large share of alerts turn out not to need action, so people stop treating each one as though it might. The term comes from clinical alarm research, where the same pattern was studied in hospitals long before it reached software. It is worth being precise about the cause: this is not carelessness. It is a reasonable adaptation to a signal-to-noise ratio that somebody else configured, and it is fixed by changing the alerts rather than by asking people to try harder.
How do you measure on-call health?
Ask after each rotation rather than in a quarterly survey, because the detail is gone within a fortnight. The useful signals are how often someone was woken versus how often they were needed, what the following day cost, who got escalated to whether or not they were on the rota, and whether the work done during the shift prevents the next page or merely ends this one. None of that appears in an incident record.
What is a healthy on-call rotation?
There is no universal number, and rotations differ too much for a benchmark to mean much. A workable test is whether you would hand the rota to a competent new joiner without warning them about its quirks. If it only functions because experienced people absorb its rough edges, the load is real, it is concentrated on your most capable engineers, and it is invisible to the schedule.
The on-call health pulse, free
Six questions, asked when a rotation ends rather than in a quarterly survey. Two weeks later the detail is gone and you get impressions instead of facts.
- How many times were you woken, and how many of those needed you specifically? Two numbers, not one. The gap between them is the part of the load that better routing or a runbook could have removed entirely.
- What did you do the next day that you would not have done rested? The cost of a bad night lands the following day and never appears in an incident record. People remember this vividly and are almost never asked.
- Which alert do you now trust least? Everyone on a rota has one. Naming it is the fastest route to the noisiest alert in the system, and it is faster than any dashboard.
- Who did you escalate to, and were they on the rota? This is how you find the shadow rota. There is usually one person who gets called whether or not it is their week, and their load is invisible to the schedule.
- What did you fix during the shift that will page someone again next month? Separates the incident from the recurrence. Work done at 3am to restore service rarely gets revisited in daylight, so the same page returns on someone else’s week.
- Would you hand this rotation to a new joiner as it stands? The summary question. A rota that only works because experienced people absorb its rough edges is a rota with a retention problem attached.
The alert noise audit
Five steps, about an hour, no tooling. It turns a feeling everyone on the rota already has into a number the team can actually move.
- Take the last thirty pages, whatever they were Not a sample chosen for a meeting. The last thirty in order, including the ones everyone already knows are noise.
- Sort each into three piles Needed a human within minutes. Needed a human eventually. Needed nobody. Do it fast and do not argue about edge cases; the proportions are the finding.
- Count the third pile as a share of the total This number is the alert fatigue problem, stated plainly. When it is high, people are not being careless by skimming alerts. They are adapting correctly to a signal someone else configured.
- Take the single most frequent alert and ask what it would take to remove it Remove, not tune. Tuning a threshold moves the noise. Ask what would have to be true for the alert to be unnecessary, and price that.
- Give it an owner and a date, then re-run the audit next quarter Without the re-run this is an interesting afternoon. With it, the share of the third pile becomes a number the team can move, which is the whole point.
Expect the first run to be uncomfortable and the number to be worse than anyone guessed. That is normal and it is not an indictment of whoever configured the alerts, who was adding them one at a time for good reasons over several years.
Not ready to change anything today? You already have the health pulse and the noise audit above, copy buttons and all. If you want one improvement loop like these in your inbox each month, leave your email.
You have the health pulse and the noise audit above, yours to keep and to run without buying anything. It is also generic, as any page has to be. What it cannot tell you is which of it applies to your team, in your situation, this quarter. That judgement is the actual work.
That judgement is what Aurora Coach does. Your team supplies the context through Coaching Sessions and check-ins, in its own words, every period. The AI works from that context rather than from a template, and recommends specific next steps with the reasoning and how you will know whether it worked. The team decides and commits. The next period shows whether it held. This page is one problem in the Operations domain. The loop runs across all six, with every team, every period.
What it costs to run: one Coaching Session per person per period, fifteen to twenty minutes, and a period defaults to four weeks with the team setting its own. What the team writes stays private to that person. Only the team-level picture rolls up to leadership.
Both are free. No signup, about two minutes.