







10 on the record this issue
Mandate Post-Rotation Recovery Days
Val NarodetskyCEO · OdesaI pulled my team off seven-day on-call cycles and moved to a paired rotation where two engineers share a five-day window, then both get a full week with zero pages. One person handles daytime alerts, the other covers nights, and they swap midweek. Coverage stays continuous, but no single person is tethered to their phone for more than about 60 hours before a handoff.
The policy change that moved the needle most was adding a mandatory post-on-call recovery day. Not optional, not "take it if you need it". It's blocked on the calendar automatically when your rotation ends. My engineers were burning out not from the volume of incidents but from jumping straight back into sprint work after a rough night of pages, and that transition killed their deep focus for the rest of the week.
Once we started treating on-call recovery the same way we treat deployment freezes, something we schedule and protect, incident response times improved. Engineers stopped dreading the rotation because they knew a buffer existed on the other side. Volunteering for extra coverage went up, and we stopped losing mid-level developers who cited on-call fatigue during exit conversations.
Page Only Urgent Actionable Failures
Callum GracieFounder · Otto MediaAn on-call rota protects people only when a page means someone must act now. My policy would page engineers solely for urgent, actionable events that threaten users or an agreed reliability target. Warnings, capacity trends and non-urgent errors should enter a daytime queue with a named owner. Every page should include impact, escalation steps and a runbook, so the engineer does not begin with detective work. The team should review every overnight page and remove alerts that did not require immediate human action. This reduces needless interruptions while preserving coverage for the failures that genuinely cannot wait.
Shield Responders From Feature Work
Abhishek PareekFounder & Director · Coders.devTo safeguard engineering objectives while ensuring coverage of the systems, the segregation of incident response from the specific deadlines of delivering features is a must. In my role of supervising engineering groups distributed across numerous customer engagements, I noticed that burnout comes not so much from the pager, but from the unreasonable expectation imposed on the engineers to stick to the flow state together with being alert. There has been only one change made to policy, which reduces employees' turnover and improves reliability at a systemic level - it is putting into practice the zero-feature-work policy for the engineer on call. While on call, the engineer does not participate in any sprint activities and any new features. The only thing they need to focus on is the well-being of the system, the elimination of the technical debts, and development of the documentation explaining the monitoring process. The engineers do not need to be concerned about their performance in story points while on call.
The resulting shift allows applying the principles of Site Reliability Engineering meaning that the incidents are analyzed with consideration of the entire system rather than blaming a person involved in the incident. The engineer, whose alert went off, has an opportunity to find out the real reasons for the emergence of the alarm and to work out an automatic way of handling the alarm. The on-call period is treated as a carefully organized operational exercise instead of a careless rush to resolve issues as they arise. This solution helps protecting financial and human investments in engineering talent and ensures that the engineers do not get exhausted by maintaining the system.
Enforce Error Budget Deployment Gates
Hasan Can SoygökFounder · RemotifyWhen setting on-call rotations I protect focus and health by treating stability as the product and enforcing an error budget gate on deployments. The rule is simple: if a release is consuming our error budget or pushing change failure rate up, we stop shipping new features and focus on fixes and monitoring. That approach keeps on-call shifts predictable and reduces firefighting because engineers are not expected to absorb churn from risky deployments while managing incidents. Formalizing the error budget gate was the single policy change that most reduced burnout for our team while keeping reliability steady.
Replace Solo Escalation With Shared Response
The policy change with the biggest benefit was banning one-to-one incident escalation. A page opened a shared incident channel by default, with the primary responder, secondary responder, and service owner visible from the first minute. That sounds procedural, but it prevents the familiar pattern in which one tired engineer quietly carries diagnosis, customer pressure, and deployment decisions alone.
Coverage became stronger because the secondary could absorb work early, while the service owner supplied context and accepted tradeoffs. I require an incident record and a second set of eyes for overnight production changes. Shared visibility turns on-call from solitary endurance into controlled practice, reducing fatigue without asking people to be less responsive.
Send Noncritical Events to Morning Tickets
Victor SmushkevichFounder · Mold Scanner AIThe one policy I'd defend on a call: a page has to mean a user is hurt right now. Everything else becomes a ticket that waits until morning. Most of what wrecks sleep on a small team is alert noise, and a rule like that cuts it faster than any schedule tweak.
I build a consumer mobile app with a small team, so there's no follow-the-sun bench to lean on. When the same two or three people catch every alert, a 2am page for a slow background job is pure cost. Coverage doesn't drop when you stop paging for it. Nobody was going to fix it at 2am anyway, and the fix is the same at 9am.
The second piece is scope. Whoever is on call owns interrupts that week and doesn't also carry a feature, which leaves everyone else with long uninterrupted blocks. Rotate weekly and let people swap freely. After a rough night, the person gets a lighter next day without having to ask.
Then every page that did fire gets a short review. If it didn't need a human, the alert is deleted or downgraded to a ticket, and the page count should shrink month over month.
Audit Paging Noise Before Roster Changes
Kamyar ShahFractional COO · World Consulting GroupOn-call burns people out through unpredictability more than through volume. A rotation with few pages can still ruin a week if the engineer cannot plan around it. The damage I see comes from losing control of time, not from the count of alerts.
Fix the predictability before touching the schedule. Publish the rotation far enough ahead that people can plan, keep shift boundaries fixed, and never extend a shift to cover a gap. Pair the primary with a named secondary so one person is never the only path to resolution.
The policy change that reduces burnout most is a hard actionability rule for paging alerts. A standing review deletes or downgrades anything that woke someone without requiring action. Noise falls, sleep holds, and reliability improves because the remaining alerts get taken seriously. Alert quality protects both coverage and people, so audit the alerts before adjusting the roster.
Purge Pages That Need No Intervention
Siim KostabiCEO · PagelootOn-call nearly broke our small team in 2021. We had four engineers sharing a rotation and no rule about what qualified as a real alert. Everything paged. A slow dashboard query at 2am, a retry that self-healed in 90 seconds, a monitoring false positive we'd seen 40 times. Engineers were waking up three nights a week for things that didn't need a human.
The single policy change that fixed it: we audited every alert from the previous 90 days and deleted anything that had never required a human action. Cut our page volume by around 60%. Response quality went up immediately because the alerts that fired actually meant something.
The second piece, which costs nothing: the person on-call gets the following day protected. No standups, no sprint commitments, no Slack expectations. If they had a rough night, they recover. We're bootstrapped, so we couldn't hire our way out of the problem. We had to make the rotation survivable with the people we had.
Coverage didn't weaken. If anything it got more reliable, because engineers stopped dreading their week and started treating alerts as signal instead of noise.
Pilot Policies Before Permanent Adoption
Ronan LeonardFounder · Intelligent ResourcingThe single policy change was to require that any on-call process, schedule, or automation be piloted and validated by the team before it is formalized. Piloting lets us keep rotations lean, preserve predictable focus time, and avoid adding needless procedural overhead that drives fatigue. We only make a practice permanent after it proves it maintains coverage and reduces interrupting work. That approach protected engineers' health while keeping reliability steady.
Cut Sprint Commitments for Coverage Weeks
Mangesh GothankarChief Technology Officer · Your Team in IndiaI've found that on-call becomes a burnout problem when it is treated as an additional responsibility rather than part of the engineer's workload. Someone may receive only two or three alerts during a week, but those interruptions can break up deep work and affect the next day's productivity.
The policy change that made the biggest difference was reducing planned delivery work during an on-call week. For example, if an engineer would normally take on five meaningful tasks in a sprint, we would plan fewer when they were covering production. It gave them room to respond to incidents without having to make up the lost time later.
We kept coverage steady through a clear primary and secondary rotation. A critical production issue could immediately move to the secondary engineer, while lower-severity alerts could wait for the next working window. We also reviewed noisy alerts regularly. If the same alert kept waking people up without requiring action, we fixed the underlying monitoring rather than simply accepting the interruption.
I also think the handoff matters. Once an engineer's rotation ends, they should be able to disconnect unless they are actively involved in a major incident.
The principle is straightforward: on-call needs to be reflected in delivery planning. When teams have room for incident response, reliability does not have to come at the expense of people's ability to work sustainably.

