A small team can operate a useful on-call system if its promises match its staffing and tooling.

A small team can operate a useful on-call system if its promises match its staffing and tooling. The goal is to connect a meaningful service problem with someone who can act, while making escalation and recovery predictable. An informal expectation that a founder is always available creates hidden risk and burnout rather than dependable coverage.
Define the supported service and hours
Identify critical journeys, customer commitments and the periods when response is required. Distinguish monitoring from a staffed response promise. Agree what happens outside coverage and how customers are informed. If continuous coverage is necessary, account for primary and backup responders, leave, illness and time to recover after overnight work.
Page on actionable consequences
Select alerts that indicate a customer-impacting failure or a credible imminent problem. Route lower-urgency signals to normal work. Every page should name the service, impact, dashboard and relevant runbook. Review repeated false alarms and duplicate notifications; an overloaded responder cannot reliably recognise the important signal.
Build a supported rotation
Give responders the access, practice and authority needed to act. Name a backup and an escalation route for unfamiliar or dangerous situations. Rehearse a few likely incidents in a safe environment and verify that contact paths work. Keep production changes coordinated with whoever is on call so new behaviour is not a surprise.
Use incidents to reduce future load
Track pages, after-hours interruptions, unresolved operational work and recurring causes. Allocate time to fix noisy alerts, automate safe repetitive steps and improve recovery. Review customer promises when workload exceeds staffing rather than silently increasing personal availability. A dedicated SRE hire may eventually help, but ownership, meaningful signals and tested procedures are necessary regardless of team title.
- Related service
- Incident severity levels: a practical escalation matrix
- Rollback or hotfix during a production incident
Frequently asked questions
Do we need an SRE before launching?
Not necessarily. You do need explicit operating ownership, suitable coverage and the ability to recover critical journeys.
Can one person provide continuous coverage?
That creates a fragile dependency. Plan backup, leave and sustainable response expectations.
Should every error send a page?
No. Page when timely action is required; route non-urgent issues to normal prioritisation.
What should the first rotation include?
Primary and backup owners, coverage hours, access, escalation contacts and tested runbooks for likely failures.
How do we know the system is improving?
Look for fewer repeated actionable incidents, lower unnecessary interruption and demonstrated recovery capability.
Bring the scope. We will help make it buildable.
Share the user journey, integrations and launch constraints. We can clarify the scope and prepare an estimate with assumptions and exclusions.
Further reading
Incident severity levels: a practical escalation matrix
Incident severity should describe current or credible business impact, not how alarming a log message looks.
Rollback or hotfix during a production incident
During an incident, choose the action most likely to restore acceptable service with controlled risk.
Application disaster recovery: test RTO and RPO
A disaster recovery plan is credible when the team can demonstrate restoration of a useful service.