OutsourcingVN is operated by Netbase JSC, which runs managed operations and would like to be the team your alerts reach, so treat this as a supplier's checklist. It applies to an in-house team or any provider. It draws on Netbase's scoped recovery of a compromised Magento store and on the support coverage Netbase actually offers.
What does incident readiness mean in practice?
Most outages are not caused by exotic failures. They are ordinary failures that nobody noticed, that reached the wrong inbox, or that the person on duty had no permission or instructions to fix. A certificate expires, a disk fills, a payment provider changes a callback, a scheduled job stops after a server move. Readiness is the unglamorous work that turns each of those from a crisis into a routine task.
Reliability and readiness are related but different. Reliability is how often the system fails and how badly. Readiness is how quickly and calmly the organisation recovers when it does. A new or inherited system usually needs readiness first, because the causes of unreliability are not yet known.
How ready is your system today?
Score each capability honestly. A capability counts as ready only when it has been tested, not when it exists on paper.
| Capability | Not ready | Ready | How to test it |
|---|---|---|---|
| Journey monitoring | Server metrics only | The journeys users notice are checked end to end | Break a test journey and see whether anyone hears |
| Alert routing | Alerts go to a shared inbox | Each alert reaches a named person with a backup | Send a test alert out of hours |
| Runbooks | Knowledge in one engineer's head | A first-response page per critical alert | Ask someone new to follow one |
| Backup and restore | Backups run; nobody restores | A restore has been performed and timed in a test environment | Restore last night's backup |
| Rollback | Rollback means a new release | The previous version can be redeployed quickly | Roll back in a test environment |
| Emergency access | Access requested during the incident | On-duty people hold the access they need, with MFA | Check access for each on-duty person |
| Incident roles | Everyone joins the call | One person leads; others have named jobs | Run a short tabletop exercise |
| Third-party dependencies | Discovered during outages | Listed with status pages and support contacts | Review the list against the running system |
| Post-incident review | Blame or silence | A written review with actions and owners | Check the last incident's actions were done |
Who does what during an incident?
Google's Site Reliability Engineering book describes separate incident roles: an incident commander who holds the overall state and coordinates, an operations lead who changes the system, a communications lead who keeps stakeholders informed, and a live incident document that records status and actions as they happen. A small team cannot staff four people, but it can still name who leads. One person may both fix and communicate; what fails is three people fixing at once while nobody tells the business what is happening.
Write the business side into the roles too. Someone on the buyer's side must be able to decide, at short notice, whether to take a feature offline, send a customer message or accept a degraded mode.
How do you build readiness in six weeks?
-
Week 1: pick the critical journeys
Two or three paths whose failure the business feels first, such as checkout, order sync or payroll export.
-
Week 1: route alerts
Each journey check alerts a named person and a backup, and the route is tested.
-
Week 2: write first-response runbooks
One page per critical alert: what it means, what to check, how to make the system safe, whom to call.
-
Week 3: rehearse recovery
Restore a backup and roll back a release in a test environment, and record how long each took.
-
Week 4: fix emergency access
On-duty people hold the access they need, with role-based control and MFA, before an incident rather than during one.
-
Week 5: run a tabletop exercise
Walk through a realistic failure with the named roles and note every point of confusion.
-
Week 6: agree the review habit
Every significant incident gets a written review within a week, and its actions go into the normal work queue.
After the six weeks, the weekly review is where incident actions are followed up. Netbase delivery runs with weekly reviews and a named project manager, which gives those actions an owner rather than leaving them in the incident document.
What about incidents outside support hours?
The Hanoi office works Monday to Saturday, 9:00-18:15 Vietnam time (UTC+7), and support coverage follows those days. Anything wider is agreed per engagement. Whatever the coverage, design for the hours nobody is watching: queue work rather than dropping it, fail closed on payments, show users an honest message rather than a broken page, and send an alert that a named person will see first thing. A system that fails safely overnight can be fixed calmly in the morning. The managed software operations scope guide explains how to write coverage into the contract.
Worked scenario: a payment callback that stopped overnight
A subscription business sells monthly plans online. Its payment provider sends a callback when a renewal succeeds, and the application activates the account.
-
02:10
- What happens
- The provider changes a callback signature format; callbacks start failing validation
- State
- Renewals charged but accounts not extended
- Owner
- —
-
02:15
- What happens
- The journey check "renewal extends account" fails and alerts the on-duty contact and the backup
- State
- Detected
- Owner
- Monitoring
-
08:05
- What happens
- On-duty engineer opens the runbook, confirms failures are callbacks only, and enables the queue that stores rejected callbacks
- State
- Contained; no data lost
- Owner
- Provider
-
08:40
- What happens
- Business owner approves a short customer message to affected users
- State
- Communicated
- Owner
- Buyer
-
10:30
- What happens
- Fix released through the normal change route; queued callbacks replayed
- State
- Recovered
- Owner
- Provider, with buyer approval
-
Within a week
- What happens
- Written review: add provider change notices to the dependency list and a test for signature formats
- State
- Learning recorded
- Owner
- Both
Without the journey check, the first signal would have been customer complaints about access, hours later and in larger numbers. The release readiness checklist covers the checks that stop the fix itself from causing a second incident.
What changes when the incident is a security compromise?
A compromise needs investigation before repair, because cleaning first destroys the evidence of how the attacker got in. For an existing client store, Netbase scoped a three-phase recovery of a compromised Magento site: investigation and security scan with cause analysis and a damage assessment, delivering an infected-file report and a recovery plan; cleaning and system recovery with malware and backdoor removal; and file restore, recoding and hardening with security patches applied and the administrator area hardened. It typically takes 6 to 13 days in total depending on the damage, and recovery is not guaranteed. The Magento recovery and hardening record shows that order of work.
Preparation limits the damage before it happens. At Netbase, security practices include secure code review and version control, role-based access control, MFA for admin dashboards, contributors under NDA, and NDAs and DPAs on request. The data security and compliance guide covers how access is granted to an outside team.
Which questions should you ask your operations provider?
- Which journeys will you monitor end to end? A good answer names business journeys, not only servers.
- Who receives an alert, and who is the backup? Expect names and a tested route, not a mailbox.
- When did you last restore a backup of this system? Look for a date and a recorded duration.
- What happens to an alert outside your coverage hours? Expect a written answer that matches the contract.
- Who leads during an incident, and who talks to our business? Look for named roles, even in a small team.
- What does your post-incident review contain? Expect a timeline, causes, actions with owners and a follow-up date, without blame.
- How do you handle a suspected compromise? Look for investigation before cleaning, and no promise of guaranteed recovery.
What usually goes wrong?
- Monitoring servers, not journeys. Signal: dashboards are green while customers complain. Owner: the service owner, who adds end-to-end checks.
- Alerts nobody owns. Signal: an alert fired hours before anyone acted. Owner: the operations lead, who names recipients and backups.
- Untested restores. Signal: a restore fails or takes far longer than expected during an outage. Owner: the provider, who rehearses on a schedule.
- Too many people fixing at once. Signal: conflicting changes during an incident. Owner: the incident lead, who controls changes.
- Reviews that assign blame. Signal: people stop reporting near-misses. Owner: the engineering leader, who keeps reviews blameless.
- Actions that never land. Signal: the same incident repeats. Owner: the project manager, who tracks review actions weekly.
How this guide is sourced and where it stops
This guide draws on the incident management and postmortem chapters of Google's Site Reliability Engineering book and on Netbase records approved in the OutsourcingVN claim register: the scoped Magento recovery, office coverage, the review rhythm and security practices. The methodology explains how those records are sourced. Netbase publishes no uptime figures or response-time commitments, and the payment-callback scenario is illustrative rather than a client record.
Plan the next step for your project
Common questions
Choose the one journey the business cannot lose, add an end-to-end check for it, and route the alert to a named person with a backup. That single step turns the most expensive failure from a customer complaint into an alert, and it gives the rest of the readiness work a starting point.
Yes, even if one person holds several. Name who leads and who informs the business before an incident happens. In a two-person team the leader may also fix, but the other person should not change the system without telling the leader, and someone must keep a short running record.
Often enough that the last successful restore is recent and its duration is known, and always after a change to the database, hosting or backup tooling. A backup that has never been restored is an assumption, and incidents are the worst time to test assumptions.
A timeline, what users experienced, the contributing causes, what helped and what slowed recovery, and a short list of actions with owners and dates. Keep it blameless so people report honestly, and check at the next review whether the actions were completed.
It should be written into the managed operations scope: which journeys are monitored, who receives alerts, the coverage window, the restore rehearsal schedule and the review habit. Without that, the provider may run the system without being ready for its failures.
Make your next incident boring
Start with the critical journey, the current alert route and the date of your last restore. Managed Operations is the service route for running a system with this readiness in place, and the software takeover and maintenance guide covers inheriting a system before readiness work begins.
OutsourcingVN is Netbase's own outsourcing-services platform. Submit a project with the journeys you cannot afford to lose, and a person will reply with what a readiness review would check first.
Related services and solutions
Application maintenance and managed operations, agreed in an assessment
Managed Operations keeps a production system maintained and improving under an agreement written after a paid assessment. The assessment defines which systems are covered, what "covered" means, the service levels and how you leave. No response time or service level is published: those numbers go into a contract after the assessment, or they do not exist.
Learn More