Incidents
What happens when something goes wrong, on managed hosting and on your own.
Severity
Section titled “Severity”| Level | Means | Example |
|---|---|---|
| SEV1 | Down, or data at risk | The deployment is unreachable. A suspected breach |
| SEV2 | Badly degraded | Email is not sending. A connector the desk depends on is failing |
| SEV3 | Working, with a problem | A report is wrong. A screen is slow |
Severity describes impact, not cause.
The first fifteen minutes
Section titled “The first fifteen minutes”Work in this order, and leave the cause until last.
- Confirm the scope. One tenant or all of them. One feature or everything.
/api/health/readyand/api/health/versionanswer the basic questions without a database. - Say something. An acknowledgement with what you know beats a polished statement twenty minutes later. On managed hosting we tell the affected customers; self-hosted, tell your own school.
- Stop it getting worse. Roll back if a change is implicated. The update script prints its own rollback command.
- Preserve evidence. Take a copy of the logs before restarting anything. Restarting is often the fix, and it destroys the state you would want to look at. This is the step people skip.
- Then investigate.
Communication
Section titled “Communication”Tell people what you know. “The API is returning 500 on every request and we are rolling back this morning’s release” is more useful than “we are investigating an issue”.
Use more than one channel. The status page is one, and it does not help if Plugboard itself is down. Add email or a chat channel.
Say when you will next update, and then do, even if the update is that you know no more.
Data breaches
Section titled “Data breaches”A suspected breach is SEV1, however small it looks operationally.
Notification obligations depend on jurisdiction and are generally measured in hours rather than days. GDPR is 72 hours to the supervisory authority. Australia’s notifiable data breaches scheme requires an assessment within 30 days and notification as soon as practicable once it is established. Others differ.
Know your obligation before you need it. Whoever owns privacy at your school should be part of the response, not told afterwards.
Plugboard’s audit log is the primary evidence source: it records who accessed what and when, and it cannot be edited or deleted through the application by anybody.
Managed hosting
Section titled “Managed hosting”We tell you about anything affecting your deployment, with what we know at the time.
Your audit log is yours during an incident as at any other time, including vendor access entries.
Restores are requested as described in support, restores and exits. We confirm what will be lost before starting one.
Self-hosted
Section titled “Self-hosted”The response is yours. Write these down before you need them:
- Who decides to roll back, and who they call at 7am.
- Where the backups are, and who can restore one.
- Where
SECRETS_MASTER_KEYis, and who has it. Not on the server. - How you tell your community, if the desk is down during a school day.
A page in your ICT runbook, reviewed once a year, is enough.
Afterwards
Section titled “Afterwards”Write down what happened. The same class of failure recurs, and the note is what makes it shorter the second time.
Include the timeline. When it started, when it was noticed, when it was fixed. The gap between the first two is usually a monitoring problem, and usually the part you can act on.
Change one thing. A retrospective with twelve actions gets none of them done.
Reporting a vulnerability
Section titled “Reporting a vulnerability”If you have found a security issue in Plugboard rather than an outage, email [email protected]. We confirm receipt within one business day.