Skip to content

Incidents

What happens when something goes wrong, on managed hosting and on your own.

Level Means Example
SEV1 Down, or data at risk The deployment is unreachable. A suspected breach
SEV2 Badly degraded Email is not sending. A connector the desk depends on is failing
SEV3 Working, with a problem A report is wrong. A screen is slow

Severity describes impact, not cause.

Work in this order, and leave the cause until last.

  1. Confirm the scope. One tenant or all of them. One feature or everything. /api/health/ready and /api/health/version answer the basic questions without a database.
  2. Say something. An acknowledgement with what you know beats a polished statement twenty minutes later. On managed hosting we tell the affected customers; self-hosted, tell your own school.
  3. Stop it getting worse. Roll back if a change is implicated. The update script prints its own rollback command.
  4. Preserve evidence. Take a copy of the logs before restarting anything. Restarting is often the fix, and it destroys the state you would want to look at. This is the step people skip.
  5. Then investigate.

Tell people what you know. “The API is returning 500 on every request and we are rolling back this morning’s release” is more useful than “we are investigating an issue”.

Use more than one channel. The status page is one, and it does not help if Plugboard itself is down. Add email or a chat channel.

Say when you will next update, and then do, even if the update is that you know no more.

A suspected breach is SEV1, however small it looks operationally.

Notification obligations depend on jurisdiction and are generally measured in hours rather than days. GDPR is 72 hours to the supervisory authority. Australia’s notifiable data breaches scheme requires an assessment within 30 days and notification as soon as practicable once it is established. Others differ.

Know your obligation before you need it. Whoever owns privacy at your school should be part of the response, not told afterwards.

Plugboard’s audit log is the primary evidence source: it records who accessed what and when, and it cannot be edited or deleted through the application by anybody.

We tell you about anything affecting your deployment, with what we know at the time.

Your audit log is yours during an incident as at any other time, including vendor access entries.

Restores are requested as described in support, restores and exits. We confirm what will be lost before starting one.

The response is yours. Write these down before you need them:

  • Who decides to roll back, and who they call at 7am.
  • Where the backups are, and who can restore one.
  • Where SECRETS_MASTER_KEY is, and who has it. Not on the server.
  • How you tell your community, if the desk is down during a school day.

A page in your ICT runbook, reviewed once a year, is enough.

Write down what happened. The same class of failure recurs, and the note is what makes it shorter the second time.

Include the timeline. When it started, when it was noticed, when it was fixed. The gap between the first two is usually a monitoring problem, and usually the part you can act on.

Change one thing. A retrospective with twelve actions gets none of them done.

If you have found a security issue in Plugboard rather than an outage, email [email protected]. We confirm receipt within one business day.