Skip to content

Keeping an eye on it

Three endpoints, one metrics scrape, and a handful of alerts. This page is short on purpose: there is not much to watch, and watching the wrong things is worse than watching nothing.

These need no tenant and no database, so they still answer when the instance is unhealthy.

Endpoint Answers
/api/health Liveness plus build information. Is the process up
/api/health/ready Readiness. Is it able to serve, including the database
/api/health/version Version and release channel
Terminal window
curl -s https://helpdesk.yourschool.org/api/health/version
{ "version": "0.1.0", "channel": "stable" }

They also need no credential, which is why they say this much and no more. The version names a release; the git commit and the build timestamp name one exact image, and since the release notes say which security fixes are in which release, publishing both to anybody who asks would answer whether your instance still carries a given bug.

To read the exact build, set BUILD_INFO_TOKEN to a long random string, restart, and present it:

Terminal window
curl -s -H "Authorization: Bearer $BUILD_INFO_TOKEN" \
https://helpdesk.yourschool.org/api/health/version
{ "version": "0.1.0", "channel": "stable", "commit": "9f2c1ab44e01", "builtAt": "2026-08-01T09:14:22Z" }

With no token configured — the default — nobody gets the commit. A missing or wrong token is answered with the short form rather than an error, so setting this can never stop a health check working.

Point your uptime checker at /api/health/ready. Point your load balancer at the same. /api/health is for “is the process alive”, which is a different question and a less useful one.

Prometheus metrics are exposed for scraping, including plugboard_build_info, which carries the version as a label. That last one is worth graphing: it makes a fleet-wide “who is on what version” answerable at a glance, and it makes a failed update visible as a version that did not change.

/api/metrics is open unless you set METRICS_TOKEN, so it carries the same restraint as the health endpoints: plugboard_build_info is labelled with the version and the channel, and gains a commit label only once a token is required to scrape it.

Use the row for the deployment actually installed. Native release packaging currently targets Linux x64; Windows/macOS paths are reference-launcher instructions and do not imply available packages.

Install Where
systemd journalctl -u plugboard -f
Validated reference bundle on Windows or macOS logs/ inside the install directory

In the current release candidate, native startup restricts retained logs to the installation owner (plus SYSTEM and administrators on Windows). Setup, tracking and feedback link paths are redacted in application logs and native diagnostics. Do not use old logs to recover an administrator account: use the private recovery flow below.

Errors can also be sent to your own collector. Set SENTRY_DSN to a self-hosted Sentry or GlitchTip. There is no default destination, and nothing is sent anywhere unless you set it.

In rough order of how much it will matter at 8:40 on a Monday.

Alert Condition Why
Instance down /api/health/ready fails twice in a row Everything else is downstream of this
Certificate expiring Under 14 days The most common self-inflicted outage
Disk above 85 percent On the volume holding the database and backups A full disk stops PostgreSQL and stops backups, in that order
Backup failed The backup.failed event You find out now, not during a restore
Version unchanged after an update plugboard_build_info still shows the old version A tag that resolved to the old image
Database connections near the limit Depends on your max_connections Usually a sign something is holding transactions open

Plugboard has its own service monitor for the things around it: the intranet, the print server, the Wi-Fi controller, a connector agent’s heartbeat. That is for the services your users depend on.

Do not use it to monitor Plugboard itself. A monitor that lives inside the thing it is watching cannot tell you the thing is down.

  • External uptime check on /api/health/ready, from outside the school network if the desk is reachable from outside.
  • Plugboard’s own monitors for everything else you run, with the ones your community cares about marked public so they land on the status page.
  • Email alerts going to a monitored mailbox, not a personal one. The monitor.down and backup.failed messages are the two that matter.
Task Frequency
Confirm a backup completed and verify one Weekly, automated
Restore a backup somewhere and sign in Annually, by hand
Review the audit log for accounts that should be gone Termly
Check for a new release and read its changelog Monthly
Review connectors still in demo mode After any onboarding
Confirm certificate renewal happened Whenever it renews

Plugboard is not demanding, and the numbers that grow are predictable.

  • The audit log is the largest table on any mature install. Retention is configurable; see audit and retention.
  • Submissions and tickets grow linearly with the desk’s workload and are small.
  • Realtime connections cost one open HTTP connection per signed-in console tab. A desk with twelve technicians is not a load problem.

If the database is slow, check that endpoint protection is not scanning every write to the data directory. That is the usual answer on Windows.

Source and container administrator recovery

Section titled “Source and container administrator recovery”

For a built source installation, run node scripts/plugboard.mjs setup-link <admin-email> under the account that owns its protected .env. Add --embedded-db for an installation using the built-in database. Inside an API container, use --from-environment to select its injected configuration.

The named address must identify one active school administrator. The command writes owner-private .setup-link.html, prints only that file’s path, and audits the operation. Open the file to claim the account. The link works once for four hours and replaces any previous setup link; creating it does not change the password. Remove the file after use. Never send it with diagnostics.