← Security

Business continuity and backup plan

Last updated 2026-09-25. Owner: founder.

Status: the restore procedure below is written but has NOT been tested in production. It must be run against a non-production copy before the first live firm, then every 6 months, and the results recorded in the test log at the end of this file.


1. What we protect and how

Data Primary copy Backup / protection Recovery point Notes
Structured data (blotter, versions, audit log, users) Neon Postgres, US Neon point-in-time restore (PITR) within the project's history-retention window Any moment inside the window Set the window to the plan maximum and record it here: ____ days (not yet confirmed).
Images, attachments, import files, exam packages AWS S3, one US region S3 versioning + Object Lock compliance mode; S3 stores data redundantly across multiple facilities in the region No loss expected: objects cannot be deleted or overwritten No second-region copy (not done; needs replication with Object Lock on the destination).
KMS key AWS KMS Automatic yearly rotation keeps old key material; 30-day deletion waiting period; Terraform prevent_destroy n/a Losing the key = losing access to every S3 object. Alarm on deletion requests is not done.
Field-encryption data key KMS_DATA_KEY_CIPHERTEXT in Netlify Second copy of the ciphertext in the founder's password manager n/a Without it, encrypted fields (client account numbers, 2FA secrets, SSO secrets) cannot be read.
Application code Git repository Remote on GitHub (to be confirmed) + founder's laptop Last commit Netlify keeps previous deploys for instant rollback.
Infrastructure Terraform files in the repo Re-apply with Terraform n/a Bucket and key are never recreated; they are protected from destroy.
Secrets Netlify environment variables Founder's password manager n/a Each can be rotated at its provider.

2. Targets

These are targets, not yet proven by a test:

  • Recovery point objective (RPO): database: minutes (PITR); files: zero.
  • Recovery time objective (RTO): 8 business hours for a database restore; 1 hour for a bad deploy (rollback); for a full AWS-region or Neon-region outage, CheckTrail waits for the provider to recover (no second region today).

During an outage, firms continue to log checks on paper per their own written procedures and enter them afterwards; the app records the true receipt time and flags late logging.

3. Scenarios

Scenario Response
Bad deploy Roll back in Netlify to the previous deploy.
Data wrongly changed or corrupted in Postgres Record tables are insert-only, so this mainly affects cached/state columns. Recompute state (engine sync); if needed, restore a PITR branch and compare (procedure A).
Postgres project lost or unusable Procedure A into a new branch or new project.
S3 object unreadable Check the KMS key is enabled; check IAM; the object itself cannot have been deleted.
KMS key disabled Re-enable it. If deletion was scheduled, cancel it within the 30-day window.
AWS region outage Wait for AWS. App can run read-only against Postgres; images unavailable until AWS returns.
Netlify outage Wait, or deploy the same code to another Node host from the repository (not rehearsed).
Resend / Stripe / Sentry outage Non-critical. Alerts are stored as in-app notifications first, so they are not lost. An email that fails to send is logged and not retried (by design, to avoid duplicates); users still see the alert in the app. Billing waits.
Founder unavailable See the break-glass procedure in access-control-policy.md. No backup person named yet.

4. Procedure A — restore Postgres from point in time

Run as the founder with MFA. Never restore over production until a verified copy exists.

  1. Freeze writes if production is affected. Put the site in maintenance mode (Netlify: deploy the maintenance page or disable functions) so nothing new is written during restore.
  2. Pick the target time. Use the incident log and audit log to choose a time just before the problem, in UTC.
  3. Create a restore branch. Neon console → project → Branches → Create branch → from the production branch "at a point in time" → the target time. Name it restore-YYYYMMDD-HHMM. (The CLI equivalent is neonctl branches create --parent <prod-branch>@<timestamp>.)
  4. Check roles. On the new branch, confirm the checktrail_app role exists and has the expected grants: SELECT has_table_privilege('checktrail_app', '"AuditLog"', 'UPDATE'); → must be false. SELECT has_table_privilege('checktrail_app', '"AuditLog"', 'INSERT'); → must be true. If grants are missing, run SELECT ct_apply_grants(); as the owner.
  5. Check triggers. SELECT tgname, tgrelid::regclass FROM pg_trigger WHERE tgname LIKE 'ct_%'; → triggers present on every insert-only table.
  6. Verify the audit chain. Point DATABASE_URL at the branch and run npm run audit:verify. Every firm must report OK.
  7. Spot-check data. Count rows in CheckEntry, CheckEntryVersion, AuditLog per firm and compare to the latest known counts. Open three entries across firms in a local build pointed at the branch; check their images load from S3 and their hashes match.
  8. Decide. Either (a) promote: in Neon, restore the production branch from the restore branch (or make the restore branch primary) — Neon keeps a backup of the replaced state; or (b) copy specific rows back. Record which, and why.
  9. Reconnect and reopen. Update Netlify DATABASE_URL if the endpoint changed, redeploy, leave maintenance mode.
  10. Post-checks. Sign in as a test user; log a test check in a sample firm; run an exam export for a sample date range and confirm its manifest hashes.
  11. Record the result in the test log below (or the incident log).

5. Procedure B — confirm S3 records are intact

  1. Pick 10 random CheckImage rows and 2 ExportPackage rows.
  2. For each, aws s3api head-object and aws s3api get-object-retention: object present, mode COMPLIANCE, retain-until date as expected.
  3. Download and compare the SHA-256 with the database value (the app does this on every export).
  4. Try aws s3api delete-object with the app profile: must fail with AccessDenied.

6. Test log

Record every test. Use a non-production Neon branch and fake data only.

Date Tester Scenario / procedure Target time restored to Time taken (start → verified) Audit chain OK? Grants/triggers OK? S3 checks OK? Problems found Follow-up
not yet run

Checklist for each test:

  • Restore branch created from a point in time
  • checktrail_app role present; UPDATE on AuditLog = false; INSERT = true
  • All ct_ triggers present
  • npm run audit:verify OK for every firm
  • Row counts match expectations
  • Images load and hashes match (Procedure B)
  • Exam export generated and manifest hashes verified
  • Delete attempt on S3 denied
  • Time taken recorded and compared with the 8-hour target
  • This document updated with anything that was wrong