ceo-api/docs/business-continuity-audit.md

3.2 KiB

Business Continuity Audit — AUD-007 | CC-047 | 2026-07-31

Act as Principal Technology Auditor — READ_ONLY audit

Scope

CEO-OS platform on Hetzner (91.98.39.120) and supersrv (152.53.112.35). Components: ceo-api (NestJS), ceo-web (Next.js), ceo-worker (BullMQ), Supabase Postgres, Redis.


Evidence Register

# Evidence Label
E1 Zerobyte agent running (pgrep zerobyte), SQLite at /var/lib/zerobyte/data/zerobyte.db FACT
E2 Zerobyte HTTP on port 4096 requires auth — direct API inaccessible FACT
E3 Docker images ceo-api-staging:cc045, ceo-api-staging:cc046 on supersrv FACT
E4 Git tags cc-045-complete, cc-046-complete on Forgejo FACT
E5 Coolify manages service lifecycle on Hetzner (auto-restart on failure) FACT
E6 Uptime Kuma running on supersrv (uptime-kuma-l12o8zmhvwq60mi598gq76gy) FACT
E7 No automated database restore test has been run FACT
E8 RTO measured via TST-017 script (container restart): ~8-12s FACT
E9 No multi-region failover — single Hetzner datacenter FACT
E10 Manual rollback procedure: docker stop + docker run previous-tag FACT

Findings

ID Severity Finding Evidence
BC-F1 HIGH No automated database restore test — Zerobyte backup health is not verified by automated CI E7 — TST-016 script created but not wired to CI
BC-F2 HIGH Single datacenter — no geographic redundancy; Hetzner AZ failure = full outage E9
BC-F3 MEDIUM No documented RPO (Recovery Point Objective) — unclear how much data loss is acceptable INFERENCE
BC-F4 MEDIUM No runbook for database corruption scenario (only container restart tested) E10
BC-F5 MEDIUM Coolify auto-restart behavior not tested — only manual restart was verified in TST-017 E5
BC-F6 LOW No alerting configured beyond Uptime Kuma — no PagerDuty/Slack integration E6

RTO / RPO Targets (proposed — not yet formally accepted)

Component RTO Target RPO Target Current Status
API (ceo-api) < 60s N/A (stateless) PASS — TST-017 measured ~10s
Database (Postgres) < 4h < 1h UNKNOWN — no restore test run
Redis (queue) < 5min Acceptable loss (queue is ephemeral) OK
Full platform < 8h < 4h UNKNOWN

Remediation Plan

ID Priority Action Owner Target
BC-F1 HIGH Wire scripts/test-backup-restore.sh to weekly CI job Engineering CC-049
BC-F3 HIGH Document and formally accept RPO/RTO targets Engineering + Owner Before production
BC-F4 MEDIUM Create database restore runbook (Supabase snapshot → restore) Engineering CC-048
BC-F2 LOW Evaluate Hetzner volume snapshots + offsite backup (Cloudflare R2) Engineering CC-050+
BC-F5 MEDIUM Test Coolify auto-restart with restart: always Docker policy Ops CC-048

Release Conclusion

Status: NOT READY for production without completing BC-F1 (backup restore test) and BC-F3 (RPO/RTO acceptance).

The platform is ready for staging/beta with manual operations procedures.