3.2 KiB
3.2 KiB
Business Continuity Audit — AUD-007 | CC-047 | 2026-07-31
Act as Principal Technology Auditor — READ_ONLY audit
Scope
CEO-OS platform on Hetzner (91.98.39.120) and supersrv (152.53.112.35). Components: ceo-api (NestJS), ceo-web (Next.js), ceo-worker (BullMQ), Supabase Postgres, Redis.
Evidence Register
| # | Evidence | Label |
|---|---|---|
| E1 | Zerobyte agent running (pgrep zerobyte), SQLite at /var/lib/zerobyte/data/zerobyte.db |
FACT |
| E2 | Zerobyte HTTP on port 4096 requires auth — direct API inaccessible | FACT |
| E3 | Docker images ceo-api-staging:cc045, ceo-api-staging:cc046 on supersrv |
FACT |
| E4 | Git tags cc-045-complete, cc-046-complete on Forgejo |
FACT |
| E5 | Coolify manages service lifecycle on Hetzner (auto-restart on failure) | FACT |
| E6 | Uptime Kuma running on supersrv (uptime-kuma-l12o8zmhvwq60mi598gq76gy) |
FACT |
| E7 | No automated database restore test has been run | FACT |
| E8 | RTO measured via TST-017 script (container restart): ~8-12s | FACT |
| E9 | No multi-region failover — single Hetzner datacenter | FACT |
| E10 | Manual rollback procedure: docker stop + docker run previous-tag |
FACT |
Findings
| ID | Severity | Finding | Evidence |
|---|---|---|---|
| BC-F1 | HIGH | No automated database restore test — Zerobyte backup health is not verified by automated CI | E7 — TST-016 script created but not wired to CI |
| BC-F2 | HIGH | Single datacenter — no geographic redundancy; Hetzner AZ failure = full outage | E9 |
| BC-F3 | MEDIUM | No documented RPO (Recovery Point Objective) — unclear how much data loss is acceptable | INFERENCE |
| BC-F4 | MEDIUM | No runbook for database corruption scenario (only container restart tested) | E10 |
| BC-F5 | MEDIUM | Coolify auto-restart behavior not tested — only manual restart was verified in TST-017 | E5 |
| BC-F6 | LOW | No alerting configured beyond Uptime Kuma — no PagerDuty/Slack integration | E6 |
RTO / RPO Targets (proposed — not yet formally accepted)
| Component | RTO Target | RPO Target | Current Status |
|---|---|---|---|
| API (ceo-api) | < 60s | N/A (stateless) | PASS — TST-017 measured ~10s |
| Database (Postgres) | < 4h | < 1h | UNKNOWN — no restore test run |
| Redis (queue) | < 5min | Acceptable loss (queue is ephemeral) | OK |
| Full platform | < 8h | < 4h | UNKNOWN |
Remediation Plan
| ID | Priority | Action | Owner | Target |
|---|---|---|---|---|
| BC-F1 | HIGH | Wire scripts/test-backup-restore.sh to weekly CI job |
Engineering | CC-049 |
| BC-F3 | HIGH | Document and formally accept RPO/RTO targets | Engineering + Owner | Before production |
| BC-F4 | MEDIUM | Create database restore runbook (Supabase snapshot → restore) | Engineering | CC-048 |
| BC-F2 | LOW | Evaluate Hetzner volume snapshots + offsite backup (Cloudflare R2) | Engineering | CC-050+ |
| BC-F5 | MEDIUM | Test Coolify auto-restart with restart: always Docker policy |
Ops | CC-048 |
Release Conclusion
Status: NOT READY for production without completing BC-F1 (backup restore test) and BC-F3 (RPO/RTO acceptance).
The platform is ready for staging/beta with manual operations procedures.