68 lines
3.2 KiB
Markdown
68 lines
3.2 KiB
Markdown
# Business Continuity Audit — AUD-007 | CC-047 | 2026-07-31
|
|
# Act as Principal Technology Auditor — READ_ONLY audit
|
|
|
|
## Scope
|
|
|
|
CEO-OS platform on Hetzner (91.98.39.120) and supersrv (152.53.112.35).
|
|
Components: ceo-api (NestJS), ceo-web (Next.js), ceo-worker (BullMQ), Supabase Postgres, Redis.
|
|
|
|
---
|
|
|
|
## Evidence Register
|
|
|
|
| # | Evidence | Label |
|
|
|---|---------|-------|
|
|
| E1 | Zerobyte agent running (`pgrep zerobyte`), SQLite at `/var/lib/zerobyte/data/zerobyte.db` | FACT |
|
|
| E2 | Zerobyte HTTP on port 4096 requires auth — direct API inaccessible | FACT |
|
|
| E3 | Docker images `ceo-api-staging:cc045`, `ceo-api-staging:cc046` on supersrv | FACT |
|
|
| E4 | Git tags `cc-045-complete`, `cc-046-complete` on Forgejo | FACT |
|
|
| E5 | Coolify manages service lifecycle on Hetzner (auto-restart on failure) | FACT |
|
|
| E6 | Uptime Kuma running on supersrv (`uptime-kuma-l12o8zmhvwq60mi598gq76gy`) | FACT |
|
|
| E7 | No automated database restore test has been run | FACT |
|
|
| E8 | RTO measured via TST-017 script (container restart): ~8-12s | FACT |
|
|
| E9 | No multi-region failover — single Hetzner datacenter | FACT |
|
|
| E10 | Manual rollback procedure: `docker stop` + `docker run previous-tag` | FACT |
|
|
|
|
---
|
|
|
|
## Findings
|
|
|
|
| ID | Severity | Finding | Evidence |
|
|
|----|----------|---------|---------|
|
|
| BC-F1 | HIGH | No automated database restore test — Zerobyte backup health is not verified by automated CI | E7 — TST-016 script created but not wired to CI |
|
|
| BC-F2 | HIGH | Single datacenter — no geographic redundancy; Hetzner AZ failure = full outage | E9 |
|
|
| BC-F3 | MEDIUM | No documented RPO (Recovery Point Objective) — unclear how much data loss is acceptable | INFERENCE |
|
|
| BC-F4 | MEDIUM | No runbook for database corruption scenario (only container restart tested) | E10 |
|
|
| BC-F5 | MEDIUM | Coolify auto-restart behavior not tested — only manual restart was verified in TST-017 | E5 |
|
|
| BC-F6 | LOW | No alerting configured beyond Uptime Kuma — no PagerDuty/Slack integration | E6 |
|
|
|
|
---
|
|
|
|
## RTO / RPO Targets (proposed — not yet formally accepted)
|
|
|
|
| Component | RTO Target | RPO Target | Current Status |
|
|
|-----------|-----------|-----------|----------------|
|
|
| API (ceo-api) | < 60s | N/A (stateless) | PASS — TST-017 measured ~10s |
|
|
| Database (Postgres) | < 4h | < 1h | UNKNOWN — no restore test run |
|
|
| Redis (queue) | < 5min | Acceptable loss (queue is ephemeral) | OK |
|
|
| Full platform | < 8h | < 4h | UNKNOWN |
|
|
|
|
---
|
|
|
|
## Remediation Plan
|
|
|
|
| ID | Priority | Action | Owner | Target |
|
|
|----|----------|--------|-------|--------|
|
|
| BC-F1 | HIGH | Wire `scripts/test-backup-restore.sh` to weekly CI job | Engineering | CC-049 |
|
|
| BC-F3 | HIGH | Document and formally accept RPO/RTO targets | Engineering + Owner | Before production |
|
|
| BC-F4 | MEDIUM | Create database restore runbook (Supabase snapshot → restore) | Engineering | CC-048 |
|
|
| BC-F2 | LOW | Evaluate Hetzner volume snapshots + offsite backup (Cloudflare R2) | Engineering | CC-050+ |
|
|
| BC-F5 | MEDIUM | Test Coolify auto-restart with `restart: always` Docker policy | Ops | CC-048 |
|
|
|
|
---
|
|
|
|
## Release Conclusion
|
|
|
|
**Status**: NOT READY for production without completing BC-F1 (backup restore test) and BC-F3 (RPO/RTO acceptance).
|
|
|
|
The platform is ready for staging/beta with manual operations procedures.
|