Files
monlet/docs/ops/backup-restore.md
2026-05-28 14:19:27 +04:00

3.4 KiB
Raw Permalink Blame History

Backup and restore

All persistent state lives in PostgreSQL (agents, checks, events, incidents, notification_outbox). The server is stateless; the agent local spool is best-effort and not part of the canonical state (ADR-0006).

What to back up

  • The Monlet PostgreSQL database (logical or physical).
  • The server MONLET_AUTH_TOKEN and .env (kept in your secret store, not in backups of the DB).
  • The agent config.toml per host and local state_dir/agent.key if you need to preserve accepted identity without re-acceptance.
PGPASSWORD=monlet pg_dump \
  -h $PGHOST -U monlet -d monlet \
  --format=custom --no-owner --no-privileges \
  --file=monlet-$(date -u +%Y%m%dT%H%M%SZ).dump

Restore into an empty database:

createdb -h $PGHOST -U postgres monlet
PGPASSWORD=monlet pg_restore \
  -h $PGHOST -U monlet -d monlet \
  --no-owner --no-privileges \
  monlet-YYYYMMDDTHHMMSSZ.dump

The schema is owned by Alembic. Restoring a logical dump from the same Monlet version is safe. After restore, run alembic upgrade head only if you restored from an older version.

Retention

  • events is partitioned by received_at (one monthly partition per range). Retention is controlled by MONLET_EVENTS_RETENTION_MONTHS (default 36). The partition maintenance loop drops old partitions on MONLET_PARTITION_MAINTENANCE_INTERVAL_SEC — disk is reclaimed by DROP PARTITION, not by row-level DELETE, so reclaim is bulk-fast and lock-light. Do NOT run row-level DELETE FROM events — it bloats the table and conflicts with partition rotation.
  • event_ingest_dedup is the idempotency table; TTL MONLET_EVENT_DEDUP_RETENTION_DAYS (default 30) prunes rows in bounded batches each detector tick. After dedup TTL elapses an agent replaying a very old spooled event will be accepted again (no double-side effects because the corresponding events partition is already dropped on the same horizon).
  • notification_outbox retention is bounded by MONLET_OUTBOX_RETENTION_MAX_ROWS (default 50000) and pruned in batches of MONLET_OUTBOX_PRUNE_BATCH_SIZE per tick (PH-012). Only terminal states (sent, failed, discarded) are eligible — active rows (pending, sending, retry) are never deleted.
  • incidents is small and not pruned automatically; size scales with operator activity rather than ingestion rate.

Disaster recovery

  • Restore the database.
  • Restart the server. Agents will resume sending; spooled events on each agent are replayed and de-duplicated by event_id (ADR-0006).
  • The agent spool is bounded (10 000 events / 50 MiB FIFO, ADR-0005). During a long outage the oldest events get dropped to make room for new ones, so observations from the start of the outage are lost first. Plan recovery time accordingly — for an agent producing ~1 event/s per check, ~10 000 events is roughly 23 hours of buffering for a handful of checks.
  • Spool overflow and non-retryable (4xx) server rejections are observable via monlet_agent_events_dropped_total{reason} with reason ∈ spool_overflow_events | spool_overflow_bytes | send_non_retryable. Alert on any non-zero rate. Operator action: investigate the agent log (one warning per ~30s drop burst), the server (for 4xx), and whether spool limits need raising in state_dir.
  • If agents were also lost, only events that were not yet persisted to the spool file are lost in addition.