From 1d027f71f6110f093692635fb8f3adfd9fa5aaf9 Mon Sep 17 00:00:00 2001 From: root Date: Sat, 18 Jul 2026 22:02:45 +0000 Subject: [PATCH 1/3] feat(contracts): add infrastructure-maintenance contract New responsibility contract consolidating host-level system maintenance and Docker image lifecycle management, filling the gap left by infrastructure-update (which owns cluster-wide apt waves). Scope: - OS package updates on primary host with pre-update snapshot/backup check - Docker image pulls for LiteLLM, SearXNG, and other running containers - Container restarts with per-stack health verification - Post-update verification: LiteLLM, SearXNG, Zulip, Gitea, PM2, Hermes gateways - Rollback on failure (image/apt/config restore) with circuit breaker Owner: ops (firstmate secondmate). Trigger: weekly Sunday 2am ET. Escalation: warning/critical->abiba+mumuni, fatal->abiba+mumuni+kwame. circuit_breaker: max_retries 2, window 7200, trip_action escalate_to_fatal. depends_on: infrastructure-monitoring (pre-update health baseline). Registry: - Add infrastructure-maintenance to by_category.maintenance, by_domain.infrastructure, by_owner.ops (new), by_trigger.scheduled, by_sensitivity.high - Add 'ops' to owners list - Move infrastructure-update owner abiba -> ops (in contracts entry + by_owner index) Also adds '## Maintaining this file' section to AGENTS.md per fm-ensure-agents-md. --- AGENTS.md | 7 + CLAUDE.md | 1 + contract-registry.yaml | 106 ++++++++++++++- infrastructure-maintenance.prose.md | 195 ++++++++++++++++++++++++++++ 4 files changed, 307 insertions(+), 2 deletions(-) create mode 120000 CLAUDE.md create mode 100644 infrastructure-maintenance.prose.md diff --git a/AGENTS.md b/AGENTS.md index ea7e0d9..8f41e82 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -128,3 +128,10 @@ safe-mutate --verify "CMD" [--expect "PATTERN"] --mutate "CMD" [--reason "WHY"] Read the [Authoring Guide](docs/AUTHORING-GUIDE.md) before writing any new contract. It covers the full process: verify → draft → lint → review → ship, with templates and style rules. + +## Maintaining this file + +Keep this file for knowledge useful to almost every future agent session in this project. +Do not repeat what the codebase already shows; point to the authoritative file or command instead. +Prefer rewriting or pruning existing entries over appending new ones. +When updating this file, preserve this bar for all agents and keep entries concise. diff --git a/CLAUDE.md b/CLAUDE.md new file mode 120000 index 0000000..47dc3e3 --- /dev/null +++ b/CLAUDE.md @@ -0,0 +1 @@ +AGENTS.md \ No newline at end of file diff --git a/contract-registry.yaml b/contract-registry.yaml index 9f7af92..94d899e 100644 --- a/contract-registry.yaml +++ b/contract-registry.yaml @@ -22,6 +22,7 @@ owners: - abiba - mumuni - kwame +- ops trigger_types: - scheduled - event_driven @@ -58,6 +59,7 @@ index: - memory-audit-maintenance - gpu-fleet - infrastructure-update + - infrastructure-maintenance reference: - infrastructure-control - ra-h-os-custodianship-contract @@ -90,6 +92,7 @@ index: - infrastructure-control - infrastructure-monitoring - infrastructure-update + - infrastructure-maintenance - pm2-self-heal - disk-gc-threat-response gpu: @@ -130,7 +133,6 @@ index: - build-zulip-plugin - stirling-pdf-agent-access - gpu-fleet - - infrastructure-update - infrastructure-control - zulip-adapter-lessons - pi-approval-architecture @@ -144,6 +146,9 @@ index: - mumuni-delegation kwame: - hello-world + ops: + - infrastructure-maintenance + - infrastructure-update by_trigger: scheduled: - hermes-key-enforcement @@ -156,6 +161,7 @@ index: - litellm-health - memory-audit-maintenance - infrastructure-update + - infrastructure-maintenance event_driven: - litellm-self-heal - pm2-self-heal @@ -197,6 +203,7 @@ index: - hermes-zulip-plugin - build-zulip-plugin - infrastructure-update + - infrastructure-maintenance - ra-h-os-custodianship-contract - mumuni-delegation normal: @@ -1256,13 +1263,108 @@ contracts: last_run: null last_status: null drift_alerts: [] +- name: infrastructure-maintenance + file: infrastructure-maintenance.prose.md + kind: responsibility + category: maintenance + sensitivity: high + status: active + owner: ops + version: 1.0.0 + trigger: + type: scheduled + cadence: 0 2 * * 0 + description: Weekly host-level maintenance Sunday at 2am ET (replaces infrastructure-update + build-phase role; infra-update moves to ops) + cron_job_id: null + execution: + agent: ops + timeout: 3600 + requires: + - infrastructure-monitoring run within last 30 minutes (pre-update health baseline) + - Proxmox snapshot of primary host OR /tmp backup dir created this run + - LiteLLM master key from Infisical vault for health verification + protocol: + - Load contract from prose-contracts/main + - Phase 0 preflight — capture health baseline, backup check, record image baseline + - Phase 1 apt update && apt upgrade -y on primary host + - Phase 2 docker compose pull for LiteLLM, SearXNG, and other running containers + - Phase 3 restart stacks one at a time with per-stack health verification + - Phase 4 verify every critical service (LiteLLM, SearXNG, Zulip, Gitea, PM2, Hermes gateways) + - On failure — rollback per protocol, escalate, do not loop beyond circuit breaker + - Log actions to ~/.hermes/runs/infrastructure-maintenance/ + verification: + postconditions: + - check: all critical services running after update + verify: 'curl -sf http://192.168.68.116/litellm/v1/models && curl -sf https://chat.sysloggh.net/api/v1/server_settings && curl -sf https://git.sysloggh.net/api/v1/version && curl -sf http://192.168.68.7:8888 && pm2 jlist' + expect: all probes 200 OK / processes online + - check: no regressions from pre-update health baseline + verify: diff Phase 0 health-baseline against Phase 4 results + expect: no GREEN service turned RED + - check: docker containers on latest stable tags + verify: docker inspect --format '{{.Config.Image}}' per service matches image-baseline.pulled_tag + expect: all containers running pulled tags + - check: APT packages up to date with no held broken packages + verify: apt list --upgradable 2>/dev/null | wc -l and apt-get -s upgrade | grep -ci broken + expect: upgradable == 0, broken == 0 + artifact: maintenance run report with phase results and any rollback/escalation + verify_commands: + - curl -sf http://192.168.68.116/litellm/v1/models + - curl -sf http://192.168.68.7:8888 + - curl -sf https://chat.sysloggh.net/api/v1/server_settings + - curl -sf https://git.sysloggh.net/api/v1/version + - pm2 jlist + receipt: + format: json + storage: ~/.hermes/runs/infrastructure-maintenance/ + graph_node: true + schema: + contract: string + run_id: string + timestamp: ISO 8601 + agent: string + status: pass|fail|escalated + phase: preflight|apt|images|restarts|verify|rollback|done|failed + actions_taken: array + postconditions: array + drift_alerts: array + evidence_path: string + escalation: + info: + action: log_to_receipt + notify: [] + warning: + action: relay_alert + notify: + - abiba + - mumuni + critical: + action: relay_alert + notify: + - abiba + - mumuni + fatal: + action: relay_alert + pause + human_required + notify: + - abiba + - mumuni + - kwame + circuit_breaker: + max_retries: 2 + window: 7200 + trip_action: escalate_to_fatal + depends_on: + - infrastructure-monitoring + last_run: null + last_status: null + drift_alerts: [] - name: infrastructure-update file: infrastructure-update.prose.md kind: responsibility category: maintenance sensitivity: high status: active - owner: abiba + owner: ops version: 1.0.0 trigger: type: scheduled diff --git a/infrastructure-maintenance.prose.md b/infrastructure-maintenance.prose.md new file mode 100644 index 0000000..e9a3b1a --- /dev/null +++ b/infrastructure-maintenance.prose.md @@ -0,0 +1,195 @@ +--- +kind: responsibility +name: infrastructure-maintenance +description: > + Weekly system-level maintenance for the Syslog inference fleet: OS package + updates on the primary host, Docker image pulls for LiteLLM/SearXNG and other + running containers, container restarts with health verification, post-update + verification of every critical service (LiteLLM proxy, SearXNG, Zulip, Gitea, + PM2 processes, Hermes gateways), and rollback on failure. Consolidates the + raw shell scripts that previously did this piecemeal and closes the Docker + image management gap that infrastructure-update left open (infra-update owns + cluster-wide apt waves; this contract owns the host-level maintenance loop + on the primary host + Docker image lifecycle). Runs Sunday 2am ET. Owner: + ops (firstmate secondmate). Blast radius: an unverified image pull can break + LiteLLM (all agents lose inference) or SearXNG (search-stack down); a bad apt + upgrade can leave the host in a half-upgraded state. Pre-update backup check + and rollback are mandatory for this reason. +agent: ops +triggers: + - weekly (Sunday 02:00 ET) via cron + - on demand when ops/abiba triggers "infra maintenance" +version: 1.0.0 +--- + +## Maintains + +- maintenance-status: { phase: idle|preflight|apt|images|restarts|verify|rollback|done|failed, host, step, result, timestamp } +- image-baseline: { service, current_tag, pulled_tag, digest, updated_at } — last known-good image per container +- apt-state: { upgradable_before, upgradable_after, held_broken, kernel_reboot_required } +- health-baseline: snapshot of critical-service health captured pre-update (used for regression check post-update) +- rollback-snapshot: { backup_path, configs, image_digests, timestamp } — restore point created in preflight +- maintenance-history: array of past runs with phase results and any escalations + +## Scope + +Primary host is the maintenance host where apt updates apply. Docker image pulls +span the two Docker ecosystems that run critical services. Topology, CT IDs, and +IPs are live-state fields — verify against `infrastructure-control.prose.md` +(the source of truth) and the live system before mutating. + +| Host | IP | Role | Trust | +|------|----|------|-------| +| CT 116 (syslog-api) | 192.168.68.116 | LiteLLM proxy + Grafana + Prometheus (inference harness) | ⚠️ VERIFY-BEFORE-USE | +| VM 109 (docker-vm) | 192.168.68.7 | SearXNG + Firecrawl + home stack (Docker host) | ⚠️ VERIFY-BEFORE-USE | +| CT 117 (zulip) | 192.168.68.19 | Zulip (storepve bridge IP .19) | ⚠️ VERIFY-BEFORE-USE | +| Gitea | https://git.sysloggh.net | Prose-contracts + agent configs source control | ⚠️ VERIFY-BEFORE-USE | +| CT 100 (abiba/pi) | 192.168.68.24 | PM2 processes (pi agent harness) | ⚠️ VERIFY-BEFORE-USE | + +> "Primary host" for the apt phase is the host the ops agent runs maintenance +> from. Confirm which host that is against infrastructure-control before +> running; do not assume. If the ops agent is containerized/CT-based, apt runs +> inside that CT. + +## Requires + +- SSH/exec access to CT 116 (.116) and VM 109 (.7) for Docker operations +- `apt`, `docker`, `docker compose` available on target hosts +- LiteLLM master key available (Infisical vault, `LITELLM_API_KEY`) for health verification +- `infrastructure-monitoring` run completed within the last 30 minutes — provides the pre-update health baseline used by the regression check +- Writable backup directory `/tmp/infra-maintenance-backup-/` on each mutated host +- Proxmox snapshot of the primary host available (or confirmed not required) before apt phase + +## Continuity + +- Self-driven: weekly cron `0 2 * * 0` (Sunday 02:00 ET) +- Also wakes on: explicit "infra maintenance" trigger from ops/abiba +- Depends on `infrastructure-monitoring` for the pre-update health baseline — do not run if the last monitoring run is stale (>30 min) or RED; abort and escalate instead + +## Execution + +### Phase 0 — Preflight (snapshot/backup check + health baseline) + +1. **Capture health baseline** — run the `infrastructure-monitoring` postcondition checks (LiteLLM, Zulip, Gitea, SearXNG, Proxmox API) and record results as `health-baseline`. If any critical service is already down, **abort**: maintenance must not run on a degraded fleet. +2. **Backup check** — confirm a Proxmox snapshot of the primary host exists OR `/tmp/infra-maintenance-backup-/` was created this run. Snapshot critical config files into the backup dir: + - `/opt/inference-harness/docker-compose.yml`, `/opt/inference-harness/litellm_config.yaml` (CT 116) + - `/opt/search-stack/searxng/docker-compose.yml`, `/opt/search-stack/firecrawl-source/docker-compose.yaml` (VM 109) +3. **Record image baseline** — `docker inspect --format '{{.Image}} {{.Config.Image}}' ` for every running container on .116 and .7; store digests in `image-baseline` so rollback can restore them. +4. **Disk check** — `df -h` on each mutated host; abort if free space <20% (apt upgrade + image pulls need headroom). + +### Phase 1 — OS package updates (primary host) + +```bash +# On the primary host only (VERIFY host against infrastructure-control first) +apt update +apt upgrade -y +``` + +- Capture `apt list --upgradable` before and after → store in `apt-state`. +- If apt reports held/broken packages (`apt-get -s upgrade | grep -i broken`, or non-zero exit), **stop** — do not force. Record `held_broken` and go to rollback/escalate. +- If `/var/run/reboot-required` exists after upgrade, flag `kernel_reboot_required: true` in `apt-state` but **do not reboot automatically** — that's a separate coordinated action (see infra-update Wave 4). Note it in the report. + +### Phase 2 — Docker image pulls + +Pull latest stable tags for every running container. Do NOT pin to `:main`/`:nightly` — use stable tags where the compose file specifies them; otherwise `latest`. + +```bash +# CT 116 (.116) — inference harness +cd /opt/inference-harness && docker compose pull + +# VM 109 (.7) — search + home stacks +cd /opt/search-stack/searxng && docker compose pull +cd /opt/search-stack/firecrawl-source && docker compose pull +# any other running stacks on .7 (home stack, audiobookshelf) — pull per their compose files +``` + +- LiteLLM and SearXNG are the two explicitly required pulls; "any other running containers" means every stack with a compose file on .116 and .7. +- Record pulled tag + digest per service in `image-baseline`. + +### Phase 3 — Container restarts with health verification + +Restart one stack at a time, verify health before moving to the next. Do not restart everything at once — a failure mid-wave must leave the rest running. + +```bash +# CT 116 +cd /opt/inference-harness && docker compose up -d +# VM 109 +cd /opt/search-stack/searxng && docker compose up -d +cd /opt/search-stack/firecrawl-source && docker compose up -d +``` + +After each stack comes up, wait for health (max 120s): +- `docker ps` shows the container `Up` (and `healthy` if a healthcheck is defined) +- Service-specific probe passes (see Phase 4 probes) + +If a stack fails to come up within 120s, **stop the wave** and go to rollback for that stack only; do not proceed to the next. + +### Phase 4 — Post-update service verification + +After ALL updates (apt + images + restarts), verify every critical service is back up and matches the pre-update baseline. This is the regression gate. + +| Service | Probe | Expect | +|---------|-------|--------| +| LiteLLM proxy | `curl -sf http://192.168.68.116/litellm/v1/models` | 200 OK, models returned | +| LiteLLM MCP gateway | `curl -sf http://192.168.68.116:4000/mcp-rest/tools/list -H "Authorization: Bearer $MASTER_KEY"` | 90 tools (23 RA-H OS + 67 GitHub) | +| SearXNG | `curl -sf http://192.168.68.7:8888` | 200 OK | +| Zulip | `curl -sf https://chat.sysloggh.net/api/v1/server_settings` | 200 OK | +| Gitea | `curl -sf https://git.sysloggh.net/api/v1/version` | 200 OK | +| PM2 processes | `pm2 jlist` (CT 100) | all pi-agent processes `online` | +| Hermes gateways | SSH to Mumuni CT 114, Tanko CT 112; `systemctl is-active hermes-gateway` | `active` for each | + +Regression check: every service that was GREEN in `health-baseline` must still be GREEN. A service that was already RED (and caused a preflight abort) is excluded — but Phase 0 should have aborted before we got here. + +## Rollback Protocol + +If ANY service in Phase 4 fails to come back up (or regresses vs baseline): + +1. **Image rollback** — for the failing stack, restore the previous image: + ```bash + # Restore from recorded image-baseline digest + docker compose down + docker compose pull # or set image: in compose and up -d + docker compose up -d + ``` +2. **APT rollback** — restore the primary host from the Proxmox snapshot taken/confirmed in Phase 0. If no snapshot, `apt install =` per package using apt history (`/var/log/apt/history.log`). +3. **Config rollback** — restore configs from `/tmp/infra-maintenance-backup-/`. +4. **Re-verify** — re-run the Phase 4 probes on the rolled-back service. If still failing, escalate (do not loop — circuit breaker below). +5. **Escalate** — send a Zulip DM to abiba + mumuni with: failing service, phase, baseline vs current, rollback actions taken, backup path. + +## Circuit Breaker + +- `max_retries: 2` per failing phase — after 2 rollback attempts on the same service, stop and escalate. +- `window: 7200` seconds — no more than 2 retries within a 2-hour window. +- `trip_action: escalate_to_fatal` — when tripped, escalate to fatal (abiba + mumuni + kwame) and pause; a human must clear before the next scheduled run. + +## Report + +After completion (or on abort), emit a receipt (JSON) to `~/.hermes/runs/infrastructure-maintenance/` and send a Zulip DM summary: + +``` +🛠 Infrastructure Maintenance — YYYY-MM-DD + +Phase: apt | images | restarts | verify | rollback +Primary host: +APT: packages upgraded, held/broken, kernel_reboot_required= +Images pulled: LiteLLM , SearXNG , +Services: all GREEN | FAILED (rolled back) +Baseline regression: none |
+Backup: /tmp/infra-maintenance-backup-YYYYMMDD/ +Escalation: none | warning | critical | fatal +``` + +## Verification Postconditions + +- All critical services running after update (Phase 4 all GREEN) +- No regressions from pre-update health baseline (Phase 0 baseline) +- Docker containers on latest stable tags (`image-baseline.pulled_tag` recorded) +- APT packages up to date with no held broken packages (`apt-state.held_broken == 0`) + +## Related Contracts + +- `infrastructure-update.prose.md` — cluster-wide apt waves across the PVE cluster + CTs/VMs (the build-phase update); this contract handles the host-level maintenance loop and Docker image lifecycle. +- `infrastructure-monitoring.prose.md` — provides the pre-update health baseline (depends_on). +- `infrastructure-control.prose.md` — topology source of truth (CT IDs, IPs, hostnames). +- `litellm-health.prose.md` — LiteLLM probe details. +- `proxmox-monitor.prose.md` — Docker stats + monitoring stack health. From a68879e904d091d1236fd3e3ce83afca605e505b Mon Sep 17 00:00:00 2001 From: root Date: Sat, 18 Jul 2026 22:07:15 +0000 Subject: [PATCH 2/3] no-mistakes(review): fix docker rollback command to pin digest then compose up --- infrastructure-maintenance.prose.md | 5 +++-- 1 file changed, 3 insertions(+), 2 deletions(-) diff --git a/infrastructure-maintenance.prose.md b/infrastructure-maintenance.prose.md index e9a3b1a..e0dbf07 100644 --- a/infrastructure-maintenance.prose.md +++ b/infrastructure-maintenance.prose.md @@ -148,8 +148,9 @@ If ANY service in Phase 4 fails to come back up (or regresses vs baseline): ```bash # Restore from recorded image-baseline digest docker compose down - docker compose pull # or set image: in compose and up -d - docker compose up -d + # Pin the service image to the recorded digest in compose, then recreate + # image: @sha256: + docker compose pull && docker compose up -d ``` 2. **APT rollback** — restore the primary host from the Proxmox snapshot taken/confirmed in Phase 0. If no snapshot, `apt install =` per package using apt history (`/var/log/apt/history.log`). 3. **Config rollback** — restore configs from `/tmp/infra-maintenance-backup-/`. From 99789a00a14eff626983601164977c4bd70ccbdb Mon Sep 17 00:00:00 2001 From: root Date: Sat, 18 Jul 2026 22:16:09 +0000 Subject: [PATCH 3/3] no-mistakes(document): Reframed infra-maintenance contract as deliberate partition --- infrastructure-maintenance.prose.md | 20 ++++++++++++-------- 1 file changed, 12 insertions(+), 8 deletions(-) diff --git a/infrastructure-maintenance.prose.md b/infrastructure-maintenance.prose.md index e0dbf07..96c12bf 100644 --- a/infrastructure-maintenance.prose.md +++ b/infrastructure-maintenance.prose.md @@ -7,10 +7,11 @@ description: > running containers, container restarts with health verification, post-update verification of every critical service (LiteLLM proxy, SearXNG, Zulip, Gitea, PM2 processes, Hermes gateways), and rollback on failure. Consolidates the - raw shell scripts that previously did this piecemeal and closes the Docker - image management gap that infrastructure-update left open (infra-update owns - cluster-wide apt waves; this contract owns the host-level maintenance loop - on the primary host + Docker image lifecycle). Runs Sunday 2am ET. Owner: + raw shell scripts that previously did this piecemeal. This contract owns the + HOST-LEVEL weekly maintenance loop on the primary host plus Docker image + pulls ONLY for .116 and .7, while infrastructure-update owns the FULL-FLEET + cluster-wide wave (apt across the full PVE cluster + CTs/VMs AND its Docker + image Wave 3 across all stacks). Runs Sunday 2am ET. Owner: ops (firstmate secondmate). Blast radius: an unverified image pull can break LiteLLM (all agents lose inference) or SearXNG (search-stack down); a bad apt upgrade can leave the host in a half-upgraded state. Pre-update backup check @@ -34,9 +35,12 @@ version: 1.0.0 ## Scope Primary host is the maintenance host where apt updates apply. Docker image pulls -span the two Docker ecosystems that run critical services. Topology, CT IDs, and -IPs are live-state fields — verify against `infrastructure-control.prose.md` -(the source of truth) and the live system before mutating. +span the two Docker ecosystems that run critical services. infrastructure-update +runs the full-fleet cluster-wide wave (including its Wave 3 Docker pulls across +all stacks/hosts); this contract runs a narrower host-level weekly pull limited +to .116 and .7. Topology, CT IDs, and IPs are live-state fields — verify against +`infrastructure-control.prose.md` (the source of truth) and the live system +before mutating. | Host | IP | Role | Trust | |------|----|------|-------| @@ -189,7 +193,7 @@ Escalation: none | warning | critical | fatal ## Related Contracts -- `infrastructure-update.prose.md` — cluster-wide apt waves across the PVE cluster + CTs/VMs (the build-phase update); this contract handles the host-level maintenance loop and Docker image lifecycle. +- `infrastructure-update.prose.md` — owns the full-fleet cluster-wide wave INCLUDING its Wave 3 Docker image updates across all stacks (SearXNG, Firecrawl, Inference Harness on .116, home stack, audiobookshelf); infrastructure-maintenance is a deliberately narrower host-level weekly pull scoped to .116 and .7. - `infrastructure-monitoring.prose.md` — provides the pre-update health baseline (depends_on). - `infrastructure-control.prose.md` — topology source of truth (CT IDs, IPs, hostnames). - `litellm-health.prose.md` — LiteLLM probe details.