feat(contracts): add infrastructure-maintenance contract
New responsibility contract consolidating host-level system maintenance and Docker image lifecycle management, filling the gap left by infrastructure-update (which owns cluster-wide apt waves). Scope: - OS package updates on primary host with pre-update snapshot/backup check - Docker image pulls for LiteLLM, SearXNG, and other running containers - Container restarts with per-stack health verification - Post-update verification: LiteLLM, SearXNG, Zulip, Gitea, PM2, Hermes gateways - Rollback on failure (image/apt/config restore) with circuit breaker Owner: ops (firstmate secondmate). Trigger: weekly Sunday 2am ET. Escalation: warning/critical->abiba+mumuni, fatal->abiba+mumuni+kwame. circuit_breaker: max_retries 2, window 7200, trip_action escalate_to_fatal. depends_on: infrastructure-monitoring (pre-update health baseline). Registry: - Add infrastructure-maintenance to by_category.maintenance, by_domain.infrastructure, by_owner.ops (new), by_trigger.scheduled, by_sensitivity.high - Add 'ops' to owners list - Move infrastructure-update owner abiba -> ops (in contracts entry + by_owner index) Also adds '## Maintaining this file' section to AGENTS.md per fm-ensure-agents-md.
This commit is contained in:
+104
-2
@@ -22,6 +22,7 @@ owners:
|
||||
- abiba
|
||||
- mumuni
|
||||
- kwame
|
||||
- ops
|
||||
trigger_types:
|
||||
- scheduled
|
||||
- event_driven
|
||||
@@ -58,6 +59,7 @@ index:
|
||||
- memory-audit-maintenance
|
||||
- gpu-fleet
|
||||
- infrastructure-update
|
||||
- infrastructure-maintenance
|
||||
reference:
|
||||
- infrastructure-control
|
||||
- ra-h-os-custodianship-contract
|
||||
@@ -90,6 +92,7 @@ index:
|
||||
- infrastructure-control
|
||||
- infrastructure-monitoring
|
||||
- infrastructure-update
|
||||
- infrastructure-maintenance
|
||||
- pm2-self-heal
|
||||
- disk-gc-threat-response
|
||||
gpu:
|
||||
@@ -130,7 +133,6 @@ index:
|
||||
- build-zulip-plugin
|
||||
- stirling-pdf-agent-access
|
||||
- gpu-fleet
|
||||
- infrastructure-update
|
||||
- infrastructure-control
|
||||
- zulip-adapter-lessons
|
||||
- pi-approval-architecture
|
||||
@@ -144,6 +146,9 @@ index:
|
||||
- mumuni-delegation
|
||||
kwame:
|
||||
- hello-world
|
||||
ops:
|
||||
- infrastructure-maintenance
|
||||
- infrastructure-update
|
||||
by_trigger:
|
||||
scheduled:
|
||||
- hermes-key-enforcement
|
||||
@@ -156,6 +161,7 @@ index:
|
||||
- litellm-health
|
||||
- memory-audit-maintenance
|
||||
- infrastructure-update
|
||||
- infrastructure-maintenance
|
||||
event_driven:
|
||||
- litellm-self-heal
|
||||
- pm2-self-heal
|
||||
@@ -197,6 +203,7 @@ index:
|
||||
- hermes-zulip-plugin
|
||||
- build-zulip-plugin
|
||||
- infrastructure-update
|
||||
- infrastructure-maintenance
|
||||
- ra-h-os-custodianship-contract
|
||||
- mumuni-delegation
|
||||
normal:
|
||||
@@ -1256,13 +1263,108 @@ contracts:
|
||||
last_run: null
|
||||
last_status: null
|
||||
drift_alerts: []
|
||||
- name: infrastructure-maintenance
|
||||
file: infrastructure-maintenance.prose.md
|
||||
kind: responsibility
|
||||
category: maintenance
|
||||
sensitivity: high
|
||||
status: active
|
||||
owner: ops
|
||||
version: 1.0.0
|
||||
trigger:
|
||||
type: scheduled
|
||||
cadence: 0 2 * * 0
|
||||
description: Weekly host-level maintenance Sunday at 2am ET (replaces infrastructure-update
|
||||
build-phase role; infra-update moves to ops)
|
||||
cron_job_id: null
|
||||
execution:
|
||||
agent: ops
|
||||
timeout: 3600
|
||||
requires:
|
||||
- infrastructure-monitoring run within last 30 minutes (pre-update health baseline)
|
||||
- Proxmox snapshot of primary host OR /tmp backup dir created this run
|
||||
- LiteLLM master key from Infisical vault for health verification
|
||||
protocol:
|
||||
- Load contract from prose-contracts/main
|
||||
- Phase 0 preflight — capture health baseline, backup check, record image baseline
|
||||
- Phase 1 apt update && apt upgrade -y on primary host
|
||||
- Phase 2 docker compose pull for LiteLLM, SearXNG, and other running containers
|
||||
- Phase 3 restart stacks one at a time with per-stack health verification
|
||||
- Phase 4 verify every critical service (LiteLLM, SearXNG, Zulip, Gitea, PM2, Hermes gateways)
|
||||
- On failure — rollback per protocol, escalate, do not loop beyond circuit breaker
|
||||
- Log actions to ~/.hermes/runs/infrastructure-maintenance/
|
||||
verification:
|
||||
postconditions:
|
||||
- check: all critical services running after update
|
||||
verify: 'curl -sf http://192.168.68.116/litellm/v1/models && curl -sf https://chat.sysloggh.net/api/v1/server_settings && curl -sf https://git.sysloggh.net/api/v1/version && curl -sf http://192.168.68.7:8888 && pm2 jlist'
|
||||
expect: all probes 200 OK / processes online
|
||||
- check: no regressions from pre-update health baseline
|
||||
verify: diff Phase 0 health-baseline against Phase 4 results
|
||||
expect: no GREEN service turned RED
|
||||
- check: docker containers on latest stable tags
|
||||
verify: docker inspect --format '{{.Config.Image}}' <container> per service matches image-baseline.pulled_tag
|
||||
expect: all containers running pulled tags
|
||||
- check: APT packages up to date with no held broken packages
|
||||
verify: apt list --upgradable 2>/dev/null | wc -l and apt-get -s upgrade | grep -ci broken
|
||||
expect: upgradable == 0, broken == 0
|
||||
artifact: maintenance run report with phase results and any rollback/escalation
|
||||
verify_commands:
|
||||
- curl -sf http://192.168.68.116/litellm/v1/models
|
||||
- curl -sf http://192.168.68.7:8888
|
||||
- curl -sf https://chat.sysloggh.net/api/v1/server_settings
|
||||
- curl -sf https://git.sysloggh.net/api/v1/version
|
||||
- pm2 jlist
|
||||
receipt:
|
||||
format: json
|
||||
storage: ~/.hermes/runs/infrastructure-maintenance/
|
||||
graph_node: true
|
||||
schema:
|
||||
contract: string
|
||||
run_id: string
|
||||
timestamp: ISO 8601
|
||||
agent: string
|
||||
status: pass|fail|escalated
|
||||
phase: preflight|apt|images|restarts|verify|rollback|done|failed
|
||||
actions_taken: array
|
||||
postconditions: array
|
||||
drift_alerts: array
|
||||
evidence_path: string
|
||||
escalation:
|
||||
info:
|
||||
action: log_to_receipt
|
||||
notify: []
|
||||
warning:
|
||||
action: relay_alert
|
||||
notify:
|
||||
- abiba
|
||||
- mumuni
|
||||
critical:
|
||||
action: relay_alert
|
||||
notify:
|
||||
- abiba
|
||||
- mumuni
|
||||
fatal:
|
||||
action: relay_alert + pause + human_required
|
||||
notify:
|
||||
- abiba
|
||||
- mumuni
|
||||
- kwame
|
||||
circuit_breaker:
|
||||
max_retries: 2
|
||||
window: 7200
|
||||
trip_action: escalate_to_fatal
|
||||
depends_on:
|
||||
- infrastructure-monitoring
|
||||
last_run: null
|
||||
last_status: null
|
||||
drift_alerts: []
|
||||
- name: infrastructure-update
|
||||
file: infrastructure-update.prose.md
|
||||
kind: responsibility
|
||||
category: maintenance
|
||||
sensitivity: high
|
||||
status: active
|
||||
owner: abiba
|
||||
owner: ops
|
||||
version: 1.0.0
|
||||
trigger:
|
||||
type: scheduled
|
||||
|
||||
Reference in New Issue
Block a user