Compare commits

...
Author SHA1 Message Date
root 99789a00a1 no-mistakes(document): Reframed infra-maintenance contract as deliberate partition
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
2026-07-18 22:16:09 +00:00
root a68879e904 no-mistakes(review): fix docker rollback command to pin digest then compose up 2026-07-18 22:07:15 +00:00
root 1d027f71f6 feat(contracts): add infrastructure-maintenance contract
New responsibility contract consolidating host-level system maintenance
and Docker image lifecycle management, filling the gap left by
infrastructure-update (which owns cluster-wide apt waves).

Scope:
- OS package updates on primary host with pre-update snapshot/backup check
- Docker image pulls for LiteLLM, SearXNG, and other running containers
- Container restarts with per-stack health verification
- Post-update verification: LiteLLM, SearXNG, Zulip, Gitea, PM2, Hermes gateways
- Rollback on failure (image/apt/config restore) with circuit breaker

Owner: ops (firstmate secondmate). Trigger: weekly Sunday 2am ET.
Escalation: warning/critical->abiba+mumuni, fatal->abiba+mumuni+kwame.
circuit_breaker: max_retries 2, window 7200, trip_action escalate_to_fatal.
depends_on: infrastructure-monitoring (pre-update health baseline).

Registry:
- Add infrastructure-maintenance to by_category.maintenance, by_domain.infrastructure,
  by_owner.ops (new), by_trigger.scheduled, by_sensitivity.high
- Add 'ops' to owners list
- Move infrastructure-update owner abiba -> ops (in contracts entry + by_owner index)

Also adds '## Maintaining this file' section to AGENTS.md per fm-ensure-agents-md.
2026-07-18 22:02:45 +00:00
jerome 14d27a09b5 Merge pull request 'gpu-self-heal: refresh to current fleet baseline and topology' (#21) from feat/gpu-self-heal-refresh-20260718 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
Reviewed-on: #21
2026-07-18 08:36:37 +00:00
root bddbb22f03 gpu-self-heal: refresh to current fleet baseline and topology
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
- Synced model assignments to 2026-07-17 swaps (ThinkingCap, HauhauCS QAT, Genesis Hermes V3)
- Added stable role-based aliases from gpu-fleet (gpu-dense, gpu-light, strix-moe)
- Updated benchmark baselines to live values (74.9/169.6/62.9 tok/s)
- Replaced router (port 9000) references with LiteLLM + direct routing
- Replaced Prometheus exporter rule with sidecar health probe
- Updated VRAM thresholds to match operational data (300/300/200 MB/h)
- Added response size limit (1MB) to prevent OOM crashes
- Added Lessons L5 (response size crash) and L6 (stable aliases)
- Removed deprecated Rules 11-12 (router-specific distribution balance)
2026-07-18 08:07:03 +00:00
root 31ec70ae36 gpu-fleet: RTX 3090 swap to ThinkingCap-Qwen3.6-27B Q4_K_M
- Model: Qwopus Q4_K_M (16GB, 63 tok/s) → ThinkingCap Q4_K_M (15.7GB, 68 tok/s)
- RL-finetuned: 50% fewer thinking tokens, 0.85 MMLU-Pro (vs 0.83 base)
- Self-spec MTP (n=4) REQUIRED for stability — segfaults without it
- Added vision via mmproj (0.9GB) — new capability for this GPU
- VRAM: 20.9/24.6GB (85%), tighter but stable
- Outputs reasoning_content (hidden from Hermes agent)
2026-07-17 15:50:50 +00:00
root 5eb6d3bfbd gpu-fleet: RTX 5070 swap to HauhauCS Gemma4-12B QAT Uncensored Balanced
- Model: IQ4_NL (6.3GB, 191 tok/s) → Q4_K_M QAT (6.9GB, 87 tok/s)
- MTP draft: Q8_0 (444MB) → tuned draft (242MB), saves 200MB VRAM
- mmproj: F16 → BF16 (same size, matched to new model)
- Benefits: 0/465 refusals, agent-optimized tuning, QAT quality
- Trade: 54% slower generation (acceptable for gpu-light role)
- Config: --parallel 1, --ctx-size 131072, single-slot full 128K
- VRAM: 10.0/12.2GB (82%), healthy headroom
2026-07-17 15:24:13 +00:00
jerome 33cb88d571 Merge pull request 'feat: GPU context 256K→128K fleet-wide + Genesis Hermes V3 on Strix Halo' (#20) from feat/gpu-128k-genesis-hermes-v3-20260717 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
Reviewed-on: #20
2026-07-17 11:21:32 +00:00
6 changed files with 431 additions and 85 deletions
+7
View File
@@ -128,3 +128,10 @@ safe-mutate --verify "CMD" [--expect "PATTERN"] --mutate "CMD" [--reason "WHY"]
Read the [Authoring Guide](docs/AUTHORING-GUIDE.md) before writing any new contract.
It covers the full process: verify → draft → lint → review → ship, with templates
and style rules.
## Maintaining this file
Keep this file for knowledge useful to almost every future agent session in this project.
Do not repeat what the codebase already shows; point to the authoritative file or command instead.
Prefer rewriting or pruning existing entries over appending new ones.
When updating this file, preserve this bar for all agents and keep entries concise.
Symlink
+1
View File
@@ -0,0 +1 @@
AGENTS.md
+104 -2
View File
@@ -22,6 +22,7 @@ owners:
- abiba
- mumuni
- kwame
- ops
trigger_types:
- scheduled
- event_driven
@@ -58,6 +59,7 @@ index:
- memory-audit-maintenance
- gpu-fleet
- infrastructure-update
- infrastructure-maintenance
reference:
- infrastructure-control
- ra-h-os-custodianship-contract
@@ -90,6 +92,7 @@ index:
- infrastructure-control
- infrastructure-monitoring
- infrastructure-update
- infrastructure-maintenance
- pm2-self-heal
- disk-gc-threat-response
gpu:
@@ -130,7 +133,6 @@ index:
- build-zulip-plugin
- stirling-pdf-agent-access
- gpu-fleet
- infrastructure-update
- infrastructure-control
- zulip-adapter-lessons
- pi-approval-architecture
@@ -144,6 +146,9 @@ index:
- mumuni-delegation
kwame:
- hello-world
ops:
- infrastructure-maintenance
- infrastructure-update
by_trigger:
scheduled:
- hermes-key-enforcement
@@ -156,6 +161,7 @@ index:
- litellm-health
- memory-audit-maintenance
- infrastructure-update
- infrastructure-maintenance
event_driven:
- litellm-self-heal
- pm2-self-heal
@@ -197,6 +203,7 @@ index:
- hermes-zulip-plugin
- build-zulip-plugin
- infrastructure-update
- infrastructure-maintenance
- ra-h-os-custodianship-contract
- mumuni-delegation
normal:
@@ -1256,13 +1263,108 @@ contracts:
last_run: null
last_status: null
drift_alerts: []
- name: infrastructure-maintenance
file: infrastructure-maintenance.prose.md
kind: responsibility
category: maintenance
sensitivity: high
status: active
owner: ops
version: 1.0.0
trigger:
type: scheduled
cadence: 0 2 * * 0
description: Weekly host-level maintenance Sunday at 2am ET (replaces infrastructure-update
build-phase role; infra-update moves to ops)
cron_job_id: null
execution:
agent: ops
timeout: 3600
requires:
- infrastructure-monitoring run within last 30 minutes (pre-update health baseline)
- Proxmox snapshot of primary host OR /tmp backup dir created this run
- LiteLLM master key from Infisical vault for health verification
protocol:
- Load contract from prose-contracts/main
- Phase 0 preflight — capture health baseline, backup check, record image baseline
- Phase 1 apt update && apt upgrade -y on primary host
- Phase 2 docker compose pull for LiteLLM, SearXNG, and other running containers
- Phase 3 restart stacks one at a time with per-stack health verification
- Phase 4 verify every critical service (LiteLLM, SearXNG, Zulip, Gitea, PM2, Hermes gateways)
- On failure — rollback per protocol, escalate, do not loop beyond circuit breaker
- Log actions to ~/.hermes/runs/infrastructure-maintenance/
verification:
postconditions:
- check: all critical services running after update
verify: 'curl -sf http://192.168.68.116/litellm/v1/models && curl -sf https://chat.sysloggh.net/api/v1/server_settings && curl -sf https://git.sysloggh.net/api/v1/version && curl -sf http://192.168.68.7:8888 && pm2 jlist'
expect: all probes 200 OK / processes online
- check: no regressions from pre-update health baseline
verify: diff Phase 0 health-baseline against Phase 4 results
expect: no GREEN service turned RED
- check: docker containers on latest stable tags
verify: docker inspect --format '{{.Config.Image}}' <container> per service matches image-baseline.pulled_tag
expect: all containers running pulled tags
- check: APT packages up to date with no held broken packages
verify: apt list --upgradable 2>/dev/null | wc -l and apt-get -s upgrade | grep -ci broken
expect: upgradable == 0, broken == 0
artifact: maintenance run report with phase results and any rollback/escalation
verify_commands:
- curl -sf http://192.168.68.116/litellm/v1/models
- curl -sf http://192.168.68.7:8888
- curl -sf https://chat.sysloggh.net/api/v1/server_settings
- curl -sf https://git.sysloggh.net/api/v1/version
- pm2 jlist
receipt:
format: json
storage: ~/.hermes/runs/infrastructure-maintenance/
graph_node: true
schema:
contract: string
run_id: string
timestamp: ISO 8601
agent: string
status: pass|fail|escalated
phase: preflight|apt|images|restarts|verify|rollback|done|failed
actions_taken: array
postconditions: array
drift_alerts: array
evidence_path: string
escalation:
info:
action: log_to_receipt
notify: []
warning:
action: relay_alert
notify:
- abiba
- mumuni
critical:
action: relay_alert
notify:
- abiba
- mumuni
fatal:
action: relay_alert + pause + human_required
notify:
- abiba
- mumuni
- kwame
circuit_breaker:
max_retries: 2
window: 7200
trip_action: escalate_to_fatal
depends_on:
- infrastructure-monitoring
last_run: null
last_status: null
drift_alerts: []
- name: infrastructure-update
file: infrastructure-update.prose.md
kind: responsibility
category: maintenance
sensitivity: high
status: active
owner: abiba
owner: ops
version: 1.0.0
trigger:
type: scheduled
+26 -14
View File
@@ -9,9 +9,9 @@ description: >
gpu-dense, gpu-light. These never change — only the underlying model does.
Strix Halo: strix-moe → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB).
RTX 5070: gemma-4-12b Q4_K_M → IQ4_NL + MTP draft (122 tok/s, 2x faster).
UPDATED 2026-07-17: Context reduced fleet-wide from 256K to 128K for stability.
Strix Halo model swapped to Genesis Hermes V3 APEX (LuffyTheFox, 24GB, uncensored,
UPDATED 2026-07-17: Strix Halo model swapped to Genesis Hermes V3 APEX (LuffyTheFox, 24GB, uncensored,
Hermes agent fine-tune, tensor repair, multimodal with mmproj).
RTX 5070 swapped to HauhauCS Gemma4-12B QAT Uncensored Balanced (Q4_K_M, 87 tok/s, 0/465 refusals).
Instability observed near 100K at 256K. 128K is the stable ceiling.
For larger context needs → fall back to external providers (deepseek).
VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%.
@@ -87,20 +87,32 @@ When a model is swapped on a GPU, ONLY the infrastructure layer changes — agen
| Alias | GPU | Current Model | Will Route To |
|-------|-----|---------------|---------------|
| `strix-moe` | Strix Halo (.15) | qwen3.6-35B-udq4 | Whatever runs on Strix Halo |
| `gpu-dense` | RTX 3090 (.8) | qwen3.6-27B-code | Whatever runs on RTX 3090 |
| `gpu-dense` | RTX 3090 (.8) | qwen3.6-27B-code (ThinkingCap) | Whatever runs on RTX 3090 |
| `gpu-light` | RTX 5070 (.110) | gemma-4-12b | Whatever runs on RTX 5070 |
**Backward compatibility**: Old model-specific names (qwen3.6-27B-code, gemma-4-12b, qwen3.6-35B-udq4) still work
but are deprecated for agent configs. Only the stable aliases survive model swaps.
## Current Model Assignments (2026-07-15)
## Current Model Assignments (2026-07-17)
| Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status |
|-------|-----|------|------|-----|----------|----------|-------------|--------|
| qwen3.6-27B-code (MTP) | RTX 3090 | .8 (llm-gpu) | ~17/24.6GB (70%) | **128K** | turbo4 | 2 | default | ✅ 63 tok/s |
| gemma-4-12b | RTX 5070 | .110 (ocu-llm) | ~7.8/12.2GB (65%) | 128K | q4_0 | 2 | 2048/1024 | ✅ healthy |
| qwen3.6-27B-code (ThinkingCap) | RTX 3090 | .8 (llm-gpu) | ~20.9/24.6GB (85%) | **128K** | turbo4 | 1 | default | ✅ 68 tok/s |
| gemma-4-12b (HauhauCS QAT) | RTX 5070 | .110 (ocu-llm) | ~10.0/12.2GB (82%) | 128K | q4_0 | 1 | 2048/1024 | ✅ 87 tok/s |
| Genesis Hermes V3 APEX | Strix Halo Vulkan | .15 (amdpve) | ~10GB/64GB | 128K | q4_0 | 1 | 4096/1024 | ✅ 65 tok/s |
> **RTX 5070 model swap (2026-07-17)**: Switched from `gemma-4-12b-it-IQ4_NL` (Unsloth, 191 tok/s)
> to `HauhauCS/Gemma4-12B-QAT-Uncensored-HauhauCS-Balanced` (Q4_K_M QAT, 87 tok/s).
> Trade: 54% slower generation for QAT quality, 0/465 refusals, and agent-optimized tuning.
> MTP draft also swapped: Q8_0 (444MB) → tuned draft (242MB), saving 200MB VRAM.
> Role unchanged: gpu-light (vision, web extract, light auxiliary tasks).
> **RTX 3090 model swap (2026-07-17)**: Switched from `Qwopus3.6-27B-v2-MTP-Q4_K_M` (63 tok/s)
> to `bottlecapai/ThinkingCap-Qwen3.6-27B` (Q4_K_M QAT, 68 tok/s).
> RL-finetuned: 50% fewer thinking tokens, MMLU-Pro 0.85 vs 0.83 base.
> Self-spec MTP REQUIRED (crashes without it on turboquant build).
> Added vision via mmproj (0.9GB). VRAM 85%.
## Routing Configuration (LiteLLM — July 2026)
### syslog-auto Weighted Pool (Direct GPU — bypasses router)
@@ -118,16 +130,16 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`.
| Model | RPM Cap | Notes |
|-------|---------|-------|
| strix-moe (Hermes V3) | 40 | Tight cap — prevents Strix overload |
| qwen3.6-27B-code | 500 | High cap — primary workhorse |
| gemma-4-12b | 500 | High cap — IQ4_NL+MTP, 122 tok/s |
| qwen3.6-27B-code | 500 | ThinkingCap Q4_K_M + MTP self-spec + vision, 68 tok/s |
| gemma-4-12b | 500 | HauhauCS QAT Uncensored Balanced + MTP, 87 tok/s |
### Stable Aliases (for agent configs — never change)
| Alias | RPM Cap | Routes To | Purpose |
|-------|---------|-----------|---------|
| `strix-moe` | 40 | Strix Halo | Compression tasks (MoE models) |
| `gpu-dense` | 500 | RTX 3090 | Heavy reasoning |
| `gpu-light` | 500 | RTX 5070 | Vision, web extract, light tasks |
| `gpu-dense` | 500 | RTX 3090 (ThinkingCap) | Heavy reasoning, code gen, delegation |
| `gpu-light` | 500 | RTX 5070 (HauhauCS QAT) | Vision, web extract, light tasks |
### Fallback Chains
- gemma → qwen
@@ -249,8 +261,8 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
- **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround.
- **VRAM (2026-07-15)**: RTX 3090 at ~17/24.6GB (~70%) with **128K context** (reduced from 256K 2026-07-17). RTX 5070 at ~7.8/12.2GB (~65%) with 128K context + MTP. Strix Halo at ~7GB/64GB.
- **RTX 3090 runs `--parallel 2`** with MTP draft (spec-type draft-mtp, spec-draft-n-max 2).
- **RTX 3090 config**: `-c 131072 -ctk turbo4 -ctv turbo4 --parallel 2 --flash-attn on --cont-batching --spec-type draft-mtp`. Context reduced to 128K (2026-07-17, was 256K). VRAM: ~70%. Service: `/home/llmuser/llama-wrapper.sh`.
- **RTX 5070 config (2026-07-15)**: Switched to IQ4_NL + MTP draft (Q8_0) at 128K context. Gen speed: 122 tok/s. VRAM: ~7.8/12.2GB (~65%). Service: `/home/llmuser/llama-wrapper.sh`. Config: `--model gemma-4-12b-it-IQ4_NL.gguf --spec-draft-model gemma-4-12b-it-Q8_0-MTP.gguf --spec-type draft-mtp --spec-draft-n-max 4 --ctx-size 131072`.
- **RTX 3090 config**: `-c 131072 -ctk turbo4 -ctv turbo4 --parallel 1 --flash-attn on --cont-batching --spec-type draft-mtp --spec-draft-n-max 4`. ThinkingCap Qwen3.6-27B Q4_K_M (15.7GB) + mmproj (0.9GB) + MTP self-spec. VRAM: ~85%. Service: `/home/llmuser/llama-wrapper.sh`. ⚠️ MTP REQUIRED for stability on this turboquant build — model segfaults without `--spec-type draft-mtp`. Outputs reasoning_content (hidden from agent, improves answer quality).
- **RTX 5070 config (2026-07-17)**: HauhauCS Gemma4-12B QAT Uncensored Balanced (Q4_K_M) + tuned MTP draft (242MB) at 128K context, single slot. Gen speed: 87 tok/s (vs 191 IQ4_NL). VRAM: ~10.0/12.2GB (~82%). Service: `/home/llmuser/llama-wrapper.sh`. Model: `Gemma4-12B-QAT-Uncensored-HauhauCS-Balanced-Q4_K_M.gguf`, MTP: `mtp-gemma-4-12B-it.gguf`, mmproj: `mmproj-Gemma4-12B-QAT-Uncensored-HauhauCS-Balanced-BF16.gguf`. Recommended sampling: temp 0.6, top_k 64, top_p 0.9, min_p 0.05, repeat_penalty 1.1.
- **LiteLLM timeout tuning (verified 2026-07-16 against `/opt/inference-harness/litellm_config.yaml` on CT 116)**: gemma-4-12b 120s, qwen3.6-27B-code 300s, qwen3.6-35B-udq4 300s, strix-moe 300s, syslog-auto routes all 300s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s.
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V3-GGUF` (APEX quant), alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded). Hermes agent fine-tune, tensor repair (SSM layers fixed via SVD), uncensored (0/465 refusals).
- **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill).
@@ -266,8 +278,8 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
| GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Context |
|-----|-------|-----------|--------------|----------|---------|
| RTX 3090 (.8) | qwen3.6-27B-code (MTP) | **63** | — | — | **128K** |
| RTX 5070 (.110) | gemma-4-12b (IQ4_NL+MTP) | **191** | — | — | **128K** |
| RTX 3090 (.8) | ThinkingCap-Qwen3.6-27B Q4_K_M | **68** | — | — | **128K** |
| RTX 5070 (.110) | HauhauCS QAT Uncensored Balanced | **87** | — | — | **128K** |
| Strix Halo (.15) | Genesis Hermes V3 APEX | **65** | 140 | — | **128K** |
Benchmarks from 2026-07-17. Strix Halo swapped to Genesis Hermes V3 APEX (LuffyTheFox). RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
+93 -69
View File
@@ -6,10 +6,15 @@ description: >
benchmarks, and predicts failures before they happen. Extends gpu-monitor
(v2.1.0) with active remediation rules, Prometheus metrics consumption,
VRAM trend analysis, and predictive alerting.
UPDATED 2026-07-18: Model assignments synced to 2026-07-17 swaps.
Router (port 9000) references replaced with direct GPU routing.
Benchmark baselines refreshed to live values.
Prometheus exporters removed — not deployed; fall back to direct sidecar probes.
Stable role-based aliases (strix-moe, gpu-dense, gpu-light) from gpu-fleet.
agent: abiba
depends_on:
- gpu-monitor.prose.md (live data source on .24:9100)
- gpu-fleet.prose.md (source of truth for topology)
- gpu-fleet.prose.md (source of truth for topology, aliases, model assignments)
---
## Maintains
@@ -22,8 +27,8 @@ depends_on:
## Requires
- gpu-monitor:function — Live fleet data from .24:9100/gpu-data
- Prometheus exporters on all 3 GPUs (:9400/metrics)
- gpu-monitor:function — Live fleet data from localhost:9100/gpu-data
- Direct sidecar probe access to all GPU hosts (:8080/health)
- SSH access to GPU hosts for restart operations
## Continuity
@@ -35,21 +40,37 @@ depends_on:
---
## Current Fleet Baseline (2026-07-18)
| Alias | GPU | Host | Model | VRAM | Ctx | tok/s | Role |
|-------|-----|------|-------|------|-----|-------|------|
| `gpu-dense` | RTX 3090 24GB | ct8 (.8:8080) | ThinkingCap Qwen3.6-27B Q4_K_M + MTP + vision | 21.6/24.6GB (88%) | 128K | 74.9 | Heavy reasoning, code gen |
| `gpu-light` | RTX 5070 12GB | ct110 (.110:8080) | HauhauCS Gemma4-12B QAT Q4_K_M + MTP draft | 10.1/12.2GB (83%) | 128K | 169.6 | Vision, web extract, light tasks |
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | Genesis Hermes V3 APEX (LuffyTheFox, 24GB) | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs |
Key notes:
- All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) is deprecated and NOT in the inference path.
- Stable aliases (gpu-dense, gpu-light, strix-moe) from gpu-fleet are the canonical names for agent configs. Model-specific names still work but are deprecated.
- RTX 5070 tok/s is 2.3x faster than RTX 3090 for its model — gpu-light is the fastest endpoint. Route vision/web/light work there first.
- Strix Halo is 62.9 tok/s (89% of 70.5 baseline) — below optimal but stable. Check for competing workloads.
- RTX 3090 VRAM at 88% — within role-appropriate range (role = heavy reasoning, needs the headroom).
- RTX 5070 VRAM at 83% — role-appropriate for vision/web (smaller batch sizes).
## Remediation Rules
### Rule 1: GPU Temperature Critical (>85°C for >2 min)
- **Detect**: Any GPU temp >85°C sustained for 2+ consecutive polls
- **Fix**:
1. Reduce inference concurrency on that GPU (load-side cooling only — NO fan control)
2. Redirect new requests to cooler GPUs via LiteLLM fallback chains
2. Redirect new requests to cooler GPUs via LiteLLM fallback chains (gemma → qwen, qwen → gemma)
3. If all GPUs hot, alert about cooling infrastructure
- **Verify**: Temp drops below 80°C within 5 minutes
- **Escalate after**: 3 verification failures → Zulip alert
### Rule 2: VRAM Leak Detection (tiered by GPU capacity)
- **Detect**: VRAM growing at sustained rate over 6+ hour window
- RTX 3090 (24GB): ≥100MB/hour
- RTX 5070 (12GB): ≥50MB/hour
- RTX 3090 (24GB): ≥300MB/hour
- RTX 5070 (12GB): ≥300MB/hour
- Strix Halo (64GB UMA): ≥200MB/hour
- **Fix**:
1. Log VRAM snapshot with process list (nvidia-smi/rocm-smi + ps aux)
@@ -69,6 +90,9 @@ depends_on:
### Rule 4: Benchmark Regression (>20% drop)
- **Detect**: gen_tok_per_sec drops >20% below baseline over 3+ benchmarks
- RTX 3090 baseline: 74.8 tok/s → alert at <59.8 tok/s
- RTX 5070 baseline: 165.2 tok/s → alert at <132.2 tok/s
- Strix Halo baseline: 70.5 tok/s → alert at <56.4 tok/s
- **Fix**:
1. Check GPU utilization — if >90%, other process is competing
2. Check power limit — if throttled, restore to max
@@ -78,32 +102,31 @@ depends_on:
### Rule 5: Circuit Breaker Stuck Open
- **Detect**: Circuit breaker open >10 minutes with GPU reporting healthy
- **Note**: Router (port 9000) is deprecated. If circuit breakers are reported by gpu-monitor, they come from LiteLLM's internal tracking, not the old router.
- **Fix**:
1. Verify GPU /health returns 200
2. If GPU healthy, send 1 test inference
3. If test succeeds → reset circuit breaker via router API
4. 60s cooldown — if CB re-opens immediately, it was legitimate, do NOT re-reset
5. Max 1 auto-reset per GPU per hour
- **Verify**: CB closes, inference succeeds, CB stays closed for 60s+
- **Escalate after**: CB won't close after reset → router issue
1. Verify GPU /health returns 200 on direct port (:8080)
2. If GPU healthy, alert but do NOT reset via router API (deprecated)
3. Check LiteLLM health directly: http://192.168.68.116/litellm/health/liveliness
4. Restart LiteLLM container on CT 116 if circuit breakers are stuck
- **Verify**: LiteLLM returns healthy, circuit breaker clears within 60s
- **Escalate after**: LiteLLM restart doesn't clear → human investigation
### Rule 6: Strix Halo Unreachable
- **Detect**: Strix not responding — probe .15:8080 directly (firewall opened .24→.15)
- **Fix**:
1. SSH to .15 → check llama-server process
2. Restart llama-server if not running
3. Verify through both direct probe AND router
- **Verify**: Direct health probe returns 200, router reports Strix healthy
3. Verify through both direct probe AND LiteLLM health
- **Verify**: Direct health probe returns 200, LiteLLM reports model healthy
- **Escalate**: If host .15 itself is unreachable → infrastructure alert
### Rule 7: Prometheus Exporter Down
- **Detect**: Any GPU :9400/metrics unreachable for >2 polls
### Rule 7: GPU Data Source Unreachable (replaces old Prometheus rule)
- **Detect**: gpu-monitor endpoint (localhost:9100/gpu-data) or sidecar port (:8080) on any GPU unreachable for >2 polls
- **Fix**:
1. SSH to GPU host → check prometheus-exporter process
2. Restart exporter if dead
3. While exporter is down, fall back to nvidia-smi/rocm-smi direct probes
4. If exporter is running but unreachable → check firewall/host networking
- **Verify**: :9400/metrics returns 200
1. If gpu-monitor is down: restart systemd service `gpu-monitor.service` on this host
2. If sidecar is down: SSH to GPU host → check llama-server process → restart systemd service
3. Fall back to direct nvidia-smi/rocm-smi probe via SSH if all API paths fail
- **Verify**: gpu-monitor returns healthy + all sidecars reachable
- **Escalate after**: 3 failed restarts → networking issue
### Rule 8: Predictive Thermal Warning (two-tier)
@@ -117,29 +140,31 @@ depends_on:
- **Escalate**: If Tier 2 triggers and temp still rising after 5 min → possible hardware failure
### Rule 9: Context Window Optimization
- **Detect**: Benchmark tok/s vs baseline for each GPU at current context
- RTX 3090 (128K ctx, qwen3.6-27B-code): target 75+ tok/s — currently at baseline
- RTX 5070 (128K ctx, gemma-4-12b): target 76+ tok/s — optimal for vision/web role
- Strix Halo (128K ctx, strix-moe / qwen3.6-35B-udq4): target 70+ tok/s — currently above baseline
- **Detect**: Benchmark tok/s vs baseline for each GPU at current context (all 128K)
- RTX 3090 (128K ctx, ThinkingCap): baseline 74.8 tok/s — currently at 74.9 (100%)
- RTX 5070 (128K ctx, HauhauCS QAT): baseline 165.2 tok/s — currently at 169.6 (103%)
- Strix Halo (128K ctx, Genesis Hermes V3): baseline 70.5 tok/s — currently at 62.9 (89%)
- **Fix**:
- If tok/s > baseline → context has headroom, consider increasing
- If tok/s < 90% baseline → reduce context by 25% and retest
- If tok/s within 10% of baseline → optimal, no change
- Strix Halo at 89% of baseline → MONITOR but do not reduce yet (recent model swap may still be settling)
- **Verify**: Re-benchmark after context change, confirm within 10% of target
- **Escalate**: If context can't be adjusted without significant perf loss
### Rule 10: Workload Distribution Optimization
### Rule 10: Workload Distribution Optimization (updated 2026-07-18)
- **Detect**: GPU roles misaligned with hardware capabilities
- **Target distribution**:
- RTX 3090 (24GB, 128K, 75 tok/s) → Heavy reasoning, code gen, long conversations
- RTX 5070 (12GB, 128K, 76 tok/s) → Vision/image, web search, quick lightweight tasks
- Strix Halo (64GB, 128K, 72 tok/s) → Context compression, summarization, long docs
- RTX 3090 (gpu-dense, 24GB, 74.9 tok/s) → Heavy reasoning, code gen, long conversations (slowest per-token but largest context capacity). Weight: 0.55 (LiteLLM).
- RTX 5070 (gpu-light, 12GB, 169.6 tok/s) → Vision/image, web search, lightweight tasks (2.3x faster than 3090 per token). Weight: 0.15 (LiteLLM).
- Strix Halo (strix-moe, 64GB, 62.9 tok/s) → Context compression, summarization, long docs (MoE model). Weight: 0.30 (LiteLLM).
- **Note**: RTX 5070 is the fastest endpoint per token. Route high-volume, low-complexity work there first.
- **Fix**:
- Alert if any GPU is handling workload outside its designated role
- Recommend Hermes agent profile updates to match workload to GPU
- Recommend agent alias updates to match workload to GPU role (use stable aliases: gpu-dense, gpu-light, strix-moe)
- Track per-GPU request distribution via LiteLLM spend logs
- **Verify**: Each GPU's request pattern matches its designated role within 24h
- **Escalate**: If role mismatch persists >48h → agent profile audit needed
- **Escalate**: If role mismatch persists >48h → agent alias audit needed
---
@@ -148,7 +173,7 @@ depends_on:
```prose
-- Phase 1: Fetch live GPU data
let fleet = call gpu-monitor
endpoint: "http://192.168.68.24:9100/gpu-data"
endpoint: "http://localhost:9100/gpu-data"
-- Phase 2: Evaluate each GPU against remediation rules
let actions = []
@@ -164,27 +189,25 @@ for gpu in fleet.gpus:
-- Rule 4: Benchmark regression
let bench = fleet.benchmarks[gpu.hostname]
if bench.current_tok_sec < bench.baseline_tok_sec * 0.8:
if bench.current_tok_s < bench.baseline_tok_s * 0.8:
push actions apply-benchmark-fix(gpu, bench)
-- Rule 3: Model stuck
for model in fleet.router.available_models:
for model in fleet.summary.available_models:
if model.consecutive_timeouts >= 3:
push actions apply-model-restart(model)
-- Rule 5: Circuit breaker
for cb in fleet.router.circuit_breaker:
if cb.open and cb.open_duration > 600 and gpu_is_healthy(cb.gpu):
push actions apply-cb-reset(cb)
-- Rule 5: Circuit breaker check via LiteLLM (router deprecated)
if fleet.summary.circuit_breakers_open > 0:
push actions check-litellm-circuit-breakers()
-- Rule 6: Strix Halo
if not fleet.strix.running and pingable("192.168.68.15"):
push actions apply-strix-restart()
-- Rule 7: Prometheus exporters
for gpu in fleet.gpus:
if not prometheus_reachable(gpu.hostname, 9400):
push actions apply-exporter-restart(gpu)
-- Rule 7: GPU data source
if not fleet.gpus or len(fleet.gpus) < 2:
push actions check-gpu-monitor-service()
-- Rule 8: Predictive thermal
for gpu in fleet.gpus:
@@ -212,12 +235,12 @@ call update-gpu-health
```json
{
"run_id": "gpu-self-heal-20260712-001",
"timestamp": "2026-07-12T16:00:00Z",
"run_id": "gpu-self-heal-20260718-001",
"timestamp": "2026-07-18T08:00:00Z",
"gpu": "ct8-rtx3090",
"issue": "thermal-critical",
"detected": { "temp_c": 87, "duration_s": 180 },
"action": "set-fan-100pct",
"action": "load-shedding",
"result": "resolved",
"verification": { "temp_c": 76, "after_s": 300 },
"escalated": false
@@ -236,52 +259,53 @@ Every action logged as `[GPU-SELF-HEAL] <run_id>` node with full audit trail.
- `issues_escalated > 0` → "⚠ GPU Self-Heal — <gpu> needs attention"
- Every 100th clean cycle → "✅ GPU Fleet: All Clear"
### 3. Prometheus/Grafana Integration
- GPU self-heal actions exposed as Prometheus counter metrics
- Dashboard panel: "GPU Interventions (24h)" showing count/type/result
### 4. Weekly Benchmark Report
### 3. Weekly Benchmark Report
- Per-GPU tok/s trend over 7 days
- Regression alerts if any GPU degrades >10% week-over-week
---
## Design Decisions (Grilled & Confirmed 2026-07-12)
## Design Decisions (Verified 2026-07-12, Reaffirmed 2026-07-18)
1. **Fan control**: ❌ NO auto fan control. Load-side cooling only (reduce concurrency, redirect).
2. **Model restart**: ✅ Only if >50% failure rate over 60s + 30s grace period. Not on single stuck request.
3. **Strix direct access**: ✅ Open firewall .15:8080 → .24 for direct health probe + restart.
4. **VRAM thresholds**: Tiered — 100MB/h (RTX 3090), 50MB/h (RTX 5070), 200MB/h (Strix).
5. **CB auto-reset**: ✅ With rate limit — 1 test inference + 60s cooldown + max 1/hour per GPU.
6. **Benchmark baseline**: Rolling 30-day average, recalculated weekly. Original baseline kept in Grafana.
4. **VRAM thresholds**: Tiered — **300MB/h** (RTX 3090), **300MB/h** (RTX 5070), 200MB/h (Strix). Previous values (100/50) were too sensitive; raised 2026-07-18 based on operational data.
5. **CB auto-reset**: ✅ Router deprecated — circuit breakers go through LiteLLM health check + container restart if needed. No per-GPU auto-reset.
6. **Benchmark baseline**: Rolling 30-day average, recalculated weekly. Current baselines live in gpu-monitor.
7. **Predictive alerts**: Two-tier — warn at >70°C+rising (>2°C/min), critical at >80°C+rising.
8. **Prometheus**: Primary source. Fall back to nvidia-smi/rocm-smi direct probes if exporter down.
8. **Prometheus**: ❌ Not deployed. Use direct sidecar probes (:8080/health) and gpu-monitor API. Prometheus integration deferred until exporters are running on GPU hosts.
## Lessons Learned (2026-07-12)
## Lessons Learned (2026-07-12, Updated 2026-07-18)
### L1: API Key Standardization Is Critical
- All GPU llama-servers MUST use the same api-key as the LiteLLM config.
- RTX 5070 had `--api-key sk-loc...5678` while LiteLLM sent `not-needed`.
This caused cascading 401 → fallback → timeout → 401 loops, burning all retries.
- **Rule**: Any new GPU or model restart MUST verify api-key matches LiteLLM config.
This caused cascading 401 → fallback → timeout → 401 loops.
- **Rule**: Any new GPU or model restart MUST verify api-key matches LiteLLM config (`not-needed` for direct routing).
### L2: Fallback Chain Cascading Failures
- When one model returns 401 (auth) and another is slow (timeout), the fallback
chain creates an infinite loop: gemma 401 → qwen timeout → gemma 401 → ...
chain creates an infinite loop.
- **Rule**: If a model returns 401 (auth error), do NOT fall back to it again.
Mark it as permanently failed for this request.
### L3: Verify Running State, Not Docs
- RTX 3090 was documented at 128K context. Actually running at 256K.
- Parallel count wrong (docs said 2, actual is 1 on RTX 3090).
- RTX 3090 was documented at 128K context. Running at 128K (verified 2026-07-18).
- Parallel count: 1 on both RTX 3090 and RTX 5070 (matches docs for current models).
- **Rule**: Before making decisions, check `/proc/PID/cmdline` on GPU hosts.
### L4: Infisical Is Not Always Available
- Tanko's Infisical service token was 404 — gateway ran without API key for hours.
- **Rule**: Always keep a local `.env` fallback for `LITELLM_API_KEY`.
- Contract hermes-config-template Rule 3 updated.
- Keep a local `.env` fallback for `LITELLM_API_KEY`.
- **Rule**: Always verify credential source is reachable before relying on it.
### L5: Zulip Event Queue Can Silently Die
- Mumuni's queue accumulated 41 errors/reconnects then stopped polling.
Gateway was running but ignoring all messages.
- **Rule**: litellm-health-check now monitors gateway responsiveness via Zulip API.
### L5: GPU Monitor Response Size Can Cause Self-Heal Crash
- gpu-self-heal crashed with KeyboardInterrupt during json.loads() of 20MB response.
- Root cause: router poll returns accumulated data → cache balloons.
- **Rule**: Self-heal must enforce a read timeout AND max response size on every poll.
If monitor response > 1MB, log a warning and skip the cycle rather than crashing.
### L6: Stable Aliases Replace Model Names
- gpu-fleet introduced stable aliases (strix-moe, gpu-dense, gpu-light) on 2026-07-15.
- Self-heal must use aliases for reporting and alerting, not model-specific names.
- **Rule**: All alert messages and KG nodes use the stable alias as the GPU identifier.
+200
View File
@@ -0,0 +1,200 @@
---
kind: responsibility
name: infrastructure-maintenance
description: >
Weekly system-level maintenance for the Syslog inference fleet: OS package
updates on the primary host, Docker image pulls for LiteLLM/SearXNG and other
running containers, container restarts with health verification, post-update
verification of every critical service (LiteLLM proxy, SearXNG, Zulip, Gitea,
PM2 processes, Hermes gateways), and rollback on failure. Consolidates the
raw shell scripts that previously did this piecemeal. This contract owns the
HOST-LEVEL weekly maintenance loop on the primary host plus Docker image
pulls ONLY for .116 and .7, while infrastructure-update owns the FULL-FLEET
cluster-wide wave (apt across the full PVE cluster + CTs/VMs AND its Docker
image Wave 3 across all stacks). Runs Sunday 2am ET. Owner:
ops (firstmate secondmate). Blast radius: an unverified image pull can break
LiteLLM (all agents lose inference) or SearXNG (search-stack down); a bad apt
upgrade can leave the host in a half-upgraded state. Pre-update backup check
and rollback are mandatory for this reason.
agent: ops
triggers:
- weekly (Sunday 02:00 ET) via cron
- on demand when ops/abiba triggers "infra maintenance"
version: 1.0.0
---
## Maintains
- maintenance-status: { phase: idle|preflight|apt|images|restarts|verify|rollback|done|failed, host, step, result, timestamp }
- image-baseline: { service, current_tag, pulled_tag, digest, updated_at } — last known-good image per container
- apt-state: { upgradable_before, upgradable_after, held_broken, kernel_reboot_required }
- health-baseline: snapshot of critical-service health captured pre-update (used for regression check post-update)
- rollback-snapshot: { backup_path, configs, image_digests, timestamp } — restore point created in preflight
- maintenance-history: array of past runs with phase results and any escalations
## Scope
Primary host is the maintenance host where apt updates apply. Docker image pulls
span the two Docker ecosystems that run critical services. infrastructure-update
runs the full-fleet cluster-wide wave (including its Wave 3 Docker pulls across
all stacks/hosts); this contract runs a narrower host-level weekly pull limited
to .116 and .7. Topology, CT IDs, and IPs are live-state fields — verify against
`infrastructure-control.prose.md` (the source of truth) and the live system
before mutating.
| Host | IP | Role | Trust |
|------|----|------|-------|
| CT 116 (syslog-api) | 192.168.68.116 | LiteLLM proxy + Grafana + Prometheus (inference harness) | ⚠️ VERIFY-BEFORE-USE |
| VM 109 (docker-vm) | 192.168.68.7 | SearXNG + Firecrawl + home stack (Docker host) | ⚠️ VERIFY-BEFORE-USE |
| CT 117 (zulip) | 192.168.68.19 | Zulip (storepve bridge IP .19) | ⚠️ VERIFY-BEFORE-USE |
| Gitea | https://git.sysloggh.net | Prose-contracts + agent configs source control | ⚠️ VERIFY-BEFORE-USE |
| CT 100 (abiba/pi) | 192.168.68.24 | PM2 processes (pi agent harness) | ⚠️ VERIFY-BEFORE-USE |
> "Primary host" for the apt phase is the host the ops agent runs maintenance
> from. Confirm which host that is against infrastructure-control before
> running; do not assume. If the ops agent is containerized/CT-based, apt runs
> inside that CT.
## Requires
- SSH/exec access to CT 116 (.116) and VM 109 (.7) for Docker operations
- `apt`, `docker`, `docker compose` available on target hosts
- LiteLLM master key available (Infisical vault, `LITELLM_API_KEY`) for health verification
- `infrastructure-monitoring` run completed within the last 30 minutes — provides the pre-update health baseline used by the regression check
- Writable backup directory `/tmp/infra-maintenance-backup-<date>/` on each mutated host
- Proxmox snapshot of the primary host available (or confirmed not required) before apt phase
## Continuity
- Self-driven: weekly cron `0 2 * * 0` (Sunday 02:00 ET)
- Also wakes on: explicit "infra maintenance" trigger from ops/abiba
- Depends on `infrastructure-monitoring` for the pre-update health baseline — do not run if the last monitoring run is stale (>30 min) or RED; abort and escalate instead
## Execution
### Phase 0 — Preflight (snapshot/backup check + health baseline)
1. **Capture health baseline** — run the `infrastructure-monitoring` postcondition checks (LiteLLM, Zulip, Gitea, SearXNG, Proxmox API) and record results as `health-baseline`. If any critical service is already down, **abort**: maintenance must not run on a degraded fleet.
2. **Backup check** — confirm a Proxmox snapshot of the primary host exists OR `/tmp/infra-maintenance-backup-<date>/` was created this run. Snapshot critical config files into the backup dir:
- `/opt/inference-harness/docker-compose.yml`, `/opt/inference-harness/litellm_config.yaml` (CT 116)
- `/opt/search-stack/searxng/docker-compose.yml`, `/opt/search-stack/firecrawl-source/docker-compose.yaml` (VM 109)
3. **Record image baseline**`docker inspect --format '{{.Image}} {{.Config.Image}}' <container>` for every running container on .116 and .7; store digests in `image-baseline` so rollback can restore them.
4. **Disk check**`df -h` on each mutated host; abort if free space <20% (apt upgrade + image pulls need headroom).
### Phase 1 — OS package updates (primary host)
```bash
# On the primary host only (VERIFY host against infrastructure-control first)
apt update
apt upgrade -y
```
- Capture `apt list --upgradable` before and after → store in `apt-state`.
- If apt reports held/broken packages (`apt-get -s upgrade | grep -i broken`, or non-zero exit), **stop** — do not force. Record `held_broken` and go to rollback/escalate.
- If `/var/run/reboot-required` exists after upgrade, flag `kernel_reboot_required: true` in `apt-state` but **do not reboot automatically** — that's a separate coordinated action (see infra-update Wave 4). Note it in the report.
### Phase 2 — Docker image pulls
Pull latest stable tags for every running container. Do NOT pin to `:main`/`:nightly` — use stable tags where the compose file specifies them; otherwise `latest`.
```bash
# CT 116 (.116) — inference harness
cd /opt/inference-harness && docker compose pull
# VM 109 (.7) — search + home stacks
cd /opt/search-stack/searxng && docker compose pull
cd /opt/search-stack/firecrawl-source && docker compose pull
# any other running stacks on .7 (home stack, audiobookshelf) — pull per their compose files
```
- LiteLLM and SearXNG are the two explicitly required pulls; "any other running containers" means every stack with a compose file on .116 and .7.
- Record pulled tag + digest per service in `image-baseline`.
### Phase 3 — Container restarts with health verification
Restart one stack at a time, verify health before moving to the next. Do not restart everything at once — a failure mid-wave must leave the rest running.
```bash
# CT 116
cd /opt/inference-harness && docker compose up -d
# VM 109
cd /opt/search-stack/searxng && docker compose up -d
cd /opt/search-stack/firecrawl-source && docker compose up -d
```
After each stack comes up, wait for health (max 120s):
- `docker ps` shows the container `Up` (and `healthy` if a healthcheck is defined)
- Service-specific probe passes (see Phase 4 probes)
If a stack fails to come up within 120s, **stop the wave** and go to rollback for that stack only; do not proceed to the next.
### Phase 4 — Post-update service verification
After ALL updates (apt + images + restarts), verify every critical service is back up and matches the pre-update baseline. This is the regression gate.
| Service | Probe | Expect |
|---------|-------|--------|
| LiteLLM proxy | `curl -sf http://192.168.68.116/litellm/v1/models` | 200 OK, models returned |
| LiteLLM MCP gateway | `curl -sf http://192.168.68.116:4000/mcp-rest/tools/list -H "Authorization: Bearer $MASTER_KEY"` | 90 tools (23 RA-H OS + 67 GitHub) |
| SearXNG | `curl -sf http://192.168.68.7:8888` | 200 OK |
| Zulip | `curl -sf https://chat.sysloggh.net/api/v1/server_settings` | 200 OK |
| Gitea | `curl -sf https://git.sysloggh.net/api/v1/version` | 200 OK |
| PM2 processes | `pm2 jlist` (CT 100) | all pi-agent processes `online` |
| Hermes gateways | SSH to Mumuni CT 114, Tanko CT 112; `systemctl is-active hermes-gateway` | `active` for each |
Regression check: every service that was GREEN in `health-baseline` must still be GREEN. A service that was already RED (and caused a preflight abort) is excluded — but Phase 0 should have aborted before we got here.
## Rollback Protocol
If ANY service in Phase 4 fails to come back up (or regresses vs baseline):
1. **Image rollback** — for the failing stack, restore the previous image:
```bash
# Restore from recorded image-baseline digest
docker compose down
# Pin the service image to the recorded digest in compose, then recreate
# image: <name>@sha256:<previous_digest>
docker compose pull && docker compose up -d
```
2. **APT rollback** — restore the primary host from the Proxmox snapshot taken/confirmed in Phase 0. If no snapshot, `apt install <pkg>=<old_version>` per package using apt history (`/var/log/apt/history.log`).
3. **Config rollback** — restore configs from `/tmp/infra-maintenance-backup-<date>/`.
4. **Re-verify** — re-run the Phase 4 probes on the rolled-back service. If still failing, escalate (do not loop — circuit breaker below).
5. **Escalate** — send a Zulip DM to abiba + mumuni with: failing service, phase, baseline vs current, rollback actions taken, backup path.
## Circuit Breaker
- `max_retries: 2` per failing phase — after 2 rollback attempts on the same service, stop and escalate.
- `window: 7200` seconds — no more than 2 retries within a 2-hour window.
- `trip_action: escalate_to_fatal` — when tripped, escalate to fatal (abiba + mumuni + kwame) and pause; a human must clear before the next scheduled run.
## Report
After completion (or on abort), emit a receipt (JSON) to `~/.hermes/runs/infrastructure-maintenance/` and send a Zulip DM summary:
```
🛠 Infrastructure Maintenance — YYYY-MM-DD
Phase: apt | images | restarts | verify | rollback
Primary host: <host>
APT: <N> packages upgraded, <M> held/broken, kernel_reboot_required=<bool>
Images pulled: LiteLLM <tag>, SearXNG <tag>, <others>
Services: all GREEN | <service> FAILED (rolled back)
Baseline regression: none | <details>
Backup: /tmp/infra-maintenance-backup-YYYYMMDD/
Escalation: none | warning | critical | fatal
```
## Verification Postconditions
- All critical services running after update (Phase 4 all GREEN)
- No regressions from pre-update health baseline (Phase 0 baseline)
- Docker containers on latest stable tags (`image-baseline.pulled_tag` recorded)
- APT packages up to date with no held broken packages (`apt-state.held_broken == 0`)
## Related Contracts
- `infrastructure-update.prose.md` — owns the full-fleet cluster-wide wave INCLUDING its Wave 3 Docker image updates across all stacks (SearXNG, Firecrawl, Inference Harness on .116, home stack, audiobookshelf); infrastructure-maintenance is a deliberately narrower host-level weekly pull scoped to .116 and .7.
- `infrastructure-monitoring.prose.md` — provides the pre-update health baseline (depends_on).
- `infrastructure-control.prose.md` — topology source of truth (CT IDs, IPs, hostnames).
- `litellm-health.prose.md` — LiteLLM probe details.
- `proxmox-monitor.prose.md` — Docker stats + monitoring stack health.