Compare commits

..
Author SHA1 Message Date
root 6cc1b7d492 no-mistakes(document): Update 5→6 node references and Mumuni placement across all docs 2026-07-20 17:27:28 +00:00
root 35d09ba69a no-mistakes(review): Fix 5→6 node examples and parameterize restart commands 2026-07-20 17:19:18 +00:00
root df836b115c no-mistakes(review): fix review findings: 5→6 node labels, parameterize dual-gateway restart 2026-07-20 17:13:55 +00:00
root 9f0e04f22c fix: update contracts for hwepve 6th node + Mumuni migration (lxc/114 on hwepve)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
- proxmox-monitor: 5→6 nodes, add hwepve (.4, Huawei Matebook 16), note Mumuni (lxc/114) placement
- zulip-health: add dual-gateway collision detection to B2, note Mumuni on hwepve
- hermes-agent-baseline: fix Mumuni node minipve→hwepve
- hermes-zulip-restore: add hwepve Proxmox column for Mumuni
- mumuni-delegation: fix CT 118/storepve/.6 → lxc/114/hwepve/.123, 5→6 nodes
- pm2-self-heal: verified — no Mumuni references, no changes needed

Post-migration discoveries:
- Mumuni container migrated from minipve to hwepve (6th node, 192.168.68.4)
- IP preserved (192.168.68.123), hostname 'mumuni' unchanged
- Dual gateway collision (PID 1442 --replace + PID 2924 --force) resolved by
  container restart at 10:07; gateway restarted clean at 13:06 with fresh Zulip
  queue (errors=0, reconnects=0)
2026-07-20 17:08:58 +00:00
jerome fc88265e76 Merge pull request 'tune: switch compression model from strix-moe to syslog-auto' (#25) from tune/compression-syslog-auto into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
Reviewed-on: #25
2026-07-19 01:32:44 +00:00
jerome 51da19d92d Merge pull request 'feat(contracts): add infrastructure-maintenance contract' (#24) from fm/infra-maint-contract into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
Reviewed-on: #24
2026-07-19 01:32:29 +00:00
root 6570fd60e7 tune: switch compression model from strix-moe to syslog-auto
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
Relieve Strix Halo pressure by distributing compression across the
syslog-auto weighted pool (55% RTX 3090, 30% Strix Halo, 15% RTX 5070).

Updates:
- compression.model: strix-moe -> syslog-auto
- auxiliary.compression.model: strix-moe -> syslog-auto
- Rule 7: Updated for syslog-auto compression, removed aux prohibition
- Rule 8: Updated GPU workload distribution
- All docs/comments updated to reflect the change
2026-07-18 23:28:07 +00:00
jerome 5d2ecbace6 Merge branch 'master' into fm/infra-maint-contract
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
2026-07-18 23:12:20 +00:00
root 99789a00a1 no-mistakes(document): Reframed infra-maintenance contract as deliberate partition
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
2026-07-18 22:16:09 +00:00
root a68879e904 no-mistakes(review): fix docker rollback command to pin digest then compose up 2026-07-18 22:07:15 +00:00
root 1d027f71f6 feat(contracts): add infrastructure-maintenance contract
New responsibility contract consolidating host-level system maintenance
and Docker image lifecycle management, filling the gap left by
infrastructure-update (which owns cluster-wide apt waves).

Scope:
- OS package updates on primary host with pre-update snapshot/backup check
- Docker image pulls for LiteLLM, SearXNG, and other running containers
- Container restarts with per-stack health verification
- Post-update verification: LiteLLM, SearXNG, Zulip, Gitea, PM2, Hermes gateways
- Rollback on failure (image/apt/config restore) with circuit breaker

Owner: ops (firstmate secondmate). Trigger: weekly Sunday 2am ET.
Escalation: warning/critical->abiba+mumuni, fatal->abiba+mumuni+kwame.
circuit_breaker: max_retries 2, window 7200, trip_action escalate_to_fatal.
depends_on: infrastructure-monitoring (pre-update health baseline).

Registry:
- Add infrastructure-maintenance to by_category.maintenance, by_domain.infrastructure,
  by_owner.ops (new), by_trigger.scheduled, by_sensitivity.high
- Add 'ops' to owners list
- Move infrastructure-update owner abiba -> ops (in contracts entry + by_owner index)

Also adds '## Maintaining this file' section to AGENTS.md per fm-ensure-agents-md.
2026-07-18 22:02:45 +00:00
18 changed files with 389 additions and 70 deletions
+7
View File
@@ -128,3 +128,10 @@ safe-mutate --verify "CMD" [--expect "PATTERN"] --mutate "CMD" [--reason "WHY"]
Read the [Authoring Guide](docs/AUTHORING-GUIDE.md) before writing any new contract.
It covers the full process: verify → draft → lint → review → ship, with templates
and style rules.
## Maintaining this file
Keep this file for knowledge useful to almost every future agent session in this project.
Do not repeat what the codebase already shows; point to the authoritative file or command instead.
Prefer rewriting or pruning existing entries over appending new ones.
When updating this file, preserve this bar for all agents and keep entries concise.
Symlink
+1
View File
@@ -0,0 +1 @@
AGENTS.md
+1 -1
View File
@@ -137,7 +137,7 @@ Encode topology, architectural decisions, and lessons. You read them before plan
| Contract | Description |
|---|---|
| `infrastructure-control` | Full topology and control pattern: 5-node Proxmox cluster, 3 Docker ecosystems, NFS storage, network verification, IP-first configuration doctrine. Live-state fields marked `VERIFY-BEFORE-USE`. |
| `infrastructure-control` | Full topology and control pattern: 6-node Proxmox cluster, 3 Docker ecosystems, NFS storage, network verification, IP-first configuration doctrine. Live-state fields marked `VERIFY-BEFORE-USE`. |
| `pi-approval-architecture` | pi's approval model vs Hermes, available commands, architectural constraints. |
| `zulip-adapter-lessons` | Failure modes, fixes, and patterns from building the pi Zulip extension and Hermes Zulip plugin. |
+104 -2
View File
@@ -22,6 +22,7 @@ owners:
- abiba
- mumuni
- kwame
- ops
trigger_types:
- scheduled
- event_driven
@@ -58,6 +59,7 @@ index:
- memory-audit-maintenance
- gpu-fleet
- infrastructure-update
- infrastructure-maintenance
reference:
- infrastructure-control
- ra-h-os-custodianship-contract
@@ -90,6 +92,7 @@ index:
- infrastructure-control
- infrastructure-monitoring
- infrastructure-update
- infrastructure-maintenance
- pm2-self-heal
- disk-gc-threat-response
gpu:
@@ -130,7 +133,6 @@ index:
- build-zulip-plugin
- stirling-pdf-agent-access
- gpu-fleet
- infrastructure-update
- infrastructure-control
- zulip-adapter-lessons
- pi-approval-architecture
@@ -144,6 +146,9 @@ index:
- mumuni-delegation
kwame:
- hello-world
ops:
- infrastructure-maintenance
- infrastructure-update
by_trigger:
scheduled:
- hermes-key-enforcement
@@ -156,6 +161,7 @@ index:
- litellm-health
- memory-audit-maintenance
- infrastructure-update
- infrastructure-maintenance
event_driven:
- litellm-self-heal
- pm2-self-heal
@@ -197,6 +203,7 @@ index:
- hermes-zulip-plugin
- build-zulip-plugin
- infrastructure-update
- infrastructure-maintenance
- ra-h-os-custodianship-contract
- mumuni-delegation
normal:
@@ -1256,13 +1263,108 @@ contracts:
last_run: null
last_status: null
drift_alerts: []
- name: infrastructure-maintenance
file: infrastructure-maintenance.prose.md
kind: responsibility
category: maintenance
sensitivity: high
status: active
owner: ops
version: 1.0.0
trigger:
type: scheduled
cadence: 0 2 * * 0
description: Weekly host-level maintenance Sunday at 2am ET (replaces infrastructure-update
build-phase role; infra-update moves to ops)
cron_job_id: null
execution:
agent: ops
timeout: 3600
requires:
- infrastructure-monitoring run within last 30 minutes (pre-update health baseline)
- Proxmox snapshot of primary host OR /tmp backup dir created this run
- LiteLLM master key from Infisical vault for health verification
protocol:
- Load contract from prose-contracts/main
- Phase 0 preflight — capture health baseline, backup check, record image baseline
- Phase 1 apt update && apt upgrade -y on primary host
- Phase 2 docker compose pull for LiteLLM, SearXNG, and other running containers
- Phase 3 restart stacks one at a time with per-stack health verification
- Phase 4 verify every critical service (LiteLLM, SearXNG, Zulip, Gitea, PM2, Hermes gateways)
- On failure — rollback per protocol, escalate, do not loop beyond circuit breaker
- Log actions to ~/.hermes/runs/infrastructure-maintenance/
verification:
postconditions:
- check: all critical services running after update
verify: 'curl -sf http://192.168.68.116/litellm/v1/models && curl -sf https://chat.sysloggh.net/api/v1/server_settings && curl -sf https://git.sysloggh.net/api/v1/version && curl -sf http://192.168.68.7:8888 && pm2 jlist'
expect: all probes 200 OK / processes online
- check: no regressions from pre-update health baseline
verify: diff Phase 0 health-baseline against Phase 4 results
expect: no GREEN service turned RED
- check: docker containers on latest stable tags
verify: docker inspect --format '{{.Config.Image}}' <container> per service matches image-baseline.pulled_tag
expect: all containers running pulled tags
- check: APT packages up to date with no held broken packages
verify: apt list --upgradable 2>/dev/null | wc -l and apt-get -s upgrade | grep -ci broken
expect: upgradable == 0, broken == 0
artifact: maintenance run report with phase results and any rollback/escalation
verify_commands:
- curl -sf http://192.168.68.116/litellm/v1/models
- curl -sf http://192.168.68.7:8888
- curl -sf https://chat.sysloggh.net/api/v1/server_settings
- curl -sf https://git.sysloggh.net/api/v1/version
- pm2 jlist
receipt:
format: json
storage: ~/.hermes/runs/infrastructure-maintenance/
graph_node: true
schema:
contract: string
run_id: string
timestamp: ISO 8601
agent: string
status: pass|fail|escalated
phase: preflight|apt|images|restarts|verify|rollback|done|failed
actions_taken: array
postconditions: array
drift_alerts: array
evidence_path: string
escalation:
info:
action: log_to_receipt
notify: []
warning:
action: relay_alert
notify:
- abiba
- mumuni
critical:
action: relay_alert
notify:
- abiba
- mumuni
fatal:
action: relay_alert + pause + human_required
notify:
- abiba
- mumuni
- kwame
circuit_breaker:
max_retries: 2
window: 7200
trip_action: escalate_to_fatal
depends_on:
- infrastructure-monitoring
last_run: null
last_status: null
drift_alerts: []
- name: infrastructure-update
file: infrastructure-update.prose.md
kind: responsibility
category: maintenance
sensitivity: high
status: active
owner: abiba
owner: ops
version: 1.0.0
trigger:
type: scheduled
+1 -1
View File
@@ -286,7 +286,7 @@ one-off GPU builds. No automated post-migration cleanup was in place.
| 111 | tdunna | amdpve | ✅ reachable |
| 112 | tanko | amdpve | ✅ reachable |
| 113 | baggy | amdpve | ✅ reachable |
| 114 | mumuni | minipve | ✅ reachable |
| 114 | mumuni | hwepve | ✅ reachable |
| 115 | scottdenya | amdpve | ✅ reachable |
| 116 | syslog-api | minipve | ✅ reachable |
| 117 | zulip | storepve | ✅ reachable |
+1 -1
View File
@@ -185,7 +185,7 @@ what, and why should I care?
```
❌ "Monitors infrastructure health"
✅ "Scans all 5 Proxmox nodes and 19 CTs for disk pressure, checks Docker
✅ "Scans all 6 Proxmox nodes and 19 CTs for disk pressure, checks Docker
container health on .7/.116/.17, alerts via Telegram DM on RED/CRITICAL"
```
+1 -1
View File
@@ -25,7 +25,7 @@ done
| Agent | CT | Node | IP | LiteLLM Alias | Key Source | Platform |
|-------|-----|------|-----|---------------|------------|----------|
| Tanko | 112 | amdpve | .122 | `tanko` | Infisical vault | Hermes |
| Mumuni | 114 | minipve | .123 | `mumuni` | Infisical vault | Hermes |
| Mumuni | 114 | hwepve | .123 | `mumuni` | Infisical vault | Hermes |
| Koby | 129 | amdpve | srv1079750 | `koby` | Infisical vault | **Hermes** |
| Koonimo | 114 | amdpve | ? | `koonimo` | Infisical vault | Hermes |
| Shumba | — | 192.168.68.119 | N/A | N/A (DeepSeek) | Hermes (RETIRED — CT119 now Infisical vault) |
+29 -22
View File
@@ -5,11 +5,14 @@ description: >
Standard Hermes configuration template for Syslog Solution LLC agents.
Enforces shared infrastructure setup (Firecrawl, SearXNG, local models,
RA-H OS MCP) while keeping agent-specific API keys and model choices.
UPDATED 2026-07-16: Compression model is the stable alias `strix-moe` (NOT `ornith-1.0-35b`,
UPDATED 2026-07-18: Compression model switched to `syslog-auto` (was `strix-moe`)
to relieve Strix Halo pressure. syslog-auto distributes compression across the
weighted pool (55% RTX 3090, 30% Strix Halo, 15% RTX 5070).
UPDATED 2026-07-16: Compression model was the stable alias `strix-moe` (NOT `ornith-1.0-35b`,
which LiteLLM does not serve). All 3 GPUs verified at 128K (reduced from 256K 2026-07-17 for stability).
Added Rule 12 (Context-Issue Diagnostic) + Rule 13 (.env fallback enforcement) from the
2026-07-16 Mumuni root-cause investigation (WAL #1300).
UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo. RTX 3090 context verified at 128K. Infisical .env fallback required (Rule 3/13).
UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo (later switched to syslog-auto 2026-07-18). RTX 3090 context verified at 128K. Infisical .env fallback required (Rule 3/13).
---
## Maintains
@@ -130,7 +133,9 @@ mcp_servers:
# ─── Compression ───
compression:
enabled: true
model: strix-moe # ⚠️ Must match auxiliary.compression.model. Stable alias (gpu-fleet § Stable Role-Based Aliases). NOT ornith-1.0-35b (LiteLLM does not serve that name).
model: syslog-auto # ⚠️ Switched from strix-moe 2026-07-18 to relieve Strix Halo.
# syslog-auto distributes across weighted pool (55% RTX 3090,
# 30% Strix Halo, 15% RTX 5070). All GPUs at 128K.
provider: harness
max_context_window: 131072 # MUST match actual GPU capacity. All 3 GPUs are 128K (Jul 17).
threshold: 0.65 # Fires at ~170K for 262K window, ~85K for 128K
@@ -145,8 +150,9 @@ compression:
# model: gpu-light # stable alias (NOT raw "gemma-4-12b")
# base_url: http://192.168.68.116/v1
# api_key_env: LITELLM_API_KEY
# Do NOT use syslog-auto for auxiliary tasks — it routes to the primary GPU.
# gpu-light = RTX 5070 (12B), freeing the Strix Halo for agent reasoning.
# Compression uses syslog-auto (switched from strix-moe 2026-07-18) to distribute
# load across the weighted pool and relieve Strix Halo pressure.
# Vision and web_extract use gpu-light = RTX 5070 (12B).
# Heavy aux (delegation, x_search) use gpu-dense (RTX 3090) instead.
# NEVER use raw model names (gemma-4-12b, qwen3.6-27B-code, qwen3.6-35B-udq4)
# in agent configs — use the stable aliases so model swaps don't break agents.
@@ -166,7 +172,7 @@ auxiliary:
timeout: 30
compression:
provider: harness
model: strix-moe # MUST match compression.model above. Stable alias for Strix Halo.
model: syslog-auto # Switched from strix-moe 2026-07-18. Relieves Strix Halo pressure.
base_url: http://192.168.68.116/v1 # Rule 5: /v1 NOT /litellm/v1
api_key_env: LITELLM_API_KEY
timeout: 300 # gpu-fleet: 300s for large-history summarization (was 60)
@@ -247,29 +253,30 @@ The following MUST be identical across ALL profiles:
- Apply to BOTH main config AND all sub-agent profiles
- For agents needing longer outputs: raise to 8192, but never omit
### Rule 7: Auxiliary Model Consistency (UPDATED 2026-07-16)
- Vision and web_extract use `gemma-4-12b` (RTX 5070 — 12GB, vision-optimized)
- Compression uses `strix-moe` (stable alias for Strix Halo — 64GB, 128K ctx, compression-optimized)
- **`strix-moe` is the only valid compression model name** — LiteLLM does NOT serve `ornith-1.0-35b`
(it serves `strix-moe`, `qwen3.6-35B-udq4`, `gpu-dense`, `gpu-light`, `syslog-auto`, `gemma-4-12b`, `qwen3.6-27B-code`). Old configs with `ornith-1.0-35b` cause 403/model-not-found on compression calls.
### Rule 7: Auxiliary Model Consistency (UPDATED 2026-07-18)
- Vision and web_extract use `gpu-light` (stable alias, RTX 5070 — 12GB, vision-optimized)
- Compression now uses `syslog-auto` (switched from `strix-moe` 2026-07-18) to distribute
compression load across the weighted pool (55% RTX 3090, 30% Strix Halo, 15% RTX 5070).
This relieves Strix Halo pressure while keeping compression functional on all GPUs.
- **`syslog-auto` is the valid compression model** — LiteLLM serves it as the weighted pool.
Old configs with `strix-moe` for compression should be updated to `syslog-auto`.
- All auxiliary services MUST use identical routing:
- `base_url: http://192.168.68.116/v1` (Rule 5: `/v1`, NOT `/litellm/v1`)
- `api_key_env: LITELLM_API_KEY`
- **Do NOT use `syslog-auto`** for auxiliary tasks — it routes unpredictably
- **Compression on Strix Halo**: The strix-moe alias routes to Strix Halo
(64GB UMA, 128K context) — the designated compression GPU. This frees the
RTX 5070 for vision and web search, and the RTX 3090 for heavy reasoning.
- **Compression via syslog-auto**: Routes through the weighted pool. Strix Halo still handles
~30% of compression calls (at 60 RPM via pool vs 40 RPM direct), but the bulk (55%)
goes to RTX 3090 which has ample spare capacity.
- The `compression:` block's `model` MUST match `auxiliary: compression: model`
- The `compression: max_context_window: 131072` MUST match actual GPU capacity (128K)
### Rule 8: GPU Workload Distribution (UPDATED 2026-07-16)
- **RTX 3090 (24GB, 128K ctx, qwen3.6-27B-code)**: Heavy reasoning, code gen, long conversations
- **RTX 5070 (12GB, 128K ctx, gemma-4-12b)**: Vision, web search, quick tasks, web_extract (IQ4_NL+MTP, ~65% VRAM at 128K)
- **Strix Halo (64GB, 128K ctx, strix-moe)**: Context compression, summarization, long docs
### Rule 8: GPU Workload Distribution (UPDATED 2026-07-18)
- **RTX 3090 (24GB, 128K ctx, qwen3.6-27B-code)**: Heavy reasoning, code gen, long conversations — also handles ~55% of compression via syslog-auto pool
- **RTX 5070 (12GB, 128K ctx, gemma-4-12b)**: Vision, web search, quick tasks, web_extract — handles ~15% of compression via syslog-auto pool
- **Strix Halo (64GB, 128K ctx, Geneis Hermes V3 APEX)**: Agent reasoning, compression (~30% via syslog-auto pool), fallback for other GPUs
- Agent profiles MUST route auxiliary tasks to the correct GPU:
- `auxiliary.vision.model: gemma-4-12b` (RTX 5070)
- `auxiliary.web_extract.model: gemma-4-12b` (RTX 5070)
- `auxiliary.compression.model: strix-moe` (Strix Halo)
- `auxiliary.vision.model: gpu-light` (RTX 5070)
- `auxiliary.web_extract.model: gpu-light` (RTX 5070)
- `auxiliary.compression.model: syslog-auto` (distributed pool, switched from strix-moe 2026-07-18)
- Default model (`model.default`) and custom_provider remain `syslog-auto` for auto-routing
- For 128K context window: `threshold: 0.65` (fires at ~85K tokens)
- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss
+1 -1
View File
@@ -51,7 +51,7 @@ gateway restart, and connection validation.
| Host | CT | Proxmox | IP (direct) | Hermes Home | User |
|------|-----|---------|-------------|-------------|------|
| Mumuni | CT114 | | 192.168.68.123 | /root/.hermes | root |
| Mumuni | CT114 | hwepve | 192.168.68.123 | /root/.hermes | root |
| Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome |
| Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root |
| Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky |
+9 -8
View File
@@ -3,7 +3,7 @@ kind: pattern
name: infrastructure-control
description: >
Full infrastructure monitoring and control pattern covering the
5-node Proxmox cluster, 3 Docker ecosystems (22 containers),
6-node Proxmox cluster, 3 Docker ecosystems (22 containers),
NFS storage, and network services. Defines monitors, remediations,
and the access matrix for all environments.
@@ -42,8 +42,8 @@ description: >
│ │ │ │ │
┌────┘ ┌────┘ ┌────┘ ┌────┘ ┌────┘
▼ ▼ ▼ ▼ ▼
minipve amdpve storepve acerpve ocupve
(.12) (.15) (.6) (.9) (.5)
minipve amdpve storepve acerpve ocupve hwepve
(.12) (.15) (.6) (.9) (.5) (.4)
┌─────────────────────────────────────────────┐
@@ -95,15 +95,16 @@ description: >
## Section 2: Proxmox Cluster — Monitoring
### Nodes (5)
### Nodes (6)
| Node | IP | CPU | RAM | VMs/CTs | Role |
|------|----|-----|-----|---------|------|
| minipve | .12 | 16C | 30GB | authentik, gitea, mumuni, syslog-api, jitsi | Auth, git, messaging |
| minipve | .12 | 16C | 30GB | authentik, gitea, syslog-api, jitsi | Auth, git, messaging |
| amdpve | .15 | 32C | 62GB | abiba, kagentz, tanko, tdunna, baggy, scottdenya | Agents, compute |
| storepve | .6 | 28C | 31GB | docker-vm, ra-h-os, PBS, media, zulip | Docker, storage, chat |
| acerpve | .9 | 28C | 31GB | llm-gpu, adguard | GPU VMs |
| ocupve | .5 | 12C | 14GB | ocu-llm | GPU VMs |
| hwepve | .4 | 12C | 15GB | mumuni | Huawei Matebook 16, agent host |
### Checks (every 5 min)
@@ -117,7 +118,7 @@ description: >
## Checks
- For each node in [minipve, amdpve, storepve, acerpve, ocupve]:
- For each node in [minipve, amdpve, storepve, acerpve, ocupve, hwepve]:
- GET /api2/json/nodes/{node}/status → check status == "online"
- GET /api2/json/nodes/{node}/status → cpu < 0.80
- GET /api2/json/nodes/{node}/status → free_mem > 10%
@@ -589,7 +590,7 @@ ssh root@192.168.68.110 "systemctl restart llama-server"
| 111 | tdunna | amdpve | .129 | Hermes agent | ✅ |
| 112 | tanko | amdpve | .122 | Hermes agent | ✅ |
| 113 | baggy | amdpve | ? | Hermes agent | ✅ |
| 114 | mumuni | minipve | .123 | Hermes agent | ✅ |
| 114 | mumuni | hwepve | .123 | Hermes agent | ✅ |
| 115 | scottdenya | amdpve | — | ? | ❌ |
| 116 | syslog-api | minipve | .116 | LiteLLM + Grafana | ❌ |
| 117 | zulip | storepve | — | Chat | ❌ |
@@ -621,7 +622,7 @@ Source of truth: `/root/scripts/pct-run.sh` or `prose-contracts/scripts/pct-run.
| 115 | scottdenya | amdpve | `pct-run 115` |
| 104 | authentik | minipve | `pct-run 104` |
| 110 | gitea | minipve | `pct-run 110` |
| 114 | mumuni | minipve | `pct-run 114` |
| 114 | mumuni | hwepve | `pct-run 114` |
| 116 | syslog-api | minipve | `pct-run 116` |
| 106 | ra-h-os | storepve | `pct-run 106` |
| 107 | proxmox-backup | storepve | `pct-run 107` |
+200
View File
@@ -0,0 +1,200 @@
---
kind: responsibility
name: infrastructure-maintenance
description: >
Weekly system-level maintenance for the Syslog inference fleet: OS package
updates on the primary host, Docker image pulls for LiteLLM/SearXNG and other
running containers, container restarts with health verification, post-update
verification of every critical service (LiteLLM proxy, SearXNG, Zulip, Gitea,
PM2 processes, Hermes gateways), and rollback on failure. Consolidates the
raw shell scripts that previously did this piecemeal. This contract owns the
HOST-LEVEL weekly maintenance loop on the primary host plus Docker image
pulls ONLY for .116 and .7, while infrastructure-update owns the FULL-FLEET
cluster-wide wave (apt across the full PVE cluster + CTs/VMs AND its Docker
image Wave 3 across all stacks). Runs Sunday 2am ET. Owner:
ops (firstmate secondmate). Blast radius: an unverified image pull can break
LiteLLM (all agents lose inference) or SearXNG (search-stack down); a bad apt
upgrade can leave the host in a half-upgraded state. Pre-update backup check
and rollback are mandatory for this reason.
agent: ops
triggers:
- weekly (Sunday 02:00 ET) via cron
- on demand when ops/abiba triggers "infra maintenance"
version: 1.0.0
---
## Maintains
- maintenance-status: { phase: idle|preflight|apt|images|restarts|verify|rollback|done|failed, host, step, result, timestamp }
- image-baseline: { service, current_tag, pulled_tag, digest, updated_at } — last known-good image per container
- apt-state: { upgradable_before, upgradable_after, held_broken, kernel_reboot_required }
- health-baseline: snapshot of critical-service health captured pre-update (used for regression check post-update)
- rollback-snapshot: { backup_path, configs, image_digests, timestamp } — restore point created in preflight
- maintenance-history: array of past runs with phase results and any escalations
## Scope
Primary host is the maintenance host where apt updates apply. Docker image pulls
span the two Docker ecosystems that run critical services. infrastructure-update
runs the full-fleet cluster-wide wave (including its Wave 3 Docker pulls across
all stacks/hosts); this contract runs a narrower host-level weekly pull limited
to .116 and .7. Topology, CT IDs, and IPs are live-state fields — verify against
`infrastructure-control.prose.md` (the source of truth) and the live system
before mutating.
| Host | IP | Role | Trust |
|------|----|------|-------|
| CT 116 (syslog-api) | 192.168.68.116 | LiteLLM proxy + Grafana + Prometheus (inference harness) | ⚠️ VERIFY-BEFORE-USE |
| VM 109 (docker-vm) | 192.168.68.7 | SearXNG + Firecrawl + home stack (Docker host) | ⚠️ VERIFY-BEFORE-USE |
| CT 117 (zulip) | 192.168.68.19 | Zulip (storepve bridge IP .19) | ⚠️ VERIFY-BEFORE-USE |
| Gitea | https://git.sysloggh.net | Prose-contracts + agent configs source control | ⚠️ VERIFY-BEFORE-USE |
| CT 100 (abiba/pi) | 192.168.68.24 | PM2 processes (pi agent harness) | ⚠️ VERIFY-BEFORE-USE |
> "Primary host" for the apt phase is the host the ops agent runs maintenance
> from. Confirm which host that is against infrastructure-control before
> running; do not assume. If the ops agent is containerized/CT-based, apt runs
> inside that CT.
## Requires
- SSH/exec access to CT 116 (.116) and VM 109 (.7) for Docker operations
- `apt`, `docker`, `docker compose` available on target hosts
- LiteLLM master key available (Infisical vault, `LITELLM_API_KEY`) for health verification
- `infrastructure-monitoring` run completed within the last 30 minutes — provides the pre-update health baseline used by the regression check
- Writable backup directory `/tmp/infra-maintenance-backup-<date>/` on each mutated host
- Proxmox snapshot of the primary host available (or confirmed not required) before apt phase
## Continuity
- Self-driven: weekly cron `0 2 * * 0` (Sunday 02:00 ET)
- Also wakes on: explicit "infra maintenance" trigger from ops/abiba
- Depends on `infrastructure-monitoring` for the pre-update health baseline — do not run if the last monitoring run is stale (>30 min) or RED; abort and escalate instead
## Execution
### Phase 0 — Preflight (snapshot/backup check + health baseline)
1. **Capture health baseline** — run the `infrastructure-monitoring` postcondition checks (LiteLLM, Zulip, Gitea, SearXNG, Proxmox API) and record results as `health-baseline`. If any critical service is already down, **abort**: maintenance must not run on a degraded fleet.
2. **Backup check** — confirm a Proxmox snapshot of the primary host exists OR `/tmp/infra-maintenance-backup-<date>/` was created this run. Snapshot critical config files into the backup dir:
- `/opt/inference-harness/docker-compose.yml`, `/opt/inference-harness/litellm_config.yaml` (CT 116)
- `/opt/search-stack/searxng/docker-compose.yml`, `/opt/search-stack/firecrawl-source/docker-compose.yaml` (VM 109)
3. **Record image baseline**`docker inspect --format '{{.Image}} {{.Config.Image}}' <container>` for every running container on .116 and .7; store digests in `image-baseline` so rollback can restore them.
4. **Disk check**`df -h` on each mutated host; abort if free space <20% (apt upgrade + image pulls need headroom).
### Phase 1 — OS package updates (primary host)
```bash
# On the primary host only (VERIFY host against infrastructure-control first)
apt update
apt upgrade -y
```
- Capture `apt list --upgradable` before and after → store in `apt-state`.
- If apt reports held/broken packages (`apt-get -s upgrade | grep -i broken`, or non-zero exit), **stop** — do not force. Record `held_broken` and go to rollback/escalate.
- If `/var/run/reboot-required` exists after upgrade, flag `kernel_reboot_required: true` in `apt-state` but **do not reboot automatically** — that's a separate coordinated action (see infra-update Wave 4). Note it in the report.
### Phase 2 — Docker image pulls
Pull latest stable tags for every running container. Do NOT pin to `:main`/`:nightly` — use stable tags where the compose file specifies them; otherwise `latest`.
```bash
# CT 116 (.116) — inference harness
cd /opt/inference-harness && docker compose pull
# VM 109 (.7) — search + home stacks
cd /opt/search-stack/searxng && docker compose pull
cd /opt/search-stack/firecrawl-source && docker compose pull
# any other running stacks on .7 (home stack, audiobookshelf) — pull per their compose files
```
- LiteLLM and SearXNG are the two explicitly required pulls; "any other running containers" means every stack with a compose file on .116 and .7.
- Record pulled tag + digest per service in `image-baseline`.
### Phase 3 — Container restarts with health verification
Restart one stack at a time, verify health before moving to the next. Do not restart everything at once — a failure mid-wave must leave the rest running.
```bash
# CT 116
cd /opt/inference-harness && docker compose up -d
# VM 109
cd /opt/search-stack/searxng && docker compose up -d
cd /opt/search-stack/firecrawl-source && docker compose up -d
```
After each stack comes up, wait for health (max 120s):
- `docker ps` shows the container `Up` (and `healthy` if a healthcheck is defined)
- Service-specific probe passes (see Phase 4 probes)
If a stack fails to come up within 120s, **stop the wave** and go to rollback for that stack only; do not proceed to the next.
### Phase 4 — Post-update service verification
After ALL updates (apt + images + restarts), verify every critical service is back up and matches the pre-update baseline. This is the regression gate.
| Service | Probe | Expect |
|---------|-------|--------|
| LiteLLM proxy | `curl -sf http://192.168.68.116/litellm/v1/models` | 200 OK, models returned |
| LiteLLM MCP gateway | `curl -sf http://192.168.68.116:4000/mcp-rest/tools/list -H "Authorization: Bearer $MASTER_KEY"` | 90 tools (23 RA-H OS + 67 GitHub) |
| SearXNG | `curl -sf http://192.168.68.7:8888` | 200 OK |
| Zulip | `curl -sf https://chat.sysloggh.net/api/v1/server_settings` | 200 OK |
| Gitea | `curl -sf https://git.sysloggh.net/api/v1/version` | 200 OK |
| PM2 processes | `pm2 jlist` (CT 100) | all pi-agent processes `online` |
| Hermes gateways | SSH to Mumuni CT 114, Tanko CT 112; `systemctl is-active hermes-gateway` | `active` for each |
Regression check: every service that was GREEN in `health-baseline` must still be GREEN. A service that was already RED (and caused a preflight abort) is excluded — but Phase 0 should have aborted before we got here.
## Rollback Protocol
If ANY service in Phase 4 fails to come back up (or regresses vs baseline):
1. **Image rollback** — for the failing stack, restore the previous image:
```bash
# Restore from recorded image-baseline digest
docker compose down
# Pin the service image to the recorded digest in compose, then recreate
# image: <name>@sha256:<previous_digest>
docker compose pull && docker compose up -d
```
2. **APT rollback** — restore the primary host from the Proxmox snapshot taken/confirmed in Phase 0. If no snapshot, `apt install <pkg>=<old_version>` per package using apt history (`/var/log/apt/history.log`).
3. **Config rollback** — restore configs from `/tmp/infra-maintenance-backup-<date>/`.
4. **Re-verify** — re-run the Phase 4 probes on the rolled-back service. If still failing, escalate (do not loop — circuit breaker below).
5. **Escalate** — send a Zulip DM to abiba + mumuni with: failing service, phase, baseline vs current, rollback actions taken, backup path.
## Circuit Breaker
- `max_retries: 2` per failing phase — after 2 rollback attempts on the same service, stop and escalate.
- `window: 7200` seconds — no more than 2 retries within a 2-hour window.
- `trip_action: escalate_to_fatal` — when tripped, escalate to fatal (abiba + mumuni + kwame) and pause; a human must clear before the next scheduled run.
## Report
After completion (or on abort), emit a receipt (JSON) to `~/.hermes/runs/infrastructure-maintenance/` and send a Zulip DM summary:
```
🛠 Infrastructure Maintenance — YYYY-MM-DD
Phase: apt | images | restarts | verify | rollback
Primary host: <host>
APT: <N> packages upgraded, <M> held/broken, kernel_reboot_required=<bool>
Images pulled: LiteLLM <tag>, SearXNG <tag>, <others>
Services: all GREEN | <service> FAILED (rolled back)
Baseline regression: none | <details>
Backup: /tmp/infra-maintenance-backup-YYYYMMDD/
Escalation: none | warning | critical | fatal
```
## Verification Postconditions
- All critical services running after update (Phase 4 all GREEN)
- No regressions from pre-update health baseline (Phase 0 baseline)
- Docker containers on latest stable tags (`image-baseline.pulled_tag` recorded)
- APT packages up to date with no held broken packages (`apt-state.held_broken == 0`)
## Related Contracts
- `infrastructure-update.prose.md` — owns the full-fleet cluster-wide wave INCLUDING its Wave 3 Docker image updates across all stacks (SearXNG, Firecrawl, Inference Harness on .116, home stack, audiobookshelf); infrastructure-maintenance is a deliberately narrower host-level weekly pull scoped to .116 and .7.
- `infrastructure-monitoring.prose.md` — provides the pre-update health baseline (depends_on).
- `infrastructure-control.prose.md` — topology source of truth (CT IDs, IPs, hostnames).
- `litellm-health.prose.md` — LiteLLM probe details.
- `proxmox-monitor.prose.md` — Docker stats + monitoring stack health.
+5 -4
View File
@@ -2,7 +2,7 @@
kind: responsibility
name: infrastructure-update
description: >
Autonomous system-wide update contract covering all 5 Proxmox nodes,
Autonomous system-wide update contract covering all 6 Proxmox nodes,
15+ containers/VMs, and 4 Docker ecosystems. Updates apt packages,
Docker images, and container stacks in safe waves with health checks
and automatic rollback on failure.
@@ -56,10 +56,11 @@ Before ANY update wave:
| amdpve (.15) | Proxmox node | `apt update && apt upgrade -y` | 5 min |
| acerpve (.9) | Proxmox node | `apt update && apt upgrade -y` | 5 min |
| ocupve (.5) | Proxmox node | `apt update && apt upgrade -y` | 5 min |
| hwepve (.4) | Proxmox node | `apt update && apt upgrade -y` | 5 min |
| CT 100 (.24) | Abiba (pi) | `apt update && apt upgrade -y` | 3 min |
| CT 116 (.116) | syslog-api (LiteLLM host) | `apt update && apt upgrade -y` | 3 min |
| CT 112 (tanko, amdpve) | Tanko | `apt update && apt upgrade -y` | 3 min |
| CT 114 (mumuni, minipve) | Mumuni | `apt update && apt upgrade -y` | 3 min |
| CT 114 (mumuni, hwepve) | Mumuni | `apt update && apt upgrade -y` | 3 min |
| VM 101 (.8) | llm-gpu (RTX 3090) | `apt update && apt upgrade -y` | 3 min |
| VM 103 (.110) | ocu-llm (RTX 5070) | `apt update && apt upgrade -y` | 3 min |
@@ -183,7 +184,7 @@ When LiteLLM is upgraded to a version supporting per-key MCP grants:
## Success Criteria
- [ ] All 5 PVE nodes updated, no reboot-loop
- [ ] All 6 PVE nodes updated, no reboot-loop
- [ ] All VMs/CTs running post-update
- [ ] All Docker containers healthy (VM 109 + CT 116 + CT 117)
- [ ] LiteLLM inference passing (syslog-auto test)
@@ -199,7 +200,7 @@ After completion, send Zulip DM:
```
📋 Infrastructure Update — YYYY-MM-DD
Updated: 5 PVE nodes, 12 CTs/VMs, 30+ containers
Updated: 6 PVE nodes, 12 CTs/VMs, 30+ containers
Security fixes: N CVEs patched
Downtime: <service> <duration>
Failures: none / <details>
+7 -7
View File
@@ -6,7 +6,7 @@ description: >
delegation, verification, and delivery. Defines when to delegate, which
worker to use for what, how to handle failures, and the kanban board
protocol. Enforces context-window discipline and separation of concerns.
Runs on Mumuni (CT 118, storepve, .6) via Hermes agent.
Runs on Mumuni (lxc/114, hwepve, .123) via Hermes agent.
version: 1.0.0
---
@@ -19,8 +19,8 @@ version: 1.0.0
## Topology
**Cluster:** 5 Proxmox nodes (ocupve, acerpve, minipve, amdpve, storepve)
**Manager:** Mumuni (CT 118, storepve, .6) via Hermes agent
**Cluster:** 6 Proxmox nodes (ocupve, acerpve, minipve, amdpve, storepve, hwepve)
**Manager:** Mumuni (lxc/114, hwepve, .123) via Hermes agent
**Workers:** 6 profiles, all running on the same agent — no separate hosts needed
This contract is infrastructure-agnostic in terms of which nodes are used.
@@ -31,7 +31,7 @@ Workers execute tasks on whatever infrastructure they're given — SSH to .6,
## Why This Matters
Without enforced delegation, the manager consumes the full iteration budget
(60 calls) on single-turn tasks — SSH to 5 nodes, check each VM, read logs —
(60 calls) on single-turn tasks — SSH to 6 nodes, check each VM, read logs —
leaving no capacity for actual coordination. The result: context overflow
(59K tokens in system prompt), iteration exhaustion, and degraded response
quality. This contract exists because I blew through my budget checking
@@ -82,7 +82,7 @@ it asks the manager (via relay) — it doesn't go find it on its own.
**This is a hard rule, not a recommendation.** Violating it produces the exact
type of discrepancy the kanban pipeline exists to prevent: a review worker finds
"5 nodes present" in the raw data but "5/5 online" in the report — even though
"6 nodes present" in the raw data but "5/6 online" in the report — even though
one of those nodes was unreachable. The report lied because it used data the
raw data never provided.
@@ -137,7 +137,7 @@ delegate_task(
```
delegate_task(
tasks=[
{"goal": "Check all 5 Proxmox nodes for VM status", "context": "SSH to each node via 192.168.68.x, run 'qm list'"},
{"goal": "Check all 6 Proxmox nodes for VM status", "context": "SSH to each node via 192.168.68.x, run 'qm list'"},
{"goal": "Check Docker container health on .7/.116/.17", "context": "SSH to each host, check container status"},
]
)
@@ -191,7 +191,7 @@ Only verified results reach Kwame. Format per channel:
{
"lane_id": "devops-check",
"worker": "syslog-devops",
"goal": "Check all 5 Proxmox nodes",
"goal": "Check all 6 Proxmox nodes",
"status": "dispatched|completed|failed",
"output_file": "/tmp/node-report.md"
}
+9 -8
View File
@@ -5,7 +5,7 @@ description: >
Proxmox cluster + Docker monitoring via the existing Grafana/Prometheus stack
on CT 116. Replaces Pulse with file-provisioned Grafana dashboards. Three
exporters feed Prometheus: prometheus-pve-exporter (cluster-aware, single
instance), node_exporter (all 5 PVE nodes), and a custom docker-stats-exporter
instance), node_exporter (all 6 PVE nodes), and a custom docker-stats-exporter
(Docker 29 / containerd image-store compatible, since cAdvisor cannot resolve
the layerdb). Dashboards exposed at http://192.168.68.116:3001/ (direct LAN, not behind nginx).
agent: abiba
@@ -20,7 +20,7 @@ agent: abiba
└───────┬──────────────┬──────────────┬───────────────────────┘
│ │ │
▼ ▼ ▼
pve-exporter docker-stats (scrapes 5x node_exporter)
pve-exporter docker-stats (scrapes 6x node_exporter)
:9221 :9324
│ │
▼ ▼
@@ -33,8 +33,8 @@ agent: abiba
| Exporter | Host:Port | Scope | Notes |
|----------|-----------|-------|-------|
| prometheus-pve-exporter | .116:9221 (container) | All 5 nodes + 14 guests + 36 storage pools | Single instance, cluster-aware via amdpve API. Config `/opt/monitoring/pve.yml` (token `monitoring@pve!prometheus`, PVEAuditor role). Metric schema is label-based (`id=node/amdpve`, `id=lxc/100`). |
| node_exporter | .5/.6/.9/.12/.15:9100 (systemd) | Per-node CPU/mem/disk/net/temp | Installed via apt on all 5 PVE nodes, enabled (reboot-persistent). Collectors: textfile, systemd, tcpstat, ethtool. |
| prometheus-pve-exporter | .116:9221 (container) | All 6 nodes + guests + 36 storage pools | Single instance, cluster-aware via amdpve API. Config `/opt/monitoring/pve.yml` (token `monitoring@pve!prometheus`, PVEAuditor role). Metric schema is label-based (`id=node/amdpve`, `id=lxc/100`). |
| node_exporter | .5/.6/.9/.12/.15/.4:9100 (systemd) | Per-node CPU/mem/disk/net/temp | Installed via apt on all 6 PVE nodes, enabled (reboot-persistent). Collectors: textfile, systemd, tcpstat, ethtool. hwepve (.4) added 2026-07-19. |
| docker-stats-exporter | .116:9324 (container) | 10 Docker containers on .116 | **Custom** (cAdvisor v0.51 incompatible with Docker 29 containerd image store — layerdb gone). Uses Docker Engine API over unix socket. Script `/opt/monitoring/docker-stats-exporter.py`. |
## PVE API Token
@@ -48,7 +48,7 @@ agent: abiba
| UID | Title | Panels | Source |
|-----|-------|--------|--------|
| proxmox-cluster | Proxmox Cluster Overview | 16 | cluster status, 5-node CPU/mem/disk/load gauges, guests table, storage pools, guest CPU/mem timeseries |
| proxmox-cluster | Proxmox Cluster Overview | 16 | cluster status, 6-node CPU/mem/disk/load gauges, guests table, storage pools, guest CPU/mem timeseries |
| proxmox-node | Proxmox Node Detail | 13 | per-node CPU per-core, memory, network, disk IO/IOPS/latency, temperature, disk space (variable: $node) |
| docker-containers | Docker Containers | 10 | per-container CPU/mem/network, restarts, memory limit ratio (variable: $container) |
| gpu-fleet | GPU Fleet | 7 | (existing, preserved in DB, not provisioned) |
@@ -71,15 +71,15 @@ agent: abiba
| File | Host | Purpose |
|------|------|---------|
| `/opt/monitoring/docker-compose.yml` | .116 | monitoring stack (prometheus, grafana, pve-exporter, docker-stats) |
| `/opt/monitoring/prometheus.yml` | .116 | 6 scrape jobs (3 GPU, pve, node x5, docker-stats) |
| `/opt/monitoring/prometheus.yml` | .116 | 6 scrape jobs (3 GPU, pve, node x6, docker-stats) |
| `/opt/monitoring/pve.yml` | .116 | PVE API credentials (chmod 644, contains token) |
| `/opt/monitoring/docker-stats-exporter.py` | .116 | custom Docker metrics exporter |
| `/opt/monitoring/grafana/dashboards/build-dashboards.py` | .116 | dashboard JSON generator |
| `/opt/monitoring/grafana/dashboards/json/*.json` | .116 | provisioned dashboard definitions |
| `/opt/monitoring/grafana/datasources/prometheus.yml` | .116 | datasource provisioning |
| `/etc/default/prometheus-node-exporter` | .5/.6/.9/.12/.15 | node_exporter collector config |
| `/etc/default/prometheus-node-exporter` | .5/.6/.9/.12/.15/.4 | node_exporter collector config |
## Cluster "Tabiri" — 5 Nodes
## Cluster "Tabiri" — 6 Nodes
| Node | IP | Role |
|------|----|----|
@@ -88,6 +88,7 @@ agent: abiba
| acerpve | 192.168.68.9 | PVE (hosts llm-gpu qemu/101) |
| minipve | 192.168.68.12 | PVE |
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (ornith) |
| hwepve | 192.168.68.4 | PVE (Huawei Matebook 16, 12C/15GB) — hosts Mumuni (lxc/114) migrated from minipve 2026-07-20 |
## Operations
+3 -1
View File
@@ -17,10 +17,11 @@ declare -A CT_NODES=(
[112]=amdpve # tanko
[113]=amdpve # baggy
[115]=amdpve # scottdenya
# hwepve (192.168.68.4) — Huawei Matebook 16
[114]=hwepve # mumuni (migrated from minipve 2026-07-20)
# minipve (192.168.68.12)
[104]=minipve # authentik
[110]=minipve # gitea
[114]=minipve # mumuni
[116]=minipve # syslog-api
# storepve (192.168.68.6)
[106]=storepve # ra-h-os
@@ -48,6 +49,7 @@ declare -A NODE_IPS=(
[storepve]=192.168.68.6
[acerpve]=192.168.68.9
[ocupve]=192.168.68.5
[hwepve]=192.168.68.4
)
resolve_node() {
+3 -2
View File
@@ -46,12 +46,13 @@ You are a code reviewer for OpenProse infrastructure contracts in the Syslog Sol
The infrastructure-control.prose.md contract is the canonical reference for the cluster topology:
**Proxmox Cluster "Tabiri" (5 nodes):**
**Proxmox Cluster "Tabiri" (6 nodes):**
- amdpve (192.168.68.15): abiba, kagentz, tanko, tdunna, baggy, scottdenya
- minipve (192.168.68.12): authentik, gitea, mumuni, syslog-api, jitsi
- minipve (192.168.68.12): authentik, gitea, syslog-api, jitsi
- storepve (192.168.68.6): docker-vm, ra-h-os, PBS, media, zulip
- acerpve (192.168.68.9): llm-gpu, adguard
- ocupve (192.168.68.5): ocu-llm
- hwepve (192.168.68.4): mumuni (migrated from minipve 2026-07-20, Huawei Matebook 16, 12C/15GB)
**CT IDs (verified 2026-07-04 against PVE API):**
100:abiba 102:adguard 104:authentik 105:kagentz 106:ra-h-os
+4 -8
View File
@@ -90,16 +90,12 @@ echo "── 3. Cross-contract consistency ──"
# Check that contracts referencing each other have correct names
if [ -f "infrastructure-control.prose.md" ]; then
# Any contract that claims to check "all 5 PVE nodes" should name them
# Any contract that claims to check "all 6 PVE nodes" should name them
for f in *.prose.md; do
[ -f "$f" ] || continue
if grep -q "5-node\|5 node\|5 Proxmox\|all.*PVE.*node" "$f" 2>/dev/null; then
for node in amdpve minipve storepve acerpve ocupve; do
grep -q "$node" "$f" || {
echo " ⚠️ $f: references 5 nodes but '$node' not mentioned"
WARNINGS=$((WARNINGS + 1))
}
done
if grep -qE "\b5-node\b|\b5 node\b|\b5 Proxmox\b|all 5 PVE" "$f" 2>/dev/null; then
echo " ⚠️ $f: still references 5-node cluster (migrated to 6 nodes 2026-07-20)"
WARNINGS=$((WARNINGS + 1))
fi
done
fi
+3 -3
View File
@@ -16,7 +16,7 @@ Runs every 15 minutes in the background. Also triggers on session start.
## Requires
- **Zulip API key** for `abiba-bot@chat.sysloggh.net` in `$ZULIP_API_KEY`
- **SSH access** to Tanko (192.168.68.122), Mumuni (192.168.68.123), and Agent Zero Docker host (192.168.68.14)
- **SSH access** to Tanko (192.168.68.122), Mumuni (192.168.68.123, lxc/114 on hwepve since 2026-07-20), and Agent Zero Docker host (192.168.68.14)
- **PM2** on localhost for pi process management
- **Network access** to `chat.sysloggh.net`, `localhost:9200`
- **Write access** to `/root/zulip-health-monitor.log` and `/tmp/zulip-monitor-debounce`
@@ -199,7 +199,7 @@ Check `platforms.zulip.state`: `connected` ✅ | `disconnected` ❌ | `error`
ssh root@<CT> "ps aux | grep 'gateway run' | grep -v grep"
```
Gateway PID should exist with uptime > 60s.
Gateway PID should exist with uptime > 60s. **Dual-gateway detection**: if more than one `gateway run` process is found, the gateway has a collision (typically one `--force` and one `--replace` process). Kill the newer/duplicate process, then restart the remaining gateway: for Mumuni (PM2-managed), use `pm2 restart mumuni-zulip`; for Tanko (systemd/non-PM2), use `hermes gateway restart`. Check gateway log for "Gateway running with 2 platform(s)" (not 1) to confirm Zulip reloaded.
**B3: Heartbeat Verification**
@@ -222,7 +222,7 @@ ssh root@<CT> "grep -E 'Finalized|Failed to finalize|Replied to' ~/.hermes/logs/
| Condition | Action |
|-----------|--------|
| `zulip.state != "connected"` | `ssh root@<CT> "pkill -f 'gateway run'; sleep 2; hermes gateway restart"` |
| `zulip.state != "connected"` | Mumuni (PM2): `ssh root@<CT> "pm2 restart mumuni-zulip"`<br>Tanko (systemd): `ssh root@<CT> "hermes gateway restart"` |
| No heartbeat in 10min | Same as above |
| `Failed to finalize` > 50% | Check PATCH API, Zulip server |
| Response empty/short | Check A2A endpoint / LiteLLM model |