Compare commits

..
Author SHA1 Message Date
root 2b9b545ca9 no-mistakes(document): Fixed 2 stale model-name references in doc files
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-07-23 09:10:43 +00:00
root e6c52bf071 no-mistakes(review): Fix F1 model rename and F2 128K compaction math in contracts 2026-07-23 09:04:51 +00:00
root 1c44bf1259 contract updates: ornith decommissioned, 256K→128K context, Mumuni Discord disabled
Change 1: Strix Halo — ornith decommissioned
- gpu-fleet: Genesis Hermes V3 APEX → qwen3.6-35B-udq4 throughout
- inference-optimization: ornith→strix-moe/qwen3.6-35B-udq4
- gpu-monitor: ornith status → Strix Halo status
- infrastructure-control: Strix Halo — ornith → qwen3.6-35B-udq4
- infrastructure-update: ornith→strix-moe via router
- proxmox-monitor: Strix Halo LLM (ornith) → (qwen3.6-35B-udq4, strix-moe)

Change 2: GPU context 256K→128K fleet-wide
- hermes-agent-baseline: frontmatter description updated
- litellm-health: GPU Fleet Topology table 256K→128K
- litellm-self-heal: GPU Fleet Topology, engine flags, VRAM alert
- inference-optimization: compress threshold 256K→128K compact at 85K
- gpu-fleet: instability note updated

Change 3: Mumuni Discord platform disabled
- gpu-fleet: Mumuni platforms: removed discord
2026-07-23 09:00:42 +00:00
jerome fc88265e76 Merge pull request 'tune: switch compression model from strix-moe to syslog-auto' (#25) from tune/compression-syslog-auto into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
Reviewed-on: #25
2026-07-19 01:32:44 +00:00
jerome 51da19d92d Merge pull request 'feat(contracts): add infrastructure-maintenance contract' (#24) from fm/infra-maint-contract into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
Reviewed-on: #24
2026-07-19 01:32:29 +00:00
root 6570fd60e7 tune: switch compression model from strix-moe to syslog-auto
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
Relieve Strix Halo pressure by distributing compression across the
syslog-auto weighted pool (55% RTX 3090, 30% Strix Halo, 15% RTX 5070).

Updates:
- compression.model: strix-moe -> syslog-auto
- auxiliary.compression.model: strix-moe -> syslog-auto
- Rule 7: Updated for syslog-auto compression, removed aux prohibition
- Rule 8: Updated GPU workload distribution
- All docs/comments updated to reflect the change
2026-07-18 23:28:07 +00:00
jerome 5d2ecbace6 Merge branch 'master' into fm/infra-maint-contract
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
2026-07-18 23:12:20 +00:00
root 99789a00a1 no-mistakes(document): Reframed infra-maintenance contract as deliberate partition
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
2026-07-18 22:16:09 +00:00
root a68879e904 no-mistakes(review): fix docker rollback command to pin digest then compose up 2026-07-18 22:07:15 +00:00
jerome 06d2bcbc9e Merge pull request 'fix: fleet config issues from 2026-07-18 relay review' (#23) from fix/fleet-config-issues-20260718 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 0s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
Reviewed-on: #23
2026-07-18 22:06:48 +00:00
jerome c4a8c45835 Merge pull request 'zulip-resilience: fleet-wide audit findings and fixes 2026-07-18' (#22) from feat/zulip-resilience-audit-20260718 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
Reviewed-on: #22
2026-07-18 22:06:34 +00:00
root 1d027f71f6 feat(contracts): add infrastructure-maintenance contract
New responsibility contract consolidating host-level system maintenance
and Docker image lifecycle management, filling the gap left by
infrastructure-update (which owns cluster-wide apt waves).

Scope:
- OS package updates on primary host with pre-update snapshot/backup check
- Docker image pulls for LiteLLM, SearXNG, and other running containers
- Container restarts with per-stack health verification
- Post-update verification: LiteLLM, SearXNG, Zulip, Gitea, PM2, Hermes gateways
- Rollback on failure (image/apt/config restore) with circuit breaker

Owner: ops (firstmate secondmate). Trigger: weekly Sunday 2am ET.
Escalation: warning/critical->abiba+mumuni, fatal->abiba+mumuni+kwame.
circuit_breaker: max_retries 2, window 7200, trip_action escalate_to_fatal.
depends_on: infrastructure-monitoring (pre-update health baseline).

Registry:
- Add infrastructure-maintenance to by_category.maintenance, by_domain.infrastructure,
  by_owner.ops (new), by_trigger.scheduled, by_sensitivity.high
- Add 'ops' to owners list
- Move infrastructure-update owner abiba -> ops (in contracts entry + by_owner index)

Also adds '## Maintaining this file' section to AGENTS.md per fm-ensure-agents-md.
2026-07-18 22:02:45 +00:00
jerome a74229ee74 Merge branch 'master' into fix/fleet-config-issues-20260718
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-07-18 18:06:19 +00:00
root aebc98ead6 fix: remove trailing whitespace from litellm-api-keys.prose.md
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
2026-07-18 17:48:52 +00:00
root 17a77e6b3f fix: fleet config issues from 2026-07-18 relay review
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
- hermes-agent-baseline: add gpu-dense and gpu-light to models list
- hermes-config-template: fix pgrep traps (exclude infisical wrapper) + add
  vault empty-key guard documentation in Rule 13
- litellm-api-keys: fix pgrep pattern in auditable check
- scripts/agent-health-check: fix pgrep to exclude infisical bash wrapper

Addresses issues found during relay inbox resolution session:
1. pgrep -f 'hermes_cli.main gateway run' matches both the real python
   gateway and the infisical bash wrapper, causing false health readings
2. infisical vault stores empty key silently — no guard/monitoring
3. gpu-dense/gpu-light stable aliases missing from baseline config
2026-07-18 17:39:46 +00:00
jerome 14d27a09b5 Merge pull request 'gpu-self-heal: refresh to current fleet baseline and topology' (#21) from feat/gpu-self-heal-refresh-20260718 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
Reviewed-on: #21
2026-07-18 08:36:37 +00:00
17 changed files with 411 additions and 95 deletions
+7
View File
@@ -128,3 +128,10 @@ safe-mutate --verify "CMD" [--expect "PATTERN"] --mutate "CMD" [--reason "WHY"]
Read the [Authoring Guide](docs/AUTHORING-GUIDE.md) before writing any new contract.
It covers the full process: verify → draft → lint → review → ship, with templates
and style rules.
## Maintaining this file
Keep this file for knowledge useful to almost every future agent session in this project.
Do not repeat what the codebase already shows; point to the authoritative file or command instead.
Prefer rewriting or pruning existing entries over appending new ones.
When updating this file, preserve this bar for all agents and keep entries concise.
Symlink
+1
View File
@@ -0,0 +1 @@
AGENTS.md
+104 -2
View File
@@ -22,6 +22,7 @@ owners:
- abiba
- mumuni
- kwame
- ops
trigger_types:
- scheduled
- event_driven
@@ -58,6 +59,7 @@ index:
- memory-audit-maintenance
- gpu-fleet
- infrastructure-update
- infrastructure-maintenance
reference:
- infrastructure-control
- ra-h-os-custodianship-contract
@@ -90,6 +92,7 @@ index:
- infrastructure-control
- infrastructure-monitoring
- infrastructure-update
- infrastructure-maintenance
- pm2-self-heal
- disk-gc-threat-response
gpu:
@@ -130,7 +133,6 @@ index:
- build-zulip-plugin
- stirling-pdf-agent-access
- gpu-fleet
- infrastructure-update
- infrastructure-control
- zulip-adapter-lessons
- pi-approval-architecture
@@ -144,6 +146,9 @@ index:
- mumuni-delegation
kwame:
- hello-world
ops:
- infrastructure-maintenance
- infrastructure-update
by_trigger:
scheduled:
- hermes-key-enforcement
@@ -156,6 +161,7 @@ index:
- litellm-health
- memory-audit-maintenance
- infrastructure-update
- infrastructure-maintenance
event_driven:
- litellm-self-heal
- pm2-self-heal
@@ -197,6 +203,7 @@ index:
- hermes-zulip-plugin
- build-zulip-plugin
- infrastructure-update
- infrastructure-maintenance
- ra-h-os-custodianship-contract
- mumuni-delegation
normal:
@@ -1256,13 +1263,108 @@ contracts:
last_run: null
last_status: null
drift_alerts: []
- name: infrastructure-maintenance
file: infrastructure-maintenance.prose.md
kind: responsibility
category: maintenance
sensitivity: high
status: active
owner: ops
version: 1.0.0
trigger:
type: scheduled
cadence: 0 2 * * 0
description: Weekly host-level maintenance Sunday at 2am ET (replaces infrastructure-update
build-phase role; infra-update moves to ops)
cron_job_id: null
execution:
agent: ops
timeout: 3600
requires:
- infrastructure-monitoring run within last 30 minutes (pre-update health baseline)
- Proxmox snapshot of primary host OR /tmp backup dir created this run
- LiteLLM master key from Infisical vault for health verification
protocol:
- Load contract from prose-contracts/main
- Phase 0 preflight — capture health baseline, backup check, record image baseline
- Phase 1 apt update && apt upgrade -y on primary host
- Phase 2 docker compose pull for LiteLLM, SearXNG, and other running containers
- Phase 3 restart stacks one at a time with per-stack health verification
- Phase 4 verify every critical service (LiteLLM, SearXNG, Zulip, Gitea, PM2, Hermes gateways)
- On failure — rollback per protocol, escalate, do not loop beyond circuit breaker
- Log actions to ~/.hermes/runs/infrastructure-maintenance/
verification:
postconditions:
- check: all critical services running after update
verify: 'curl -sf http://192.168.68.116/litellm/v1/models && curl -sf https://chat.sysloggh.net/api/v1/server_settings && curl -sf https://git.sysloggh.net/api/v1/version && curl -sf http://192.168.68.7:8888 && pm2 jlist'
expect: all probes 200 OK / processes online
- check: no regressions from pre-update health baseline
verify: diff Phase 0 health-baseline against Phase 4 results
expect: no GREEN service turned RED
- check: docker containers on latest stable tags
verify: docker inspect --format '{{.Config.Image}}' <container> per service matches image-baseline.pulled_tag
expect: all containers running pulled tags
- check: APT packages up to date with no held broken packages
verify: apt list --upgradable 2>/dev/null | wc -l and apt-get -s upgrade | grep -ci broken
expect: upgradable == 0, broken == 0
artifact: maintenance run report with phase results and any rollback/escalation
verify_commands:
- curl -sf http://192.168.68.116/litellm/v1/models
- curl -sf http://192.168.68.7:8888
- curl -sf https://chat.sysloggh.net/api/v1/server_settings
- curl -sf https://git.sysloggh.net/api/v1/version
- pm2 jlist
receipt:
format: json
storage: ~/.hermes/runs/infrastructure-maintenance/
graph_node: true
schema:
contract: string
run_id: string
timestamp: ISO 8601
agent: string
status: pass|fail|escalated
phase: preflight|apt|images|restarts|verify|rollback|done|failed
actions_taken: array
postconditions: array
drift_alerts: array
evidence_path: string
escalation:
info:
action: log_to_receipt
notify: []
warning:
action: relay_alert
notify:
- abiba
- mumuni
critical:
action: relay_alert
notify:
- abiba
- mumuni
fatal:
action: relay_alert + pause + human_required
notify:
- abiba
- mumuni
- kwame
circuit_breaker:
max_retries: 2
window: 7200
trip_action: escalate_to_fatal
depends_on:
- infrastructure-monitoring
last_run: null
last_status: null
drift_alerts: []
- name: infrastructure-update
file: infrastructure-update.prose.md
kind: responsibility
category: maintenance
sensitivity: high
status: active
owner: abiba
owner: ops
version: 1.0.0
trigger:
type: scheduled
+22 -35
View File
@@ -9,10 +9,9 @@ description: >
gpu-dense, gpu-light. These never change — only the underlying model does.
Strix Halo: strix-moe → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB).
RTX 5070: gemma-4-12b Q4_K_M → IQ4_NL + MTP draft (122 tok/s, 2x faster).
UPDATED 2026-07-17: Strix Halo model swapped to Genesis Hermes V3 APEX (LuffyTheFox, 24GB, uncensored,
Hermes agent fine-tune, tensor repair, multimodal with mmproj).
RTX 5070 swapped to HauhauCS Gemma4-12B QAT Uncensored Balanced (Q4_K_M, 87 tok/s, 0/465 refusals).
Instability observed near 100K at 256K. 128K is the stable ceiling.
UPDATED 2026-07-17: Context reduced fleet-wide from 256K to 128K for stability.
Strix Halo model swapped to qwen3.6-35B-udq4 (22GB, strix-moe alias).
Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling.
For larger context needs → fall back to external providers (deepseek).
VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%.
agent: abiba
@@ -87,31 +86,19 @@ When a model is swapped on a GPU, ONLY the infrastructure layer changes — agen
| Alias | GPU | Current Model | Will Route To |
|-------|-----|---------------|---------------|
| `strix-moe` | Strix Halo (.15) | qwen3.6-35B-udq4 | Whatever runs on Strix Halo |
| `gpu-dense` | RTX 3090 (.8) | qwen3.6-27B-code (ThinkingCap) | Whatever runs on RTX 3090 |
| `gpu-dense` | RTX 3090 (.8) | qwen3.6-27B-code | Whatever runs on RTX 3090 |
| `gpu-light` | RTX 5070 (.110) | gemma-4-12b | Whatever runs on RTX 5070 |
**Backward compatibility**: Old model-specific names (qwen3.6-27B-code, gemma-4-12b, qwen3.6-35B-udq4) still work
but are deprecated for agent configs. Only the stable aliases survive model swaps.
## Current Model Assignments (2026-07-17)
## Current Model Assignments (2026-07-15)
| Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status |
|-------|-----|------|------|-----|----------|----------|-------------|--------|
| qwen3.6-27B-code (ThinkingCap) | RTX 3090 | .8 (llm-gpu) | ~20.9/24.6GB (85%) | **128K** | turbo4 | 1 | default | ✅ 68 tok/s |
| gemma-4-12b (HauhauCS QAT) | RTX 5070 | .110 (ocu-llm) | ~10.0/12.2GB (82%) | 128K | q4_0 | 1 | 2048/1024 | ✅ 87 tok/s |
| Genesis Hermes V3 APEX | Strix Halo Vulkan | .15 (amdpve) | ~10GB/64GB | 128K | q4_0 | 1 | 4096/1024 | ✅ 65 tok/s |
> **RTX 5070 model swap (2026-07-17)**: Switched from `gemma-4-12b-it-IQ4_NL` (Unsloth, 191 tok/s)
> to `HauhauCS/Gemma4-12B-QAT-Uncensored-HauhauCS-Balanced` (Q4_K_M QAT, 87 tok/s).
> Trade: 54% slower generation for QAT quality, 0/465 refusals, and agent-optimized tuning.
> MTP draft also swapped: Q8_0 (444MB) → tuned draft (242MB), saving 200MB VRAM.
> Role unchanged: gpu-light (vision, web extract, light auxiliary tasks).
> **RTX 3090 model swap (2026-07-17)**: Switched from `Qwopus3.6-27B-v2-MTP-Q4_K_M` (63 tok/s)
> to `bottlecapai/ThinkingCap-Qwen3.6-27B` (Q4_K_M QAT, 68 tok/s).
> RL-finetuned: 50% fewer thinking tokens, MMLU-Pro 0.85 vs 0.83 base.
> Self-spec MTP REQUIRED (crashes without it on turboquant build).
> Added vision via mmproj (0.9GB). VRAM 85%.
| qwen3.6-27B-code (MTP) | RTX 3090 | .8 (llm-gpu) | ~17/24.6GB (70%) | **128K** | turbo4 | 2 | default | ✅ 63 tok/s |
| gemma-4-12b | RTX 5070 | .110 (ocu-llm) | ~7.8/12.2GB (65%) | 128K | q4_0 | 2 | 2048/1024 | ✅ healthy |
| qwen3.6-35B-udq4 | Strix Halo Vulkan | .15 (amdpve) | ~22GB/64GB | 128K | q4_0 | 1 | 4096/1024 | ✅ 65 tok/s |
## Routing Configuration (LiteLLM — July 2026)
@@ -120,7 +107,7 @@ but are deprecated for agent configs. Only the stable aliases survive model swap
| Model | GPU | Weight | RPM Cap | Timeout |
|-------|-----|--------|---------|---------|
| qwen3.6-27B-code | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** |
| Genesis Hermes V3 APEX | Strix Halo (.15:8080) | **0.30** | 60 | **300s** |
| qwen3.6-35B-udq4 | Strix Halo (.15:8080) | **0.30** | 60 | **300s** |
| gemma-4-12b | RTX 5070 (.110:8080) | **0.15** | 200 | **120s** |
Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) is NOT in the inference path.
@@ -129,17 +116,17 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`.
| Model | RPM Cap | Notes |
|-------|---------|-------|
| strix-moe (Hermes V3) | 40 | Tight cap — prevents Strix overload |
| qwen3.6-27B-code | 500 | ThinkingCap Q4_K_M + MTP self-spec + vision, 68 tok/s |
| gemma-4-12b | 500 | HauhauCS QAT Uncensored Balanced + MTP, 87 tok/s |
| strix-moe (qwen3.6-35B-udq4) | 40 | Tight cap — prevents Strix overload |
| qwen3.6-27B-code | 500 | High cap — primary workhorse |
| gemma-4-12b | 500 | High cap — IQ4_NL+MTP, 122 tok/s |
### Stable Aliases (for agent configs — never change)
| Alias | RPM Cap | Routes To | Purpose |
|-------|---------|-----------|---------|
| `strix-moe` | 40 | Strix Halo | Compression tasks (MoE models) |
| `gpu-dense` | 500 | RTX 3090 (ThinkingCap) | Heavy reasoning, code gen, delegation |
| `gpu-light` | 500 | RTX 5070 (HauhauCS QAT) | Vision, web extract, light tasks |
| `gpu-dense` | 500 | RTX 3090 | Heavy reasoning |
| `gpu-light` | 500 | RTX 5070 | Vision, web extract, light tasks |
### Fallback Chains
- gemma → qwen
@@ -261,10 +248,10 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
- **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround.
- **VRAM (2026-07-15)**: RTX 3090 at ~17/24.6GB (~70%) with **128K context** (reduced from 256K 2026-07-17). RTX 5070 at ~7.8/12.2GB (~65%) with 128K context + MTP. Strix Halo at ~7GB/64GB.
- **RTX 3090 runs `--parallel 2`** with MTP draft (spec-type draft-mtp, spec-draft-n-max 2).
- **RTX 3090 config**: `-c 131072 -ctk turbo4 -ctv turbo4 --parallel 1 --flash-attn on --cont-batching --spec-type draft-mtp --spec-draft-n-max 4`. ThinkingCap Qwen3.6-27B Q4_K_M (15.7GB) + mmproj (0.9GB) + MTP self-spec. VRAM: ~85%. Service: `/home/llmuser/llama-wrapper.sh`. ⚠️ MTP REQUIRED for stability on this turboquant build — model segfaults without `--spec-type draft-mtp`. Outputs reasoning_content (hidden from agent, improves answer quality).
- **RTX 5070 config (2026-07-17)**: HauhauCS Gemma4-12B QAT Uncensored Balanced (Q4_K_M) + tuned MTP draft (242MB) at 128K context, single slot. Gen speed: 87 tok/s (vs 191 IQ4_NL). VRAM: ~10.0/12.2GB (~82%). Service: `/home/llmuser/llama-wrapper.sh`. Model: `Gemma4-12B-QAT-Uncensored-HauhauCS-Balanced-Q4_K_M.gguf`, MTP: `mtp-gemma-4-12B-it.gguf`, mmproj: `mmproj-Gemma4-12B-QAT-Uncensored-HauhauCS-Balanced-BF16.gguf`. Recommended sampling: temp 0.6, top_k 64, top_p 0.9, min_p 0.05, repeat_penalty 1.1.
- **RTX 3090 config**: `-c 131072 -ctk turbo4 -ctv turbo4 --parallel 2 --flash-attn on --cont-batching --spec-type draft-mtp`. Context reduced to 128K (2026-07-17, was 256K). VRAM: ~70%. Service: `/home/llmuser/llama-wrapper.sh`.
- **RTX 5070 config (2026-07-15)**: Switched to IQ4_NL + MTP draft (Q8_0) at 128K context. Gen speed: 122 tok/s. VRAM: ~7.8/12.2GB (~65%). Service: `/home/llmuser/llama-wrapper.sh`. Config: `--model gemma-4-12b-it-IQ4_NL.gguf --spec-draft-model gemma-4-12b-it-Q8_0-MTP.gguf --spec-type draft-mtp --spec-draft-n-max 4 --ctx-size 131072`.
- **LiteLLM timeout tuning (verified 2026-07-16 against `/opt/inference-harness/litellm_config.yaml` on CT 116)**: gemma-4-12b 120s, qwen3.6-27B-code 300s, qwen3.6-35B-udq4 300s, strix-moe 300s, syslog-auto routes all 300s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s.
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V3-GGUF` (APEX quant), alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded). Hermes agent fine-tune, tensor repair (SSM layers fixed via SVD), uncensored (0/465 refusals).
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `qwen3.6-35B-udq4`, alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded).
- **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill).
- **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (Mumuni) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`.
- **Port 8080 firewall**: amdpve iptables restricts 8080 to 192.168.68.116 (LiteLLM/router host) only. All inbound connections are from .116 (LiteLLM proxied via nginx). Localhost curls hang (SYN dropped). Always test from .116.
@@ -278,11 +265,11 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
| GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Context |
|-----|-------|-----------|--------------|----------|---------|
| RTX 3090 (.8) | ThinkingCap-Qwen3.6-27B Q4_K_M | **68** | — | — | **128K** |
| RTX 5070 (.110) | HauhauCS QAT Uncensored Balanced | **87** | — | — | **128K** |
| Strix Halo (.15) | Genesis Hermes V3 APEX | **65** | 140 | — | **128K** |
| RTX 3090 (.8) | qwen3.6-27B-code (MTP) | **63** | — | — | **128K** |
| RTX 5070 (.110) | gemma-4-12b (IQ4_NL+MTP) | **191** | — | — | **128K** |
| Strix Halo (.15) | qwen3.6-35B-udq4 | **65** | 140 | — | **128K** |
Benchmarks from 2026-07-17. Strix Halo swapped to Genesis Hermes V3 APEX (LuffyTheFox). RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
Benchmarks from 2026-07-17. Strix Halo model: qwen3.6-35B-udq4. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
All 3 GPUs now at 128K context (2026-07-17, reduced from 256K for stability).
Benchmarks run through LiteLLM proxy (192.168.68.116:4001) every 5 minutes.
@@ -328,7 +315,7 @@ Mumuni (CT114, 192.168.68.123) is the primary business assistant. This profile i
| `compression.protect_last_n` | 40 | Preserves last 40 messages |
| `memory.memory_char_limit` | 800 | Brief memory entries |
| `personalities` | `creative` | Creative assistant personality |
| Platforms | cli, discord, homeassistant, signal, telegram, zulip | All Hermes platforms |
| Platforms | cli, homeassistant, signal, telegram, zulip | All Hermes platforms |
| Main model timeout | 300s | LiteLLM global timeout |
| Compression model timeout | 300s | strix-moe timeout increased from 120s |
+1 -1
View File
@@ -47,7 +47,7 @@ poll .15:8080 directly; must go through router on .116.
| Router (unified) | `http://192.168.68.116/health/unified` | 15s | models, CB, scores, GPU status |
| Router (basic) | `http://192.168.68.116/health` | 15s | basic aliveness |
| LiteLLM | `http://192.168.68.116/litellm/health` | 15s | proxy health, model count |
| Strix Halo | `http://192.168.68.116/health/unified` (router) | 15s | ornith status via router — cannot poll .15:8080 directly (firewalled to .116 only) |
| Strix Halo | `http://192.168.68.116/health/unified` (router) | 15s | Strix Halo status via router — cannot poll .15:8080 directly (firewalled to .116 only) |
| Dashboard | `http://192.168.68.116/dashboard/` | 15s | harness-dashboard aliveness |
### Alert Delivery
+2 -2
View File
@@ -46,7 +46,7 @@ depends_on:
|-------|-----|------|-------|------|-----|-------|------|
| `gpu-dense` | RTX 3090 24GB | ct8 (.8:8080) | ThinkingCap Qwen3.6-27B Q4_K_M + MTP + vision | 21.6/24.6GB (88%) | 128K | 74.9 | Heavy reasoning, code gen |
| `gpu-light` | RTX 5070 12GB | ct110 (.110:8080) | HauhauCS Gemma4-12B QAT Q4_K_M + MTP draft | 10.1/12.2GB (83%) | 128K | 169.6 | Vision, web extract, light tasks |
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | Genesis Hermes V3 APEX (LuffyTheFox, 24GB) | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs |
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | qwen3.6-35B-udq4 | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs |
Key notes:
- All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) is deprecated and NOT in the inference path.
@@ -143,7 +143,7 @@ Key notes:
- **Detect**: Benchmark tok/s vs baseline for each GPU at current context (all 128K)
- RTX 3090 (128K ctx, ThinkingCap): baseline 74.8 tok/s — currently at 74.9 (100%)
- RTX 5070 (128K ctx, HauhauCS QAT): baseline 165.2 tok/s — currently at 169.6 (103%)
- Strix Halo (128K ctx, Genesis Hermes V3): baseline 70.5 tok/s — currently at 62.9 (89%)
- Strix Halo (128K ctx, qwen3.6-35B-udq4): baseline 70.5 tok/s — currently at 62.9 (89%)
- **Fix**:
- If tok/s > baseline → context has headroom, consider increasing
- If tok/s < 90% baseline → reduce context by 25% and retest
+3 -1
View File
@@ -5,7 +5,7 @@ version: 1.0.0
description: >
Canonical known-good baseline for all Syslog Hermes agents. Captures the exact
configuration state, keys, workarounds, and audit procedure. When an agent's
configuration goes sideways, restore from this baseline. Last verified 2026-07-16. All GPUs 256K context (RTX 3090 .8, RTX 5070 .110, Strix Halo .15). Parallel 1 fleet-wide (Strix Halo handles compression solo).
configuration goes sideways, restore from this baseline. Last verified 2026-07-16. All GPUs 128K context (reduced from 256K for stability Jul 2026) (RTX 3090 .8, RTX 5070 .110, Strix Halo .15). Parallel 1 fleet-wide (Strix Halo handles compression solo).
author: Abiba (pi agent)
---
@@ -193,6 +193,8 @@ Key is injected via `infisical run --` wrapper at PM2 startup:
"models": [
{ "id": "syslog-auto" },
{ "id": "strix-moe" },
{ "id": "gpu-dense" },
{ "id": "gpu-light" },
{ "id": "qwen3.6-27B-code" },
{ "id": "gemma-4-12b" }
]
+42 -24
View File
@@ -5,11 +5,14 @@ description: >
Standard Hermes configuration template for Syslog Solution LLC agents.
Enforces shared infrastructure setup (Firecrawl, SearXNG, local models,
RA-H OS MCP) while keeping agent-specific API keys and model choices.
UPDATED 2026-07-16: Compression model is the stable alias `strix-moe` (NOT `ornith-1.0-35b`,
UPDATED 2026-07-18: Compression model switched to `syslog-auto` (was `strix-moe`)
to relieve Strix Halo pressure. syslog-auto distributes compression across the
weighted pool (55% RTX 3090, 30% Strix Halo, 15% RTX 5070).
UPDATED 2026-07-16: Compression model was the stable alias `strix-moe` (NOT `ornith-1.0-35b`,
which LiteLLM does not serve). All 3 GPUs verified at 128K (reduced from 256K 2026-07-17 for stability).
Added Rule 12 (Context-Issue Diagnostic) + Rule 13 (.env fallback enforcement) from the
2026-07-16 Mumuni root-cause investigation (WAL #1300).
UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo. RTX 3090 context verified at 128K. Infisical .env fallback required (Rule 3/13).
UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo (later switched to syslog-auto 2026-07-18). RTX 3090 context verified at 128K. Infisical .env fallback required (Rule 3/13).
---
## Maintains
@@ -130,7 +133,9 @@ mcp_servers:
# ─── Compression ───
compression:
enabled: true
model: strix-moe # ⚠️ Must match auxiliary.compression.model. Stable alias (gpu-fleet § Stable Role-Based Aliases). NOT ornith-1.0-35b (LiteLLM does not serve that name).
model: syslog-auto # ⚠️ Switched from strix-moe 2026-07-18 to relieve Strix Halo.
# syslog-auto distributes across weighted pool (55% RTX 3090,
# 30% Strix Halo, 15% RTX 5070). All GPUs at 128K.
provider: harness
max_context_window: 131072 # MUST match actual GPU capacity. All 3 GPUs are 128K (Jul 17).
threshold: 0.65 # Fires at ~170K for 262K window, ~85K for 128K
@@ -145,8 +150,9 @@ compression:
# model: gpu-light # stable alias (NOT raw "gemma-4-12b")
# base_url: http://192.168.68.116/v1
# api_key_env: LITELLM_API_KEY
# Do NOT use syslog-auto for auxiliary tasks — it routes to the primary GPU.
# gpu-light = RTX 5070 (12B), freeing the Strix Halo for agent reasoning.
# Compression uses syslog-auto (switched from strix-moe 2026-07-18) to distribute
# load across the weighted pool and relieve Strix Halo pressure.
# Vision and web_extract use gpu-light = RTX 5070 (12B).
# Heavy aux (delegation, x_search) use gpu-dense (RTX 3090) instead.
# NEVER use raw model names (gemma-4-12b, qwen3.6-27B-code, qwen3.6-35B-udq4)
# in agent configs — use the stable aliases so model swaps don't break agents.
@@ -166,7 +172,7 @@ auxiliary:
timeout: 30
compression:
provider: harness
model: strix-moe # MUST match compression.model above. Stable alias for Strix Halo.
model: syslog-auto # Switched from strix-moe 2026-07-18. Relieves Strix Halo pressure.
base_url: http://192.168.68.116/v1 # Rule 5: /v1 NOT /litellm/v1
api_key_env: LITELLM_API_KEY
timeout: 300 # gpu-fleet: 300s for large-history summarization (was 60)
@@ -247,29 +253,30 @@ The following MUST be identical across ALL profiles:
- Apply to BOTH main config AND all sub-agent profiles
- For agents needing longer outputs: raise to 8192, but never omit
### Rule 7: Auxiliary Model Consistency (UPDATED 2026-07-16)
- Vision and web_extract use `gemma-4-12b` (RTX 5070 — 12GB, vision-optimized)
- Compression uses `strix-moe` (stable alias for Strix Halo — 64GB, 128K ctx, compression-optimized)
- **`strix-moe` is the only valid compression model name** — LiteLLM does NOT serve `ornith-1.0-35b`
(it serves `strix-moe`, `qwen3.6-35B-udq4`, `gpu-dense`, `gpu-light`, `syslog-auto`, `gemma-4-12b`, `qwen3.6-27B-code`). Old configs with `ornith-1.0-35b` cause 403/model-not-found on compression calls.
### Rule 7: Auxiliary Model Consistency (UPDATED 2026-07-18)
- Vision and web_extract use `gpu-light` (stable alias, RTX 5070 — 12GB, vision-optimized)
- Compression now uses `syslog-auto` (switched from `strix-moe` 2026-07-18) to distribute
compression load across the weighted pool (55% RTX 3090, 30% Strix Halo, 15% RTX 5070).
This relieves Strix Halo pressure while keeping compression functional on all GPUs.
- **`syslog-auto` is the valid compression model** — LiteLLM serves it as the weighted pool.
Old configs with `strix-moe` for compression should be updated to `syslog-auto`.
- All auxiliary services MUST use identical routing:
- `base_url: http://192.168.68.116/v1` (Rule 5: `/v1`, NOT `/litellm/v1`)
- `api_key_env: LITELLM_API_KEY`
- **Do NOT use `syslog-auto`** for auxiliary tasks — it routes unpredictably
- **Compression on Strix Halo**: The strix-moe alias routes to Strix Halo
(64GB UMA, 128K context) — the designated compression GPU. This frees the
RTX 5070 for vision and web search, and the RTX 3090 for heavy reasoning.
- **Compression via syslog-auto**: Routes through the weighted pool. Strix Halo still handles
~30% of compression calls (at 60 RPM via pool vs 40 RPM direct), but the bulk (55%)
goes to RTX 3090 which has ample spare capacity.
- The `compression:` block's `model` MUST match `auxiliary: compression: model`
- The `compression: max_context_window: 131072` MUST match actual GPU capacity (128K)
### Rule 8: GPU Workload Distribution (UPDATED 2026-07-16)
- **RTX 3090 (24GB, 128K ctx, qwen3.6-27B-code)**: Heavy reasoning, code gen, long conversations
- **RTX 5070 (12GB, 128K ctx, gemma-4-12b)**: Vision, web search, quick tasks, web_extract (IQ4_NL+MTP, ~65% VRAM at 128K)
- **Strix Halo (64GB, 128K ctx, strix-moe)**: Context compression, summarization, long docs
### Rule 8: GPU Workload Distribution (UPDATED 2026-07-18)
- **RTX 3090 (24GB, 128K ctx, qwen3.6-27B-code)**: Heavy reasoning, code gen, long conversations — also handles ~55% of compression via syslog-auto pool
- **RTX 5070 (12GB, 128K ctx, gemma-4-12b)**: Vision, web search, quick tasks, web_extract — handles ~15% of compression via syslog-auto pool
- **Strix Halo (64GB, 128K ctx, qwen3.6-35B-udq4)**: Agent reasoning, compression (~30% via syslog-auto pool), fallback for other GPUs
- Agent profiles MUST route auxiliary tasks to the correct GPU:
- `auxiliary.vision.model: gemma-4-12b` (RTX 5070)
- `auxiliary.web_extract.model: gemma-4-12b` (RTX 5070)
- `auxiliary.compression.model: strix-moe` (Strix Halo)
- `auxiliary.vision.model: gpu-light` (RTX 5070)
- `auxiliary.web_extract.model: gpu-light` (RTX 5070)
- `auxiliary.compression.model: syslog-auto` (distributed pool, switched from strix-moe 2026-07-18)
- Default model (`model.default`) and custom_provider remain `syslog-auto` for auto-routing
- For 128K context window: `threshold: 0.65` (fires at ~85K tokens)
- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss
@@ -325,7 +332,8 @@ verify ALL FOUR of these against the live config. They are the only root causes
One-line agent health check (run on the agent host):
```bash
PID=$(pgrep -f "python -m hermes_cli.main gateway run" | head -1)
# Use grep -v infisical to avoid matching the bash wrapper that contains the same string
PID=$(pgrep -f "python -m hermes_cli.main gateway run" | grep -v infisical | head -1)
cat /proc/$PID/environ | tr '\0' '\n' | grep ^LITELLM_API_KEY= | sed 's/=.*/<set>/'
curl -s -o /dev/null -w 'key_health: %{http_code}\n' -H "Authorization: Bearer $(cat /proc/$PID/environ | tr '\0' '\n' | grep ^LITELLM_API_KEY= | cut -d= -f2)" http://192.168.68.116/v1/models
```
@@ -351,9 +359,19 @@ directly (no infisical). Apply with `systemctl daemon-reload && systemctl restar
The wrapper sources `~/.hermes/.env` then exports `LITELLM_API_KEY="$<AGENT>_LITELLM_API_KEY"`.
See litellm-api-keys.prose.md § Machine Identity for Vault Writes for vault sync.
**⚠️ Vault empty-key guard:** If the vault stores the secret as an empty string,
the wrapper will inject an empty key and the gateway will silently get 401 errors
on all LiteLLM requests (triggering silent DeepSeek fallback). The `.env` fallback
is present but the vault takes precedence when the secret key exists (even if empty).
**Fix:** The wrapper MUST validate the key length after injection. If LITELLM_API_KEY
is empty or shorter than 20 chars, log a warning and either fail with a clear error
message or fall back to the `.env` value before starting the gateway.
**Verification (all agents):**
```bash
GP=$(pgrep -f "python -m hermes_cli.main gateway run" | head -1)
# Use grep -v infisical to avoid matching the bash wrapper that contains the same string
GP=$(pgrep -f "python -m hermes_cli.main gateway run" | grep -v infisical | head -1)
K=$(cat /proc/$GP/environ | tr '\0' '\n' | grep '^LITELLM_API_KEY=' | cut -d= -f2)
curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $K" http://192.168.68.116/v1/models # must be 200
```
+9 -10
View File
@@ -22,14 +22,14 @@ management, and prompt caching — without sacrificing agent capability.
- `agent-configs`: current config.yaml from each active Hermes agent (Mumuni
.123, any others on .129/.122) including compression, model, context_window,
prompt_caching, memory settings
- `gpu-health`: health check response from all 3 GPU backends (ornith .15:8080,
- `gpu-health`: health check response from all 3 GPU backends (strix-moe .15:8080,
qwen .8:8080, gemma .110:8080)
### Maintains
The optimized inference stack configuration — every change is applied and
verified end-to-end. Postcondition: avg request_duration_ms ≤ 15000 for 90% of
non-ornith traffic; ≤ 30000 for ornith-bound agentic calls.
inference calls.
#### liteLLM-routing
The syslog-auto routing weights, model-specific timeouts, RPM limits, and
@@ -53,18 +53,17 @@ duration.
### Strategies
**Context is the root cause.** Every ~46K prompt token costs ~87s of ornith
**Context is the root cause.** Every ~46K prompt token costs ~87s of
prefill time at 532 tok/s. Fix context first, routing second.
- **Route by task**: ornith for multi-step reasoning only; qwen for code/standard
queries; gemma for compression/auxiliary. Never send simple completion to a
35B MoE.
- **Compress aggressively**: threshold at 40% (not 65%) — a 256K window should
compact at 102K, not 166K. Target 15% tail (not 30%).
- **Route by task**: qwen for code/standard queries; gemma for
compression/auxiliary; strix-moe for compression tasks.
- **Compress aggressively**: threshold at 40% (not 65%) — a 128K window should
compact at 51K, not 85K. Target 15% tail (not 30%).
- **Cache everything repeated**: system prompts, skill docs, AGENTS.md — these
never change between turns. Single-digit cache hit rate is unacceptable.
- **Lower context ceiling**: 128K window is the stable ceiling for agent conversations.
GPUs reduced from 256K to 128K (2026-07-17). For larger contexts, route to external providers.
GPUs reduced from 256K to 128K (2026-07-17). 128K window should compact at 85K (0.65 threshold). For larger contexts, route to external providers.
### Shape
@@ -100,5 +99,5 @@ call enable-prompt-caching
call verify-latency
host: 192.168.68.116
models: [syslog-auto, qwen3.6-27B-code, gemma-4-12b, ornith-1.0-35b]
models: [syslog-auto, qwen3.6-27B-code, gemma-4-12b]
```
+1 -1
View File
@@ -199,7 +199,7 @@ description: >
**Prometheus targets**:
- 192.168.68.8:9400 (RTX 3090 — qwen)
- 192.168.68.110:9400 (RTX 5070 — gemma)
- 192.168.68.15:9400 (Strix Halo — ornith)
- 192.168.68.15:9400 (Strix Halo — qwen3.6-35B-udq4)
- 192.168.68.24:9401 (Router metrics exporter)
- harness-litellm:4000 (LiteLLM health)
+200
View File
@@ -0,0 +1,200 @@
---
kind: responsibility
name: infrastructure-maintenance
description: >
Weekly system-level maintenance for the Syslog inference fleet: OS package
updates on the primary host, Docker image pulls for LiteLLM/SearXNG and other
running containers, container restarts with health verification, post-update
verification of every critical service (LiteLLM proxy, SearXNG, Zulip, Gitea,
PM2 processes, Hermes gateways), and rollback on failure. Consolidates the
raw shell scripts that previously did this piecemeal. This contract owns the
HOST-LEVEL weekly maintenance loop on the primary host plus Docker image
pulls ONLY for .116 and .7, while infrastructure-update owns the FULL-FLEET
cluster-wide wave (apt across the full PVE cluster + CTs/VMs AND its Docker
image Wave 3 across all stacks). Runs Sunday 2am ET. Owner:
ops (firstmate secondmate). Blast radius: an unverified image pull can break
LiteLLM (all agents lose inference) or SearXNG (search-stack down); a bad apt
upgrade can leave the host in a half-upgraded state. Pre-update backup check
and rollback are mandatory for this reason.
agent: ops
triggers:
- weekly (Sunday 02:00 ET) via cron
- on demand when ops/abiba triggers "infra maintenance"
version: 1.0.0
---
## Maintains
- maintenance-status: { phase: idle|preflight|apt|images|restarts|verify|rollback|done|failed, host, step, result, timestamp }
- image-baseline: { service, current_tag, pulled_tag, digest, updated_at } — last known-good image per container
- apt-state: { upgradable_before, upgradable_after, held_broken, kernel_reboot_required }
- health-baseline: snapshot of critical-service health captured pre-update (used for regression check post-update)
- rollback-snapshot: { backup_path, configs, image_digests, timestamp } — restore point created in preflight
- maintenance-history: array of past runs with phase results and any escalations
## Scope
Primary host is the maintenance host where apt updates apply. Docker image pulls
span the two Docker ecosystems that run critical services. infrastructure-update
runs the full-fleet cluster-wide wave (including its Wave 3 Docker pulls across
all stacks/hosts); this contract runs a narrower host-level weekly pull limited
to .116 and .7. Topology, CT IDs, and IPs are live-state fields — verify against
`infrastructure-control.prose.md` (the source of truth) and the live system
before mutating.
| Host | IP | Role | Trust |
|------|----|------|-------|
| CT 116 (syslog-api) | 192.168.68.116 | LiteLLM proxy + Grafana + Prometheus (inference harness) | ⚠️ VERIFY-BEFORE-USE |
| VM 109 (docker-vm) | 192.168.68.7 | SearXNG + Firecrawl + home stack (Docker host) | ⚠️ VERIFY-BEFORE-USE |
| CT 117 (zulip) | 192.168.68.19 | Zulip (storepve bridge IP .19) | ⚠️ VERIFY-BEFORE-USE |
| Gitea | https://git.sysloggh.net | Prose-contracts + agent configs source control | ⚠️ VERIFY-BEFORE-USE |
| CT 100 (abiba/pi) | 192.168.68.24 | PM2 processes (pi agent harness) | ⚠️ VERIFY-BEFORE-USE |
> "Primary host" for the apt phase is the host the ops agent runs maintenance
> from. Confirm which host that is against infrastructure-control before
> running; do not assume. If the ops agent is containerized/CT-based, apt runs
> inside that CT.
## Requires
- SSH/exec access to CT 116 (.116) and VM 109 (.7) for Docker operations
- `apt`, `docker`, `docker compose` available on target hosts
- LiteLLM master key available (Infisical vault, `LITELLM_API_KEY`) for health verification
- `infrastructure-monitoring` run completed within the last 30 minutes — provides the pre-update health baseline used by the regression check
- Writable backup directory `/tmp/infra-maintenance-backup-<date>/` on each mutated host
- Proxmox snapshot of the primary host available (or confirmed not required) before apt phase
## Continuity
- Self-driven: weekly cron `0 2 * * 0` (Sunday 02:00 ET)
- Also wakes on: explicit "infra maintenance" trigger from ops/abiba
- Depends on `infrastructure-monitoring` for the pre-update health baseline — do not run if the last monitoring run is stale (>30 min) or RED; abort and escalate instead
## Execution
### Phase 0 — Preflight (snapshot/backup check + health baseline)
1. **Capture health baseline** — run the `infrastructure-monitoring` postcondition checks (LiteLLM, Zulip, Gitea, SearXNG, Proxmox API) and record results as `health-baseline`. If any critical service is already down, **abort**: maintenance must not run on a degraded fleet.
2. **Backup check** — confirm a Proxmox snapshot of the primary host exists OR `/tmp/infra-maintenance-backup-<date>/` was created this run. Snapshot critical config files into the backup dir:
- `/opt/inference-harness/docker-compose.yml`, `/opt/inference-harness/litellm_config.yaml` (CT 116)
- `/opt/search-stack/searxng/docker-compose.yml`, `/opt/search-stack/firecrawl-source/docker-compose.yaml` (VM 109)
3. **Record image baseline**`docker inspect --format '{{.Image}} {{.Config.Image}}' <container>` for every running container on .116 and .7; store digests in `image-baseline` so rollback can restore them.
4. **Disk check**`df -h` on each mutated host; abort if free space <20% (apt upgrade + image pulls need headroom).
### Phase 1 — OS package updates (primary host)
```bash
# On the primary host only (VERIFY host against infrastructure-control first)
apt update
apt upgrade -y
```
- Capture `apt list --upgradable` before and after → store in `apt-state`.
- If apt reports held/broken packages (`apt-get -s upgrade | grep -i broken`, or non-zero exit), **stop** — do not force. Record `held_broken` and go to rollback/escalate.
- If `/var/run/reboot-required` exists after upgrade, flag `kernel_reboot_required: true` in `apt-state` but **do not reboot automatically** — that's a separate coordinated action (see infra-update Wave 4). Note it in the report.
### Phase 2 — Docker image pulls
Pull latest stable tags for every running container. Do NOT pin to `:main`/`:nightly` — use stable tags where the compose file specifies them; otherwise `latest`.
```bash
# CT 116 (.116) — inference harness
cd /opt/inference-harness && docker compose pull
# VM 109 (.7) — search + home stacks
cd /opt/search-stack/searxng && docker compose pull
cd /opt/search-stack/firecrawl-source && docker compose pull
# any other running stacks on .7 (home stack, audiobookshelf) — pull per their compose files
```
- LiteLLM and SearXNG are the two explicitly required pulls; "any other running containers" means every stack with a compose file on .116 and .7.
- Record pulled tag + digest per service in `image-baseline`.
### Phase 3 — Container restarts with health verification
Restart one stack at a time, verify health before moving to the next. Do not restart everything at once — a failure mid-wave must leave the rest running.
```bash
# CT 116
cd /opt/inference-harness && docker compose up -d
# VM 109
cd /opt/search-stack/searxng && docker compose up -d
cd /opt/search-stack/firecrawl-source && docker compose up -d
```
After each stack comes up, wait for health (max 120s):
- `docker ps` shows the container `Up` (and `healthy` if a healthcheck is defined)
- Service-specific probe passes (see Phase 4 probes)
If a stack fails to come up within 120s, **stop the wave** and go to rollback for that stack only; do not proceed to the next.
### Phase 4 — Post-update service verification
After ALL updates (apt + images + restarts), verify every critical service is back up and matches the pre-update baseline. This is the regression gate.
| Service | Probe | Expect |
|---------|-------|--------|
| LiteLLM proxy | `curl -sf http://192.168.68.116/litellm/v1/models` | 200 OK, models returned |
| LiteLLM MCP gateway | `curl -sf http://192.168.68.116:4000/mcp-rest/tools/list -H "Authorization: Bearer $MASTER_KEY"` | 90 tools (23 RA-H OS + 67 GitHub) |
| SearXNG | `curl -sf http://192.168.68.7:8888` | 200 OK |
| Zulip | `curl -sf https://chat.sysloggh.net/api/v1/server_settings` | 200 OK |
| Gitea | `curl -sf https://git.sysloggh.net/api/v1/version` | 200 OK |
| PM2 processes | `pm2 jlist` (CT 100) | all pi-agent processes `online` |
| Hermes gateways | SSH to Mumuni CT 114, Tanko CT 112; `systemctl is-active hermes-gateway` | `active` for each |
Regression check: every service that was GREEN in `health-baseline` must still be GREEN. A service that was already RED (and caused a preflight abort) is excluded — but Phase 0 should have aborted before we got here.
## Rollback Protocol
If ANY service in Phase 4 fails to come back up (or regresses vs baseline):
1. **Image rollback** — for the failing stack, restore the previous image:
```bash
# Restore from recorded image-baseline digest
docker compose down
# Pin the service image to the recorded digest in compose, then recreate
# image: <name>@sha256:<previous_digest>
docker compose pull && docker compose up -d
```
2. **APT rollback** — restore the primary host from the Proxmox snapshot taken/confirmed in Phase 0. If no snapshot, `apt install <pkg>=<old_version>` per package using apt history (`/var/log/apt/history.log`).
3. **Config rollback** — restore configs from `/tmp/infra-maintenance-backup-<date>/`.
4. **Re-verify** — re-run the Phase 4 probes on the rolled-back service. If still failing, escalate (do not loop — circuit breaker below).
5. **Escalate** — send a Zulip DM to abiba + mumuni with: failing service, phase, baseline vs current, rollback actions taken, backup path.
## Circuit Breaker
- `max_retries: 2` per failing phase — after 2 rollback attempts on the same service, stop and escalate.
- `window: 7200` seconds — no more than 2 retries within a 2-hour window.
- `trip_action: escalate_to_fatal` — when tripped, escalate to fatal (abiba + mumuni + kwame) and pause; a human must clear before the next scheduled run.
## Report
After completion (or on abort), emit a receipt (JSON) to `~/.hermes/runs/infrastructure-maintenance/` and send a Zulip DM summary:
```
🛠 Infrastructure Maintenance — YYYY-MM-DD
Phase: apt | images | restarts | verify | rollback
Primary host: <host>
APT: <N> packages upgraded, <M> held/broken, kernel_reboot_required=<bool>
Images pulled: LiteLLM <tag>, SearXNG <tag>, <others>
Services: all GREEN | <service> FAILED (rolled back)
Baseline regression: none | <details>
Backup: /tmp/infra-maintenance-backup-YYYYMMDD/
Escalation: none | warning | critical | fatal
```
## Verification Postconditions
- All critical services running after update (Phase 4 all GREEN)
- No regressions from pre-update health baseline (Phase 0 baseline)
- Docker containers on latest stable tags (`image-baseline.pulled_tag` recorded)
- APT packages up to date with no held broken packages (`apt-state.held_broken == 0`)
## Related Contracts
- `infrastructure-update.prose.md` — owns the full-fleet cluster-wide wave INCLUDING its Wave 3 Docker image updates across all stacks (SearXNG, Firecrawl, Inference Harness on .116, home stack, audiobookshelf); infrastructure-maintenance is a deliberately narrower host-level weekly pull scoped to .116 and .7.
- `infrastructure-monitoring.prose.md` — provides the pre-update health baseline (depends_on).
- `infrastructure-control.prose.md` — topology source of truth (CT IDs, IPs, hostnames).
- `litellm-health.prose.md` — LiteLLM probe details.
- `proxmox-monitor.prose.md` — Docker stats + monitoring stack health.
+2 -2
View File
@@ -67,7 +67,7 @@ Before ANY update wave:
- All VMs/CTs running: check via Proxmox API
- LiteLLM healthy: `curl localhost:4000/health/liveliness` (via CT 116)
- LiteLLM MCP tools: `curl localhost:4000/mcp-rest/tools/list -H "Authorization: Bearer $MASTER_KEY"` → 90 tools
- GPU servers responding: check :8080 on VM 101, VM 103; check ornith via router (http://192.168.68.116/health/unified — .15:8080 is firewalled to .116 only)
- GPU servers responding: check :8080 on VM 101, VM 103; check strix-moe via router (http://192.168.68.116/health/unified — .15:8080 is firewalled to .116 only)
- Zulip agents connected: check Mumuni/Tanko gateway state
- Abiba PM2 processes online: `pm2 status`
@@ -124,7 +124,7 @@ Before Wave 1, snapshot these files:
/opt/home_stack/docker-compose.yml (VM 109 .7)
/opt/audiobookshelf/docker-compose.yml (VM 109 .7)
/root/.pi/agent/extensions/config.yaml (CT 100 .24)
/etc/systemd/system/ornith-server.service (amdpve .15)
/etc/systemd/system/ornith-server.service (amdpve .15 — strix-moe)
/etc/systemd/system/llama-server.service (VM 101 .8, VM 103 .110)
# Hermes agent configs (key enforcement — 2026-07-10)
/root/.hermes/config.yaml (Mumuni CT 114, Tanko CT 112, etc.)
+6 -6
View File
@@ -8,18 +8,18 @@ description: >
Ensures agents never use the master key directly. Rotation is event-driven,
not calendar-driven — rotate only on compromise, personnel change, or
periodic security hygiene (quarterly/annually).
UPDATED 2026-07-12: Keys are stored in Infisical vault (project=agents, env=production)
BUT each agent host MUST keep a local .env fallback. Infisical service tokens can
expire/404. The .env fallback prevents agents from running without keys.
Tanko incident: token 404 → gateway had no LITELLM_API_KEY for hours.
UPDATED 2026-07-16: Vault is SYNCED (session-13 keys written to vault via abiba service
token, all validate 200). Koby/Koonimo migrated from hardcoded drop-ins to the
infisical-gateway.sh wrapper (live vault injection). 4/5 agents now vault-backed.
Canonical process: see § Production Vault Access Process. Tanko (user jerome) pending.
Abiba's key is now a proper agent key (NOT the master key — stale note removed).
UPDATED 2026-07-17: FLEET-WIDE STANDARDIZATION. All 4 agents (Mumuni, Tanko, Koby, Koonimo)
standardized on a single pattern: systemd drop-in (ExecStart= reset + wrapper path) →
infisical-gateway.sh while-true loop → /usr/bin/infisical run --token → bash -c key
@@ -30,7 +30,7 @@ description: >
Critical lessons: (1) NEVER use shell variables inside single-quoted bash -c in wrappers
— hardcode absolute paths. (2) Drop-ins override unit file ExecStart permanently.
(3) Capture /proc/<pid>/environ before gateway restarts to preserve running env set.
Current key inventory and agent list: see gpu-fleet.prose.md § Agent Keys.
Source of truth for LiteLLM config: /opt/inference-harness/litellm_config.yaml
on CT 116. Last verified: 2026-07-17.
@@ -154,7 +154,7 @@ through its agent wrapper.
The `ExecStart=` (empty reset) clears any ExecStart from the main unit file,
then the second `ExecStart=` sets the wrapper. This drop-in **survives unit file
regeneration** by `hermes gateway install` — the drop-in always wins.
**Why a drop-in instead of editing the unit file:** `hermes gateway install`
(called during Hermes updates and some self-heal operations) regenerates the
systemd unit file with `ExecStart=/path/to/python -m hermes_cli.main gateway run`.
@@ -171,7 +171,7 @@ through its agent wrapper.
- **Survives gateway crash**: the wrapper's `while true` + systemd `Restart=always` revive the gateway. Two-layer defense.
- **Survives Hermes updates**: systemd drop-in overrides unit file ExecStart — `hermes gateway install` cannot break the vault injection.
- **Survives reboot**: systemd user service + `loginctl enable-linger` ensures gateway starts at boot without a login session.
- **Auditable**: `cat /proc/$(pgrep hermes_cli)/environ` shows all injected keys; `infisical secrets` shows the vault source.
- **Auditable**: `cat /proc/$(pgrep -f 'python.*hermes_cli.main.gateway.run' | grep -v infisical | head -1)/environ` shows all injected keys (note: pipe through grep -v infisical to avoid matching the bash wrapper); `infisical secrets` shows the vault source.
### Migration status (2026-07-17)
+4 -4
View File
@@ -42,7 +42,7 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
**What changed (v3.2.0 → v4.0.0 — 2026-07-08)**:
- Router REMOVED from request path — LiteLLM proxies directly to GPU
- All GPUs at parallel 2 (was parallel 1)
- NVIDIA context reduced 256K→128K to free VRAM (SUPERSEDED 2026-07-16: all GPUs back to 256K — see litellm-self-heal)
- NVIDIA context reduced 256K→128K to free VRAM — now the stable ceiling across all GPUs (2026-07-17)
- LiteLLM timeouts tuned: gemma 25→120s, qwen 40→90s (SUPERSEDED 2026-07-16: qwen 300s, gemma 120s, strix 300s — see litellm-self-heal)
- nginx proxy_read_timeout: 600s, LiteLLM request_timeout: 300s
@@ -72,9 +72,9 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
| Host | IP | Hardware | Models Served | Engine | Context | Parallel |
|------|-----|----------|---------------|--------|---------|----------|
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd | **256K** | 2 |
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd | **256K** | 2 |
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: strix-moe) | llama-server systemd (Vulkan) | 256K | 2 |
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd | **128K** | 2 |
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd | **128K** | 2 |
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: strix-moe) | llama-server systemd (Vulkan) | 128K | 2 |
## Model Fallback Chains (LiteLLM)
+4 -4
View File
@@ -57,9 +57,9 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
| Host | IP | Hardware | Models Served | Engine | Context | Parallel |
|------|-----|----------|---------------|--------|---------|----------|
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `-c 262144 --parallel 2 --ngl 99`) | **256K** | 2 |
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `--ctx-size 262144 --parallel 2`, IQ4_NL + MTP draft) | **256K** | 2 |
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: `strix-moe`) | llama-server systemd (Vulkan) | 256K | 2 |
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `-c 131072 --parallel 2 --ngl 99`) | **128K** | 2 |
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `--ctx-size 131072 --parallel 2`, IQ4_NL + MTP draft) | **128K** | 2 |
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: `strix-moe`) | llama-server systemd (Vulkan) | 128K | 2 |
> Verified on ground 2026-07-16 via `curl /v1/models` on each host + `llama-wrapper.sh`. The AMD host's underlying model is `qwen3.6-35B-udq4`; LiteLLM exposes it under two `model_name`s: `qwen3.6-35B-udq4` and `strix-moe` (rpm 40). The legacy name `ornith-1.0-35b` does NOT exist in LiteLLM and must not be referenced.
@@ -101,7 +101,7 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
## Script Operations (synced 2026-07-16)
- **Health-check script** (`/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116): `gpu-fleet` check fails only on **critical** alerts (warnings are informational). Tests `strix-moe` (not `ornith-1.0-35b`).
- **GPU monitor** (`/root/scripts/gpu-monitor-server.py` on pi .24): runs as **systemd unit `gpu-monitor.service`** (was bare `&` process). `gpu_count` includes Strix Halo (was 2, now 3). VRAM alert thresholds: warning 93%, critical 97% (raised from 90/95 — 256K context steady-state is ~96% on RTX 3090, not a fault).
- **GPU monitor** (`/root/scripts/gpu-monitor-server.py` on pi .24): runs as **systemd unit `gpu-monitor.service`** (was bare `&` process). `gpu_count` includes Strix Halo (was 2, now 3). VRAM alert thresholds: warning 93%, critical 97% (raised from 90/95 — 128K context steady-state is ~70% on RTX 3090, not a fault).
- **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): reads each agent's **live** `LITELLM_API_KEY` from its gateway process env via SSH — never hardcodes keys (hardcoded keys rot on rotation and caused 9×401/30min). Fleet roster: abiba, tanko, mumuni, koby, koonimo (legacy `tdunna`/`baggy` removed — never existed).
- **Stale keys cleaned**: `daily-infra-report.py` SYNTHETIC_API_KEY was stale (`sk-U_ydi3B` → 401); now reads `LITELLM_MASTER_KEY` from env. Deprecated scripts (`router-original.py`, `router-phase0-backup.py`, `apply-fixes.py`) still reference `sk-syslog-local-master-key` but do not actively poll LiteLLM.
+1 -1
View File
@@ -87,7 +87,7 @@ agent: abiba
| storepve | 192.168.68.6 | PVE |
| acerpve | 192.168.68.9 | PVE (hosts llm-gpu qemu/101) |
| minipve | 192.168.68.12 | PVE |
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (ornith) |
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (qwen3.6-35B-udq4, strix-moe) |
## Operations
+2 -2
View File
@@ -184,8 +184,8 @@ def check_agents():
print(f"{name} (CT {ct}): cannot SSH — skip liveness check")
continue
# Gateway process
pid = ssh(host, "pgrep -f 'hermes_cli.main gateway run' | head -1", user=user)
# Gateway process (exclude the infisical bash wrapper that contains the same string)
pid = ssh(host, "pgrep -f 'hermes_cli.main gateway run' | grep -v infisical | head -1", user=user)
if not pid:
print(f"{name}: GATEWAY NOT RUNNING")
FAIL.append(f"gateway-down:{name}")