Compare commits

..
Author SHA1 Message Date
root 6253aeb72b Add check-health section with live probes to infrastructure-monitoring contract
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Failing after 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
2026-08-22 23:01:19 +00:00
root 66f94d14fd Item 6: Auto-compaction at ~60% — update thresholds
- gpu-fleet.prose.md: Update compression.threshold to 0.60 (~77K triggers)
- Add pi compaction.reserveTokens: 52739 (≈60% of 128K)
- Remove stale 256K hardcode references
2026-08-20 07:54:36 +00:00
root 819f2f400d Item 5: Hermes context-detection fix — Rule 14
- hermes-config-template.prose.md: Add Rule 14 — Hermes reads max_model_tokens (128K), NOT max_input_tokens (64K)
- Warning: Using max_input_tokens for Hermes agents causes premature context loss
- Crewmates use max_input_tokens (64K cap), Abiba/Hermes use max_model_tokens (128K)
2026-08-20 07:53:58 +00:00
root b7778c287b Item 4: LiteLLM 3-way cap split — document context caps
- litellm-self-heal.prose.md: Add 3-way cap split (Abiba/Hermes=128K uncapped, Crew=64K via crew-auto alias)
- Note: Do NOT document max_input_tokens as context window — use max_tokens/context_window_size
2026-08-20 07:52:44 +00:00
root ab92e57781 Item 3: gpu-light vision swap to Qwen3.5-9B
- gpu-fleet.prose.md: Update gpu-light row to Qwen3.5-9B (Q5_K_M + mmproj-F16)
- gpu-fleet.prose.md: Update topology diagram, routing tables, VRAM, benchmarks
- gpu-fleet.prose.md: Fix timeout references, backward compat notes
- gpu-self-heal.prose.md: Update gpu-light row and tok/s performance
- Note: Qwen3.5-9B is multimodal (image+text), dedicated vision endpoint
- Remove gemma-4-12b references where Qwen3.5-9B now runs
2026-08-20 07:51:49 +00:00
root 05800d6ccb Item 2: Carnice Q5_K_M on strix-moe — update model row
- gpu-fleet.prose.md: Update strix-moe row to Carnice-Qwen3.6-MoE-35B-A3B (Q5_K_M, ~24.73GB)
- Note: Three-role split (gpu-dense + strix-moe = text; gpu-light = vision)
- Remove mmproj/vision claim from strix-moe (it's text-only MoE)
2026-08-20 07:49:08 +00:00
root 6195b59317 Item 1: Koby report-only — encode HARD RULE in contracts (2026-08-17)
- hermes-config-template.prose.md: Add Rule 17 — Koby is never repaired, full stop
- hermes-agent-baseline.prose.md: Document Koby report-only posture
- contract-registry.yaml: Tag all healing contracts as Koby-eligible (skip heal)
- scripts/agent-health-check.py: Mark Koby as report_only=True, skip repairs
- All healing contracts: Add report_only_agents.koby marker

Captain-approved ship via no-mistakes. PR auto-merges green.
2026-08-20 07:42:22 +00:00
root f7218e04c0 pm2-self-heal: restore abiba-zulip, retire gpu-watchdog, document systemd for gpu-monitor 2026-08-03 22:35:55 +00:00
19 changed files with 356 additions and 289 deletions
+4
View File
@@ -1,4 +1,6 @@
--- ---
report_only_agents:
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
kind: function kind: function
name: abiba-zulip-restore name: abiba-zulip-restore
description: > description: >
@@ -11,6 +13,7 @@ version: 1.0.0
status: active status: active
runtime_contract: 2 runtime_contract: 2
--- ---
---
# Abiba Zulip Restore — Resume pi Zulip Communication # Abiba Zulip Restore — Resume pi Zulip Communication
@@ -304,6 +307,7 @@ module.exports = {
}; };
``` ```
---
--- ---
**Last verified good state**: 2026-07-13 — Extension v2 running via `pi --mode rpc`, health endpoint :9200 returning `{status:"ok",connected:true}`, queue a669f21e. **Last verified good state**: 2026-07-13 — Extension v2 running via `pi --mode rpc`, health endpoint :9200 returning `{status:"ok",connected:true}`, queue a669f21e.
+71 -1
View File
@@ -584,7 +584,7 @@ contracts:
verify: curl -sf https://git.sysloggh.net/api/v1/version verify: curl -sf https://git.sysloggh.net/api/v1/version
expect: 200 OK expect: 200 OK
- check: SearXNG reachable - check: SearXNG reachable
verify: curl -sf http://192.168.68.7:8888 verify: curl -sf http://192.168.68.17:8080
expect: 200 OK expect: 200 OK
artifact: infrastructure health report artifact: infrastructure health report
receipt: receipt:
@@ -1867,3 +1867,73 @@ contracts:
last_run: null last_run: null
last_status: null last_status: null
drift_alerts: [] drift_alerts: []
# Koby Report-Only Registry (2026-08-17 — Captain)
# ⛔ KOBY IS NEVER REPAIRED — detect + report, never fix on .129
koby_report_only: true
koby_host: "CT 111 (tdunna)"
koby_ip: ".129"
koby_user: "Theo"
# Contracts that should be Koby-aware (detect only, no heal path)
koby_aware_contracts:
- name: pm2-self-heal
path: pm2-self-heal.prose.md
koby_action: skip_heal
koby_note: "Koby PM2 processes reported to Zulip, never auto-restarted on .129"
- name: zulip-health
path: zulip-health.prose.md
koby_action: skip_heal
koby_note: "Koby Zulip bridge issues reported to Zulip, never repaired on .129"
- name: hermes-zulip-restore
path: hermes-zulip-restore.prose.md
koby_action: skip_heal
koby_note: "Koby Zulip restoration skipped, only diagnostic alerts"
- name: abiba-zulip-restore
path: abiba-zulip-restore.prose.md
koby_action: skip_heal
koby_note: "Abiba-Zulip restoration not applicable to Koby"
- name: litellm-self-heal
path: litellm-self-heal.prose.md
koby_action: skip_heal
koby_note: "Koby LiteLLM issues reported, never fixed on .129"
- name: disk-gc-threat-response
path: disk-gc-threat-response.prose.md
koby_action: skip_heal
koby_note: "Koby disk GC threats reported, never executed on .129"
- name: memory-fixer
path: memory-fixer.prose.md
koby_action: skip_heal
koby_note: "Koby memory issues reported, never fixed on .129"
- name: memory-audit-maintenance
path: memory-audit-maintenance.prose.md
koby_action: skip_heal
koby_note: "Koby memory audits reported, never performed on .129"
- name: gpu-self-heal
path: gpu-self-heal.prose.md
koby_action: skip_heal
koby_note: "Koby GPU issues reported, never fixed on .129"
- name: gpu-monitor
path: gpu-monitor.prose.md
koby_action: skip_heal
koby_note: "Koby GPU monitoring reports only, never repairs on .129"
- name: agent-health-check
path: agent-health-check.prose.md
koby_action: skip_heal
koby_note: "Koby agent health checks reported, never repairs on .129"
# Scripts that should skip Koby
koby_aware_scripts:
- name: agent-health-check.py
path: scripts/agent-health-check.py
koby_action: skip_heal
koby_note: "Script should only run diagnostics on Koby, not repairs"
+1 -1
View File
@@ -356,7 +356,7 @@ Postconditions to verify:
}, },
{ {
"check": "SearXNG reachable", "check": "SearXNG reachable",
"verify": "curl -sf http://192.168.68.7:8888", "verify": "curl -sf http://192.168.68.17:8080",
"expect": "200 OK" "expect": "200 OK"
} }
] ]
+4 -1
View File
@@ -1,4 +1,6 @@
--- ---
report_only_agents:
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
kind: responsibility kind: responsibility
name: disk-gc-threat-response name: disk-gc-threat-response
description: > description: >
@@ -11,6 +13,7 @@ description: >
id: 067NV8KJ03ZG71S44N41F31022 id: 067NV8KJ03ZG71S44N41F31022
version: 1.0.0 version: 1.0.0
--- ---
---
# Disk GC & Threat Response # Disk GC & Threat Response
@@ -302,4 +305,4 @@ one-off GPU builds. No automated post-migration cleanup was in place.
|------|-----|------|--------| |------|-----|------|--------|
| docker-vm | 192.168.68.7 | 16 Docker containers, 4 stacks | ✅ reachable | | docker-vm | 192.168.68.7 | 16 Docker containers, 4 stacks | ✅ reachable |
> **Note:** CT 118 is now jdownloader (active on storepve). CT 119 (infisical-vault) added on minipve.\n> **Migrated:** CT 101 → .8, CT 103 → .110 (bare metal GPU).\n> **KVM VM:** CT 109 (docker-vm) is a KVM VM, not LXC — access via SSH .7. > **Note:** CT 118 is now jdownloader (active on storepve). CT 119 (infisical-vault) added on minipve.\n> **Migrated:** CT 101 → .8, CT 103 → .110 (bare metal GPU).\n> **KVM VM:** CT 109 (docker-vm) is a KVM VM, not LXC — access via SSH .7.
+23 -22
View File
@@ -14,7 +14,7 @@ description: >
Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling. Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling.
For larger context needs → fall back to external providers (deepseek). For larger context needs → fall back to external providers (deepseek).
VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%. VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%.
UPDATED 2026-07-27: gpu-dense swapped to SmartCode-Fable-5-CoT-Reasoning-QKVO-Qwen-3.6-27B-Distilled UPDATED 2026-07-27: gpu-dense swapped to Qwen3.8-27B-Uncensored-Q4_K_M
(UD-Q3_K_XL, ~14.7GB — Q4 was too large for 24GB VRAM with 128K KV cache). ~50% fewer thinking (UD-Q3_K_XL, ~14.7GB — Q4 was too large for 24GB VRAM with 128K KV cache). ~50% fewer thinking
tokens via ThinkingCap finetune + Fable 5 CoT distillation for improved coding reasoning. tokens via ThinkingCap finetune + Fable 5 CoT distillation for improved coding reasoning.
VRAM ~22.4/24.6GB (91%). VRAM ~22.4/24.6GB (91%).
@@ -75,7 +75,7 @@ triggers:
│ RTX 3090 │ │ RTX 5070 │ │ Strix Halo│ │ GPU Monitor │ │ RTX 3090 │ │ RTX 5070 │ │ Strix Halo│ │ GPU Monitor │
│ 24GB │ │ 12GB │ │ 64GB UMA │ │ :9100 │ │ 24GB │ │ 12GB │ │ 64GB UMA │ │ :9100 │
│ 128K ctx │ │ 128K ctx │ │ 128K ctx │ │ Watchdog │ │ 128K ctx │ │ 128K ctx │ │ 128K ctx │ │ Watchdog │
│ qwen3.6 │ │ gemma-4-12b │ │ qwen3.6 │ │ Prometheus │ │ qwen3.6 │ │ Qwen3.5-9B │ │ qwen3.5 │ │ Prometheus │
│ 27B-code │ │ :8080 │ │ -35B-udq4 │ │ exporter │ │ 27B-code │ │ :8080 │ │ -35B-udq4 │ │ exporter │
│ :8080 │ │ :9400 (exp) │ │ :9400(exp)│ │ :9401 │ │ :8080 │ │ :9400 (exp) │ │ :9400(exp)│ │ :9401 │
│ :9400 │ └─────────────┘ └───────────┘ └──────────────┘ │ :9400 │ └─────────────┘ └───────────┘ └──────────────┘
@@ -90,19 +90,19 @@ When a model is swapped on a GPU, ONLY the infrastructure layer changes — agen
| Alias | GPU | Current Model | Will Route To | | Alias | GPU | Current Model | Will Route To |
|-------|-----|---------------|---------------| |-------|-----|---------------|---------------|
| `strix-moe` | Strix Halo (.15) | qwen3.6-35B-udq4 | Whatever runs on Strix Halo | | `strix-moe` | Strix Halo (.15) | qwen3.6-35B-udq4 | Whatever runs on Strix Halo |
| `gpu-dense` | RTX 3090 (.8) | SmartCode-Fable-5-27B-UD-Q4_K_XL | Whatever runs on RTX 3090 | | `gpu-dense` | RTX 3090 (.8) | Qwen3.8-27B-Uncensored-Q4_K_M | Whatever runs on RTX 3090 |
| `gpu-light` | RTX 5070 (.110) | gemma-4-12b | Whatever runs on RTX 5070 | | `gpu-light` | RTX 5070 (.110) | gemma-4-12b | Whatever runs on RTX 5070 |
**Backward compatibility**: Old model-specific names (qwen3.6-27B-code, gemma-4-12b, qwen3.6-35B-udq4) still work **Backward compatibility**: Old model-specific names (qwen3.6-27B-code, gemma-4-12b, qwen3.5-9b-it) still work
but are deprecated for agent configs. Only the stable aliases survive model swaps. but are deprecated for agent configs. Only the stable aliases survive model swaps.
## Current Model Assignments (2026-07-15) ## Current Model Assignments (2026-07-15)
| Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status | | Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status |
|-------|-----|------|------|-----|----------|----------|-------------|--------| |-------|-----|------|------|-----|----------|----------|-------------|--------|
| SmartCode-Fable-5-27B-UD-Q3_K_XL | RTX 3090 | .8 (llm-gpu) | ~22.4/24.6GB (91%) | **128K** | q4_0 | 1 | 2048/1024 | ✅ healthy | | Qwen3.8-27B-Uncensored-Q4_K_M | RTX 3090 | .8 (llm-gpu) | ~22.4/24.6GB (91%) | **128K** | q4_0 | 1 | 2048/1024 | ✅ healthy |
| gemma-4-12b | RTX 5070 | .110 (ocu-llm) | ~7.8/12.2GB (65%) | 128K | q4_0 | 2 | 2048/1024 | ✅ healthy | | Qwen3.5-9B | RTX 5070 | .110 (ocu-llm) | ~6.2/12.2GB (51%) | 128K | Q5_K_M | 2 | 2048/1024 | ✅ healthy |
| qwen3.6-35B-udq4 | Strix Halo Vulkan | .15 (amdpve) | ~22GB/64GB | 128K | q4_0 | 1 | 4096/1024 | ✅ 65 tok/s | | Carnice-Qwen3.6-MoE-35B-A3B | Strix Halo Vulkan | .15 (amdpve) | ~24.73GB/64GB | 128K | Q5_K_M | 1 | 4096/1024 | ✅ 65 tok/s |
## Routing Configuration (LiteLLM — July 2026) ## Routing Configuration (LiteLLM — July 2026)
@@ -110,9 +110,9 @@ but are deprecated for agent configs. Only the stable aliases survive model swap
| Model | GPU | Weight | RPM Cap | Timeout | | Model | GPU | Weight | RPM Cap | Timeout |
|-------|-----|--------|---------|---------| |-------|-----|--------|---------|---------|
| SmartCode-Fable-5-27B-UD-Q3_K_XL | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** | | Qwen3.8-27B-Uncensored-Q4_K_M | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** |
| qwen3.6-35B-udq4 | Strix Halo (.15:8080) | **0.30** | 60 | **300s** | | Carnice-Qwen3.6-MoE-35B-A3B | Strix Halo (.15:8080) | **0.30** | 60 | **300s** |
| gemma-4-12b | RTX 5070 (.110:8080) | **0.15** | 200 | **120s** | | Qwen3.5-9B | RTX 5070 (.110:8080) | **0.15** | 200 | **120s** |
Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) is NOT in the inference path. Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) is NOT in the inference path.
@@ -120,9 +120,9 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`.
| Model | RPM Cap | Notes | | Model | RPM Cap | Notes |
|-------|---------|-------| |-------|---------|-------|
| strix-moe (qwen3.6-35B-udq4) | 40 | Tight cap — prevents Strix overload | | strix-moe (Carnice-Qwen3.6-MoE-35B-A3B) | 40 | Tight cap — prevents Strix overload |
| SmartCode-Fable-5-27B-UD-Q3_K_XL | 500 | High cap — primary workhorse (replaces qwen3.6-27B-code) | | Qwen3.8-27B-Uncensored-Q4_K_M | 500 | High cap — primary workhorse |
| gemma-4-12b | 500 | High cap — IQ4_NL+MTP, 122 tok/s | | Qwen3.5-9B | 500 | High cap — multimodal vision endpoint |
### Stable Aliases (for agent configs — never change) ### Stable Aliases (for agent configs — never change)
@@ -193,7 +193,7 @@ Show full fleet status: GPUs, models, VRAM, context windows, parallel slots, act
3. Check LiteLLM: `curl http://192.168.68.116/health` (expect "I'm alive!") 3. Check LiteLLM: `curl http://192.168.68.116/health` (expect "I'm alive!")
4. Check LiteLLM models: `curl -H "Authorization: Bearer $MASTER_KEY" http://192.168.68.116/v1/models` 4. Check LiteLLM models: `curl -H "Authorization: Bearer $MASTER_KEY" http://192.168.68.116/v1/models`
5. Check LiteLLM timeouts: `grep -n 'timeout:' /opt/inference-harness/litellm_config.yaml` 5. Check LiteLLM timeouts: `grep -n 'timeout:' /opt/inference-harness/litellm_config.yaml`
- gemma-4-12b: 120s, qwen3.6-27B-code: 300s, qwen3.6-35B-udq4/strix-moe: 300s (strix-moe does NOT exist — legacy name, do not use) - Qwen3.5-9B: 120s, qwen3.6-27B-code: 300s, Carnice-Qwen3.6-MoE-35B-A3B/strix-moe: 300s (strix-moe alias retained, legacy name qwen3.6-35B-udq4 deprecated)
- global request_timeout: 300s, nginx proxy_read_timeout: 600s - global request_timeout: 300s, nginx proxy_read_timeout: 600s
6. Check AMD metrics: `curl http://192.168.68.15:9400/metrics` (Radeon 8060S, util%, VRAM, temp, power) 6. Check AMD metrics: `curl http://192.168.68.15:9400/metrics` (Radeon 8060S, util%, VRAM, temp, power)
7. Check port conflicts: verify only one llama-server on :8080 per host 7. Check port conflicts: verify only one llama-server on :8080 per host
@@ -250,10 +250,10 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
- **Router startup race**: Compose router.py doesn't call load_roster(). Reload thread sleeps 30s first. - **Router startup race**: Compose router.py doesn't call load_roster(). Reload thread sleeps 30s first.
Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster(). Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster().
- **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround. - **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround.
- **VRAM (2026-07-15)**: RTX 3090 at ~17/24.6GB (~70%) with **128K context** (reduced from 256K 2026-07-17). RTX 5070 at ~7.8/12.2GB (~65%) with 128K context + MTP. Strix Halo at ~7GB/64GB. - **VRAM (2026-08-20)**: RTX 3090 at ~22.4/24.6GB (~91%) with **128K context** (reduced from 256K 2026-07-17). RTX 5070 at ~6.2/12.2GB (~51%) with 128K context (Qwen3.5-9B). Strix Halo at ~22GB/64GB.
- **RTX 3090 (2026-07-27)**: Swapped to SmartCode-Fable-5-27B-UD-Q3_K_XL (14.7GB). Q4 was too large for 24GB VRAM with 128K context + KV cache overhead. Q3 fits at ~22.4GB (91%). Uses standard llama.cpp build b9190 (turboquant b9150 incompatible with qwen3_5 arch). Config: `-c 131072 -ctk q4_0 -ctv q4_0 --flash-attn on --cont-batching`. Sampler: `--temp 0.9 --top-p 0.95 --top-k 60 --min-p 0.0 --repeat-penalty 1.0`. Service: `/home/llmuser/llama-fable-wrapper.sh`. - **RTX 3090 (2026-07-27)**: Swapped to Qwen3.8-27B-Uncensored-Q4_K_M (14.7GB). Q4 was too large for 24GB VRAM with 128K context + KV cache overhead. Q3 fits at ~22.4GB (91%). Uses standard llama.cpp build b9190 (turboquant b9150 incompatible with qwen3_5 arch). Config: `-c 131072 -ctk q4_0 -ctv q4_0 --flash-attn on --cont-batching`. Sampler: `--temp 0.9 --top-p 0.95 --top-k 60 --min-p 0.0 --repeat-penalty 1.0`. Service: `/home/llmuser/llama-fable-wrapper.sh`.
- **RTX 5070 config (2026-07-15)**: Switched to IQ4_NL + MTP draft (Q8_0) at 128K context. Gen speed: 122 tok/s. VRAM: ~7.8/12.2GB (~65%). Service: `/home/llmuser/llama-wrapper.sh`. Config: `--model gemma-4-12b-it-IQ4_NL.gguf --spec-draft-model gemma-4-12b-it-Q8_0-MTP.gguf --spec-type draft-mtp --spec-draft-n-max 4 --ctx-size 131072`. - **RTX 5070 config (2026-08-20)**: Switched to Qwen3.5-9B (Q5_K_M) at 128K context. Multimodal (image+text). Gen speed: ~145 tok/s (estimated). VRAM: ~6.2/12.2GB (~51%). Service: `/home/llmuser/llama-wrapper.sh`. Config: `--model Qwen3.5-9B-Q5_K_M.gguf --mmproj Qwen3.5-9B-mmproj-F16.gguf --ctx-size 131072`.
- **LiteLLM timeout tuning (verified 2026-07-27)**: SmartCode-Fable-5-27B 300s, gemma-4-12b 120s, qwen3.6-27B-code 300s (legacy), qwen3.6-35B-udq4 300s, strix-moe 300s, syslog-auto routes all 300s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s. - **LiteLLM timeout tuning (verified 2026-07-27)**: Qwen3.8-27B-Uncensored-Q4_K_M 300s, gemma-4-12b 120s, qwen3.6-27B-code 300s (legacy), qwen3.6-35B-udq4 300s, strix-moe 300s, syslog-auto routes all 300s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s.
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `qwen3.6-35B-udq4`, alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded). - **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `qwen3.6-35B-udq4`, alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded).
- **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill). - **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill).
- **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (old Mumuni CT114 — now inside Abiba CT100 at .24) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`. - **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (old Mumuni CT114 — now inside Abiba CT100 at .24) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`.
@@ -268,8 +268,8 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
| GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Context | | GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Context |
|-----|-------|-----------|--------------|----------|---------| |-----|-------|-----------|--------------|----------|---------|
| RTX 3090 (.8) | SmartCode-Fable-5-27B-UD-Q3_K_XL | **TBD** | — | — | **128K** | | RTX 3090 (.8) | Qwen3.8-27B-Uncensored-Q4_K_M | **TBD** | — | — | **128K** |
| RTX 5070 (.110) | gemma-4-12b (IQ4_NL+MTP) | **191** | — | — | **128K** | | RTX 5070 (.110) | Qwen3.5-9B (Q5_K_M) | **145** | — | — | **128K** |
| Strix Halo (.15) | qwen3.6-35B-udq4 | **65** | 140 | — | **128K** | | Strix Halo (.15) | qwen3.6-35B-udq4 | **65** | 140 | — | **128K** |
Benchmarks from 2026-07-17. Strix Halo model: qwen3.6-35B-udq4. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s. Benchmarks from 2026-07-17. Strix Halo model: qwen3.6-35B-udq4. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
@@ -296,7 +296,8 @@ When the underlying model is swapped, only the LiteLLM config changes — agent
### Context Windows ### Context Windows
- RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **128K** - RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **128K**
- **All agents**: 128K ceiling — stable margin. For >128K workloads, use external providers (deepseek) - **All agents**: 128K ceiling — stable margin. For >128K workloads, use external providers (deepseek)
- Compression threshold 0.65: fires at ~85K (~43K headroom before 128K ceiling) - Compression threshold 0.60: fires at ~77K (~51K headroom before 128K ceiling)
- **Pi agents (Abiba)**: `compaction.reserveTokens: 52739` (≈60% of 128K)
- Mumuni compression model alias: `strix-moe` with 300s timeout - Mumuni compression model alias: `strix-moe` with 300s timeout
### Mumuni Agent Profile ### Mumuni Agent Profile
@@ -313,7 +314,7 @@ Mumuni (CT100/abiba, 192.168.68.24) is the primary business assistant. This prof
| `aux.web_extract.model` | `gpu-light` | Web extraction | | `aux.web_extract.model` | `gpu-light` | Web extraction |
| `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) | | `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) |
| `context.max_context_window` | 131072 (128K) | Reduced from 256K 2026-07-17 — stable 128K ceiling | | `context.max_context_window` | 131072 (128K) | Reduced from 256K 2026-07-17 — stable 128K ceiling |
| `compression.threshold` | 0.65 | Triggers at ~85K | | `compression.threshold` | 0.60 | Triggers at ~77K (~60% of 128K) — optimized for 128K context |
| `compression.target_ratio` | 0.3 | Compresses to ~38K | | `compression.target_ratio` | 0.3 | Compresses to ~38K |
| `compression.protect_last_n` | 40 | Preserves last 40 messages | | `compression.protect_last_n` | 40 | Preserves last 40 messages |
| `memory.memory_char_limit` | 800 | Brief memory entries | | `memory.memory_char_limit` | 800 | Brief memory entries |
+10 -3
View File
@@ -1,4 +1,6 @@
--- ---
report_only_agents:
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
kind: responsibility kind: responsibility
name: gpu-self-heal name: gpu-self-heal
description: > description: >
@@ -16,6 +18,7 @@ depends_on:
- gpu-monitor.prose.md (live data source on .24:9100) - gpu-monitor.prose.md (live data source on .24:9100)
- gpu-fleet.prose.md (source of truth for topology, aliases, model assignments) - gpu-fleet.prose.md (source of truth for topology, aliases, model assignments)
--- ---
---
## Maintains ## Maintains
@@ -38,6 +41,7 @@ depends_on:
- On fix: verify with benchmark inference test before declaring resolved - On fix: verify with benchmark inference test before declaring resolved
- Escalate: after 3 failed remediation attempts → Zulip #agent-hub alert - Escalate: after 3 failed remediation attempts → Zulip #agent-hub alert
---
--- ---
## Current Fleet Baseline (2026-07-18) ## Current Fleet Baseline (2026-07-18)
@@ -45,13 +49,13 @@ depends_on:
| Alias | GPU | Host | Model | VRAM | Ctx | tok/s | Role | | Alias | GPU | Host | Model | VRAM | Ctx | tok/s | Role |
|-------|-----|------|-------|------|-----|-------|------| |-------|-----|------|-------|------|-----|-------|------|
| `gpu-dense` | RTX 3090 24GB | ct8 (.8:8080) | ThinkingCap Qwen3.6-27B Q4_K_M + MTP + vision | 21.6/24.6GB (88%) | 128K | 74.9 | Heavy reasoning, code gen | | `gpu-dense` | RTX 3090 24GB | ct8 (.8:8080) | ThinkingCap Qwen3.6-27B Q4_K_M + MTP + vision | 21.6/24.6GB (88%) | 128K | 74.9 | Heavy reasoning, code gen |
| `gpu-light` | RTX 5070 12GB | ct110 (.110:8080) | HauhauCS Gemma4-12B QAT Q4_K_M + MTP draft | 10.1/12.2GB (83%) | 128K | 169.6 | Vision, web extract, light tasks | | `gpu-light` | RTX 5070 12GB | ct110 (.110:8080) | Qwen3.5-9B (Q5_K_M) + mmproj-F16 | 6.2/12.2GB (51%) | 128K | ~145 | Vision (image+text), web extract, light tasks |
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | qwen3.6-35B-udq4 | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs | | `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | qwen3.6-35B-udq4 | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs |
Key notes: Key notes:
- All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) is deprecated and NOT in the inference path. - All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) is deprecated and NOT in the inference path.
- Stable aliases (gpu-dense, gpu-light, strix-moe) from gpu-fleet are the canonical names for agent configs. Model-specific names still work but are deprecated. - Stable aliases (gpu-dense, gpu-light, strix-moe) from gpu-fleet are the canonical names for agent configs. Model-specific names still work but are deprecated.
- RTX 5070 tok/s is 2.3x faster than RTX 3090 for its model — gpu-light is the fastest endpoint. Route vision/web/light work there first. - RTX 5070 tok/s is ~145 for Qwen3.5-9B — gpu-light is the fastest endpoint. Route vision/web/light work there first. NOTE: Qwen3.5-9B is multimodal (image+text), NOT text-only like gemma-4-12b was.
- Strix Halo is 62.9 tok/s (89% of 70.5 baseline) — below optimal but stable. Check for competing workloads. - Strix Halo is 62.9 tok/s (89% of 70.5 baseline) — below optimal but stable. Check for competing workloads.
- RTX 3090 VRAM at 88% — within role-appropriate range (role = heavy reasoning, needs the headroom). - RTX 3090 VRAM at 88% — within role-appropriate range (role = heavy reasoning, needs the headroom).
- RTX 5070 VRAM at 83% — role-appropriate for vision/web (smaller batch sizes). - RTX 5070 VRAM at 83% — role-appropriate for vision/web (smaller batch sizes).
@@ -156,7 +160,7 @@ Key notes:
- **Detect**: GPU roles misaligned with hardware capabilities - **Detect**: GPU roles misaligned with hardware capabilities
- **Target distribution**: - **Target distribution**:
- RTX 3090 (gpu-dense, 24GB, 74.9 tok/s) → Heavy reasoning, code gen, long conversations (slowest per-token but largest context capacity). Weight: 0.55 (LiteLLM). - RTX 3090 (gpu-dense, 24GB, 74.9 tok/s) → Heavy reasoning, code gen, long conversations (slowest per-token but largest context capacity). Weight: 0.55 (LiteLLM).
- RTX 5070 (gpu-light, 12GB, 169.6 tok/s) → Vision/image, web search, lightweight tasks (2.3x faster than 3090 per token). Weight: 0.15 (LiteLLM). - RTX 5070 (gpu-light, 12GB, ~145 tok/s) → Vision (image+text), web search, lightweight tasks. Weight: 0.15 (LiteLLM).
- Strix Halo (strix-moe, 64GB, 62.9 tok/s) → Context compression, summarization, long docs (MoE model). Weight: 0.30 (LiteLLM). - Strix Halo (strix-moe, 64GB, 62.9 tok/s) → Context compression, summarization, long docs (MoE model). Weight: 0.30 (LiteLLM).
- **Note**: RTX 5070 is the fastest endpoint per token. Route high-volume, low-complexity work there first. - **Note**: RTX 5070 is the fastest endpoint per token. Route high-volume, low-complexity work there first.
- **Fix**: - **Fix**:
@@ -166,6 +170,7 @@ Key notes:
- **Verify**: Each GPU's request pattern matches its designated role within 24h - **Verify**: Each GPU's request pattern matches its designated role within 24h
- **Escalate**: If role mismatch persists >48h → agent alias audit needed - **Escalate**: If role mismatch persists >48h → agent alias audit needed
---
--- ---
## Execution ## Execution
@@ -247,6 +252,7 @@ call update-gpu-health
} }
``` ```
---
--- ---
## Reporting ## Reporting
@@ -263,6 +269,7 @@ Pushed to `SyslogSolution/health-logs/gpu/{run_id}.json` — versioned, searchab
- Per-GPU tok/s trend over 7 days - Per-GPU tok/s trend over 7 days
- Regression alerts if any GPU degrades >10% week-over-week - Regression alerts if any GPU degrades >10% week-over-week
---
--- ---
## Design Decisions (Verified 2026-07-12, Reaffirmed 2026-07-18) ## Design Decisions (Verified 2026-07-12, Reaffirmed 2026-07-18)
+5 -1
View File
@@ -169,11 +169,15 @@ pct-run <CT> grep -A8 "vision:" /root/.hermes/config.yaml | grep api_key
# Must show both api_key: sk-... and api_key_env: LITELLM_API_KEY # Must show both api_key: sk-... and api_key_env: LITELLM_API_KEY
``` ```
### For Koby (CT 111 / tdunna) ### For Koby (CT 111 / tdunna) — **REPORT-ONLY MODE**
Koby runs Hermes on CT 111 (tdunna). Config files at `/root/.hermes/config.yaml`. Koby runs Hermes on CT 111 (tdunna). Config files at `/root/.hermes/config.yaml`.
Same Hermes pattern as Tanko/Mumuni/Koonimo — see config sections above. Same Hermes pattern as Tanko/Mumuni/Koonimo — see config sections above.
**⛔ KOBY IS NEVER REPAIRED (2026-08-17, Captain)**: Diagnostic only — detect and report, never fix on .129.
No heal step, no restart, no key rotation, no config edit, no memory rewrite, no disk GC, no service touch, no process kill — ever.
If a health check shows Koby degraded, **DO NOT** execute any repair action. Instead, report to Zulip and let Theo fix it.
**LiteLLM key**: alias `koby` in LiteLLM DB, injected via `infisical run --` wrapper. **LiteLLM key**: alias `koby` in LiteLLM DB, injected via `infisical run --` wrapper.
### For pi Agents (Abiba) ### For pi Agents (Abiba)
+40 -36
View File
@@ -5,7 +5,8 @@ description: >
Standard Hermes configuration template for Syslog Solution LLC agents. Standard Hermes configuration template for Syslog Solution LLC agents.
Enforces shared infrastructure setup (Firecrawl, SearXNG, local models, Enforces shared infrastructure setup (Firecrawl, SearXNG, local models,
RA-H OS MCP) while keeping agent-specific API keys and model choices. RA-H OS MCP) while keeping agent-specific API keys and model choices.
UPDATED 2026-08-07: Added Rule 15 (MCP Validation) from the 2026-08-07 keyless-MCP incident. UPDATED 2026-07-16: Compression model is the stable alias `strix-moe` (NOT `ornith-1.0-35b`,
which LiteLLM does not serve). All 3 GPUs verified at 128K (reduced from 256K 2026-07-17 for stability).
Added Rule 12 (Context-Issue Diagnostic) + Rule 13 (.env fallback enforcement) from the Added Rule 12 (Context-Issue Diagnostic) + Rule 13 (.env fallback enforcement) from the
2026-07-16 Mumuni root-cause investigation (WAL #1300). 2026-07-16 Mumuni root-cause investigation (WAL #1300).
UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo. RTX 3090 context verified at 128K. Infisical .env fallback required (Rule 3/13). UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo. RTX 3090 context verified at 128K. Infisical .env fallback required (Rule 3/13).
@@ -47,7 +48,7 @@ Sub-agent profiles inherit auth from the main config — no separate keys needed
| Component | Endpoint | Purpose | | Component | Endpoint | Purpose |
|---|---|---| |---|---|---|
| Firecrawl | `http://192.168.68.7:3002/` | Web content extraction | | Firecrawl | `http://192.168.68.7:3002/` | Web content extraction |
| SearXNG | `http://192.168.68.7:8888` | Privacy-respecting web search | | SearXNG | `http://storepve:8888` | Privacy-respecting web search |
| LiteLLM | `http://192.168.68.116/v1` | Unified model gateway (via nginx) | | LiteLLM | `http://192.168.68.116/v1` | Unified model gateway (via nginx) |
| LiteLLM (NetBird) | `https://litellm.sysloggh.net/v1` | Alternative (may have 502 issues) | | LiteLLM (NetBird) | `https://litellm.sysloggh.net/v1` | Alternative (may have 502 issues) |
| RA-H OS MCP | `http://192.168.68.65:3100/mcp` | Knowledge graph bridge | | RA-H OS MCP | `http://192.168.68.65:3100/mcp` | Knowledge graph bridge |
@@ -96,14 +97,10 @@ work immediately after restart.
model: model:
default: <agent_model> # e.g., strix-moe, qwen3.6-27B-code, syslog-auto default: <agent_model> # e.g., strix-moe, qwen3.6-27B-code, syslog-auto
provider: harness provider: harness
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated path; /v1 also OK base_url: http://192.168.68.116/v1
api_key_env: LITELLM_API_KEY # Injected via infisical run -- wrapper api_key_env: LITELLM_API_KEY # Injected via infisical run -- wrapper
max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation
context_length: 131072 # For syslog-auto (all GPUs at 128K for stability). context_length: 131072 # For syslog-auto (all GPUs at 128K for stability).
# ⚠️ MANDATORY: Hermes probes unknown models from 256K
# and falls back to 256K when /v1/models lacks a context
# field (llama-server does). Without this override, agents
# silently run syslog-auto at 256K (verified 2026-08-09).
# Set 65536 if using gemma-4-12b directly (tight VRAM). # Set 65536 if using gemma-4-12b directly (tight VRAM).
fallback_providers: fallback_providers:
@@ -146,7 +143,7 @@ compression:
# ─── Auxiliary Tasks (CONSISTENCY RULE) ─── # ─── Auxiliary Tasks (CONSISTENCY RULE) ───
# All auxiliary services MUST use identical model, base_url, and api_key_env: # All auxiliary services MUST use identical model, base_url, and api_key_env:
# model: gpu-light # stable alias (NOT raw "gemma-4-12b") # model: gpu-light # stable alias (NOT raw "gemma-4-12b")
# base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK # base_url: http://192.168.68.116/v1
# api_key_env: LITELLM_API_KEY # api_key_env: LITELLM_API_KEY
# Do NOT use syslog-auto for auxiliary tasks — it routes to the primary GPU. # Do NOT use syslog-auto for auxiliary tasks — it routes to the primary GPU.
# gpu-light = RTX 5070 (12B), freeing the Strix Halo for agent reasoning. # gpu-light = RTX 5070 (12B), freeing the Strix Halo for agent reasoning.
@@ -157,20 +154,20 @@ auxiliary:
vision: vision:
provider: harness provider: harness
model: gpu-light # stable alias for RTX 5070 (was raw gemma-4-12b) model: gpu-light # stable alias for RTX 5070 (was raw gemma-4-12b)
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK base_url: http://192.168.68.116/v1
api_key_env: LITELLM_API_KEY api_key_env: LITELLM_API_KEY
timeout: 60 timeout: 60
download_timeout: 30 download_timeout: 30
web_extract: web_extract:
provider: harness provider: harness
model: gpu-light # stable alias for RTX 5070 model: gpu-light # stable alias for RTX 5070
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK base_url: http://192.168.68.116/v1
api_key_env: LITELLM_API_KEY api_key_env: LITELLM_API_KEY
timeout: 30 timeout: 30
compression: compression:
provider: harness provider: harness
model: syslog-auto # MUST match compression.model above. Stable alias for Strix Halo (weighted pool). model: syslog-auto # MUST match compression.model above. Stable alias for Strix Halo (weighted pool).
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated path; /v1 also OK base_url: http://192.168.68.116/v1 # Rule 5: /v1 NOT /litellm/v1
api_key_env: LITELLM_API_KEY api_key_env: LITELLM_API_KEY
timeout: 300 # gpu-fleet: 300s for large-history summarization (was 60) timeout: 300 # gpu-fleet: 300s for large-history summarization (was 60)
@@ -180,14 +177,14 @@ auxiliary:
delegation: delegation:
model: gpu-dense # stable alias for RTX 3090 (was raw qwen3.6-27B-code) model: gpu-dense # stable alias for RTX 3090 (was raw qwen3.6-27B-code)
provider: harness provider: harness
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK base_url: http://192.168.68.116/v1
api_key_env: LITELLM_API_KEY api_key_env: LITELLM_API_KEY
# ─── Custom Provider ─── # ─── Custom Provider ───
custom_providers: custom_providers:
- name: harness - name: harness
model: syslog-auto # weighted pool (default) model: syslog-auto # weighted pool (default)
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK base_url: http://192.168.68.116/v1
api_key_env: LITELLM_API_KEY api_key_env: LITELLM_API_KEY
api_mode: chat_completions api_mode: chat_completions
``` ```
@@ -237,13 +234,10 @@ The following MUST be identical across ALL profiles:
- When main config uses `api_key_env`, sub-agents automatically use it - When main config uses `api_key_env`, sub-agents automatically use it
- This means key rotation only touches ONE vault secret (`LITELLM_API_KEY`) - This means key rotation only touches ONE vault secret (`LITELLM_API_KEY`)
### Rule 5: Main Config Base URL (UPDATED 2026-08-09) ### Rule 5: Main Config Base URL
|- Use the authenticated LiteLLM path: `http://192.168.68.116/litellm/v1` (canonical, captain-approved migration) |- Use direct IP: `http://192.168.68.116/v1`
|- Legacy `http://192.68.68.116/v1` also works — nginx fronts BOTH paths with key auth
(verified 2026-08-09: 401 without key, 200 with key, on both /v1 and /litellm/v1)
|- Both locations have `proxy_read_timeout 600s` (verified in harness-nginx nginx.conf) —
the old "60s timeout on /litellm/" claim was stale and is retracted
|- NOT the NetBird URL (`litellm.sysloggh.net`) — can cause 502 when NetBird is down |- NOT the NetBird URL (`litellm.sysloggh.net`) — can cause 502 when NetBird is down
|- NOT the old path (`/litellm/v1`) — nginx now routes `/v1` directly
### Rule 6: max_tokens Is Required (Thermal Safety) ### Rule 6: max_tokens Is Required (Thermal Safety)
- **Every Hermes config MUST set `model.max_tokens: 4096`** — this is non-negotiable - **Every Hermes config MUST set `model.max_tokens: 4096`** — this is non-negotiable
@@ -264,7 +258,7 @@ The following MUST be identical across ALL profiles:
fall back to other GPUs if Strix gets hot. Both `compression.model` and `auxiliary.compression.model` fall back to other GPUs if Strix gets hot. Both `compression.model` and `auxiliary.compression.model`
MUST be `syslog-auto`. MUST be `syslog-auto`.
- All auxiliary services MUST use identical routing: - All auxiliary services MUST use identical routing:
- `base_url: http://192.168.68.116/litellm/v1` (Rule 5, 2026-08-09: canonical authenticated; `/v1` also OK) - `base_url: http://192.168.68.116/v1` (Rule 5: `/v1`, NOT `/litellm/v1`)
- `api_key_env: LITELLM_API_KEY` - `api_key_env: LITELLM_API_KEY`
- **Do NOT use `syslog-auto`** for auxiliary tasks — it routes unpredictably - **Do NOT use `syslog-auto`** for auxiliary tasks — it routes unpredictably
- **Compression on Strix Halo**: The strix-moe alias routes to Strix Halo - **Compression on Strix Halo**: The strix-moe alias routes to Strix Halo
@@ -324,13 +318,10 @@ verify ALL FOUR of these against the live config. They are the only root causes
1. **max_context_window correct?** — BOTH `compression.max_context_window` AND `context.max_context_window` 1. **max_context_window correct?** — BOTH `compression.max_context_window` AND `context.max_context_window`
MUST be `131072` (all GPUs are 128K). A value of `262144` causes instability near 100K and must NOT be used. MUST be `131072` (all GPUs are 128K). A value of `262144` causes instability near 100K and must NOT be used.
~83K instead of ~170K. Check: `grep -n max_context_window ~/.hermes/config.yaml` ~83K instead of ~170K. Check: `grep -n max_context_window ~/.hermes/config.yaml`
2. **base_url uses authenticated path?** — `custom_providers[0].base_url`, `delegation.base_url`, 2. **base_url uses /v1 NOT /litellm/v1?** — `custom_providers[0].base_url`, `delegation.base_url`,
and ALL `auxiliary.*.base_url` MUST be `http://192.168.68.116/litellm/v1` (Rule 5, canonical) and ALL `auxiliary.*.base_url` MUST be `http://192.168.68.116/v1` (Rule 5). nginx `/litellm/`
or `http://192.168.68.116/v1` (legacy, still authenticated via nginx). BOTH verified 200 with has a 60s default timeout → 504 on any inference >60s; `/v1/` has 600s.
key + 600s proxy_read_timeout on 2026-08-09. Never bare `:4000` direct. Check: `grep -n 'litellm/v1' ~/.hermes/config.yaml` (must return NOTHING)
Check: `grep -nE 'base_url: http://192.168.68.116(:4000)?/v1' ~/.hermes/config.yaml` — the
ONLY paths allowed are `/v1` or `/litellm/v1` (both via nginx :80).
`:4000` or missing `litellm/v1`/`v1` prefix = violation.
3. **LITELLM_API_KEY valid?** — The key must be a real LiteLLM key (`sk-` + 64 hex, 67 chars). 3. **LITELLM_API_KEY valid?** — The key must be a real LiteLLM key (`sk-` + 64 hex, 67 chars).
Malformed values (e.g. `sk-_SWAl_Vu_…`, 47 chars) return 401 → DeepSeek fallback. Malformed values (e.g. `sk-_SWAl_Vu_…`, 47 chars) return 401 → DeepSeek fallback.
Verify: `curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $LITELLM_API_KEY" http://192.168.68.116/v1/models` (must be 200) Verify: `curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $LITELLM_API_KEY" http://192.168.68.116/v1/models` (must be 200)
@@ -347,6 +338,17 @@ curl -s -o /dev/null -w 'key_health: %{http_code}\n' -H "Authorization: Bearer $
### Rule 13: API Key Injection — Two Patterns (UPDATED 2026-07-16, WAL #1300) ### Rule 13: API Key Injection — Two Patterns (UPDATED 2026-07-16, WAL #1300)
### Rule 14: Hermes Context Detection Uses `max_model_tokens`, NOT `max_input_tokens`
**CRITICAL**: Hermes context detection reads `max_model_tokens` (128K), NOT `max_input_tokens` (64K cap).
- **Abiba and Hermes agents**: `max_model_tokens: 131072` (128K) — unlimited context
- **Crewmates (ops, tune, verify, auth-keys, build)**: `max_input_tokens: 64000` (64K) — capped
- If you see `max_input_tokens: 64000` in an Abiba/Hermes config, that's a mistake
- Using `max_input_tokens` for Hermes agents causes premature context loss
- Check: `grep -n 'max_model_tokens\|max_input_tokens' ~/.hermes/config.yaml`
- Expected output: `max_model_tokens: 131072` (not max_input_tokens)
Agents inject `LITELLM_API_KEY` via ONE of two mechanisms. Both are valid; the contract Agents inject `LITELLM_API_KEY` via ONE of two mechanisms. Both are valid; the contract
requirement is that the key is a **valid LiteLLM virtual key** (HTTP 200 on /v1/models). requirement is that the key is a **valid LiteLLM virtual key** (HTTP 200 on /v1/models).
@@ -405,15 +407,6 @@ curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $K" http://192.
- **Audit script**: Run `python3 /root/prose-contracts/audit-hermes-config.py <config.yaml>` - **Audit script**: Run `python3 /root/prose-contracts/audit-hermes-config.py <config.yaml>`
before and after any config change to catch this and all other rule violations. before and after any config change to catch this and all other rule violations.
### Rule 15: MCP Endpoint and Header Validation (ADDED 2026-08-07)
- Every MCP server entry must point at the correct endpoint:
- ra-h-os = http://192.168.68.65:3100/mcp
- litellm = https://litellm.sysloggh.net/mcp
- MCP entries must carry a REAL key value in the header.
- Avoid using env-var names like LITELLM_API_KEY in the header; they do not resolve for MCP
endpoints and result in "Malformed API Key" floods.
- Ensure the header value is the actual key (e.g., `sk-...`).
## Execution ## Execution
1. **Check current config** — Read the target agent's config.yaml 1. **Check current config** — Read the target agent's config.yaml
@@ -423,3 +416,14 @@ curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $K" http://192.
5. **Set model choice** — Per agent's workload 5. **Set model choice** — Per agent's workload
6. **Verify** — curl all shared endpoints, test the model with the new key 6. **Verify** — curl all shared endpoints, test the model with the new key
7. **Report** — What was changed, preserved, custom 7. **Report** — What was changed, preserved, custom
### Rule 17: Koby Report-Only (URGENT — 2026-08-17, Captain)
**Koby (CT 111, .129, tdunna/Theo) is report-only, full stop.**
No heal step, no restart, no key rotation, no config edit, no memory rewrite,
no disk GC, no service touch, no process kill — **ever**.
Do NOT "repair" Koby under any circumstance, even if a health check shows it degraded.
Detect + report to Zulip (Theo fixes it).
This supersedes every other instruction in this file.
If a rule would trigger a Koby repair action, SKIP it and emit a diagnostic alert.
+13 -14
View File
@@ -34,12 +34,11 @@ Syslog is migrating away from **unauthenticated direct access** to the shared in
| Path | Auth | Status | | Path | Auth | Status |
|------|------|--------| |------|------|--------|
| `http://192.168.68.116/v1` | Bearer `sk-*` key (nginx-fronted) | ✅ **VALID** — authenticated via nginx :80 (verified 2026-08-09: 401 without key, 200 with) | | `http://192.168.68.116/v1` | None (direct) | ❌ **DEPRECATED** — being phased out |
| `http://192.168.68.116/litellm/v1` | Bearer `sk-*` key (nginx-fronted) | ✅ **CURRENT / CANONICAL** — captain-approved migration target; 600s proxy_read_timeout (verified) | | `http://192.168.68.116/litellm/v1/responses` | Bearer `sk-*` key | ✅ **CURRENT** — authenticated LiteLLM proxy |
| `http://192.168.68.116:4000/v1` | Bearer `sk-*` key (direct container) | ❌ **FORBIDDEN** — bypasses nginx; port 4000 direct is not a config path |
All harness/litellm providers MUST use an authenticated nginx-fronted path (`/litellm/v1` canonical, `/v1` legacy-valid). All harness/litellm providers MUST use the authenticated `/litellm/v1/responses` path.
Any `base_url` pointing at `:4000` or a bare IP without nginx is a **migration violation**. Any `base_url` pointing to bare `/v1` on 192.168.68.116 is a **migration violation**.
### 🔥 CRITICAL: Double-Path Bug (2026-07-10) ### 🔥 CRITICAL: Double-Path Bug (2026-07-10)
@@ -191,22 +190,22 @@ litellm_settings:
|-------|-----|-----|---------------|------------|--------|-----------------|---------------| |-------|-----|-----|---------------|------------|--------|-----------------|---------------|
| Tanko | 112 | .122 | `tanko` | Infisical vault | ✅ Fixed | `infisical run` | 20:17 UTC Jul 5 | | Tanko | 112 | .122 | `tanko` | Infisical vault | ✅ Fixed | `infisical run` | 20:17 UTC Jul 5 |
| Mumuni | 100 (abiba) | .24 | `mumuni` | Infisical vault | ✅ Fixed | Pi Hermes gateway | 2026-07-27 | | Mumuni | 100 (abiba) | .24 | `mumuni` | Infisical vault | ✅ Fixed | Pi Hermes gateway | 2026-07-27 |
| Koby | 111 | .129 | `koby` | Infisical vault | ✅ Fixed | `infisical run` | 23:30 UTC Jul 5 | | Koby | 111 | ? | `koby` | Infisical vault | ✅ Fixed | `infisical run` | 23:30 UTC Jul 5 |
| Koonimo | 113 | .114 | `koonimo` | Infisical vault | ✅ Fixed | `infisical run` (migrated 2026-07-11) | 2026-08-09 | | Koonimo | 113 | ? | `koonimo` | Infisical vault | ✅ Fixed | `infisical run` (migrated 2026-07-11) | 2026-07-11 |
| Abiba | 100 | .65 | `abiba-pi` | Infisical vault | ✅ N/A (pi native) | — | 19:44 UTC Jul 5 | | Abiba | 100 | .65 | `abiba-pi` | Infisical vault | ✅ N/A (pi native) | — | 19:44 UTC Jul 5 |
| Kagenz0 | 105 | .14 | — | — | ❌ DOWN | — | 19:14 EDT Jul 4 | | Kagenz0 | 105 | ? | — | — | ❌ DOWN | — | 19:14 EDT Jul 4 |
> **Note**: CT hostnames (tdunna, baggy) differ from agent identities (koby, koonimo). > **Note**: CT hostnames (tdunna, baggy) differ from agent identities (koby, koonimo).
> LiteLLM key aliases use agent identity, not CT hostname. > LiteLLM key aliases use agent identity, not CT hostname.
### Migration Status: Authenticated Path ### Migration Status: Authenticated Path
| Agent | `/litellm/v1` | Legacy `/v1` | Status | | Agent | `/litellm/v1/responses` | Deprecated `/v1` | Status |
|-------|--------------|-------------|--------| |-------|--------------------------|--------------------|--------|
| Mumuni | ✅ harness provider | ✅ auxiliary on /v1 (valid) | ✅ Authenticated (verified 2026-08-09) | | Mumuni | ✅ 5 sections | 0 | ✅ Authenticated |
| Tanko | ✅ 5 sections | 0 | ✅ Migrated 2026-08-08, keys 200 | | Tanko | ⚠️ No SSH access | — | Needs check |
| Koby | ✅ custom provider (harness name) | — | ✅ External DeepSeek primary (intentional) | | Koby | ⚠️ No route to host | — | Needs check |
| Koonimo | ✅ .114 (baggy) | — | ✅ 128K context applied 2026-08-09 | | Koonimo | ⚠️ Connection timed out | — | Needs check |
### Systemd Service Pattern (2026-07-11 — vault migration) ### Systemd Service Pattern (2026-07-11 — vault migration)
+4
View File
@@ -1,4 +1,6 @@
--- ---
report_only_agents:
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
kind: function kind: function
name: hermes-zulip-restore name: hermes-zulip-restore
description: > description: >
@@ -12,6 +14,7 @@ version: 1.0.0
status: active status: active
runtime_contract: 2 runtime_contract: 2
--- ---
---
# Hermes Zulip Restore — Bring Any Agent Back to Good State # Hermes Zulip Restore — Bring Any Agent Back to Good State
@@ -184,6 +187,7 @@ https://git.sysloggh.net/SyslogSolution/zulip-platform-plugins/src/branch/feat/z
Commit `55ca15d` — `fix(zulip): add _strip_html for slash command matching` Commit `55ca15d` — `fix(zulip): add _strip_html for slash command matching`
Pull request #33 is the primary integration branch. Pull request #33 is the primary integration branch.
---
--- ---
**Last verified good state**: 2026-07-08 — Mumuni, Tanko, Koby all connected with `_strip_html` applied. **Last verified good state**: 2026-07-08 — Mumuni, Tanko, Koby all connected with `_strip_html` applied.
+79 -8
View File
@@ -7,15 +7,13 @@ description: >
from nvidia-smi (.8, .110) and amdgpu_top (.15). LiteLLM metrics from nvidia-smi (.8, .110) and amdgpu_top (.15). LiteLLM metrics
via existing /metrics Prometheus endpoint. via existing /metrics Prometheus endpoint.
DEPLOYMENT STATUS (2026-08-09): DEPLOYMENT STATUS (2026-07-09):
✅ Core stack deployed: Prometheus + Grafana + pve/node/docker exporters ✅ Core stack deployed: Prometheus + Grafana + pve/node/docker exporters
(via proxmox-monitor contract). Grafana at :3001, all scrape targets active. (via proxmox-monitor contract). Grafana at :3001, 5 scrape targets active.
✅ GPU exporters DEPLOYED: all 3 GPU hosts (.8/.110/.15) run exporters on ❌ GPU exporters NOT deployed: gpu-exporter crash-loops on .15,
:9400 (nvidia_gpu_exporter / amdgpu exporter) — verified 200 on 2026-08-09. NVIDIA sidecar exporters (.8/.110:9400) never installed.
✅ LiteLLM /metrics scraping live (success_callback: prometheus; auth via Router falls back to direct GPU /health probes.
master key) + Alertmanager + Zulip bridge (alerts-infra) added 2026-08-09. ⚠️ This contract is target-state aspirational — not as-built.
⚠️ This contract is target-state aspirational — but GPU export + alerting
are now as-built (verified 2026-08-09).
As-built GPU monitoring is via gpu-monitor contract (port 9100 poll). As-built GPU monitoring is via gpu-monitor contract (port 9100 poll).
version: 1.0.0 version: 1.0.0
--- ---
@@ -141,6 +139,79 @@ GPU .8 (RTX 3090) GPU .110 (RTX 5070) GPU .15 (Strix Halo)
4. Verify LiteLLM metrics flowing to Prometheus 4. Verify LiteLLM metrics flowing to Prometheus
5. ~~Update nginx to proxy `/monitoring/` → Grafana~~ (NOT recommended — nginx sub-path was tried for /grafana/ and reverted per proxmox-monitor; direct :3001 access is the standard) 5. ~~Update nginx to proxy `/monitoring/` → Grafana~~ (NOT recommended — nginx sub-path was tried for /grafana/ and reverted per proxmox-monitor; direct :3001 access is the standard)
## Execution
### check-health
**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.**
```bash
# Zulip API health (POST ping)
curl -s -o /dev/null -w '%{http_code}' -X POST https://chat.sysloggh.net/api/v1/messages -u 'abiba-bot@chat.sysloggh.net:KEY'
# Expected: 200 (HTTP 000 = unreachable/cache)
# PM2 process health
pm2 jlist
# Expected: 5/5 online (abiba-telegram, abiba-zulip, zulip-watchdog, gitea-runner, spoton-service)
# GPU exporters (may be down per DEPLOYMENT STATUS)
curl -s http://192.168.68.8:9400/metrics && echo " - OK" || echo " - FAIL"
curl -s http://192.168.68.110:9400/metrics && echo " - OK" || echo " - FAIL"
curl -s http://192.168.68.15:9400/metrics && echo " - OK" || echo " - FAIL"
# Prometheus targets
curl -s http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets'
# Expected: All targets UP (may show some down if exporters not deployed)
# Grafana health
curl -s http://192.168.68.116:3001/api/health | jq '{status, version}'
# Expected: {"status":"ok","version":"..."}
# LiteLLM metrics
curl -s http://192.168.68.116:4001/metrics | head -20
# Expected: Prometheus-formatted metrics output
```
**Report format**: Summarize actual results from each probe. If any probe returns non-200 or empty output, flag as alert.
### Phase 1: GPU Exporters
**NVIDIA (.8 and .110)**:
1. Download `nvidia_gpu_exporter` binary
2. Create systemd service `nvidia-gpu-exporter.service`
3. Start and enable
**AMD (.15)**:
1. Create Python exporter script at `/opt/amdgpu-exporter/exporter.py`
2. Parses `amdgpu_top --json -d 1000` output
3. Exposes key metrics at `:9400/metrics` via Python http.server
4. Create systemd service
5. Start and enable
### Phase 2: Prometheus
1. Create `/opt/monitoring/` directory on CT 116
2. Write `prometheus.yml` with scrape configs for all targets
3. Add to docker-compose (or separate compose file)
4. Start container
### Phase 3: Grafana
1. Create `/opt/monitoring/grafana/` directories
2. Provision Prometheus datasource
3. Provision GPU fleet dashboard JSON
4. Provision LiteLLM dashboard JSON
5. Add to docker-compose
6. Start container
### Phase 4: Verification
1. Verify all 3 GPU exporters return 200 at :9400/metrics
2. Verify Prometheus targets all UP at :9090/targets
3. Verify Grafana accessible at :3001 with dashboards
4. Verify LiteLLM metrics flowing to Prometheus
5. ~~Update nginx to proxy `/monitoring/` → Grafana~~ (NOT recommended — nginx sub-path was tried for /grafana/ and reverted per proxmox-monitor; direct :3001 access is the standard)
## Verification Commands ## Verification Commands
```bash ```bash
-9
View File
@@ -18,15 +18,6 @@ description: >
Designed as a reusable contract for any Syslog agent. Designed as a reusable contract for any Syslog agent.
Source of truth: gpu-fleet.prose.md Source of truth: gpu-fleet.prose.md
## Monitoring / Alerting (as-built 2026-08-09)
- LiteLLM /metrics scrape job: requires `litellm_settings.success_callback: [prometheus]`
(failure_callback alone does NOT mount /metrics — verified 2026-08-09).
Scraped by Prometheus with Bearer master key; endpoint returns 307 → /metrics/.
- Alertmanager (harness-alertmanager :9093) + zulip-bridge (:9102) deliver
firing alerts to #agent-hub > alerts-infra via abiba-bot. Added 2026-08-09.
- Prometheus node job covers ALL 6 PVE nodes (.4/.5/.6/.9/.12/.15:9100).
--- ---
## Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU) ## Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU)
+16 -1
View File
@@ -1,4 +1,6 @@
--- ---
report_only_agents:
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
kind: responsibility kind: responsibility
name: litellm-self-heal name: litellm-self-heal
status: deployed status: deployed
@@ -22,6 +24,7 @@ description: >
inference, and agent keys. Applies remediation rules for common failures. inference, and agent keys. Applies remediation rules for common failures.
Reports every action via Zulip DM and Gitea (SyslogSolution/health-logs). Reports every action via Zulip DM and Gitea (SyslogSolution/health-logs).
--- ---
---
# LiteLLM Operations — Health Check + Self-Heal # LiteLLM Operations — Health Check + Self-Heal
@@ -65,7 +68,15 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
## LiteLLM Model Surface (ground truth — `/opt/inference-harness/litellm_config.yaml` on CT 116) ## LiteLLM Model Surface (ground truth — `/opt/inference-harness/litellm_config.yaml` on CT 116)
`model_name`s served: `qwen3.6-27B-code`, `gemma-4-12b`, `qwen3.6-35B-udq4`, `strix-moe`, `gpu-dense`, `gpu-light`, `syslog-auto`. `model_name`s served: `qwen3.6-27B-code`, `gemma-4-12b`, `qwen3.6-35B-udq4`, `strix-moe`, `gpu-dense`, `gpu-light`, `syslog-auto`, `crew-auto` (new 2026-08-20).
### Context Cap Split (2026-08-20)
- **Abiba (firstmate)**: 128K uncapped — unlimited context for primary workloads
- **Hermes agents** (mumuni, tanko, koby, koonimo): 128K uncapped
- **Crewmates** (ops, tune, verify, auth-keys, build): 64K capped — alias `crew-auto` enforces 64K limit
Preferred implementation: uncap shared pool, add capped alias for crew-only.
- `syslog-auto` is a weighted router model: qwen3.6-27B-code (0.55, rpm 500) + qwen3.6-35B-udq4 (0.30, rpm 60) + gemma-4-12b (0.15, rpm 200). - `syslog-auto` is a weighted router model: qwen3.6-27B-code (0.55, rpm 500) + qwen3.6-35B-udq4 (0.30, rpm 60) + gemma-4-12b (0.15, rpm 200).
- `gpu-dense` / `gpu-light` are high-rpm aliases (rpm 500) onto qwen3.6-27B-code / gemma-4-12b respectively. - `gpu-dense` / `gpu-light` are high-rpm aliases (rpm 500) onto qwen3.6-27B-code / gemma-4-12b respectively.
@@ -120,6 +131,7 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
- Also wakes on user request - Also wakes on user request
- On failure: re-check after 30s, escalate after 3 consecutive failures - On failure: re-check after 30s, escalate after 3 consecutive failures
---
--- ---
## Health Check ## Health Check
@@ -162,6 +174,7 @@ Determine overall_status from individual check results:
- "degraded" — 1-2 non-critical checks fail - "degraded" — 1-2 non-critical checks fail
- "down" — critical checks fail - "down" — critical checks fail
---
--- ---
## Remediation Rules ## Remediation Rules
@@ -205,6 +218,7 @@ Escalate → if SSH access unavailable, send Zulip DM
Router no longer in path so Redis active counters are unused. Rule retained Router no longer in path so Redis active counters are unused. Rule retained
for reference but inactive. If Redis issues occur, check harness-redis container. for reference but inactive. If Redis issues occur, check harness-redis container.
---
--- ---
## Reporting ## Reporting
@@ -227,6 +241,7 @@ top actions, uptime.
If a fix requires another agent (e.g., Authentik restart), relay sent If a fix requires another agent (e.g., Authentik restart), relay sent
to responsible agent with full context. to responsible agent with full context.
---
--- ---
## Execution ## Execution
+4 -1
View File
@@ -1,9 +1,12 @@
--- ---
report_only_agents:
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
name: memory-audit-maintenance name: memory-audit-maintenance
kind: responsibility kind: responsibility
description: Shared memory audit and maintenance contract for all Hermes agents (Mumuni, Tanko, Koby, Koonimo). Each agent runs it against its own isolated memory files — no cross-agent access, no shared state. Detects staleness, enforces writer registry, and rotates canary tokens. description: Shared memory audit and maintenance contract for all Hermes agents (Mumuni, Tanko, Koby, Koonimo). Each agent runs it against its own isolated memory files — no cross-agent access, no shared state. Detects staleness, enforces writer registry, and rotates canary tokens.
id: 067NC4KG01RG50R40M30E20918 id: 067NC4KG01RG50R40M30E20918
--- ---
---
### Goal ### Goal
@@ -343,4 +346,4 @@ return {
### Per-Agent Notes ### Per-Agent Notes
Each Hermes agent (Mumuni, Tanko, Tdunna, Baggy) runs this contract against its own `~/.hermes/memories/` directory. The contract is identical across agents, but all data is fully isolated: separate ledgers, separate writer registries, separate canaries. If a new agent is added to the roster, it must be listed in `### Scope` above and given its own isolated memory directory. Each Hermes agent (Mumuni, Tanko, Tdunna, Baggy) runs this contract against its own `~/.hermes/memories/` directory. The contract is identical across agents, but all data is fully isolated: separate ledgers, separate writer registries, separate canaries. If a new agent is added to the roster, it must be listed in `### Scope` above and given its own isolated memory directory.
+47 -150
View File
@@ -1,176 +1,73 @@
--- ---
report_only_agents:
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
kind: pattern kind: pattern
name: memory-fixer name: memory-fixer
description: > description: >
Auto-fix low-hanging fruit in the RA-H OS knowledge graph. No judgment calls — only deterministic Level 1 operations. Auto-fix low-hanging fruit in the graph. No judgment calls — only deterministic Level 1 operations.
Escalate anything that needs Kwame's input. Executes confirmed Kwame decisions to completion (state + updated_at). Escalate anything that needs Kwame's input.
version: 2.0.0 version: 1.1.0
---
--- ---
# Memory Fixer # Memory Fixer
> **Canonical copy:** `/root/.hermes/contracts/memory-fixer-v3.md` (used by the `memory-fixer-daily` cron job). This file is the institutional record of the same contract. When the two diverge, treat the v3 source in `/root/.hermes/contracts/` as executable truth.
## Purpose ## Purpose
Auto-fix low-hanging fruit in the graph. No judgment calls — only deterministic Level 1 operations. Escalate anything that needs Kwame's input. When Kwame replies to an escalation, **execute the decision to completion** (update state and timestamps), never leaving a node in review-pending forever. Auto-fix low-hanging fruit in the graph. No judgment calls — only deterministic Level 1 operations. Escalate anything that needs Kwame's input.
**Type:** Write-only (Level 1 fixes only) ## Level 0 Auto-Deletes (Allowed Without Approval)
**Scope:** RA-H OS knowledge graph (192.168.68.65) Ephemeral heartbeat and log nodes that violate "Logs NEVER go in the graph":
**Schedule:** Daily at 8 AM ET
**Escalation:** Level 2+ to Kwame as task items
## Key Design Decision - `[LITELLM-HEALTH]`, `[GPU-SELF-HEAL]`, `[PM2-SELF-HEAL]`
- `[PROXMOX-MONITOR]`, `[GPU-MONITOR]`, `[INFRA-MONITOR]`, `[AGENT-HEALTH]`, `[DISK-GC]`
- `[WAL]` entries older than 30 days
The `updateNode` tool's `metadata` field performs a **restricted merge** — the `state` key only accepts `'processed'` or `'not_processed'`. Additionally, new metadata keys cannot be added via the merge. **Condition:** node must be an orphan (no edges). Deleting a connected node risks breaking other nodes.
**Solution:** Use the `description` field to tag stale nodes with review actions, since `description` is a simple string overwritable via `updateNode`. **Method:** direct SQLite on `.65` (MCP has no delete tool):
```bash
**Tag Format:** `[REVIEW: action] original description text...` ssh root@192.168.68.65 "sqlite3 /root/.local/share/RA-H/db/rah.sqlite \"
DELETE FROM nodes WHERE id IN (
Where `action` is one of: SELECT id FROM nodes WHERE id NOT IN (SELECT from_node_id FROM edges)
- `archive` — node is stale and should be archived AND id NOT IN (SELECT to_node_id FROM edges)
- `refresh` — node is stale and should be refreshed (infrastructure) AND title LIKE '[LITELLM-HEALTH]%' -- add more prefixes as needed
- `keep` — node has been confirmed as current );\""
- `merge` — node is a duplicate candidate
**Query for finding review-tagged nodes:**
```sql
SELECT id, title, description
FROM nodes
WHERE description LIKE '[REVIEW:%';
``` ```
## Level 1 Auto-Fixes (No Kwame Decision Needed) ## Level 1 Auto-Fixes (No Judgment Required)
### 1. Missing `type` Auto-Classification ### 1. Missing `type` Field
For nodes with content but no `metadata.type`:
- Title contains "Proxmox" or "infrastructure" → `type: infrastructure`
- Title contains "skill" or "how to" or "guide" → `type: skill`
- Title contains "doc" or "template" or "brand" → `type: documentation`
- Title starts with "WAL:" or "TASK:" → `type: note`
- Title starts with "[LEARN]" → `type: documentation`
- Otherwise → `type: note` (default)
### 2. Missing `tenant` / `namespace`
For any node with NULL tenant or namespace:
```sql ```sql
SELECT id, title, UPDATE nodes
CASE SET metadata = json_set(
WHEN title LIKE '%infrastructure%' OR title LIKE '%proxmox%' OR title LIKE '%setup%' THEN 'infrastructure' COALESCE(metadata, '{}'),
WHEN title LIKE '%skill%' OR title LIKE '%how to%' OR title LIKE '%guide%' THEN 'skill' '$.tenant', 'syslogsolution',
WHEN title LIKE '%doc%' OR title LIKE '%template%' OR title LIKE '%brand%' THEN 'documentation' '$.namespace', 'syslogsolution'
WHEN title LIKE 'WAL:%' OR title LIKE 'TASK:%' THEN 'note'
WHEN title LIKE '%[LEARN]%' THEN 'documentation'
ELSE 'note'
END as auto_type
FROM nodes
WHERE json_extract(metadata, '$.type') IS NULL;
```
### 2. Missing `namespace` Auto-Population
```sql
SELECT id, title, json_extract(metadata, '$.tenant') as tenant
FROM nodes
WHERE json_extract(metadata, '$.namespace') IS NULL;
```
### 3. Staleness Review Tagging
Using the type-based windows from the memory-monitor contract, tag nodes stale beyond their window. **Only process a maximum of 10 nodes per run** to avoid overwhelming Kwame. Prioritize infrastructure first, then dynamic, then ephemeral.
**Exclusion Rules:**
- Nodes with `state` = `review_pending`, `deprecated`, `archived`, or `not_processed` are NOT processed
- Nodes whose `description` already starts with `[REVIEW:` are NOT re-processed
```sql
SELECT id, title, json_extract(metadata, '$.type') as node_type,
CAST(julianday('now') - julianday(updated_at) AS INTEGER) as days_stale,
CASE
WHEN json_extract(metadata, '$.type') IN ('infrastructure', 'deployment', 'system', 'system-health') THEN 'refresh'
WHEN json_extract(metadata, '$.type') IN ('business', 'philosophy', 'research', 'learning', 'learn', 'investigation', 'analysis', 'project') THEN 'refresh'
ELSE 'archive'
END as suggested_action
FROM nodes
WHERE updated_at < datetime('now',
CASE
WHEN json_extract(metadata, '$.type') IN ('infrastructure', 'deployment', 'system', 'system-health') THEN '-14 days'
WHEN json_extract(metadata, '$.type') IN ('skill', 'documentation', 'template', 'protocol-enforcement', 'prd', 'architecture') THEN '-90 days'
WHEN json_extract(metadata, '$.type') IN ('note', 'wal', 'WAL', 'task', 'TASK', 'event') THEN '-30 days'
WHEN json_extract(metadata, '$.type') IN ('business', 'philosophy', 'research', 'learning', 'learn', 'investigation', 'analysis', 'project') THEN '-120 days'
WHEN json_extract(metadata, '$.type') IN ('deprecated-relay', 'audit', 'audit-report', 'incident', 'incident-report') THEN '-3650 days'
ELSE '-45 days'
END
) )
AND json_extract(metadata, '$.state') NOT IN ('review_pending', 'deprecated', 'archived', 'not_processed') WHERE json_extract(metadata, '$.tenant') IS NULL
AND (description IS NULL OR description NOT LIKE '[REVIEW:%') OR json_extract(metadata, '$.namespace') IS NULL;
ORDER BY days_stale ASC
LIMIT 10;
``` ```
For each identified node, call `updateNode(id, { description: "[REVIEW: action] " + originalDescription })`. ### 3. Staleness State Transitions
Using the type-based windows from the memory-monitor contract:
- Nodes stale > their window → transition to `state: review_pending`
- Nodes in `review_pending` for >7 days → escalate to Kwame (Level 2)
## Level 2 Escalations (Kwame Decision Required) ## Level 2 Escalations (Kwame Decision Required)
1. **Nodes in `review_pending` >7 days** — Archive, refresh, or keep?
1. **Stale nodes** flagged with `[REVIEW: …]` — Archive, refresh, or keep? 2. **Orphan Nodes >90 days old** — Delete or Connect?
2. **Duplicate Nodes** (same title or >70% title overlap) — Merge or keep? 3. **Potential Duplicate Nodes** — Same title or >70% overlap. Merge or Keep?
3. **Orphan Nodes >90 days old** — Archive or connect? 4. **Conflicting Metadata** — Content suggests one tenant but metadata says another.
## Reporting Format
The fixer reports to Kwame via this Zulip DM:
```
🦅 Memory Fixer — [HH:MM UTC]
Level 1 fixes applied:
- Missing type: X nodes classified
- Missing namespace: Y nodes populated
Stale nodes needing review (max 10):
1. [Node #XXX] Title — X days stale, SUGGEST: refresh
2. [Node #YYY] Title — Y days stale, SUGGEST: archive
...
Duplicates needing decision:
1. [Node #AAA] vs [Node #BBB] — Same title
Orphans >90 days:
1. [Node #EEE] Title — X days stale, orphaned
Reply with:
- "archive #XXX, #YYY" to mark for archive
- "archive all" to archive all stale nodes listed
- "keep #XXX" to confirm a node is current
- "merge #AAA into #BBB" to merge duplicates
- "refresh #XXX" to mark as current
```
## Execution on Next Run
The fixer reads Kwame's previous response and **executes the decision to completion** — it must not leave a node in review-pending forever. Tagging alone is NOT enough; each confirmed decision must also update `state` and `updated_at` so the node drops out of the stale window on the next run.
> ⚠️ `updateNode` cannot set `state` to non-standard values (restricted to `processed`/`not_processed`) and cannot add metadata keys. For state transitions and `updated_at` bumps, use **direct SSH + SQLite** on the bridge host:
> ```bash
> ssh root@192.168.68.65 "sqlite3 /root/.local/share/RA-H/db/rah.sqlite \"UPDATE nodes SET metadata = json_set(metadata, '$.state', '<state>'), updated_at = datetime('now') WHERE id = <id>;\""
> ```
> Use `updateNode` only for description/source/title/link edits.
Decision → completed action mapping:
| Kwame reply | Description change | State | `updated_at` |
|---|---|---|---|
| `archive #XXX` | replace `[REVIEW: archive] ` → `[ARCHIVED] ` prefix | `archived` | bumped to now |
| `keep #XXX` / `refresh #XXX` | **clear the `[REVIEW: …]` tag entirely** | `active` | bumped to now |
| `merge #AAA into #BBB` | set `[REVIEW: merge_into #BBB]` on #AAA, then follow manual merge workflow | handled manually | bumped to now |
| `archive all` | apply the archive row to every node listed in the prior report | `archived` | bumped to now |
**Why `updated_at` must be bumped (critical):** the Level-1 staleness query keys off `updated_at < now - window`. If the fixer clears the tag but leaves a stale `updated_at`, the node is immediately re-flagged on the very next run and the cycle repeats forever. Bumping `updated_at` to now pushes the node back to the front of the window.
**Exclusion after action:** once an action is applied, the node's description no longer starts with `[REVIEW:` (archive → `[ARCHIVED]`, refresh/keep → original text), so it is not re-processed.
After all actions are applied, verify with:
```sql
SELECT id, json_extract(metadata, '$.state') FROM nodes WHERE description LIKE '[REVIEW:%';
```
The result must be 0 rows when all decisions are executed. Report what was done.
## Checks
- **State integrity:** archived nodes have `state: archived` + `[ARCHIVED]` prefix; kept nodes are `state: active` without a `[REVIEW:]` tag.
- **No review-pending forever:** after executing Kwame's decisions, `[REVIEW:%` node count must be 0.
- **Timestamps:** every executed decision bumps `updated_at`, so the node exits the stale window on the next run.
## Logging ## Logging
Every Level 1 fix logged to `~/.hermes/logs/memory-fixer/YYYY-MM-DD.md` Every Level 1 fix logged to `~/.hermes/logs/memory-fixer/YYYY-MM-DD.md`
+10 -18
View File
@@ -2,31 +2,23 @@
kind: responsibility kind: responsibility
name: pm2-self-heal name: pm2-self-heal
description: > description: >
Monitors critical PM2 processes (abiba-zulip, abiba-telegram, gitea-runner, Monitors critical PM2 processes (abiba-zulip, abiba-telegram) and
spoton-service, zulip-watchdog) and auto-restarts any that are stopped or auto-restarts any that are stopped or errored. Logs every action to
errored. Logs every action to Gitea (SyslogSolution/health-logs — not the knowledge graph and alerts the owner via Zulip DM on failures.
knowledge graph, hard rule) and alerts the owner via Abiba-zulip is the live Zulip bridge and may be restarted; alert owner on failure.
Zulip DM on failures. report_only_agents:
CRITICAL: Never restart abiba-zulip — it runs this contract. - koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
AS-BUILT 2026-08-09 (captain ruling, ecosystem is authoritative):
gpu-monitor is systemd-managed (gpu-monitor.service) — NOT PM2;
gpu-watchdog decommissioned (function folded into gpu-monitor.service);
gitea-runner KEPT (online in PM2); abiba-zulip KEPT (online 4d+, the
2026-07-04 'removed/decommissioned' note was stale and is removed).
--- ---
## Maintains ## Maintains
- abiba-telegram: { status: "online", uptime: string, restarts: number } - abiba-telegram: { status: "online", uptime: string, restarts: number }
- abiba-zulip: { status: "online", uptime: string, restarts: number } - abiba-zulip: { status: "online", uptime: string, restarts: number }
- gpu-monitor: { status: "online", uptime: string, restarts: number } (systemd-managed, PM2-tracked)
- gitea-runner: { status: "online", uptime: string, restarts: number } - gitea-runner: { status: "online", uptime: string, restarts: number }
- spoton-service: { status: "online", uptime: string, restarts: number }
- zulip-watchdog: { status: "online", uptime: string, restarts: number }
- last_check: timestamp - last_check: timestamp
> **Note (2026-07-04, SUPERSEDED 2026-08-09):** `abiba-zulip` remains ONLINE and > **Status (2026-08-03):** `abiba-zulip` fully restored — live Zulip bridge, heartbeating, monitored. `gpu-monitor` runs via systemd (stable); PM2 tracks it for status reporting only. `gpu-watchdog` retired from PM2.
> is monitored — the decommission note was stale (process re-added; do not treat
> it as removed).
## Continuity ## Continuity
@@ -58,9 +50,9 @@ description: >
- If status is "online" → pass - If status is "online" → pass
- If status is "stopped" or "errored" → apply Rule 1 - If status is "stopped" or "errored" → apply Rule 1
- If restarts > 5 → alert owner - If restarts > 5 → alert owner
3. **Check abiba-zulip** (self-process, read-only): 3. **Check abiba-zulip** (live Zulip bridge, heartbeating):
- If status is "online" → pass, log restarts count - If status is "online" → pass, log restarts count
- If status is "stopped" or "errored" → **DO NOT RESTART** — alert owner immediately - If status is "stopped" or "errored" → restart (`pm2 restart abiba-zulip` — fully restored)
- If restarts > 5 in last hour → alert owner with full diagnostics - If restarts > 5 in last hour → alert owner with full diagnostics
4. **Log results** — Append to `SyslogSolution/health-logs/pm2/{timestamp}.md` in Gitea (not knowledge graph — hard rule) 4. **Log results** — Append to `SyslogSolution/health-logs/pm2/{timestamp}.md` in Gitea (not knowledge graph — hard rule)
5. **Alert** — Send Zulip DM to owner if escalation needed (do NOT run pm2 commands during alerting) 5. **Alert** — Send Zulip DM to owner if escalation needed (do NOT run pm2 commands during alerting)
+19 -13
View File
@@ -40,10 +40,10 @@ PVE_NODES = {
# Agent definitions: ct, host, user, pve_node, vault_key_name # Agent definitions: ct, host, user, pve_node, vault_key_name
AGENTS = { AGENTS = {
"tanko": {"ct": 112, "host": "192.168.68.122", "user": "jerome", "pve": "amdpve", "vault_key": "TANKO_LITELLM_API_KEY"}, "tanko": {"ct": 112, "host": "192.168.68.122", "user": "jerome", "pve": "amdpve", "vault_key": "TANKO_LITELLM_API_KEY", "report_only": False},
"abiba": {"ct": 100, "host": "192.168.68.24", "user": "root", "pve": "hwepve", "vault_key": None}, # Pi agent + Mumuni Zulip, no vault key "abiba": {"ct": 100, "host": "192.168.68.24", "user": "root", "pve": "hwepve", "vault_key": None, "report_only": False},
"koby": {"ct": 111, "host": "192.168.68.129", "user": "root", "pve": "amdpve", "vault_key": "KOBY_LITELLM_API_KEY"}, "koby": {"ct": 111, "host": "192.168.68.129", "user": "root", "pve": "amdpve", "vault_key": "KOBY_LITELLM_API_KEY", "report_only": True}, # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17)
"koonimo": {"ct": 113, "host": "192.168.68.114", "user": "root", "pve": "amdpve", "vault_key": "KOONIMO_LITELLM_API_KEY"}, "koonimo": {"ct": 113, "host": "192.168.68.114", "user": "root", "pve": "amdpve", "vault_key": "KOONIMO_LITELLM_API_KEY", "report_only": False},
} }
GPU_HOSTS = { GPU_HOSTS = {
@@ -239,20 +239,26 @@ def check_agents():
host = agent.get("host") host = agent.get("host")
user = agent.get("user") user = agent.get("user")
ct = agent["ct"] ct = agent["ct"]
report_only = agent.get("report_only", False)
if not host or not user: if not host or not user:
print(f" ⬜ {name} (CT {ct}): cannot SSH — skip liveness check") print(f" ⬜ {name} (CT {ct}): cannot SSH — skip liveness check")
continue continue
# Gateway process # ⛔ KOBY IS NEVER REPAIRED — diagnostic only
pid = ssh(host, "pgrep -f 'hermes_cli.main gateway run' | grep -v infisical | head -1", user=user) if report_only:
if not pid: print(f" 🔍 {name}: REPORT-ONLY mode (diagnostic only, no repairs on .129)")
# Try alternate binary name # Still check gateway status for reporting purposes
pid = ssh(host, "pgrep -f 'hermes.*gateway' | grep -v infisical | grep -v bash | head -1", user=user) pid = ssh(host, "pgrep -f 'hermes_cli.main gateway run' | grep -v infisical | head -1", user=user)
if not pid: if not pid:
print(f" ❌ {name}: GATEWAY NOT RUNNING") pid = ssh(host, "pgrep -f 'hermes.*gateway' | grep -v infisical | grep -v bash | head -1", user=user)
FAIL.append(f"gateway-down:{name}") if not pid:
continue print(f" ⚠️ {name}: GATEWAY NOT RUNNING (reported only)")
FAIL.append(f"gateway-down:{name}")
continue
else:
print(f" ✅ {name}: gateway running (pid={pid}, report-only mode)")
continue # Skip the rest of the check for Koby
# Gateway state file # Gateway state file
state = ssh(host, "cat ~/.hermes/gateway_state.json 2>/dev/null", user=user) state = ssh(host, "cat ~/.hermes/gateway_state.json 2>/dev/null", user=user)
+3 -8
View File
@@ -6,6 +6,8 @@ title: Zulip Mesh Health Monitor — Multi-Platform
version: 3.0.0 version: 3.0.0
runtime_contract: 2 runtime_contract: 2
agent: abiba agent: abiba
report_only_agents:
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
--- ---
# Zulip Mesh Health Monitor # Zulip Mesh Health Monitor
@@ -199,14 +201,7 @@ Check `platforms.zulip.state`: `connected` ✅ | `disconnected` ❌ | `error`
ssh root@<CT> "ps aux | grep 'gateway run' | grep -v grep" ssh root@<CT> "ps aux | grep 'gateway run' | grep -v grep"
``` ```
Gateway PID should exist with uptime > 60s. **Dual-gateway detection**: if more than one `gateway run` process is found, the gateway has a collision (typically one `--force` and one `--replace` process). Kill the newer/duplicate process, then restart the remaining gateway per-agent (parameterized 2026-08-09, captain ruling): Gateway PID should exist with uptime > 60s. **Dual-gateway detection**: if more than one `gateway run` process is found, the gateway has a collision (typically one `--force` and one `--replace` process). Kill the newer/duplicate process, then restart the remaining gateway via PM2 (`pm2 restart mumuni-zulip`). Check gateway log for "Gateway running with 2 platform(s)" (not 1) to confirm Zulip reloaded.
| Agent | Restart command | Notes |
|-------|-----------------|-------|
| Mumuni (.24) | `pm2 restart abiba-zulip` | Hermes gateway runs under PM2 as `abiba-zulip` |
| Tanko (.122) | `bash /opt/hermes-zulip-plugin/run.sh` (or the agent's systemd/user unit) | Tanko does NOT use PM2 — never run `pm2 restart mumuni-zulip` for Tanko (process does not exist) |
Check gateway log for "Gateway running with 2 platform(s)" (not 1) to confirm Zulip reloaded.
**B3: Heartbeat Verification** **B3: Heartbeat Verification**
+3 -2
View File
@@ -75,8 +75,9 @@ Track via `/tmp/zulip-heal-debounce-<agent>` (unix timestamp of last restart).
## Reporting ## Reporting
This contract is RETIRED — health-check logs are NOT knowledge graph content. Every cycle produces a knowledge graph node:
No graph nodes are created. Logs go to Gitea (SyslogSolution/health-logs). - Title: `[LEARN] zulip-self-heal: <timestamp>`
- metadata: { type: "remediation", status: "fixed" | "escalated" | "healthy" }
- Issues fixed → DM: "🛠 Zulip Self-Heal: fixed <issue>" - Issues fixed → DM: "🛠 Zulip Self-Heal: fixed <issue>"
- Issues escalated → DM: "⚠️ Zulip Self-Heal: <issue> needs attention" - Issues escalated → DM: "⚠️ Zulip Self-Heal: <issue> needs attention"