Compare commits

...
Author SHA1 Message Date
jerome 14d27a09b5 Merge pull request 'gpu-self-heal: refresh to current fleet baseline and topology' (#21) from feat/gpu-self-heal-refresh-20260718 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
Reviewed-on: #21
2026-07-18 08:36:37 +00:00
root bddbb22f03 gpu-self-heal: refresh to current fleet baseline and topology
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
- Synced model assignments to 2026-07-17 swaps (ThinkingCap, HauhauCS QAT, Genesis Hermes V3)
- Added stable role-based aliases from gpu-fleet (gpu-dense, gpu-light, strix-moe)
- Updated benchmark baselines to live values (74.9/169.6/62.9 tok/s)
- Replaced router (port 9000) references with LiteLLM + direct routing
- Replaced Prometheus exporter rule with sidecar health probe
- Updated VRAM thresholds to match operational data (300/300/200 MB/h)
- Added response size limit (1MB) to prevent OOM crashes
- Added Lessons L5 (response size crash) and L6 (stable aliases)
- Removed deprecated Rules 11-12 (router-specific distribution balance)
2026-07-18 08:07:03 +00:00
root 31ec70ae36 gpu-fleet: RTX 3090 swap to ThinkingCap-Qwen3.6-27B Q4_K_M
- Model: Qwopus Q4_K_M (16GB, 63 tok/s) → ThinkingCap Q4_K_M (15.7GB, 68 tok/s)
- RL-finetuned: 50% fewer thinking tokens, 0.85 MMLU-Pro (vs 0.83 base)
- Self-spec MTP (n=4) REQUIRED for stability — segfaults without it
- Added vision via mmproj (0.9GB) — new capability for this GPU
- VRAM: 20.9/24.6GB (85%), tighter but stable
- Outputs reasoning_content (hidden from Hermes agent)
2026-07-17 15:50:50 +00:00
root 5eb6d3bfbd gpu-fleet: RTX 5070 swap to HauhauCS Gemma4-12B QAT Uncensored Balanced
- Model: IQ4_NL (6.3GB, 191 tok/s) → Q4_K_M QAT (6.9GB, 87 tok/s)
- MTP draft: Q8_0 (444MB) → tuned draft (242MB), saves 200MB VRAM
- mmproj: F16 → BF16 (same size, matched to new model)
- Benefits: 0/465 refusals, agent-optimized tuning, QAT quality
- Trade: 54% slower generation (acceptable for gpu-light role)
- Config: --parallel 1, --ctx-size 131072, single-slot full 128K
- VRAM: 10.0/12.2GB (82%), healthy headroom
2026-07-17 15:24:13 +00:00
jerome 33cb88d571 Merge pull request 'feat: GPU context 256K→128K fleet-wide + Genesis Hermes V3 on Strix Halo' (#20) from feat/gpu-128k-genesis-hermes-v3-20260717 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
Reviewed-on: #20
2026-07-17 11:21:32 +00:00
root 4a1f476623 fix: add missing description to inference-optimization frontmatter
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
2026-07-17 11:11:27 +00:00
root ba2c55c7e6 ci: re-trigger pipeline for PR #20
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Failing after 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Has been skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Has been skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Has been skipped
2026-07-17 11:10:29 +00:00
root b9149bce47 feat: GPU context 256K→128K fleet-wide + Genesis Hermes V3 on Strix Halo
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Failing after 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Has been skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Has been skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Has been skipped
GPU Changes:
- All 3 GPUs reduced from 256K (-c 262144) to 128K (-c 131072) for stability
- Observed instability near 100K at 256K — 128K is the stable ceiling
- VRAM improved: RTX 3090 ~70% (was 90%), RTX 5070 ~65% (was 88%)
- Strix Halo swapped to LuffyTheFox/Genesis Hermes V3 APEX
  - Hermes agent fine-tune, tensor repair (3 SSM layers, 76% W1 improvement)
  - Uncensored (0/465 refusals), multimodal (mmproj F16)
  - Speed: 65 tok/s gen, 140 tok/s prompt
  - Alias strix-moe maintained

Agent Updates:
- Mumuni: max_context_window 262144→131072, already aligned on strix-moe/0.65
- Tanko: max_context_window 262144→131072
- Koonimo: max_context_window + context_length 262144→131072
- CT114 SSH access confirmed (was 'Zulip only')

LiteLLM (CT116):
- Updated backend model references qwen3.6-35B-udq4→strix-moe
- Removed stale ornith-1.0-35b from model_cost
- Fallback chains updated

Contracts Updated:
- gpu-fleet.prose.md: topology, VRAM, benchmarks, config lines, model assignments
- gpu-self-heal.prose.md: Rule 9/10 context targets
- hermes-config-template.prose.md: template values, Rules 7-9, compression thresholds
- inference-optimization.prose.md: added to repo, 128K recommendation

Compression: 0.65 fires at ~85K (~43K headroom before 128K ceiling)
For >128K workloads: route to external providers (deepseek)
2026-07-17 10:59:45 +00:00
root 65c99dab50 contract: update litellm-api-keys to v2026-07-17 — fleet standardization
- Document canonical systemd drop-in pattern (ExecStart= reset + wrapper)
- Hardcode venv paths — never use variables in single-quoted bash -c
- Update migration status: all 4 agents on while-true wrapper + st.8e848433
- Infisical CLI update procedure (0.38.0 → 0.43.109 via artifacts-cli)
- Service token inventory + .env fallback inventory
- Key rotation log: fleet standardize + tanko fix-zulip entries
- Remove git merge conflict artifacts
- Add fleet-wide standardization lessons section
2026-07-17 10:17:23 +00:00
jerome 20cbb96e2d Merge pull request 'Vault cleanup + contract sync (WAL #1316)' (#19) from fix/vault-cleanup-contract-sync into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
Reviewed-on: #19
2026-07-16 21:24:59 +00:00
root 8215b84f88 merge: resolve conflicts with master (delegation delete + gpu-fleet)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
2026-07-16 21:22:34 +00:00
root b7e23e2592 vault: Infisical cleanup + contract sync (WAL #1316)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
- litellm-api-keys: vault audit, tanko/koby/koonimo migration, key rotation log
- gpu-fleet: ornith-1.0-35b→strix-moe (remaining refs)
- hermes-agent-baseline: 256K all GPUs, CT IPs updated, Shumba retired
- delegation-prose-contract: removed (renamed to mumuni-delegation)
- contract-registry.yaml + cron-prompts-review.md: from feat/contract-registry
2026-07-16 21:17:09 +00:00
9 changed files with 2798 additions and 464 deletions
File diff suppressed because it is too large Load Diff
+612
View File
@@ -0,0 +1,612 @@
# Cron Prompts Review — All 10 Scheduled Contracts
Generated: 2026-07-13 20:59:18 ET
---
## hermes-key-enforcement
**Category:** compliance | **Domain:** hermes-agent | **Owner:** abiba | **Schedule:** 0 6 * * *
```
Contract Enforcement: hermes-key-enforcement
Category: compliance
Domain: hermes-agent
Owner: abiba
Schedule: Daily compliance scan at 6am ET
This is a compliance contract. Verify that the contract enforces the required standards and policies. Report any violations found.
Steps:
1. Load contract from prose-contracts/main (file: hermes-key-enforcement.prose.md)
2. Verify prerequisites (connectivity, tools, deps)
3. Execute contract per SOP
4. Run postconditions from contract registry
5. Generate receipt with status (pass/fail/escalated)
6. If any postcondition fails, escalate per contract escalation tiers
7. Log to ~/.hermes/runs/hermes-key-enforcement/
Postconditions to verify:
[
{
"check": "no plaintext API keys in config",
"verify": "grep -rc 'api_key: sk-' /root/.hermes/config.yaml",
"expect": "0 matches"
},
{
"check": "api_key_env used for harness/litellm providers",
"verify": "grep -c 'api_key_env.*LITELLM_API_KEY' /root/.hermes/config.yaml",
"expect": "count > 0"
}
]
Escalation Tiers:
- INFO: notify nobody → action: log_to_receipt
- WARNING: notify abiba → action: relay_alert
- CRITICAL: notify abiba, mumuni → action: relay_alert + pause
- FATAL: notify abiba, mumuni, kwame → action: relay_alert + pause + human_required
Circuit Breaker:
- Max retries: 3 per 60.0min window
- On trip: escalate_to_fatal
Receipt format: JSON with contract, run_id, timestamp, agent, status, actions_taken, postconditions, drift_alerts, evidence_path
Receipt storage: ~/.hermes/runs/hermes-key-enforcement/receipt-{timestamp}.json
Graph node: Create RA-H OS node for receipt with metadata: type=receipt, contract=hermes-key-enforcement, status=<status>
If the contract has no postconditions defined (e.g., reference/pattern contracts), log that it was loaded and skip execution.
IMPORTANT: If the contract file does not exist in prose-contracts/main, report failure and do NOT hallucinate forward.
```
---
## hermes-config-template
**Category:** compliance | **Domain:** hermes-agent | **Owner:** abiba | **Schedule:** 0 4 * * 1
```
Contract Enforcement: hermes-config-template
Category: compliance
Domain: hermes-agent
Owner: abiba
Schedule: Weekly config drift check Monday at 4am ET
This is a compliance contract. Verify that the contract enforces the required standards and policies. Report any violations found.
Steps:
1. Load contract from prose-contracts/main (file: hermes-config-template.prose.md)
2. Verify prerequisites (connectivity, tools, deps)
3. Execute contract per SOP
4. Run postconditions from contract registry
5. Generate receipt with status (pass/fail/escalated)
6. If any postcondition fails, escalate per contract escalation tiers
7. Log to ~/.hermes/runs/hermes-config-template/
Postconditions to verify:
[
{
"check": "agent config template_version matches template file",
"verify": "grep -q 'template_version' /root/.hermes/config.yaml && diff <(grep 'template_version' /root/.hermes/config.yaml | cut -d: -f2 | xargs) <(grep 'template_version' /root/prose-contracts/hermes-config-template.prose.md | cut -d: -f2 | xargs) && echo match || echo mismatch",
"expect": "match"
},
{
"check": "config file is valid YAML",
"verify": "python3 -c 'import yaml; yaml.safe_load(open(\"/root/.hermes/config.yaml\"))' && echo valid || echo invalid",
"expect": "valid"
}
]
Escalation Tiers:
- INFO: notify nobody → action: log_to_receipt
- WARNING: notify abiba → action: relay_alert
- CRITICAL: notify abiba, mumuni → action: relay_alert
- FATAL: notify abiba, mumuni, kwame → action: relay_alert + human_required
Circuit Breaker:
- Max retries: 3 per 60.0min window
- On trip: escalate_to_fatal
Receipt format: JSON with contract, run_id, timestamp, agent, status, actions_taken, postconditions, drift_alerts, evidence_path
Receipt storage: ~/.hermes/runs/hermes-config-template/receipt-{timestamp}.json
Graph node: Create RA-H OS node for receipt with metadata: type=receipt, contract=hermes-config-template, status=<status>
If the contract has no postconditions defined (e.g., reference/pattern contracts), log that it was loaded and skip execution.
IMPORTANT: If the contract file does not exist in prose-contracts/main, report failure and do NOT hallucinate forward.
```
---
## hermes-agent-baseline
**Category:** compliance | **Domain:** hermes-agent | **Owner:** abiba | **Schedule:** 0 5 * * 1
```
Contract Enforcement: hermes-agent-baseline
Category: compliance
Domain: hermes-agent
Owner: abiba
Schedule: Weekly baseline verification Monday at 5am ET
This is a compliance contract. Verify that the contract enforces the required standards and policies. Report any violations found.
Steps:
1. Load contract from prose-contracts/main (file: hermes-agent-baseline.prose.md)
2. Verify prerequisites (connectivity, tools, deps)
3. Execute contract per SOP
4. Run postconditions from contract registry
5. Generate receipt with status (pass/fail/escalated)
6. If any postcondition fails, escalate per contract escalation tiers
7. Log to ~/.hermes/runs/hermes-agent-baseline/
Postconditions to verify:
[
{
"check": "Hermes agent process running",
"verify": "pgrep -f 'hermes' > /dev/null && echo running || echo stopped",
"expect": "running"
},
{
"check": "agent config file exists and valid YAML",
"verify": "test -f /root/.hermes/config.yaml && python3 -c 'import yaml; yaml.safe_load(open(\"/root/.hermes/config.yaml\"))' && echo valid || echo invalid",
"expect": "valid"
},
{
"check": "no uncommitted changes in hermes directory",
"verify": "cd /root/.hermes && git status --porcelain | wc -l",
"expect": "0"
}
]
Escalation Tiers:
- INFO: notify nobody → action: log_to_receipt
- WARNING: notify abiba → action: relay_alert
- CRITICAL: notify abiba, mumuni → action: relay_alert
- FATAL: notify abiba, mumuni, kwame → action: relay_alert + human_required
Circuit Breaker:
- Max retries: 3 per 60.0min window
- On trip: escalate_to_fatal
Receipt format: JSON with contract, run_id, timestamp, agent, status, actions_taken, postconditions, drift_alerts, evidence_path
Receipt storage: ~/.hermes/runs/hermes-agent-baseline/receipt-{timestamp}.json
Graph node: Create RA-H OS node for receipt with metadata: type=receipt, contract=hermes-agent-baseline, status=<status>
If the contract has no postconditions defined (e.g., reference/pattern contracts), log that it was loaded and skip execution.
IMPORTANT: If the contract file does not exist in prose-contracts/main, report failure and do NOT hallucinate forward.
```
---
## proxmox-monitor
**Category:** monitoring | **Domain:** proxmox | **Owner:** abiba | **Schedule:** */15 * * * *
```
Contract Enforcement: proxmox-monitor
Category: monitoring
Domain: proxmox
Owner: abiba
Schedule: Every 15 minutes
This is a monitoring contract. Execute the monitoring checks defined in the contract. Report any deviations from expected state.
Steps:
1. Load contract from prose-contracts/main (file: proxmox-monitor.prose.md)
2. Verify prerequisites (connectivity, tools, deps)
3. Execute contract per SOP
4. Run postconditions from contract registry
5. Generate receipt with status (pass/fail/escalated)
6. If any postcondition fails, escalate per contract escalation tiers
7. Log to ~/.hermes/runs/proxmox-monitor/
Postconditions to verify:
[
{
"check": "all Proxmox nodes reachable",
"verify": "curl -sf http://192.168.68.10:8006/api2/json/status | jq '.status'",
"expect": "healthy"
},
{
"check": "no VMs in crashed state",
"verify": "pvesh get /nodes -output-format=json | jq '.[] | select(.status==\"Crashed\")'",
"expect": "empty"
},
{
"check": "backups running on schedule",
"verify": "pbs-info --check",
"expect": "last_backup < 24h ago"
}
]
Escalation Tiers:
- INFO: notify nobody → action: log_to_receipt
- WARNING: notify abiba, mumuni → action: relay_alert
- CRITICAL: notify abiba, mumuni → action: relay_alert + trigger_remediation
- FATAL: notify abiba, mumuni, kwame → action: relay_alert + pause + human_required
Circuit Breaker:
- Max retries: 3 per 60.0min window
- On trip: escalate_to_fatal
Receipt format: JSON with contract, run_id, timestamp, agent, status, actions_taken, postconditions, drift_alerts, evidence_path
Receipt storage: ~/.hermes/runs/proxmox-monitor/receipt-{timestamp}.json
Graph node: Create RA-H OS node for receipt with metadata: type=receipt, contract=proxmox-monitor, status=<status>
If the contract has no postconditions defined (e.g., reference/pattern contracts), log that it was loaded and skip execution.
IMPORTANT: If the contract file does not exist in prose-contracts/main, report failure and do NOT hallucinate forward.
```
---
## gpu-monitor
**Category:** monitoring | **Domain:** gpu | **Owner:** abiba | **Schedule:** */15 * * * *
```
Contract Enforcement: gpu-monitor
Category: monitoring
Domain: gpu
Owner: abiba
Schedule: Every 15 minutes — polls all GPU subsystems
This is a monitoring contract. Execute the monitoring checks defined in the contract. Report any deviations from expected state.
Steps:
1. Load contract from prose-contracts/main (file: gpu-monitor.prose.md)
2. Verify prerequisites (connectivity, tools, deps)
3. Execute contract per SOP
4. Run postconditions from contract registry
5. Generate receipt with status (pass/fail/escalated)
6. If any postcondition fails, escalate per contract escalation tiers
7. Log to ~/.hermes/runs/gpu-monitor/
Postconditions to verify:
[
{
"check": "GPU metrics accessible",
"verify": "curl -sf http://localhost:9100/gpu-data",
"expect": "200 OK, populated data"
},
{
"check": "dashboard serving",
"verify": "curl -sf http://localhost:9100/gpu-fleet.html",
"expect": "200 OK, HTML returned"
},
{
"check": "health endpoint responsive",
"verify": "curl -sf http://localhost:9100/health",
"expect": "200 OK"
}
]
Escalation Tiers:
- INFO: notify nobody → action: log_to_receipt
- WARNING: notify abiba, mumuni → action: relay_alert
- CRITICAL: notify abiba, mumuni → action: relay_alert + trigger_remediation
- FATAL: notify abiba, mumuni, kwame → action: relay_alert + pause + human_required
Circuit Breaker:
- Max retries: 3 per 60.0min window
- On trip: escalate_to_fatal
Receipt format: JSON with contract, run_id, timestamp, agent, status, actions_taken, postconditions, drift_alerts, evidence_path
Receipt storage: ~/.hermes/runs/gpu-monitor/receipt-{timestamp}.json
Graph node: Create RA-H OS node for receipt with metadata: type=receipt, contract=gpu-monitor, status=<status>
If the contract has no postconditions defined (e.g., reference/pattern contracts), log that it was loaded and skip execution.
IMPORTANT: If the contract file does not exist in prose-contracts/main, report failure and do NOT hallucinate forward.
```
---
## infrastructure-monitoring
**Category:** monitoring | **Domain:** infrastructure | **Owner:** abiba | **Schedule:** */30 * * * *
```
Contract Enforcement: infrastructure-monitoring
Category: monitoring
Domain: infrastructure
Owner: abiba
Schedule: Every 30 minutes
This is a monitoring contract. Execute the monitoring checks defined in the contract. Report any deviations from expected state.
Steps:
1. Load contract from prose-contracts/main (file: infrastructure-monitoring.prose.md)
2. Verify prerequisites (connectivity, tools, deps)
3. Execute contract per SOP
4. Run postconditions from contract registry
5. Generate receipt with status (pass/fail/escalated)
6. If any postcondition fails, escalate per contract escalation tiers
7. Log to ~/.hermes/runs/infrastructure-monitoring/
Postconditions to verify:
[
{
"check": "Proxmox API reachable",
"verify": "curl -sf http://192.168.68.10:8006/api2/json",
"expect": "200 OK"
},
{
"check": "Zulip API reachable",
"verify": "curl -sf https://chat.sysloggh.net/api/v1/me",
"expect": "200 OK"
},
{
"check": "LiteLLM proxy reachable",
"verify": "curl -sf http://192.168.68.116/litellm/v1/models",
"expect": "200 OK"
},
{
"check": "Gitea API reachable",
"verify": "curl -sf https://git.sysloggh.net/api/v1/version",
"expect": "200 OK"
},
{
"check": "SearXNG reachable",
"verify": "curl -sf http://192.168.68.17:8080",
"expect": "200 OK"
}
]
Escalation Tiers:
- INFO: notify nobody → action: log_to_receipt
- WARNING: notify abiba, mumuni → action: relay_alert
- CRITICAL: notify abiba, mumuni → action: relay_alert + trigger_remediation
- FATAL: notify abiba, mumuni, kwame → action: relay_alert + pause + human_required
Circuit Breaker:
- Max retries: 3 per 60.0min window
- On trip: escalate_to_fatal
Receipt format: JSON with contract, run_id, timestamp, agent, status, actions_taken, postconditions, drift_alerts, evidence_path
Receipt storage: ~/.hermes/runs/infrastructure-monitoring/receipt-{timestamp}.json
Graph node: Create RA-H OS node for receipt with metadata: type=receipt, contract=infrastructure-monitoring, status=<status>
If the contract has no postconditions defined (e.g., reference/pattern contracts), log that it was loaded and skip execution.
IMPORTANT: If the contract file does not exist in prose-contracts/main, report failure and do NOT hallucinate forward.
```
---
## zulip-health
**Category:** monitoring | **Domain:** zulip | **Owner:** abiba | **Schedule:** */15 * * * *
```
Contract Enforcement: zulip-health
Category: monitoring
Domain: zulip
Owner: abiba
Schedule: Every 15 minutes — monitors all Zulip-connected agents
This is a monitoring contract. Execute the monitoring checks defined in the contract. Report any deviations from expected state.
Steps:
1. Load contract from prose-contracts/main (file: zulip-health.prose.md)
2. Verify prerequisites (connectivity, tools, deps)
3. Execute contract per SOP
4. Run postconditions from contract registry
5. Generate receipt with status (pass/fail/escalated)
6. If any postcondition fails, escalate per contract escalation tiers
7. Log to ~/.hermes/runs/zulip-health/
Postconditions to verify:
[
{
"check": "bot registration active",
"verify": "curl -sf https://chat.sysloggh.net/api/v1/me | jq '.user_id'",
"expect": "bot_id present"
},
{
"check": "DM delivery working",
"verify": "curl -sf https://chat.sysloggh.net/api/v1/users/me/is-online",
"expect": "online: true"
}
]
Escalation Tiers:
- INFO: notify nobody → action: log_to_receipt
- WARNING: notify abiba, mumuni → action: relay_alert
- CRITICAL: notify abiba, mumuni → action: relay_alert + trigger_remediation
- FATAL: notify abiba, mumuni, kwame → action: relay_alert + pause + human_required
Circuit Breaker:
- Max retries: 3 per 60.0min window
- On trip: escalate_to_fatal
Receipt format: JSON with contract, run_id, timestamp, agent, status, actions_taken, postconditions, drift_alerts, evidence_path
Receipt storage: ~/.hermes/runs/zulip-health/receipt-{timestamp}.json
Graph node: Create RA-H OS node for receipt with metadata: type=receipt, contract=zulip-health, status=<status>
If the contract has no postconditions defined (e.g., reference/pattern contracts), log that it was loaded and skip execution.
IMPORTANT: If the contract file does not exist in prose-contracts/main, report failure and do NOT hallucinate forward.
```
---
## litellm-health
**Category:** monitoring | **Domain:** litellm | **Owner:** abiba | **Schedule:** */10 * * * *
```
Contract Enforcement: litellm-health
Category: monitoring
Domain: litellm
Owner: abiba
Schedule: Every 10 minutes — LiteLLM proxy health
This is a monitoring contract. Execute the monitoring checks defined in the contract. Report any deviations from expected state.
Steps:
1. Load contract from prose-contracts/main (file: litellm-health.prose.md)
2. Verify prerequisites (connectivity, tools, deps)
3. Execute contract per SOP
4. Run postconditions from contract registry
5. Generate receipt with status (pass/fail/escalated)
6. If any postcondition fails, escalate per contract escalation tiers
7. Log to ~/.hermes/runs/litellm-health/
Postconditions to verify:
[
{
"check": "LiteLLM proxy reachable",
"verify": "curl -sf http://192.168.68.116/litellm/v1/models",
"expect": "200 OK, models returned"
},
{
"check": "router deprecated, nginx routes work",
"verify": "curl -sf https://litellm.sysloggh.net/v1/models",
"expect": "200 OK (via nginx)"
}
]
Escalation Tiers:
- INFO: notify nobody → action: log_to_receipt
- WARNING: notify abiba, mumuni → action: relay_alert
- CRITICAL: notify abiba, mumuni → action: relay_alert + trigger_remediation
- FATAL: notify abiba, mumuni, kwame → action: relay_alert + pause + human_required
Circuit Breaker:
- Max retries: 3 per 60.0min window
- On trip: escalate_to_fatal
Receipt format: JSON with contract, run_id, timestamp, agent, status, actions_taken, postconditions, drift_alerts, evidence_path
Receipt storage: ~/.hermes/runs/litellm-health/receipt-{timestamp}.json
Graph node: Create RA-H OS node for receipt with metadata: type=receipt, contract=litellm-health, status=<status>
If the contract has no postconditions defined (e.g., reference/pattern contracts), log that it was loaded and skip execution.
IMPORTANT: If the contract file does not exist in prose-contracts/main, report failure and do NOT hallucinate forward.
```
---
## memory-audit-maintenance
**Category:** maintenance | **Domain:** memory | **Owner:** mumuni | **Schedule:** 0 3 * * *
```
Contract Enforcement: memory-audit-maintenance
Category: maintenance
Domain: memory
Owner: mumuni
Schedule: Daily at 3am ET
This is a maintenance contract. Execute the maintenance tasks defined in the contract. Report any issues found.
Steps:
1. Load contract from prose-contracts/main (file: memory-audit-maintenance.prose.md)
2. Verify prerequisites (connectivity, tools, deps)
3. Execute contract per SOP
4. Run postconditions from contract registry
5. Generate receipt with status (pass/fail/escalated)
6. If any postcondition fails, escalate per contract escalation tiers
7. Log to ~/.hermes/runs/memory-audit-maintenance/
Postconditions to verify:
[
{
"check": "memory files below 80% capacity",
"verify": "wc -l ~/.hermes/memories/*.md",
"expect": "total lines < threshold"
},
{
"check": "no stale entries",
"verify": "grep -r 'STALE' ~/.hermes/memories/",
"expect": "0 matches"
}
]
Escalation Tiers:
- INFO: notify nobody → action: log_to_receipt
- WARNING: notify mumuni → action: relay_alert
- CRITICAL: notify mumuni, abiba → action: relay_alert
- FATAL: notify mumuni, abiba, kwame → action: relay_alert + human_required
Circuit Breaker:
- Max retries: 3 per 60.0min window
- On trip: escalate_to_fatal
Receipt format: JSON with contract, run_id, timestamp, agent, status, actions_taken, postconditions, drift_alerts, evidence_path
Receipt storage: ~/.hermes/runs/memory-audit-maintenance/receipt-{timestamp}.json
Graph node: Create RA-H OS node for receipt with metadata: type=receipt, contract=memory-audit-maintenance, status=<status>
If the contract has no postconditions defined (e.g., reference/pattern contracts), log that it was loaded and skip execution.
IMPORTANT: If the contract file does not exist in prose-contracts/main, report failure and do NOT hallucinate forward.
```
---
## infrastructure-update
**Category:** maintenance | **Domain:** infrastructure | **Owner:** abiba | **Schedule:** 0 2 * * 0
```
Contract Enforcement: infrastructure-update
Category: maintenance
Domain: infrastructure
Owner: abiba
Schedule: Weekly system updates Sunday at 2am ET
This is a maintenance contract. Execute the maintenance tasks defined in the contract. Report any issues found.
Steps:
1. Load contract from prose-contracts/main (file: infrastructure-update.prose.md)
2. Verify prerequisites (connectivity, tools, deps)
3. Execute contract per SOP
4. Run postconditions from contract registry
5. Generate receipt with status (pass/fail/escalated)
6. If any postcondition fails, escalate per contract escalation tiers
7. Log to ~/.hermes/runs/infrastructure-update/
Postconditions to verify:
[
{
"check": "all services running after update",
"verify": "systemctl list-units --state=running",
"expect": "all critical services"
}
]
Escalation Tiers:
- INFO: notify nobody → action: log_to_receipt
- WARNING: notify abiba → action: relay_alert
- CRITICAL: notify abiba, mumuni → action: relay_alert
- FATAL: notify abiba, mumuni, kwame → action: relay_alert + human_required
Circuit Breaker:
- Max retries: 1 per 120.0min window
- On trip: escalate_to_fatal
Receipt format: JSON with contract, run_id, timestamp, agent, status, actions_taken, postconditions, drift_alerts, evidence_path
Receipt storage: ~/.hermes/runs/infrastructure-update/receipt-{timestamp}.json
Graph node: Create RA-H OS node for receipt with metadata: type=receipt, contract=infrastructure-update, status=<status>
If the contract has no postconditions defined (e.g., reference/pattern contracts), log that it was loaded and skip execution.
IMPORTANT: If the contract file does not exist in prose-contracts/main, report failure and do NOT hallucinate forward.
```
---
*End of review*
-256
View File
@@ -1,256 +0,0 @@
---
kind: pattern
name: delegation-prose-contract
description: >
Manager (Mumuni) operating doctrine for task decomposition, worker
delegation, verification, and delivery. Defines when to delegate, which
worker to use for what, how to handle failures, and the kanban board
protocol. Enforces context-window discipline and separation of concerns.
Runs on Mumuni (CT 118, storepve, .6) via Hermes agent.
version: 1.0.0
---
## Maintains
- Worker roster: 6 profiles (`syslog-code`, `syslog-devops`, `syslog-email`,
`syslog-research`, `syslog-review`, `syslog-writer`)
- Kanban board state at `~/.hermes/kanban/kanban.json`
- Context window budget: ~65K tokens per request (131K total, 60% threshold)
## Topology
**Cluster:** 5 Proxmox nodes (ocupve, acerpve, minipve, amdpve, storepve)
**Manager:** Mumuni (CT 118, storepve, .6) via Hermes agent
**Workers:** 6 profiles, all running on the same agent — no separate hosts needed
This contract is infrastructure-agnostic in terms of which nodes are used.
Workers execute tasks on whatever infrastructure they're given — SSH to .6,
.pm, .9, .12, or .15 depending on the task. The contract defines the
**who** and **when** — not the **where**.
## Why This Matters
Without enforced delegation, the manager consumes the full iteration budget
(60 calls) on single-turn tasks — SSH to 5 nodes, check each VM, read logs —
leaving no capacity for actual coordination. The result: context overflow
(59K tokens in system prompt), iteration exhaustion, and degraded response
quality. This contract exists because I blew through my budget checking
Proxmox node status instead of delegating to `syslog-devops`.
## Context Window Discipline
**The system prompt is ~6.5K tokens (stable: ~4.5K tool schemas + ~2K other guidance).**
**Volatile (MEMORY.md + USER.md): ~300 tokens.**
**Total base: ~6,800 tokens per request.**
The remaining budget is the conversation. Every tool call result adds to it.
If a single call returns >10K tokens (e.g., `grep` on a large file, SSH output
from multiple nodes), the context fills fast. That's why we delegate: workers
process in isolation and return compact results.
## Trigger Conditions
Delegation is **mandatory** when any of these apply:
| Condition | Threshold | Example |
|-----------|-----------|---------|
| Multiple tool calls needed | 2+ calls with intermediate logic | Read file → analyze → write report |
| Large data retrieval | Output >5K tokens | `grep -r "pattern" /path` on large dirs |
| Cross-domain work | Spans 2+ worker specialties | Infra check + email filter |
| Infrastructure changes | Any mutating operation | `qm set`, `systemctl restart`, `git push` |
| Research/analysis | Needs browser or deep reading | Web research, code review, data analysis |
| Code builds or changes | Writing or modifying code | Scripts, configs, patches |
| Sequential dependencies | Worker B needs Worker A's output | Code → Review → Deliver |
**Single tool calls stay at manager level.** Quick `grep`, `ls`, `cat`,
`curl`, `hermes tools list` — these are decision-making tools. The manager
reads them directly.
## Worker Selection Matrix
| Worker | Model | Toolsets | Role | Use When |
|--------|-------|----------|------|----------|
| `syslog-code` | qwen3.6-27B-code | terminal, file, web, memory, skills | Code patches, automation, scripts | Writing/modifying code, creating scripts, debugging, reading/writing files |
| `syslog-devops` | qwen3.6-27B-code | terminal, file, web, memory, skills | Infrastructure, DB, bridge, Proxmox | Server ops, SSH, Docker, Proxmox, DB queries, hardware checks |
| `syslog-email` | strix-moe | terminal, file, web, memory, skills | Email automation, mail operations | Sending/receiving email, inbox management, SMTP operations |
| `syslog-research` | strix-moe | terminal, file, web, memory, skills, **browser** | Analysis, classification, data processing | Web research, browser tasks, data analysis, classification, reading docs |
| `syslog-review` | strix-moe | terminal, file, web, memory, skills | Verification, QA, audit validation | **ALWAYS** verify worker output before delivery — especially for infra changes, code builds, and research findings |
| `syslog-writer` | strix-moe | terminal, file, web, memory, skills | Docs, content, branding, reports | Writing docs, reports, proposals, content, markdown formatting |
### Selection Rules
1. **Match specialty first.** A code task → `syslog-code`. An infra task →
`syslog-devops`. Don't put a `syslog-email` worker on a code review.
2. **Research tasks with browser needs → `syslog-research`.** Other workers
don't have the browser toolset.
3. **Verification → `syslog-review`.** Never deliver raw worker output.
4. **Documentation/content → `syslog-writer`.** Let them own the prose.
5. **If unsure, delegate to `syslog-research`** — it has the broadest toolset
(includes browser) and high reasoning effort.
## Delegation Protocol
### Step 1: Decompose
Break the task into lanes. Each lane does ONE thing. Workers are independent —
no lane depends on another's output mid-flight. If lanes depend on each other,
dispatch sequentially.
### Step 2: Dispatch
Fire workers via `delegate_task`:
**Parallel (independent lanes):**
```
delegate_task(
tasks=[
{"goal": "Check all 5 Proxmox nodes for VM status", "context": "SSH to each node via 192.168.68.x, run 'qm list'"},
{"goal": "Check Docker container health on .7/.116/.17", "context": "SSH to each host, check container status"},
]
)
```
**Sequential (dependent lanes):**
Dispatch lane 1 → wait for result → dispatch lane 2.
### Step 3: Verify
**MANDATORY for:**
- Infrastructure changes (any `qm`, `pct`, `systemctl`, `git push`)
- Code builds and modifications
- Research findings (web data, external sources)
- Any output that will reach the user
**Fire `syslog-review` to verify:**
```
delegate_task(
goal="Review the output of the devops worker. Verify the node status
report is accurate, check for inconsistencies, confirm all nodes were
reachable.",
context="Worker was syslog-devops. Output is at /tmp/node-report.md.
Verify against live system."
)
```
**If verification fails:**
1. Send work back to original worker with review feedback
2. Re-verify
3. Max 2 re-verify cycles before escalating to Kwame
### Step 4: Deliver
Only verified results reach Kwame. Format per channel:
- Telegram: Use `telegram-formatting` skill
- Zulip: Use Zulip Markdown (CommonMark)
- Email: Use `syslog-email` skill
## Kanban Board Protocol
**File:** `~/.hermes/kanban/kanban.json`
```json
{
"task_id": "unique-id",
"title": "Task description",
"created": "2026-07-09T01:00:00",
"status": "backlog|in_progress|review|done",
"lanes": [
{
"lane_id": "devops-check",
"worker": "syslog-devops",
"goal": "Check all 5 Proxmox nodes",
"status": "dispatched|completed|failed",
"output_file": "/tmp/node-report.md"
}
]
}
```
**Update the board on every state change.**
## Failure Handling
### Worker Timeouts
- Child timeout: **900 seconds** (15 minutes)
- Worker model `syslog-auto` is slow — it can hit the timeout limit with
22+ API calls
- **If a worker times out:** Re-dispatch with a narrower scope. Break the
task into smaller pieces that fit in the timeout window.
- **Avoid delegating sequential SSH hops** — each SSH connection adds latency
that compounds quickly. Prefer API-based or local approaches when possible.
### Worker Selection Failures
- `syslog-devops` is best for infrastructure tasks (SSH, Proxmox, Docker)
- `syslog-code` is best for code-level work (reading files, writing scripts)
- `syslog-research` has the browser toolset — use for web research
- `syslog-review` is the QA gate — always fire before delivery
- **Never fire more than 3 parallel workers** (max_concurrent_children: 3)
- **Never nest delegation** (max_spawn_depth: 1)
### Context Overflow
- If a task requires >10K tokens of output, delegate the processing
- Workers return compact summaries, not raw data dumps
- Pass file paths and concrete goals — never dump raw data into context
## Anti-patterns
- ❌ Reading large files into your own context before deciding → delegate the read
- ❌ Carrying SSH/grep/output results in your context → delegate the analysis
- ❌ Doing work yourself and then "pretending" to delegate → the user can tell
- ❌ Skipping verification → raw worker output never reaches the user
- ❌ Delegating single tool calls → keep quick reads/writes at manager level
- ❌ Firing more than 3 workers in parallel → hard limit
## Emergency Exception
**In an emergency (server down, service must be restored immediately):**
- Delegate the diagnosis (find the problem)
- Execute the fix yourself (minimize handoff latency)
- Verify the fix after delivery
- Log the exception in the kanban board
The emergency exception exists because the user needs the service back NOW,
not after three worker round-trips. But it's an exception — not the rule.
## What This Contract Doesn't Cover
1. **Worker profile configuration** — covered by `hermes-config-template.prose.md`
2. **SSH key management** — covered by existing SSH/Proxmox contracts
3. **Git workflow** — covered by `AGENTS.md` in the prose-contracts repo
4. **Cron job management** — covered by individual cron contracts
5. **Infra verification** — covered by `verify-before-mutate` protocol
## Verification
Run `scripts/worker-audit.py` to verify all 6 profiles are aligned:
```bash
python3 /root/.hermes/skills/kanban-orchestrator/scripts/worker-audit.py
```
## References
- `kanban-orchestrator` skill: The operational playbook (detailed execution steps)
- `worker-profile-audit.md` (skill reference): Worker configuration audit notes
- `delegation-timeout-patterns.md` (skill reference): Timeout handling patterns
- `verify-before-mutate` protocol: Infrastructure change verification
- `hermes-config-template.prose.md`: Worker profile configuration
## Success Criteria
This contract succeeds when:
1. **No context overflow** — single-turn tasks don't exhaust the iteration budget
2. **Workers do the work** — manager coordinates, doesn't execute
3. **Verification before delivery** — all output passes through `syslog-review`
4. **Kanban board is current** — every task has a lane, every lane has a status
5. **User gets verified results** — raw worker output never reaches Kwame
---
**Last updated:** 2026-07-09
**Author:** Mumuni (with Kwame's input on triggers and exception criteria)
**Status:** Draft — awaiting PR review and merge to prose-contracts main
+48 -32
View File
@@ -7,10 +7,14 @@ description: >
saturation watchdog, Prometheus/Grafana monitoring, and self-healing.
UPDATED 2026-07-15: Stable role-based aliases introduced: strix-moe,
gpu-dense, gpu-light. These never change — only the underlying model does.
Strix Halo: ornith-1.0-35b → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB).
Strix Halo: strix-moe → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB).
RTX 5070: gemma-4-12b Q4_K_M → IQ4_NL + MTP draft (122 tok/s, 2x faster).
RTX 5070 context: 131K → 256K. VRAM: 88% (10.8/12.2GB).
Compression timeout: 300s (was 120s). Mumuni context: 128K (was 256K).
UPDATED 2026-07-17: Strix Halo model swapped to Genesis Hermes V3 APEX (LuffyTheFox, 24GB, uncensored,
Hermes agent fine-tune, tensor repair, multimodal with mmproj).
RTX 5070 swapped to HauhauCS Gemma4-12B QAT Uncensored Balanced (Q4_K_M, 87 tok/s, 0/465 refusals).
Instability observed near 100K at 256K. 128K is the stable ceiling.
For larger context needs → fall back to external providers (deepseek).
VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%.
agent: abiba
triggers:
- on model add/remove
@@ -67,7 +71,7 @@ triggers:
│ CT 8 │ │ CT 110 │ │ CT 15 │ │ pi (.24) │
│ RTX 3090 │ │ RTX 5070 │ │ Strix Halo│ │ GPU Monitor │
│ 24GB │ │ 12GB │ │ 64GB UMA │ │ :9100 │
256K ctx │ │ 256K ctx │ │ 256K ctx │ │ Watchdog │
128K ctx │ │ 128K ctx │ │ 128K ctx │ │ Watchdog │
│ qwen3.6 │ │ gemma-4-12b │ │ qwen3.6 │ │ Prometheus │
│ 27B-code │ │ :8080 │ │ -35B-udq4 │ │ exporter │
│ :8080 │ │ :9400 (exp) │ │ :9400(exp)│ │ :9401 │
@@ -83,19 +87,31 @@ When a model is swapped on a GPU, ONLY the infrastructure layer changes — agen
| Alias | GPU | Current Model | Will Route To |
|-------|-----|---------------|---------------|
| `strix-moe` | Strix Halo (.15) | qwen3.6-35B-udq4 | Whatever runs on Strix Halo |
| `gpu-dense` | RTX 3090 (.8) | qwen3.6-27B-code | Whatever runs on RTX 3090 |
| `gpu-dense` | RTX 3090 (.8) | qwen3.6-27B-code (ThinkingCap) | Whatever runs on RTX 3090 |
| `gpu-light` | RTX 5070 (.110) | gemma-4-12b | Whatever runs on RTX 5070 |
**Backward compatibility**: Old model-specific names (qwen3.6-27B-code, gemma-4-12b, qwen3.6-35B-udq4) still work
but are deprecated for agent configs. Only the stable aliases survive model swaps.
## Current Model Assignments (2026-07-15)
## Current Model Assignments (2026-07-17)
| Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status |
|-------|-----|------|------|-----|----------|----------|-------------|--------|
| qwen3.6-27B-code (MTP) | RTX 3090 | .8 (llm-gpu) | 22.2/24.6GB (90%) | **256K** 🚀 | turbo4 | 2 | default | ✅ 63 tok/s |
| gemma-4-12b | RTX 5070 | .110 (ocu-llm) | 10.0/12.2GB (82%) | 256K | q4_0 | 2 | 2048/1024 | ✅ healthy |
| qwen3.6-35B-udq4 | Strix Halo Vulkan | .15 (amdpve) | ~9GB/64GB | 256K | q8_0 | 2 | 2048/512 | ✅ healthy |
| qwen3.6-27B-code (ThinkingCap) | RTX 3090 | .8 (llm-gpu) | ~20.9/24.6GB (85%) | **128K** | turbo4 | 1 | default | ✅ 68 tok/s |
| gemma-4-12b (HauhauCS QAT) | RTX 5070 | .110 (ocu-llm) | ~10.0/12.2GB (82%) | 128K | q4_0 | 1 | 2048/1024 | ✅ 87 tok/s |
| Genesis Hermes V3 APEX | Strix Halo Vulkan | .15 (amdpve) | ~10GB/64GB | 128K | q4_0 | 1 | 4096/1024 | ✅ 65 tok/s |
> **RTX 5070 model swap (2026-07-17)**: Switched from `gemma-4-12b-it-IQ4_NL` (Unsloth, 191 tok/s)
> to `HauhauCS/Gemma4-12B-QAT-Uncensored-HauhauCS-Balanced` (Q4_K_M QAT, 87 tok/s).
> Trade: 54% slower generation for QAT quality, 0/465 refusals, and agent-optimized tuning.
> MTP draft also swapped: Q8_0 (444MB) → tuned draft (242MB), saving 200MB VRAM.
> Role unchanged: gpu-light (vision, web extract, light auxiliary tasks).
> **RTX 3090 model swap (2026-07-17)**: Switched from `Qwopus3.6-27B-v2-MTP-Q4_K_M` (63 tok/s)
> to `bottlecapai/ThinkingCap-Qwen3.6-27B` (Q4_K_M QAT, 68 tok/s).
> RL-finetuned: 50% fewer thinking tokens, MMLU-Pro 0.85 vs 0.83 base.
> Self-spec MTP REQUIRED (crashes without it on turboquant build).
> Added vision via mmproj (0.9GB). VRAM 85%.
## Routing Configuration (LiteLLM — July 2026)
@@ -104,7 +120,7 @@ but are deprecated for agent configs. Only the stable aliases survive model swap
| Model | GPU | Weight | RPM Cap | Timeout |
|-------|-----|--------|---------|---------|
| qwen3.6-27B-code | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** |
| qwen3.6-35B-udq4 | Strix Halo (.15:8080) | **0.30** | 60 | **300s** |
| Genesis Hermes V3 APEX | Strix Halo (.15:8080) | **0.30** | 60 | **300s** |
| gemma-4-12b | RTX 5070 (.110:8080) | **0.15** | 200 | **120s** |
Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) is NOT in the inference path.
@@ -113,26 +129,26 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`.
| Model | RPM Cap | Notes |
|-------|---------|-------|
| qwen3.6-35B-udq4 | 40 | Tight cap — prevents Strix overload |
| qwen3.6-27B-code | 500 | High cap — primary workhorse |
| gemma-4-12b | 500 | High cap — IQ4_NL+MTP, 122 tok/s |
| strix-moe (Hermes V3) | 40 | Tight cap — prevents Strix overload |
| qwen3.6-27B-code | 500 | ThinkingCap Q4_K_M + MTP self-spec + vision, 68 tok/s |
| gemma-4-12b | 500 | HauhauCS QAT Uncensored Balanced + MTP, 87 tok/s |
### Stable Aliases (for agent configs — never change)
| Alias | RPM Cap | Routes To | Purpose |
|-------|---------|-----------|---------|
| `strix-moe` | 40 | Strix Halo | Compression tasks (MoE models) |
| `gpu-dense` | 500 | RTX 3090 | Heavy reasoning |
| `gpu-light` | 500 | RTX 5070 | Vision, web extract, light tasks |
| `gpu-dense` | 500 | RTX 3090 (ThinkingCap) | Heavy reasoning, code gen, delegation |
| `gpu-light` | 500 | RTX 5070 (HauhauCS QAT) | Vision, web extract, light tasks |
### Fallback Chains
- gemma → qwen
- qwen → gemma
- qwen3.6-35B-udq4 → qwen → gemma
- strix-moe → qwen → gemma
- syslog-auto → qwen → gemma → qwen3.6-35B-udq4
### Why Strix Halo RPM Is Capped
- Direct (qwen3.6-35B-udq4): 40 RPM (tight) — Strix Halo is shared with compression tasks
- Direct (strix-moe): 40 RPM (tight) — Strix Halo is shared with compression tasks
- Via syslog-auto: 60 RPM (moderate) — prevents flooding when multiple agents use syslog-auto simultaneously
- Combined max: ~100 RPM across both paths — Strix Halo can sustain this at 80°C
@@ -186,7 +202,7 @@ Show full fleet status: GPUs, models, VRAM, context windows, parallel slots, act
3. Check LiteLLM: `curl http://192.168.68.116/health` (expect "I'm alive!")
4. Check LiteLLM models: `curl -H "Authorization: Bearer $MASTER_KEY" http://192.168.68.116/v1/models`
5. Check LiteLLM timeouts: `grep -n 'timeout:' /opt/inference-harness/litellm_config.yaml`
- gemma-4-12b: 120s, qwen3.6-27B-code: 300s, qwen3.6-35B-udq4/strix-moe: 300s (ornith-1.0-35b does NOT exist — legacy name, do not use)
- gemma-4-12b: 120s, qwen3.6-27B-code: 300s, qwen3.6-35B-udq4/strix-moe: 300s (strix-moe does NOT exist — legacy name, do not use)
- global request_timeout: 300s, nginx proxy_read_timeout: 600s
6. Check AMD metrics: `curl http://192.168.68.15:9400/metrics` (Radeon 8060S, util%, VRAM, temp, power)
7. Check port conflicts: verify only one llama-server on :8080 per host
@@ -243,12 +259,12 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
- **Router startup race**: Compose router.py doesn't call load_roster(). Reload thread sleeps 30s first.
Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster().
- **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround.
- **VRAM (2026-07-15)**: RTX 3090 at 22.2/24.6GB (90%) with **256K context** (corrected from 131K). RTX 5070 at 10.8/12.2GB (88%) with 256K context + MTP. Strix Halo at ~9GB/64GB.
- **VRAM (2026-07-15)**: RTX 3090 at ~17/24.6GB (~70%) with **128K context** (reduced from 256K 2026-07-17). RTX 5070 at ~7.8/12.2GB (~65%) with 128K context + MTP. Strix Halo at ~7GB/64GB.
- **RTX 3090 runs `--parallel 2`** with MTP draft (spec-type draft-mtp, spec-draft-n-max 2).
- **RTX 3090 config**: `-c 262144 -ctk turbo4 -ctv turbo4 --parallel 2 --flash-attn on --cont-batching --spec-type draft-mtp`. Context corrected to 256K (2026-07-15). VRAM: 90%. Service: `/home/llmuser/llama-wrapper.sh`.
- **RTX 5070 config (2026-07-15)**: Switched to IQ4_NL + MTP draft (Q8_0) at 256K context. Gen speed: 122 tok/s (was 70). VRAM: 10.8/12.2GB (88%). No draft model pre-upgrade due to VRAM constraints. Service: `/home/llmuser/llama-wrapper.sh`. Config: `--model gemma-4-12b-it-IQ4_NL.gguf --spec-draft-model gemma-4-12b-it-Q8_0-MTP.gguf --spec-type draft-mtp --spec-draft-n-max 4 --ctx-size 262144`.
- **RTX 3090 config**: `-c 131072 -ctk turbo4 -ctv turbo4 --parallel 1 --flash-attn on --cont-batching --spec-type draft-mtp --spec-draft-n-max 4`. ThinkingCap Qwen3.6-27B Q4_K_M (15.7GB) + mmproj (0.9GB) + MTP self-spec. VRAM: ~85%. Service: `/home/llmuser/llama-wrapper.sh`. ⚠️ MTP REQUIRED for stability on this turboquant build — model segfaults without `--spec-type draft-mtp`. Outputs reasoning_content (hidden from agent, improves answer quality).
- **RTX 5070 config (2026-07-17)**: HauhauCS Gemma4-12B QAT Uncensored Balanced (Q4_K_M) + tuned MTP draft (242MB) at 128K context, single slot. Gen speed: 87 tok/s (vs 191 IQ4_NL). VRAM: ~10.0/12.2GB (~82%). Service: `/home/llmuser/llama-wrapper.sh`. Model: `Gemma4-12B-QAT-Uncensored-HauhauCS-Balanced-Q4_K_M.gguf`, MTP: `mtp-gemma-4-12B-it.gguf`, mmproj: `mmproj-Gemma4-12B-QAT-Uncensored-HauhauCS-Balanced-BF16.gguf`. Recommended sampling: temp 0.6, top_k 64, top_p 0.9, min_p 0.05, repeat_penalty 1.1.
- **LiteLLM timeout tuning (verified 2026-07-16 against `/opt/inference-harness/litellm_config.yaml` on CT 116)**: gemma-4-12b 120s, qwen3.6-27B-code 300s, qwen3.6-35B-udq4 300s, strix-moe 300s, syslog-auto routes all 300s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s.
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080 (was `ornith-server.service`), model changed to `unsloth/Qwen3.6-35B-A3B-MTP-GGUF` (UD-Q4_K_M), alias `qwen3.6-35B-udq4`, 256K context, flash-attn + q8 KV. MTP support enabled for 1.4-2.2x faster inference.
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V3-GGUF` (APEX quant), alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded). Hermes agent fine-tune, tensor repair (SSM layers fixed via SVD), uncensored (0/465 refusals).
- **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill).
- **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (Mumuni) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`.
- **Port 8080 firewall**: amdpve iptables restricts 8080 to 192.168.68.116 (LiteLLM/router host) only. All inbound connections are from .116 (LiteLLM proxied via nginx). Localhost curls hang (SYN dropped). Always test from .116.
@@ -262,12 +278,12 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
| GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Context |
|-----|-------|-----------|--------------|----------|---------|
| RTX 3090 (.8) | qwen3.6-27B-code (MTP) | **63** | — | — | **256K** |
| RTX 5070 (.110) | gemma-4-12b (IQ4_NL+MTP) | **191** | — | — | **256K** |
| Strix Halo (.15) | qwen3.6-35B-udq4 | **71** | | — | **256K** |
| RTX 3090 (.8) | ThinkingCap-Qwen3.6-27B Q4_K_M | **68** | — | — | **128K** |
| RTX 5070 (.110) | HauhauCS QAT Uncensored Balanced | **87** | — | — | **128K** |
| Strix Halo (.15) | Genesis Hermes V3 APEX | **65** | 140 | — | **128K** |
Benchmarks from 2026-07-15 verification run. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
All 3 GPUs now at 256K context (2026-07-15).
Benchmarks from 2026-07-17. Strix Halo swapped to Genesis Hermes V3 APEX (LuffyTheFox). RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
All 3 GPUs now at 128K context (2026-07-17, reduced from 256K for stability).
Benchmarks run through LiteLLM proxy (192.168.68.116:4001) every 5 minutes.
Degradation alerts fire at 30% (warning) and 50% (critical) below baseline.
@@ -288,9 +304,9 @@ All agent configs MUST use stable role-based aliases, never model-specific names
When the underlying model is swapped, only the LiteLLM config changes — agent configs are untouched.
### Context Windows
- RTX 3090: **256K** (was 131K, bumped 2026-07-15) | RTX 5070: **256K** (up from 131K) | Strix Halo: **256K**
- **Mumuni compression context**: 128K (down from 256K) — ensures compression model doesn't timeout
- Compression threshold 0.65: fires at ~85K for 128K context window
- RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **128K**
- **All agents**: 128K ceiling — stable margin. For >128K workloads, use external providers (deepseek)
- Compression threshold 0.65: fires at ~85K (~43K headroom before 128K ceiling)
- Mumuni compression model alias: `strix-moe` with 300s timeout
### Mumuni Agent Profile
@@ -306,7 +322,7 @@ Mumuni (CT114, 192.168.68.123) is the primary business assistant. This profile i
| `aux.vision.model` | `gpu-light` | Vision tasks (RTX 5070) |
| `aux.web_extract.model` | `gpu-light` | Web extraction |
| `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) |
| `context.max_context_window` | 262144 (256K) | Fixed 2026-07-16 (was 131072 — caused premature compression, WAL #1300) |
| `context.max_context_window` | 131072 (128K) | Reduced from 256K 2026-07-17 — stable 128K ceiling |
| `compression.threshold` | 0.65 | Triggers at ~85K |
| `compression.target_ratio` | 0.3 | Compresses to ~38K |
| `compression.protect_last_n` | 40 | Preserves last 40 messages |
@@ -314,7 +330,7 @@ Mumuni (CT114, 192.168.68.123) is the primary business assistant. This profile i
| `personalities` | `creative` | Creative assistant personality |
| Platforms | cli, discord, homeassistant, signal, telegram, zulip | All Hermes platforms |
| Main model timeout | 300s | LiteLLM global timeout |
| Compression model timeout | 300s | ornith timeout increased from 120s |
| Compression model timeout | 300s | strix-moe timeout increased from 120s |
### Agent Update Status (2026-07-15)
+93 -69
View File
@@ -6,10 +6,15 @@ description: >
benchmarks, and predicts failures before they happen. Extends gpu-monitor
(v2.1.0) with active remediation rules, Prometheus metrics consumption,
VRAM trend analysis, and predictive alerting.
UPDATED 2026-07-18: Model assignments synced to 2026-07-17 swaps.
Router (port 9000) references replaced with direct GPU routing.
Benchmark baselines refreshed to live values.
Prometheus exporters removed — not deployed; fall back to direct sidecar probes.
Stable role-based aliases (strix-moe, gpu-dense, gpu-light) from gpu-fleet.
agent: abiba
depends_on:
- gpu-monitor.prose.md (live data source on .24:9100)
- gpu-fleet.prose.md (source of truth for topology)
- gpu-fleet.prose.md (source of truth for topology, aliases, model assignments)
---
## Maintains
@@ -22,8 +27,8 @@ depends_on:
## Requires
- gpu-monitor:function — Live fleet data from .24:9100/gpu-data
- Prometheus exporters on all 3 GPUs (:9400/metrics)
- gpu-monitor:function — Live fleet data from localhost:9100/gpu-data
- Direct sidecar probe access to all GPU hosts (:8080/health)
- SSH access to GPU hosts for restart operations
## Continuity
@@ -35,21 +40,37 @@ depends_on:
---
## Current Fleet Baseline (2026-07-18)
| Alias | GPU | Host | Model | VRAM | Ctx | tok/s | Role |
|-------|-----|------|-------|------|-----|-------|------|
| `gpu-dense` | RTX 3090 24GB | ct8 (.8:8080) | ThinkingCap Qwen3.6-27B Q4_K_M + MTP + vision | 21.6/24.6GB (88%) | 128K | 74.9 | Heavy reasoning, code gen |
| `gpu-light` | RTX 5070 12GB | ct110 (.110:8080) | HauhauCS Gemma4-12B QAT Q4_K_M + MTP draft | 10.1/12.2GB (83%) | 128K | 169.6 | Vision, web extract, light tasks |
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | Genesis Hermes V3 APEX (LuffyTheFox, 24GB) | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs |
Key notes:
- All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) is deprecated and NOT in the inference path.
- Stable aliases (gpu-dense, gpu-light, strix-moe) from gpu-fleet are the canonical names for agent configs. Model-specific names still work but are deprecated.
- RTX 5070 tok/s is 2.3x faster than RTX 3090 for its model — gpu-light is the fastest endpoint. Route vision/web/light work there first.
- Strix Halo is 62.9 tok/s (89% of 70.5 baseline) — below optimal but stable. Check for competing workloads.
- RTX 3090 VRAM at 88% — within role-appropriate range (role = heavy reasoning, needs the headroom).
- RTX 5070 VRAM at 83% — role-appropriate for vision/web (smaller batch sizes).
## Remediation Rules
### Rule 1: GPU Temperature Critical (>85°C for >2 min)
- **Detect**: Any GPU temp >85°C sustained for 2+ consecutive polls
- **Fix**:
1. Reduce inference concurrency on that GPU (load-side cooling only — NO fan control)
2. Redirect new requests to cooler GPUs via LiteLLM fallback chains
2. Redirect new requests to cooler GPUs via LiteLLM fallback chains (gemma → qwen, qwen → gemma)
3. If all GPUs hot, alert about cooling infrastructure
- **Verify**: Temp drops below 80°C within 5 minutes
- **Escalate after**: 3 verification failures → Zulip alert
### Rule 2: VRAM Leak Detection (tiered by GPU capacity)
- **Detect**: VRAM growing at sustained rate over 6+ hour window
- RTX 3090 (24GB): ≥100MB/hour
- RTX 5070 (12GB): ≥50MB/hour
- RTX 3090 (24GB): ≥300MB/hour
- RTX 5070 (12GB): ≥300MB/hour
- Strix Halo (64GB UMA): ≥200MB/hour
- **Fix**:
1. Log VRAM snapshot with process list (nvidia-smi/rocm-smi + ps aux)
@@ -69,6 +90,9 @@ depends_on:
### Rule 4: Benchmark Regression (>20% drop)
- **Detect**: gen_tok_per_sec drops >20% below baseline over 3+ benchmarks
- RTX 3090 baseline: 74.8 tok/s → alert at <59.8 tok/s
- RTX 5070 baseline: 165.2 tok/s → alert at <132.2 tok/s
- Strix Halo baseline: 70.5 tok/s → alert at <56.4 tok/s
- **Fix**:
1. Check GPU utilization — if >90%, other process is competing
2. Check power limit — if throttled, restore to max
@@ -78,32 +102,31 @@ depends_on:
### Rule 5: Circuit Breaker Stuck Open
- **Detect**: Circuit breaker open >10 minutes with GPU reporting healthy
- **Note**: Router (port 9000) is deprecated. If circuit breakers are reported by gpu-monitor, they come from LiteLLM's internal tracking, not the old router.
- **Fix**:
1. Verify GPU /health returns 200
2. If GPU healthy, send 1 test inference
3. If test succeeds → reset circuit breaker via router API
4. 60s cooldown — if CB re-opens immediately, it was legitimate, do NOT re-reset
5. Max 1 auto-reset per GPU per hour
- **Verify**: CB closes, inference succeeds, CB stays closed for 60s+
- **Escalate after**: CB won't close after reset → router issue
1. Verify GPU /health returns 200 on direct port (:8080)
2. If GPU healthy, alert but do NOT reset via router API (deprecated)
3. Check LiteLLM health directly: http://192.168.68.116/litellm/health/liveliness
4. Restart LiteLLM container on CT 116 if circuit breakers are stuck
- **Verify**: LiteLLM returns healthy, circuit breaker clears within 60s
- **Escalate after**: LiteLLM restart doesn't clear → human investigation
### Rule 6: Strix Halo Unreachable
- **Detect**: Strix not responding — probe .15:8080 directly (firewall opened .24→.15)
- **Fix**:
1. SSH to .15 → check llama-server process
2. Restart llama-server if not running
3. Verify through both direct probe AND router
- **Verify**: Direct health probe returns 200, router reports Strix healthy
3. Verify through both direct probe AND LiteLLM health
- **Verify**: Direct health probe returns 200, LiteLLM reports model healthy
- **Escalate**: If host .15 itself is unreachable → infrastructure alert
### Rule 7: Prometheus Exporter Down
- **Detect**: Any GPU :9400/metrics unreachable for >2 polls
### Rule 7: GPU Data Source Unreachable (replaces old Prometheus rule)
- **Detect**: gpu-monitor endpoint (localhost:9100/gpu-data) or sidecar port (:8080) on any GPU unreachable for >2 polls
- **Fix**:
1. SSH to GPU host → check prometheus-exporter process
2. Restart exporter if dead
3. While exporter is down, fall back to nvidia-smi/rocm-smi direct probes
4. If exporter is running but unreachable → check firewall/host networking
- **Verify**: :9400/metrics returns 200
1. If gpu-monitor is down: restart systemd service `gpu-monitor.service` on this host
2. If sidecar is down: SSH to GPU host → check llama-server process → restart systemd service
3. Fall back to direct nvidia-smi/rocm-smi probe via SSH if all API paths fail
- **Verify**: gpu-monitor returns healthy + all sidecars reachable
- **Escalate after**: 3 failed restarts → networking issue
### Rule 8: Predictive Thermal Warning (two-tier)
@@ -117,29 +140,31 @@ depends_on:
- **Escalate**: If Tier 2 triggers and temp still rising after 5 min → possible hardware failure
### Rule 9: Context Window Optimization
- **Detect**: Benchmark tok/s vs baseline for each GPU at current context
- RTX 3090 (256K ctx, qwen3.6-27B-code): target 75+ tok/s — currently at baseline
- RTX 5070 (131K ctx, gemma-4-12b): target 76+ tok/s — optimal for vision/web role
- Strix Halo (256K ctx, strix-moe / qwen3.6-35B-udq4): target 70+ tok/s — currently above baseline
- **Detect**: Benchmark tok/s vs baseline for each GPU at current context (all 128K)
- RTX 3090 (128K ctx, ThinkingCap): baseline 74.8 tok/s — currently at 74.9 (100%)
- RTX 5070 (128K ctx, HauhauCS QAT): baseline 165.2 tok/s — currently at 169.6 (103%)
- Strix Halo (128K ctx, Genesis Hermes V3): baseline 70.5 tok/s — currently at 62.9 (89%)
- **Fix**:
- If tok/s > baseline → context has headroom, consider increasing
- If tok/s < 90% baseline → reduce context by 25% and retest
- If tok/s within 10% of baseline → optimal, no change
- Strix Halo at 89% of baseline → MONITOR but do not reduce yet (recent model swap may still be settling)
- **Verify**: Re-benchmark after context change, confirm within 10% of target
- **Escalate**: If context can't be adjusted without significant perf loss
### Rule 10: Workload Distribution Optimization
### Rule 10: Workload Distribution Optimization (updated 2026-07-18)
- **Detect**: GPU roles misaligned with hardware capabilities
- **Target distribution**:
- RTX 3090 (24GB, 256K, 75 tok/s) → Heavy reasoning, code gen, long conversations
- RTX 5070 (12GB, 131K, 76 tok/s) → Vision/image, web search, quick lightweight tasks
- Strix Halo (64GB, 256K, 72 tok/s) → Context compression, summarization, long docs
- RTX 3090 (gpu-dense, 24GB, 74.9 tok/s) → Heavy reasoning, code gen, long conversations (slowest per-token but largest context capacity). Weight: 0.55 (LiteLLM).
- RTX 5070 (gpu-light, 12GB, 169.6 tok/s) → Vision/image, web search, lightweight tasks (2.3x faster than 3090 per token). Weight: 0.15 (LiteLLM).
- Strix Halo (strix-moe, 64GB, 62.9 tok/s) → Context compression, summarization, long docs (MoE model). Weight: 0.30 (LiteLLM).
- **Note**: RTX 5070 is the fastest endpoint per token. Route high-volume, low-complexity work there first.
- **Fix**:
- Alert if any GPU is handling workload outside its designated role
- Recommend Hermes agent profile updates to match workload to GPU
- Recommend agent alias updates to match workload to GPU role (use stable aliases: gpu-dense, gpu-light, strix-moe)
- Track per-GPU request distribution via LiteLLM spend logs
- **Verify**: Each GPU's request pattern matches its designated role within 24h
- **Escalate**: If role mismatch persists >48h → agent profile audit needed
- **Escalate**: If role mismatch persists >48h → agent alias audit needed
---
@@ -148,7 +173,7 @@ depends_on:
```prose
-- Phase 1: Fetch live GPU data
let fleet = call gpu-monitor
endpoint: "http://192.168.68.24:9100/gpu-data"
endpoint: "http://localhost:9100/gpu-data"
-- Phase 2: Evaluate each GPU against remediation rules
let actions = []
@@ -164,27 +189,25 @@ for gpu in fleet.gpus:
-- Rule 4: Benchmark regression
let bench = fleet.benchmarks[gpu.hostname]
if bench.current_tok_sec < bench.baseline_tok_sec * 0.8:
if bench.current_tok_s < bench.baseline_tok_s * 0.8:
push actions apply-benchmark-fix(gpu, bench)
-- Rule 3: Model stuck
for model in fleet.router.available_models:
for model in fleet.summary.available_models:
if model.consecutive_timeouts >= 3:
push actions apply-model-restart(model)
-- Rule 5: Circuit breaker
for cb in fleet.router.circuit_breaker:
if cb.open and cb.open_duration > 600 and gpu_is_healthy(cb.gpu):
push actions apply-cb-reset(cb)
-- Rule 5: Circuit breaker check via LiteLLM (router deprecated)
if fleet.summary.circuit_breakers_open > 0:
push actions check-litellm-circuit-breakers()
-- Rule 6: Strix Halo
if not fleet.strix.running and pingable("192.168.68.15"):
push actions apply-strix-restart()
-- Rule 7: Prometheus exporters
for gpu in fleet.gpus:
if not prometheus_reachable(gpu.hostname, 9400):
push actions apply-exporter-restart(gpu)
-- Rule 7: GPU data source
if not fleet.gpus or len(fleet.gpus) < 2:
push actions check-gpu-monitor-service()
-- Rule 8: Predictive thermal
for gpu in fleet.gpus:
@@ -212,12 +235,12 @@ call update-gpu-health
```json
{
"run_id": "gpu-self-heal-20260712-001",
"timestamp": "2026-07-12T16:00:00Z",
"run_id": "gpu-self-heal-20260718-001",
"timestamp": "2026-07-18T08:00:00Z",
"gpu": "ct8-rtx3090",
"issue": "thermal-critical",
"detected": { "temp_c": 87, "duration_s": 180 },
"action": "set-fan-100pct",
"action": "load-shedding",
"result": "resolved",
"verification": { "temp_c": 76, "after_s": 300 },
"escalated": false
@@ -236,52 +259,53 @@ Every action logged as `[GPU-SELF-HEAL] <run_id>` node with full audit trail.
- `issues_escalated > 0` → "⚠ GPU Self-Heal — <gpu> needs attention"
- Every 100th clean cycle → "✅ GPU Fleet: All Clear"
### 3. Prometheus/Grafana Integration
- GPU self-heal actions exposed as Prometheus counter metrics
- Dashboard panel: "GPU Interventions (24h)" showing count/type/result
### 4. Weekly Benchmark Report
### 3. Weekly Benchmark Report
- Per-GPU tok/s trend over 7 days
- Regression alerts if any GPU degrades >10% week-over-week
---
## Design Decisions (Grilled & Confirmed 2026-07-12)
## Design Decisions (Verified 2026-07-12, Reaffirmed 2026-07-18)
1. **Fan control**: ❌ NO auto fan control. Load-side cooling only (reduce concurrency, redirect).
2. **Model restart**: ✅ Only if >50% failure rate over 60s + 30s grace period. Not on single stuck request.
3. **Strix direct access**: ✅ Open firewall .15:8080 → .24 for direct health probe + restart.
4. **VRAM thresholds**: Tiered — 100MB/h (RTX 3090), 50MB/h (RTX 5070), 200MB/h (Strix).
5. **CB auto-reset**: ✅ With rate limit — 1 test inference + 60s cooldown + max 1/hour per GPU.
6. **Benchmark baseline**: Rolling 30-day average, recalculated weekly. Original baseline kept in Grafana.
4. **VRAM thresholds**: Tiered — **300MB/h** (RTX 3090), **300MB/h** (RTX 5070), 200MB/h (Strix). Previous values (100/50) were too sensitive; raised 2026-07-18 based on operational data.
5. **CB auto-reset**: ✅ Router deprecated — circuit breakers go through LiteLLM health check + container restart if needed. No per-GPU auto-reset.
6. **Benchmark baseline**: Rolling 30-day average, recalculated weekly. Current baselines live in gpu-monitor.
7. **Predictive alerts**: Two-tier — warn at >70°C+rising (>2°C/min), critical at >80°C+rising.
8. **Prometheus**: Primary source. Fall back to nvidia-smi/rocm-smi direct probes if exporter down.
8. **Prometheus**: ❌ Not deployed. Use direct sidecar probes (:8080/health) and gpu-monitor API. Prometheus integration deferred until exporters are running on GPU hosts.
## Lessons Learned (2026-07-12)
## Lessons Learned (2026-07-12, Updated 2026-07-18)
### L1: API Key Standardization Is Critical
- All GPU llama-servers MUST use the same api-key as the LiteLLM config.
- RTX 5070 had `--api-key sk-loc...5678` while LiteLLM sent `not-needed`.
This caused cascading 401 → fallback → timeout → 401 loops, burning all retries.
- **Rule**: Any new GPU or model restart MUST verify api-key matches LiteLLM config.
This caused cascading 401 → fallback → timeout → 401 loops.
- **Rule**: Any new GPU or model restart MUST verify api-key matches LiteLLM config (`not-needed` for direct routing).
### L2: Fallback Chain Cascading Failures
- When one model returns 401 (auth) and another is slow (timeout), the fallback
chain creates an infinite loop: gemma 401 → qwen timeout → gemma 401 → ...
chain creates an infinite loop.
- **Rule**: If a model returns 401 (auth error), do NOT fall back to it again.
Mark it as permanently failed for this request.
### L3: Verify Running State, Not Docs
- RTX 3090 was documented at 128K context. Actually running at 256K.
- Parallel count wrong (docs said 2, actual is 1 on RTX 3090).
- RTX 3090 was documented at 128K context. Running at 128K (verified 2026-07-18).
- Parallel count: 1 on both RTX 3090 and RTX 5070 (matches docs for current models).
- **Rule**: Before making decisions, check `/proc/PID/cmdline` on GPU hosts.
### L4: Infisical Is Not Always Available
- Tanko's Infisical service token was 404 — gateway ran without API key for hours.
- **Rule**: Always keep a local `.env` fallback for `LITELLM_API_KEY`.
- Contract hermes-config-template Rule 3 updated.
- Keep a local `.env` fallback for `LITELLM_API_KEY`.
- **Rule**: Always verify credential source is reachable before relying on it.
### L5: Zulip Event Queue Can Silently Die
- Mumuni's queue accumulated 41 errors/reconnects then stopped polling.
Gateway was running but ignoring all messages.
- **Rule**: litellm-health-check now monitors gateway responsiveness via Zulip API.
### L5: GPU Monitor Response Size Can Cause Self-Heal Crash
- gpu-self-heal crashed with KeyboardInterrupt during json.loads() of 20MB response.
- Root cause: router poll returns accumulated data → cache balloons.
- **Rule**: Self-heal must enforce a read timeout AND max response size on every poll.
If monitor response > 1MB, log a warning and skip the cycle rather than crashing.
### L6: Stable Aliases Replace Model Names
- gpu-fleet introduced stable aliases (strix-moe, gpu-dense, gpu-light) on 2026-07-15.
- Self-heal must use aliases for reporting and alerting, not model-specific names.
- **Rule**: All alert messages and KG nodes use the stable alias as the GPU identifier.
+7 -7
View File
@@ -5,7 +5,7 @@ version: 1.0.0
description: >
Canonical known-good baseline for all Syslog Hermes agents. Captures the exact
configuration state, keys, workarounds, and audit procedure. When an agent's
configuration goes sideways, restore from this baseline. Last verified 2026-07-08. GPU context reduced to 128K on .8/.110, parallel 2 fleet-wide.
configuration goes sideways, restore from this baseline. Last verified 2026-07-16. All GPUs 256K context (RTX 3090 .8, RTX 5070 .110, Strix Halo .15). Parallel 1 fleet-wide (Strix Halo handles compression solo).
author: Abiba (pi agent)
---
@@ -26,11 +26,11 @@ done
|-------|-----|------|-----|---------------|------------|----------|
| Tanko | 112 | amdpve | .122 | `tanko` | Infisical vault | Hermes |
| Mumuni | 114 | minipve | .123 | `mumuni` | Infisical vault | Hermes |
| Koby | 111 | amdpve | srv1079750 | `koby` | Infisical vault | **Hermes** |
| Koonimo | 113 | amdpve | ? | `koonimo` | Infisical vault | Hermes |
| Shumba | — | 192.168.68.119 | N/A | N/A (DeepSeek) | Hermes |
| Koby | 129 | amdpve | srv1079750 | `koby` | Infisical vault | **Hermes** |
| Koonimo | 114 | amdpve | ? | `koonimo` | Infisical vault | Hermes |
| Shumba | — | 192.168.68.119 | N/A | N/A (DeepSeek) | Hermes (RETIRED — CT119 now Infisical vault) |
> **Note**: CT hostnames (tdunna→CT111, baggy→CT113) differ from agent identities (koby, koonimo).
> **Note**: CT hostnames (tdunna→CT129, baggy→CT114) differ from agent identities (koby, koonimo).
Access: `pct-run <CT_ID> <command>` — no IPs needed. GPU hosts (.8, .110, .15) use SSH.
Keys are stored in Infisical vault (project=agents, env=production) and injected at
@@ -169,9 +169,9 @@ pct-run <CT> grep -A8 "vision:" /root/.hermes/config.yaml | grep api_key
# Must show both api_key: sk-... and api_key_env: LITELLM_API_KEY
```
### For Koby (CT 111 / tdunna)
### For Koby (CT 129 / tdunna)
Koby runs Hermes on CT 111 (tdunna). Config files at `/root/.hermes/config.yaml`.
Koby runs Hermes on CT 129 (tdunna). Config files at `/root/.hermes/config.yaml`.
Same Hermes pattern as Tanko/Mumuni/Koonimo — see config sections above.
**LiteLLM key**: alias `koby` in LiteLLM DB, injected via `infisical run --` wrapper.
+20 -21
View File
@@ -6,11 +6,10 @@ description: >
Enforces shared infrastructure setup (Firecrawl, SearXNG, local models,
RA-H OS MCP) while keeping agent-specific API keys and model choices.
UPDATED 2026-07-16: Compression model is the stable alias `strix-moe` (NOT `ornith-1.0-35b`,
which LiteLLM does not serve). All 3 GPUs verified 256K (RTX 5070 bumped 131K→256K on Jul 15).
which LiteLLM does not serve). All 3 GPUs verified at 128K (reduced from 256K 2026-07-17 for stability).
Added Rule 12 (Context-Issue Diagnostic) + Rule 13 (.env fallback enforcement) from the
2026-07-16 Mumuni root-cause investigation (WAL #1300).
UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo. RTX 3090 context
verified at 256K. Infisical .env fallback required (Rule 3/13).
UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo. RTX 3090 context verified at 128K. Infisical .env fallback required (Rule 3/13).
---
## Maintains
@@ -36,7 +35,7 @@ Sub-agent profiles inherit auth from the main config — no separate keys needed
| Mumuni | `mumuni` | 192.168.68.123 | root@.123 | 6 profiles ✱ |
| Abiba | `abiba-pi` | 192.168.68.24 | local | — |
| Koby | `koby` | CT 111 (tdunna) | Zulip | — |
| Koonimo | `koonimo` | CT 113 (baggy) | Zulip | — |
| Koonimo | `koonimo` | CT 114 (baggy) | SSH root | — |
| Kagenz0 | `kagenz0-*` | ? | Zulip | — |
> CT hostnames (tdunna, baggy) differ from agent identities (koby, koonimo).
@@ -101,8 +100,8 @@ model:
base_url: http://192.168.68.116/v1
api_key_env: LITELLM_API_KEY # Injected via infisical run -- wrapper
max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation
context_length: 262144 # For syslog-auto (all GPUs support 256K).
# Set 131072 if using gemma-4-12b directly (12GB VRAM constraint).
context_length: 131072 # For syslog-auto (all GPUs at 128K for stability).
# Set 65536 if using gemma-4-12b directly (tight VRAM).
fallback_providers:
provider: deepseek
@@ -133,7 +132,7 @@ compression:
enabled: true
model: strix-moe # ⚠️ Must match auxiliary.compression.model. Stable alias (gpu-fleet § Stable Role-Based Aliases). NOT ornith-1.0-35b (LiteLLM does not serve that name).
provider: harness
max_context_window: 262144 # MUST match actual GPU capacity. All 3 GPUs are 256K (Jul 15).
max_context_window: 131072 # MUST match actual GPU capacity. All 3 GPUs are 128K (Jul 17).
threshold: 0.65 # Fires at ~170K for 262K window, ~85K for 128K
target_ratio: 0.30
protect_last_n: 40
@@ -250,7 +249,7 @@ The following MUST be identical across ALL profiles:
### Rule 7: Auxiliary Model Consistency (UPDATED 2026-07-16)
- Vision and web_extract use `gemma-4-12b` (RTX 5070 — 12GB, vision-optimized)
- Compression uses `strix-moe` (stable alias for Strix Halo — 64GB, 256K ctx, compression-optimized)
- Compression uses `strix-moe` (stable alias for Strix Halo — 64GB, 128K ctx, compression-optimized)
- **`strix-moe` is the only valid compression model name** — LiteLLM does NOT serve `ornith-1.0-35b`
(it serves `strix-moe`, `qwen3.6-35B-udq4`, `gpu-dense`, `gpu-light`, `syslog-auto`, `gemma-4-12b`, `qwen3.6-27B-code`). Old configs with `ornith-1.0-35b` cause 403/model-not-found on compression calls.
- All auxiliary services MUST use identical routing:
@@ -258,31 +257,31 @@ The following MUST be identical across ALL profiles:
- `api_key_env: LITELLM_API_KEY`
- **Do NOT use `syslog-auto`** for auxiliary tasks — it routes unpredictably
- **Compression on Strix Halo**: The strix-moe alias routes to Strix Halo
(64GB UMA, 256K context) — the designated compression GPU. This frees the
(64GB UMA, 128K context) — the designated compression GPU. This frees the
RTX 5070 for vision and web search, and the RTX 3090 for heavy reasoning.
- The `compression:` block's `model` MUST match `auxiliary: compression: model`
- The `compression: max_context_window: 262144` MUST match Strix Halo's actual capacity
- The `compression: max_context_window: 131072` MUST match actual GPU capacity (128K)
### Rule 8: GPU Workload Distribution (UPDATED 2026-07-16)
- **RTX 3090 (24GB, 256K ctx, qwen3.6-27B-code)**: Heavy reasoning, code gen, long conversations
- **RTX 5070 (12GB, 256K ctx, gemma-4-12b)**: Vision, web search, quick tasks, web_extract (bumped 131K→256K Jul 15; IQ4_NL+MTP, 88% VRAM)
- **Strix Halo (64GB, 256K ctx, strix-moe)**: Context compression, summarization, long docs
- **RTX 3090 (24GB, 128K ctx, qwen3.6-27B-code)**: Heavy reasoning, code gen, long conversations
- **RTX 5070 (12GB, 128K ctx, gemma-4-12b)**: Vision, web search, quick tasks, web_extract (IQ4_NL+MTP, ~65% VRAM at 128K)
- **Strix Halo (64GB, 128K ctx, strix-moe)**: Context compression, summarization, long docs
- Agent profiles MUST route auxiliary tasks to the correct GPU:
- `auxiliary.vision.model: gemma-4-12b` (RTX 5070)
- `auxiliary.web_extract.model: gemma-4-12b` (RTX 5070)
- `auxiliary.compression.model: strix-moe` (Strix Halo)
- Default model (`model.default`) and custom_provider remain `syslog-auto` for auto-routing
- For 262K context window: `threshold: 0.65` (fires at ~170K tokens)
- For 128K context window: `threshold: 0.65` (fires at ~85K tokens)
- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss
- Do NOT use `threshold: 0.80` — this delays until 209K, risking the gateway hygiene layer
- `max_context_window: 262144` MUST match the model's actual capacity
- Do NOT use `threshold: 0.80` — this delays until ~105K, leaving only 23K margin
- `max_context_window: 131072` MUST match the model's actual capacity (128K)
- See `devops-hermes-compression` skill for full reference
### Rule 9: Compression Threshold for 256K Models
- For 262K context window: `threshold: 0.65` (fires at ~170K tokens)
### Rule 9: Compression Threshold for 128K Models
- For 128K context window: `threshold: 0.65` (fires at ~85K tokens)
- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss
- Do NOT use `threshold: 0.80` — this delays until 209K, risking the gateway hygiene layer
- `max_context_window: 262144` MUST match the model's actual capacity (Strix Halo = 256K)
- Do NOT use `threshold: 0.80` — this delays until ~105K, leaving only 23K margin
- `max_context_window: 131072` MUST match the model's actual capacity (all GPUs = 128K)
- See `devops-hermes-compression` skill for full reference
### Rule 10: Default Model Must Be `syslog-auto` (All Agents)
@@ -312,7 +311,7 @@ When an agent shows "context issues" (premature compression, 401s, 504s, DeepSee
verify ALL FOUR of these against the live config. They are the only root causes found in production:
1. **max_context_window correct?** — BOTH `compression.max_context_window` AND `context.max_context_window`
MUST be `262144` (all GPUs are 256K). A value of `131072` causes premature compression at
MUST be `131072` (all GPUs are 128K). A value of `262144` causes instability near 100K and must NOT be used.
~83K instead of ~170K. Check: `grep -n max_context_window ~/.hermes/config.yaml`
2. **base_url uses /v1 NOT /litellm/v1?** — `custom_providers[0].base_url`, `delegation.base_url`,
and ALL `auxiliary.*.base_url` MUST be `http://192.168.68.116/v1` (Rule 5). nginx `/litellm/`
+104
View File
@@ -0,0 +1,104 @@
---
name: inference-optimization
kind: responsibility
description: >
Optimizes the full Syslog inference stack — LiteLLM routing weights, GPU model
assignments, agent context management, and prompt caching — to reduce response
times to sub-15s average. All GPUs now at 128K context (stable ceiling).
id: 067NC6KP02RG60S50M40E30928
---
### Goal
Syslog inference response times reduced to sub-15s average by optimizing the full
stack: LiteLLM routing weights, GPU model assignments, Hermes agent context
management, and prompt caching — without sacrificing agent capability.
### Requires
- `inference-metrics`: current SpendLogs from CT116 LiteLLM Postgres — avg
request_duration_ms, prompt_tokens, completion_tokens, model_group breakdown,
cache_hit rate over the last 3 hours
- `agent-configs`: current config.yaml from each active Hermes agent (Mumuni
.123, any others on .129/.122) including compression, model, context_window,
prompt_caching, memory settings
- `gpu-health`: health check response from all 3 GPU backends (ornith .15:8080,
qwen .8:8080, gemma .110:8080)
### Maintains
The optimized inference stack configuration — every change is applied and
verified end-to-end. Postcondition: avg request_duration_ms ≤ 15000 for 90% of
non-ornith traffic; ≤ 30000 for ornith-bound agentic calls.
#### liteLLM-routing
The syslog-auto routing weights, model-specific timeouts, RPM limits, and
model_list entries on CT116 `/opt/inference-harness/litellm_config.yaml`.
#### agent-compression
Each Hermes agent's `~/.hermes/config.yaml` compression, context_window,
prompt_caching, and model sections.
#### prompt-caching
LiteLLM cache configuration and llama.cpp `--cache-prompt` flag on GPU hosts.
#### verification
End-to-end latency measurements after changes applied — at least 3 test
inference calls per model path measuring ttft (time-to-first-token) and total
duration.
### Continuity
- input-driven
### Strategies
**Context is the root cause.** Every ~46K prompt token costs ~87s of ornith
prefill time at 532 tok/s. Fix context first, routing second.
- **Route by task**: ornith for multi-step reasoning only; qwen for code/standard
queries; gemma for compression/auxiliary. Never send simple completion to a
35B MoE.
- **Compress aggressively**: threshold at 40% (not 65%) — a 256K window should
compact at 102K, not 166K. Target 15% tail (not 30%).
- **Cache everything repeated**: system prompts, skill docs, AGENTS.md — these
never change between turns. Single-digit cache hit rate is unacceptable.
- **Lower context ceiling**: 128K window is the stable ceiling for agent conversations.
GPUs reduced from 256K to 128K (2026-07-17). For larger contexts, route to external providers.
### Shape
- `self`: analyze metrics, compute optimal configs, apply changes, verify
- `delegates`:
- `apply-liteLLM`: update litellm_config.yaml and reload
- `apply-agent-config`: update hermes config.yaml per agent
- `verify-latency`: run test inference calls and measure response
### Execution
```prose
-- Phase 1: Analyze current state (already complete)
-- Phase 2: Apply LiteLLM routing optimization
call apply-liteLLM-routing
config_path: /opt/inference-harness/litellm_config.yaml
host: 192.168.68.116
-- Phase 3: Apply agent context compression optimization
call apply-agent-compression
agent: mumuni
host: 192.168.68.123
config_path: /root/.hermes/config.yaml
-- Phase 4: Enable llama.cpp prompt caching on GPU hosts
call enable-prompt-caching
hosts: [192.168.68.15, 192.168.68.8, 192.168.68.110]
-- Phase 5: Verify end-to-end latency
call verify-latency
host: 192.168.68.116
models: [syslog-auto, qwen3.6-27B-code, gemma-4-12b, ornith-1.0-35b]
```
+147 -79
View File
@@ -20,9 +20,20 @@ description: >
Canonical process: see § Production Vault Access Process. Tanko (user jerome) pending.
Abiba's key is now a proper agent key (NOT the master key — stale note removed).
UPDATED 2026-07-17: FLEET-WIDE STANDARDIZATION. All 4 agents (Mumuni, Tanko, Koby, Koonimo)
standardized on a single pattern: systemd drop-in (ExecStart= reset + wrapper path) →
infisical-gateway.sh while-true loop → /usr/bin/infisical run --token → bash -c key
injection → .env fallback → exec python. Systemd drop-ins are IMMUNE to hermes gateway
install which overwrites the unit file ExecStart. Infisical CLI updated to 0.43.109 on
all agents (was 0.38.0). Service token st.8e848433 shared across fleet (st.353699cd
for tanko was deleted). .env fallback on every agent protects against token loss.
Critical lessons: (1) NEVER use shell variables inside single-quoted bash -c in wrappers
— hardcode absolute paths. (2) Drop-ins override unit file ExecStart permanently.
(3) Capture /proc/<pid>/environ before gateway restarts to preserve running env set.
Current key inventory and agent list: see gpu-fleet.prose.md § Agent Keys.
Source of truth for LiteLLM config: /opt/inference-harness/litellm_config.yaml
on CT 116. Last verified: 2026-07-16.
on CT 116. Last verified: 2026-07-17.
---
## Parameters
@@ -74,84 +85,147 @@ description: >
- Confirm key alias matches agent_name in LiteLLM key list
- Verify agent gateway uses vault wrapper: `cat /proc/<pid>/cmdline` shows `infisical run`
## Production Vault Access Process (canonical, 2026-07-16)
## Production Vault Access Process (canonical, 2026-07-17)
The non-fail approach to agentic vault access. Deployed on 4/5 agents (tanko pending —
runs as user `jerome`, not systemd root, needs user-scope adaptation).
The non-fail approach to agentic vault access. Deployed on all 4 Hermes agents
(Mumuni, Tanko, Koby, Koonimo) as of 2026-07-17. Abiba (pi) uses a similar pattern
through its agent wrapper.
### The canonical pattern
1. **infisical CLI** installed on the host (`/usr/local/bin/infisical` or `/usr/bin/infisical`).
2. **Service token** (Infisical Machine Identity, `st.…`) stored root-only at `/root/.infisical-token` (`chmod 600`).
- Interim: the shared `abiba` service token (`st.8e848433…`) has READ+WRITE on the `agents` project.
- Proper: one machine identity per agent (create in Infisical UI → Project Settings → Machine Identities).
3. **`infisical-gateway.sh` wrapper** at `/root/.hermes/infisical-gateway.sh` (`chmod 700`):
1. **infisical CLI** installed on the host at `/usr/bin/infisical` (v0.43.109+, from
artifacts-cli.infisical.com apt repo). Update procedure:
```bash
curl -1sLf 'https://artifacts-cli.infisical.com/setup.deb.sh' | sudo -E bash
sudo apt-get update && sudo apt-get install -y infisical
# Remove stale old binary if present
rm -f /usr/local/bin/infisical /bin/infisical
```
Wrappers use absolute path `/usr/bin/infisical run`. Never rely on PATH resolution.
2. **Service token** (Infisical Machine Identity, `st.…`) stored at `~/.infisical-token`
(`chmod 600`). Current: shared `st.8e848433…` (abiba, READ+WRITE on agents project).
Tanko's `st.353699cd…` (tanko-agent) was deleted — reverted to shared token.
Proper: one machine identity per agent (create in Infisical UI → Project Settings →
Machine Identities).
3. **`infisical-gateway.sh` wrapper** at `~/.hermes/infisical-gateway.sh` (`chmod 700`):
```bash
#!/bin/bash
export INFISICAL_API_URL="https://vault.sysloggh.net"
TOKEN=$(cat /root/.infisical-token)
LOG=/root/.hermes/logs/gateway.log; mkdir -p /root/.hermes/logs
TOKEN=$(cat $HOME/.infisical-token)
LOG=$HOME/.hermes/logs/gateway.log; mkdir -p $HOME/.hermes/logs
while true; do
infisical run --token="$TOKEN" --projectId=322fceab-39da-4854-a55a-568e76c0f13f \
echo "[$(date -Iseconds)] Starting gateway with Infisical injection..." >> $LOG
/usr/bin/infisical run --token="$TOKEN" \
--projectId=322fceab-39da-4854-a55a-568e76c0f13f \
--env=prod --domain=https://vault.sysloggh.net -- bash -c '
. /root/.hermes/.env 2>/dev/null # [FALLBACK Rule 3] safety net only
export LITELLM_API_KEY="$<AGENT>_LITELLM_API_KEY"
exec <HERMES_VENV>/bin/python -m hermes_cli.main gateway run
. $HOME/.hermes/.env 2>/dev/null # [FALLBACK Rule 3]
export LITELLM_API_KEY="${<AGENT>_LITELLM_API_KEY}"
export ZULIP_API_KEY="${<AGENT>_ZULIP_API_KEY}"
export ZULIP_SITE="https://chat.sysloggh.net"
export ZULIP_EMAIL="<agent>-bot@chat.sysloggh.net"
export SEARXNG_URL="http://192.168.68.7:8888"
# ⚠️ HARDCODE the full venv path. NEVER use $VENV inside single quotes.
exec /root/.hermes/hermes-agent/venv/bin/python -m hermes_cli.main gateway run
' >> $LOG 2>&1
sleep 5 # restart on exit
EXIT_CODE=$?
echo "[$(date -Iseconds)] Gateway exited with code $EXIT_CODE — restarting in 5s..." >> $LOG
sleep 5
done
```
4. **Agent key in vault** as `<AGENT>_LITELLM_API_KEY` (e.g. `KOBY_LITELLM_API_KEY`). Vault = source of truth.
5. **`.env` fallback** at `/root/.hermes/.env` (`chmod 600`) with the same key — safety net ONLY for vault outage (Rule 3/13). Must be kept in sync on rotation.
6. **systemd service** `hermes-gateway.service` with `ExecStart=/root/.hermes/infisical-gateway.sh`. NO `litellm-key.conf` drop-in (those hardcode keys and rot).
7. **NEVER hardcode** LiteLLM keys in systemd drop-ins, config.yaml, or /etc/environment. The wrapper injects live from vault.
**CRITICAL: VENV PATH.** The inner `bash -c '...'` uses single quotes. Shell
variables set in the outer wrapper are NOT expanded inside single quotes.
`$VENV/bin/python` resolves to `/bin/python` (file not found). Always hardcode
the absolute path to the venv python binary.
4. **Agent key in vault** as `<AGENT>_LITELLM_API_KEY` and `<AGENT>_ZULIP_API_KEY`.
Vault = source of truth for ALL platform credentials.
5. **`.env` fallback** at `~/.hermes/.env` (`chmod 600`) with agent-specific keys —
safety net for vault outage or token revocation. Must be kept in sync on rotation.
Example:
```bash
MUMUNI_LITELLM_API_KEY=sk-OzuWsoX22Hmb3Ps3JY01gw
MUMUNI_ZULIP_API_KEY=H8dY6V7aHmWNcfgNtJaDBPZ1dGWn0Ttt
```
6. **systemd drop-in** at `~/.config/systemd/user/hermes-gateway.service.d/50-vault-wrapper.conf`:
```ini
[Service]
ExecStart=
ExecStart=/root/.hermes/infisical-gateway.sh
```
The `ExecStart=` (empty reset) clears any ExecStart from the main unit file,
then the second `ExecStart=` sets the wrapper. This drop-in **survives unit file
regeneration** by `hermes gateway install` — the drop-in always wins.
**Why a drop-in instead of editing the unit file:** `hermes gateway install`
(called during Hermes updates and some self-heal operations) regenerates the
systemd unit file with `ExecStart=/path/to/python -m hermes_cli.main gateway run`.
Editing the unit file directly is futile — it will be overwritten. The drop-in
approach explicitly resets ExecStart and sets the wrapper regardless of what the
main unit file says.
7. **NEVER hardcode** API keys in systemd drop-ins, config.yaml, or /etc/environment.
The wrapper injects live from vault at every start.
### Why this is non-fail
- **No rot**: keys pulled live from vault at every gateway start. Rotation = one `infisical secrets set` + `systemctl restart`. No per-host file edits.
- **Survives vault outage**: the `.env` fallback (Rule 3) keeps the gateway running if Infisical is unreachable.
- **Survives gateway crash**: the wrapper's `while true` + systemd `Restart=on-failure` revive the gateway.
- **Auditable**: `cat /proc/$(pgrep hermes_cli)/environ` shows the live key; `infisical secrets` shows the vault source.
- **Survives vault outage**: the `.env` fallback (Rule 3) keeps the gateway running if Infisical is unreachable or the service token is revoked.
- **Survives gateway crash**: the wrapper's `while true` + systemd `Restart=always` revive the gateway. Two-layer defense.
- **Survives Hermes updates**: systemd drop-in overrides unit file ExecStart — `hermes gateway install` cannot break the vault injection.
- **Survives reboot**: systemd user service + `loginctl enable-linger` ensures gateway starts at boot without a login session.
- **Auditable**: `cat /proc/$(pgrep hermes_cli)/environ` shows all injected keys; `infisical secrets` shows the vault source.
### Migration status (2026-07-16)
### Migration status (2026-07-17)
| Agent | Host | Pattern | Vault key | Status |
|-------|------|---------|-----------|--------|
| abiba | .24 | `infisical run` (pi agent wrapper, service token) | ABIBA_LITELLM_API_KEY | ✅ vault-backed |
| mumuni | .123 | infisical-gateway.sh + user-login machine identity | MUMUNI_LITELLM_API_KEY | ✅ vault-backed |
| koby | .129 | infisical-gateway.sh + service token (migrated 2026-07-16) | KOBY_LITELLM_API_KEY | ✅ vault-backed, Zulip (tanko-bot@) + Telegram |
| koonimo | .114 | infisical-gateway.sh + service token (migrated 2026-07-16) | KOONIMO_LITELLM_API_KEY | ✅ vault-backed |
> **Baggy = Koonimo (CT 113).** Deleted `BAGGY_LITELLM_API_KEY` from vault 2026-07-16. Only `KOONIMO_LITELLM_API_KEY` exists — one secret per agent.
| tanko | .122 | **hardcoded in config.yaml** (runs as user jerome, not systemd) | TANKO_LITELLM_API_KEY | ⚠️ TODO: migrate to user-scope wrapper |
| Agent | Host | Pattern | Keys | Status |
|-------|------|---------|------|--------|
| abiba | .24 | pi agent wrapper | ABIBA_LITELLM_API_KEY + ABIBA_ZULIP_API_KEY | ✅ vault-backed |
| mumuni | .123 | systemd drop-in + while-true wrapper + st.8e848433 | MUMUNI_LITELLM_API_KEY + MUMUNI_ZULIP_API_KEY | ✅ vault-backed + .env fallback |
| tanko | .122 | systemd drop-in + while-true wrapper + st.8e848433 (user jerome) | TANKO_LITELLM_API_KEY + TANKO_ZULIP_API_KEY | ✅ vault-backed + .env fallback |
| koby | .129 | systemd drop-in + while-true wrapper + st.8e848433 | KOBY_LITELLM_API_KEY, shares TANKO_ZULIP_API_KEY (tanko-bot) | ✅ vault-backed |
| koonimo | .114 | systemd drop-in + while-true wrapper + st.8e848433 | KOONIMO_LITELLM_API_KEY + KOONIMO_ZULIP_API_KEY | ✅ vault-backed |
### Tanko migration (pending)
> Tanko runs as user `jerome` — wrapper/token at `~/.hermes/infisical-gateway.sh` and
> `~/.infisical-token`. Linger enabled (`loginctl enable-linger jerome`) for boot startup.
Tanko runs the gateway as user `jerome` (not root/systemd), with the key hardcoded in
`/home/jerome/.hermes/config.yaml` (`api_key: sk-CggiHWlamQy…`, valid but not vault-sourced).
Migration: create a user-scope systemd service (`~/.config/systemd/user/hermes-gateway.service`)
with `infisical-gateway.sh` wrapper in jerome's home, token at `~/.infisical-token`, lingering
enabled (`loginctl enable-linger jerome`) so the user service runs without a login session.
### Tanko migration (COMPLETED 2026-07-17)
### Koby migration lessons (2026-07-16)
Tanko was the last agent migrated from hardcoded keys to vault wrapper.
Previously: key hardcoded in `/home/jerome/.hermes/config.yaml` (`api_key: sk-CggiHWlamQy…`)
and `zulip-env.conf` systemd drop-in. Now: user-scope systemd service with drop-in
`50-vault-wrapper.conf`, `infisical-gateway.sh` wrapper with while-true loop, token at
`~/.infisical-token`, `.env` fallback at `~/.hermes/.env`. Keys injected live from vault.
### Koby migration lessons (2026-07-16, updated 2026-07-17)
Migrated Koby from hardcoded systemd drop-in → `infisical-gateway.sh` wrapper.
**Two mistakes I made that broke the agent:**
**Three mistakes made:**
1. **Overwrote `/root/.hermes/.env`** without backing it up. The Zulip API key only existed
in the running process memory — the old .env was minimal (just LiteLLM key). Zulip creds were
inherited from the pre-migration gateway env, not stored in any file. Lost on restart.
2. **Only injected `LITELLM_API_KEY`** in the wrapper — forgot Zulip + Telegram credentials.
Agents need ALL their platform env vars. Missing vars cause silent adapter failures.
3. (2026-07-17 fix) **VENV variable in single-quoted bash -c**: `exec "$VENV/bin/python"`
inside single quotes resolved to `exec "/bin/python"` (file not found). Hardcoded full path.
**How Koby actually connects (2026-07-16):**
**How Koby actually connects:**
- Zulip: shares **Tanko's bot** (`tanko-bot@chat.sysloggh.net`, `TANKO_ZULIP_API_KEY=5PeD6f3zo…`).
Koby doesn't have its own Zulip bot (koby-bot@ doesn't exist in the swarm config).
- Telegram: token `828640…` recovered from `.env.bak-20260603` (18KB backup from June 2026).
Allowed users: 6679773481. Home channel: 6679773481.
- Telegram: token from `.env` fallback. Allowed users: 6679773481.
- Both platforms now connect through the wrapper's env injection.
**Golden rule for gateway restarts:** always `cat /proc/<pid>/environ` before killing the old
process — captures the live env set. Especially important when migrating gateways between
injection mechanisms.
**Golden rules for gateway restarts:**
1. Always `cat /proc/<pid>/environ` before killing the old process — captures the live env set.
2. Hardcode venv python path in wrapper — never use variables inside single-quoted bash -c.
3. Use systemd drop-ins (not unit file edits) to override ExecStart — survives Hermes updates.
### Fleet-wide standardization lessons (2026-07-17)
After auditing all 4 agents, five systemic patterns caused repeated failures:
1. **Three incompatible startup patterns** coexisted (systemd drop-in, direct python, orphaned wrapper)
2. **Systemd unit files reverted** by `hermes gateway install` during updates
3. **VENV variable scoping** broke wrappers on Koby and Mumuni (single-quote bash -c)
4. **Service token expiry** — Tanko's `st.353699cd` was deleted from Infisical
5. **No ZULIP_API_KEY** in env on Tanko — wrapper bypassed by systemd direct python
All resolved by the canonical drop-in + while-true wrapper pattern documented above.
### Key rotation procedure (one vault operation with this standard)
@@ -161,47 +235,41 @@ injection mechanisms.
4. Restart: `systemctl restart hermes-gateway`. The wrapper pulls the new key live.
5. Verify: `curl -H "Authorization: Bearer sk-NEW" http://192.168.68.116/v1/models` → 200.
## Machine Identity for Vault Writes (ADDED 2026-07-16, WAL #1300)
## Machine Identity for Vault Writes (UPDATED 2026-07-17)
**Problem:** The infisical CLI on agent hosts is logged in as a user session (jerome@sysloggh.com).
In CLI v0.38.0, `infisical secrets set` / `infisical export` fail with "project id missing" / "workspace
key 404" — a known bug where user-session auth works for `run` but NOT for `secrets set`. The apt
repo only ships 0.38.0, so `apt upgrade` does not help.
**Current state:** Infisical CLI updated to v0.43.109 on all agents (from v0.38.0).
The v0.38.0 bug (user-session auth fails for `secrets set`/`export`) is resolved.
Service token `st.8e848433…` (abiba, READ+WRITE) can write to vault from CLI.
**Proper fix — Machine Identity (Infisical automation best practice):**
Create a machine identity with READ+WRITE scope on the `agents` project (project_id=
`322fceab-39da-4854-a55a-568e76c0f13f`, env `prod`). Store client_id + client_secret securely.
Then vault writes work from any host:
```bash
# Get a machine-identity access token
TOKEN=$(curl -fsSL -X POST https://vault.sysloggh.net/api/v1/auth/universal-auth/login \
-H 'Content-Type: application/json' \
-d '{"clientId":"<CLIENT_ID>","clientSecret":"<CLIENT_SECRET>"}' | jq -r .accessToken)
# Write a secret via REST API v3
curl -fsSL -X PATCH https://vault.sysloggh.net/api/v3/secrets/MUMUNI_LITELLM_API_KEY \
-H "Authorization: Bearer $TOKEN" -H 'Content-Type: application/json' \
-d '{"environment":"prod","secretValue":"sk-<NEW_KEY>","workspaceId":"<WORKSPACE_ID>","type":"shared"}'
# OR via CLI: infisical secrets set --token=$TOKEN --projectId=322fceab... --env=prod ...
```
Creation requires the Infisical web UI (https://vault.sysloggh.net) under Project Settings →
Machine Identities, or an admin API call. **TODO: create `abiba-automation` machine identity
and store its credentials in the vault itself (or a root-only file).**
**Proper fix — per-agent Machine Identities:**
Create machine identities in Infisical UI → Project Settings → Machine Identities
for each agent with READ-only scope on the `agents` project. Store client_id +
client_secret per agent. Then vault writes use the shared abiba identity, and
reads use per-agent identities. This eliminates the single shared token risk.
**Interim (working now):** the `.env` fallback (hermes-config-template Rule 3/13). The
infisical-gateway.sh wrapper sources `~/.hermes/.env`, so its `<AGENT>_LITELLM_API_KEY`
overrides a stale vault value.
**Service Token Inventory (2026-07-17):**
| Token ID | Name | Permissions | Used By | Status |
|----------|------|-------------|---------|--------|
| `st.8e848433…` | tanko-gateway | READ+WRITE | Mumuni, Tanko, Koby, Koonimo, Abiba | ✅ Active |
| `st.353699cd…` | tanko-agent | READ-only | — | ❌ Deleted from Infisical |
**2026-07-16 UPDATE — vault is now SYNCED.** The abiba service token (`st.8e848433…`, READ+WRITE)
can write to the vault, so the session-13 rotated keys (mumuni `sk-OzuWsoX2…`, koby `sk-BqRRMboTI…`,
koonimo `sk-OEK7z26n6E…`) are now in the vault as `MUMUNI_LITELLM_API_KEY` / `KOBY_LITELLM_API_KEY` /
`KOONIMO_LITELLM_API_KEY` and validate 200 against LiteLLM. The vault is the source of truth again.
Creating a dedicated `abiba-automation` machine identity (via UI) is still the proper long-term fix
so the shared service token isn't reused across hosts — but it is no longer blocking.
**Per-agent .env fallback inventory (2026-07-17):**
| Agent | .env Keys |
|-------|-----------|
| Mumuni | MUMUNI_LITELLM_API_KEY, MUMUNI_ZULIP_API_KEY |
| Tanko | TANKO_LITELLM_API_KEY, TANKO_ZULIP_API_KEY |
| Koby | (wrapper injects from vault — .env has Telegram token) |
| Koonimo | KOONIMO_LITELLM_API_KEY, KOONIMO_ZULIP_API_KEY |
## Key Rotation Log
| Date | Agent | Action | Notes |
|------|-------|--------|-------|
| 2026-07-17 | fleet | standardize | All 4 agents standardized on systemd drop-in + while-true wrapper + infisical v0.43.109. Removed conflicting zulip-env.conf + litellm-key.conf drop-ins. Added .env fallbacks with ZULIP keys. WAL #1322. |
| 2026-07-17 | tanko | fix-zulip | Added ZULIP_API_KEY to env (was missing — systemd bypassed vault). Updated wrapper from exec to while-true. Created .env fallback. Removed hardcoded zulip-env.conf drop-in. WAL #1321. |
| 2026-07-16 | vault | cleanup | 4 stale secrets deprecated. 5 personal creds flagged. |
| 2026-07-16 | koonimo | add-zulip | Added KOONIMO_ZULIP_API_KEY to vault. Wrapper injects ZULIP_API_KEY + ZULIP_EMAIL. 3 platforms. |
| 2026-07-16 | tanko | migrate | Migrated from hardcoded config.yaml to infisical-gateway.sh + st.353699cd. NOTE: st.353699cd later deleted — reverted to st.8e848433 on 2026-07-17. |
| 2026-07-16 | mumuni | rotate | Old key malformed (sk-_SWAl_Vu_, 47 chars, not LiteLLM format) → 401. Deleted old `mumuni` key (token 15cbca18…), generated fresh (alias `mumuni`, 7 models: syslog-auto, qwen3.6-27B-code, gemma-4-12b, strix-moe, gpu-dense, gpu-light, qwen3.6-35B-udq4). New key sk-OzuWsoX2… written to /root/.hermes/.env (Rule 3/13 fallback). Vault sync PENDING (needs machine identity). WAL #1300. |
| 2026-07-16 | koby | rotate | Old key sk-6sbCNjz (401, stale in /etc/environment). Deleted old `koby` key, generated fresh (alias `koby`). New key sk-BqRRMboTI… in systemd drop-in `hermes-gateway.service.d/litellm-key.conf` + /etc/environment. Created `hermes-gateway.service` unit (was missing — gateway wasn't persistent) with `--replace`. Verified HTTP 200, Telegram connected. |
| 2026-07-16 | baggy (koonimo) | rotate | Old key sk-krnw_zGB (401, hardcoded in systemd drop-in). Deleted old `baggy` key, generated fresh (alias `baggy`, metadata agent=koonimo). New key sk-OEK7z26n6E… in drop-in `hermes-gateway.service.d/litellm-key.conf`. CT113 IP changed .113→.114. Verified HTTP 200, Zulip connected. |