Compare commits

...
Author SHA1 Message Date
root 5eb6d3bfbd gpu-fleet: RTX 5070 swap to HauhauCS Gemma4-12B QAT Uncensored Balanced
- Model: IQ4_NL (6.3GB, 191 tok/s) → Q4_K_M QAT (6.9GB, 87 tok/s)
- MTP draft: Q8_0 (444MB) → tuned draft (242MB), saves 200MB VRAM
- mmproj: F16 → BF16 (same size, matched to new model)
- Benefits: 0/465 refusals, agent-optimized tuning, QAT quality
- Trade: 54% slower generation (acceptable for gpu-light role)
- Config: --parallel 1, --ctx-size 131072, single-slot full 128K
- VRAM: 10.0/12.2GB (82%), healthy headroom
2026-07-17 15:24:13 +00:00
jerome 33cb88d571 Merge pull request 'feat: GPU context 256K→128K fleet-wide + Genesis Hermes V3 on Strix Halo' (#20) from feat/gpu-128k-genesis-hermes-v3-20260717 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
Reviewed-on: #20
2026-07-17 11:21:32 +00:00
root 4a1f476623 fix: add missing description to inference-optimization frontmatter
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
2026-07-17 11:11:27 +00:00
root ba2c55c7e6 ci: re-trigger pipeline for PR #20
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Failing after 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Has been skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Has been skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Has been skipped
2026-07-17 11:10:29 +00:00
root b9149bce47 feat: GPU context 256K→128K fleet-wide + Genesis Hermes V3 on Strix Halo
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Failing after 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Has been skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Has been skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Has been skipped
GPU Changes:
- All 3 GPUs reduced from 256K (-c 262144) to 128K (-c 131072) for stability
- Observed instability near 100K at 256K — 128K is the stable ceiling
- VRAM improved: RTX 3090 ~70% (was 90%), RTX 5070 ~65% (was 88%)
- Strix Halo swapped to LuffyTheFox/Genesis Hermes V3 APEX
  - Hermes agent fine-tune, tensor repair (3 SSM layers, 76% W1 improvement)
  - Uncensored (0/465 refusals), multimodal (mmproj F16)
  - Speed: 65 tok/s gen, 140 tok/s prompt
  - Alias strix-moe maintained

Agent Updates:
- Mumuni: max_context_window 262144→131072, already aligned on strix-moe/0.65
- Tanko: max_context_window 262144→131072
- Koonimo: max_context_window + context_length 262144→131072
- CT114 SSH access confirmed (was 'Zulip only')

LiteLLM (CT116):
- Updated backend model references qwen3.6-35B-udq4→strix-moe
- Removed stale ornith-1.0-35b from model_cost
- Fallback chains updated

Contracts Updated:
- gpu-fleet.prose.md: topology, VRAM, benchmarks, config lines, model assignments
- gpu-self-heal.prose.md: Rule 9/10 context targets
- hermes-config-template.prose.md: template values, Rules 7-9, compression thresholds
- inference-optimization.prose.md: added to repo, 128K recommendation

Compression: 0.65 fires at ~85K (~43K headroom before 128K ceiling)
For >128K workloads: route to external providers (deepseek)
2026-07-17 10:59:45 +00:00
root 65c99dab50 contract: update litellm-api-keys to v2026-07-17 — fleet standardization
- Document canonical systemd drop-in pattern (ExecStart= reset + wrapper)
- Hardcode venv paths — never use variables in single-quoted bash -c
- Update migration status: all 4 agents on while-true wrapper + st.8e848433
- Infisical CLI update procedure (0.38.0 → 0.43.109 via artifacts-cli)
- Service token inventory + .env fallback inventory
- Key rotation log: fleet standardize + tanko fix-zulip entries
- Remove git merge conflict artifacts
- Add fleet-wide standardization lessons section
2026-07-17 10:17:23 +00:00
jerome 20cbb96e2d Merge pull request 'Vault cleanup + contract sync (WAL #1316)' (#19) from fix/vault-cleanup-contract-sync into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
Reviewed-on: #19
2026-07-16 21:24:59 +00:00
root 8215b84f88 merge: resolve conflicts with master (delegation delete + gpu-fleet)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
2026-07-16 21:22:34 +00:00
jerome 24cd7f72ac Merge branch 'master' into feat/zulip-v3-resilience 2026-07-16 21:20:24 +00:00
jerome d6e85e68e9 Merge pull request 'UPDATE 2026-07-15: Full GPU fleet rebuild + stable aliases' (#18) from feat/gpu-fleet-rebuild-20260715 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
Reviewed-on: #18
2026-07-16 21:17:12 +00:00
root b7e23e2592 vault: Infisical cleanup + contract sync (WAL #1316)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
- litellm-api-keys: vault audit, tanko/koby/koonimo migration, key rotation log
- gpu-fleet: ornith-1.0-35b→strix-moe (remaining refs)
- hermes-agent-baseline: 256K all GPUs, CT IPs updated, Shumba retired
- delegation-prose-contract: removed (renamed to mumuni-delegation)
- contract-registry.yaml + cron-prompts-review.md: from feat/contract-registry
2026-07-16 21:17:09 +00:00
9 changed files with 2702 additions and 398 deletions
File diff suppressed because it is too large Load Diff
+612
View File
@@ -0,0 +1,612 @@
# Cron Prompts Review — All 10 Scheduled Contracts
Generated: 2026-07-13 20:59:18 ET
---
## hermes-key-enforcement
**Category:** compliance | **Domain:** hermes-agent | **Owner:** abiba | **Schedule:** 0 6 * * *
```
Contract Enforcement: hermes-key-enforcement
Category: compliance
Domain: hermes-agent
Owner: abiba
Schedule: Daily compliance scan at 6am ET
This is a compliance contract. Verify that the contract enforces the required standards and policies. Report any violations found.
Steps:
1. Load contract from prose-contracts/main (file: hermes-key-enforcement.prose.md)
2. Verify prerequisites (connectivity, tools, deps)
3. Execute contract per SOP
4. Run postconditions from contract registry
5. Generate receipt with status (pass/fail/escalated)
6. If any postcondition fails, escalate per contract escalation tiers
7. Log to ~/.hermes/runs/hermes-key-enforcement/
Postconditions to verify:
[
{
"check": "no plaintext API keys in config",
"verify": "grep -rc 'api_key: sk-' /root/.hermes/config.yaml",
"expect": "0 matches"
},
{
"check": "api_key_env used for harness/litellm providers",
"verify": "grep -c 'api_key_env.*LITELLM_API_KEY' /root/.hermes/config.yaml",
"expect": "count > 0"
}
]
Escalation Tiers:
- INFO: notify nobody → action: log_to_receipt
- WARNING: notify abiba → action: relay_alert
- CRITICAL: notify abiba, mumuni → action: relay_alert + pause
- FATAL: notify abiba, mumuni, kwame → action: relay_alert + pause + human_required
Circuit Breaker:
- Max retries: 3 per 60.0min window
- On trip: escalate_to_fatal
Receipt format: JSON with contract, run_id, timestamp, agent, status, actions_taken, postconditions, drift_alerts, evidence_path
Receipt storage: ~/.hermes/runs/hermes-key-enforcement/receipt-{timestamp}.json
Graph node: Create RA-H OS node for receipt with metadata: type=receipt, contract=hermes-key-enforcement, status=<status>
If the contract has no postconditions defined (e.g., reference/pattern contracts), log that it was loaded and skip execution.
IMPORTANT: If the contract file does not exist in prose-contracts/main, report failure and do NOT hallucinate forward.
```
---
## hermes-config-template
**Category:** compliance | **Domain:** hermes-agent | **Owner:** abiba | **Schedule:** 0 4 * * 1
```
Contract Enforcement: hermes-config-template
Category: compliance
Domain: hermes-agent
Owner: abiba
Schedule: Weekly config drift check Monday at 4am ET
This is a compliance contract. Verify that the contract enforces the required standards and policies. Report any violations found.
Steps:
1. Load contract from prose-contracts/main (file: hermes-config-template.prose.md)
2. Verify prerequisites (connectivity, tools, deps)
3. Execute contract per SOP
4. Run postconditions from contract registry
5. Generate receipt with status (pass/fail/escalated)
6. If any postcondition fails, escalate per contract escalation tiers
7. Log to ~/.hermes/runs/hermes-config-template/
Postconditions to verify:
[
{
"check": "agent config template_version matches template file",
"verify": "grep -q 'template_version' /root/.hermes/config.yaml && diff <(grep 'template_version' /root/.hermes/config.yaml | cut -d: -f2 | xargs) <(grep 'template_version' /root/prose-contracts/hermes-config-template.prose.md | cut -d: -f2 | xargs) && echo match || echo mismatch",
"expect": "match"
},
{
"check": "config file is valid YAML",
"verify": "python3 -c 'import yaml; yaml.safe_load(open(\"/root/.hermes/config.yaml\"))' && echo valid || echo invalid",
"expect": "valid"
}
]
Escalation Tiers:
- INFO: notify nobody → action: log_to_receipt
- WARNING: notify abiba → action: relay_alert
- CRITICAL: notify abiba, mumuni → action: relay_alert
- FATAL: notify abiba, mumuni, kwame → action: relay_alert + human_required
Circuit Breaker:
- Max retries: 3 per 60.0min window
- On trip: escalate_to_fatal
Receipt format: JSON with contract, run_id, timestamp, agent, status, actions_taken, postconditions, drift_alerts, evidence_path
Receipt storage: ~/.hermes/runs/hermes-config-template/receipt-{timestamp}.json
Graph node: Create RA-H OS node for receipt with metadata: type=receipt, contract=hermes-config-template, status=<status>
If the contract has no postconditions defined (e.g., reference/pattern contracts), log that it was loaded and skip execution.
IMPORTANT: If the contract file does not exist in prose-contracts/main, report failure and do NOT hallucinate forward.
```
---
## hermes-agent-baseline
**Category:** compliance | **Domain:** hermes-agent | **Owner:** abiba | **Schedule:** 0 5 * * 1
```
Contract Enforcement: hermes-agent-baseline
Category: compliance
Domain: hermes-agent
Owner: abiba
Schedule: Weekly baseline verification Monday at 5am ET
This is a compliance contract. Verify that the contract enforces the required standards and policies. Report any violations found.
Steps:
1. Load contract from prose-contracts/main (file: hermes-agent-baseline.prose.md)
2. Verify prerequisites (connectivity, tools, deps)
3. Execute contract per SOP
4. Run postconditions from contract registry
5. Generate receipt with status (pass/fail/escalated)
6. If any postcondition fails, escalate per contract escalation tiers
7. Log to ~/.hermes/runs/hermes-agent-baseline/
Postconditions to verify:
[
{
"check": "Hermes agent process running",
"verify": "pgrep -f 'hermes' > /dev/null && echo running || echo stopped",
"expect": "running"
},
{
"check": "agent config file exists and valid YAML",
"verify": "test -f /root/.hermes/config.yaml && python3 -c 'import yaml; yaml.safe_load(open(\"/root/.hermes/config.yaml\"))' && echo valid || echo invalid",
"expect": "valid"
},
{
"check": "no uncommitted changes in hermes directory",
"verify": "cd /root/.hermes && git status --porcelain | wc -l",
"expect": "0"
}
]
Escalation Tiers:
- INFO: notify nobody → action: log_to_receipt
- WARNING: notify abiba → action: relay_alert
- CRITICAL: notify abiba, mumuni → action: relay_alert
- FATAL: notify abiba, mumuni, kwame → action: relay_alert + human_required
Circuit Breaker:
- Max retries: 3 per 60.0min window
- On trip: escalate_to_fatal
Receipt format: JSON with contract, run_id, timestamp, agent, status, actions_taken, postconditions, drift_alerts, evidence_path
Receipt storage: ~/.hermes/runs/hermes-agent-baseline/receipt-{timestamp}.json
Graph node: Create RA-H OS node for receipt with metadata: type=receipt, contract=hermes-agent-baseline, status=<status>
If the contract has no postconditions defined (e.g., reference/pattern contracts), log that it was loaded and skip execution.
IMPORTANT: If the contract file does not exist in prose-contracts/main, report failure and do NOT hallucinate forward.
```
---
## proxmox-monitor
**Category:** monitoring | **Domain:** proxmox | **Owner:** abiba | **Schedule:** */15 * * * *
```
Contract Enforcement: proxmox-monitor
Category: monitoring
Domain: proxmox
Owner: abiba
Schedule: Every 15 minutes
This is a monitoring contract. Execute the monitoring checks defined in the contract. Report any deviations from expected state.
Steps:
1. Load contract from prose-contracts/main (file: proxmox-monitor.prose.md)
2. Verify prerequisites (connectivity, tools, deps)
3. Execute contract per SOP
4. Run postconditions from contract registry
5. Generate receipt with status (pass/fail/escalated)
6. If any postcondition fails, escalate per contract escalation tiers
7. Log to ~/.hermes/runs/proxmox-monitor/
Postconditions to verify:
[
{
"check": "all Proxmox nodes reachable",
"verify": "curl -sf http://192.168.68.10:8006/api2/json/status | jq '.status'",
"expect": "healthy"
},
{
"check": "no VMs in crashed state",
"verify": "pvesh get /nodes -output-format=json | jq '.[] | select(.status==\"Crashed\")'",
"expect": "empty"
},
{
"check": "backups running on schedule",
"verify": "pbs-info --check",
"expect": "last_backup < 24h ago"
}
]
Escalation Tiers:
- INFO: notify nobody → action: log_to_receipt
- WARNING: notify abiba, mumuni → action: relay_alert
- CRITICAL: notify abiba, mumuni → action: relay_alert + trigger_remediation
- FATAL: notify abiba, mumuni, kwame → action: relay_alert + pause + human_required
Circuit Breaker:
- Max retries: 3 per 60.0min window
- On trip: escalate_to_fatal
Receipt format: JSON with contract, run_id, timestamp, agent, status, actions_taken, postconditions, drift_alerts, evidence_path
Receipt storage: ~/.hermes/runs/proxmox-monitor/receipt-{timestamp}.json
Graph node: Create RA-H OS node for receipt with metadata: type=receipt, contract=proxmox-monitor, status=<status>
If the contract has no postconditions defined (e.g., reference/pattern contracts), log that it was loaded and skip execution.
IMPORTANT: If the contract file does not exist in prose-contracts/main, report failure and do NOT hallucinate forward.
```
---
## gpu-monitor
**Category:** monitoring | **Domain:** gpu | **Owner:** abiba | **Schedule:** */15 * * * *
```
Contract Enforcement: gpu-monitor
Category: monitoring
Domain: gpu
Owner: abiba
Schedule: Every 15 minutes — polls all GPU subsystems
This is a monitoring contract. Execute the monitoring checks defined in the contract. Report any deviations from expected state.
Steps:
1. Load contract from prose-contracts/main (file: gpu-monitor.prose.md)
2. Verify prerequisites (connectivity, tools, deps)
3. Execute contract per SOP
4. Run postconditions from contract registry
5. Generate receipt with status (pass/fail/escalated)
6. If any postcondition fails, escalate per contract escalation tiers
7. Log to ~/.hermes/runs/gpu-monitor/
Postconditions to verify:
[
{
"check": "GPU metrics accessible",
"verify": "curl -sf http://localhost:9100/gpu-data",
"expect": "200 OK, populated data"
},
{
"check": "dashboard serving",
"verify": "curl -sf http://localhost:9100/gpu-fleet.html",
"expect": "200 OK, HTML returned"
},
{
"check": "health endpoint responsive",
"verify": "curl -sf http://localhost:9100/health",
"expect": "200 OK"
}
]
Escalation Tiers:
- INFO: notify nobody → action: log_to_receipt
- WARNING: notify abiba, mumuni → action: relay_alert
- CRITICAL: notify abiba, mumuni → action: relay_alert + trigger_remediation
- FATAL: notify abiba, mumuni, kwame → action: relay_alert + pause + human_required
Circuit Breaker:
- Max retries: 3 per 60.0min window
- On trip: escalate_to_fatal
Receipt format: JSON with contract, run_id, timestamp, agent, status, actions_taken, postconditions, drift_alerts, evidence_path
Receipt storage: ~/.hermes/runs/gpu-monitor/receipt-{timestamp}.json
Graph node: Create RA-H OS node for receipt with metadata: type=receipt, contract=gpu-monitor, status=<status>
If the contract has no postconditions defined (e.g., reference/pattern contracts), log that it was loaded and skip execution.
IMPORTANT: If the contract file does not exist in prose-contracts/main, report failure and do NOT hallucinate forward.
```
---
## infrastructure-monitoring
**Category:** monitoring | **Domain:** infrastructure | **Owner:** abiba | **Schedule:** */30 * * * *
```
Contract Enforcement: infrastructure-monitoring
Category: monitoring
Domain: infrastructure
Owner: abiba
Schedule: Every 30 minutes
This is a monitoring contract. Execute the monitoring checks defined in the contract. Report any deviations from expected state.
Steps:
1. Load contract from prose-contracts/main (file: infrastructure-monitoring.prose.md)
2. Verify prerequisites (connectivity, tools, deps)
3. Execute contract per SOP
4. Run postconditions from contract registry
5. Generate receipt with status (pass/fail/escalated)
6. If any postcondition fails, escalate per contract escalation tiers
7. Log to ~/.hermes/runs/infrastructure-monitoring/
Postconditions to verify:
[
{
"check": "Proxmox API reachable",
"verify": "curl -sf http://192.168.68.10:8006/api2/json",
"expect": "200 OK"
},
{
"check": "Zulip API reachable",
"verify": "curl -sf https://chat.sysloggh.net/api/v1/me",
"expect": "200 OK"
},
{
"check": "LiteLLM proxy reachable",
"verify": "curl -sf http://192.168.68.116/litellm/v1/models",
"expect": "200 OK"
},
{
"check": "Gitea API reachable",
"verify": "curl -sf https://git.sysloggh.net/api/v1/version",
"expect": "200 OK"
},
{
"check": "SearXNG reachable",
"verify": "curl -sf http://192.168.68.17:8080",
"expect": "200 OK"
}
]
Escalation Tiers:
- INFO: notify nobody → action: log_to_receipt
- WARNING: notify abiba, mumuni → action: relay_alert
- CRITICAL: notify abiba, mumuni → action: relay_alert + trigger_remediation
- FATAL: notify abiba, mumuni, kwame → action: relay_alert + pause + human_required
Circuit Breaker:
- Max retries: 3 per 60.0min window
- On trip: escalate_to_fatal
Receipt format: JSON with contract, run_id, timestamp, agent, status, actions_taken, postconditions, drift_alerts, evidence_path
Receipt storage: ~/.hermes/runs/infrastructure-monitoring/receipt-{timestamp}.json
Graph node: Create RA-H OS node for receipt with metadata: type=receipt, contract=infrastructure-monitoring, status=<status>
If the contract has no postconditions defined (e.g., reference/pattern contracts), log that it was loaded and skip execution.
IMPORTANT: If the contract file does not exist in prose-contracts/main, report failure and do NOT hallucinate forward.
```
---
## zulip-health
**Category:** monitoring | **Domain:** zulip | **Owner:** abiba | **Schedule:** */15 * * * *
```
Contract Enforcement: zulip-health
Category: monitoring
Domain: zulip
Owner: abiba
Schedule: Every 15 minutes — monitors all Zulip-connected agents
This is a monitoring contract. Execute the monitoring checks defined in the contract. Report any deviations from expected state.
Steps:
1. Load contract from prose-contracts/main (file: zulip-health.prose.md)
2. Verify prerequisites (connectivity, tools, deps)
3. Execute contract per SOP
4. Run postconditions from contract registry
5. Generate receipt with status (pass/fail/escalated)
6. If any postcondition fails, escalate per contract escalation tiers
7. Log to ~/.hermes/runs/zulip-health/
Postconditions to verify:
[
{
"check": "bot registration active",
"verify": "curl -sf https://chat.sysloggh.net/api/v1/me | jq '.user_id'",
"expect": "bot_id present"
},
{
"check": "DM delivery working",
"verify": "curl -sf https://chat.sysloggh.net/api/v1/users/me/is-online",
"expect": "online: true"
}
]
Escalation Tiers:
- INFO: notify nobody → action: log_to_receipt
- WARNING: notify abiba, mumuni → action: relay_alert
- CRITICAL: notify abiba, mumuni → action: relay_alert + trigger_remediation
- FATAL: notify abiba, mumuni, kwame → action: relay_alert + pause + human_required
Circuit Breaker:
- Max retries: 3 per 60.0min window
- On trip: escalate_to_fatal
Receipt format: JSON with contract, run_id, timestamp, agent, status, actions_taken, postconditions, drift_alerts, evidence_path
Receipt storage: ~/.hermes/runs/zulip-health/receipt-{timestamp}.json
Graph node: Create RA-H OS node for receipt with metadata: type=receipt, contract=zulip-health, status=<status>
If the contract has no postconditions defined (e.g., reference/pattern contracts), log that it was loaded and skip execution.
IMPORTANT: If the contract file does not exist in prose-contracts/main, report failure and do NOT hallucinate forward.
```
---
## litellm-health
**Category:** monitoring | **Domain:** litellm | **Owner:** abiba | **Schedule:** */10 * * * *
```
Contract Enforcement: litellm-health
Category: monitoring
Domain: litellm
Owner: abiba
Schedule: Every 10 minutes — LiteLLM proxy health
This is a monitoring contract. Execute the monitoring checks defined in the contract. Report any deviations from expected state.
Steps:
1. Load contract from prose-contracts/main (file: litellm-health.prose.md)
2. Verify prerequisites (connectivity, tools, deps)
3. Execute contract per SOP
4. Run postconditions from contract registry
5. Generate receipt with status (pass/fail/escalated)
6. If any postcondition fails, escalate per contract escalation tiers
7. Log to ~/.hermes/runs/litellm-health/
Postconditions to verify:
[
{
"check": "LiteLLM proxy reachable",
"verify": "curl -sf http://192.168.68.116/litellm/v1/models",
"expect": "200 OK, models returned"
},
{
"check": "router deprecated, nginx routes work",
"verify": "curl -sf https://litellm.sysloggh.net/v1/models",
"expect": "200 OK (via nginx)"
}
]
Escalation Tiers:
- INFO: notify nobody → action: log_to_receipt
- WARNING: notify abiba, mumuni → action: relay_alert
- CRITICAL: notify abiba, mumuni → action: relay_alert + trigger_remediation
- FATAL: notify abiba, mumuni, kwame → action: relay_alert + pause + human_required
Circuit Breaker:
- Max retries: 3 per 60.0min window
- On trip: escalate_to_fatal
Receipt format: JSON with contract, run_id, timestamp, agent, status, actions_taken, postconditions, drift_alerts, evidence_path
Receipt storage: ~/.hermes/runs/litellm-health/receipt-{timestamp}.json
Graph node: Create RA-H OS node for receipt with metadata: type=receipt, contract=litellm-health, status=<status>
If the contract has no postconditions defined (e.g., reference/pattern contracts), log that it was loaded and skip execution.
IMPORTANT: If the contract file does not exist in prose-contracts/main, report failure and do NOT hallucinate forward.
```
---
## memory-audit-maintenance
**Category:** maintenance | **Domain:** memory | **Owner:** mumuni | **Schedule:** 0 3 * * *
```
Contract Enforcement: memory-audit-maintenance
Category: maintenance
Domain: memory
Owner: mumuni
Schedule: Daily at 3am ET
This is a maintenance contract. Execute the maintenance tasks defined in the contract. Report any issues found.
Steps:
1. Load contract from prose-contracts/main (file: memory-audit-maintenance.prose.md)
2. Verify prerequisites (connectivity, tools, deps)
3. Execute contract per SOP
4. Run postconditions from contract registry
5. Generate receipt with status (pass/fail/escalated)
6. If any postcondition fails, escalate per contract escalation tiers
7. Log to ~/.hermes/runs/memory-audit-maintenance/
Postconditions to verify:
[
{
"check": "memory files below 80% capacity",
"verify": "wc -l ~/.hermes/memories/*.md",
"expect": "total lines < threshold"
},
{
"check": "no stale entries",
"verify": "grep -r 'STALE' ~/.hermes/memories/",
"expect": "0 matches"
}
]
Escalation Tiers:
- INFO: notify nobody → action: log_to_receipt
- WARNING: notify mumuni → action: relay_alert
- CRITICAL: notify mumuni, abiba → action: relay_alert
- FATAL: notify mumuni, abiba, kwame → action: relay_alert + human_required
Circuit Breaker:
- Max retries: 3 per 60.0min window
- On trip: escalate_to_fatal
Receipt format: JSON with contract, run_id, timestamp, agent, status, actions_taken, postconditions, drift_alerts, evidence_path
Receipt storage: ~/.hermes/runs/memory-audit-maintenance/receipt-{timestamp}.json
Graph node: Create RA-H OS node for receipt with metadata: type=receipt, contract=memory-audit-maintenance, status=<status>
If the contract has no postconditions defined (e.g., reference/pattern contracts), log that it was loaded and skip execution.
IMPORTANT: If the contract file does not exist in prose-contracts/main, report failure and do NOT hallucinate forward.
```
---
## infrastructure-update
**Category:** maintenance | **Domain:** infrastructure | **Owner:** abiba | **Schedule:** 0 2 * * 0
```
Contract Enforcement: infrastructure-update
Category: maintenance
Domain: infrastructure
Owner: abiba
Schedule: Weekly system updates Sunday at 2am ET
This is a maintenance contract. Execute the maintenance tasks defined in the contract. Report any issues found.
Steps:
1. Load contract from prose-contracts/main (file: infrastructure-update.prose.md)
2. Verify prerequisites (connectivity, tools, deps)
3. Execute contract per SOP
4. Run postconditions from contract registry
5. Generate receipt with status (pass/fail/escalated)
6. If any postcondition fails, escalate per contract escalation tiers
7. Log to ~/.hermes/runs/infrastructure-update/
Postconditions to verify:
[
{
"check": "all services running after update",
"verify": "systemctl list-units --state=running",
"expect": "all critical services"
}
]
Escalation Tiers:
- INFO: notify nobody → action: log_to_receipt
- WARNING: notify abiba → action: relay_alert
- CRITICAL: notify abiba, mumuni → action: relay_alert
- FATAL: notify abiba, mumuni, kwame → action: relay_alert + human_required
Circuit Breaker:
- Max retries: 1 per 120.0min window
- On trip: escalate_to_fatal
Receipt format: JSON with contract, run_id, timestamp, agent, status, actions_taken, postconditions, drift_alerts, evidence_path
Receipt storage: ~/.hermes/runs/infrastructure-update/receipt-{timestamp}.json
Graph node: Create RA-H OS node for receipt with metadata: type=receipt, contract=infrastructure-update, status=<status>
If the contract has no postconditions defined (e.g., reference/pattern contracts), log that it was loaded and skip execution.
IMPORTANT: If the contract file does not exist in prose-contracts/main, report failure and do NOT hallucinate forward.
```
---
*End of review*
-256
View File
@@ -1,256 +0,0 @@
---
kind: pattern
name: delegation-prose-contract
description: >
Manager (Mumuni) operating doctrine for task decomposition, worker
delegation, verification, and delivery. Defines when to delegate, which
worker to use for what, how to handle failures, and the kanban board
protocol. Enforces context-window discipline and separation of concerns.
Runs on Mumuni (CT 118, storepve, .6) via Hermes agent.
version: 1.0.0
---
## Maintains
- Worker roster: 6 profiles (`syslog-code`, `syslog-devops`, `syslog-email`,
`syslog-research`, `syslog-review`, `syslog-writer`)
- Kanban board state at `~/.hermes/kanban/kanban.json`
- Context window budget: ~65K tokens per request (131K total, 60% threshold)
## Topology
**Cluster:** 5 Proxmox nodes (ocupve, acerpve, minipve, amdpve, storepve)
**Manager:** Mumuni (CT 118, storepve, .6) via Hermes agent
**Workers:** 6 profiles, all running on the same agent — no separate hosts needed
This contract is infrastructure-agnostic in terms of which nodes are used.
Workers execute tasks on whatever infrastructure they're given — SSH to .6,
.pm, .9, .12, or .15 depending on the task. The contract defines the
**who** and **when** — not the **where**.
## Why This Matters
Without enforced delegation, the manager consumes the full iteration budget
(60 calls) on single-turn tasks — SSH to 5 nodes, check each VM, read logs —
leaving no capacity for actual coordination. The result: context overflow
(59K tokens in system prompt), iteration exhaustion, and degraded response
quality. This contract exists because I blew through my budget checking
Proxmox node status instead of delegating to `syslog-devops`.
## Context Window Discipline
**The system prompt is ~6.5K tokens (stable: ~4.5K tool schemas + ~2K other guidance).**
**Volatile (MEMORY.md + USER.md): ~300 tokens.**
**Total base: ~6,800 tokens per request.**
The remaining budget is the conversation. Every tool call result adds to it.
If a single call returns >10K tokens (e.g., `grep` on a large file, SSH output
from multiple nodes), the context fills fast. That's why we delegate: workers
process in isolation and return compact results.
## Trigger Conditions
Delegation is **mandatory** when any of these apply:
| Condition | Threshold | Example |
|-----------|-----------|---------|
| Multiple tool calls needed | 2+ calls with intermediate logic | Read file → analyze → write report |
| Large data retrieval | Output >5K tokens | `grep -r "pattern" /path` on large dirs |
| Cross-domain work | Spans 2+ worker specialties | Infra check + email filter |
| Infrastructure changes | Any mutating operation | `qm set`, `systemctl restart`, `git push` |
| Research/analysis | Needs browser or deep reading | Web research, code review, data analysis |
| Code builds or changes | Writing or modifying code | Scripts, configs, patches |
| Sequential dependencies | Worker B needs Worker A's output | Code → Review → Deliver |
**Single tool calls stay at manager level.** Quick `grep`, `ls`, `cat`,
`curl`, `hermes tools list` — these are decision-making tools. The manager
reads them directly.
## Worker Selection Matrix
| Worker | Model | Toolsets | Role | Use When |
|--------|-------|----------|------|----------|
| `syslog-code` | qwen3.6-27B-code | terminal, file, web, memory, skills | Code patches, automation, scripts | Writing/modifying code, creating scripts, debugging, reading/writing files |
| `syslog-devops` | qwen3.6-27B-code | terminal, file, web, memory, skills | Infrastructure, DB, bridge, Proxmox | Server ops, SSH, Docker, Proxmox, DB queries, hardware checks |
| `syslog-email` | strix-moe | terminal, file, web, memory, skills | Email automation, mail operations | Sending/receiving email, inbox management, SMTP operations |
| `syslog-research` | strix-moe | terminal, file, web, memory, skills, **browser** | Analysis, classification, data processing | Web research, browser tasks, data analysis, classification, reading docs |
| `syslog-review` | strix-moe | terminal, file, web, memory, skills | Verification, QA, audit validation | **ALWAYS** verify worker output before delivery — especially for infra changes, code builds, and research findings |
| `syslog-writer` | strix-moe | terminal, file, web, memory, skills | Docs, content, branding, reports | Writing docs, reports, proposals, content, markdown formatting |
### Selection Rules
1. **Match specialty first.** A code task → `syslog-code`. An infra task →
`syslog-devops`. Don't put a `syslog-email` worker on a code review.
2. **Research tasks with browser needs → `syslog-research`.** Other workers
don't have the browser toolset.
3. **Verification → `syslog-review`.** Never deliver raw worker output.
4. **Documentation/content → `syslog-writer`.** Let them own the prose.
5. **If unsure, delegate to `syslog-research`** — it has the broadest toolset
(includes browser) and high reasoning effort.
## Delegation Protocol
### Step 1: Decompose
Break the task into lanes. Each lane does ONE thing. Workers are independent —
no lane depends on another's output mid-flight. If lanes depend on each other,
dispatch sequentially.
### Step 2: Dispatch
Fire workers via `delegate_task`:
**Parallel (independent lanes):**
```
delegate_task(
tasks=[
{"goal": "Check all 5 Proxmox nodes for VM status", "context": "SSH to each node via 192.168.68.x, run 'qm list'"},
{"goal": "Check Docker container health on .7/.116/.17", "context": "SSH to each host, check container status"},
]
)
```
**Sequential (dependent lanes):**
Dispatch lane 1 → wait for result → dispatch lane 2.
### Step 3: Verify
**MANDATORY for:**
- Infrastructure changes (any `qm`, `pct`, `systemctl`, `git push`)
- Code builds and modifications
- Research findings (web data, external sources)
- Any output that will reach the user
**Fire `syslog-review` to verify:**
```
delegate_task(
goal="Review the output of the devops worker. Verify the node status
report is accurate, check for inconsistencies, confirm all nodes were
reachable.",
context="Worker was syslog-devops. Output is at /tmp/node-report.md.
Verify against live system."
)
```
**If verification fails:**
1. Send work back to original worker with review feedback
2. Re-verify
3. Max 2 re-verify cycles before escalating to Kwame
### Step 4: Deliver
Only verified results reach Kwame. Format per channel:
- Telegram: Use `telegram-formatting` skill
- Zulip: Use Zulip Markdown (CommonMark)
- Email: Use `syslog-email` skill
## Kanban Board Protocol
**File:** `~/.hermes/kanban/kanban.json`
```json
{
"task_id": "unique-id",
"title": "Task description",
"created": "2026-07-09T01:00:00",
"status": "backlog|in_progress|review|done",
"lanes": [
{
"lane_id": "devops-check",
"worker": "syslog-devops",
"goal": "Check all 5 Proxmox nodes",
"status": "dispatched|completed|failed",
"output_file": "/tmp/node-report.md"
}
]
}
```
**Update the board on every state change.**
## Failure Handling
### Worker Timeouts
- Child timeout: **900 seconds** (15 minutes)
- Worker model `syslog-auto` is slow — it can hit the timeout limit with
22+ API calls
- **If a worker times out:** Re-dispatch with a narrower scope. Break the
task into smaller pieces that fit in the timeout window.
- **Avoid delegating sequential SSH hops** — each SSH connection adds latency
that compounds quickly. Prefer API-based or local approaches when possible.
### Worker Selection Failures
- `syslog-devops` is best for infrastructure tasks (SSH, Proxmox, Docker)
- `syslog-code` is best for code-level work (reading files, writing scripts)
- `syslog-research` has the browser toolset — use for web research
- `syslog-review` is the QA gate — always fire before delivery
- **Never fire more than 3 parallel workers** (max_concurrent_children: 3)
- **Never nest delegation** (max_spawn_depth: 1)
### Context Overflow
- If a task requires >10K tokens of output, delegate the processing
- Workers return compact summaries, not raw data dumps
- Pass file paths and concrete goals — never dump raw data into context
## Anti-patterns
- ❌ Reading large files into your own context before deciding → delegate the read
- ❌ Carrying SSH/grep/output results in your context → delegate the analysis
- ❌ Doing work yourself and then "pretending" to delegate → the user can tell
- ❌ Skipping verification → raw worker output never reaches the user
- ❌ Delegating single tool calls → keep quick reads/writes at manager level
- ❌ Firing more than 3 workers in parallel → hard limit
## Emergency Exception
**In an emergency (server down, service must be restored immediately):**
- Delegate the diagnosis (find the problem)
- Execute the fix yourself (minimize handoff latency)
- Verify the fix after delivery
- Log the exception in the kanban board
The emergency exception exists because the user needs the service back NOW,
not after three worker round-trips. But it's an exception — not the rule.
## What This Contract Doesn't Cover
1. **Worker profile configuration** — covered by `hermes-config-template.prose.md`
2. **SSH key management** — covered by existing SSH/Proxmox contracts
3. **Git workflow** — covered by `AGENTS.md` in the prose-contracts repo
4. **Cron job management** — covered by individual cron contracts
5. **Infra verification** — covered by `verify-before-mutate` protocol
## Verification
Run `scripts/worker-audit.py` to verify all 6 profiles are aligned:
```bash
python3 /root/.hermes/skills/kanban-orchestrator/scripts/worker-audit.py
```
## References
- `kanban-orchestrator` skill: The operational playbook (detailed execution steps)
- `worker-profile-audit.md` (skill reference): Worker configuration audit notes
- `delegation-timeout-patterns.md` (skill reference): Timeout handling patterns
- `verify-before-mutate` protocol: Infrastructure change verification
- `hermes-config-template.prose.md`: Worker profile configuration
## Success Criteria
This contract succeeds when:
1. **No context overflow** — single-turn tasks don't exhaust the iteration budget
2. **Workers do the work** — manager coordinates, doesn't execute
3. **Verification before delivery** — all output passes through `syslog-review`
4. **Kanban board is current** — every task has a lane, every lane has a status
5. **User gets verified results** — raw worker output never reaches Kwame
---
**Last updated:** 2026-07-09
**Author:** Mumuni (with Kwame's input on triggers and exception criteria)
**Status:** Draft — awaiting PR review and merge to prose-contracts main
+39 -29
View File
@@ -7,10 +7,14 @@ description: >
saturation watchdog, Prometheus/Grafana monitoring, and self-healing.
UPDATED 2026-07-15: Stable role-based aliases introduced: strix-moe,
gpu-dense, gpu-light. These never change — only the underlying model does.
Strix Halo: ornith-1.0-35b → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB).
Strix Halo: strix-moe → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB).
RTX 5070: gemma-4-12b Q4_K_M → IQ4_NL + MTP draft (122 tok/s, 2x faster).
RTX 5070 context: 131K → 256K. VRAM: 88% (10.8/12.2GB).
Compression timeout: 300s (was 120s). Mumuni context: 128K (was 256K).
UPDATED 2026-07-17: Strix Halo model swapped to Genesis Hermes V3 APEX (LuffyTheFox, 24GB, uncensored,
Hermes agent fine-tune, tensor repair, multimodal with mmproj).
RTX 5070 swapped to HauhauCS Gemma4-12B QAT Uncensored Balanced (Q4_K_M, 87 tok/s, 0/465 refusals).
Instability observed near 100K at 256K. 128K is the stable ceiling.
For larger context needs → fall back to external providers (deepseek).
VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%.
agent: abiba
triggers:
- on model add/remove
@@ -67,7 +71,7 @@ triggers:
│ CT 8 │ │ CT 110 │ │ CT 15 │ │ pi (.24) │
│ RTX 3090 │ │ RTX 5070 │ │ Strix Halo│ │ GPU Monitor │
│ 24GB │ │ 12GB │ │ 64GB UMA │ │ :9100 │
256K ctx │ │ 256K ctx │ │ 256K ctx │ │ Watchdog │
128K ctx │ │ 128K ctx │ │ 128K ctx │ │ Watchdog │
│ qwen3.6 │ │ gemma-4-12b │ │ qwen3.6 │ │ Prometheus │
│ 27B-code │ │ :8080 │ │ -35B-udq4 │ │ exporter │
│ :8080 │ │ :9400 (exp) │ │ :9400(exp)│ │ :9401 │
@@ -89,13 +93,19 @@ When a model is swapped on a GPU, ONLY the infrastructure layer changes — agen
**Backward compatibility**: Old model-specific names (qwen3.6-27B-code, gemma-4-12b, qwen3.6-35B-udq4) still work
but are deprecated for agent configs. Only the stable aliases survive model swaps.
## Current Model Assignments (2026-07-15)
## Current Model Assignments (2026-07-17)
| Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status |
|-------|-----|------|------|-----|----------|----------|-------------|--------|
| qwen3.6-27B-code (MTP) | RTX 3090 | .8 (llm-gpu) | 22.2/24.6GB (90%) | **256K** 🚀 | turbo4 | 2 | default | ✅ 63 tok/s |
| gemma-4-12b | RTX 5070 | .110 (ocu-llm) | 10.0/12.2GB (82%) | 256K | q4_0 | 2 | 2048/1024 | ✅ healthy |
| qwen3.6-35B-udq4 | Strix Halo Vulkan | .15 (amdpve) | ~9GB/64GB | 256K | q8_0 | 2 | 2048/512 | ✅ healthy |
| qwen3.6-27B-code (MTP) | RTX 3090 | .8 (llm-gpu) | ~17/24.6GB (70%) | **128K** | turbo4 | 2 | default | ✅ 63 tok/s |
| gemma-4-12b (HauhauCS QAT) | RTX 5070 | .110 (ocu-llm) | ~10.0/12.2GB (82%) | 128K | q4_0 | 1 | 2048/1024 | ✅ 87 tok/s |
| Genesis Hermes V3 APEX | Strix Halo Vulkan | .15 (amdpve) | ~10GB/64GB | 128K | q4_0 | 1 | 4096/1024 | ✅ 65 tok/s |
> **RTX 5070 model swap (2026-07-17)**: Switched from `gemma-4-12b-it-IQ4_NL` (Unsloth, 191 tok/s)
> to `HauhauCS/Gemma4-12B-QAT-Uncensored-HauhauCS-Balanced` (Q4_K_M QAT, 87 tok/s).
> Trade: 54% slower generation for QAT quality, 0/465 refusals, and agent-optimized tuning.
> MTP draft also swapped: Q8_0 (444MB) → tuned draft (242MB), saving 200MB VRAM.
> Role unchanged: gpu-light (vision, web extract, light auxiliary tasks).
## Routing Configuration (LiteLLM — July 2026)
@@ -104,7 +114,7 @@ but are deprecated for agent configs. Only the stable aliases survive model swap
| Model | GPU | Weight | RPM Cap | Timeout |
|-------|-----|--------|---------|---------|
| qwen3.6-27B-code | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** |
| qwen3.6-35B-udq4 | Strix Halo (.15:8080) | **0.30** | 60 | **300s** |
| Genesis Hermes V3 APEX | Strix Halo (.15:8080) | **0.30** | 60 | **300s** |
| gemma-4-12b | RTX 5070 (.110:8080) | **0.15** | 200 | **120s** |
Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) is NOT in the inference path.
@@ -113,9 +123,9 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`.
| Model | RPM Cap | Notes |
|-------|---------|-------|
| qwen3.6-35B-udq4 | 40 | Tight cap — prevents Strix overload |
| strix-moe (Hermes V3) | 40 | Tight cap — prevents Strix overload |
| qwen3.6-27B-code | 500 | High cap — primary workhorse |
| gemma-4-12b | 500 | High cap — IQ4_NL+MTP, 122 tok/s |
| gemma-4-12b | 500 | HauhauCS QAT Uncensored Balanced + MTP, 87 tok/s |
### Stable Aliases (for agent configs — never change)
@@ -123,16 +133,16 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`.
|-------|---------|-----------|---------|
| `strix-moe` | 40 | Strix Halo | Compression tasks (MoE models) |
| `gpu-dense` | 500 | RTX 3090 | Heavy reasoning |
| `gpu-light` | 500 | RTX 5070 | Vision, web extract, light tasks |
| `gpu-light` | 500 | RTX 5070 (HauhauCS QAT) | Vision, web extract, light tasks |
### Fallback Chains
- gemma → qwen
- qwen → gemma
- qwen3.6-35B-udq4 → qwen → gemma
- strix-moe → qwen → gemma
- syslog-auto → qwen → gemma → qwen3.6-35B-udq4
### Why Strix Halo RPM Is Capped
- Direct (qwen3.6-35B-udq4): 40 RPM (tight) — Strix Halo is shared with compression tasks
- Direct (strix-moe): 40 RPM (tight) — Strix Halo is shared with compression tasks
- Via syslog-auto: 60 RPM (moderate) — prevents flooding when multiple agents use syslog-auto simultaneously
- Combined max: ~100 RPM across both paths — Strix Halo can sustain this at 80°C
@@ -186,7 +196,7 @@ Show full fleet status: GPUs, models, VRAM, context windows, parallel slots, act
3. Check LiteLLM: `curl http://192.168.68.116/health` (expect "I'm alive!")
4. Check LiteLLM models: `curl -H "Authorization: Bearer $MASTER_KEY" http://192.168.68.116/v1/models`
5. Check LiteLLM timeouts: `grep -n 'timeout:' /opt/inference-harness/litellm_config.yaml`
- gemma-4-12b: 120s, qwen3.6-27B-code: 300s, qwen3.6-35B-udq4/strix-moe: 300s (ornith-1.0-35b does NOT exist — legacy name, do not use)
- gemma-4-12b: 120s, qwen3.6-27B-code: 300s, qwen3.6-35B-udq4/strix-moe: 300s (strix-moe does NOT exist — legacy name, do not use)
- global request_timeout: 300s, nginx proxy_read_timeout: 600s
6. Check AMD metrics: `curl http://192.168.68.15:9400/metrics` (Radeon 8060S, util%, VRAM, temp, power)
7. Check port conflicts: verify only one llama-server on :8080 per host
@@ -243,12 +253,12 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
- **Router startup race**: Compose router.py doesn't call load_roster(). Reload thread sleeps 30s first.
Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster().
- **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround.
- **VRAM (2026-07-15)**: RTX 3090 at 22.2/24.6GB (90%) with **256K context** (corrected from 131K). RTX 5070 at 10.8/12.2GB (88%) with 256K context + MTP. Strix Halo at ~9GB/64GB.
- **VRAM (2026-07-15)**: RTX 3090 at ~17/24.6GB (~70%) with **128K context** (reduced from 256K 2026-07-17). RTX 5070 at ~7.8/12.2GB (~65%) with 128K context + MTP. Strix Halo at ~7GB/64GB.
- **RTX 3090 runs `--parallel 2`** with MTP draft (spec-type draft-mtp, spec-draft-n-max 2).
- **RTX 3090 config**: `-c 262144 -ctk turbo4 -ctv turbo4 --parallel 2 --flash-attn on --cont-batching --spec-type draft-mtp`. Context corrected to 256K (2026-07-15). VRAM: 90%. Service: `/home/llmuser/llama-wrapper.sh`.
- **RTX 5070 config (2026-07-15)**: Switched to IQ4_NL + MTP draft (Q8_0) at 256K context. Gen speed: 122 tok/s (was 70). VRAM: 10.8/12.2GB (88%). No draft model pre-upgrade due to VRAM constraints. Service: `/home/llmuser/llama-wrapper.sh`. Config: `--model gemma-4-12b-it-IQ4_NL.gguf --spec-draft-model gemma-4-12b-it-Q8_0-MTP.gguf --spec-type draft-mtp --spec-draft-n-max 4 --ctx-size 262144`.
- **RTX 3090 config**: `-c 131072 -ctk turbo4 -ctv turbo4 --parallel 1 --flash-attn on --cont-batching --spec-type draft-mtp`. Context 128K, single slot for full 128K per-request. VRAM: ~85%. Service: `/home/llmuser/llama-wrapper.sh`.
- **RTX 5070 config (2026-07-17)**: HauhauCS Gemma4-12B QAT Uncensored Balanced (Q4_K_M) + tuned MTP draft (242MB) at 128K context, single slot. Gen speed: 87 tok/s (vs 191 IQ4_NL). VRAM: ~10.0/12.2GB (~82%). Service: `/home/llmuser/llama-wrapper.sh`. Model: `Gemma4-12B-QAT-Uncensored-HauhauCS-Balanced-Q4_K_M.gguf`, MTP: `mtp-gemma-4-12B-it.gguf`, mmproj: `mmproj-Gemma4-12B-QAT-Uncensored-HauhauCS-Balanced-BF16.gguf`. Recommended sampling: temp 0.6, top_k 64, top_p 0.9, min_p 0.05, repeat_penalty 1.1.
- **LiteLLM timeout tuning (verified 2026-07-16 against `/opt/inference-harness/litellm_config.yaml` on CT 116)**: gemma-4-12b 120s, qwen3.6-27B-code 300s, qwen3.6-35B-udq4 300s, strix-moe 300s, syslog-auto routes all 300s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s.
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080 (was `ornith-server.service`), model changed to `unsloth/Qwen3.6-35B-A3B-MTP-GGUF` (UD-Q4_K_M), alias `qwen3.6-35B-udq4`, 256K context, flash-attn + q8 KV. MTP support enabled for 1.4-2.2x faster inference.
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V3-GGUF` (APEX quant), alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded). Hermes agent fine-tune, tensor repair (SSM layers fixed via SVD), uncensored (0/465 refusals).
- **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill).
- **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (Mumuni) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`.
- **Port 8080 firewall**: amdpve iptables restricts 8080 to 192.168.68.116 (LiteLLM/router host) only. All inbound connections are from .116 (LiteLLM proxied via nginx). Localhost curls hang (SYN dropped). Always test from .116.
@@ -262,12 +272,12 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
| GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Context |
|-----|-------|-----------|--------------|----------|---------|
| RTX 3090 (.8) | qwen3.6-27B-code (MTP) | **63** | — | — | **256K** |
| RTX 5070 (.110) | gemma-4-12b (IQ4_NL+MTP) | **191** | — | — | **256K** |
| Strix Halo (.15) | qwen3.6-35B-udq4 | **71** | | — | **256K** |
| RTX 3090 (.8) | qwen3.6-27B-code (MTP) | **63** | — | — | **128K** |
| RTX 5070 (.110) | HauhauCS QAT Uncensored Balanced | **87** | — | — | **128K** |
| Strix Halo (.15) | Genesis Hermes V3 APEX | **65** | 140 | — | **128K** |
Benchmarks from 2026-07-15 verification run. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
All 3 GPUs now at 256K context (2026-07-15).
Benchmarks from 2026-07-17. Strix Halo swapped to Genesis Hermes V3 APEX (LuffyTheFox). RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
All 3 GPUs now at 128K context (2026-07-17, reduced from 256K for stability).
Benchmarks run through LiteLLM proxy (192.168.68.116:4001) every 5 minutes.
Degradation alerts fire at 30% (warning) and 50% (critical) below baseline.
@@ -288,9 +298,9 @@ All agent configs MUST use stable role-based aliases, never model-specific names
When the underlying model is swapped, only the LiteLLM config changes — agent configs are untouched.
### Context Windows
- RTX 3090: **256K** (was 131K, bumped 2026-07-15) | RTX 5070: **256K** (up from 131K) | Strix Halo: **256K**
- **Mumuni compression context**: 128K (down from 256K) — ensures compression model doesn't timeout
- Compression threshold 0.65: fires at ~85K for 128K context window
- RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **128K**
- **All agents**: 128K ceiling — stable margin. For >128K workloads, use external providers (deepseek)
- Compression threshold 0.65: fires at ~85K (~43K headroom before 128K ceiling)
- Mumuni compression model alias: `strix-moe` with 300s timeout
### Mumuni Agent Profile
@@ -306,7 +316,7 @@ Mumuni (CT114, 192.168.68.123) is the primary business assistant. This profile i
| `aux.vision.model` | `gpu-light` | Vision tasks (RTX 5070) |
| `aux.web_extract.model` | `gpu-light` | Web extraction |
| `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) |
| `context.max_context_window` | 262144 (256K) | Fixed 2026-07-16 (was 131072 — caused premature compression, WAL #1300) |
| `context.max_context_window` | 131072 (128K) | Reduced from 256K 2026-07-17 — stable 128K ceiling |
| `compression.threshold` | 0.65 | Triggers at ~85K |
| `compression.target_ratio` | 0.3 | Compresses to ~38K |
| `compression.protect_last_n` | 40 | Preserves last 40 messages |
@@ -314,7 +324,7 @@ Mumuni (CT114, 192.168.68.123) is the primary business assistant. This profile i
| `personalities` | `creative` | Creative assistant personality |
| Platforms | cli, discord, homeassistant, signal, telegram, zulip | All Hermes platforms |
| Main model timeout | 300s | LiteLLM global timeout |
| Compression model timeout | 300s | ornith timeout increased from 120s |
| Compression model timeout | 300s | strix-moe timeout increased from 120s |
### Agent Update Status (2026-07-15)
+6 -6
View File
@@ -118,9 +118,9 @@ depends_on:
### Rule 9: Context Window Optimization
- **Detect**: Benchmark tok/s vs baseline for each GPU at current context
- RTX 3090 (256K ctx, qwen3.6-27B-code): target 75+ tok/s — currently at baseline
- RTX 5070 (131K ctx, gemma-4-12b): target 76+ tok/s — optimal for vision/web role
- Strix Halo (256K ctx, strix-moe / qwen3.6-35B-udq4): target 70+ tok/s — currently above baseline
- RTX 3090 (128K ctx, qwen3.6-27B-code): target 75+ tok/s — currently at baseline
- RTX 5070 (128K ctx, gemma-4-12b): target 76+ tok/s — optimal for vision/web role
- Strix Halo (128K ctx, strix-moe / qwen3.6-35B-udq4): target 70+ tok/s — currently above baseline
- **Fix**:
- If tok/s > baseline → context has headroom, consider increasing
- If tok/s < 90% baseline → reduce context by 25% and retest
@@ -131,9 +131,9 @@ depends_on:
### Rule 10: Workload Distribution Optimization
- **Detect**: GPU roles misaligned with hardware capabilities
- **Target distribution**:
- RTX 3090 (24GB, 256K, 75 tok/s) → Heavy reasoning, code gen, long conversations
- RTX 5070 (12GB, 131K, 76 tok/s) → Vision/image, web search, quick lightweight tasks
- Strix Halo (64GB, 256K, 72 tok/s) → Context compression, summarization, long docs
- RTX 3090 (24GB, 128K, 75 tok/s) → Heavy reasoning, code gen, long conversations
- RTX 5070 (12GB, 128K, 76 tok/s) → Vision/image, web search, quick lightweight tasks
- Strix Halo (64GB, 128K, 72 tok/s) → Context compression, summarization, long docs
- **Fix**:
- Alert if any GPU is handling workload outside its designated role
- Recommend Hermes agent profile updates to match workload to GPU
+7 -7
View File
@@ -5,7 +5,7 @@ version: 1.0.0
description: >
Canonical known-good baseline for all Syslog Hermes agents. Captures the exact
configuration state, keys, workarounds, and audit procedure. When an agent's
configuration goes sideways, restore from this baseline. Last verified 2026-07-08. GPU context reduced to 128K on .8/.110, parallel 2 fleet-wide.
configuration goes sideways, restore from this baseline. Last verified 2026-07-16. All GPUs 256K context (RTX 3090 .8, RTX 5070 .110, Strix Halo .15). Parallel 1 fleet-wide (Strix Halo handles compression solo).
author: Abiba (pi agent)
---
@@ -26,11 +26,11 @@ done
|-------|-----|------|-----|---------------|------------|----------|
| Tanko | 112 | amdpve | .122 | `tanko` | Infisical vault | Hermes |
| Mumuni | 114 | minipve | .123 | `mumuni` | Infisical vault | Hermes |
| Koby | 111 | amdpve | srv1079750 | `koby` | Infisical vault | **Hermes** |
| Koonimo | 113 | amdpve | ? | `koonimo` | Infisical vault | Hermes |
| Shumba | — | 192.168.68.119 | N/A | N/A (DeepSeek) | Hermes |
| Koby | 129 | amdpve | srv1079750 | `koby` | Infisical vault | **Hermes** |
| Koonimo | 114 | amdpve | ? | `koonimo` | Infisical vault | Hermes |
| Shumba | — | 192.168.68.119 | N/A | N/A (DeepSeek) | Hermes (RETIRED — CT119 now Infisical vault) |
> **Note**: CT hostnames (tdunna→CT111, baggy→CT113) differ from agent identities (koby, koonimo).
> **Note**: CT hostnames (tdunna→CT129, baggy→CT114) differ from agent identities (koby, koonimo).
Access: `pct-run <CT_ID> <command>` — no IPs needed. GPU hosts (.8, .110, .15) use SSH.
Keys are stored in Infisical vault (project=agents, env=production) and injected at
@@ -169,9 +169,9 @@ pct-run <CT> grep -A8 "vision:" /root/.hermes/config.yaml | grep api_key
# Must show both api_key: sk-... and api_key_env: LITELLM_API_KEY
```
### For Koby (CT 111 / tdunna)
### For Koby (CT 129 / tdunna)
Koby runs Hermes on CT 111 (tdunna). Config files at `/root/.hermes/config.yaml`.
Koby runs Hermes on CT 129 (tdunna). Config files at `/root/.hermes/config.yaml`.
Same Hermes pattern as Tanko/Mumuni/Koonimo — see config sections above.
**LiteLLM key**: alias `koby` in LiteLLM DB, injected via `infisical run --` wrapper.
+20 -21
View File
@@ -6,11 +6,10 @@ description: >
Enforces shared infrastructure setup (Firecrawl, SearXNG, local models,
RA-H OS MCP) while keeping agent-specific API keys and model choices.
UPDATED 2026-07-16: Compression model is the stable alias `strix-moe` (NOT `ornith-1.0-35b`,
which LiteLLM does not serve). All 3 GPUs verified 256K (RTX 5070 bumped 131K→256K on Jul 15).
which LiteLLM does not serve). All 3 GPUs verified at 128K (reduced from 256K 2026-07-17 for stability).
Added Rule 12 (Context-Issue Diagnostic) + Rule 13 (.env fallback enforcement) from the
2026-07-16 Mumuni root-cause investigation (WAL #1300).
UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo. RTX 3090 context
verified at 256K. Infisical .env fallback required (Rule 3/13).
UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo. RTX 3090 context verified at 128K. Infisical .env fallback required (Rule 3/13).
---
## Maintains
@@ -36,7 +35,7 @@ Sub-agent profiles inherit auth from the main config — no separate keys needed
| Mumuni | `mumuni` | 192.168.68.123 | root@.123 | 6 profiles ✱ |
| Abiba | `abiba-pi` | 192.168.68.24 | local | — |
| Koby | `koby` | CT 111 (tdunna) | Zulip | — |
| Koonimo | `koonimo` | CT 113 (baggy) | Zulip | — |
| Koonimo | `koonimo` | CT 114 (baggy) | SSH root | — |
| Kagenz0 | `kagenz0-*` | ? | Zulip | — |
> CT hostnames (tdunna, baggy) differ from agent identities (koby, koonimo).
@@ -101,8 +100,8 @@ model:
base_url: http://192.168.68.116/v1
api_key_env: LITELLM_API_KEY # Injected via infisical run -- wrapper
max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation
context_length: 262144 # For syslog-auto (all GPUs support 256K).
# Set 131072 if using gemma-4-12b directly (12GB VRAM constraint).
context_length: 131072 # For syslog-auto (all GPUs at 128K for stability).
# Set 65536 if using gemma-4-12b directly (tight VRAM).
fallback_providers:
provider: deepseek
@@ -133,7 +132,7 @@ compression:
enabled: true
model: strix-moe # ⚠️ Must match auxiliary.compression.model. Stable alias (gpu-fleet § Stable Role-Based Aliases). NOT ornith-1.0-35b (LiteLLM does not serve that name).
provider: harness
max_context_window: 262144 # MUST match actual GPU capacity. All 3 GPUs are 256K (Jul 15).
max_context_window: 131072 # MUST match actual GPU capacity. All 3 GPUs are 128K (Jul 17).
threshold: 0.65 # Fires at ~170K for 262K window, ~85K for 128K
target_ratio: 0.30
protect_last_n: 40
@@ -250,7 +249,7 @@ The following MUST be identical across ALL profiles:
### Rule 7: Auxiliary Model Consistency (UPDATED 2026-07-16)
- Vision and web_extract use `gemma-4-12b` (RTX 5070 — 12GB, vision-optimized)
- Compression uses `strix-moe` (stable alias for Strix Halo — 64GB, 256K ctx, compression-optimized)
- Compression uses `strix-moe` (stable alias for Strix Halo — 64GB, 128K ctx, compression-optimized)
- **`strix-moe` is the only valid compression model name** — LiteLLM does NOT serve `ornith-1.0-35b`
(it serves `strix-moe`, `qwen3.6-35B-udq4`, `gpu-dense`, `gpu-light`, `syslog-auto`, `gemma-4-12b`, `qwen3.6-27B-code`). Old configs with `ornith-1.0-35b` cause 403/model-not-found on compression calls.
- All auxiliary services MUST use identical routing:
@@ -258,31 +257,31 @@ The following MUST be identical across ALL profiles:
- `api_key_env: LITELLM_API_KEY`
- **Do NOT use `syslog-auto`** for auxiliary tasks — it routes unpredictably
- **Compression on Strix Halo**: The strix-moe alias routes to Strix Halo
(64GB UMA, 256K context) — the designated compression GPU. This frees the
(64GB UMA, 128K context) — the designated compression GPU. This frees the
RTX 5070 for vision and web search, and the RTX 3090 for heavy reasoning.
- The `compression:` block's `model` MUST match `auxiliary: compression: model`
- The `compression: max_context_window: 262144` MUST match Strix Halo's actual capacity
- The `compression: max_context_window: 131072` MUST match actual GPU capacity (128K)
### Rule 8: GPU Workload Distribution (UPDATED 2026-07-16)
- **RTX 3090 (24GB, 256K ctx, qwen3.6-27B-code)**: Heavy reasoning, code gen, long conversations
- **RTX 5070 (12GB, 256K ctx, gemma-4-12b)**: Vision, web search, quick tasks, web_extract (bumped 131K→256K Jul 15; IQ4_NL+MTP, 88% VRAM)
- **Strix Halo (64GB, 256K ctx, strix-moe)**: Context compression, summarization, long docs
- **RTX 3090 (24GB, 128K ctx, qwen3.6-27B-code)**: Heavy reasoning, code gen, long conversations
- **RTX 5070 (12GB, 128K ctx, gemma-4-12b)**: Vision, web search, quick tasks, web_extract (IQ4_NL+MTP, ~65% VRAM at 128K)
- **Strix Halo (64GB, 128K ctx, strix-moe)**: Context compression, summarization, long docs
- Agent profiles MUST route auxiliary tasks to the correct GPU:
- `auxiliary.vision.model: gemma-4-12b` (RTX 5070)
- `auxiliary.web_extract.model: gemma-4-12b` (RTX 5070)
- `auxiliary.compression.model: strix-moe` (Strix Halo)
- Default model (`model.default`) and custom_provider remain `syslog-auto` for auto-routing
- For 262K context window: `threshold: 0.65` (fires at ~170K tokens)
- For 128K context window: `threshold: 0.65` (fires at ~85K tokens)
- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss
- Do NOT use `threshold: 0.80` — this delays until 209K, risking the gateway hygiene layer
- `max_context_window: 262144` MUST match the model's actual capacity
- Do NOT use `threshold: 0.80` — this delays until ~105K, leaving only 23K margin
- `max_context_window: 131072` MUST match the model's actual capacity (128K)
- See `devops-hermes-compression` skill for full reference
### Rule 9: Compression Threshold for 256K Models
- For 262K context window: `threshold: 0.65` (fires at ~170K tokens)
### Rule 9: Compression Threshold for 128K Models
- For 128K context window: `threshold: 0.65` (fires at ~85K tokens)
- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss
- Do NOT use `threshold: 0.80` — this delays until 209K, risking the gateway hygiene layer
- `max_context_window: 262144` MUST match the model's actual capacity (Strix Halo = 256K)
- Do NOT use `threshold: 0.80` — this delays until ~105K, leaving only 23K margin
- `max_context_window: 131072` MUST match the model's actual capacity (all GPUs = 128K)
- See `devops-hermes-compression` skill for full reference
### Rule 10: Default Model Must Be `syslog-auto` (All Agents)
@@ -312,7 +311,7 @@ When an agent shows "context issues" (premature compression, 401s, 504s, DeepSee
verify ALL FOUR of these against the live config. They are the only root causes found in production:
1. **max_context_window correct?** — BOTH `compression.max_context_window` AND `context.max_context_window`
MUST be `262144` (all GPUs are 256K). A value of `131072` causes premature compression at
MUST be `131072` (all GPUs are 128K). A value of `262144` causes instability near 100K and must NOT be used.
~83K instead of ~170K. Check: `grep -n max_context_window ~/.hermes/config.yaml`
2. **base_url uses /v1 NOT /litellm/v1?** — `custom_providers[0].base_url`, `delegation.base_url`,
and ALL `auxiliary.*.base_url` MUST be `http://192.168.68.116/v1` (Rule 5). nginx `/litellm/`
+104
View File
@@ -0,0 +1,104 @@
---
name: inference-optimization
kind: responsibility
description: >
Optimizes the full Syslog inference stack — LiteLLM routing weights, GPU model
assignments, agent context management, and prompt caching — to reduce response
times to sub-15s average. All GPUs now at 128K context (stable ceiling).
id: 067NC6KP02RG60S50M40E30928
---
### Goal
Syslog inference response times reduced to sub-15s average by optimizing the full
stack: LiteLLM routing weights, GPU model assignments, Hermes agent context
management, and prompt caching — without sacrificing agent capability.
### Requires
- `inference-metrics`: current SpendLogs from CT116 LiteLLM Postgres — avg
request_duration_ms, prompt_tokens, completion_tokens, model_group breakdown,
cache_hit rate over the last 3 hours
- `agent-configs`: current config.yaml from each active Hermes agent (Mumuni
.123, any others on .129/.122) including compression, model, context_window,
prompt_caching, memory settings
- `gpu-health`: health check response from all 3 GPU backends (ornith .15:8080,
qwen .8:8080, gemma .110:8080)
### Maintains
The optimized inference stack configuration — every change is applied and
verified end-to-end. Postcondition: avg request_duration_ms ≤ 15000 for 90% of
non-ornith traffic; ≤ 30000 for ornith-bound agentic calls.
#### liteLLM-routing
The syslog-auto routing weights, model-specific timeouts, RPM limits, and
model_list entries on CT116 `/opt/inference-harness/litellm_config.yaml`.
#### agent-compression
Each Hermes agent's `~/.hermes/config.yaml` compression, context_window,
prompt_caching, and model sections.
#### prompt-caching
LiteLLM cache configuration and llama.cpp `--cache-prompt` flag on GPU hosts.
#### verification
End-to-end latency measurements after changes applied — at least 3 test
inference calls per model path measuring ttft (time-to-first-token) and total
duration.
### Continuity
- input-driven
### Strategies
**Context is the root cause.** Every ~46K prompt token costs ~87s of ornith
prefill time at 532 tok/s. Fix context first, routing second.
- **Route by task**: ornith for multi-step reasoning only; qwen for code/standard
queries; gemma for compression/auxiliary. Never send simple completion to a
35B MoE.
- **Compress aggressively**: threshold at 40% (not 65%) — a 256K window should
compact at 102K, not 166K. Target 15% tail (not 30%).
- **Cache everything repeated**: system prompts, skill docs, AGENTS.md — these
never change between turns. Single-digit cache hit rate is unacceptable.
- **Lower context ceiling**: 128K window is the stable ceiling for agent conversations.
GPUs reduced from 256K to 128K (2026-07-17). For larger contexts, route to external providers.
### Shape
- `self`: analyze metrics, compute optimal configs, apply changes, verify
- `delegates`:
- `apply-liteLLM`: update litellm_config.yaml and reload
- `apply-agent-config`: update hermes config.yaml per agent
- `verify-latency`: run test inference calls and measure response
### Execution
```prose
-- Phase 1: Analyze current state (already complete)
-- Phase 2: Apply LiteLLM routing optimization
call apply-liteLLM-routing
config_path: /opt/inference-harness/litellm_config.yaml
host: 192.168.68.116
-- Phase 3: Apply agent context compression optimization
call apply-agent-compression
agent: mumuni
host: 192.168.68.123
config_path: /root/.hermes/config.yaml
-- Phase 4: Enable llama.cpp prompt caching on GPU hosts
call enable-prompt-caching
hosts: [192.168.68.15, 192.168.68.8, 192.168.68.110]
-- Phase 5: Verify end-to-end latency
call verify-latency
host: 192.168.68.116
models: [syslog-auto, qwen3.6-27B-code, gemma-4-12b, ornith-1.0-35b]
```
+147 -79
View File
@@ -20,9 +20,20 @@ description: >
Canonical process: see § Production Vault Access Process. Tanko (user jerome) pending.
Abiba's key is now a proper agent key (NOT the master key — stale note removed).
UPDATED 2026-07-17: FLEET-WIDE STANDARDIZATION. All 4 agents (Mumuni, Tanko, Koby, Koonimo)
standardized on a single pattern: systemd drop-in (ExecStart= reset + wrapper path) →
infisical-gateway.sh while-true loop → /usr/bin/infisical run --token → bash -c key
injection → .env fallback → exec python. Systemd drop-ins are IMMUNE to hermes gateway
install which overwrites the unit file ExecStart. Infisical CLI updated to 0.43.109 on
all agents (was 0.38.0). Service token st.8e848433 shared across fleet (st.353699cd
for tanko was deleted). .env fallback on every agent protects against token loss.
Critical lessons: (1) NEVER use shell variables inside single-quoted bash -c in wrappers
— hardcode absolute paths. (2) Drop-ins override unit file ExecStart permanently.
(3) Capture /proc/<pid>/environ before gateway restarts to preserve running env set.
Current key inventory and agent list: see gpu-fleet.prose.md § Agent Keys.
Source of truth for LiteLLM config: /opt/inference-harness/litellm_config.yaml
on CT 116. Last verified: 2026-07-16.
on CT 116. Last verified: 2026-07-17.
---
## Parameters
@@ -74,84 +85,147 @@ description: >
- Confirm key alias matches agent_name in LiteLLM key list
- Verify agent gateway uses vault wrapper: `cat /proc/<pid>/cmdline` shows `infisical run`
## Production Vault Access Process (canonical, 2026-07-16)
## Production Vault Access Process (canonical, 2026-07-17)
The non-fail approach to agentic vault access. Deployed on 4/5 agents (tanko pending —
runs as user `jerome`, not systemd root, needs user-scope adaptation).
The non-fail approach to agentic vault access. Deployed on all 4 Hermes agents
(Mumuni, Tanko, Koby, Koonimo) as of 2026-07-17. Abiba (pi) uses a similar pattern
through its agent wrapper.
### The canonical pattern
1. **infisical CLI** installed on the host (`/usr/local/bin/infisical` or `/usr/bin/infisical`).
2. **Service token** (Infisical Machine Identity, `st.…`) stored root-only at `/root/.infisical-token` (`chmod 600`).
- Interim: the shared `abiba` service token (`st.8e848433…`) has READ+WRITE on the `agents` project.
- Proper: one machine identity per agent (create in Infisical UI → Project Settings → Machine Identities).
3. **`infisical-gateway.sh` wrapper** at `/root/.hermes/infisical-gateway.sh` (`chmod 700`):
1. **infisical CLI** installed on the host at `/usr/bin/infisical` (v0.43.109+, from
artifacts-cli.infisical.com apt repo). Update procedure:
```bash
curl -1sLf 'https://artifacts-cli.infisical.com/setup.deb.sh' | sudo -E bash
sudo apt-get update && sudo apt-get install -y infisical
# Remove stale old binary if present
rm -f /usr/local/bin/infisical /bin/infisical
```
Wrappers use absolute path `/usr/bin/infisical run`. Never rely on PATH resolution.
2. **Service token** (Infisical Machine Identity, `st.…`) stored at `~/.infisical-token`
(`chmod 600`). Current: shared `st.8e848433…` (abiba, READ+WRITE on agents project).
Tanko's `st.353699cd…` (tanko-agent) was deleted — reverted to shared token.
Proper: one machine identity per agent (create in Infisical UI → Project Settings →
Machine Identities).
3. **`infisical-gateway.sh` wrapper** at `~/.hermes/infisical-gateway.sh` (`chmod 700`):
```bash
#!/bin/bash
export INFISICAL_API_URL="https://vault.sysloggh.net"
TOKEN=$(cat /root/.infisical-token)
LOG=/root/.hermes/logs/gateway.log; mkdir -p /root/.hermes/logs
TOKEN=$(cat $HOME/.infisical-token)
LOG=$HOME/.hermes/logs/gateway.log; mkdir -p $HOME/.hermes/logs
while true; do
infisical run --token="$TOKEN" --projectId=322fceab-39da-4854-a55a-568e76c0f13f \
echo "[$(date -Iseconds)] Starting gateway with Infisical injection..." >> $LOG
/usr/bin/infisical run --token="$TOKEN" \
--projectId=322fceab-39da-4854-a55a-568e76c0f13f \
--env=prod --domain=https://vault.sysloggh.net -- bash -c '
. /root/.hermes/.env 2>/dev/null # [FALLBACK Rule 3] safety net only
export LITELLM_API_KEY="$<AGENT>_LITELLM_API_KEY"
exec <HERMES_VENV>/bin/python -m hermes_cli.main gateway run
. $HOME/.hermes/.env 2>/dev/null # [FALLBACK Rule 3]
export LITELLM_API_KEY="${<AGENT>_LITELLM_API_KEY}"
export ZULIP_API_KEY="${<AGENT>_ZULIP_API_KEY}"
export ZULIP_SITE="https://chat.sysloggh.net"
export ZULIP_EMAIL="<agent>-bot@chat.sysloggh.net"
export SEARXNG_URL="http://192.168.68.7:8888"
# ⚠️ HARDCODE the full venv path. NEVER use $VENV inside single quotes.
exec /root/.hermes/hermes-agent/venv/bin/python -m hermes_cli.main gateway run
' >> $LOG 2>&1
sleep 5 # restart on exit
EXIT_CODE=$?
echo "[$(date -Iseconds)] Gateway exited with code $EXIT_CODE — restarting in 5s..." >> $LOG
sleep 5
done
```
4. **Agent key in vault** as `<AGENT>_LITELLM_API_KEY` (e.g. `KOBY_LITELLM_API_KEY`). Vault = source of truth.
5. **`.env` fallback** at `/root/.hermes/.env` (`chmod 600`) with the same key — safety net ONLY for vault outage (Rule 3/13). Must be kept in sync on rotation.
6. **systemd service** `hermes-gateway.service` with `ExecStart=/root/.hermes/infisical-gateway.sh`. NO `litellm-key.conf` drop-in (those hardcode keys and rot).
7. **NEVER hardcode** LiteLLM keys in systemd drop-ins, config.yaml, or /etc/environment. The wrapper injects live from vault.
**CRITICAL: VENV PATH.** The inner `bash -c '...'` uses single quotes. Shell
variables set in the outer wrapper are NOT expanded inside single quotes.
`$VENV/bin/python` resolves to `/bin/python` (file not found). Always hardcode
the absolute path to the venv python binary.
4. **Agent key in vault** as `<AGENT>_LITELLM_API_KEY` and `<AGENT>_ZULIP_API_KEY`.
Vault = source of truth for ALL platform credentials.
5. **`.env` fallback** at `~/.hermes/.env` (`chmod 600`) with agent-specific keys —
safety net for vault outage or token revocation. Must be kept in sync on rotation.
Example:
```bash
MUMUNI_LITELLM_API_KEY=sk-OzuWsoX22Hmb3Ps3JY01gw
MUMUNI_ZULIP_API_KEY=H8dY6V7aHmWNcfgNtJaDBPZ1dGWn0Ttt
```
6. **systemd drop-in** at `~/.config/systemd/user/hermes-gateway.service.d/50-vault-wrapper.conf`:
```ini
[Service]
ExecStart=
ExecStart=/root/.hermes/infisical-gateway.sh
```
The `ExecStart=` (empty reset) clears any ExecStart from the main unit file,
then the second `ExecStart=` sets the wrapper. This drop-in **survives unit file
regeneration** by `hermes gateway install` — the drop-in always wins.
**Why a drop-in instead of editing the unit file:** `hermes gateway install`
(called during Hermes updates and some self-heal operations) regenerates the
systemd unit file with `ExecStart=/path/to/python -m hermes_cli.main gateway run`.
Editing the unit file directly is futile — it will be overwritten. The drop-in
approach explicitly resets ExecStart and sets the wrapper regardless of what the
main unit file says.
7. **NEVER hardcode** API keys in systemd drop-ins, config.yaml, or /etc/environment.
The wrapper injects live from vault at every start.
### Why this is non-fail
- **No rot**: keys pulled live from vault at every gateway start. Rotation = one `infisical secrets set` + `systemctl restart`. No per-host file edits.
- **Survives vault outage**: the `.env` fallback (Rule 3) keeps the gateway running if Infisical is unreachable.
- **Survives gateway crash**: the wrapper's `while true` + systemd `Restart=on-failure` revive the gateway.
- **Auditable**: `cat /proc/$(pgrep hermes_cli)/environ` shows the live key; `infisical secrets` shows the vault source.
- **Survives vault outage**: the `.env` fallback (Rule 3) keeps the gateway running if Infisical is unreachable or the service token is revoked.
- **Survives gateway crash**: the wrapper's `while true` + systemd `Restart=always` revive the gateway. Two-layer defense.
- **Survives Hermes updates**: systemd drop-in overrides unit file ExecStart — `hermes gateway install` cannot break the vault injection.
- **Survives reboot**: systemd user service + `loginctl enable-linger` ensures gateway starts at boot without a login session.
- **Auditable**: `cat /proc/$(pgrep hermes_cli)/environ` shows all injected keys; `infisical secrets` shows the vault source.
### Migration status (2026-07-16)
### Migration status (2026-07-17)
| Agent | Host | Pattern | Vault key | Status |
|-------|------|---------|-----------|--------|
| abiba | .24 | `infisical run` (pi agent wrapper, service token) | ABIBA_LITELLM_API_KEY | ✅ vault-backed |
| mumuni | .123 | infisical-gateway.sh + user-login machine identity | MUMUNI_LITELLM_API_KEY | ✅ vault-backed |
| koby | .129 | infisical-gateway.sh + service token (migrated 2026-07-16) | KOBY_LITELLM_API_KEY | ✅ vault-backed, Zulip (tanko-bot@) + Telegram |
| koonimo | .114 | infisical-gateway.sh + service token (migrated 2026-07-16) | KOONIMO_LITELLM_API_KEY | ✅ vault-backed |
> **Baggy = Koonimo (CT 113).** Deleted `BAGGY_LITELLM_API_KEY` from vault 2026-07-16. Only `KOONIMO_LITELLM_API_KEY` exists — one secret per agent.
| tanko | .122 | **hardcoded in config.yaml** (runs as user jerome, not systemd) | TANKO_LITELLM_API_KEY | ⚠️ TODO: migrate to user-scope wrapper |
| Agent | Host | Pattern | Keys | Status |
|-------|------|---------|------|--------|
| abiba | .24 | pi agent wrapper | ABIBA_LITELLM_API_KEY + ABIBA_ZULIP_API_KEY | ✅ vault-backed |
| mumuni | .123 | systemd drop-in + while-true wrapper + st.8e848433 | MUMUNI_LITELLM_API_KEY + MUMUNI_ZULIP_API_KEY | ✅ vault-backed + .env fallback |
| tanko | .122 | systemd drop-in + while-true wrapper + st.8e848433 (user jerome) | TANKO_LITELLM_API_KEY + TANKO_ZULIP_API_KEY | ✅ vault-backed + .env fallback |
| koby | .129 | systemd drop-in + while-true wrapper + st.8e848433 | KOBY_LITELLM_API_KEY, shares TANKO_ZULIP_API_KEY (tanko-bot) | ✅ vault-backed |
| koonimo | .114 | systemd drop-in + while-true wrapper + st.8e848433 | KOONIMO_LITELLM_API_KEY + KOONIMO_ZULIP_API_KEY | ✅ vault-backed |
### Tanko migration (pending)
> Tanko runs as user `jerome` — wrapper/token at `~/.hermes/infisical-gateway.sh` and
> `~/.infisical-token`. Linger enabled (`loginctl enable-linger jerome`) for boot startup.
Tanko runs the gateway as user `jerome` (not root/systemd), with the key hardcoded in
`/home/jerome/.hermes/config.yaml` (`api_key: sk-CggiHWlamQy…`, valid but not vault-sourced).
Migration: create a user-scope systemd service (`~/.config/systemd/user/hermes-gateway.service`)
with `infisical-gateway.sh` wrapper in jerome's home, token at `~/.infisical-token`, lingering
enabled (`loginctl enable-linger jerome`) so the user service runs without a login session.
### Tanko migration (COMPLETED 2026-07-17)
### Koby migration lessons (2026-07-16)
Tanko was the last agent migrated from hardcoded keys to vault wrapper.
Previously: key hardcoded in `/home/jerome/.hermes/config.yaml` (`api_key: sk-CggiHWlamQy…`)
and `zulip-env.conf` systemd drop-in. Now: user-scope systemd service with drop-in
`50-vault-wrapper.conf`, `infisical-gateway.sh` wrapper with while-true loop, token at
`~/.infisical-token`, `.env` fallback at `~/.hermes/.env`. Keys injected live from vault.
### Koby migration lessons (2026-07-16, updated 2026-07-17)
Migrated Koby from hardcoded systemd drop-in → `infisical-gateway.sh` wrapper.
**Two mistakes I made that broke the agent:**
**Three mistakes made:**
1. **Overwrote `/root/.hermes/.env`** without backing it up. The Zulip API key only existed
in the running process memory — the old .env was minimal (just LiteLLM key). Zulip creds were
inherited from the pre-migration gateway env, not stored in any file. Lost on restart.
2. **Only injected `LITELLM_API_KEY`** in the wrapper — forgot Zulip + Telegram credentials.
Agents need ALL their platform env vars. Missing vars cause silent adapter failures.
3. (2026-07-17 fix) **VENV variable in single-quoted bash -c**: `exec "$VENV/bin/python"`
inside single quotes resolved to `exec "/bin/python"` (file not found). Hardcoded full path.
**How Koby actually connects (2026-07-16):**
**How Koby actually connects:**
- Zulip: shares **Tanko's bot** (`tanko-bot@chat.sysloggh.net`, `TANKO_ZULIP_API_KEY=5PeD6f3zo…`).
Koby doesn't have its own Zulip bot (koby-bot@ doesn't exist in the swarm config).
- Telegram: token `828640…` recovered from `.env.bak-20260603` (18KB backup from June 2026).
Allowed users: 6679773481. Home channel: 6679773481.
- Telegram: token from `.env` fallback. Allowed users: 6679773481.
- Both platforms now connect through the wrapper's env injection.
**Golden rule for gateway restarts:** always `cat /proc/<pid>/environ` before killing the old
process — captures the live env set. Especially important when migrating gateways between
injection mechanisms.
**Golden rules for gateway restarts:**
1. Always `cat /proc/<pid>/environ` before killing the old process — captures the live env set.
2. Hardcode venv python path in wrapper — never use variables inside single-quoted bash -c.
3. Use systemd drop-ins (not unit file edits) to override ExecStart — survives Hermes updates.
### Fleet-wide standardization lessons (2026-07-17)
After auditing all 4 agents, five systemic patterns caused repeated failures:
1. **Three incompatible startup patterns** coexisted (systemd drop-in, direct python, orphaned wrapper)
2. **Systemd unit files reverted** by `hermes gateway install` during updates
3. **VENV variable scoping** broke wrappers on Koby and Mumuni (single-quote bash -c)
4. **Service token expiry** — Tanko's `st.353699cd` was deleted from Infisical
5. **No ZULIP_API_KEY** in env on Tanko — wrapper bypassed by systemd direct python
All resolved by the canonical drop-in + while-true wrapper pattern documented above.
### Key rotation procedure (one vault operation with this standard)
@@ -161,47 +235,41 @@ injection mechanisms.
4. Restart: `systemctl restart hermes-gateway`. The wrapper pulls the new key live.
5. Verify: `curl -H "Authorization: Bearer sk-NEW" http://192.168.68.116/v1/models` → 200.
## Machine Identity for Vault Writes (ADDED 2026-07-16, WAL #1300)
## Machine Identity for Vault Writes (UPDATED 2026-07-17)
**Problem:** The infisical CLI on agent hosts is logged in as a user session (jerome@sysloggh.com).
In CLI v0.38.0, `infisical secrets set` / `infisical export` fail with "project id missing" / "workspace
key 404" — a known bug where user-session auth works for `run` but NOT for `secrets set`. The apt
repo only ships 0.38.0, so `apt upgrade` does not help.
**Current state:** Infisical CLI updated to v0.43.109 on all agents (from v0.38.0).
The v0.38.0 bug (user-session auth fails for `secrets set`/`export`) is resolved.
Service token `st.8e848433…` (abiba, READ+WRITE) can write to vault from CLI.
**Proper fix — Machine Identity (Infisical automation best practice):**
Create a machine identity with READ+WRITE scope on the `agents` project (project_id=
`322fceab-39da-4854-a55a-568e76c0f13f`, env `prod`). Store client_id + client_secret securely.
Then vault writes work from any host:
```bash
# Get a machine-identity access token
TOKEN=$(curl -fsSL -X POST https://vault.sysloggh.net/api/v1/auth/universal-auth/login \
-H 'Content-Type: application/json' \
-d '{"clientId":"<CLIENT_ID>","clientSecret":"<CLIENT_SECRET>"}' | jq -r .accessToken)
# Write a secret via REST API v3
curl -fsSL -X PATCH https://vault.sysloggh.net/api/v3/secrets/MUMUNI_LITELLM_API_KEY \
-H "Authorization: Bearer $TOKEN" -H 'Content-Type: application/json' \
-d '{"environment":"prod","secretValue":"sk-<NEW_KEY>","workspaceId":"<WORKSPACE_ID>","type":"shared"}'
# OR via CLI: infisical secrets set --token=$TOKEN --projectId=322fceab... --env=prod ...
```
Creation requires the Infisical web UI (https://vault.sysloggh.net) under Project Settings →
Machine Identities, or an admin API call. **TODO: create `abiba-automation` machine identity
and store its credentials in the vault itself (or a root-only file).**
**Proper fix — per-agent Machine Identities:**
Create machine identities in Infisical UI → Project Settings → Machine Identities
for each agent with READ-only scope on the `agents` project. Store client_id +
client_secret per agent. Then vault writes use the shared abiba identity, and
reads use per-agent identities. This eliminates the single shared token risk.
**Interim (working now):** the `.env` fallback (hermes-config-template Rule 3/13). The
infisical-gateway.sh wrapper sources `~/.hermes/.env`, so its `<AGENT>_LITELLM_API_KEY`
overrides a stale vault value.
**Service Token Inventory (2026-07-17):**
| Token ID | Name | Permissions | Used By | Status |
|----------|------|-------------|---------|--------|
| `st.8e848433…` | tanko-gateway | READ+WRITE | Mumuni, Tanko, Koby, Koonimo, Abiba | ✅ Active |
| `st.353699cd…` | tanko-agent | READ-only | — | ❌ Deleted from Infisical |
**2026-07-16 UPDATE — vault is now SYNCED.** The abiba service token (`st.8e848433…`, READ+WRITE)
can write to the vault, so the session-13 rotated keys (mumuni `sk-OzuWsoX2…`, koby `sk-BqRRMboTI…`,
koonimo `sk-OEK7z26n6E…`) are now in the vault as `MUMUNI_LITELLM_API_KEY` / `KOBY_LITELLM_API_KEY` /
`KOONIMO_LITELLM_API_KEY` and validate 200 against LiteLLM. The vault is the source of truth again.
Creating a dedicated `abiba-automation` machine identity (via UI) is still the proper long-term fix
so the shared service token isn't reused across hosts — but it is no longer blocking.
**Per-agent .env fallback inventory (2026-07-17):**
| Agent | .env Keys |
|-------|-----------|
| Mumuni | MUMUNI_LITELLM_API_KEY, MUMUNI_ZULIP_API_KEY |
| Tanko | TANKO_LITELLM_API_KEY, TANKO_ZULIP_API_KEY |
| Koby | (wrapper injects from vault — .env has Telegram token) |
| Koonimo | KOONIMO_LITELLM_API_KEY, KOONIMO_ZULIP_API_KEY |
## Key Rotation Log
| Date | Agent | Action | Notes |
|------|-------|--------|-------|
| 2026-07-17 | fleet | standardize | All 4 agents standardized on systemd drop-in + while-true wrapper + infisical v0.43.109. Removed conflicting zulip-env.conf + litellm-key.conf drop-ins. Added .env fallbacks with ZULIP keys. WAL #1322. |
| 2026-07-17 | tanko | fix-zulip | Added ZULIP_API_KEY to env (was missing — systemd bypassed vault). Updated wrapper from exec to while-true. Created .env fallback. Removed hardcoded zulip-env.conf drop-in. WAL #1321. |
| 2026-07-16 | vault | cleanup | 4 stale secrets deprecated. 5 personal creds flagged. |
| 2026-07-16 | koonimo | add-zulip | Added KOONIMO_ZULIP_API_KEY to vault. Wrapper injects ZULIP_API_KEY + ZULIP_EMAIL. 3 platforms. |
| 2026-07-16 | tanko | migrate | Migrated from hardcoded config.yaml to infisical-gateway.sh + st.353699cd. NOTE: st.353699cd later deleted — reverted to st.8e848433 on 2026-07-17. |
| 2026-07-16 | mumuni | rotate | Old key malformed (sk-_SWAl_Vu_, 47 chars, not LiteLLM format) → 401. Deleted old `mumuni` key (token 15cbca18…), generated fresh (alias `mumuni`, 7 models: syslog-auto, qwen3.6-27B-code, gemma-4-12b, strix-moe, gpu-dense, gpu-light, qwen3.6-35B-udq4). New key sk-OzuWsoX2… written to /root/.hermes/.env (Rule 3/13 fallback). Vault sync PENDING (needs machine identity). WAL #1300. |
| 2026-07-16 | koby | rotate | Old key sk-6sbCNjz (401, stale in /etc/environment). Deleted old `koby` key, generated fresh (alias `koby`). New key sk-BqRRMboTI… in systemd drop-in `hermes-gateway.service.d/litellm-key.conf` + /etc/environment. Created `hermes-gateway.service` unit (was missing — gateway wasn't persistent) with `--replace`. Verified HTTP 200, Telegram connected. |
| 2026-07-16 | baggy (koonimo) | rotate | Old key sk-krnw_zGB (401, hardcoded in systemd drop-in). Deleted old `baggy` key, generated fresh (alias `baggy`, metadata agent=koonimo). New key sk-OEK7z26n6E… in drop-in `hermes-gateway.service.d/litellm-key.conf`. CT113 IP changed .113→.114. Verified HTTP 200, Zulip connected. |