Compare commits
11
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
38a32f8b32 | ||
|
|
96b0caa0b3 | ||
|
|
ea17f64bb4 | ||
|
|
c13a15acad | ||
|
|
68f9ebff74 | ||
|
|
93fb0d8e1e | ||
|
|
cb6efa81c3 | ||
|
|
2dc77ee7da | ||
|
|
561c4d98c9 | ||
|
|
c243fcddc3 | ||
|
|
50114f32c0 |
@@ -584,7 +584,7 @@ contracts:
|
|||||||
verify: curl -sf https://git.sysloggh.net/api/v1/version
|
verify: curl -sf https://git.sysloggh.net/api/v1/version
|
||||||
expect: 200 OK
|
expect: 200 OK
|
||||||
- check: SearXNG reachable
|
- check: SearXNG reachable
|
||||||
verify: curl -sf http://192.168.68.17:8080
|
verify: curl -sf http://192.168.68.7:8888
|
||||||
expect: 200 OK
|
expect: 200 OK
|
||||||
artifact: infrastructure health report
|
artifact: infrastructure health report
|
||||||
receipt:
|
receipt:
|
||||||
|
|||||||
@@ -356,7 +356,7 @@ Postconditions to verify:
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"check": "SearXNG reachable",
|
"check": "SearXNG reachable",
|
||||||
"verify": "curl -sf http://192.168.68.17:8080",
|
"verify": "curl -sf http://192.168.68.7:8888",
|
||||||
"expect": "200 OK"
|
"expect": "200 OK"
|
||||||
}
|
}
|
||||||
]
|
]
|
||||||
|
|||||||
@@ -5,8 +5,7 @@ description: >
|
|||||||
Standard Hermes configuration template for Syslog Solution LLC agents.
|
Standard Hermes configuration template for Syslog Solution LLC agents.
|
||||||
Enforces shared infrastructure setup (Firecrawl, SearXNG, local models,
|
Enforces shared infrastructure setup (Firecrawl, SearXNG, local models,
|
||||||
RA-H OS MCP) while keeping agent-specific API keys and model choices.
|
RA-H OS MCP) while keeping agent-specific API keys and model choices.
|
||||||
UPDATED 2026-07-16: Compression model is the stable alias `strix-moe` (NOT `ornith-1.0-35b`,
|
UPDATED 2026-08-07: Added Rule 15 (MCP Validation) from the 2026-08-07 keyless-MCP incident.
|
||||||
which LiteLLM does not serve). All 3 GPUs verified at 128K (reduced from 256K 2026-07-17 for stability).
|
|
||||||
Added Rule 12 (Context-Issue Diagnostic) + Rule 13 (.env fallback enforcement) from the
|
Added Rule 12 (Context-Issue Diagnostic) + Rule 13 (.env fallback enforcement) from the
|
||||||
2026-07-16 Mumuni root-cause investigation (WAL #1300).
|
2026-07-16 Mumuni root-cause investigation (WAL #1300).
|
||||||
UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo. RTX 3090 context verified at 128K. Infisical .env fallback required (Rule 3/13).
|
UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo. RTX 3090 context verified at 128K. Infisical .env fallback required (Rule 3/13).
|
||||||
@@ -48,7 +47,7 @@ Sub-agent profiles inherit auth from the main config — no separate keys needed
|
|||||||
| Component | Endpoint | Purpose |
|
| Component | Endpoint | Purpose |
|
||||||
|---|---|---|
|
|---|---|---|
|
||||||
| Firecrawl | `http://192.168.68.7:3002/` | Web content extraction |
|
| Firecrawl | `http://192.168.68.7:3002/` | Web content extraction |
|
||||||
| SearXNG | `http://storepve:8888` | Privacy-respecting web search |
|
| SearXNG | `http://192.168.68.7:8888` | Privacy-respecting web search |
|
||||||
| LiteLLM | `http://192.168.68.116/v1` | Unified model gateway (via nginx) |
|
| LiteLLM | `http://192.168.68.116/v1` | Unified model gateway (via nginx) |
|
||||||
| LiteLLM (NetBird) | `https://litellm.sysloggh.net/v1` | Alternative (may have 502 issues) |
|
| LiteLLM (NetBird) | `https://litellm.sysloggh.net/v1` | Alternative (may have 502 issues) |
|
||||||
| RA-H OS MCP | `http://192.168.68.65:3100/mcp` | Knowledge graph bridge |
|
| RA-H OS MCP | `http://192.168.68.65:3100/mcp` | Knowledge graph bridge |
|
||||||
@@ -97,10 +96,14 @@ work immediately after restart.
|
|||||||
model:
|
model:
|
||||||
default: <agent_model> # e.g., strix-moe, qwen3.6-27B-code, syslog-auto
|
default: <agent_model> # e.g., strix-moe, qwen3.6-27B-code, syslog-auto
|
||||||
provider: harness
|
provider: harness
|
||||||
base_url: http://192.168.68.116/v1
|
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated path; /v1 also OK
|
||||||
api_key_env: LITELLM_API_KEY # Injected via infisical run -- wrapper
|
api_key_env: LITELLM_API_KEY # Injected via infisical run -- wrapper
|
||||||
max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation
|
max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation
|
||||||
context_length: 131072 # For syslog-auto (all GPUs at 128K for stability).
|
context_length: 131072 # For syslog-auto (all GPUs at 128K for stability).
|
||||||
|
# ⚠️ MANDATORY: Hermes probes unknown models from 256K
|
||||||
|
# and falls back to 256K when /v1/models lacks a context
|
||||||
|
# field (llama-server does). Without this override, agents
|
||||||
|
# silently run syslog-auto at 256K (verified 2026-08-09).
|
||||||
# Set 65536 if using gemma-4-12b directly (tight VRAM).
|
# Set 65536 if using gemma-4-12b directly (tight VRAM).
|
||||||
|
|
||||||
fallback_providers:
|
fallback_providers:
|
||||||
@@ -143,7 +146,7 @@ compression:
|
|||||||
# ─── Auxiliary Tasks (CONSISTENCY RULE) ───
|
# ─── Auxiliary Tasks (CONSISTENCY RULE) ───
|
||||||
# All auxiliary services MUST use identical model, base_url, and api_key_env:
|
# All auxiliary services MUST use identical model, base_url, and api_key_env:
|
||||||
# model: gpu-light # stable alias (NOT raw "gemma-4-12b")
|
# model: gpu-light # stable alias (NOT raw "gemma-4-12b")
|
||||||
# base_url: http://192.168.68.116/v1
|
# base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
|
||||||
# api_key_env: LITELLM_API_KEY
|
# api_key_env: LITELLM_API_KEY
|
||||||
# Do NOT use syslog-auto for auxiliary tasks — it routes to the primary GPU.
|
# Do NOT use syslog-auto for auxiliary tasks — it routes to the primary GPU.
|
||||||
# gpu-light = RTX 5070 (12B), freeing the Strix Halo for agent reasoning.
|
# gpu-light = RTX 5070 (12B), freeing the Strix Halo for agent reasoning.
|
||||||
@@ -154,20 +157,20 @@ auxiliary:
|
|||||||
vision:
|
vision:
|
||||||
provider: harness
|
provider: harness
|
||||||
model: gpu-light # stable alias for RTX 5070 (was raw gemma-4-12b)
|
model: gpu-light # stable alias for RTX 5070 (was raw gemma-4-12b)
|
||||||
base_url: http://192.168.68.116/v1
|
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
|
||||||
api_key_env: LITELLM_API_KEY
|
api_key_env: LITELLM_API_KEY
|
||||||
timeout: 60
|
timeout: 60
|
||||||
download_timeout: 30
|
download_timeout: 30
|
||||||
web_extract:
|
web_extract:
|
||||||
provider: harness
|
provider: harness
|
||||||
model: gpu-light # stable alias for RTX 5070
|
model: gpu-light # stable alias for RTX 5070
|
||||||
base_url: http://192.168.68.116/v1
|
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
|
||||||
api_key_env: LITELLM_API_KEY
|
api_key_env: LITELLM_API_KEY
|
||||||
timeout: 30
|
timeout: 30
|
||||||
compression:
|
compression:
|
||||||
provider: harness
|
provider: harness
|
||||||
model: syslog-auto # MUST match compression.model above. Stable alias for Strix Halo (weighted pool).
|
model: syslog-auto # MUST match compression.model above. Stable alias for Strix Halo (weighted pool).
|
||||||
base_url: http://192.168.68.116/v1 # Rule 5: /v1 NOT /litellm/v1
|
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated path; /v1 also OK
|
||||||
api_key_env: LITELLM_API_KEY
|
api_key_env: LITELLM_API_KEY
|
||||||
timeout: 300 # gpu-fleet: 300s for large-history summarization (was 60)
|
timeout: 300 # gpu-fleet: 300s for large-history summarization (was 60)
|
||||||
|
|
||||||
@@ -177,14 +180,14 @@ auxiliary:
|
|||||||
delegation:
|
delegation:
|
||||||
model: gpu-dense # stable alias for RTX 3090 (was raw qwen3.6-27B-code)
|
model: gpu-dense # stable alias for RTX 3090 (was raw qwen3.6-27B-code)
|
||||||
provider: harness
|
provider: harness
|
||||||
base_url: http://192.168.68.116/v1
|
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
|
||||||
api_key_env: LITELLM_API_KEY
|
api_key_env: LITELLM_API_KEY
|
||||||
|
|
||||||
# ─── Custom Provider ───
|
# ─── Custom Provider ───
|
||||||
custom_providers:
|
custom_providers:
|
||||||
- name: harness
|
- name: harness
|
||||||
model: syslog-auto # weighted pool (default)
|
model: syslog-auto # weighted pool (default)
|
||||||
base_url: http://192.168.68.116/v1
|
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
|
||||||
api_key_env: LITELLM_API_KEY
|
api_key_env: LITELLM_API_KEY
|
||||||
api_mode: chat_completions
|
api_mode: chat_completions
|
||||||
```
|
```
|
||||||
@@ -234,10 +237,13 @@ The following MUST be identical across ALL profiles:
|
|||||||
- When main config uses `api_key_env`, sub-agents automatically use it
|
- When main config uses `api_key_env`, sub-agents automatically use it
|
||||||
- This means key rotation only touches ONE vault secret (`LITELLM_API_KEY`)
|
- This means key rotation only touches ONE vault secret (`LITELLM_API_KEY`)
|
||||||
|
|
||||||
### Rule 5: Main Config Base URL
|
### Rule 5: Main Config Base URL (UPDATED 2026-08-09)
|
||||||
|- Use direct IP: `http://192.168.68.116/v1`
|
|- Use the authenticated LiteLLM path: `http://192.168.68.116/litellm/v1` (canonical, captain-approved migration)
|
||||||
|
|- Legacy `http://192.68.68.116/v1` also works — nginx fronts BOTH paths with key auth
|
||||||
|
(verified 2026-08-09: 401 without key, 200 with key, on both /v1 and /litellm/v1)
|
||||||
|
|- Both locations have `proxy_read_timeout 600s` (verified in harness-nginx nginx.conf) —
|
||||||
|
the old "60s timeout on /litellm/" claim was stale and is retracted
|
||||||
|- NOT the NetBird URL (`litellm.sysloggh.net`) — can cause 502 when NetBird is down
|
|- NOT the NetBird URL (`litellm.sysloggh.net`) — can cause 502 when NetBird is down
|
||||||
|- NOT the old path (`/litellm/v1`) — nginx now routes `/v1` directly
|
|
||||||
|
|
||||||
### Rule 6: max_tokens Is Required (Thermal Safety)
|
### Rule 6: max_tokens Is Required (Thermal Safety)
|
||||||
- **Every Hermes config MUST set `model.max_tokens: 4096`** — this is non-negotiable
|
- **Every Hermes config MUST set `model.max_tokens: 4096`** — this is non-negotiable
|
||||||
@@ -258,7 +264,7 @@ The following MUST be identical across ALL profiles:
|
|||||||
fall back to other GPUs if Strix gets hot. Both `compression.model` and `auxiliary.compression.model`
|
fall back to other GPUs if Strix gets hot. Both `compression.model` and `auxiliary.compression.model`
|
||||||
MUST be `syslog-auto`.
|
MUST be `syslog-auto`.
|
||||||
- All auxiliary services MUST use identical routing:
|
- All auxiliary services MUST use identical routing:
|
||||||
- `base_url: http://192.168.68.116/v1` (Rule 5: `/v1`, NOT `/litellm/v1`)
|
- `base_url: http://192.168.68.116/litellm/v1` (Rule 5, 2026-08-09: canonical authenticated; `/v1` also OK)
|
||||||
- `api_key_env: LITELLM_API_KEY`
|
- `api_key_env: LITELLM_API_KEY`
|
||||||
- **Do NOT use `syslog-auto`** for auxiliary tasks — it routes unpredictably
|
- **Do NOT use `syslog-auto`** for auxiliary tasks — it routes unpredictably
|
||||||
- **Compression on Strix Halo**: The strix-moe alias routes to Strix Halo
|
- **Compression on Strix Halo**: The strix-moe alias routes to Strix Halo
|
||||||
@@ -318,10 +324,13 @@ verify ALL FOUR of these against the live config. They are the only root causes
|
|||||||
1. **max_context_window correct?** — BOTH `compression.max_context_window` AND `context.max_context_window`
|
1. **max_context_window correct?** — BOTH `compression.max_context_window` AND `context.max_context_window`
|
||||||
MUST be `131072` (all GPUs are 128K). A value of `262144` causes instability near 100K and must NOT be used.
|
MUST be `131072` (all GPUs are 128K). A value of `262144` causes instability near 100K and must NOT be used.
|
||||||
~83K instead of ~170K. Check: `grep -n max_context_window ~/.hermes/config.yaml`
|
~83K instead of ~170K. Check: `grep -n max_context_window ~/.hermes/config.yaml`
|
||||||
2. **base_url uses /v1 NOT /litellm/v1?** — `custom_providers[0].base_url`, `delegation.base_url`,
|
2. **base_url uses authenticated path?** — `custom_providers[0].base_url`, `delegation.base_url`,
|
||||||
and ALL `auxiliary.*.base_url` MUST be `http://192.168.68.116/v1` (Rule 5). nginx `/litellm/`
|
and ALL `auxiliary.*.base_url` MUST be `http://192.168.68.116/litellm/v1` (Rule 5, canonical)
|
||||||
has a 60s default timeout → 504 on any inference >60s; `/v1/` has 600s.
|
or `http://192.168.68.116/v1` (legacy, still authenticated via nginx). BOTH verified 200 with
|
||||||
Check: `grep -n 'litellm/v1' ~/.hermes/config.yaml` (must return NOTHING)
|
key + 600s proxy_read_timeout on 2026-08-09. Never bare `:4000` direct.
|
||||||
|
Check: `grep -nE 'base_url: http://192.168.68.116(:4000)?/v1' ~/.hermes/config.yaml` — the
|
||||||
|
ONLY paths allowed are `/v1` or `/litellm/v1` (both via nginx :80).
|
||||||
|
`:4000` or missing `litellm/v1`/`v1` prefix = violation.
|
||||||
3. **LITELLM_API_KEY valid?** — The key must be a real LiteLLM key (`sk-` + 64 hex, 67 chars).
|
3. **LITELLM_API_KEY valid?** — The key must be a real LiteLLM key (`sk-` + 64 hex, 67 chars).
|
||||||
Malformed values (e.g. `sk-_SWAl_Vu_…`, 47 chars) return 401 → DeepSeek fallback.
|
Malformed values (e.g. `sk-_SWAl_Vu_…`, 47 chars) return 401 → DeepSeek fallback.
|
||||||
Verify: `curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $LITELLM_API_KEY" http://192.168.68.116/v1/models` (must be 200)
|
Verify: `curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $LITELLM_API_KEY" http://192.168.68.116/v1/models` (must be 200)
|
||||||
@@ -396,6 +405,15 @@ curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $K" http://192.
|
|||||||
- **Audit script**: Run `python3 /root/prose-contracts/audit-hermes-config.py <config.yaml>`
|
- **Audit script**: Run `python3 /root/prose-contracts/audit-hermes-config.py <config.yaml>`
|
||||||
before and after any config change to catch this and all other rule violations.
|
before and after any config change to catch this and all other rule violations.
|
||||||
|
|
||||||
|
### Rule 15: MCP Endpoint and Header Validation (ADDED 2026-08-07)
|
||||||
|
- Every MCP server entry must point at the correct endpoint:
|
||||||
|
- ra-h-os = http://192.168.68.65:3100/mcp
|
||||||
|
- litellm = https://litellm.sysloggh.net/mcp
|
||||||
|
- MCP entries must carry a REAL key value in the header.
|
||||||
|
- Avoid using env-var names like LITELLM_API_KEY in the header; they do not resolve for MCP
|
||||||
|
endpoints and result in "Malformed API Key" floods.
|
||||||
|
- Ensure the header value is the actual key (e.g., `sk-...`).
|
||||||
|
|
||||||
## Execution
|
## Execution
|
||||||
|
|
||||||
1. **Check current config** — Read the target agent's config.yaml
|
1. **Check current config** — Read the target agent's config.yaml
|
||||||
|
|||||||
@@ -34,11 +34,12 @@ Syslog is migrating away from **unauthenticated direct access** to the shared in
|
|||||||
|
|
||||||
| Path | Auth | Status |
|
| Path | Auth | Status |
|
||||||
|------|------|--------|
|
|------|------|--------|
|
||||||
| `http://192.168.68.116/v1` | None (direct) | ❌ **DEPRECATED** — being phased out |
|
| `http://192.168.68.116/v1` | Bearer `sk-*` key (nginx-fronted) | ✅ **VALID** — authenticated via nginx :80 (verified 2026-08-09: 401 without key, 200 with) |
|
||||||
| `http://192.168.68.116/litellm/v1/responses` | Bearer `sk-*` key | ✅ **CURRENT** — authenticated LiteLLM proxy |
|
| `http://192.168.68.116/litellm/v1` | Bearer `sk-*` key (nginx-fronted) | ✅ **CURRENT / CANONICAL** — captain-approved migration target; 600s proxy_read_timeout (verified) |
|
||||||
|
| `http://192.168.68.116:4000/v1` | Bearer `sk-*` key (direct container) | ❌ **FORBIDDEN** — bypasses nginx; port 4000 direct is not a config path |
|
||||||
|
|
||||||
All harness/litellm providers MUST use the authenticated `/litellm/v1/responses` path.
|
All harness/litellm providers MUST use an authenticated nginx-fronted path (`/litellm/v1` canonical, `/v1` legacy-valid).
|
||||||
Any `base_url` pointing to bare `/v1` on 192.168.68.116 is a **migration violation**.
|
Any `base_url` pointing at `:4000` or a bare IP without nginx is a **migration violation**.
|
||||||
|
|
||||||
### 🔥 CRITICAL: Double-Path Bug (2026-07-10)
|
### 🔥 CRITICAL: Double-Path Bug (2026-07-10)
|
||||||
|
|
||||||
@@ -190,22 +191,22 @@ litellm_settings:
|
|||||||
|-------|-----|-----|---------------|------------|--------|-----------------|---------------|
|
|-------|-----|-----|---------------|------------|--------|-----------------|---------------|
|
||||||
| Tanko | 112 | .122 | `tanko` | Infisical vault | ✅ Fixed | `infisical run` | 20:17 UTC Jul 5 |
|
| Tanko | 112 | .122 | `tanko` | Infisical vault | ✅ Fixed | `infisical run` | 20:17 UTC Jul 5 |
|
||||||
| Mumuni | 100 (abiba) | .24 | `mumuni` | Infisical vault | ✅ Fixed | Pi Hermes gateway | 2026-07-27 |
|
| Mumuni | 100 (abiba) | .24 | `mumuni` | Infisical vault | ✅ Fixed | Pi Hermes gateway | 2026-07-27 |
|
||||||
| Koby | 111 | ? | `koby` | Infisical vault | ✅ Fixed | `infisical run` | 23:30 UTC Jul 5 |
|
| Koby | 111 | .129 | `koby` | Infisical vault | ✅ Fixed | `infisical run` | 23:30 UTC Jul 5 |
|
||||||
| Koonimo | 113 | ? | `koonimo` | Infisical vault | ✅ Fixed | `infisical run` (migrated 2026-07-11) | 2026-07-11 |
|
| Koonimo | 113 | .114 | `koonimo` | Infisical vault | ✅ Fixed | `infisical run` (migrated 2026-07-11) | 2026-08-09 |
|
||||||
| Abiba | 100 | .65 | `abiba-pi` | Infisical vault | ✅ N/A (pi native) | — | 19:44 UTC Jul 5 |
|
| Abiba | 100 | .65 | `abiba-pi` | Infisical vault | ✅ N/A (pi native) | — | 19:44 UTC Jul 5 |
|
||||||
| Kagenz0 | 105 | ? | — | — | ❌ DOWN | — | 19:14 EDT Jul 4 |
|
| Kagenz0 | 105 | .14 | — | — | ❌ DOWN | — | 19:14 EDT Jul 4 |
|
||||||
|
|
||||||
> **Note**: CT hostnames (tdunna, baggy) differ from agent identities (koby, koonimo).
|
> **Note**: CT hostnames (tdunna, baggy) differ from agent identities (koby, koonimo).
|
||||||
> LiteLLM key aliases use agent identity, not CT hostname.
|
> LiteLLM key aliases use agent identity, not CT hostname.
|
||||||
|
|
||||||
### Migration Status: Authenticated Path
|
### Migration Status: Authenticated Path
|
||||||
|
|
||||||
| Agent | `/litellm/v1/responses` | Deprecated `/v1` | Status |
|
| Agent | `/litellm/v1` | Legacy `/v1` | Status |
|
||||||
|-------|--------------------------|--------------------|--------|
|
|-------|--------------|-------------|--------|
|
||||||
| Mumuni | ✅ 5 sections | 0 | ✅ Authenticated |
|
| Mumuni | ✅ harness provider | ✅ auxiliary on /v1 (valid) | ✅ Authenticated (verified 2026-08-09) |
|
||||||
| Tanko | ⚠️ No SSH access | — | Needs check |
|
| Tanko | ✅ 5 sections | 0 | ✅ Migrated 2026-08-08, keys 200 |
|
||||||
| Koby | ⚠️ No route to host | — | Needs check |
|
| Koby | ✅ custom provider (harness name) | — | ✅ External DeepSeek primary (intentional) |
|
||||||
| Koonimo | ⚠️ Connection timed out | — | Needs check |
|
| Koonimo | ✅ .114 (baggy) | — | ✅ 128K context applied 2026-08-09 |
|
||||||
|
|
||||||
### Systemd Service Pattern (2026-07-11 — vault migration)
|
### Systemd Service Pattern (2026-07-11 — vault migration)
|
||||||
|
|
||||||
|
|||||||
@@ -7,13 +7,15 @@ description: >
|
|||||||
from nvidia-smi (.8, .110) and amdgpu_top (.15). LiteLLM metrics
|
from nvidia-smi (.8, .110) and amdgpu_top (.15). LiteLLM metrics
|
||||||
via existing /metrics Prometheus endpoint.
|
via existing /metrics Prometheus endpoint.
|
||||||
|
|
||||||
DEPLOYMENT STATUS (2026-07-09):
|
DEPLOYMENT STATUS (2026-08-09):
|
||||||
✅ Core stack deployed: Prometheus + Grafana + pve/node/docker exporters
|
✅ Core stack deployed: Prometheus + Grafana + pve/node/docker exporters
|
||||||
(via proxmox-monitor contract). Grafana at :3001, 5 scrape targets active.
|
(via proxmox-monitor contract). Grafana at :3001, all scrape targets active.
|
||||||
❌ GPU exporters NOT deployed: gpu-exporter crash-loops on .15,
|
✅ GPU exporters DEPLOYED: all 3 GPU hosts (.8/.110/.15) run exporters on
|
||||||
NVIDIA sidecar exporters (.8/.110:9400) never installed.
|
:9400 (nvidia_gpu_exporter / amdgpu exporter) — verified 200 on 2026-08-09.
|
||||||
Router falls back to direct GPU /health probes.
|
✅ LiteLLM /metrics scraping live (success_callback: prometheus; auth via
|
||||||
⚠️ This contract is target-state aspirational — not as-built.
|
master key) + Alertmanager + Zulip bridge (alerts-infra) added 2026-08-09.
|
||||||
|
⚠️ This contract is target-state aspirational — but GPU export + alerting
|
||||||
|
are now as-built (verified 2026-08-09).
|
||||||
As-built GPU monitoring is via gpu-monitor contract (port 9100 poll).
|
As-built GPU monitoring is via gpu-monitor contract (port 9100 poll).
|
||||||
version: 1.0.0
|
version: 1.0.0
|
||||||
---
|
---
|
||||||
|
|||||||
@@ -18,6 +18,15 @@ description: >
|
|||||||
Designed as a reusable contract for any Syslog agent.
|
Designed as a reusable contract for any Syslog agent.
|
||||||
|
|
||||||
Source of truth: gpu-fleet.prose.md
|
Source of truth: gpu-fleet.prose.md
|
||||||
|
|
||||||
|
## Monitoring / Alerting (as-built 2026-08-09)
|
||||||
|
|
||||||
|
- LiteLLM /metrics scrape job: requires `litellm_settings.success_callback: [prometheus]`
|
||||||
|
(failure_callback alone does NOT mount /metrics — verified 2026-08-09).
|
||||||
|
Scraped by Prometheus with Bearer master key; endpoint returns 307 → /metrics/.
|
||||||
|
- Alertmanager (harness-alertmanager :9093) + zulip-bridge (:9102) deliver
|
||||||
|
firing alerts to #agent-hub > alerts-infra via abiba-bot. Added 2026-08-09.
|
||||||
|
- Prometheus node job covers ALL 6 PVE nodes (.4/.5/.6/.9/.12/.15:9100).
|
||||||
---
|
---
|
||||||
|
|
||||||
## Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU)
|
## Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU)
|
||||||
|
|||||||
+152
-46
@@ -2,69 +2,175 @@
|
|||||||
kind: pattern
|
kind: pattern
|
||||||
name: memory-fixer
|
name: memory-fixer
|
||||||
description: >
|
description: >
|
||||||
Auto-fix low-hanging fruit in the graph. No judgment calls — only deterministic Level 1 operations.
|
Auto-fix low-hanging fruit in the RA-H OS knowledge graph. No judgment calls — only deterministic Level 1 operations.
|
||||||
Escalate anything that needs Kwame's input.
|
Escalate anything that needs Kwame's input. Executes confirmed Kwame decisions to completion (state + updated_at).
|
||||||
version: 1.1.0
|
version: 2.0.0
|
||||||
---
|
---
|
||||||
|
|
||||||
# Memory Fixer
|
# Memory Fixer
|
||||||
|
|
||||||
|
> **Canonical copy:** `/root/.hermes/contracts/memory-fixer-v3.md` (used by the `memory-fixer-daily` cron job). This file is the institutional record of the same contract. When the two diverge, treat the v3 source in `/root/.hermes/contracts/` as executable truth.
|
||||||
|
|
||||||
## Purpose
|
## Purpose
|
||||||
Auto-fix low-hanging fruit in the graph. No judgment calls — only deterministic Level 1 operations. Escalate anything that needs Kwame's input.
|
Auto-fix low-hanging fruit in the graph. No judgment calls — only deterministic Level 1 operations. Escalate anything that needs Kwame's input. When Kwame replies to an escalation, **execute the decision to completion** (update state and timestamps), never leaving a node in review-pending forever.
|
||||||
|
|
||||||
## Level 0 Auto-Deletes (Allowed Without Approval)
|
**Type:** Write-only (Level 1 fixes only)
|
||||||
Ephemeral heartbeat and log nodes that violate "Logs NEVER go in the graph":
|
**Scope:** RA-H OS knowledge graph (192.168.68.65)
|
||||||
|
**Schedule:** Daily at 8 AM ET
|
||||||
|
**Escalation:** Level 2+ to Kwame as task items
|
||||||
|
|
||||||
- `[LITELLM-HEALTH]`, `[GPU-SELF-HEAL]`, `[PM2-SELF-HEAL]`
|
## Key Design Decision
|
||||||
- `[PROXMOX-MONITOR]`, `[GPU-MONITOR]`, `[INFRA-MONITOR]`, `[AGENT-HEALTH]`, `[DISK-GC]`
|
|
||||||
- `[WAL]` entries older than 30 days
|
|
||||||
|
|
||||||
**Condition:** node must be an orphan (no edges). Deleting a connected node risks breaking other nodes.
|
The `updateNode` tool's `metadata` field performs a **restricted merge** — the `state` key only accepts `'processed'` or `'not_processed'`. Additionally, new metadata keys cannot be added via the merge.
|
||||||
|
|
||||||
**Method:** direct SQLite on `.65` (MCP has no delete tool):
|
**Solution:** Use the `description` field to tag stale nodes with review actions, since `description` is a simple string overwritable via `updateNode`.
|
||||||
```bash
|
|
||||||
ssh root@192.168.68.65 "sqlite3 /root/.local/share/RA-H/db/rah.sqlite \"
|
|
||||||
DELETE FROM nodes WHERE id IN (
|
|
||||||
SELECT id FROM nodes WHERE id NOT IN (SELECT from_node_id FROM edges)
|
|
||||||
AND id NOT IN (SELECT to_node_id FROM edges)
|
|
||||||
AND title LIKE '[LITELLM-HEALTH]%' -- add more prefixes as needed
|
|
||||||
);\""
|
|
||||||
```
|
|
||||||
|
|
||||||
## Level 1 Auto-Fixes (No Judgment Required)
|
**Tag Format:** `[REVIEW: action] original description text...`
|
||||||
|
|
||||||
### 1. Missing `type` Field
|
Where `action` is one of:
|
||||||
For nodes with content but no `metadata.type`:
|
- `archive` — node is stale and should be archived
|
||||||
- Title contains "Proxmox" or "infrastructure" → `type: infrastructure`
|
- `refresh` — node is stale and should be refreshed (infrastructure)
|
||||||
- Title contains "skill" or "how to" or "guide" → `type: skill`
|
- `keep` — node has been confirmed as current
|
||||||
- Title contains "doc" or "template" or "brand" → `type: documentation`
|
- `merge` — node is a duplicate candidate
|
||||||
- Title starts with "WAL:" or "TASK:" → `type: note`
|
|
||||||
- Title starts with "[LEARN]" → `type: documentation`
|
|
||||||
- Otherwise → `type: note` (default)
|
|
||||||
|
|
||||||
### 2. Missing `tenant` / `namespace`
|
**Query for finding review-tagged nodes:**
|
||||||
For any node with NULL tenant or namespace:
|
|
||||||
```sql
|
```sql
|
||||||
UPDATE nodes
|
SELECT id, title, description
|
||||||
SET metadata = json_set(
|
FROM nodes
|
||||||
COALESCE(metadata, '{}'),
|
WHERE description LIKE '[REVIEW:%';
|
||||||
'$.tenant', 'syslogsolution',
|
|
||||||
'$.namespace', 'syslogsolution'
|
|
||||||
)
|
|
||||||
WHERE json_extract(metadata, '$.tenant') IS NULL
|
|
||||||
OR json_extract(metadata, '$.namespace') IS NULL;
|
|
||||||
```
|
```
|
||||||
|
|
||||||
### 3. Staleness State Transitions
|
## Level 1 Auto-Fixes (No Kwame Decision Needed)
|
||||||
Using the type-based windows from the memory-monitor contract:
|
|
||||||
- Nodes stale > their window → transition to `state: review_pending`
|
### 1. Missing `type` Auto-Classification
|
||||||
- Nodes in `review_pending` for >7 days → escalate to Kwame (Level 2)
|
|
||||||
|
```sql
|
||||||
|
SELECT id, title,
|
||||||
|
CASE
|
||||||
|
WHEN title LIKE '%infrastructure%' OR title LIKE '%proxmox%' OR title LIKE '%setup%' THEN 'infrastructure'
|
||||||
|
WHEN title LIKE '%skill%' OR title LIKE '%how to%' OR title LIKE '%guide%' THEN 'skill'
|
||||||
|
WHEN title LIKE '%doc%' OR title LIKE '%template%' OR title LIKE '%brand%' THEN 'documentation'
|
||||||
|
WHEN title LIKE 'WAL:%' OR title LIKE 'TASK:%' THEN 'note'
|
||||||
|
WHEN title LIKE '%[LEARN]%' THEN 'documentation'
|
||||||
|
ELSE 'note'
|
||||||
|
END as auto_type
|
||||||
|
FROM nodes
|
||||||
|
WHERE json_extract(metadata, '$.type') IS NULL;
|
||||||
|
```
|
||||||
|
|
||||||
|
### 2. Missing `namespace` Auto-Population
|
||||||
|
|
||||||
|
```sql
|
||||||
|
SELECT id, title, json_extract(metadata, '$.tenant') as tenant
|
||||||
|
FROM nodes
|
||||||
|
WHERE json_extract(metadata, '$.namespace') IS NULL;
|
||||||
|
```
|
||||||
|
|
||||||
|
### 3. Staleness Review Tagging
|
||||||
|
|
||||||
|
Using the type-based windows from the memory-monitor contract, tag nodes stale beyond their window. **Only process a maximum of 10 nodes per run** to avoid overwhelming Kwame. Prioritize infrastructure first, then dynamic, then ephemeral.
|
||||||
|
|
||||||
|
**Exclusion Rules:**
|
||||||
|
- Nodes with `state` = `review_pending`, `deprecated`, `archived`, or `not_processed` are NOT processed
|
||||||
|
- Nodes whose `description` already starts with `[REVIEW:` are NOT re-processed
|
||||||
|
|
||||||
|
```sql
|
||||||
|
SELECT id, title, json_extract(metadata, '$.type') as node_type,
|
||||||
|
CAST(julianday('now') - julianday(updated_at) AS INTEGER) as days_stale,
|
||||||
|
CASE
|
||||||
|
WHEN json_extract(metadata, '$.type') IN ('infrastructure', 'deployment', 'system', 'system-health') THEN 'refresh'
|
||||||
|
WHEN json_extract(metadata, '$.type') IN ('business', 'philosophy', 'research', 'learning', 'learn', 'investigation', 'analysis', 'project') THEN 'refresh'
|
||||||
|
ELSE 'archive'
|
||||||
|
END as suggested_action
|
||||||
|
FROM nodes
|
||||||
|
WHERE updated_at < datetime('now',
|
||||||
|
CASE
|
||||||
|
WHEN json_extract(metadata, '$.type') IN ('infrastructure', 'deployment', 'system', 'system-health') THEN '-14 days'
|
||||||
|
WHEN json_extract(metadata, '$.type') IN ('skill', 'documentation', 'template', 'protocol-enforcement', 'prd', 'architecture') THEN '-90 days'
|
||||||
|
WHEN json_extract(metadata, '$.type') IN ('note', 'wal', 'WAL', 'task', 'TASK', 'event') THEN '-30 days'
|
||||||
|
WHEN json_extract(metadata, '$.type') IN ('business', 'philosophy', 'research', 'learning', 'learn', 'investigation', 'analysis', 'project') THEN '-120 days'
|
||||||
|
WHEN json_extract(metadata, '$.type') IN ('deprecated-relay', 'audit', 'audit-report', 'incident', 'incident-report') THEN '-3650 days'
|
||||||
|
ELSE '-45 days'
|
||||||
|
END
|
||||||
|
)
|
||||||
|
AND json_extract(metadata, '$.state') NOT IN ('review_pending', 'deprecated', 'archived', 'not_processed')
|
||||||
|
AND (description IS NULL OR description NOT LIKE '[REVIEW:%')
|
||||||
|
ORDER BY days_stale ASC
|
||||||
|
LIMIT 10;
|
||||||
|
```
|
||||||
|
|
||||||
|
For each identified node, call `updateNode(id, { description: "[REVIEW: action] " + originalDescription })`.
|
||||||
|
|
||||||
## Level 2 Escalations (Kwame Decision Required)
|
## Level 2 Escalations (Kwame Decision Required)
|
||||||
1. **Nodes in `review_pending` >7 days** — Archive, refresh, or keep?
|
|
||||||
2. **Orphan Nodes >90 days old** — Delete or Connect?
|
1. **Stale nodes** flagged with `[REVIEW: …]` — Archive, refresh, or keep?
|
||||||
3. **Potential Duplicate Nodes** — Same title or >70% overlap. Merge or Keep?
|
2. **Duplicate Nodes** (same title or >70% title overlap) — Merge or keep?
|
||||||
4. **Conflicting Metadata** — Content suggests one tenant but metadata says another.
|
3. **Orphan Nodes >90 days old** — Archive or connect?
|
||||||
|
|
||||||
|
## Reporting Format
|
||||||
|
|
||||||
|
The fixer reports to Kwame via this Zulip DM:
|
||||||
|
|
||||||
|
```
|
||||||
|
🦅 Memory Fixer — [HH:MM UTC]
|
||||||
|
|
||||||
|
Level 1 fixes applied:
|
||||||
|
- Missing type: X nodes classified
|
||||||
|
- Missing namespace: Y nodes populated
|
||||||
|
|
||||||
|
Stale nodes needing review (max 10):
|
||||||
|
1. [Node #XXX] Title — X days stale, SUGGEST: refresh
|
||||||
|
2. [Node #YYY] Title — Y days stale, SUGGEST: archive
|
||||||
|
...
|
||||||
|
|
||||||
|
Duplicates needing decision:
|
||||||
|
1. [Node #AAA] vs [Node #BBB] — Same title
|
||||||
|
|
||||||
|
Orphans >90 days:
|
||||||
|
1. [Node #EEE] Title — X days stale, orphaned
|
||||||
|
|
||||||
|
Reply with:
|
||||||
|
- "archive #XXX, #YYY" to mark for archive
|
||||||
|
- "archive all" to archive all stale nodes listed
|
||||||
|
- "keep #XXX" to confirm a node is current
|
||||||
|
- "merge #AAA into #BBB" to merge duplicates
|
||||||
|
- "refresh #XXX" to mark as current
|
||||||
|
```
|
||||||
|
|
||||||
|
## Execution on Next Run
|
||||||
|
|
||||||
|
The fixer reads Kwame's previous response and **executes the decision to completion** — it must not leave a node in review-pending forever. Tagging alone is NOT enough; each confirmed decision must also update `state` and `updated_at` so the node drops out of the stale window on the next run.
|
||||||
|
|
||||||
|
> ⚠️ `updateNode` cannot set `state` to non-standard values (restricted to `processed`/`not_processed`) and cannot add metadata keys. For state transitions and `updated_at` bumps, use **direct SSH + SQLite** on the bridge host:
|
||||||
|
> ```bash
|
||||||
|
> ssh root@192.168.68.65 "sqlite3 /root/.local/share/RA-H/db/rah.sqlite \"UPDATE nodes SET metadata = json_set(metadata, '$.state', '<state>'), updated_at = datetime('now') WHERE id = <id>;\""
|
||||||
|
> ```
|
||||||
|
> Use `updateNode` only for description/source/title/link edits.
|
||||||
|
|
||||||
|
Decision → completed action mapping:
|
||||||
|
|
||||||
|
| Kwame reply | Description change | State | `updated_at` |
|
||||||
|
|---|---|---|---|
|
||||||
|
| `archive #XXX` | replace `[REVIEW: archive] ` → `[ARCHIVED] ` prefix | `archived` | bumped to now |
|
||||||
|
| `keep #XXX` / `refresh #XXX` | **clear the `[REVIEW: …]` tag entirely** | `active` | bumped to now |
|
||||||
|
| `merge #AAA into #BBB` | set `[REVIEW: merge_into #BBB]` on #AAA, then follow manual merge workflow | handled manually | bumped to now |
|
||||||
|
| `archive all` | apply the archive row to every node listed in the prior report | `archived` | bumped to now |
|
||||||
|
|
||||||
|
**Why `updated_at` must be bumped (critical):** the Level-1 staleness query keys off `updated_at < now - window`. If the fixer clears the tag but leaves a stale `updated_at`, the node is immediately re-flagged on the very next run and the cycle repeats forever. Bumping `updated_at` to now pushes the node back to the front of the window.
|
||||||
|
|
||||||
|
**Exclusion after action:** once an action is applied, the node's description no longer starts with `[REVIEW:` (archive → `[ARCHIVED]`, refresh/keep → original text), so it is not re-processed.
|
||||||
|
|
||||||
|
After all actions are applied, verify with:
|
||||||
|
```sql
|
||||||
|
SELECT id, json_extract(metadata, '$.state') FROM nodes WHERE description LIKE '[REVIEW:%';
|
||||||
|
```
|
||||||
|
The result must be 0 rows when all decisions are executed. Report what was done.
|
||||||
|
|
||||||
|
## Checks
|
||||||
|
|
||||||
|
- **State integrity:** archived nodes have `state: archived` + `[ARCHIVED]` prefix; kept nodes are `state: active` without a `[REVIEW:]` tag.
|
||||||
|
- **No review-pending forever:** after executing Kwame's decisions, `[REVIEW:%` node count must be 0.
|
||||||
|
- **Timestamps:** every executed decision bumps `updated_at`, so the node exits the stale window on the next run.
|
||||||
|
|
||||||
## Logging
|
## Logging
|
||||||
Every Level 1 fix logged to `~/.hermes/logs/memory-fixer/YYYY-MM-DD.md`
|
Every Level 1 fix logged to `~/.hermes/logs/memory-fixer/YYYY-MM-DD.md`
|
||||||
|
|||||||
+15
-6
@@ -2,21 +2,30 @@
|
|||||||
kind: responsibility
|
kind: responsibility
|
||||||
name: pm2-self-heal
|
name: pm2-self-heal
|
||||||
description: >
|
description: >
|
||||||
Monitors critical PM2 processes (abiba-zulip, abiba-telegram) and
|
Monitors critical PM2 processes (abiba-zulip, abiba-telegram, gitea-runner,
|
||||||
auto-restarts any that are stopped or errored. Logs every action to
|
spoton-service, zulip-watchdog) and auto-restarts any that are stopped or
|
||||||
the knowledge graph and alerts the owner via Zulip DM on failures.
|
errored. Logs every action to the knowledge graph and alerts the owner via
|
||||||
|
Zulip DM on failures.
|
||||||
CRITICAL: Never restart abiba-zulip — it runs this contract.
|
CRITICAL: Never restart abiba-zulip — it runs this contract.
|
||||||
|
AS-BUILT 2026-08-09 (captain ruling, ecosystem is authoritative):
|
||||||
|
gpu-monitor is systemd-managed (gpu-monitor.service) — NOT PM2;
|
||||||
|
gpu-watchdog decommissioned (function folded into gpu-monitor.service);
|
||||||
|
gitea-runner KEPT (online in PM2); abiba-zulip KEPT (online 4d+, the
|
||||||
|
2026-07-04 'removed/decommissioned' note was stale and is removed).
|
||||||
---
|
---
|
||||||
|
|
||||||
## Maintains
|
## Maintains
|
||||||
|
|
||||||
- abiba-telegram: { status: "online", uptime: string, restarts: number }
|
- abiba-telegram: { status: "online", uptime: string, restarts: number }
|
||||||
- gpu-watchdog: { status: "online", uptime: string, restarts: number }
|
- abiba-zulip: { status: "online", uptime: string, restarts: number }
|
||||||
- gpu-monitor: { status: "online", uptime: string, restarts: number }
|
|
||||||
- gitea-runner: { status: "online", uptime: string, restarts: number }
|
- gitea-runner: { status: "online", uptime: string, restarts: number }
|
||||||
|
- spoton-service: { status: "online", uptime: string, restarts: number }
|
||||||
|
- zulip-watchdog: { status: "online", uptime: string, restarts: number }
|
||||||
- last_check: timestamp
|
- last_check: timestamp
|
||||||
|
|
||||||
> **Note (2026-07-04):** `abiba-zulip` removed — Zulip extension decommissioned.
|
> **Note (2026-07-04, SUPERSEDED 2026-08-09):** `abiba-zulip` remains ONLINE and
|
||||||
|
> is monitored — the decommission note was stale (process re-added; do not treat
|
||||||
|
> it as removed).
|
||||||
|
|
||||||
## Continuity
|
## Continuity
|
||||||
|
|
||||||
|
|||||||
@@ -199,7 +199,14 @@ Check `platforms.zulip.state`: `connected` ✅ | `disconnected` ❌ | `error`
|
|||||||
ssh root@<CT> "ps aux | grep 'gateway run' | grep -v grep"
|
ssh root@<CT> "ps aux | grep 'gateway run' | grep -v grep"
|
||||||
```
|
```
|
||||||
|
|
||||||
Gateway PID should exist with uptime > 60s. **Dual-gateway detection**: if more than one `gateway run` process is found, the gateway has a collision (typically one `--force` and one `--replace` process). Kill the newer/duplicate process, then restart the remaining gateway via PM2 (`pm2 restart mumuni-zulip`). Check gateway log for "Gateway running with 2 platform(s)" (not 1) to confirm Zulip reloaded.
|
Gateway PID should exist with uptime > 60s. **Dual-gateway detection**: if more than one `gateway run` process is found, the gateway has a collision (typically one `--force` and one `--replace` process). Kill the newer/duplicate process, then restart the remaining gateway per-agent (parameterized 2026-08-09, captain ruling):
|
||||||
|
|
||||||
|
| Agent | Restart command | Notes |
|
||||||
|
|-------|-----------------|-------|
|
||||||
|
| Mumuni (.24) | `pm2 restart abiba-zulip` | Hermes gateway runs under PM2 as `abiba-zulip` |
|
||||||
|
| Tanko (.122) | `bash /opt/hermes-zulip-plugin/run.sh` (or the agent's systemd/user unit) | Tanko does NOT use PM2 — never run `pm2 restart mumuni-zulip` for Tanko (process does not exist) |
|
||||||
|
|
||||||
|
Check gateway log for "Gateway running with 2 platform(s)" (not 1) to confirm Zulip reloaded.
|
||||||
|
|
||||||
**B3: Heartbeat Verification**
|
**B3: Heartbeat Verification**
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user