Compare commits

..
Author SHA1 Message Date
mumuni-bot b799d46596 fix: redirect pm2/zulip health-check logging to Gitea, never graph (hard rule)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
- pm2-self-heal.prose.md: description said 'logs every action to the
  knowledge graph', contradicting its own execution step (line ~60) that
  says '(not knowledge graph - hard rule)'. Align description to Gitea.
- zulip-self-heal.prose.md: Reporting section said 'every cycle produces
  a knowledge graph node'. Contract is RETIRED; remove graph-node directive.
- Both now consistently log to SyslogSolution/health-logs, never the graph.
2026-08-13 00:09:40 +00:00
jerome 1d4f6c8ebb Merge pull request 'Ship: enforcement-rules-reality-fix (captain directive 2026-08-09)' (#45) from ship/enforcement-rules-reality-fix into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
Reviewed-on: #45
2026-08-10 16:38:05 +00:00
root 38a32f8b32 Ship: enforcement-rules-reality-fix (captain directive 2026-08-09)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-08-10 13:28:53 +00:00
jerome 96b0caa0b3 Merge pull request 'Ship: pm2-zulip contract fixes (captain ruling 2026-08-09)' (#44) from ship/pm2-zulip-contract-fixes into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
Reviewed-on: #44
2026-08-10 01:10:13 +00:00
jerome ea17f64bb4 Merge pull request 'Ship: monitoring contract fixes (2026-08-09)' (#43) from ship/monitoring-contract-fixes into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
Reviewed-on: #43
2026-08-10 01:08:07 +00:00
root c13a15acad Ship: pm2-zulip contract fixes (captain ruling 2026-08-09)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-08-09 22:11:27 +00:00
root 68f9ebff74 Ship: monitoring contract fixes (2026-08-09)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-08-09 21:57:44 +00:00
root 93fb0d8e1e feat(contracts): add MCP URL validation and verify virtual keys on LiteLLM
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 2s
2026-08-07 21:33:59 +00:00
mumuni-bot cb6efa81c3 Merge pull request 'fix(contracts): memory-fixer executes decisions to completion (v2.0.0)' (#42) from fix/memory-fixer-execution-v3 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-08-04 12:32:56 +00:00
7 changed files with 92 additions and 46 deletions
+35 -17
View File
@@ -5,8 +5,7 @@ description: >
Standard Hermes configuration template for Syslog Solution LLC agents. Standard Hermes configuration template for Syslog Solution LLC agents.
Enforces shared infrastructure setup (Firecrawl, SearXNG, local models, Enforces shared infrastructure setup (Firecrawl, SearXNG, local models,
RA-H OS MCP) while keeping agent-specific API keys and model choices. RA-H OS MCP) while keeping agent-specific API keys and model choices.
UPDATED 2026-07-16: Compression model is the stable alias `strix-moe` (NOT `ornith-1.0-35b`, UPDATED 2026-08-07: Added Rule 15 (MCP Validation) from the 2026-08-07 keyless-MCP incident.
which LiteLLM does not serve). All 3 GPUs verified at 128K (reduced from 256K 2026-07-17 for stability).
Added Rule 12 (Context-Issue Diagnostic) + Rule 13 (.env fallback enforcement) from the Added Rule 12 (Context-Issue Diagnostic) + Rule 13 (.env fallback enforcement) from the
2026-07-16 Mumuni root-cause investigation (WAL #1300). 2026-07-16 Mumuni root-cause investigation (WAL #1300).
UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo. RTX 3090 context verified at 128K. Infisical .env fallback required (Rule 3/13). UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo. RTX 3090 context verified at 128K. Infisical .env fallback required (Rule 3/13).
@@ -97,10 +96,14 @@ work immediately after restart.
model: model:
default: <agent_model> # e.g., strix-moe, qwen3.6-27B-code, syslog-auto default: <agent_model> # e.g., strix-moe, qwen3.6-27B-code, syslog-auto
provider: harness provider: harness
base_url: http://192.168.68.116/v1 base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated path; /v1 also OK
api_key_env: LITELLM_API_KEY # Injected via infisical run -- wrapper api_key_env: LITELLM_API_KEY # Injected via infisical run -- wrapper
max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation
context_length: 131072 # For syslog-auto (all GPUs at 128K for stability). context_length: 131072 # For syslog-auto (all GPUs at 128K for stability).
# ⚠️ MANDATORY: Hermes probes unknown models from 256K
# and falls back to 256K when /v1/models lacks a context
# field (llama-server does). Without this override, agents
# silently run syslog-auto at 256K (verified 2026-08-09).
# Set 65536 if using gemma-4-12b directly (tight VRAM). # Set 65536 if using gemma-4-12b directly (tight VRAM).
fallback_providers: fallback_providers:
@@ -143,7 +146,7 @@ compression:
# ─── Auxiliary Tasks (CONSISTENCY RULE) ─── # ─── Auxiliary Tasks (CONSISTENCY RULE) ───
# All auxiliary services MUST use identical model, base_url, and api_key_env: # All auxiliary services MUST use identical model, base_url, and api_key_env:
# model: gpu-light # stable alias (NOT raw "gemma-4-12b") # model: gpu-light # stable alias (NOT raw "gemma-4-12b")
# base_url: http://192.168.68.116/v1 # base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
# api_key_env: LITELLM_API_KEY # api_key_env: LITELLM_API_KEY
# Do NOT use syslog-auto for auxiliary tasks — it routes to the primary GPU. # Do NOT use syslog-auto for auxiliary tasks — it routes to the primary GPU.
# gpu-light = RTX 5070 (12B), freeing the Strix Halo for agent reasoning. # gpu-light = RTX 5070 (12B), freeing the Strix Halo for agent reasoning.
@@ -154,20 +157,20 @@ auxiliary:
vision: vision:
provider: harness provider: harness
model: gpu-light # stable alias for RTX 5070 (was raw gemma-4-12b) model: gpu-light # stable alias for RTX 5070 (was raw gemma-4-12b)
base_url: http://192.168.68.116/v1 base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
api_key_env: LITELLM_API_KEY api_key_env: LITELLM_API_KEY
timeout: 60 timeout: 60
download_timeout: 30 download_timeout: 30
web_extract: web_extract:
provider: harness provider: harness
model: gpu-light # stable alias for RTX 5070 model: gpu-light # stable alias for RTX 5070
base_url: http://192.168.68.116/v1 base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
api_key_env: LITELLM_API_KEY api_key_env: LITELLM_API_KEY
timeout: 30 timeout: 30
compression: compression:
provider: harness provider: harness
model: syslog-auto # MUST match compression.model above. Stable alias for Strix Halo (weighted pool). model: syslog-auto # MUST match compression.model above. Stable alias for Strix Halo (weighted pool).
base_url: http://192.168.68.116/v1 # Rule 5: /v1 NOT /litellm/v1 base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated path; /v1 also OK
api_key_env: LITELLM_API_KEY api_key_env: LITELLM_API_KEY
timeout: 300 # gpu-fleet: 300s for large-history summarization (was 60) timeout: 300 # gpu-fleet: 300s for large-history summarization (was 60)
@@ -177,14 +180,14 @@ auxiliary:
delegation: delegation:
model: gpu-dense # stable alias for RTX 3090 (was raw qwen3.6-27B-code) model: gpu-dense # stable alias for RTX 3090 (was raw qwen3.6-27B-code)
provider: harness provider: harness
base_url: http://192.168.68.116/v1 base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
api_key_env: LITELLM_API_KEY api_key_env: LITELLM_API_KEY
# ─── Custom Provider ─── # ─── Custom Provider ───
custom_providers: custom_providers:
- name: harness - name: harness
model: syslog-auto # weighted pool (default) model: syslog-auto # weighted pool (default)
base_url: http://192.168.68.116/v1 base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
api_key_env: LITELLM_API_KEY api_key_env: LITELLM_API_KEY
api_mode: chat_completions api_mode: chat_completions
``` ```
@@ -234,10 +237,13 @@ The following MUST be identical across ALL profiles:
- When main config uses `api_key_env`, sub-agents automatically use it - When main config uses `api_key_env`, sub-agents automatically use it
- This means key rotation only touches ONE vault secret (`LITELLM_API_KEY`) - This means key rotation only touches ONE vault secret (`LITELLM_API_KEY`)
### Rule 5: Main Config Base URL ### Rule 5: Main Config Base URL (UPDATED 2026-08-09)
|- Use direct IP: `http://192.168.68.116/v1` |- Use the authenticated LiteLLM path: `http://192.168.68.116/litellm/v1` (canonical, captain-approved migration)
|- Legacy `http://192.68.68.116/v1` also works — nginx fronts BOTH paths with key auth
(verified 2026-08-09: 401 without key, 200 with key, on both /v1 and /litellm/v1)
|- Both locations have `proxy_read_timeout 600s` (verified in harness-nginx nginx.conf) —
the old "60s timeout on /litellm/" claim was stale and is retracted
|- NOT the NetBird URL (`litellm.sysloggh.net`) — can cause 502 when NetBird is down |- NOT the NetBird URL (`litellm.sysloggh.net`) — can cause 502 when NetBird is down
|- NOT the old path (`/litellm/v1`) — nginx now routes `/v1` directly
### Rule 6: max_tokens Is Required (Thermal Safety) ### Rule 6: max_tokens Is Required (Thermal Safety)
- **Every Hermes config MUST set `model.max_tokens: 4096`** — this is non-negotiable - **Every Hermes config MUST set `model.max_tokens: 4096`** — this is non-negotiable
@@ -258,7 +264,7 @@ The following MUST be identical across ALL profiles:
fall back to other GPUs if Strix gets hot. Both `compression.model` and `auxiliary.compression.model` fall back to other GPUs if Strix gets hot. Both `compression.model` and `auxiliary.compression.model`
MUST be `syslog-auto`. MUST be `syslog-auto`.
- All auxiliary services MUST use identical routing: - All auxiliary services MUST use identical routing:
- `base_url: http://192.168.68.116/v1` (Rule 5: `/v1`, NOT `/litellm/v1`) - `base_url: http://192.168.68.116/litellm/v1` (Rule 5, 2026-08-09: canonical authenticated; `/v1` also OK)
- `api_key_env: LITELLM_API_KEY` - `api_key_env: LITELLM_API_KEY`
- **Do NOT use `syslog-auto`** for auxiliary tasks — it routes unpredictably - **Do NOT use `syslog-auto`** for auxiliary tasks — it routes unpredictably
- **Compression on Strix Halo**: The strix-moe alias routes to Strix Halo - **Compression on Strix Halo**: The strix-moe alias routes to Strix Halo
@@ -318,10 +324,13 @@ verify ALL FOUR of these against the live config. They are the only root causes
1. **max_context_window correct?** — BOTH `compression.max_context_window` AND `context.max_context_window` 1. **max_context_window correct?** — BOTH `compression.max_context_window` AND `context.max_context_window`
MUST be `131072` (all GPUs are 128K). A value of `262144` causes instability near 100K and must NOT be used. MUST be `131072` (all GPUs are 128K). A value of `262144` causes instability near 100K and must NOT be used.
~83K instead of ~170K. Check: `grep -n max_context_window ~/.hermes/config.yaml` ~83K instead of ~170K. Check: `grep -n max_context_window ~/.hermes/config.yaml`
2. **base_url uses /v1 NOT /litellm/v1?** — `custom_providers[0].base_url`, `delegation.base_url`, 2. **base_url uses authenticated path?** — `custom_providers[0].base_url`, `delegation.base_url`,
and ALL `auxiliary.*.base_url` MUST be `http://192.168.68.116/v1` (Rule 5). nginx `/litellm/` and ALL `auxiliary.*.base_url` MUST be `http://192.168.68.116/litellm/v1` (Rule 5, canonical)
has a 60s default timeout → 504 on any inference >60s; `/v1/` has 600s. or `http://192.168.68.116/v1` (legacy, still authenticated via nginx). BOTH verified 200 with
Check: `grep -n 'litellm/v1' ~/.hermes/config.yaml` (must return NOTHING) key + 600s proxy_read_timeout on 2026-08-09. Never bare `:4000` direct.
Check: `grep -nE 'base_url: http://192.168.68.116(:4000)?/v1' ~/.hermes/config.yaml` — the
ONLY paths allowed are `/v1` or `/litellm/v1` (both via nginx :80).
`:4000` or missing `litellm/v1`/`v1` prefix = violation.
3. **LITELLM_API_KEY valid?** — The key must be a real LiteLLM key (`sk-` + 64 hex, 67 chars). 3. **LITELLM_API_KEY valid?** — The key must be a real LiteLLM key (`sk-` + 64 hex, 67 chars).
Malformed values (e.g. `sk-_SWAl_Vu_…`, 47 chars) return 401 → DeepSeek fallback. Malformed values (e.g. `sk-_SWAl_Vu_…`, 47 chars) return 401 → DeepSeek fallback.
Verify: `curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $LITELLM_API_KEY" http://192.168.68.116/v1/models` (must be 200) Verify: `curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $LITELLM_API_KEY" http://192.168.68.116/v1/models` (must be 200)
@@ -396,6 +405,15 @@ curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $K" http://192.
- **Audit script**: Run `python3 /root/prose-contracts/audit-hermes-config.py <config.yaml>` - **Audit script**: Run `python3 /root/prose-contracts/audit-hermes-config.py <config.yaml>`
before and after any config change to catch this and all other rule violations. before and after any config change to catch this and all other rule violations.
### Rule 15: MCP Endpoint and Header Validation (ADDED 2026-08-07)
- Every MCP server entry must point at the correct endpoint:
- ra-h-os = http://192.168.68.65:3100/mcp
- litellm = https://litellm.sysloggh.net/mcp
- MCP entries must carry a REAL key value in the header.
- Avoid using env-var names like LITELLM_API_KEY in the header; they do not resolve for MCP
endpoints and result in "Malformed API Key" floods.
- Ensure the header value is the actual key (e.g., `sk-...`).
## Execution ## Execution
1. **Check current config** — Read the target agent's config.yaml 1. **Check current config** — Read the target agent's config.yaml
+14 -13
View File
@@ -34,11 +34,12 @@ Syslog is migrating away from **unauthenticated direct access** to the shared in
| Path | Auth | Status | | Path | Auth | Status |
|------|------|--------| |------|------|--------|
| `http://192.168.68.116/v1` | None (direct) | ❌ **DEPRECATED** — being phased out | | `http://192.168.68.116/v1` | Bearer `sk-*` key (nginx-fronted) | ✅ **VALID** — authenticated via nginx :80 (verified 2026-08-09: 401 without key, 200 with) |
| `http://192.168.68.116/litellm/v1/responses` | Bearer `sk-*` key | ✅ **CURRENT** — authenticated LiteLLM proxy | | `http://192.168.68.116/litellm/v1` | Bearer `sk-*` key (nginx-fronted) | ✅ **CURRENT / CANONICAL** — captain-approved migration target; 600s proxy_read_timeout (verified) |
| `http://192.168.68.116:4000/v1` | Bearer `sk-*` key (direct container) | ❌ **FORBIDDEN** — bypasses nginx; port 4000 direct is not a config path |
All harness/litellm providers MUST use the authenticated `/litellm/v1/responses` path. All harness/litellm providers MUST use an authenticated nginx-fronted path (`/litellm/v1` canonical, `/v1` legacy-valid).
Any `base_url` pointing to bare `/v1` on 192.168.68.116 is a **migration violation**. Any `base_url` pointing at `:4000` or a bare IP without nginx is a **migration violation**.
### 🔥 CRITICAL: Double-Path Bug (2026-07-10) ### 🔥 CRITICAL: Double-Path Bug (2026-07-10)
@@ -190,22 +191,22 @@ litellm_settings:
|-------|-----|-----|---------------|------------|--------|-----------------|---------------| |-------|-----|-----|---------------|------------|--------|-----------------|---------------|
| Tanko | 112 | .122 | `tanko` | Infisical vault | ✅ Fixed | `infisical run` | 20:17 UTC Jul 5 | | Tanko | 112 | .122 | `tanko` | Infisical vault | ✅ Fixed | `infisical run` | 20:17 UTC Jul 5 |
| Mumuni | 100 (abiba) | .24 | `mumuni` | Infisical vault | ✅ Fixed | Pi Hermes gateway | 2026-07-27 | | Mumuni | 100 (abiba) | .24 | `mumuni` | Infisical vault | ✅ Fixed | Pi Hermes gateway | 2026-07-27 |
| Koby | 111 | ? | `koby` | Infisical vault | ✅ Fixed | `infisical run` | 23:30 UTC Jul 5 | | Koby | 111 | .129 | `koby` | Infisical vault | ✅ Fixed | `infisical run` | 23:30 UTC Jul 5 |
| Koonimo | 113 | ? | `koonimo` | Infisical vault | ✅ Fixed | `infisical run` (migrated 2026-07-11) | 2026-07-11 | | Koonimo | 113 | .114 | `koonimo` | Infisical vault | ✅ Fixed | `infisical run` (migrated 2026-07-11) | 2026-08-09 |
| Abiba | 100 | .65 | `abiba-pi` | Infisical vault | ✅ N/A (pi native) | — | 19:44 UTC Jul 5 | | Abiba | 100 | .65 | `abiba-pi` | Infisical vault | ✅ N/A (pi native) | — | 19:44 UTC Jul 5 |
| Kagenz0 | 105 | ? | — | — | ❌ DOWN | — | 19:14 EDT Jul 4 | | Kagenz0 | 105 | .14 | — | — | ❌ DOWN | — | 19:14 EDT Jul 4 |
> **Note**: CT hostnames (tdunna, baggy) differ from agent identities (koby, koonimo). > **Note**: CT hostnames (tdunna, baggy) differ from agent identities (koby, koonimo).
> LiteLLM key aliases use agent identity, not CT hostname. > LiteLLM key aliases use agent identity, not CT hostname.
### Migration Status: Authenticated Path ### Migration Status: Authenticated Path
| Agent | `/litellm/v1/responses` | Deprecated `/v1` | Status | | Agent | `/litellm/v1` | Legacy `/v1` | Status |
|-------|--------------------------|--------------------|--------| |-------|--------------|-------------|--------|
| Mumuni | ✅ 5 sections | 0 | ✅ Authenticated | | Mumuni | ✅ harness provider | ✅ auxiliary on /v1 (valid) | ✅ Authenticated (verified 2026-08-09) |
| Tanko | ⚠️ No SSH access | — | Needs check | | Tanko | ✅ 5 sections | 0 | ✅ Migrated 2026-08-08, keys 200 |
| Koby | ⚠️ No route to host | — | Needs check | | Koby | ✅ custom provider (harness name) | — | ✅ External DeepSeek primary (intentional) |
| Koonimo | ⚠️ Connection timed out | — | Needs check | | Koonimo | ✅ .114 (baggy) | — | ✅ 128K context applied 2026-08-09 |
### Systemd Service Pattern (2026-07-11 — vault migration) ### Systemd Service Pattern (2026-07-11 — vault migration)
+8 -6
View File
@@ -7,13 +7,15 @@ description: >
from nvidia-smi (.8, .110) and amdgpu_top (.15). LiteLLM metrics from nvidia-smi (.8, .110) and amdgpu_top (.15). LiteLLM metrics
via existing /metrics Prometheus endpoint. via existing /metrics Prometheus endpoint.
DEPLOYMENT STATUS (2026-07-09): DEPLOYMENT STATUS (2026-08-09):
✅ Core stack deployed: Prometheus + Grafana + pve/node/docker exporters ✅ Core stack deployed: Prometheus + Grafana + pve/node/docker exporters
(via proxmox-monitor contract). Grafana at :3001, 5 scrape targets active. (via proxmox-monitor contract). Grafana at :3001, all scrape targets active.
❌ GPU exporters NOT deployed: gpu-exporter crash-loops on .15, ✅ GPU exporters DEPLOYED: all 3 GPU hosts (.8/.110/.15) run exporters on
NVIDIA sidecar exporters (.8/.110:9400) never installed. :9400 (nvidia_gpu_exporter / amdgpu exporter) — verified 200 on 2026-08-09.
Router falls back to direct GPU /health probes. ✅ LiteLLM /metrics scraping live (success_callback: prometheus; auth via
⚠️ This contract is target-state aspirational — not as-built. master key) + Alertmanager + Zulip bridge (alerts-infra) added 2026-08-09.
⚠️ This contract is target-state aspirational — but GPU export + alerting
are now as-built (verified 2026-08-09).
As-built GPU monitoring is via gpu-monitor contract (port 9100 poll). As-built GPU monitoring is via gpu-monitor contract (port 9100 poll).
version: 1.0.0 version: 1.0.0
--- ---
+9
View File
@@ -18,6 +18,15 @@ description: >
Designed as a reusable contract for any Syslog agent. Designed as a reusable contract for any Syslog agent.
Source of truth: gpu-fleet.prose.md Source of truth: gpu-fleet.prose.md
## Monitoring / Alerting (as-built 2026-08-09)
- LiteLLM /metrics scrape job: requires `litellm_settings.success_callback: [prometheus]`
(failure_callback alone does NOT mount /metrics — verified 2026-08-09).
Scraped by Prometheus with Bearer master key; endpoint returns 307 → /metrics/.
- Alertmanager (harness-alertmanager :9093) + zulip-bridge (:9102) deliver
firing alerts to #agent-hub > alerts-infra via abiba-bot. Added 2026-08-09.
- Prometheus node job covers ALL 6 PVE nodes (.4/.5/.6/.9/.12/.15:9100).
--- ---
## Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU) ## Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU)
+16 -6
View File
@@ -2,21 +2,31 @@
kind: responsibility kind: responsibility
name: pm2-self-heal name: pm2-self-heal
description: > description: >
Monitors critical PM2 processes (abiba-zulip, abiba-telegram) and Monitors critical PM2 processes (abiba-zulip, abiba-telegram, gitea-runner,
auto-restarts any that are stopped or errored. Logs every action to spoton-service, zulip-watchdog) and auto-restarts any that are stopped or
the knowledge graph and alerts the owner via Zulip DM on failures. errored. Logs every action to Gitea (SyslogSolution/health-logs — not
knowledge graph, hard rule) and alerts the owner via
Zulip DM on failures.
CRITICAL: Never restart abiba-zulip — it runs this contract. CRITICAL: Never restart abiba-zulip — it runs this contract.
AS-BUILT 2026-08-09 (captain ruling, ecosystem is authoritative):
gpu-monitor is systemd-managed (gpu-monitor.service) — NOT PM2;
gpu-watchdog decommissioned (function folded into gpu-monitor.service);
gitea-runner KEPT (online in PM2); abiba-zulip KEPT (online 4d+, the
2026-07-04 'removed/decommissioned' note was stale and is removed).
--- ---
## Maintains ## Maintains
- abiba-telegram: { status: "online", uptime: string, restarts: number } - abiba-telegram: { status: "online", uptime: string, restarts: number }
- gpu-watchdog: { status: "online", uptime: string, restarts: number } - abiba-zulip: { status: "online", uptime: string, restarts: number }
- gpu-monitor: { status: "online", uptime: string, restarts: number }
- gitea-runner: { status: "online", uptime: string, restarts: number } - gitea-runner: { status: "online", uptime: string, restarts: number }
- spoton-service: { status: "online", uptime: string, restarts: number }
- zulip-watchdog: { status: "online", uptime: string, restarts: number }
- last_check: timestamp - last_check: timestamp
> **Note (2026-07-04):** `abiba-zulip` removed — Zulip extension decommissioned. > **Note (2026-07-04, SUPERSEDED 2026-08-09):** `abiba-zulip` remains ONLINE and
> is monitored — the decommission note was stale (process re-added; do not treat
> it as removed).
## Continuity ## Continuity
+8 -1
View File
@@ -199,7 +199,14 @@ Check `platforms.zulip.state`: `connected` ✅ | `disconnected` ❌ | `error`
ssh root@<CT> "ps aux | grep 'gateway run' | grep -v grep" ssh root@<CT> "ps aux | grep 'gateway run' | grep -v grep"
``` ```
Gateway PID should exist with uptime > 60s. **Dual-gateway detection**: if more than one `gateway run` process is found, the gateway has a collision (typically one `--force` and one `--replace` process). Kill the newer/duplicate process, then restart the remaining gateway via PM2 (`pm2 restart mumuni-zulip`). Check gateway log for "Gateway running with 2 platform(s)" (not 1) to confirm Zulip reloaded. Gateway PID should exist with uptime > 60s. **Dual-gateway detection**: if more than one `gateway run` process is found, the gateway has a collision (typically one `--force` and one `--replace` process). Kill the newer/duplicate process, then restart the remaining gateway per-agent (parameterized 2026-08-09, captain ruling):
| Agent | Restart command | Notes |
|-------|-----------------|-------|
| Mumuni (.24) | `pm2 restart abiba-zulip` | Hermes gateway runs under PM2 as `abiba-zulip` |
| Tanko (.122) | `bash /opt/hermes-zulip-plugin/run.sh` (or the agent's systemd/user unit) | Tanko does NOT use PM2 — never run `pm2 restart mumuni-zulip` for Tanko (process does not exist) |
Check gateway log for "Gateway running with 2 platform(s)" (not 1) to confirm Zulip reloaded.
**B3: Heartbeat Verification** **B3: Heartbeat Verification**
+2 -3
View File
@@ -75,9 +75,8 @@ Track via `/tmp/zulip-heal-debounce-<agent>` (unix timestamp of last restart).
## Reporting ## Reporting
Every cycle produces a knowledge graph node: This contract is RETIRED — health-check logs are NOT knowledge graph content.
- Title: `[LEARN] zulip-self-heal: <timestamp>` No graph nodes are created. Logs go to Gitea (SyslogSolution/health-logs).
- metadata: { type: "remediation", status: "fixed" | "escalated" | "healthy" }
- Issues fixed → DM: "🛠 Zulip Self-Heal: fixed <issue>" - Issues fixed → DM: "🛠 Zulip Self-Heal: fixed <issue>"
- Issues escalated → DM: "⚠️ Zulip Self-Heal: <issue> needs attention" - Issues escalated → DM: "⚠️ Zulip Self-Heal: <issue> needs attention"