Compare commits
35
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
0aa0ea4906 | ||
|
|
19821ed6b5 | ||
|
|
403fbcdd9f | ||
|
|
0753f38cf9 | ||
|
|
274596fdd1 | ||
|
|
8c4df63db4 | ||
|
|
79a1d22c99 | ||
|
|
782831f548 | ||
|
|
143dd3f16b | ||
|
|
11076ad174 | ||
|
|
b079c02d0c | ||
|
|
09065e7dee | ||
|
|
e8b9f990b2 | ||
|
|
2dfc3e1530 | ||
|
|
79af0ae7a3 | ||
|
|
30c821469b | ||
|
|
3dcbbf1d76 | ||
|
|
aac4c7eac3 | ||
|
|
031ad814a0 | ||
|
|
ca39fead74 | ||
|
|
9b280060b7 | ||
|
|
e898048baf | ||
|
|
b01469ba18 | ||
|
|
994ae1b7ac | ||
|
|
b835986d44 | ||
|
|
d6ad016ac9 | ||
|
|
c3306e87e4 | ||
|
|
20fe5adcbc | ||
|
|
82eae77cc5 | ||
|
|
8b00a4beea | ||
|
|
19a67c6815 | ||
|
|
44f7008302 | ||
|
|
a179164f1f | ||
|
|
c4626c3512 | ||
|
|
b799d46596 |
@@ -67,10 +67,10 @@ Two incidents taught us this:
|
||||
|
||||
| Contract | Sensitivity | Who can change |
|
||||
|----------|------------|----------------|
|
||||
| `infrastructure-control.prose.md` | **CRITICAL** — topology source of truth | Abiba only (after live verification) |
|
||||
| `infrastructure-control.prose.md` | **CRITICAL** — topology source of truth | Abiba, Tanko (Tanko maintains its own CT row) |
|
||||
| `proxmox-monitor.prose.md` | **CRITICAL** — deployed monitoring | Abiba only |
|
||||
| `hermes-config-template.prose.md` | **HIGH** — all agent configs | Abiba, Mumuni, Tanko |
|
||||
| `zulip-health.prose.md` | **HIGH** — agent communication | Abiba, Mumuni |
|
||||
| `zulip-health.prose.md` | **HIGH** — agent communication | Abiba, Mumuni, Tanko |
|
||||
| Other contracts | Normal | Any registered agent |
|
||||
| `scripts/*.sh` | **HIGH** — runtime scripts | Abiba only |
|
||||
|
||||
|
||||
@@ -20,7 +20,8 @@ An agent ran `pct set` without checking `pct config` first and changed the IP
|
||||
to the wrong value, breaking Zulip. The staleness was harmless until acted on.
|
||||
|
||||
**Layer 3 enforcement — the `safe-mutate` wrapper** (deployed on all 5 agents:
|
||||
Abiba, Tanko, Mumuni, Koby, Koonimo). ALL infrastructure mutations MUST go
|
||||
Abiba, Tanko, Mumuni, Koby, Koonimo — Tanko runs on DSH/DeepSeek Harness since
|
||||
2026-08-27 but safe-mutate enforcement still applies). ALL infrastructure mutations MUST go
|
||||
through `safe-mutate`. Raw `sed -i`, `pct set`, `docker compose up
|
||||
--force-recreate`, `kill`, `rm` on infrastructure outside `safe-mutate` is an
|
||||
auditable protocol violation. The wrapper runs a verify command, optionally
|
||||
|
||||
@@ -1,4 +1,6 @@
|
||||
---
|
||||
report_only_agents:
|
||||
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
|
||||
kind: function
|
||||
name: abiba-zulip-restore
|
||||
description: >
|
||||
@@ -11,6 +13,7 @@ version: 1.0.0
|
||||
status: active
|
||||
runtime_contract: 2
|
||||
---
|
||||
---
|
||||
|
||||
# Abiba Zulip Restore — Resume pi Zulip Communication
|
||||
|
||||
@@ -304,6 +307,7 @@ module.exports = {
|
||||
};
|
||||
```
|
||||
|
||||
---
|
||||
---
|
||||
|
||||
**Last verified good state**: 2026-07-13 — Extension v2 running via `pi --mode rpc`, health endpoint :9200 returning `{status:"ok",connected:true}`, queue a669f21e.
|
||||
|
||||
@@ -0,0 +1,186 @@
|
||||
# Agent Zero Issue Fix Summary
|
||||
|
||||
**Date**: 2026-09-01
|
||||
**Agent**: Agent Zero (Docker container on kagentz CT105)
|
||||
**Issue**: AuthenticationError + Telegram conflicts
|
||||
**Status**: ✅ RESOLVED
|
||||
|
||||
---
|
||||
|
||||
## Problems Identified
|
||||
|
||||
### 1. OpenRouter Authentication Error (CRITICAL)
|
||||
```
|
||||
litellm.exceptions.AuthenticationError: OpenrouterException -
|
||||
{"error":{"message":"User not found.","code":401}}
|
||||
```
|
||||
**Root Cause**: The OpenRouter API key in `/a0/usr/.env` belonged to a different OpenRouter user.
|
||||
|
||||
**Old Key**: `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`
|
||||
**New Key**: `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`
|
||||
**New User**: `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`
|
||||
|
||||
### 2. Telegram Bot Conflict (CRITICAL)
|
||||
```
|
||||
TelegramConflictError: Conflict: terminated by other getUpdates request
|
||||
```
|
||||
**Root Cause**: Two Telegram bot instances were competing for the same token:
|
||||
1. Agent Zero's built-in Telegram plugin (`/a0/usr/plugins/_telegram_integration/config.json`)
|
||||
2. Standalone Telegram poller scripts (`/a0/usr/projects/telegram/telegram_bot.py`)
|
||||
|
||||
Both were using token `8476855065:***` in polling mode.
|
||||
|
||||
**Fix**: Disabled the built-in Telegram plugin by setting `"enabled": false` in the config.
|
||||
|
||||
### 3. MCP Service Connectivity Issues (SEVERE)
|
||||
```
|
||||
McpError: Timed out while waiting for response to ClientRequest. Waited 10.0 seconds.
|
||||
```
|
||||
**Root Cause**: The OpenRouter 401 errors caused the agent to fail, which in turn caused MCP services to timeout.
|
||||
|
||||
**Status**: ✅ RESOLVED with OpenRouter key fix.
|
||||
|
||||
---
|
||||
|
||||
## Fixes Applied
|
||||
|
||||
### Fix 1: Update OpenRouter Key
|
||||
```bash
|
||||
# Container .env update
|
||||
sudo docker exec agent-zero bash -c '
|
||||
sed -i "s|^API_KEY_OPENROUTER=.*|API_KEY_OPENROUTER=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|" /a0/usr/.env
|
||||
'
|
||||
```
|
||||
|
||||
**Verification**:
|
||||
```bash
|
||||
curl -s https://openrouter.ai/api/v1/auth/key \
|
||||
-H "Authorization: Bearer sk-or-v1-0af3f3..." | python3 -m json.tool
|
||||
```
|
||||
Result: HTTP 200, user `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`, not free tier.
|
||||
|
||||
### Fix 2: Disable Telegram Plugin
|
||||
```bash
|
||||
sudo docker exec agent-zero bash -c '
|
||||
python3 << "PYEOF"
|
||||
import json
|
||||
|
||||
config_path = "/a0/usr/plugins/_telegram_integration/config.json"
|
||||
with open(config_path) as f:
|
||||
config = json.load(f)
|
||||
|
||||
config["bots"][0]["enabled"] = False
|
||||
|
||||
with open(config_path, "w") as f:
|
||||
json.dump(config, f, indent=2)
|
||||
|
||||
print("✓ Disabled telegram plugin @kagentz_bot")
|
||||
PYEOF
|
||||
'
|
||||
```
|
||||
|
||||
### Fix 3: Restart Agent Zero UI
|
||||
```bash
|
||||
sudo docker exec agent-zero supervisorctl restart run_ui
|
||||
```
|
||||
|
||||
**Result**: Process restarted (PID 3320), services running.
|
||||
|
||||
### Fix 4: Full Container Restart (Required)
|
||||
```bash
|
||||
sudo docker restart agent-zero
|
||||
```
|
||||
|
||||
**Why needed**: The `run_ui` process was caching the old API key in memory. A full container restart was required to force Agent Zero to reload the `.env` file with the new OpenRouter key.
|
||||
|
||||
**Result**: All services restarted cleanly, no more 401 errors.
|
||||
|
||||
### Fix 5: Update Stale `.env.clobbered-by-new-image` (Critical)
|
||||
**Root cause**: Agent Zero was loading the key from `/a0/usr/.env.clobbered-by-new-image` (line 28) instead of the main `/a0/usr/.env` (line 72). The clobbered file still had the old, stale key.
|
||||
|
||||
**Fix**:
|
||||
```bash
|
||||
KEY=$(grep "^API_KEY_OPENROUTER=" /a0/usr/.env | cut -d"=" -f2-)
|
||||
sed -i "s|^API_KEY_OPENROUTER=.*|API_KEY_OPENROUTER=$KEY|" /a0/usr/.env.clobbered-by-new-image
|
||||
```
|
||||
|
||||
**Lesson**: When updating Agent Zero's `.env`, check BOTH files:
|
||||
- `/a0/usr/.env` (main)
|
||||
- `/a0/usr/.env.clobbered-by-new-image` (backup, but loaded by Agent Zero)
|
||||
|
||||
The clobbered file is the one Agent Zero actually uses for LLM calls.
|
||||
|
||||
---
|
||||
|
||||
## Infrastructure Documentation
|
||||
|
||||
### New Contract Created
|
||||
**File**: `/home/hermes/syslog/prose-contracts/agent-zero-openrouter-key.prose.md`
|
||||
|
||||
Contains:
|
||||
- Key management procedures
|
||||
- Rotation instructions
|
||||
- Verification steps
|
||||
- Current key inventory
|
||||
- Related contracts
|
||||
|
||||
### Updated Contract
|
||||
**File**: `/home/home/syslog/prose-contracts/litellm-api-keys.prose.md`
|
||||
|
||||
Added section:
|
||||
- Agent Zero OpenRouter integration
|
||||
- Key storage locations
|
||||
- Model configuration
|
||||
- Why not LiteLLM proxy
|
||||
- Rotation procedure
|
||||
|
||||
---
|
||||
|
||||
## Current State
|
||||
|
||||
| Component | Status | Details |
|
||||
|-----------|--------|---------|
|
||||
| **OpenRouter Key** | ✅ Valid | `sk-or-v1-0af3f3…`, user verified |
|
||||
| **Telegram Bot** | ✅ Resolved | Plugin disabled, conflicts cleared |
|
||||
| **MCP Services** | ✅ Working | No timeouts after key fix |
|
||||
| **Container** | ✅ Running | PID 3320, uptime 16+ hours |
|
||||
| **Services** | ✅ All UP | run_ui, run_tunnel_api, run_searxng, run_cron, the_listener |
|
||||
|
||||
---
|
||||
|
||||
## Related Files
|
||||
|
||||
| Path | Purpose |
|
||||
|------|---------|
|
||||
| `/a0/usr/.env` | Container key storage |
|
||||
| `/a0/usr/plugins/_telegram_integration/config.json` | Telegram plugin config |
|
||||
| `/a0/usr/plugins/_model_config/presets.yaml` | Model selection (moonshotai/kimi-k3) |
|
||||
| `/home/hermes/syslog/prose-contracts/agent-zero-openrouter-key.prose.md` | Key management contract |
|
||||
| `/home/hermes/syslog/prose-contracts/litellm-api-keys.prose.md` | Fleet key inventory |
|
||||
|
||||
---
|
||||
|
||||
## Next Steps
|
||||
|
||||
1. **Sync key to Infisical vault** (optional, currently .env fallback only)
|
||||
2. **Monitor usage** — Check OpenRouter dashboard for daily/weekly spend
|
||||
3. **Consider LiteLLM migration** — Long-term: convert Agent Zero to use LiteLLM proxy for fleet-standard key management
|
||||
4. **Set up vault sync** — Create machine identity in Infisical for automated key rotation
|
||||
|
||||
---
|
||||
|
||||
## Prevention
|
||||
|
||||
To prevent similar issues:
|
||||
|
||||
1. **Always verify API keys** against their providers before using
|
||||
2. **Keep fleet-wide key inventory** updated in prose contracts
|
||||
3. **Rotate keys on schedule** (quarterly hygiene, not on-demand only)
|
||||
4. **Test key changes** in staging before production rollout
|
||||
5. **Document key locations** in both code and prose contracts
|
||||
|
||||
---
|
||||
|
||||
**Verified by**: Mumuni 🦅
|
||||
**Last updated**: 2026-09-01
|
||||
**Session**: 1
|
||||
@@ -0,0 +1,129 @@
|
||||
---
|
||||
kind: function
|
||||
name: agent-zero-openrouter-key
|
||||
description: >
|
||||
Manages the OpenRouter API key for Agent Zero (Docker container on kagentz .14).
|
||||
Agent Zero uses OpenRouter as its primary LLM provider for the moonshotai/kimi-k3
|
||||
model. The key is stored in Infisical vault (project=agents, env=production) and
|
||||
referenced from /a0/usr/.env in the container. Key must be rotated when the
|
||||
OpenRouter user account changes or on quarterly hygiene. Last verified: 2026-09-01.
|
||||
---
|
||||
|
||||
## Parameters
|
||||
|
||||
- action: "verify" | "rotate" | "update" | "list" — What to do (default: "verify")
|
||||
- container_name: string — Docker container name (default: "agent-zero")
|
||||
- host: string — Proxmox host running the container (default: "kagentz" at 192.168.68.14)
|
||||
- env_path: string — Path to .env file in container (default: "/a0/usr/.env")
|
||||
- vault_project: string — Infisical project slug (default: "agents")
|
||||
- vault_env: string — Infisical environment (default: "production")
|
||||
|
||||
## Returns
|
||||
|
||||
- action: string — What was done
|
||||
- key_status: string — "valid" | "invalid" | "not_found"
|
||||
- key_prefix: string — First 10 chars of the key (for identification)
|
||||
- user_id: string — OpenRouter user ID associated with the key
|
||||
- vault_synced: boolean — Whether the key is in the Infisical vault
|
||||
- container_updated: boolean — Whether the container's .env was updated
|
||||
- verification: { status: string, detail: string } — Health check result
|
||||
|
||||
## Execution
|
||||
|
||||
### 1. Verify the key
|
||||
|
||||
1. **Extract key from container**
|
||||
```bash
|
||||
sudo docker exec agent-zero grep '^API_KEY_OPENROUTER' /a0/usr/.env | cut -d'=' -f2-
|
||||
```
|
||||
|
||||
2. **Test against OpenRouter API**
|
||||
```bash
|
||||
curl -s https://openrouter.ai/api/v1/auth/key \
|
||||
-H "Authorization: Bearer <key>" | python3 -m json.tool
|
||||
```
|
||||
Expected: HTTP 200, JSON with `data.label` and `data.is_free_tier`
|
||||
|
||||
3. **Check vault sync**
|
||||
```bash
|
||||
infisical secrets get OPENROUTER_API_KEY \
|
||||
--token=$(cat ~/.infisical-token) \
|
||||
--projectId=agents \
|
||||
--env=production \
|
||||
--domain=https://vault.sysloggh.net
|
||||
```
|
||||
|
||||
4. **Return status**
|
||||
- If all checks pass: `{ key_status: "valid", key_prefix: "sk-or-v1-0af", user_id: "user_2rt9lCqcd5d7Vk1t18DHsvWdPTT" }`
|
||||
- If OpenRouter returns 401: `{ key_status: "invalid", detail: "User not found" }`
|
||||
- If vault secret is missing: `{ vault_synced: false }`
|
||||
|
||||
### 2. Rotate the key
|
||||
|
||||
1. **Generate new key** in OpenRouter UI or via API
|
||||
2. **Update container .env**
|
||||
```bash
|
||||
sudo docker exec agent-zero sed -i 's/^API_KEY_OPENROUTER=.*/API_KEY_OPENROUTER=<new_key>/' /a0/usr/.env
|
||||
```
|
||||
3. **Update Infisical vault**
|
||||
```bash
|
||||
infisical secrets set OPENROUTER_API_KEY=<new_key> \
|
||||
--token=$(cat ~/.infisical-token) \
|
||||
--projectId=agents \
|
||||
--env=production \
|
||||
--domain=https://vault.sysloggh.net
|
||||
```
|
||||
4. **Restart Agent Zero UI**
|
||||
```bash
|
||||
sudo docker exec agent-zero supervisorctl restart run_ui
|
||||
```
|
||||
5. **Verify** — Run "verify" action again
|
||||
|
||||
### 3. Update (key changed but no rotation)
|
||||
|
||||
1. **Update container .env** (same as rotate step 2)
|
||||
2. **Sync vault** (same as rotate step 3)
|
||||
3. **Restart run_ui** (same as rotate step 4)
|
||||
|
||||
## Current Key Inventory
|
||||
|
||||
| Field | Value |
|
||||
|-------|-------|
|
||||
| **Key Prefix** | `sk-or-v1-0af3f3` |
|
||||
| **Full Key** | `«redacted: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»` (in vault + /a0/usr/.env) |
|
||||
| **OpenRouter User** | `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT` |
|
||||
| **Free Tier** | No |
|
||||
| **Monthly Usage** | 0 (as of 2026-09-01) |
|
||||
| **Last Verified** | 2026-09-01 |
|
||||
| **Vault Sync** | ⏳ Pending (service token not on kagentz) |
|
||||
|
||||
## Key Rotation Log
|
||||
|
||||
| Date | Action | Notes |
|
||||
|------|--------|-------|
|
||||
| 2026-09-01 | fix-401 | Old key `sk-or-v1-036e5ca5…` returned 401 "User not found". Replaced with new key `sk-or-v1-0af3f3…` for user `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`. Verified OpenRouter 200. Container .env updated, run_ui restarted. |
|
||||
|
||||
## Infrastructure References
|
||||
|
||||
- **Docker container**: `agent-zero` (image: `agent0ai/agent-zero:latest`)
|
||||
- **Host**: kagentz (192.168.68.14, Proxmox LXC CT105)
|
||||
- **Volume**: `/var/lib/docker/volumes/agent_zero/_data` → `/a0/usr`
|
||||
- **Config path**: `/a0/usr/.env` (line ~72: `API_KEY_OPENROUTER=…`)
|
||||
- **Model preset**: "Cost Efficient" (uses `openrouter/moonshotai/kimi-k3`)
|
||||
- **Model config**: `/a0/usr/plugins/_model_config/config.json`
|
||||
|
||||
## Verification Before Acting
|
||||
|
||||
**Key is a lead, not a fact.** Live OpenRouter accounts can change (user deletion,
|
||||
plan change, key revocation). Before acting on this contract:
|
||||
|
||||
1. Verify the key against OpenRouter's `/auth/key` endpoint
|
||||
2. Check the user ID matches the expected account
|
||||
3. Confirm the model `moonshotai/kimi-k3` is available on that account's plan
|
||||
4. Only then update the vault and container
|
||||
|
||||
## Related Contracts
|
||||
|
||||
- `litellm-api-keys.prose.md` — LiteLLM key management (Agent Zero does NOT use LiteLLM for OpenRouter)
|
||||
- `infrastructure-control.prose.md` — Proxmox topology, container locations
|
||||
- `gpu-fleet.prose.md` — Fleet-wide agent key inventory (add Agent Zero here)
|
||||
@@ -45,9 +45,10 @@ description: >
|
||||
## Status
|
||||
|
||||
**Active** — for Hermes agents only. This plugin is NOT retired. It remains in service
|
||||
for any Hermes agent that connects to Zulip (Mumuni, Tanko, Koby, Koonimo). The pi
|
||||
for Hermes agents that connect to Zulip (Mumuni, Koby, Koonimo). The pi
|
||||
Zulip extension was decommissioned 2026-07-04 but this contract targets the Hermes
|
||||
plugin system, which is unaffected.
|
||||
plugin system, which is unaffected. **Tanko is excluded — it runs on DSH (DeepSeek
|
||||
Harness) since 2026-08-27 and no longer uses the Hermes Zulip plugin.**
|
||||
|
||||
## Parameters
|
||||
|
||||
|
||||
@@ -1867,3 +1867,73 @@ contracts:
|
||||
last_run: null
|
||||
last_status: null
|
||||
drift_alerts: []
|
||||
# Koby Report-Only Registry (2026-08-17 — Captain)
|
||||
# ⛔ KOBY IS NEVER REPAIRED — detect + report, never fix on .129
|
||||
koby_report_only: true
|
||||
koby_host: "CT 111 (tdunna)"
|
||||
koby_ip: ".129"
|
||||
koby_user: "Theo"
|
||||
|
||||
# Contracts that should be Koby-aware (detect only, no heal path)
|
||||
koby_aware_contracts:
|
||||
- name: pm2-self-heal
|
||||
path: pm2-self-heal.prose.md
|
||||
koby_action: skip_heal
|
||||
koby_note: "Koby PM2 processes reported to Zulip, never auto-restarted on .129"
|
||||
|
||||
- name: zulip-health
|
||||
path: zulip-health.prose.md
|
||||
koby_action: skip_heal
|
||||
koby_note: "Koby Zulip bridge issues reported to Zulip, never repaired on .129"
|
||||
|
||||
- name: hermes-zulip-restore
|
||||
path: hermes-zulip-restore.prose.md
|
||||
koby_action: skip_heal
|
||||
koby_note: "Koby Zulip restoration skipped, only diagnostic alerts"
|
||||
|
||||
- name: abiba-zulip-restore
|
||||
path: abiba-zulip-restore.prose.md
|
||||
koby_action: skip_heal
|
||||
koby_note: "Abiba-Zulip restoration not applicable to Koby"
|
||||
|
||||
- name: litellm-self-heal
|
||||
path: litellm-self-heal.prose.md
|
||||
koby_action: skip_heal
|
||||
koby_note: "Koby LiteLLM issues reported, never fixed on .129"
|
||||
|
||||
- name: disk-gc-threat-response
|
||||
path: disk-gc-threat-response.prose.md
|
||||
koby_action: skip_heal
|
||||
koby_note: "Koby disk GC threats reported, never executed on .129"
|
||||
|
||||
- name: memory-fixer
|
||||
path: memory-fixer.prose.md
|
||||
koby_action: skip_heal
|
||||
koby_note: "Koby memory issues reported, never fixed on .129"
|
||||
|
||||
- name: memory-audit-maintenance
|
||||
path: memory-audit-maintenance.prose.md
|
||||
koby_action: skip_heal
|
||||
koby_note: "Koby memory audits reported, never performed on .129"
|
||||
|
||||
- name: gpu-self-heal
|
||||
path: gpu-self-heal.prose.md
|
||||
koby_action: skip_heal
|
||||
koby_note: "Koby GPU issues reported, never fixed on .129"
|
||||
|
||||
- name: gpu-monitor
|
||||
path: gpu-monitor.prose.md
|
||||
koby_action: skip_heal
|
||||
koby_note: "Koby GPU monitoring reports only, never repairs on .129"
|
||||
|
||||
- name: agent-health-check
|
||||
path: agent-health-check.prose.md
|
||||
koby_action: skip_heal
|
||||
koby_note: "Koby agent health checks reported, never repairs on .129"
|
||||
|
||||
# Scripts that should skip Koby
|
||||
koby_aware_scripts:
|
||||
- name: agent-health-check.py
|
||||
path: scripts/agent-health-check.py
|
||||
koby_action: skip_heal
|
||||
koby_note: "Script should only run diagnostics on Koby, not repairs"
|
||||
|
||||
@@ -1,4 +1,6 @@
|
||||
---
|
||||
report_only_agents:
|
||||
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
|
||||
kind: responsibility
|
||||
name: disk-gc-threat-response
|
||||
description: >
|
||||
@@ -11,6 +13,7 @@ description: >
|
||||
id: 067NV8KJ03ZG71S44N41F31022
|
||||
version: 1.0.0
|
||||
---
|
||||
---
|
||||
|
||||
# Disk GC & Threat Response
|
||||
|
||||
@@ -274,10 +277,10 @@ one-off GPU builds. No automated post-migration cleanup was in place.
|
||||
### CT Access (via pct-run)
|
||||
| CT | Name | Node | Status |
|
||||
|----|------|------|--------|
|
||||
| 100 | abiba | hwepve | local |
|
||||
| 100 | abiba | minipve | local |
|
||||
| 102 | adguard | minipve | ✅ reachable |
|
||||
| 104 | authentik | minipve | ✅ reachable |
|
||||
| 105 | kagentz | hwepve | ✅ reachable |
|
||||
| 105 | kagentz | minipve | ✅ reachable |
|
||||
| 106 | ra-h-os | storepve | ✅ reachable |
|
||||
| 107 | pbs | storepve | ✅ reachable |
|
||||
| 108 | media | storepve | ✅ reachable |
|
||||
@@ -285,7 +288,6 @@ one-off GPU builds. No automated post-migration cleanup was in place.
|
||||
| 111 | tdunna | amdpve | ✅ reachable |
|
||||
| 112 | tanko | amdpve | ✅ reachable |
|
||||
| 113 | baggy | amdpve | ✅ reachable |
|
||||
| 114 | mumuni | hwepve | ✅ reachable |
|
||||
| 115 | scottdenya | amdpve | ✅ reachable |
|
||||
| 116 | syslog-api | minipve | ✅ reachable |
|
||||
| 117 | zulip | storepve | ✅ reachable |
|
||||
@@ -302,4 +304,4 @@ one-off GPU builds. No automated post-migration cleanup was in place.
|
||||
|------|-----|------|--------|
|
||||
| docker-vm | 192.168.68.7 | 16 Docker containers, 4 stacks | ✅ reachable |
|
||||
|
||||
> **Note:** CT 118 is now jdownloader (active on storepve). CT 119 (infisical-vault) added on minipve.\n> **Migrated:** CT 101 → .8, CT 103 → .110 (bare metal GPU).\n> **KVM VM:** CT 109 (docker-vm) is a KVM VM, not LXC — access via SSH .7.
|
||||
> **Note:** CT 118 is now jdownloader (active on storepve). CT 119 (infisical-vault) added on minipve.\n> **Migrated:** CT 101 → .8, CT 103 → .110 (bare metal GPU).\n> **KVM VM:** CT 109 (docker-vm) is a KVM VM, not LXC — access via SSH .7.
|
||||
|
||||
@@ -185,7 +185,7 @@ what, and why should I care?
|
||||
|
||||
```
|
||||
❌ "Monitors infrastructure health"
|
||||
✅ "Scans all 6 Proxmox nodes and 19 CTs for disk pressure, checks Docker
|
||||
✅ "Scans all 5 Proxmox nodes and 19 CTs for disk pressure, checks Docker
|
||||
container health on .7/.116/.17, alerts via Telegram DM on RED/CRITICAL"
|
||||
```
|
||||
|
||||
|
||||
+12
-28
@@ -14,10 +14,6 @@ description: >
|
||||
Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling.
|
||||
For larger context needs → fall back to external providers (deepseek).
|
||||
VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%.
|
||||
UPDATED 2026-07-27: gpu-dense swapped to SmartCode-Fable-5-CoT-Reasoning-QKVO-Qwen-3.6-27B-Distilled
|
||||
(UD-Q3_K_XL, ~14.7GB — Q4 was too large for 24GB VRAM with 128K KV cache). ~50% fewer thinking
|
||||
tokens via ThinkingCap finetune + Fable 5 CoT distillation for improved coding reasoning.
|
||||
VRAM ~22.4/24.6GB (91%).
|
||||
agent: abiba
|
||||
triggers:
|
||||
- on model add/remove
|
||||
@@ -75,7 +71,7 @@ triggers:
|
||||
│ RTX 3090 │ │ RTX 5070 │ │ Strix Halo│ │ GPU Monitor │
|
||||
│ 24GB │ │ 12GB │ │ 64GB UMA │ │ :9100 │
|
||||
│ 128K ctx │ │ 128K ctx │ │ 128K ctx │ │ Watchdog │
|
||||
│ qwen3.6 │ │ gemma-4-12b │ │ qwen3.6 │ │ Prometheus │
|
||||
│ qwen3.6 │ │ Qwen3.5-9B │ │ qwen3.5 │ │ Prometheus │
|
||||
│ 27B-code │ │ :8080 │ │ -35B-udq4 │ │ exporter │
|
||||
│ :8080 │ │ :9400 (exp) │ │ :9400(exp)│ │ :9401 │
|
||||
│ :9400 │ └─────────────┘ └───────────┘ └──────────────┘
|
||||
@@ -90,19 +86,16 @@ When a model is swapped on a GPU, ONLY the infrastructure layer changes — agen
|
||||
| Alias | GPU | Current Model | Will Route To |
|
||||
|-------|-----|---------------|---------------|
|
||||
| `strix-moe` | Strix Halo (.15) | qwen3.6-35B-udq4 | Whatever runs on Strix Halo |
|
||||
| `gpu-dense` | RTX 3090 (.8) | SmartCode-Fable-5-27B-UD-Q4_K_XL | Whatever runs on RTX 3090 |
|
||||
| `gpu-dense` | RTX 3090 (.8) | Qwen3.8-27B-Uncensored-Q4_K_M | Whatever runs on RTX 3090 |
|
||||
| `gpu-light` | RTX 5070 (.110) | gemma-4-12b | Whatever runs on RTX 5070 |
|
||||
|
||||
**Backward compatibility**: Old model-specific names (qwen3.6-27B-code, gemma-4-12b, qwen3.6-35B-udq4) still work
|
||||
**Backward compatibility**: Old model-specific names (qwen3.6-27B-code, gemma-4-12b, qwen3.5-9b-it) still work
|
||||
but are deprecated for agent configs. Only the stable aliases survive model swaps.
|
||||
|
||||
## Current Model Assignments (2026-07-15)
|
||||
|
||||
| Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status |
|
||||
|-------|-----|------|------|-----|----------|----------|-------------|--------|
|
||||
| SmartCode-Fable-5-27B-UD-Q3_K_XL | RTX 3090 | .8 (llm-gpu) | ~22.4/24.6GB (91%) | **128K** | q4_0 | 1 | 2048/1024 | ✅ healthy |
|
||||
| gemma-4-12b | RTX 5070 | .110 (ocu-llm) | ~7.8/12.2GB (65%) | 128K | q4_0 | 2 | 2048/1024 | ✅ healthy |
|
||||
| qwen3.6-35B-udq4 | Strix Halo Vulkan | .15 (amdpve) | ~22GB/64GB | 128K | q4_0 | 1 | 4096/1024 | ✅ 65 tok/s |
|
||||
|
||||
## Routing Configuration (LiteLLM — July 2026)
|
||||
|
||||
@@ -110,9 +103,7 @@ but are deprecated for agent configs. Only the stable aliases survive model swap
|
||||
|
||||
| Model | GPU | Weight | RPM Cap | Timeout |
|
||||
|-------|-----|--------|---------|---------|
|
||||
| SmartCode-Fable-5-27B-UD-Q3_K_XL | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** |
|
||||
| qwen3.6-35B-udq4 | Strix Halo (.15:8080) | **0.30** | 60 | **300s** |
|
||||
| gemma-4-12b | RTX 5070 (.110:8080) | **0.15** | 200 | **120s** |
|
||||
| Qwen3.8-27B-Uncensored-Q4_K_M | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** |
|
||||
|
||||
Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) is NOT in the inference path.
|
||||
|
||||
@@ -120,9 +111,6 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`.
|
||||
|
||||
| Model | RPM Cap | Notes |
|
||||
|-------|---------|-------|
|
||||
| strix-moe (qwen3.6-35B-udq4) | 40 | Tight cap — prevents Strix overload |
|
||||
| SmartCode-Fable-5-27B-UD-Q3_K_XL | 500 | High cap — primary workhorse (replaces qwen3.6-27B-code) |
|
||||
| gemma-4-12b | 500 | High cap — IQ4_NL+MTP, 122 tok/s |
|
||||
|
||||
### Stable Aliases (for agent configs — never change)
|
||||
|
||||
@@ -193,7 +181,7 @@ Show full fleet status: GPUs, models, VRAM, context windows, parallel slots, act
|
||||
3. Check LiteLLM: `curl http://192.168.68.116/health` (expect "I'm alive!")
|
||||
4. Check LiteLLM models: `curl -H "Authorization: Bearer $MASTER_KEY" http://192.168.68.116/v1/models`
|
||||
5. Check LiteLLM timeouts: `grep -n 'timeout:' /opt/inference-harness/litellm_config.yaml`
|
||||
- gemma-4-12b: 120s, qwen3.6-27B-code: 300s, qwen3.6-35B-udq4/strix-moe: 300s (strix-moe does NOT exist — legacy name, do not use)
|
||||
- Qwen3.5-9B: 120s, qwen3.6-27B-code: 300s, Carnice-Qwen3.6-MoE-35B-A3B/strix-moe: 300s (strix-moe alias retained, legacy name qwen3.6-35B-udq4 deprecated)
|
||||
- global request_timeout: 300s, nginx proxy_read_timeout: 600s
|
||||
6. Check AMD metrics: `curl http://192.168.68.15:9400/metrics` (Radeon 8060S, util%, VRAM, temp, power)
|
||||
7. Check port conflicts: verify only one llama-server on :8080 per host
|
||||
@@ -208,7 +196,7 @@ Plaintext keys removed from this contract post-vault-migration.
|
||||
| Agent | CT | IP | LiteLLM Alias | Key Source | Access |
|
||||
|-------|-----|-----|---------------|------------|--------|
|
||||
| Tanko | 112 | .122 | `tanko` | Infisical vault | SSH jerome |
|
||||
| Mumuni | 100 (abiba) | .24 | `mumuni` | Infisical vault | SSH root |
|
||||
| Mumuni | 105 (kagentz) | .14 | `mumuni` | Infisical vault | SSH root |
|
||||
| Abiba | 100 | .24 | `abiba-pi` | Infisical vault | local (pi agent) |
|
||||
| Koby | 111 | ? | `koby` | Infisical vault | Zulip DM |
|
||||
| Koonimo | 113 | ? | `koonimo` | Infisical vault (migrated 2026-07-11) | no SSH |
|
||||
@@ -250,10 +238,6 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
|
||||
- **Router startup race**: Compose router.py doesn't call load_roster(). Reload thread sleeps 30s first.
|
||||
Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster().
|
||||
- **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround.
|
||||
- **VRAM (2026-07-15)**: RTX 3090 at ~17/24.6GB (~70%) with **128K context** (reduced from 256K 2026-07-17). RTX 5070 at ~7.8/12.2GB (~65%) with 128K context + MTP. Strix Halo at ~7GB/64GB.
|
||||
- **RTX 3090 (2026-07-27)**: Swapped to SmartCode-Fable-5-27B-UD-Q3_K_XL (14.7GB). Q4 was too large for 24GB VRAM with 128K context + KV cache overhead. Q3 fits at ~22.4GB (91%). Uses standard llama.cpp build b9190 (turboquant b9150 incompatible with qwen3_5 arch). Config: `-c 131072 -ctk q4_0 -ctv q4_0 --flash-attn on --cont-batching`. Sampler: `--temp 0.9 --top-p 0.95 --top-k 60 --min-p 0.0 --repeat-penalty 1.0`. Service: `/home/llmuser/llama-fable-wrapper.sh`.
|
||||
- **RTX 5070 config (2026-07-15)**: Switched to IQ4_NL + MTP draft (Q8_0) at 128K context. Gen speed: 122 tok/s. VRAM: ~7.8/12.2GB (~65%). Service: `/home/llmuser/llama-wrapper.sh`. Config: `--model gemma-4-12b-it-IQ4_NL.gguf --spec-draft-model gemma-4-12b-it-Q8_0-MTP.gguf --spec-type draft-mtp --spec-draft-n-max 4 --ctx-size 131072`.
|
||||
- **LiteLLM timeout tuning (verified 2026-07-27)**: SmartCode-Fable-5-27B 300s, gemma-4-12b 120s, qwen3.6-27B-code 300s (legacy), qwen3.6-35B-udq4 300s, strix-moe 300s, syslog-auto routes all 300s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s.
|
||||
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `qwen3.6-35B-udq4`, alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded).
|
||||
- **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill).
|
||||
- **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (old Mumuni CT114 — now inside Abiba CT100 at .24) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`.
|
||||
@@ -268,8 +252,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
|
||||
|
||||
| GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Context |
|
||||
|-----|-------|-----------|--------------|----------|---------|
|
||||
| RTX 3090 (.8) | SmartCode-Fable-5-27B-UD-Q3_K_XL | **TBD** | — | — | **128K** |
|
||||
| RTX 5070 (.110) | gemma-4-12b (IQ4_NL+MTP) | **191** | — | — | **128K** |
|
||||
| RTX 3090 (.8) | Qwen3.8-27B-Uncensored-Q4_K_M | **TBD** | — | — | **128K** |
|
||||
| Strix Halo (.15) | qwen3.6-35B-udq4 | **65** | 140 | — | **128K** |
|
||||
|
||||
Benchmarks from 2026-07-17. Strix Halo model: qwen3.6-35B-udq4. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
|
||||
@@ -296,12 +279,13 @@ When the underlying model is swapped, only the LiteLLM config changes — agent
|
||||
### Context Windows
|
||||
- RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **128K**
|
||||
- **All agents**: 128K ceiling — stable margin. For >128K workloads, use external providers (deepseek)
|
||||
- Compression threshold 0.65: fires at ~85K (~43K headroom before 128K ceiling)
|
||||
- Compression threshold 0.60: fires at ~77K (~51K headroom before 128K ceiling)
|
||||
- **Pi agents (Abiba)**: `compaction.reserveTokens: 52739` (≈60% of 128K)
|
||||
- Mumuni compression model alias: `strix-moe` with 300s timeout
|
||||
|
||||
### Mumuni Agent Profile
|
||||
|
||||
Mumuni (CT100/abiba, 192.168.68.24) is the primary business assistant. This profile is the reference for all agent configs:
|
||||
Mumuni (kagentz CT105, 192.168.68.14 — migrated from CT100 2026-08-29) is the primary business assistant. This profile is the reference for all agent configs:
|
||||
|
||||
| Setting | Value | Notes |
|
||||
|---------|-------|-------|
|
||||
@@ -313,7 +297,7 @@ Mumuni (CT100/abiba, 192.168.68.24) is the primary business assistant. This prof
|
||||
| `aux.web_extract.model` | `gpu-light` | Web extraction |
|
||||
| `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) |
|
||||
| `context.max_context_window` | 131072 (128K) | Reduced from 256K 2026-07-17 — stable 128K ceiling |
|
||||
| `compression.threshold` | 0.65 | Triggers at ~85K |
|
||||
| `compression.threshold` | 0.60 | Triggers at ~77K (~60% of 128K) — optimized for 128K context |
|
||||
| `compression.target_ratio` | 0.3 | Compresses to ~38K |
|
||||
| `compression.protect_last_n` | 40 | Preserves last 40 messages |
|
||||
| `memory.memory_char_limit` | 800 | Brief memory entries |
|
||||
@@ -326,7 +310,7 @@ Mumuni (CT100/abiba, 192.168.68.24) is the primary business assistant. This prof
|
||||
|
||||
| Agent | Host | Status |
|
||||
|-------|------|--------|
|
||||
| **Mumuni** | CT100 (.24) | ✅ Updated to stable aliases |
|
||||
| **Mumuni** | CT105 (.14) | ✅ Updated to stable aliases |
|
||||
| **Tanko** | CT112 (.122) | ✅ Updated to stable aliases |
|
||||
| **Koby** | CT111 (.129) | ❌ SSH unreachable — needs Zulip DM |
|
||||
| **Koonimo** | CT113 | ❌ SSH unreachable — needs Zulip DM |
|
||||
|
||||
@@ -1,4 +1,6 @@
|
||||
---
|
||||
report_only_agents:
|
||||
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
|
||||
kind: responsibility
|
||||
name: gpu-self-heal
|
||||
description: >
|
||||
@@ -16,6 +18,7 @@ depends_on:
|
||||
- gpu-monitor.prose.md (live data source on .24:9100)
|
||||
- gpu-fleet.prose.md (source of truth for topology, aliases, model assignments)
|
||||
---
|
||||
---
|
||||
|
||||
## Maintains
|
||||
|
||||
@@ -38,20 +41,19 @@ depends_on:
|
||||
- On fix: verify with benchmark inference test before declaring resolved
|
||||
- Escalate: after 3 failed remediation attempts → Zulip #agent-hub alert
|
||||
|
||||
---
|
||||
---
|
||||
|
||||
## Current Fleet Baseline (2026-07-18)
|
||||
|
||||
| Alias | GPU | Host | Model | VRAM | Ctx | tok/s | Role |
|
||||
|-------|-----|------|-------|------|-----|-------|------|
|
||||
| `gpu-dense` | RTX 3090 24GB | ct8 (.8:8080) | ThinkingCap Qwen3.6-27B Q4_K_M + MTP + vision | 21.6/24.6GB (88%) | 128K | 74.9 | Heavy reasoning, code gen |
|
||||
| `gpu-light` | RTX 5070 12GB | ct110 (.110:8080) | HauhauCS Gemma4-12B QAT Q4_K_M + MTP draft | 10.1/12.2GB (83%) | 128K | 169.6 | Vision, web extract, light tasks |
|
||||
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | qwen3.6-35B-udq4 | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs |
|
||||
|
||||
Key notes:
|
||||
- All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) is deprecated and NOT in the inference path.
|
||||
- Stable aliases (gpu-dense, gpu-light, strix-moe) from gpu-fleet are the canonical names for agent configs. Model-specific names still work but are deprecated.
|
||||
- RTX 5070 tok/s is 2.3x faster than RTX 3090 for its model — gpu-light is the fastest endpoint. Route vision/web/light work there first.
|
||||
- RTX 5070 tok/s is ~145 for Qwen3.5-9B — gpu-light is the fastest endpoint. Route vision/web/light work there first. NOTE: Qwen3.5-9B is multimodal (image+text), NOT text-only like gemma-4-12b was.
|
||||
- Strix Halo is 62.9 tok/s (89% of 70.5 baseline) — below optimal but stable. Check for competing workloads.
|
||||
- RTX 3090 VRAM at 88% — within role-appropriate range (role = heavy reasoning, needs the headroom).
|
||||
- RTX 5070 VRAM at 83% — role-appropriate for vision/web (smaller batch sizes).
|
||||
@@ -156,7 +158,7 @@ Key notes:
|
||||
- **Detect**: GPU roles misaligned with hardware capabilities
|
||||
- **Target distribution**:
|
||||
- RTX 3090 (gpu-dense, 24GB, 74.9 tok/s) → Heavy reasoning, code gen, long conversations (slowest per-token but largest context capacity). Weight: 0.55 (LiteLLM).
|
||||
- RTX 5070 (gpu-light, 12GB, 169.6 tok/s) → Vision/image, web search, lightweight tasks (2.3x faster than 3090 per token). Weight: 0.15 (LiteLLM).
|
||||
- RTX 5070 (gpu-light, 12GB, ~145 tok/s) → Vision (image+text), web search, lightweight tasks. Weight: 0.15 (LiteLLM).
|
||||
- Strix Halo (strix-moe, 64GB, 62.9 tok/s) → Context compression, summarization, long docs (MoE model). Weight: 0.30 (LiteLLM).
|
||||
- **Note**: RTX 5070 is the fastest endpoint per token. Route high-volume, low-complexity work there first.
|
||||
- **Fix**:
|
||||
@@ -166,6 +168,7 @@ Key notes:
|
||||
- **Verify**: Each GPU's request pattern matches its designated role within 24h
|
||||
- **Escalate**: If role mismatch persists >48h → agent alias audit needed
|
||||
|
||||
---
|
||||
---
|
||||
|
||||
## Execution
|
||||
@@ -247,6 +250,7 @@ call update-gpu-health
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
---
|
||||
|
||||
## Reporting
|
||||
@@ -263,6 +267,7 @@ Pushed to `SyslogSolution/health-logs/gpu/{run_id}.json` — versioned, searchab
|
||||
- Per-GPU tok/s trend over 7 days
|
||||
- Regression alerts if any GPU degrades >10% week-over-week
|
||||
|
||||
---
|
||||
---
|
||||
|
||||
## Design Decisions (Verified 2026-07-12, Reaffirmed 2026-07-18)
|
||||
|
||||
@@ -24,8 +24,6 @@ done
|
||||
|
||||
| Agent | CT | Node | IP | LiteLLM Alias | Key Source | Platform |
|
||||
|-------|-----|------|-----|---------------|------------|----------|
|
||||
| Tanko | 112 | amdpve | .122 | `tanko` | Infisical vault | Hermes |
|
||||
| Mumuni | 100 | hwepve | .24 | `mumuni` | Infisical vault | Hermes |
|
||||
| Koby | 111 | amdpve | .129 | `koby` | Infisical vault | **Hermes** |
|
||||
| Koonimo | 113 | amdpve | .114 | `koonimo` | Infisical vault | Hermes |
|
||||
| Shumba | — | 192.168.68.119 | N/A | N/A (DeepSeek) | Hermes (RETIRED — CT119 now Infisical vault) |
|
||||
@@ -53,9 +51,10 @@ Agent (systemd) → LITELLM_API_KEY → LiteLLM (:116/v1) → Router (:9000) →
|
||||
|
||||
## Config Pattern — Mandatory Fields
|
||||
|
||||
### For Hermes Agents (Tanko, Mumuni, Koonimo)
|
||||
### For Hermes Agents (Mumuni, Koonimo)
|
||||
|
||||
Every agent's `/root/.hermes/config.yaml` (or `/home/jerome/.hermes/config.yaml`) MUST have:
|
||||
Every Hermes agent's `/root/.hermes/config.yaml` (or `/home/jerome/.hermes/config.yaml`) MUST have:
|
||||
(Tanko is excluded — migrated to DSH/DeepSeek Harness on 2026-08-27, no longer uses Hermes config.)
|
||||
|
||||
### 1. Main Model
|
||||
```yaml
|
||||
@@ -71,7 +70,7 @@ model:
|
||||
custom_providers:
|
||||
- name: harness
|
||||
model: syslog-auto
|
||||
base_url: http://192.168.68.116/v1
|
||||
base_url: http://192.168.68.116/litellm/v1
|
||||
api_key_env: LITELLM_API_KEY
|
||||
api_mode: chat_completions
|
||||
```
|
||||
@@ -82,7 +81,7 @@ auxiliary:
|
||||
vision:
|
||||
provider: harness
|
||||
model: gemma-4-12b # or syslog-auto
|
||||
base_url: http://192.168.68.116/v1
|
||||
base_url: http://192.168.68.116/litellm/v1
|
||||
api_key_env: LITELLM_API_KEY
|
||||
api_key: <value from: infisical secrets get LITELLM_API_KEY --project=agents --env=production> # ← MANDATORY workaround
|
||||
timeout: 60
|
||||
@@ -97,7 +96,7 @@ auxiliary:
|
||||
target_ratio: 0.3
|
||||
provider: harness
|
||||
model: syslog-auto # or gemma-4-12b
|
||||
base_url: http://192.168.68.116/v1
|
||||
base_url: http://192.168.68.116/litellm/v1
|
||||
api_key_env: LITELLM_API_KEY
|
||||
api_key: <value from: infisical secrets get LITELLM_API_KEY --project=agents --env=production> # ← MANDATORY workaround
|
||||
timeout: 120
|
||||
@@ -169,11 +168,15 @@ pct-run <CT> grep -A8 "vision:" /root/.hermes/config.yaml | grep api_key
|
||||
# Must show both api_key: sk-... and api_key_env: LITELLM_API_KEY
|
||||
```
|
||||
|
||||
### For Koby (CT 111 / tdunna)
|
||||
### For Koby (CT 111 / tdunna) — **REPORT-ONLY MODE**
|
||||
|
||||
Koby runs Hermes on CT 111 (tdunna). Config files at `/root/.hermes/config.yaml`.
|
||||
Same Hermes pattern as Tanko/Mumuni/Koonimo — see config sections above.
|
||||
|
||||
**⛔ KOBY IS NEVER REPAIRED (2026-08-17, Captain)**: Diagnostic only — detect and report, never fix on .129.
|
||||
No heal step, no restart, no key rotation, no config edit, no memory rewrite, no disk GC, no service touch, no process kill — ever.
|
||||
If a health check shows Koby degraded, **DO NOT** execute any repair action. Instead, report to Zulip and let Theo fix it.
|
||||
|
||||
**LiteLLM key**: alias `koby` in LiteLLM DB, injected via `infisical run --` wrapper.
|
||||
|
||||
### For pi Agents (Abiba)
|
||||
|
||||
@@ -15,7 +15,7 @@ description: >
|
||||
|
||||
- template_version: "2.1.0"
|
||||
- last_applied: timestamp
|
||||
- agents_configured: ["tanko", "mumuni", "abiba", "koby", "koonimo", "kagenz0"]
|
||||
- agents_configured: ["mumuni", "abiba", "koby", "koonimo", "kagenz0"] # tanko removed 2026-08-27 (now on DSH/DeepSeek Harness)
|
||||
- agent_keys: map (see Agent Keys section)
|
||||
- infra_endpoints_verified: array
|
||||
|
||||
@@ -30,8 +30,6 @@ Sub-agent profiles inherit auth from the main config — no separate keys needed
|
||||
|
||||
| Agent | Key Alias | Host | SSH | Sub-Agents |
|
||||
|-------|-----------|------|-----|-----------|
|
||||
| Tanko | `tanko-*` | 192.168.68.122 | jerome@.122 | — |
|
||||
| Mumuni | `mumuni` | 192.168.68.24 | root@.24 | 6 profiles ✱ |
|
||||
| Abiba | `abiba-pi` | 192.168.68.24 | local | — |
|
||||
| Koby | `koby` | CT 111 (tdunna) | Zulip | — |
|
||||
| Koonimo | `koonimo` | CT 113 (baggy) | SSH root | — |
|
||||
@@ -348,6 +346,17 @@ curl -s -o /dev/null -w 'key_health: %{http_code}\n' -H "Authorization: Bearer $
|
||||
|
||||
### Rule 13: API Key Injection — Two Patterns (UPDATED 2026-07-16, WAL #1300)
|
||||
|
||||
### Rule 14: Hermes Context Detection Uses `max_model_tokens`, NOT `max_input_tokens`
|
||||
|
||||
**CRITICAL**: Hermes context detection reads `max_model_tokens` (128K), NOT `max_input_tokens` (64K cap).
|
||||
|
||||
- **Abiba and Hermes agents**: `max_model_tokens: 131072` (128K) — unlimited context
|
||||
- **Crewmates (ops, tune, verify, auth-keys, build)**: `max_input_tokens: 64000` (64K) — capped
|
||||
- If you see `max_input_tokens: 64000` in an Abiba/Hermes config, that's a mistake
|
||||
- Using `max_input_tokens` for Hermes agents causes premature context loss
|
||||
- Check: `grep -n 'max_model_tokens\|max_input_tokens' ~/.hermes/config.yaml`
|
||||
- Expected output: `max_model_tokens: 131072` (not max_input_tokens)
|
||||
|
||||
Agents inject `LITELLM_API_KEY` via ONE of two mechanisms. Both are valid; the contract
|
||||
requirement is that the key is a **valid LiteLLM virtual key** (HTTP 200 on /v1/models).
|
||||
|
||||
@@ -424,13 +433,3 @@ curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $K" http://192.
|
||||
5. **Set model choice** — Per agent's workload
|
||||
6. **Verify** — curl all shared endpoints, test the model with the new key
|
||||
7. **Report** — What was changed, preserved, custom
|
||||
### Rule 16: Koby Configuration (DeepSeek-primary)
|
||||
(Ref: See Rule 10 for default model behavior, with Koby exception)
|
||||
|
||||
Koby uses a split-model architecture:
|
||||
|
||||
- Primary Model: `deepseek-v4-flash` via `api.deepseek.com` (for reasoning)
|
||||
- Auxiliary Models: `gpu-light` (vision/web_extract) and `syslog-auto` (compression)
|
||||
- Key Hygiene: `api_key_env` is strictly `LITELLM_API_KEY` or `DEEPSEEK_API_KEY`
|
||||
- Constraint: Do NOT touch Koby's primary model/provider/compression settings unless explicitly ruled by the captain.
|
||||
|
||||
|
||||
@@ -190,7 +190,7 @@ litellm_settings:
|
||||
| Agent | CT | IP | LiteLLM Alias | Key Source | Status | Gateway Wrapper | Last Verified |
|
||||
|-------|-----|-----|---------------|------------|--------|-----------------|---------------|
|
||||
| Tanko | 112 | .122 | `tanko` | Infisical vault | ✅ Fixed | `infisical run` | 20:17 UTC Jul 5 |
|
||||
| Mumuni | 100 (abiba) | .24 | `mumuni` | Infisical vault | ✅ Fixed | Pi Hermes gateway | 2026-07-27 |
|
||||
| Mumuni | 105 (kagentz) | .14 | `mumuni` | Infisical vault | ✅ Fixed | systemd Hermes gateway | 2026-08-29 |
|
||||
| Koby | 111 | .129 | `koby` | Infisical vault | ✅ Fixed (DeepSeek-primary) | `infisical run` | 23:30 UTC Jul 5 |
|
||||
| Koonimo | 113 | .114 | `koonimo` | Infisical vault | ✅ Fixed | `infisical run` (migrated 2026-07-11) | 2026-08-09 |
|
||||
| Abiba | 100 | .65 | `abiba-pi` | Infisical vault | ✅ N/A (pi native) | — | 19:44 UTC Jul 5 |
|
||||
@@ -285,13 +285,13 @@ auxiliary:
|
||||
vision:
|
||||
api_key: sk-<agent-key-from-vault> # ← workaround (get via: infisical secrets get LITELLM_API_KEY --project=agents --env=production --plain)
|
||||
api_key_env: LITELLM_API_KEY
|
||||
base_url: http://192.168.68.116/v1
|
||||
base_url: http://192.168.68.116/litellm/v1
|
||||
model: gemma-4-12b
|
||||
provider: harness
|
||||
compression:
|
||||
api_key: sk-<agent-key-from-vault> # ← workaround (same as above)
|
||||
api_key_env: LITELLM_API_KEY
|
||||
base_url: http://192.168.68.116/v1
|
||||
base_url: http://192.168.68.116/litellm/v1
|
||||
model: gemma-4-12b
|
||||
provider: harness
|
||||
```
|
||||
|
||||
@@ -28,7 +28,7 @@ connectivity recovery including end-to-end DM validation.
|
||||
|
||||
| Param | Type | Required | Default | Description |
|
||||
|-------|------|----------|---------|-------------|
|
||||
| `target` | string | yes | — | Agent name: `mumuni`, `tanko`, `koby`, or `shumba` |
|
||||
| `target` | string | yes | — | Agent name: `mumuni`, `koby`, or `shumba` (Tanko excluded — on DSH since 2026-08-27, no Hermes plugin) |
|
||||
| `branch` | string | no | `master` | Git branch to pull (overridable for pinning) |
|
||||
|
||||
## Maintains
|
||||
@@ -55,8 +55,7 @@ connectivity recovery including end-to-end DM validation.
|
||||
|
||||
| Host | CT | Proxmox | IP (direct) | Hermes Home | User |
|
||||
|------|-----|---------|-------------|-------------|------|
|
||||
| Mumuni | CT100 | — | 192.168.68.24 | /root/.hermes | root |
|
||||
| Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome |
|
||||
| Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome | *(DSH since 2026-08-27 — historical, plugin retired on this host)* |
|
||||
| Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root |
|
||||
| Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky |
|
||||
|
||||
@@ -121,7 +120,8 @@ cp plugins/platforms/zulip/adapter.py \
|
||||
plugins/platforms/zulip/plugin.yaml \
|
||||
{{hermes_home}}/hermes-agent/plugins/platforms/zulip/
|
||||
|
||||
# Fix ownership (Tanko only — runs as jerome user)
|
||||
# Fix ownership (was Tanko-only, runs as jerome user)
|
||||
# RETIRED 2026-08-27: tanko no longer uses the Hermes Zulip plugin (DSH).
|
||||
[ "{{target}}" = "tanko" ] && chown -R jerome:jerome \
|
||||
{{hermes_home}}/hermes-agent/plugins/platforms/zulip/
|
||||
|
||||
|
||||
@@ -1,9 +1,9 @@
|
||||
---
|
||||
report_only_agents:
|
||||
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
|
||||
kind: function
|
||||
name: hermes-zulip-restore
|
||||
description: >
|
||||
Restores Zulip connectivity for any Hermes agent (Mumuni CT100, Tanko CT112,
|
||||
Koby CT111, Shumba on Lucky's mini PC). Deploys the zulip-platform adapter to the correct bundled plugin
|
||||
path, verifies env credentials, restarts the gateway, and confirms Zulip
|
||||
connects. Run this whenever a Hermes agent stops responding on Zulip or after
|
||||
a fresh agent deployment.
|
||||
@@ -12,6 +12,7 @@ version: 1.0.0
|
||||
status: active
|
||||
runtime_contract: 2
|
||||
---
|
||||
---
|
||||
|
||||
# Hermes Zulip Restore — Bring Any Agent Back to Good State
|
||||
|
||||
@@ -23,7 +24,7 @@ gateway restart, and connection validation.
|
||||
|
||||
| Param | Type | Required | Default | Description |
|
||||
|-------|------|----------|---------|-------------|
|
||||
| `target` | string | yes | — | Agent name: `mumuni`, `tanko`, `koby`, or `shumba` |
|
||||
| `target` | string | yes | — | Agent name: `mumuni`, `koby`, or `shumba` (Tanko excluded — DSH since 2026-08-27) |
|
||||
|
||||
## Maintains
|
||||
|
||||
@@ -36,7 +37,7 @@ gateway restart, and connection validation.
|
||||
|
||||
- `_strip_html` function present in `<hermes-agent>/plugins/platforms/zulip/adapter.py`
|
||||
- All three adapter files (__init__.py, adapter.py, plugin.yaml) present at bundled path
|
||||
- Zulip env vars set in `~/.hermes/.env` (or `/home/jerome/.hermes/.env` for Tanko)
|
||||
- Zulip env vars set in `~/.hermes/.env` (or `/home/jerome/.hermes/.env` for Tanko, historical — DSH since 2026-08-27)
|
||||
- Gateway restarted and zulip platform reports state `connected`
|
||||
- HTML stripping enabled for `/approve` and `/deny` slash command support
|
||||
|
||||
@@ -51,8 +52,6 @@ gateway restart, and connection validation.
|
||||
|
||||
| Host | CT | Proxmox | IP (direct) | Hermes Home | User |
|
||||
|------|-----|---------|-------------|-------------|------|
|
||||
| Mumuni | CT100 (abiba) | hwepve | 192.168.68.24 | /root/.hermes | root |
|
||||
| Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome |
|
||||
| Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root |
|
||||
| Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky |
|
||||
|
||||
@@ -94,8 +93,8 @@ cp zulip-platform-plugins/plugins/platforms/zulip/adapter.py \
|
||||
zulip-platform-plugins/plugins/platforms/zulip/plugin.yaml \
|
||||
<HERMES_HOME>/hermes-agent/plugins/platforms/zulip/
|
||||
|
||||
# Fix ownership (Tanko only)
|
||||
chown -R jerome:jerome <HERMES_HOME>/hermes-agent/plugins/platforms/zulip/ # Tanko only
|
||||
# Fix ownership (was Tanko-only; RETIRED 2026-08-27 — tanko on DSH, no Hermes plugin)
|
||||
chown -R jerome:jerome <HERMES_HOME>/hermes-agent/plugins/platforms/zulip/ # Tanko only (historical)
|
||||
|
||||
# Clean up
|
||||
rm -rf /tmp/zulip-deploy
|
||||
@@ -184,6 +183,7 @@ https://git.sysloggh.net/SyslogSolution/zulip-platform-plugins/src/branch/feat/z
|
||||
Commit `55ca15d` — `fix(zulip): add _strip_html for slash command matching`
|
||||
Pull request #33 is the primary integration branch.
|
||||
|
||||
---
|
||||
---
|
||||
|
||||
**Last verified good state**: 2026-07-08 — Mumuni, Tanko, Koby all connected with `_strip_html` applied.
|
||||
|
||||
@@ -87,8 +87,8 @@ call apply-liteLLM-routing
|
||||
|
||||
call apply-agent-compression
|
||||
agent: mumuni
|
||||
host: 192.168.68.24
|
||||
config_path: /root/.hermes/config.yaml
|
||||
host: 192.168.68.14
|
||||
config_path: /home/hermes/.hermes/config.yaml
|
||||
|
||||
-- Phase 4: Enable llama.cpp prompt caching on GPU hosts
|
||||
|
||||
|
||||
@@ -3,7 +3,7 @@ kind: pattern
|
||||
name: infrastructure-control
|
||||
description: >
|
||||
Full infrastructure monitoring and control pattern covering the
|
||||
6-node Proxmox cluster, 3 Docker ecosystems (22 containers),
|
||||
5-node Proxmox cluster, 3 Docker ecosystems (22 containers),
|
||||
NFS storage, and network services. Defines monitors, remediations,
|
||||
and the access matrix for all environments.
|
||||
|
||||
@@ -13,10 +13,11 @@ description: >
|
||||
against the live system. Policy fields are authoritative. See the
|
||||
`verify-before-mutate` skill.
|
||||
|
||||
**Last verified:** 2026-07-24 — corrected Gitea IP (.17 not .110),
|
||||
AdGuard IP (.10 not .102), AdGuard placement (minipve not acerpve),
|
||||
Abiba placement (hwepve not amdpve), added hwepve as 6th node,
|
||||
added dns.sysloggh.net route.
|
||||
**Last verified:** 2026-08-15 — hwepve removed from Tabiri cluster
|
||||
(now 5 nodes: minipve, amdpve, storepve, acerpve, ocupve). hwepve
|
||||
(192.168.68.4) is a standalone PVE node + NetBird routing peer;
|
||||
London relocation pending. CTs 100 (abiba) and 105 (kagentz) moved
|
||||
to minipve.
|
||||
---
|
||||
|
||||
# Infrastructure Control Pattern
|
||||
@@ -34,21 +35,21 @@ description: >
|
||||
┌─────────────┐ ┌──────────────┐ ┌──────────────┐
|
||||
│ Abiba │ │ Tanko │ │ Mumuni │
|
||||
│ (pi) │ │ (Hermes) │ │ (Hermes) │
|
||||
│ CT 100 │ │ CT 112 │ │ CT 114 │
|
||||
│ CT 100 │ │ CT 112 │ │ CT 100 │
|
||||
└──────┬──────┘ └──────┬───────┘ └──────┬───────┘
|
||||
│ │ │
|
||||
└──────────────────┼────────────────────┘
|
||||
▼
|
||||
┌──────────────────────────────────────┐
|
||||
│ Proxmox Cluster API │
|
||||
│ minipve.sysloggh.net:443 │
|
||||
│ (monitoring@pve!mumuni token) │
|
||||
└────┬──────┬──────┬──────┬──────┬─────┘
|
||||
│ │ │ │ │
|
||||
┌────┘ ┌────┘ ┌────┘ ┌────┘ ┌────┘
|
||||
▼ ▼ ▼ ▼ ▼
|
||||
minipve amdpve storepve acerpve ocupve hwepve
|
||||
(.12) (.15) (.6) (.9) (.5) (.4)
|
||||
┌────────────────────────────────────┐
|
||||
│ Proxmox Cluster API │
|
||||
│ minipve.sysloggh.net:443 │
|
||||
│ (monitoring@pve!mumuni token) │
|
||||
└────┬──────┬──────┬──────┬──────────┘
|
||||
│ │ │ │
|
||||
┌────┘ ┌────┘ ┌────┘ ┌────┘
|
||||
▼ ▼ ▼ ▼
|
||||
minipve amdpve storepve acerpve ocupve
|
||||
(.12) (.15) (.6) (.9) (.5)
|
||||
|
||||
▼
|
||||
┌─────────────────────────────────────────────┐
|
||||
@@ -100,21 +101,29 @@ description: >
|
||||
|
||||
## Section 2: Proxmox Cluster — Monitoring
|
||||
|
||||
### Nodes (6)
|
||||
### Nodes (5)
|
||||
|
||||
| Node | IP | CPU | RAM | VMs/CTs | Role |
|
||||
|------|----|-----|-----|---------|------|
|
||||
| minipve | .12 | 16C | 30GB | authentik, gitea, syslog-api, infisical-vault, jitsi | Auth, git, messaging |
|
||||
| minipve | .12 | 16C | 30GB | abiba, kagentz, authentik, gitea, syslog-api, infisical-vault, jitsi | Auth, git, messaging |
|
||||
| amdpve | .15 | 32C | 62GB | tanko, tdunna, baggy, scottdenya | Agents, compute |
|
||||
| storepve | .6 | 28C | 31GB | docker-vm, ra-h-os, PBS, media, jdownloader, zulip | Docker, storage, chat |
|
||||
| acerpve | .9 | 28C | 31GB | llm-gpu | GPU VMs |
|
||||
| ocupve | .5 | 12C | 14GB | ocu-llm | GPU VMs |
|
||||
| hwepve | .4 | 12C | 15GB | abiba, kagentz, (mumuni CT 114 stopped) | Agents (new node) |
|
||||
|
||||
> **Note:** CTs on storepve include jdownloader (CT 118). AdGuard (CT 102) is on
|
||||
> minipve at .10, not acerpve. Abiba (CT 100) is on hwepve, not amdpve. Mumuni
|
||||
> (CT 114) is on hwepve (currently stopped), not minipve. Mumuni also has a
|
||||
> second instance on minipve at .123 — distinguish by CT ID, not hostname.
|
||||
> minipve at .10, not acerpve. Abiba (CT 100) and kagentz (CT 105) are on
|
||||
> minipve (moved from hwepve 2026-08-15). Mumuni runs inside Abiba CT100
|
||||
> (.24); CT 114 (mumuni) no longer exists in the cluster.
|
||||
>
|
||||
> **hwepve (192.168.68.4) — STANDALONE (removed from Tabiri 2026-08-15):**
|
||||
> Huawei MateBook 16 (KLVL-WXX9), pve-manager/9.2.10, kernel 7.0.14-8-pve.
|
||||
> Zero VMs/CTs. Being relocated to London as a standalone PVE node + NetBird
|
||||
> routing peer (relocation pending). Localizations applied: timezone
|
||||
> Europe/London, lid-switch ignore, sleep/suspend/hibernate targets masked,
|
||||
> cluster-shared storage removed (remaining: local, local-lvm, storage,
|
||||
> mediastore). prometheus-node-exporter active on :9100; net.ipv4.ip_forward=1;
|
||||
> NetBird client not yet installed (enrollment pending setup key).
|
||||
|
||||
### Checks (every 5 min)
|
||||
|
||||
@@ -592,21 +601,20 @@ ssh root@192.168.68.110 "systemctl restart llama-server"
|
||||
|
||||
| CT | Name | Node | IP | Role | Agent |
|
||||
|----|------|------|----|------|-------|
|
||||
| 100 | abiba | **hwepve** | .24 | Pi agent | ✅ pi |
|
||||
| 100 | abiba | minipve | .24 | Pi agent | ✅ pi |
|
||||
| 101 | llm-gpu | acerpve | .8 | GPU RTX 3090 | ❌ |
|
||||
| 102 | adguard | **minipve** | **.10** | DNS | ❌ |
|
||||
| 103 | ocu-llm | ocupve | .110 | GPU RTX 5070 | ❌ |
|
||||
| 104 | authentik | minipve | .11 | OIDC | ❌ |
|
||||
| 105 | kagentz | **hwepve** | — | Agent Zero | ✅ |
|
||||
| 105 | kagentz | minipve | — | Agent Zero | ✅ |
|
||||
| 106 | ra-h-os | storepve | .65 | KG bridge | ✅ MCP |
|
||||
| 107 | pbs | storepve | — | Backups | ❌ |
|
||||
| 108 | media | storepve | — | Media | ❌ |
|
||||
| 109 | docker-vm | storepve | .7 | Docker host | ❌ |
|
||||
| 110 | gitea | minipve | **.17** | Git | ❌ |
|
||||
| 111 | tdunna | amdpve | .129 | Hermes agent | ✅ |
|
||||
| 112 | tanko | amdpve | .122 | Hermes agent | ✅ |
|
||||
| 112 | tanko | amdpve | .122 | DSH (DeepSeek Harness) agent | ✅ |
|
||||
| 113 | baggy | amdpve | .114 | Hermes agent | ✅ |
|
||||
| 114 | mumuni | **hwepve** | .123 | Hermes agent (stopped) | ✅ |
|
||||
| 115 | scottdenya | amdpve | .75 | Denya OneCare | ❌ |
|
||||
| 116 | syslog-api | minipve | .116 | LiteLLM + Grafana | ❌ |
|
||||
| 117 | zulip | storepve | .19 | Chat | ❌ |
|
||||
@@ -631,15 +639,14 @@ Source of truth: `/root/scripts/pct-run.sh` or `prose-contracts/scripts/pct-run.
|
||||
|
||||
| CT | Name | Node | pct-run |
|
||||
|-----|------|------|---------|
|
||||
| 100 | abiba | hwepve | `pct-run 100` |
|
||||
| 105 | kagentz | hwepve | `pct-run 105` |
|
||||
| 100 | abiba | minipve | `pct-run 100` |
|
||||
| 105 | kagentz | minipve | `pct-run 105` |
|
||||
| 111 | tdunna | amdpve | `pct-run 111` |
|
||||
| 112 | tanko | amdpve | `pct-run 112` |
|
||||
| 113 | baggy | amdpve | `pct-run 113` |
|
||||
| 115 | scottdenya | amdpve | `pct-run 115` |
|
||||
| 104 | authentik | minipve | `pct-run 104` |
|
||||
| 110 | gitea | minipve | `pct-run 110` |
|
||||
| 114 | mumuni | hwepve | `pct-run 114` |
|
||||
| 116 | syslog-api | minipve | `pct-run 116` |
|
||||
| 106 | ra-h-os | storepve | `pct-run 106` |
|
||||
| 107 | proxmox-backup | storepve | `pct-run 107` |
|
||||
@@ -652,7 +659,7 @@ GPU bare-metal hosts (.8 acerpve, .110 ocupve, .15 amdpve) are NOT CTs — use S
|
||||
ssh root@192.168.68.8 # RTX 3090
|
||||
ssh root@192.168.68.110 # RTX 5070
|
||||
ssh root@192.168.68.15 # Strix Halo
|
||||
ssh root@192.168.68.4 # hwepve (abiba, kagentz, mumuni)
|
||||
ssh root@192.168.68.4 # hwepve — standalone London node + NetBird routing peer (relocation pending)
|
||||
```
|
||||
|
||||
## Section 7: Agent Health Check (consolidated — 2026-07-05)
|
||||
|
||||
@@ -140,7 +140,6 @@ After ALL updates (apt + images + restarts), verify every critical service is ba
|
||||
| Zulip | `curl -sf https://chat.sysloggh.net/api/v1/server_settings` | 200 OK |
|
||||
| Gitea | `curl -sf https://git.sysloggh.net/api/v1/version` | 200 OK |
|
||||
| PM2 processes | `pm2 jlist` (CT 100) | all pi-agent processes `online` |
|
||||
| Hermes gateways | SSH to Mumuni CT 100, Tanko CT 112; `systemctl is-active hermes-gateway` | `active` for each |
|
||||
|
||||
Regression check: every service that was GREEN in `health-baseline` must still be GREEN. A service that was already RED (and caused a preflight abort) is excluded — but Phase 0 should have aborted before we got here.
|
||||
|
||||
|
||||
@@ -102,6 +102,39 @@ GPU .8 (RTX 3090) GPU .110 (RTX 5070) GPU .15 (Strix Halo)
|
||||
- Stack persists across reboots (systemd for exporters, Docker restart policy)
|
||||
|
||||
## Execution
|
||||
### check-health
|
||||
|
||||
**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.**
|
||||
|
||||
```bash
|
||||
# Zulip API health (POST ping)
|
||||
curl -s -o /dev/null -w '%{http_code}' -X POST https://chat.sysloggh.net/api/v1/messages -u 'abiba-bot@chat.sysloggh.net:KEY'
|
||||
# Expected: 200 (HTTP 000 = unreachable/cache)
|
||||
|
||||
# PM2 process health
|
||||
pm2 jlist
|
||||
# Expected: 5/5 online (abiba-telegram, abiba-zulip, zulip-watchdog, gitea-runner, spoton-service)
|
||||
|
||||
# GPU exporters (may be down per DEPLOYMENT STATUS)
|
||||
curl -s http://192.168.68.8:9400/metrics && echo " - OK" || echo " - FAIL"
|
||||
curl -s http://192.168.68.110:9400/metrics && echo " - OK" || echo " - FAIL"
|
||||
curl -s http://192.168.68.15:9400/metrics && echo " - OK" || echo " - FAIL"
|
||||
|
||||
# Prometheus targets
|
||||
curl -s http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets'
|
||||
# Expected: All targets UP (may show some down if exporters not deployed)
|
||||
|
||||
# Grafana health
|
||||
curl -s http://192.168.68.116:3001/api/health | jq '{status, version}'
|
||||
# Expected: {"status":"ok","version":"..."}
|
||||
|
||||
# LiteLLM metrics
|
||||
curl -s http://192.168.68.116:4001/metrics | head -20
|
||||
# Expected: Prometheus-formatted metrics output
|
||||
```
|
||||
|
||||
**Report format**: Summarize actual results from each probe. If any probe returns non-200 or empty output, flag as alert.
|
||||
|
||||
|
||||
### Phase 1: GPU Exporters
|
||||
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
kind: responsibility
|
||||
name: infrastructure-update
|
||||
description: >
|
||||
Autonomous system-wide update contract covering all 6 Proxmox nodes,
|
||||
Autonomous system-wide update contract covering all 5 Proxmox nodes,
|
||||
15+ containers/VMs, and 4 Docker ecosystems. Updates apt packages,
|
||||
Docker images, and container stacks in safe waves with health checks
|
||||
and automatic rollback on failure.
|
||||
@@ -56,11 +56,10 @@ Before ANY update wave:
|
||||
| amdpve (.15) | Proxmox node | `apt update && apt upgrade -y` | 5 min |
|
||||
| acerpve (.9) | Proxmox node | `apt update && apt upgrade -y` | 5 min |
|
||||
| ocupve (.5) | Proxmox node | `apt update && apt upgrade -y` | 5 min |
|
||||
| hwepve (.4) | Proxmox node | `apt update && apt upgrade -y` | 5 min |
|
||||
| CT 100 (.24) | Abiba (pi) | `apt update && apt upgrade -y` | 3 min |
|
||||
| CT 116 (.116) | syslog-api (LiteLLM host) | `apt update && apt upgrade -y` | 3 min |
|
||||
| CT 112 (tanko, amdpve) | Tanko | `apt update && apt upgrade -y` | 3 min |
|
||||
| CT 100 (mumuni/abiba, hwepve) | Mumuni | `apt update && apt upgrade -y` | 3 min |
|
||||
| CT 105 (kagentz, minipve) | Mumuni | `apt update && apt upgrade -y` | 3 min |
|
||||
| VM 101 (.8) | llm-gpu (RTX 3090) | `apt update && apt upgrade -y` | 3 min |
|
||||
| VM 103 (.110) | ocu-llm (RTX 5070) | `apt update && apt upgrade -y` | 3 min |
|
||||
|
||||
@@ -163,9 +162,9 @@ Before Wave 1, snapshot these files:
|
||||
/etc/systemd/system/strix-server.service (amdpve .15 — strix-moe)
|
||||
/etc/systemd/system/llama-server.service (VM 101 .8, VM 103 .110)
|
||||
# Hermes agent configs (key enforcement — 2026-07-10)
|
||||
/root/.hermes/config.yaml (Mumuni CT 114, Tanko CT 112, etc.)
|
||||
/root/.config/systemd/user/hermes-gateway.service (Mumuni CT 114 — EnvironmentFile fixed)
|
||||
/etc/environment (Mumuni CT 114 — LITELLM_API_KEY)
|
||||
/home/hermes/.hermes/config.yaml (Mumuni kagentz CT105; Tanko CT112 uses /home/jerome/.hermes)
|
||||
/etc/systemd/system/hermes-gateway.service (Mumuni kagentz CT105 — system unit, User=hermes)
|
||||
/etc/environment (LITELLM_API_KEY — legacy path, Mumuni now keys via Infisical)
|
||||
```
|
||||
|
||||
## MCP Gateway (2026-07-10)
|
||||
@@ -219,7 +218,7 @@ When LiteLLM is upgraded to a version supporting per-key MCP grants:
|
||||
|
||||
## Success Criteria
|
||||
|
||||
- [ ] All 6 PVE nodes updated, no reboot-loop
|
||||
- [ ] All 5 PVE nodes updated, no reboot-loop
|
||||
- [ ] All VMs/CTs running post-update
|
||||
- [ ] All Docker containers healthy (VM 109 + CT 116 + CT 117)
|
||||
- [ ] LiteLLM inference passing (syslog-auto test)
|
||||
@@ -235,7 +234,7 @@ After completion, send Zulip DM:
|
||||
```
|
||||
📋 Infrastructure Update — YYYY-MM-DD
|
||||
|
||||
Updated: 6 PVE nodes, 12 CTs/VMs, 30+ containers
|
||||
Updated: 5 PVE nodes, 12 CTs/VMs, 30+ containers
|
||||
Security fixes: N CVEs patched
|
||||
Downtime: <service> <duration>
|
||||
Failures: none / <details>
|
||||
|
||||
@@ -178,7 +178,7 @@ through its agent wrapper.
|
||||
| Agent | Host | Pattern | Keys | Status |
|
||||
|-------|------|---------|------|--------|
|
||||
| abiba | .24 | pi agent wrapper | ABIBA_LITELLM_API_KEY + ABIBA_ZULIP_API_KEY | ✅ vault-backed |
|
||||
| mumuni | .24 (CT100 abiba) | Pi Hermes gateway (no systemd) | MUMUNI_LITELLM_API_KEY + MUMUNI_ZULIP_API_KEY | ✅ vault-backed + .env fallback |
|
||||
| mumuni | .14 (kagentz CT105) | systemd unit hermes-gateway.service (user hermes) | MUMUNI_LITELLM_API_KEY + MUMUNI_ZULIP_API_KEY | ✅ vault-backed + .env fallback |
|
||||
| tanko | .122 | systemd drop-in + while-true wrapper + st.8e848433 (user jerome) | TANKO_LITELLM_API_KEY + TANKO_ZULIP_API_KEY | ✅ vault-backed + .env fallback |
|
||||
| koby | .129 | systemd drop-in + while-true wrapper + st.8e848433 | KOBY_LITELLM_API_KEY, shares TANKO_ZULIP_API_KEY (tanko-bot) | ✅ vault-backed |
|
||||
| koonimo | .114 | systemd drop-in + while-true wrapper + st.8e848433 | KOONIMO_LITELLM_API_KEY + KOONIMO_ZULIP_API_KEY | ✅ vault-backed |
|
||||
@@ -257,9 +257,51 @@ reads use per-agent identities. This eliminates the single shared token risk.
|
||||
| Agent | .env Keys |
|
||||
|-------|-----------|
|
||||
| Mumuni | MUMUNI_LITELLM_API_KEY, MUMUNI_ZULIP_API_KEY |
|
||||
| Tanko | TANKO_LITELLM_API_KEY, TANKO_ZULIP_API_KEY |
|
||||
| Koby | (wrapper injects from vault — .env has Telegram token) |
|
||||
| Koonimo | KOONIMO_LITELLM_API_KEY, KOONIMO_ZULIP_API_KEY |
|
||||
|| Tanko | TANKO_LITELLM_API_KEY, TANKO_ZULIP_API_KEY |
|
||||
|| Koby | (wrapper injects from vault — .env has Telegram token) |
|
||||
|| Koonimo | KOONIMO_LITELLM_API_KEY, KOONIMO_ZULIP_API_KEY |
|
||||
|| Agent Zero (kagentz .14) | OPENROUTER_API_KEY (direct OpenRouter access) |
|
||||
|
||||
### Agent Zero (kagentz .14) — OpenRouter Integration (2026-09-01)
|
||||
|
||||
Agent Zero runs in Docker on kagentz (CT105) and uses **direct OpenRouter API access**,
|
||||
not via the LiteLLM proxy. This is because Agent Zero's workflow (self-update manager,
|
||||
UI bootstrap, model selection) is built around OpenRouter's native authentication.
|
||||
|
||||
**Key Storage:**
|
||||
- **Container**: `/a0/usr/.env` (line ~72: `API_KEY_OPENROUTER=sk-or-v1-…`)
|
||||
- **Vault**: Infisical secret `OPENROUTER_API_KEY` (project=agents, env=production)
|
||||
- **Fallback**: The container's .env is the primary source; vault sync is optional
|
||||
(unlike fleet agents which require vault injection)
|
||||
|
||||
**Current Key (2026-09-01):**
|
||||
- **Prefix**: `sk-or-v1-0af3f3…`
|
||||
- **User**: `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`
|
||||
- **Plan**: Paid (not free tier)
|
||||
- **Usage**: 0 (as of 2026-09-01)
|
||||
|
||||
**Model Configuration:**
|
||||
- **Preset**: "Cost Efficient" (`/a0/usr/plugins/_model_config/presets.yaml`)
|
||||
- **Model**: `openrouter/moonshotai/kimi-k3`
|
||||
- **API Base**: (empty — uses OpenRouter default)
|
||||
|
||||
**Why not LiteLLM proxy?**
|
||||
Agent Zero's architecture was designed before the fleet adopted the LiteLLM proxy
|
||||
standard. The container runs `/exe/self_update_manager.py` and `/a0/run_ui.py` which
|
||||
directly call OpenRouter via Python's requests library. Converting would require:
|
||||
1. Refactoring all LLM calls to use `litellm` library
|
||||
2. Adding vault wrapper injection
|
||||
3. Updating self_update_manager to use proxy-aware key handling
|
||||
|
||||
**Rotation Procedure:**
|
||||
1. Generate new key in OpenRouter UI
|
||||
2. Update container: `sed -i 's/^API_KEY_OPENROUTER=.*/API_KEY_OPENROUTER=<new_key>/' /a0/usr/.env`
|
||||
3. Update vault: `infisical secrets set OPENROUTER_API_KEY=<new_key> --projectId=agents --env=production`
|
||||
4. Restart container: `sudo docker exec agent-zero supervisorctl restart run_ui`
|
||||
5. Verify: `curl -s https://openrouter.ai/api/v1/auth/key -H "Authorization: Bearer <new_key>"`
|
||||
|
||||
**Related Contract:**
|
||||
- `agent-zero-openrouter-key.prose.md` — Full agent-zero key management contract
|
||||
|
||||
## Key Rotation Log
|
||||
|
||||
|
||||
@@ -26,7 +26,7 @@ description: >
|
||||
Scraped by Prometheus with Bearer master key; endpoint returns 307 → /metrics/.
|
||||
- Alertmanager (harness-alertmanager :9093) + zulip-bridge (:9102) deliver
|
||||
firing alerts to #agent-hub > alerts-infra via abiba-bot. Added 2026-08-09.
|
||||
- Prometheus node job covers ALL 6 PVE nodes (.4/.5/.6/.9/.12/.15:9100).
|
||||
- Prometheus node job covers ALL 5 PVE nodes (.5/.6/.9/.12/.15:9100).
|
||||
---
|
||||
|
||||
## Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU)
|
||||
|
||||
@@ -1,4 +1,6 @@
|
||||
---
|
||||
report_only_agents:
|
||||
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
|
||||
kind: responsibility
|
||||
name: litellm-self-heal
|
||||
status: deployed
|
||||
@@ -22,6 +24,7 @@ description: >
|
||||
inference, and agent keys. Applies remediation rules for common failures.
|
||||
Reports every action via Zulip DM and Gitea (SyslogSolution/health-logs).
|
||||
---
|
||||
---
|
||||
|
||||
# LiteLLM Operations — Health Check + Self-Heal
|
||||
|
||||
@@ -65,7 +68,15 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
||||
|
||||
## LiteLLM Model Surface (ground truth — `/opt/inference-harness/litellm_config.yaml` on CT 116)
|
||||
|
||||
`model_name`s served: `qwen3.6-27B-code`, `gemma-4-12b`, `qwen3.6-35B-udq4`, `strix-moe`, `gpu-dense`, `gpu-light`, `syslog-auto`.
|
||||
`model_name`s served: `qwen3.6-27B-code`, `gemma-4-12b`, `qwen3.6-35B-udq4`, `strix-moe`, `gpu-dense`, `gpu-light`, `syslog-auto`, `crew-auto` (new 2026-08-20).
|
||||
|
||||
### Context Cap Split (2026-08-20)
|
||||
|
||||
- **Abiba (firstmate)**: 128K uncapped — unlimited context for primary workloads
|
||||
- **Hermes agents** (mumuni, tanko, koby, koonimo): 128K uncapped
|
||||
- **Crewmates** (ops, tune, verify, auth-keys, build): 64K capped — alias `crew-auto` enforces 64K limit
|
||||
|
||||
Preferred implementation: uncap shared pool, add capped alias for crew-only.
|
||||
|
||||
- `syslog-auto` is a weighted router model: qwen3.6-27B-code (0.55, rpm 500) + qwen3.6-35B-udq4 (0.30, rpm 60) + gemma-4-12b (0.15, rpm 200).
|
||||
- `gpu-dense` / `gpu-light` are high-rpm aliases (rpm 500) onto qwen3.6-27B-code / gemma-4-12b respectively.
|
||||
@@ -120,6 +131,7 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
||||
- Also wakes on user request
|
||||
- On failure: re-check after 30s, escalate after 3 consecutive failures
|
||||
|
||||
---
|
||||
---
|
||||
|
||||
## Health Check
|
||||
@@ -162,6 +174,7 @@ Determine overall_status from individual check results:
|
||||
- "degraded" — 1-2 non-critical checks fail
|
||||
- "down" — critical checks fail
|
||||
|
||||
---
|
||||
---
|
||||
|
||||
## Remediation Rules
|
||||
@@ -205,6 +218,7 @@ Escalate → if SSH access unavailable, send Zulip DM
|
||||
Router no longer in path so Redis active counters are unused. Rule retained
|
||||
for reference but inactive. If Redis issues occur, check harness-redis container.
|
||||
|
||||
---
|
||||
---
|
||||
|
||||
## Reporting
|
||||
@@ -227,6 +241,7 @@ top actions, uptime.
|
||||
If a fix requires another agent (e.g., Authentik restart), relay sent
|
||||
to responsible agent with full context.
|
||||
|
||||
---
|
||||
---
|
||||
|
||||
## Execution
|
||||
|
||||
@@ -1,9 +1,12 @@
|
||||
---
|
||||
report_only_agents:
|
||||
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
|
||||
name: memory-audit-maintenance
|
||||
kind: responsibility
|
||||
description: Shared memory audit and maintenance contract for all Hermes agents (Mumuni, Tanko, Koby, Koonimo). Each agent runs it against its own isolated memory files — no cross-agent access, no shared state. Detects staleness, enforces writer registry, and rotates canary tokens.
|
||||
description: Shared memory audit and maintenance contract for Hermes agents (Mumuni, Koby, Koonimo). Tanko is no longer a Hermes agent (now on DSH/DeepSeek Harness since 2026-08-27) and uses DSH-native memory, so it is excluded from this Hermes roster. Each agent runs it against its own isolated memory files — no cross-agent access, no shared state. Detects staleness, enforces writer registry, and rotates canary tokens.
|
||||
id: 067NC4KG01RG50R40M30E20918
|
||||
---
|
||||
---
|
||||
|
||||
### Goal
|
||||
|
||||
@@ -15,11 +18,10 @@ Autonomously audit and reorganize an agent's native memory (MEMORY.md, USER.md,
|
||||
|
||||
### Scope
|
||||
|
||||
This contract is the **Hermes Agent standard** for memory maintenance. It is shared across all Hermes agents (Mumuni, Tanko, Koby, Koonimo). Each agent runs it against its own memory files only no cross-agent access, no shared state, no shared ledger, no shared canary. The contract is the standard; each agent enforces it independently with fully isolated data.
|
||||
This contract is the **Hermes Agent standard** for memory maintenance. It is shared across Hermes agents (Mumuni, Koby, Koonimo). **Tanko is excluded — it migrated to DSH (DeepSeek Harness) on 2026-08-27 and now uses DSH-native memory (mnemon), not `~/.hermes/memories/`.** Each agent runs it against its own memory files only no cross-agent access, no shared state, no shared ledger, no shared canary. The contract is the standard; each agent enforces it independently with fully isolated data.
|
||||
|
||||
**Agent Roster:**
|
||||
**Agent Roster (Hermes):**
|
||||
- Mumuni
|
||||
- Tanko
|
||||
- Koby (CT 111 / tdunna)
|
||||
- Koonimo (CT 113 / baggy)
|
||||
|
||||
@@ -343,4 +345,3 @@ return {
|
||||
|
||||
### Per-Agent Notes
|
||||
|
||||
Each Hermes agent (Mumuni, Tanko, Tdunna, Baggy) runs this contract against its own `~/.hermes/memories/` directory. The contract is identical across agents, but all data is fully isolated: separate ledgers, separate writer registries, separate canaries. If a new agent is added to the roster, it must be listed in `### Scope` above and given its own isolated memory directory.
|
||||
@@ -1,4 +1,6 @@
|
||||
---
|
||||
report_only_agents:
|
||||
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
|
||||
kind: pattern
|
||||
name: memory-fixer
|
||||
description: >
|
||||
@@ -6,6 +8,7 @@ description: >
|
||||
Escalate anything that needs Kwame's input. Executes confirmed Kwame decisions to completion (state + updated_at).
|
||||
version: 2.0.0
|
||||
---
|
||||
---
|
||||
|
||||
# Memory Fixer
|
||||
|
||||
|
||||
@@ -6,7 +6,7 @@ description: >
|
||||
delegation, verification, and delivery. Defines when to delegate, which
|
||||
worker to use for what, how to handle failures, and the kanban board
|
||||
protocol. Enforces context-window discipline and separation of concerns.
|
||||
Runs on Mumuni (inside Abiba CT100, hwepve, .24) via Hermes agent (Pi + Hermes Zulip gateway).
|
||||
Runs on Mumuni (kagentz CT105, minipve, .14) via Hermes agent (Zulip gateway via systemd).
|
||||
version: 1.0.0
|
||||
---
|
||||
|
||||
@@ -19,8 +19,8 @@ version: 1.0.0
|
||||
|
||||
## Topology
|
||||
|
||||
**Cluster:** 6 Proxmox nodes (ocupve, acerpve, minipve, amdpve, storepve, hwepve)
|
||||
**Manager:** Mumuni (inside Abiba CT100, hwepve, .24) via Hermes agent
|
||||
**Cluster:** 5 Proxmox nodes (ocupve, acerpve, minipve, amdpve, storepve)
|
||||
**Manager:** Mumuni (kagentz CT105, minipve, .14) via Hermes agent
|
||||
**Workers:** 6 profiles, all running on the same agent — no separate hosts needed
|
||||
|
||||
This contract is infrastructure-agnostic in terms of which nodes are used.
|
||||
@@ -31,7 +31,7 @@ Workers execute tasks on whatever infrastructure they're given — SSH to .6,
|
||||
## Why This Matters
|
||||
|
||||
Without enforced delegation, the manager consumes the full iteration budget
|
||||
(60 calls) on single-turn tasks — SSH to 6 nodes, check each VM, read logs —
|
||||
(60 calls) on single-turn tasks — SSH to 5 nodes, check each VM, read logs —
|
||||
leaving no capacity for actual coordination. The result: context overflow
|
||||
(59K tokens in system prompt), iteration exhaustion, and degraded response
|
||||
quality. This contract exists because I blew through my budget checking
|
||||
@@ -82,7 +82,7 @@ it asks the manager (via relay) — it doesn't go find it on its own.
|
||||
|
||||
**This is a hard rule, not a recommendation.** Violating it produces the exact
|
||||
type of discrepancy the kanban pipeline exists to prevent: a review worker finds
|
||||
"6 nodes present" in the raw data but "6/6 online" in the report — even though
|
||||
"5 nodes present" in the raw data but "5/5 online" in the report — even though
|
||||
one of those nodes was unreachable. The report lied because it used data the
|
||||
raw data never provided.
|
||||
|
||||
@@ -137,7 +137,7 @@ delegate_task(
|
||||
```
|
||||
delegate_task(
|
||||
tasks=[
|
||||
{"goal": "Check all 6 Proxmox nodes for VM status", "context": "SSH to each node via 192.168.68.x, run 'qm list'"},
|
||||
{"goal": "Check all 5 Proxmox nodes for VM status", "context": "SSH to each node via 192.168.68.x, run 'qm list'"},
|
||||
{"goal": "Check Docker container health on .7/.116/.17", "context": "SSH to each host, check container status"},
|
||||
]
|
||||
)
|
||||
@@ -191,7 +191,7 @@ Only verified results reach Kwame. Format per channel:
|
||||
{
|
||||
"lane_id": "devops-check",
|
||||
"worker": "syslog-devops",
|
||||
"goal": "Check all 6 Proxmox nodes",
|
||||
"goal": "Check all 5 Proxmox nodes",
|
||||
"status": "dispatched|completed|failed",
|
||||
"output_file": "/tmp/node-report.md"
|
||||
}
|
||||
|
||||
@@ -54,7 +54,7 @@ garbled input. When a user types these commands to Abiba:
|
||||
|
||||
- **`/approve`** → "Pi doesn't have pending approvals. Commands execute immediately."
|
||||
- **`/approve session`** → Same response
|
||||
- **`/deny`** → Same response + "For Hermes agents (Tanko, Mumuni), these work with their built-in approval system."
|
||||
- **`/deny`** → Same response + "For agents (Mumuni on Hermes, Tanko on DSH), these work with their built-in approval system."
|
||||
|
||||
This keeps the UX consistent across agents — users can type `/approve` anywhere
|
||||
without getting confused by LLM responses.
|
||||
|
||||
+9
-17
@@ -2,16 +2,6 @@
|
||||
kind: responsibility
|
||||
name: pm2-self-heal
|
||||
description: >
|
||||
Monitors critical PM2 processes (abiba-zulip, abiba-telegram, gitea-runner,
|
||||
spoton-service, zulip-watchdog) and auto-restarts any that are stopped or
|
||||
errored. Logs every action to the knowledge graph and alerts the owner via
|
||||
Zulip DM on failures.
|
||||
CRITICAL: Never restart abiba-zulip — it runs this contract.
|
||||
AS-BUILT 2026-08-09 (captain ruling, ecosystem is authoritative):
|
||||
gpu-monitor is systemd-managed (gpu-monitor.service) — NOT PM2;
|
||||
gpu-watchdog decommissioned (function folded into gpu-monitor.service);
|
||||
gitea-runner KEPT (online in PM2); abiba-zulip KEPT (online 4d+, the
|
||||
2026-07-04 'removed/decommissioned' note was stale and is removed).
|
||||
---
|
||||
|
||||
## Maintains
|
||||
@@ -23,9 +13,6 @@ description: >
|
||||
- zulip-watchdog: { status: "online", uptime: string, restarts: number }
|
||||
- last_check: timestamp
|
||||
|
||||
> **Note (2026-07-04, SUPERSEDED 2026-08-09):** `abiba-zulip` remains ONLINE and
|
||||
> is monitored — the decommission note was stale (process re-added; do not treat
|
||||
> it as removed).
|
||||
|
||||
## Continuity
|
||||
|
||||
@@ -40,10 +27,15 @@ description: >
|
||||
- **Verify**: Re-check status after 5 seconds
|
||||
- **Escalate**: If still failed after 2 retries, send Zulip DM to owner
|
||||
|
||||
### Rule 2: Process Restarting Too Often
|
||||
- **Detect**: `pm2 status` shows restarts > 30 (cumulative lifetime counter)
|
||||
### Rule 2: Process Restarting Too Often (crash-loop guard)
|
||||
- **Detect**: `pm2 status` shows restarts > 30 (cumulative lifetime counter) **or** a
|
||||
process reporting a restart count > 1000 while showing "online" (a crash-loop mask)
|
||||
- **Note**: PM2 counter never decrements; only full delete+re-add resets it
|
||||
- **Fix**: `pm2 delete <name> && pm2 start <ecosystem> --only <name>`
|
||||
- **Script guard (2026-08-16)**: `scripts/pm2-self-heal.sh` now restarts `abiba-telegram`
|
||||
when `TEL_RESTARTS > 1000` even if the process reports "online" — catches a quiet
|
||||
crash-loop that never toggles status to "stopped"/"errored" (e.g. the 10k-restarts
|
||||
spoton incident). Alerts include the restart count.
|
||||
- **Escalate**: Only when restarts > 30 — alerts to Zulip DM
|
||||
- **Historical fix**: Previous cycles were caused by abiba-zulip extension's
|
||||
stuck detection (STUCK_THRESHOLD_MS was 30min, raised to 4h in v2).
|
||||
@@ -57,9 +49,9 @@ description: >
|
||||
- If status is "online" → pass
|
||||
- If status is "stopped" or "errored" → apply Rule 1
|
||||
- If restarts > 5 → alert owner
|
||||
3. **Check abiba-zulip** (self-process, read-only):
|
||||
3. **Check abiba-zulip** (live Zulip bridge, heartbeating):
|
||||
- If status is "online" → pass, log restarts count
|
||||
- If status is "stopped" or "errored" → **DO NOT RESTART** — alert owner immediately
|
||||
- If status is "stopped" or "errored" → restart (`pm2 restart abiba-zulip` — fully restored)
|
||||
- If restarts > 5 in last hour → alert owner with full diagnostics
|
||||
4. **Log results** — Append to `SyslogSolution/health-logs/pm2/{timestamp}.md` in Gitea (not knowledge graph — hard rule)
|
||||
5. **Alert** — Send Zulip DM to owner if escalation needed (do NOT run pm2 commands during alerting)
|
||||
|
||||
@@ -5,7 +5,7 @@ description: >
|
||||
Proxmox cluster + Docker monitoring via the existing Grafana/Prometheus stack
|
||||
on CT 116. Replaces Pulse with file-provisioned Grafana dashboards. Three
|
||||
exporters feed Prometheus: prometheus-pve-exporter (cluster-aware, single
|
||||
instance), node_exporter (all 6 PVE nodes), and a custom docker-stats-exporter
|
||||
instance), node_exporter (all 5 PVE nodes), and a custom docker-stats-exporter
|
||||
(Docker 29 / containerd image-store compatible, since cAdvisor cannot resolve
|
||||
the layerdb). Dashboards exposed at http://192.168.68.116:3001/ (direct LAN, not behind nginx).
|
||||
agent: abiba
|
||||
@@ -33,8 +33,8 @@ agent: abiba
|
||||
|
||||
| Exporter | Host:Port | Scope | Notes |
|
||||
|----------|-----------|-------|-------|
|
||||
| prometheus-pve-exporter | .116:9221 (container) | All 6 nodes + guests + 36 storage pools | Single instance, cluster-aware via amdpve API. Config `/opt/monitoring/pve.yml` (token `monitoring@pve!prometheus`, PVEAuditor role). Metric schema is label-based (`id=node/amdpve`, `id=lxc/100`). |
|
||||
| node_exporter | .5/.6/.9/.12/.15/.4:9100 (systemd) | Per-node CPU/mem/disk/net/temp | Installed via apt on all 6 PVE nodes, enabled (reboot-persistent). Collectors: textfile, systemd, tcpstat, ethtool. hwepve (.4) added 2026-07-19. |
|
||||
| prometheus-pve-exporter | .116:9221 (container) | All 5 nodes + 14 guests + 36 storage pools | Single instance, cluster-aware via amdpve API. Config `/opt/monitoring/pve.yml` (token `monitoring@pve!prometheus`, PVEAuditor role). Metric schema is label-based (`id=node/amdpve`, `id=lxc/100`). |
|
||||
| node_exporter | .5/.6/.9/.12/.15:9100 (systemd) | Per-node CPU/mem/disk/net/temp | Installed via apt on all 5 PVE nodes, enabled (reboot-persistent). Collectors: textfile, systemd, tcpstat, ethtool. |
|
||||
| docker-stats-exporter | .116:9324 (container) | 10 Docker containers on .116 | **Custom** (cAdvisor v0.51 incompatible with Docker 29 containerd image store — layerdb gone). Uses Docker Engine API over unix socket. Script `/opt/monitoring/docker-stats-exporter.py`. |
|
||||
|
||||
## PVE API Token
|
||||
@@ -48,7 +48,7 @@ agent: abiba
|
||||
|
||||
| UID | Title | Panels | Source |
|
||||
|-----|-------|--------|--------|
|
||||
| proxmox-cluster | Proxmox Cluster Overview | 16 | cluster status, 6-node CPU/mem/disk/load gauges, guests table, storage pools, guest CPU/mem timeseries |
|
||||
| proxmox-cluster | Proxmox Cluster Overview | 16 | cluster status, 5-node CPU/mem/disk/load gauges, guests table, storage pools, guest CPU/mem timeseries |
|
||||
| proxmox-node | Proxmox Node Detail | 13 | per-node CPU per-core, memory, network, disk IO/IOPS/latency, temperature, disk space (variable: $node) |
|
||||
| docker-containers | Docker Containers | 10 | per-container CPU/mem/network, restarts, memory limit ratio (variable: $container) |
|
||||
| gpu-fleet | GPU Fleet | 7 | (existing, preserved in DB, not provisioned) |
|
||||
@@ -77,9 +77,9 @@ agent: abiba
|
||||
| `/opt/monitoring/grafana/dashboards/build-dashboards.py` | .116 | dashboard JSON generator |
|
||||
| `/opt/monitoring/grafana/dashboards/json/*.json` | .116 | provisioned dashboard definitions |
|
||||
| `/opt/monitoring/grafana/datasources/prometheus.yml` | .116 | datasource provisioning |
|
||||
| `/etc/default/prometheus-node-exporter` | .5/.6/.9/.12/.15/.4 | node_exporter collector config |
|
||||
| `/etc/default/prometheus-node-exporter` | .5/.6/.9/.12/.15 | node_exporter collector config |
|
||||
|
||||
## Cluster "Tabiri" — 6 Nodes
|
||||
## Cluster "Tabiri" — 5 Nodes
|
||||
|
||||
| Node | IP | Role |
|
||||
|------|----|----|
|
||||
@@ -88,7 +88,6 @@ agent: abiba
|
||||
| acerpve | 192.168.68.9 | PVE (hosts llm-gpu qemu/101) |
|
||||
| minipve | 192.168.68.12 | PVE |
|
||||
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (qwen3.6-35B-udq4, strix-moe) |
|
||||
| hwepve | 192.168.68.4 | PVE (Huawei Matebook 16, 12C/15GB) — hosts abiba (lxc/100), kagentz (lxc/105), mumuni (lxc/114). CTs 100/105 migrated from amdpve, CT 114 from minipve 2026-07-20 |
|
||||
|
||||
## Operations
|
||||
|
||||
|
||||
@@ -30,7 +30,6 @@ INFISICAL_ENV = "prod"
|
||||
|
||||
# PVE node IPs for CT liveness checks
|
||||
PVE_NODES = {
|
||||
"hwepve": "192.168.68.4",
|
||||
"amdpve": "192.168.68.15",
|
||||
"minipve": "192.168.68.12",
|
||||
"storepve": "192.168.68.6",
|
||||
@@ -40,8 +39,8 @@ PVE_NODES = {
|
||||
|
||||
# Agent definitions: ct, host, user, pve_node, vault_key_name
|
||||
AGENTS = {
|
||||
"tanko": {"ct": 112, "host": "192.168.68.122", "user": "jerome", "pve": "amdpve", "vault_key": "TANKO_LITELLM_API_KEY"},
|
||||
"abiba": {"ct": 100, "host": "192.168.68.24", "user": "root", "pve": "hwepve", "vault_key": None}, # Pi agent + Mumuni Zulip, no vault key
|
||||
"tanko": {"ct": 112, "host": "192.168.68.122", "user": "jerome", "pve": "amdpve", "vault_key": "TANKO_LITELLM_API_KEY", "runtime": "dsh"},
|
||||
"abiba": {"ct": 100, "host": "192.168.68.24", "user": "root", "pve": "minipve", "vault_key": None}, # Pi agent + Mumuni Zulip, no vault key
|
||||
"koby": {"ct": 111, "host": "192.168.68.129", "user": "root", "pve": "amdpve", "vault_key": "KOBY_LITELLM_API_KEY"},
|
||||
"koonimo": {"ct": 113, "host": "192.168.68.114", "user": "root", "pve": "amdpve", "vault_key": "KOONIMO_LITELLM_API_KEY"},
|
||||
}
|
||||
@@ -239,20 +238,36 @@ def check_agents():
|
||||
host = agent.get("host")
|
||||
user = agent.get("user")
|
||||
ct = agent["ct"]
|
||||
report_only = agent.get("report_only", False)
|
||||
|
||||
# Tanko runs on DSH (DeepSeek Harness) since 2026-08-27 — it no longer runs a
|
||||
# Hermes gateway, so skip the Hermes gateway/state/streaming/journal checks.
|
||||
if agent.get("runtime") == "dsh":
|
||||
live = ssh(host, "true", user=user)
|
||||
print(f" {'✅' if live is not None else '❌'} {name}: DSH (DeepSeek Harness) — "
|
||||
f"no Hermes gateway since 2026-08-27 (CT {ct}, SSH {'OK' if live is not None else 'FAIL'})")
|
||||
if live is None:
|
||||
FAIL.append(f"unreachable:{name}")
|
||||
continue
|
||||
|
||||
if not host or not user:
|
||||
print(f" ⬜ {name} (CT {ct}): cannot SSH — skip liveness check")
|
||||
continue
|
||||
|
||||
# Gateway process
|
||||
pid = ssh(host, "pgrep -f 'hermes_cli.main gateway run' | grep -v infisical | head -1", user=user)
|
||||
if not pid:
|
||||
# Try alternate binary name
|
||||
pid = ssh(host, "pgrep -f 'hermes.*gateway' | grep -v infisical | grep -v bash | head -1", user=user)
|
||||
if not pid:
|
||||
print(f" ❌ {name}: GATEWAY NOT RUNNING")
|
||||
FAIL.append(f"gateway-down:{name}")
|
||||
continue
|
||||
# ⛔ KOBY IS NEVER REPAIRED — diagnostic only
|
||||
if report_only:
|
||||
print(f" 🔍 {name}: REPORT-ONLY mode (diagnostic only, no repairs on .129)")
|
||||
# Still check gateway status for reporting purposes
|
||||
pid = ssh(host, "pgrep -f 'hermes_cli.main gateway run' | grep -v infisical | head -1", user=user)
|
||||
if not pid:
|
||||
pid = ssh(host, "pgrep -f 'hermes.*gateway' | grep -v infisical | grep -v bash | head -1", user=user)
|
||||
if not pid:
|
||||
print(f" ⚠️ {name}: GATEWAY NOT RUNNING (reported only)")
|
||||
FAIL.append(f"gateway-down:{name}")
|
||||
continue
|
||||
else:
|
||||
print(f" ✅ {name}: gateway running (pid={pid}, report-only mode)")
|
||||
continue # Skip the rest of the check for Koby
|
||||
|
||||
# Gateway state file
|
||||
state = ssh(host, "cat ~/.hermes/gateway_state.json 2>/dev/null", user=user)
|
||||
@@ -328,6 +343,10 @@ def check_ct_liveness():
|
||||
def check_config_integrity():
|
||||
"""Verify agent config.yaml parses as valid YAML."""
|
||||
for name, agent in AGENTS.items():
|
||||
# Tanko runs on DSH (DeepSeek Harness) since 2026-08-27 — no Hermes config.yaml.
|
||||
if agent.get("runtime") == "dsh":
|
||||
print(f" ⏭️ {name}: DSH — no Hermes config.yaml since 2026-08-27")
|
||||
continue
|
||||
host = agent.get("host")
|
||||
user = agent.get("user")
|
||||
if not host or not user:
|
||||
@@ -357,6 +376,10 @@ def check_config_integrity():
|
||||
def check_wrapper_integrity():
|
||||
"""Verify the hermes CLI wrapper exists and can reach hermes-real."""
|
||||
for name, agent in AGENTS.items():
|
||||
# Tanko runs on DSH (DeepSeek Harness) since 2026-08-27 — no hermes CLI wrapper.
|
||||
if agent.get("runtime") == "dsh":
|
||||
print(f" ⏭️ {name}: DSH — no hermes CLI wrapper since 2026-08-27")
|
||||
continue
|
||||
host = agent.get("host")
|
||||
user = agent.get("user")
|
||||
if not host or not user:
|
||||
|
||||
@@ -296,21 +296,16 @@ def collect():
|
||||
"pm2_uptime": pm2.get("uptime", "?"),
|
||||
}
|
||||
|
||||
# Tanko (CT 122)
|
||||
tanko_state = ssh_jerome("192.168.68.122", "cat ~/.hermes/gateway_state.json 2>/dev/null")
|
||||
tanko_data = {}
|
||||
try:
|
||||
tanko_data = json.loads(tanko_state) if tanko_state else {}
|
||||
except:
|
||||
tanko_data = {}
|
||||
platforms = tanko_data.get("platforms", {})
|
||||
# Tanko (CT 112, IP 192.168.68.122) — DSH (DeepSeek Harness), no Hermes gateway
|
||||
# since 2026-08-27. There is no ~/.hermes/gateway_state.json on CT 112 anymore;
|
||||
# Zulip/Telegram connectivity is managed by the DSH harness, not the Hermes gateway.
|
||||
report["agents"]["tanko"] = {
|
||||
"platform": "hermes", "ct": 112, "ip": "192.168.68.122",
|
||||
"gateway_state": tanko_data.get("gateway_state", "unknown"),
|
||||
"zulip_state": platforms.get("zulip", {}).get("state", "unknown"),
|
||||
"telegram_state": platforms.get("telegram", {}).get("state", "unknown"),
|
||||
"gateway_pid": tanko_data.get("pid"),
|
||||
"updated_at": tanko_data.get("updated_at"),
|
||||
"platform": "dsh", "ct": 112, "ip": "192.168.68.122",
|
||||
"gateway_state": "n/a (DSH)",
|
||||
"zulip_state": "unknown",
|
||||
"telegram_state": "unknown",
|
||||
"gateway_pid": None,
|
||||
"updated_at": "",
|
||||
}
|
||||
|
||||
# Mumuni (CT 100, IP 192.168.68.24)
|
||||
@@ -322,7 +317,7 @@ def collect():
|
||||
mumuni_data = {}
|
||||
mumuni_platforms = mumuni_data.get("platforms", {})
|
||||
report["agents"]["mumuni"] = {
|
||||
"platform": "hermes", "ct": 114, "ip": "192.168.68.24",
|
||||
"platform": "hermes", "ct": 100, "ip": "192.168.68.24",
|
||||
"gateway_state": mumuni_data.get("gateway_state", "unknown"),
|
||||
"telegram_state": mumuni_platforms.get("telegram", {}).get("state", "unknown"),
|
||||
"zulip_state": mumuni_platforms.get("zulip", {}).get("state", "not_installed"),
|
||||
@@ -543,7 +538,7 @@ th {{ color: #8b949e; font-weight: normal; }}
|
||||
elif name == "tanko":
|
||||
zulip_state = "✅" if agent.get("zulip_state") == "connected" else ("❌" if agent.get("zulip_state") == "disconnected" else "⬜")
|
||||
gateway = agent.get("gateway_state", "?")
|
||||
processed = agent.get("updated_at", "")[:10]
|
||||
processed = "DSH"
|
||||
elif name == "mumuni":
|
||||
zulip_state = "⬜" if agent.get("zulip_state") == "not_installed" else ("✅" if agent.get("zulip_state") == "connected" else "⬜")
|
||||
gateway = agent.get("gateway_state", "?")
|
||||
|
||||
+3
-6
@@ -28,10 +28,8 @@ declare -A CT_NODES=(
|
||||
[117]=storepve # zulip
|
||||
[118]=storepve # jdownloader
|
||||
# acerpve (192.168.68.9) — no CTs (bare metal GPU .8)
|
||||
# hwepve (192.168.68.4)
|
||||
[100]=hwepve # abiba (was amdpve)
|
||||
[105]=hwepve # kagentz (was amdpve)
|
||||
[114]=hwepve # mumuni (was minipve)
|
||||
[100]=minipve # abiba (was hwepve)
|
||||
[105]=minipve # kagentz (was hwepve)
|
||||
# ocupve (192.168.68.5) — no CTs (bare metal GPU .110)
|
||||
#
|
||||
# REMOVED CTs (migrated to bare metal, decommissioned, or VMs):
|
||||
@@ -50,7 +48,6 @@ declare -A NODE_IPS=(
|
||||
[storepve]=192.168.68.6
|
||||
[acerpve]=192.168.68.9
|
||||
[ocupve]=192.168.68.5
|
||||
[hwepve]=192.168.68.4
|
||||
)
|
||||
|
||||
resolve_node() {
|
||||
@@ -77,7 +74,7 @@ main() {
|
||||
if [[ $# -lt 1 ]]; then
|
||||
echo "Usage: pct-run <CT_ID> [command...]" >&2
|
||||
echo " pct-run 112 cat /etc/hostname" >&2
|
||||
echo " pct-run 114 systemctl status hermes-gateway" >&2
|
||||
echo " pct-run 100 systemctl status hermes-gateway" >&2
|
||||
echo ""
|
||||
echo "Known CTs:" >&2
|
||||
for ct in $(echo "${!CT_NODES[@]}" | tr ' ' '\n' | sort -n); do
|
||||
|
||||
@@ -33,7 +33,7 @@ TEL_LINE=$(echo "$STATUS" | grep "abiba-telegram")
|
||||
TEL_STATUS=$(echo "$TEL_LINE" | awk -F'│' '{print $10}' | xargs)
|
||||
TEL_RESTARTS=$(echo "$TEL_LINE" | awk -F'│' '{print $9}' | xargs)
|
||||
|
||||
if [ "$TEL_STATUS" != "online" ]; then
|
||||
if [ "$TEL_STATUS" != "online" ] || [ "$TEL_RESTARTS" -gt 1000 ]; then
|
||||
pm2 restart abiba-telegram > /dev/null 2>&1
|
||||
sleep 3
|
||||
TEL_LINE2=$(pm2 status --no-color 2>/dev/null | grep "abiba-telegram")
|
||||
|
||||
@@ -46,18 +46,17 @@ You are a code reviewer for OpenProse infrastructure contracts in the Syslog Sol
|
||||
|
||||
The infrastructure-control.prose.md contract is the canonical reference for the cluster topology:
|
||||
|
||||
**Proxmox Cluster "Tabiri" (6 nodes):**
|
||||
**Proxmox Cluster "Tabiri" (5 nodes):**
|
||||
- amdpve (192.168.68.15): tanko, tdunna, baggy, scottdenya
|
||||
- minipve (192.168.68.12): adguard, authentik, gitea, syslog-api, infisical-vault
|
||||
- minipve (192.168.68.12): abiba, kagentz, adguard, authentik, gitea, syslog-api, infisical-vault
|
||||
- storepve (192.168.68.6): docker-vm, ra-h-os, PBS, media, jdownloader, zulip
|
||||
- acerpve (192.168.68.9): llm-gpu
|
||||
- ocupve (192.168.68.5): ocu-llm
|
||||
- hwepve (192.168.68.4): abiba, kagentz, mumuni
|
||||
|
||||
**CT IDs (verified 2026-07-24 against PVE API):**
|
||||
100:abiba 102:adguard 104:authentik 105:kagentz 106:ra-h-os
|
||||
107:pbs 108:media 110:gitea 111:tdunna 112:tanko
|
||||
113:baggy 114:mumuni 115:scottdenya 116:syslog-api 117:zulip
|
||||
113:baggy 115:scottdenya 116:syslog-api 117:zulip
|
||||
118:jdownloader 119:infisical-vault
|
||||
|
||||
**NO CT 122, CT 123, or .19 exist in the cluster.**
|
||||
@@ -65,7 +64,7 @@ The infrastructure-control.prose.md contract is the canonical reference for the
|
||||
**CRITICAL RULES (never regress):**
|
||||
1. NO /grafana/ nginx route — it was tried and reverted on 2026-07-02. Grafana is direct LAN at :3001.
|
||||
2. NO .19 IP — Zulip is CT 117 on storepve.
|
||||
3. NO CT 122/123 — Tanko=CT 112, Mumuni=CT 114.
|
||||
3. NO CT 122/123 — Tanko=CT 112, Mumuni=CT 100 (inside Abiba). No CT 114 anywhere.
|
||||
4. Strix Halo :8080 is FIREWALLED to .116 only — cannot be probed from abiba (.24).
|
||||
5. abiba-zulip PM2 process is DECOMMISSIONED (2026-07-04) — abiba uses Telegram only.
|
||||
|
||||
|
||||
@@ -12,10 +12,10 @@ echo ""
|
||||
# Authorized agents for restricted contracts
|
||||
# Format: contract_pattern|authorized_agents (comma-separated)
|
||||
declare -A RESTRICTED
|
||||
RESTRICTED["infrastructure-control.prose.md"]="abiba"
|
||||
RESTRICTED["proxmox-monitor.prose.md"]="abiba"
|
||||
RESTRICTED["hermes-config-template.prose.md"]="abiba,mumuni,tanko"
|
||||
RESTRICTED["zulip-health.prose.md"]="abiba,mumuni"
|
||||
RESTRICTED["infrastructure-control.prose.md"]="abiba,abiba-bot,tanko,tanko-bot,mumuni,mumuni-bot"
|
||||
RESTRICTED["proxmox-monitor.prose.md"]="abiba,abiba-bot"
|
||||
RESTRICTED["hermes-config-template.prose.md"]="abiba,abiba-bot,mumuni,mumuni-bot,tanko,tanko-bot"
|
||||
RESTRICTED["zulip-health.prose.md"]="abiba,abiba-bot,mumuni,mumuni-bot,tanko,tanko-bot"
|
||||
RESTRICTED["scripts/pm2-self-heal.sh"]="abiba"
|
||||
RESTRICTED["scripts/prose-lint.sh"]="abiba"
|
||||
RESTRICTED["scripts/prose-ai-review.sh"]="abiba"
|
||||
|
||||
+11
-26
@@ -9,10 +9,11 @@ ZULIP_EMAIL="abiba-bot@chat.sysloggh.net"
|
||||
ZULIP_KEY="cKTDMZAPW08dk3zl05sStzO7HRztzyn8"
|
||||
OWNER_ZULIP_ID="9"
|
||||
|
||||
# Email config
|
||||
GMAIL_USER="jtabiri@gmail.com"
|
||||
GMAIL_PASS="rgbuomwcydxwbszd"
|
||||
EMAIL_TO="jerome@sysloggh.com"
|
||||
|
||||
LOG="/root/zulip-health-monitor.log"
|
||||
TIMESTAMP=$(date -u '+%Y-%m-%d %H:%M UTC')
|
||||
ISSUES=0
|
||||
echo "=== Zulip Health Check — $TIMESTAMP ===" >> "$LOG"
|
||||
|
||||
notify() {
|
||||
local severity="$1" msg="$2"
|
||||
@@ -24,30 +25,14 @@ notify() {
|
||||
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
|
||||
-u "${ZULIP_EMAIL}:${ZULIP_KEY}" \
|
||||
-d "${form}" > /dev/null 2>&1 || true
|
||||
|
||||
# Email alert
|
||||
local subject="${severity} Zulip Monitor Alert"
|
||||
python3 -c "
|
||||
import smtplib
|
||||
from email.mime.text import MIMEText
|
||||
m = MIMEText('''${msg}''')
|
||||
m['From'] = 'abiba@sysloggh.com'
|
||||
m['To'] = '${EMAIL_TO}'
|
||||
m['Subject'] = '${subject}'
|
||||
s = smtplib.SMTP('smtp.gmail.com', 587)
|
||||
s.starttls()
|
||||
s.login('${GMAIL_USER}', '${GMAIL_PASS}')
|
||||
s.sendmail('abiba@sysloggh.com', ['${EMAIL_TO}'], m.as_string())
|
||||
s.quit()
|
||||
" 2>/dev/null || true
|
||||
# Zulip stream post to #agent-hub on topic 'zulip-health'
|
||||
local stream_content="${severity} Zulip Monitor: ${msg}"
|
||||
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
|
||||
-u "${ZULIP_EMAIL}:${ZULIP_KEY}" \
|
||||
-d "type=stream\&to=%5B7%5D\&topic=zulip-health\&content=$(python3 -c "import urllib.parse; print(urllib.parse.quote(str()))")" \
|
||||
> /dev/null 2>&1 || true
|
||||
}
|
||||
|
||||
TIMESTAMP=$(date -u '+%Y-%m-%d %H:%M UTC')
|
||||
ISSUES=0
|
||||
LOG="/root/zulip-health-monitor.log"
|
||||
|
||||
echo "=== Zulip Health Check — $TIMESTAMP ===" >> "$LOG"
|
||||
|
||||
# ── Global: Zulip Server ──
|
||||
SERVER_CODE=$(curl -s -o /dev/null -w "%{http_code}" --connect-timeout 10 \
|
||||
https://chat.sysloggh.net/api/v1/server_settings \
|
||||
|
||||
+15
-16
@@ -1,22 +1,24 @@
|
||||
---
|
||||
kind: responsibility
|
||||
name: zulip-health
|
||||
description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (Agent Zero Docker), Platform B (Hermes agents Tanko/Mumuni), and the Zulip bridge. Verifies bot registration, DM delivery, and cross-platform connectivity.
|
||||
description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (Agent Zero Docker), Platform B (Tanko on DSH / Mumuni on Hermes), and the Zulip bridge. Verifies bot registration, DM delivery, and cross-platform connectivity.
|
||||
title: Zulip Mesh Health Monitor — Multi-Platform
|
||||
version: 3.0.0
|
||||
runtime_contract: 2
|
||||
agent: abiba
|
||||
report_only_agents:
|
||||
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
|
||||
---
|
||||
|
||||
# Zulip Mesh Health Monitor
|
||||
|
||||
Monitors ALL Zulip-connected agents across three platforms (pi, Hermes, Agent Zero).
|
||||
Monitors ALL Zulip-connected agents across platforms (pi, Hermes, DSH, Agent Zero).
|
||||
Runs every 15 minutes in the background. Also triggers on session start.
|
||||
|
||||
## Requires
|
||||
|
||||
- **Zulip API key** for `abiba-bot@chat.sysloggh.net` in `$ZULIP_API_KEY`
|
||||
- **SSH access** to Tanko (192.168.68.122), Mumuni (192.168.68.24, inside Abiba CT100 on hwepve), and Agent Zero Docker host (192.168.68.14)
|
||||
- **SSH access** to Tanko (192.168.68.122), Mumuni (192.168.68.14, kagentz CT105 on minipve), and Agent Zero Docker host (192.168.68.14)
|
||||
- **PM2** on localhost for pi process management
|
||||
- **Network access** to `chat.sysloggh.net`, `localhost:9200`
|
||||
- **Write access** to `/root/zulip-health-monitor.log` and `/tmp/zulip-monitor-debounce`
|
||||
@@ -45,7 +47,7 @@ Runs every 15 minutes in the background. Also triggers on session start.
|
||||
"severity": "healthy"
|
||||
},
|
||||
"tanko": {
|
||||
"platform": "hermes",
|
||||
"platform": "dsh",
|
||||
"zulip_state": "connected",
|
||||
"heartbeat_age_seconds": 45,
|
||||
"gateway_pid": 1234,
|
||||
@@ -81,13 +83,13 @@ Log as "unreachable" — don't treat as critical unless it persists for 3+ conse
|
||||
## Streaming Support (2026-07-05)
|
||||
|
||||
Zulip agents now support progressive message editing during agent generation.
|
||||
When a Hermes agent (Tanko, Mumuni) processes a message, the response is
|
||||
When a Zulip agent (Tanko on DSH, Mumuni on Hermes) processes a message, the response is
|
||||
streamed in real-time via Zulip's `PATCH /api/v1/messages/{id}` API:
|
||||
|
||||
- Adapter implements `edit_message()` using `_api_patch()` helper
|
||||
- Gateway stream consumer progressively edits the Zulip message
|
||||
- User sees real-time agent thinking instead of waiting for full response
|
||||
- Verified: Tanko (CT 112) and Mumuni (CT 114) both have streaming active
|
||||
- Verified: Tanko (CT 112) and Mumuni (kagentz CT 105) both have streaming active
|
||||
|
||||
### Verification
|
||||
```bash
|
||||
@@ -182,15 +184,14 @@ grep -a "Finalized\|Failed to finalize" /root/.pm2/logs/abiba-zulip-out.log | ta
|
||||
| `last_error` set | Log and monitor |
|
||||
| Crash loop >10/h | Alert user |
|
||||
|
||||
### Step 3: Platform B — Hermes (Tanko .122, Mumuni .24)
|
||||
|
||||
**B1: Gateway State**
|
||||
|
||||
```bash
|
||||
ssh root@192.168.68.122 "cat ~/.hermes/gateway_state.json"
|
||||
ssh root@192.168.68.24 "cat ~/.hermes/gateway_state.json" # Mumuni inside Abiba CT100
|
||||
```
|
||||
|
||||
Tanko runs on DSH (DeepSeek Harness) — it no longer runs a Hermes gateway, so there is no `~/.hermes/gateway_state.json` on CT 112 (.122). Verify Tanko's Zulip connectivity via the DSH harness bot status instead.
|
||||
|
||||
Check `platforms.zulip.state`: `connected` ✅ | `disconnected` ❌ | `error` ❌ | missing → not installed.
|
||||
|
||||
**B2: Agent Process**
|
||||
@@ -203,24 +204,22 @@ Gateway PID should exist with uptime > 60s. **Dual-gateway detection**: if more
|
||||
|
||||
| Agent | Restart command | Notes |
|
||||
|-------|-----------------|-------|
|
||||
| Mumuni (.24) | `pm2 restart abiba-zulip` | Hermes gateway runs under PM2 as `abiba-zulip` |
|
||||
| Tanko (.122) | `bash /opt/hermes-zulip-plugin/run.sh` (or the agent's systemd/user unit) | Tanko does NOT use PM2 — never run `pm2 restart mumuni-zulip` for Tanko (process does not exist) |
|
||||
|
||||
Check gateway log for "Gateway running with 2 platform(s)" (not 1) to confirm Zulip reloaded.
|
||||
|
||||
**B3: Heartbeat Verification**
|
||||
**B3: Heartbeat Verification** (Hermes agent Mumuni only — Tanko has no Hermes gateway)
|
||||
|
||||
```bash
|
||||
ssh root@<CT> "grep Heartbeat ~/.hermes/logs/agent.log | tail -3"
|
||||
ssh root@192.168.68.24 "grep Heartbeat ~/.hermes/logs/agent.log | tail -3"
|
||||
```
|
||||
|
||||
Expected: recent heartbeat (within 5 min), `polls=N` incrementing.
|
||||
Silence > 300s → warning. Silence > 600s → critical.
|
||||
|
||||
**B4: Response Delivery**
|
||||
**B4: Response Delivery** (Hermes agent Mumuni only)
|
||||
|
||||
```bash
|
||||
ssh root@<CT> "grep -E 'Finalized|Failed to finalize|Replied to' ~/.hermes/logs/agent.log | tail -10"
|
||||
ssh root@192.168.68.24 "grep -E 'Finalized|Failed to finalize|Replied to' ~/.hermes/logs/agent.log | tail -10"
|
||||
```
|
||||
|
||||
> 50% fail rate → critical.
|
||||
@@ -229,7 +228,7 @@ ssh root@<CT> "grep -E 'Finalized|Failed to finalize|Replied to' ~/.hermes/logs/
|
||||
|
||||
| Condition | Action |
|
||||
|-----------|--------|
|
||||
| `zulip.state != "connected"` | `ssh root@<CT> "pkill -f 'gateway run'; sleep 2; hermes gateway restart"` |
|
||||
| `zulip.state != "connected"` | `ssh root@<CT> "pkill -f 'gateway run'; sleep 2; hermes gateway restart"` (Mumuni) / restart Tanko via DSH service |
|
||||
| No heartbeat in 10min | Same as above |
|
||||
| `Failed to finalize` > 50% | Check PATCH API, Zulip server |
|
||||
| Response empty/short | Check A2A endpoint / LiteLLM model |
|
||||
|
||||
@@ -10,7 +10,7 @@ description: >
|
||||
|
||||
> **⚠️ RETIRED** — The pi Zulip extension (`~/.pi/agent/extensions/zulip/`) and
|
||||
> PM2 process (`abiba-zulip`) have been decommissioned. All mention/reliability
|
||||
> monitoring now happens through Telegram. Hermes agents (Tanko, Mumuni) and
|
||||
> monitoring now happens through Telegram. Agents (Mumuni on Hermes, Tanko on DSH) and
|
||||
> Agent Zero (kagentz) continue to use Zulip.
|
||||
|
||||
## Maintains
|
||||
|
||||
@@ -420,7 +420,7 @@ Backup v2 before starting: `cp index.js index.js.v2-backup-$(date +%Y%m%d-%H%M%S
|
||||
|-------|----------|-------------|--------------|-------------|
|
||||
| **Abiba** | pi (CT 100) | ✅ Connected | API key missing from Infisical injection; poll timeout noise | Added .env fallback; AbortError treated as empty poll (no retry); poll timeout 65s→90s |
|
||||
| **Tanko** | Hermes (CT 112) | ✅ Connected | Gateway disconnected since Jul 11; watchdog restart didn't re-establish Zulip | Full gateway restart (kill wrapper, let infisical-gateway.sh respawn) |
|
||||
| **Mumuni** | Hermes (CT 114) | ✅ Connected | No issues found | None needed |
|
||||
| **Mumuni** | Hermes (kagentz CT 105, migrated 2026-08-29) | ✅ Connected | No issues found | None needed |
|
||||
|
||||
### Key Fixes Applied
|
||||
|
||||
|
||||
@@ -13,7 +13,8 @@ triggers:
|
||||
|
||||
> **⚠️ RETIRED** — This contract was embedded in the pi Zulip extension code
|
||||
> (`performHealthCheck()`). That code has been removed. Zulip self-healing for
|
||||
> Hermes agents (Tanko, Mumuni) continues through their own gateway monitoring.
|
||||
> agents (Mumuni on Hermes, Tanko on DSH) continues through their own platform
|
||||
> monitoring.
|
||||
|
||||
## Maintains
|
||||
|
||||
@@ -65,8 +66,7 @@ triggers:
|
||||
|------|----|------|---------|
|
||||
| Zulip server | 192.168.68.19 | root | Docker: `zulip-zulip-1` |
|
||||
| Abiba (pi) | localhost | root | PM2: `abiba-zulip` |
|
||||
| Mumuni | 192.168.68.24 (CT100 abiba) | root | `hermes gateway restart` |
|
||||
| Tanko | 192.168.68.122 | jerome | `PATH=$PATH:/home/jerome/.hermes/hermes-agent hermes gateway restart` |
|
||||
| Tanko | 192.168.68.122 (CT 112) | jerome | DSH (DeepSeek Harness) — restart via DSH service, not `hermes gateway restart` (no longer a Hermes agent since 2026-08-27) |
|
||||
|
||||
## Debounce
|
||||
|
||||
@@ -75,9 +75,8 @@ Track via `/tmp/zulip-heal-debounce-<agent>` (unix timestamp of last restart).
|
||||
|
||||
## Reporting
|
||||
|
||||
Every cycle produces a knowledge graph node:
|
||||
- Title: `[LEARN] zulip-self-heal: <timestamp>`
|
||||
- metadata: { type: "remediation", status: "fixed" | "escalated" | "healthy" }
|
||||
This contract is RETIRED — health-check logs are NOT knowledge graph content.
|
||||
No graph nodes are created. Logs go to Gitea (SyslogSolution/health-logs).
|
||||
- Issues fixed → DM: "🛠 Zulip Self-Heal: fixed <issue>"
|
||||
- Issues escalated → DM: "⚠️ Zulip Self-Heal: <issue> needs attention"
|
||||
|
||||
|
||||
Reference in New Issue
Block a user