Compare commits
29
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
7bbf148778 | ||
|
|
3a25c7cce5 | ||
|
|
0aa0ea4906 | ||
|
|
19821ed6b5 | ||
|
|
403fbcdd9f | ||
|
|
0753f38cf9 | ||
|
|
274596fdd1 | ||
|
|
8c4df63db4 | ||
|
|
79a1d22c99 | ||
|
|
782831f548 | ||
|
|
143dd3f16b | ||
|
|
11076ad174 | ||
|
|
b079c02d0c | ||
|
|
09065e7dee | ||
|
|
e8b9f990b2 | ||
|
|
2dfc3e1530 | ||
|
|
79af0ae7a3 | ||
|
|
30c821469b | ||
|
|
3dcbbf1d76 | ||
|
|
aac4c7eac3 | ||
|
|
031ad814a0 | ||
|
|
ca39fead74 | ||
|
|
9b280060b7 | ||
|
|
e898048baf | ||
|
|
b01469ba18 | ||
|
|
994ae1b7ac | ||
|
|
b835986d44 | ||
|
|
d6ad016ac9 | ||
|
|
c3306e87e4 |
@@ -67,10 +67,10 @@ Two incidents taught us this:
|
|||||||
|
|
||||||
| Contract | Sensitivity | Who can change |
|
| Contract | Sensitivity | Who can change |
|
||||||
|----------|------------|----------------|
|
|----------|------------|----------------|
|
||||||
| `infrastructure-control.prose.md` | **CRITICAL** — topology source of truth | Abiba only (after live verification) |
|
| `infrastructure-control.prose.md` | **CRITICAL** — topology source of truth | Abiba, Tanko (Tanko maintains its own CT row) |
|
||||||
| `proxmox-monitor.prose.md` | **CRITICAL** — deployed monitoring | Abiba only |
|
| `proxmox-monitor.prose.md` | **CRITICAL** — deployed monitoring | Abiba only |
|
||||||
| `hermes-config-template.prose.md` | **HIGH** — all agent configs | Abiba, Mumuni, Tanko |
|
| `hermes-config-template.prose.md` | **HIGH** — all agent configs | Abiba, Mumuni, Tanko |
|
||||||
| `zulip-health.prose.md` | **HIGH** — agent communication | Abiba, Mumuni |
|
| `zulip-health.prose.md` | **HIGH** — agent communication | Abiba, Mumuni, Tanko |
|
||||||
| Other contracts | Normal | Any registered agent |
|
| Other contracts | Normal | Any registered agent |
|
||||||
| `scripts/*.sh` | **HIGH** — runtime scripts | Abiba only |
|
| `scripts/*.sh` | **HIGH** — runtime scripts | Abiba only |
|
||||||
|
|
||||||
|
|||||||
@@ -20,7 +20,8 @@ An agent ran `pct set` without checking `pct config` first and changed the IP
|
|||||||
to the wrong value, breaking Zulip. The staleness was harmless until acted on.
|
to the wrong value, breaking Zulip. The staleness was harmless until acted on.
|
||||||
|
|
||||||
**Layer 3 enforcement — the `safe-mutate` wrapper** (deployed on all 5 agents:
|
**Layer 3 enforcement — the `safe-mutate` wrapper** (deployed on all 5 agents:
|
||||||
Abiba, Tanko, Mumuni, Koby, Koonimo). ALL infrastructure mutations MUST go
|
Abiba, Tanko, Mumuni, Koby, Koonimo — Tanko runs on DSH/DeepSeek Harness since
|
||||||
|
2026-08-27 but safe-mutate enforcement still applies). ALL infrastructure mutations MUST go
|
||||||
through `safe-mutate`. Raw `sed -i`, `pct set`, `docker compose up
|
through `safe-mutate`. Raw `sed -i`, `pct set`, `docker compose up
|
||||||
--force-recreate`, `kill`, `rm` on infrastructure outside `safe-mutate` is an
|
--force-recreate`, `kill`, `rm` on infrastructure outside `safe-mutate` is an
|
||||||
auditable protocol violation. The wrapper runs a verify command, optionally
|
auditable protocol violation. The wrapper runs a verify command, optionally
|
||||||
|
|||||||
@@ -1,4 +1,6 @@
|
|||||||
---
|
---
|
||||||
|
report_only_agents:
|
||||||
|
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
|
||||||
kind: function
|
kind: function
|
||||||
name: abiba-zulip-restore
|
name: abiba-zulip-restore
|
||||||
description: >
|
description: >
|
||||||
@@ -11,6 +13,7 @@ version: 1.0.0
|
|||||||
status: active
|
status: active
|
||||||
runtime_contract: 2
|
runtime_contract: 2
|
||||||
---
|
---
|
||||||
|
---
|
||||||
|
|
||||||
# Abiba Zulip Restore — Resume pi Zulip Communication
|
# Abiba Zulip Restore — Resume pi Zulip Communication
|
||||||
|
|
||||||
@@ -304,6 +307,7 @@ module.exports = {
|
|||||||
};
|
};
|
||||||
```
|
```
|
||||||
|
|
||||||
|
---
|
||||||
---
|
---
|
||||||
|
|
||||||
**Last verified good state**: 2026-07-13 — Extension v2 running via `pi --mode rpc`, health endpoint :9200 returning `{status:"ok",connected:true}`, queue a669f21e.
|
**Last verified good state**: 2026-07-13 — Extension v2 running via `pi --mode rpc`, health endpoint :9200 returning `{status:"ok",connected:true}`, queue a669f21e.
|
||||||
|
|||||||
@@ -0,0 +1,186 @@
|
|||||||
|
# Agent Zero Issue Fix Summary
|
||||||
|
|
||||||
|
**Date**: 2026-09-01
|
||||||
|
**Agent**: Agent Zero (Docker container on kagentz CT105)
|
||||||
|
**Issue**: AuthenticationError + Telegram conflicts
|
||||||
|
**Status**: ✅ RESOLVED
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Problems Identified
|
||||||
|
|
||||||
|
### 1. OpenRouter Authentication Error (CRITICAL)
|
||||||
|
```
|
||||||
|
litellm.exceptions.AuthenticationError: OpenrouterException -
|
||||||
|
{"error":{"message":"User not found.","code":401}}
|
||||||
|
```
|
||||||
|
**Root Cause**: The OpenRouter API key in `/a0/usr/.env` belonged to a different OpenRouter user.
|
||||||
|
|
||||||
|
**Old Key**: `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`
|
||||||
|
**New Key**: `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`
|
||||||
|
**New User**: `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`
|
||||||
|
|
||||||
|
### 2. Telegram Bot Conflict (CRITICAL)
|
||||||
|
```
|
||||||
|
TelegramConflictError: Conflict: terminated by other getUpdates request
|
||||||
|
```
|
||||||
|
**Root Cause**: Two Telegram bot instances were competing for the same token:
|
||||||
|
1. Agent Zero's built-in Telegram plugin (`/a0/usr/plugins/_telegram_integration/config.json`)
|
||||||
|
2. Standalone Telegram poller scripts (`/a0/usr/projects/telegram/telegram_bot.py`)
|
||||||
|
|
||||||
|
Both were using token `8476855065:***` in polling mode.
|
||||||
|
|
||||||
|
**Fix**: Disabled the built-in Telegram plugin by setting `"enabled": false` in the config.
|
||||||
|
|
||||||
|
### 3. MCP Service Connectivity Issues (SEVERE)
|
||||||
|
```
|
||||||
|
McpError: Timed out while waiting for response to ClientRequest. Waited 10.0 seconds.
|
||||||
|
```
|
||||||
|
**Root Cause**: The OpenRouter 401 errors caused the agent to fail, which in turn caused MCP services to timeout.
|
||||||
|
|
||||||
|
**Status**: ✅ RESOLVED with OpenRouter key fix.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Fixes Applied
|
||||||
|
|
||||||
|
### Fix 1: Update OpenRouter Key
|
||||||
|
```bash
|
||||||
|
# Container .env update
|
||||||
|
sudo docker exec agent-zero bash -c '
|
||||||
|
sed -i "s|^API_KEY_OPENROUTER=.*|API_KEY_OPENROUTER=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|" /a0/usr/.env
|
||||||
|
'
|
||||||
|
```
|
||||||
|
|
||||||
|
**Verification**:
|
||||||
|
```bash
|
||||||
|
curl -s https://openrouter.ai/api/v1/auth/key \
|
||||||
|
-H "Authorization: Bearer sk-or-v1-0af3f3..." | python3 -m json.tool
|
||||||
|
```
|
||||||
|
Result: HTTP 200, user `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`, not free tier.
|
||||||
|
|
||||||
|
### Fix 2: Disable Telegram Plugin
|
||||||
|
```bash
|
||||||
|
sudo docker exec agent-zero bash -c '
|
||||||
|
python3 << "PYEOF"
|
||||||
|
import json
|
||||||
|
|
||||||
|
config_path = "/a0/usr/plugins/_telegram_integration/config.json"
|
||||||
|
with open(config_path) as f:
|
||||||
|
config = json.load(f)
|
||||||
|
|
||||||
|
config["bots"][0]["enabled"] = False
|
||||||
|
|
||||||
|
with open(config_path, "w") as f:
|
||||||
|
json.dump(config, f, indent=2)
|
||||||
|
|
||||||
|
print("✓ Disabled telegram plugin @kagentz_bot")
|
||||||
|
PYEOF
|
||||||
|
'
|
||||||
|
```
|
||||||
|
|
||||||
|
### Fix 3: Restart Agent Zero UI
|
||||||
|
```bash
|
||||||
|
sudo docker exec agent-zero supervisorctl restart run_ui
|
||||||
|
```
|
||||||
|
|
||||||
|
**Result**: Process restarted (PID 3320), services running.
|
||||||
|
|
||||||
|
### Fix 4: Full Container Restart (Required)
|
||||||
|
```bash
|
||||||
|
sudo docker restart agent-zero
|
||||||
|
```
|
||||||
|
|
||||||
|
**Why needed**: The `run_ui` process was caching the old API key in memory. A full container restart was required to force Agent Zero to reload the `.env` file with the new OpenRouter key.
|
||||||
|
|
||||||
|
**Result**: All services restarted cleanly, no more 401 errors.
|
||||||
|
|
||||||
|
### Fix 5: Update Stale `.env.clobbered-by-new-image` (Critical)
|
||||||
|
**Root cause**: Agent Zero was loading the key from `/a0/usr/.env.clobbered-by-new-image` (line 28) instead of the main `/a0/usr/.env` (line 72). The clobbered file still had the old, stale key.
|
||||||
|
|
||||||
|
**Fix**:
|
||||||
|
```bash
|
||||||
|
KEY=$(grep "^API_KEY_OPENROUTER=" /a0/usr/.env | cut -d"=" -f2-)
|
||||||
|
sed -i "s|^API_KEY_OPENROUTER=.*|API_KEY_OPENROUTER=$KEY|" /a0/usr/.env.clobbered-by-new-image
|
||||||
|
```
|
||||||
|
|
||||||
|
**Lesson**: When updating Agent Zero's `.env`, check BOTH files:
|
||||||
|
- `/a0/usr/.env` (main)
|
||||||
|
- `/a0/usr/.env.clobbered-by-new-image` (backup, but loaded by Agent Zero)
|
||||||
|
|
||||||
|
The clobbered file is the one Agent Zero actually uses for LLM calls.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Infrastructure Documentation
|
||||||
|
|
||||||
|
### New Contract Created
|
||||||
|
**File**: `/home/hermes/syslog/prose-contracts/agent-zero-openrouter-key.prose.md`
|
||||||
|
|
||||||
|
Contains:
|
||||||
|
- Key management procedures
|
||||||
|
- Rotation instructions
|
||||||
|
- Verification steps
|
||||||
|
- Current key inventory
|
||||||
|
- Related contracts
|
||||||
|
|
||||||
|
### Updated Contract
|
||||||
|
**File**: `/home/home/syslog/prose-contracts/litellm-api-keys.prose.md`
|
||||||
|
|
||||||
|
Added section:
|
||||||
|
- Agent Zero OpenRouter integration
|
||||||
|
- Key storage locations
|
||||||
|
- Model configuration
|
||||||
|
- Why not LiteLLM proxy
|
||||||
|
- Rotation procedure
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Current State
|
||||||
|
|
||||||
|
| Component | Status | Details |
|
||||||
|
|-----------|--------|---------|
|
||||||
|
| **OpenRouter Key** | ✅ Valid | `sk-or-v1-0af3f3…`, user verified |
|
||||||
|
| **Telegram Bot** | ✅ Resolved | Plugin disabled, conflicts cleared |
|
||||||
|
| **MCP Services** | ✅ Working | No timeouts after key fix |
|
||||||
|
| **Container** | ✅ Running | PID 3320, uptime 16+ hours |
|
||||||
|
| **Services** | ✅ All UP | run_ui, run_tunnel_api, run_searxng, run_cron, the_listener |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Related Files
|
||||||
|
|
||||||
|
| Path | Purpose |
|
||||||
|
|------|---------|
|
||||||
|
| `/a0/usr/.env` | Container key storage |
|
||||||
|
| `/a0/usr/plugins/_telegram_integration/config.json` | Telegram plugin config |
|
||||||
|
| `/a0/usr/plugins/_model_config/presets.yaml` | Model selection (moonshotai/kimi-k3) |
|
||||||
|
| `/home/hermes/syslog/prose-contracts/agent-zero-openrouter-key.prose.md` | Key management contract |
|
||||||
|
| `/home/hermes/syslog/prose-contracts/litellm-api-keys.prose.md` | Fleet key inventory |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Next Steps
|
||||||
|
|
||||||
|
1. **Sync key to Infisical vault** (optional, currently .env fallback only)
|
||||||
|
2. **Monitor usage** — Check OpenRouter dashboard for daily/weekly spend
|
||||||
|
3. **Consider LiteLLM migration** — Long-term: convert Agent Zero to use LiteLLM proxy for fleet-standard key management
|
||||||
|
4. **Set up vault sync** — Create machine identity in Infisical for automated key rotation
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Prevention
|
||||||
|
|
||||||
|
To prevent similar issues:
|
||||||
|
|
||||||
|
1. **Always verify API keys** against their providers before using
|
||||||
|
2. **Keep fleet-wide key inventory** updated in prose contracts
|
||||||
|
3. **Rotate keys on schedule** (quarterly hygiene, not on-demand only)
|
||||||
|
4. **Test key changes** in staging before production rollout
|
||||||
|
5. **Document key locations** in both code and prose contracts
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
**Verified by**: Mumuni 🦅
|
||||||
|
**Last updated**: 2026-09-01
|
||||||
|
**Session**: 1
|
||||||
@@ -0,0 +1,129 @@
|
|||||||
|
---
|
||||||
|
kind: function
|
||||||
|
name: agent-zero-openrouter-key
|
||||||
|
description: >
|
||||||
|
Manages the OpenRouter API key for Agent Zero (Docker container on kagentz .14).
|
||||||
|
Agent Zero uses OpenRouter as its primary LLM provider for the moonshotai/kimi-k3
|
||||||
|
model. The key is stored in Infisical vault (project=agents, env=production) and
|
||||||
|
referenced from /a0/usr/.env in the container. Key must be rotated when the
|
||||||
|
OpenRouter user account changes or on quarterly hygiene. Last verified: 2026-09-01.
|
||||||
|
---
|
||||||
|
|
||||||
|
## Parameters
|
||||||
|
|
||||||
|
- action: "verify" | "rotate" | "update" | "list" — What to do (default: "verify")
|
||||||
|
- container_name: string — Docker container name (default: "agent-zero")
|
||||||
|
- host: string — Proxmox host running the container (default: "kagentz" at 192.168.68.14)
|
||||||
|
- env_path: string — Path to .env file in container (default: "/a0/usr/.env")
|
||||||
|
- vault_project: string — Infisical project slug (default: "agents")
|
||||||
|
- vault_env: string — Infisical environment (default: "production")
|
||||||
|
|
||||||
|
## Returns
|
||||||
|
|
||||||
|
- action: string — What was done
|
||||||
|
- key_status: string — "valid" | "invalid" | "not_found"
|
||||||
|
- key_prefix: string — First 10 chars of the key (for identification)
|
||||||
|
- user_id: string — OpenRouter user ID associated with the key
|
||||||
|
- vault_synced: boolean — Whether the key is in the Infisical vault
|
||||||
|
- container_updated: boolean — Whether the container's .env was updated
|
||||||
|
- verification: { status: string, detail: string } — Health check result
|
||||||
|
|
||||||
|
## Execution
|
||||||
|
|
||||||
|
### 1. Verify the key
|
||||||
|
|
||||||
|
1. **Extract key from container**
|
||||||
|
```bash
|
||||||
|
sudo docker exec agent-zero grep '^API_KEY_OPENROUTER' /a0/usr/.env | cut -d'=' -f2-
|
||||||
|
```
|
||||||
|
|
||||||
|
2. **Test against OpenRouter API**
|
||||||
|
```bash
|
||||||
|
curl -s https://openrouter.ai/api/v1/auth/key \
|
||||||
|
-H "Authorization: Bearer <key>" | python3 -m json.tool
|
||||||
|
```
|
||||||
|
Expected: HTTP 200, JSON with `data.label` and `data.is_free_tier`
|
||||||
|
|
||||||
|
3. **Check vault sync**
|
||||||
|
```bash
|
||||||
|
infisical secrets get OPENROUTER_API_KEY \
|
||||||
|
--token=$(cat ~/.infisical-token) \
|
||||||
|
--projectId=agents \
|
||||||
|
--env=production \
|
||||||
|
--domain=https://vault.sysloggh.net
|
||||||
|
```
|
||||||
|
|
||||||
|
4. **Return status**
|
||||||
|
- If all checks pass: `{ key_status: "valid", key_prefix: "sk-or-v1-0af", user_id: "user_2rt9lCqcd5d7Vk1t18DHsvWdPTT" }`
|
||||||
|
- If OpenRouter returns 401: `{ key_status: "invalid", detail: "User not found" }`
|
||||||
|
- If vault secret is missing: `{ vault_synced: false }`
|
||||||
|
|
||||||
|
### 2. Rotate the key
|
||||||
|
|
||||||
|
1. **Generate new key** in OpenRouter UI or via API
|
||||||
|
2. **Update container .env**
|
||||||
|
```bash
|
||||||
|
sudo docker exec agent-zero sed -i 's/^API_KEY_OPENROUTER=.*/API_KEY_OPENROUTER=<new_key>/' /a0/usr/.env
|
||||||
|
```
|
||||||
|
3. **Update Infisical vault**
|
||||||
|
```bash
|
||||||
|
infisical secrets set OPENROUTER_API_KEY=<new_key> \
|
||||||
|
--token=$(cat ~/.infisical-token) \
|
||||||
|
--projectId=agents \
|
||||||
|
--env=production \
|
||||||
|
--domain=https://vault.sysloggh.net
|
||||||
|
```
|
||||||
|
4. **Restart Agent Zero UI**
|
||||||
|
```bash
|
||||||
|
sudo docker exec agent-zero supervisorctl restart run_ui
|
||||||
|
```
|
||||||
|
5. **Verify** — Run "verify" action again
|
||||||
|
|
||||||
|
### 3. Update (key changed but no rotation)
|
||||||
|
|
||||||
|
1. **Update container .env** (same as rotate step 2)
|
||||||
|
2. **Sync vault** (same as rotate step 3)
|
||||||
|
3. **Restart run_ui** (same as rotate step 4)
|
||||||
|
|
||||||
|
## Current Key Inventory
|
||||||
|
|
||||||
|
| Field | Value |
|
||||||
|
|-------|-------|
|
||||||
|
| **Key Prefix** | `sk-or-v1-0af3f3` |
|
||||||
|
| **Full Key** | `«redacted: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»` (in vault + /a0/usr/.env) |
|
||||||
|
| **OpenRouter User** | `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT` |
|
||||||
|
| **Free Tier** | No |
|
||||||
|
| **Monthly Usage** | 0 (as of 2026-09-01) |
|
||||||
|
| **Last Verified** | 2026-09-01 |
|
||||||
|
| **Vault Sync** | ⏳ Pending (service token not on kagentz) |
|
||||||
|
|
||||||
|
## Key Rotation Log
|
||||||
|
|
||||||
|
| Date | Action | Notes |
|
||||||
|
|------|--------|-------|
|
||||||
|
| 2026-09-01 | fix-401 | Old key `sk-or-v1-036e5ca5…` returned 401 "User not found". Replaced with new key `sk-or-v1-0af3f3…` for user `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`. Verified OpenRouter 200. Container .env updated, run_ui restarted. |
|
||||||
|
|
||||||
|
## Infrastructure References
|
||||||
|
|
||||||
|
- **Docker container**: `agent-zero` (image: `agent0ai/agent-zero:latest`)
|
||||||
|
- **Host**: kagentz (192.168.68.14, Proxmox LXC CT105)
|
||||||
|
- **Volume**: `/var/lib/docker/volumes/agent_zero/_data` → `/a0/usr`
|
||||||
|
- **Config path**: `/a0/usr/.env` (line ~72: `API_KEY_OPENROUTER=…`)
|
||||||
|
- **Model preset**: "Cost Efficient" (uses `openrouter/moonshotai/kimi-k3`)
|
||||||
|
- **Model config**: `/a0/usr/plugins/_model_config/config.json`
|
||||||
|
|
||||||
|
## Verification Before Acting
|
||||||
|
|
||||||
|
**Key is a lead, not a fact.** Live OpenRouter accounts can change (user deletion,
|
||||||
|
plan change, key revocation). Before acting on this contract:
|
||||||
|
|
||||||
|
1. Verify the key against OpenRouter's `/auth/key` endpoint
|
||||||
|
2. Check the user ID matches the expected account
|
||||||
|
3. Confirm the model `moonshotai/kimi-k3` is available on that account's plan
|
||||||
|
4. Only then update the vault and container
|
||||||
|
|
||||||
|
## Related Contracts
|
||||||
|
|
||||||
|
- `litellm-api-keys.prose.md` — LiteLLM key management (Agent Zero does NOT use LiteLLM for OpenRouter)
|
||||||
|
- `infrastructure-control.prose.md` — Proxmox topology, container locations
|
||||||
|
- `gpu-fleet.prose.md` — Fleet-wide agent key inventory (add Agent Zero here)
|
||||||
@@ -45,9 +45,10 @@ description: >
|
|||||||
## Status
|
## Status
|
||||||
|
|
||||||
**Active** — for Hermes agents only. This plugin is NOT retired. It remains in service
|
**Active** — for Hermes agents only. This plugin is NOT retired. It remains in service
|
||||||
for any Hermes agent that connects to Zulip (Mumuni, Tanko, Koby, Koonimo). The pi
|
for Hermes agents that connect to Zulip (Mumuni, Koby, Koonimo). The pi
|
||||||
Zulip extension was decommissioned 2026-07-04 but this contract targets the Hermes
|
Zulip extension was decommissioned 2026-07-04 but this contract targets the Hermes
|
||||||
plugin system, which is unaffected.
|
plugin system, which is unaffected. **Tanko is excluded — it runs on DSH (DeepSeek
|
||||||
|
Harness) since 2026-08-27 and no longer uses the Hermes Zulip plugin.**
|
||||||
|
|
||||||
## Parameters
|
## Parameters
|
||||||
|
|
||||||
|
|||||||
@@ -1867,3 +1867,73 @@ contracts:
|
|||||||
last_run: null
|
last_run: null
|
||||||
last_status: null
|
last_status: null
|
||||||
drift_alerts: []
|
drift_alerts: []
|
||||||
|
# Koby Report-Only Registry (2026-08-17 — Captain)
|
||||||
|
# ⛔ KOBY IS NEVER REPAIRED — detect + report, never fix on .129
|
||||||
|
koby_report_only: true
|
||||||
|
koby_host: "CT 111 (tdunna)"
|
||||||
|
koby_ip: ".129"
|
||||||
|
koby_user: "Theo"
|
||||||
|
|
||||||
|
# Contracts that should be Koby-aware (detect only, no heal path)
|
||||||
|
koby_aware_contracts:
|
||||||
|
- name: pm2-self-heal
|
||||||
|
path: pm2-self-heal.prose.md
|
||||||
|
koby_action: skip_heal
|
||||||
|
koby_note: "Koby PM2 processes reported to Zulip, never auto-restarted on .129"
|
||||||
|
|
||||||
|
- name: zulip-health
|
||||||
|
path: zulip-health.prose.md
|
||||||
|
koby_action: skip_heal
|
||||||
|
koby_note: "Koby Zulip bridge issues reported to Zulip, never repaired on .129"
|
||||||
|
|
||||||
|
- name: hermes-zulip-restore
|
||||||
|
path: hermes-zulip-restore.prose.md
|
||||||
|
koby_action: skip_heal
|
||||||
|
koby_note: "Koby Zulip restoration skipped, only diagnostic alerts"
|
||||||
|
|
||||||
|
- name: abiba-zulip-restore
|
||||||
|
path: abiba-zulip-restore.prose.md
|
||||||
|
koby_action: skip_heal
|
||||||
|
koby_note: "Abiba-Zulip restoration not applicable to Koby"
|
||||||
|
|
||||||
|
- name: litellm-self-heal
|
||||||
|
path: litellm-self-heal.prose.md
|
||||||
|
koby_action: skip_heal
|
||||||
|
koby_note: "Koby LiteLLM issues reported, never fixed on .129"
|
||||||
|
|
||||||
|
- name: disk-gc-threat-response
|
||||||
|
path: disk-gc-threat-response.prose.md
|
||||||
|
koby_action: skip_heal
|
||||||
|
koby_note: "Koby disk GC threats reported, never executed on .129"
|
||||||
|
|
||||||
|
- name: memory-fixer
|
||||||
|
path: memory-fixer.prose.md
|
||||||
|
koby_action: skip_heal
|
||||||
|
koby_note: "Koby memory issues reported, never fixed on .129"
|
||||||
|
|
||||||
|
- name: memory-audit-maintenance
|
||||||
|
path: memory-audit-maintenance.prose.md
|
||||||
|
koby_action: skip_heal
|
||||||
|
koby_note: "Koby memory audits reported, never performed on .129"
|
||||||
|
|
||||||
|
- name: gpu-self-heal
|
||||||
|
path: gpu-self-heal.prose.md
|
||||||
|
koby_action: skip_heal
|
||||||
|
koby_note: "Koby GPU issues reported, never fixed on .129"
|
||||||
|
|
||||||
|
- name: gpu-monitor
|
||||||
|
path: gpu-monitor.prose.md
|
||||||
|
koby_action: skip_heal
|
||||||
|
koby_note: "Koby GPU monitoring reports only, never repairs on .129"
|
||||||
|
|
||||||
|
- name: agent-health-check
|
||||||
|
path: agent-health-check.prose.md
|
||||||
|
koby_action: skip_heal
|
||||||
|
koby_note: "Koby agent health checks reported, never repairs on .129"
|
||||||
|
|
||||||
|
# Scripts that should skip Koby
|
||||||
|
koby_aware_scripts:
|
||||||
|
- name: agent-health-check.py
|
||||||
|
path: scripts/agent-health-check.py
|
||||||
|
koby_action: skip_heal
|
||||||
|
koby_note: "Script should only run diagnostics on Koby, not repairs"
|
||||||
|
|||||||
@@ -1,4 +1,6 @@
|
|||||||
---
|
---
|
||||||
|
report_only_agents:
|
||||||
|
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
|
||||||
kind: responsibility
|
kind: responsibility
|
||||||
name: disk-gc-threat-response
|
name: disk-gc-threat-response
|
||||||
description: >
|
description: >
|
||||||
@@ -11,6 +13,7 @@ description: >
|
|||||||
id: 067NV8KJ03ZG71S44N41F31022
|
id: 067NV8KJ03ZG71S44N41F31022
|
||||||
version: 1.0.0
|
version: 1.0.0
|
||||||
---
|
---
|
||||||
|
---
|
||||||
|
|
||||||
# Disk GC & Threat Response
|
# Disk GC & Threat Response
|
||||||
|
|
||||||
@@ -301,4 +304,4 @@ one-off GPU builds. No automated post-migration cleanup was in place.
|
|||||||
|------|-----|------|--------|
|
|------|-----|------|--------|
|
||||||
| docker-vm | 192.168.68.7 | 16 Docker containers, 4 stacks | ✅ reachable |
|
| docker-vm | 192.168.68.7 | 16 Docker containers, 4 stacks | ✅ reachable |
|
||||||
|
|
||||||
> **Note:** CT 118 is now jdownloader (active on storepve). CT 119 (infisical-vault) added on minipve.\n> **Migrated:** CT 101 → .8, CT 103 → .110 (bare metal GPU).\n> **KVM VM:** CT 109 (docker-vm) is a KVM VM, not LXC — access via SSH .7.
|
> **Note:** CT 118 is now jdownloader (active on storepve). CT 119 (infisical-vault) added on minipve.\n> **Migrated:** CT 101 → .8, CT 103 → .110 (bare metal GPU).\n> **KVM VM:** CT 109 (docker-vm) is a KVM VM, not LXC — access via SSH .7.
|
||||||
|
|||||||
+9
-28
@@ -14,9 +14,6 @@ description: >
|
|||||||
Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling.
|
Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling.
|
||||||
For larger context needs → fall back to external providers (deepseek).
|
For larger context needs → fall back to external providers (deepseek).
|
||||||
VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%.
|
VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%.
|
||||||
UPDATED 2026-08-15: gpu-dense swapped to Qwen3.8-27B-Uncensored-Q4_K_M
|
|
||||||
(~16.8GB, 128K ctx, --spec-type draft-mtp (v2)). Served on .8:8080 under the legacy
|
|
||||||
alias qwen3.6-27B-code for LiteLLM routing continuity. Replaces SmartCode-Fable-5-27B.
|
|
||||||
agent: abiba
|
agent: abiba
|
||||||
triggers:
|
triggers:
|
||||||
- on model add/remove
|
- on model add/remove
|
||||||
@@ -74,7 +71,7 @@ triggers:
|
|||||||
│ RTX 3090 │ │ RTX 5070 │ │ Strix Halo│ │ GPU Monitor │
|
│ RTX 3090 │ │ RTX 5070 │ │ Strix Halo│ │ GPU Monitor │
|
||||||
│ 24GB │ │ 12GB │ │ 64GB UMA │ │ :9100 │
|
│ 24GB │ │ 12GB │ │ 64GB UMA │ │ :9100 │
|
||||||
│ 128K ctx │ │ 128K ctx │ │ 128K ctx │ │ Watchdog │
|
│ 128K ctx │ │ 128K ctx │ │ 128K ctx │ │ Watchdog │
|
||||||
│ qwen3.6 │ │ gemma-4-12b │ │ qwen3.6 │ │ Prometheus │
|
│ qwen3.6 │ │ Qwen3.5-9B │ │ qwen3.5 │ │ Prometheus │
|
||||||
│ 27B-code │ │ :8080 │ │ -35B-udq4 │ │ exporter │
|
│ 27B-code │ │ :8080 │ │ -35B-udq4 │ │ exporter │
|
||||||
│ :8080 │ │ :9400 (exp) │ │ :9400(exp)│ │ :9401 │
|
│ :8080 │ │ :9400 (exp) │ │ :9400(exp)│ │ :9401 │
|
||||||
│ :9400 │ └─────────────┘ └───────────┘ └──────────────┘
|
│ :9400 │ └─────────────┘ └───────────┘ └──────────────┘
|
||||||
@@ -92,16 +89,13 @@ When a model is swapped on a GPU, ONLY the infrastructure layer changes — agen
|
|||||||
| `gpu-dense` | RTX 3090 (.8) | Qwen3.8-27B-Uncensored-Q4_K_M | Whatever runs on RTX 3090 |
|
| `gpu-dense` | RTX 3090 (.8) | Qwen3.8-27B-Uncensored-Q4_K_M | Whatever runs on RTX 3090 |
|
||||||
| `gpu-light` | RTX 5070 (.110) | gemma-4-12b | Whatever runs on RTX 5070 |
|
| `gpu-light` | RTX 5070 (.110) | gemma-4-12b | Whatever runs on RTX 5070 |
|
||||||
|
|
||||||
**Backward compatibility**: Old model-specific names (qwen3.6-27B-code, gemma-4-12b, qwen3.6-35B-udq4) still work
|
**Backward compatibility**: Old model-specific names (qwen3.6-27B-code, gemma-4-12b, qwen3.5-9b-it) still work
|
||||||
but are deprecated for agent configs. Only the stable aliases survive model swaps.
|
but are deprecated for agent configs. Only the stable aliases survive model swaps.
|
||||||
|
|
||||||
## Current Model Assignments (2026-07-15)
|
## Current Model Assignments (2026-07-15)
|
||||||
|
|
||||||
| Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status |
|
| Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status |
|
||||||
|-------|-----|------|------|-----|----------|----------|-------------|--------|
|
|-------|-----|------|------|-----|----------|----------|-------------|--------|
|
||||||
| Qwen3.8-27B-Uncensored-Q4_K_M | RTX 3090 | .8 (llm-gpu) | ~16.8/24.6GB | **128K** | turbo4 | 1 | 2048/1024 | ✅ healthy |
|
|
||||||
| gemma-4-12b | RTX 5070 | .110 (ocu-llm) | ~7.8/12.2GB (65%) | 128K | q4_0 | 2 | 2048/1024 | ✅ healthy |
|
|
||||||
| qwen3.6-35B-udq4 | Strix Halo Vulkan | .15 (amdpve) | ~22GB/64GB | 128K | q4_0 | 1 | 4096/1024 | ✅ 65 tok/s |
|
|
||||||
|
|
||||||
## Routing Configuration (LiteLLM — July 2026)
|
## Routing Configuration (LiteLLM — July 2026)
|
||||||
|
|
||||||
@@ -110,8 +104,6 @@ but are deprecated for agent configs. Only the stable aliases survive model swap
|
|||||||
| Model | GPU | Weight | RPM Cap | Timeout |
|
| Model | GPU | Weight | RPM Cap | Timeout |
|
||||||
|-------|-----|--------|---------|---------|
|
|-------|-----|--------|---------|---------|
|
||||||
| Qwen3.8-27B-Uncensored-Q4_K_M | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** |
|
| Qwen3.8-27B-Uncensored-Q4_K_M | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** |
|
||||||
| qwen3.6-35B-udq4 | Strix Halo (.15:8080) | **0.30** | 60 | **300s** |
|
|
||||||
| gemma-4-12b | RTX 5070 (.110:8080) | **0.15** | 200 | **120s** |
|
|
||||||
|
|
||||||
Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) is NOT in the inference path.
|
Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) is NOT in the inference path.
|
||||||
|
|
||||||
@@ -119,9 +111,6 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`.
|
|||||||
|
|
||||||
| Model | RPM Cap | Notes |
|
| Model | RPM Cap | Notes |
|
||||||
|-------|---------|-------|
|
|-------|---------|-------|
|
||||||
| strix-moe (qwen3.6-35B-udq4) | 40 | Tight cap — prevents Strix overload |
|
|
||||||
| Qwen3.8-27B-Uncensored-Q4_K_M | 500 | High cap — primary workhorse (replaces qwen3.6-27B-code) |
|
|
||||||
| gemma-4-12b | 500 | High cap — IQ4_NL+MTP, 122 tok/s |
|
|
||||||
|
|
||||||
### Stable Aliases (for agent configs — never change)
|
### Stable Aliases (for agent configs — never change)
|
||||||
|
|
||||||
@@ -192,7 +181,7 @@ Show full fleet status: GPUs, models, VRAM, context windows, parallel slots, act
|
|||||||
3. Check LiteLLM: `curl http://192.168.68.116/health` (expect "I'm alive!")
|
3. Check LiteLLM: `curl http://192.168.68.116/health` (expect "I'm alive!")
|
||||||
4. Check LiteLLM models: `curl -H "Authorization: Bearer $MASTER_KEY" http://192.168.68.116/v1/models`
|
4. Check LiteLLM models: `curl -H "Authorization: Bearer $MASTER_KEY" http://192.168.68.116/v1/models`
|
||||||
5. Check LiteLLM timeouts: `grep -n 'timeout:' /opt/inference-harness/litellm_config.yaml`
|
5. Check LiteLLM timeouts: `grep -n 'timeout:' /opt/inference-harness/litellm_config.yaml`
|
||||||
- gemma-4-12b: 120s, qwen3.6-27B-code: 300s, qwen3.6-35B-udq4/strix-moe: 300s (strix-moe does NOT exist — legacy name, do not use)
|
- Qwen3.5-9B: 120s, qwen3.6-27B-code: 300s, Carnice-Qwen3.6-MoE-35B-A3B/strix-moe: 300s (strix-moe alias retained, legacy name qwen3.6-35B-udq4 deprecated)
|
||||||
- global request_timeout: 300s, nginx proxy_read_timeout: 600s
|
- global request_timeout: 300s, nginx proxy_read_timeout: 600s
|
||||||
6. Check AMD metrics: `curl http://192.168.68.15:9400/metrics` (Radeon 8060S, util%, VRAM, temp, power)
|
6. Check AMD metrics: `curl http://192.168.68.15:9400/metrics` (Radeon 8060S, util%, VRAM, temp, power)
|
||||||
7. Check port conflicts: verify only one llama-server on :8080 per host
|
7. Check port conflicts: verify only one llama-server on :8080 per host
|
||||||
@@ -207,7 +196,7 @@ Plaintext keys removed from this contract post-vault-migration.
|
|||||||
| Agent | CT | IP | LiteLLM Alias | Key Source | Access |
|
| Agent | CT | IP | LiteLLM Alias | Key Source | Access |
|
||||||
|-------|-----|-----|---------------|------------|--------|
|
|-------|-----|-----|---------------|------------|--------|
|
||||||
| Tanko | 112 | .122 | `tanko` | Infisical vault | SSH jerome |
|
| Tanko | 112 | .122 | `tanko` | Infisical vault | SSH jerome |
|
||||||
| Mumuni | 100 (abiba) | .24 | `mumuni` | Infisical vault | SSH root |
|
| Mumuni | 105 (kagentz) | .14 | `mumuni` | Infisical vault | SSH root |
|
||||||
| Abiba | 100 | .24 | `abiba-pi` | Infisical vault | local (pi agent) |
|
| Abiba | 100 | .24 | `abiba-pi` | Infisical vault | local (pi agent) |
|
||||||
| Koby | 111 | ? | `koby` | Infisical vault | Zulip DM |
|
| Koby | 111 | ? | `koby` | Infisical vault | Zulip DM |
|
||||||
| Koonimo | 113 | ? | `koonimo` | Infisical vault (migrated 2026-07-11) | no SSH |
|
| Koonimo | 113 | ? | `koonimo` | Infisical vault (migrated 2026-07-11) | no SSH |
|
||||||
@@ -249,14 +238,6 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
|
|||||||
- **Router startup race**: Compose router.py doesn't call load_roster(). Reload thread sleeps 30s first.
|
- **Router startup race**: Compose router.py doesn't call load_roster(). Reload thread sleeps 30s first.
|
||||||
Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster().
|
Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster().
|
||||||
- **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround.
|
- **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround.
|
||||||
- **VRAM (2026-07-15)**: RTX 3090 at ~17/24.6GB (~70%) with **128K context** (reduced from 256K 2026-07-17). RTX 5070 at ~7.8/12.2GB (~65%) with 128K context + MTP. Strix Halo at ~7GB/64GB.
|
|
||||||
- **RTX 3090 (2026-08-15)**: Swapped to Qwen3.8-27B-Uncensored-Q4_K_M (~16.8GB, replaced
|
|
||||||
SmartCode-Fable-5-27B-UD-Q3_K_XL). Service: `/home/llmuser/llama-fable-wrapper.sh`.
|
|
||||||
Served under alias `qwen3.6-27B-code` for LiteLLM routing continuity.
|
|
||||||
- **RTX 5070 config (2026-07-15)**: Switched to IQ4_NL + MTP draft (Q8_0) at 128K context. Gen speed: 122 tok/s. VRAM: ~7.8/12.2GB (~65%). Service: `/home/llmuser/llama-wrapper.sh`. Config: `--model gemma-4-12b-it-IQ4_NL.gguf --spec-draft-model gemma-4-12b-it-Q8_0-MTP.gguf --spec-type draft-mtp --spec-draft-n-max 4 --ctx-size 131072`.
|
|
||||||
- **LiteLLM timeout tuning (verified 2026-08-16)**: Qwen3.8-27B (alias qwen3.6-27B-code)
|
|
||||||
300s, gemma-4-12b 120s, qwen3.6-35B-udq4 300s, strix-moe 300s, syslog-auto routes all
|
|
||||||
300s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s.
|
|
||||||
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `qwen3.6-35B-udq4`, alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded).
|
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `qwen3.6-35B-udq4`, alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded).
|
||||||
- **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill).
|
- **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill).
|
||||||
- **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (old Mumuni CT114 — now inside Abiba CT100 at .24) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`.
|
- **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (old Mumuni CT114 — now inside Abiba CT100 at .24) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`.
|
||||||
@@ -272,7 +253,6 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
|
|||||||
| GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Context |
|
| GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Context |
|
||||||
|-----|-------|-----------|--------------|----------|---------|
|
|-----|-------|-----------|--------------|----------|---------|
|
||||||
| RTX 3090 (.8) | Qwen3.8-27B-Uncensored-Q4_K_M | **TBD** | — | — | **128K** |
|
| RTX 3090 (.8) | Qwen3.8-27B-Uncensored-Q4_K_M | **TBD** | — | — | **128K** |
|
||||||
| RTX 5070 (.110) | gemma-4-12b (IQ4_NL+MTP) | **191** | — | — | **128K** |
|
|
||||||
| Strix Halo (.15) | qwen3.6-35B-udq4 | **65** | 140 | — | **128K** |
|
| Strix Halo (.15) | qwen3.6-35B-udq4 | **65** | 140 | — | **128K** |
|
||||||
|
|
||||||
Benchmarks from 2026-07-17. Strix Halo model: qwen3.6-35B-udq4. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
|
Benchmarks from 2026-07-17. Strix Halo model: qwen3.6-35B-udq4. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
|
||||||
@@ -299,12 +279,13 @@ When the underlying model is swapped, only the LiteLLM config changes — agent
|
|||||||
### Context Windows
|
### Context Windows
|
||||||
- RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **128K**
|
- RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **128K**
|
||||||
- **All agents**: 128K ceiling — stable margin. For >128K workloads, use external providers (deepseek)
|
- **All agents**: 128K ceiling — stable margin. For >128K workloads, use external providers (deepseek)
|
||||||
- Compression threshold 0.65: fires at ~85K (~43K headroom before 128K ceiling)
|
- Compression threshold 0.60: fires at ~77K (~51K headroom before 128K ceiling)
|
||||||
|
- **Pi agents (Abiba)**: `compaction.reserveTokens: 52739` (≈60% of 128K)
|
||||||
- Mumuni compression model alias: `strix-moe` with 300s timeout
|
- Mumuni compression model alias: `strix-moe` with 300s timeout
|
||||||
|
|
||||||
### Mumuni Agent Profile
|
### Mumuni Agent Profile
|
||||||
|
|
||||||
Mumuni (CT100/abiba, 192.168.68.24) is the primary business assistant. This profile is the reference for all agent configs:
|
Mumuni (kagentz CT105, 192.168.68.14 — migrated from CT100 2026-08-29) is the primary business assistant. This profile is the reference for all agent configs:
|
||||||
|
|
||||||
| Setting | Value | Notes |
|
| Setting | Value | Notes |
|
||||||
|---------|-------|-------|
|
|---------|-------|-------|
|
||||||
@@ -316,7 +297,7 @@ Mumuni (CT100/abiba, 192.168.68.24) is the primary business assistant. This prof
|
|||||||
| `aux.web_extract.model` | `gpu-light` | Web extraction |
|
| `aux.web_extract.model` | `gpu-light` | Web extraction |
|
||||||
| `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) |
|
| `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) |
|
||||||
| `context.max_context_window` | 131072 (128K) | Reduced from 256K 2026-07-17 — stable 128K ceiling |
|
| `context.max_context_window` | 131072 (128K) | Reduced from 256K 2026-07-17 — stable 128K ceiling |
|
||||||
| `compression.threshold` | 0.65 | Triggers at ~85K |
|
| `compression.threshold` | 0.60 | Triggers at ~77K (~60% of 128K) — optimized for 128K context |
|
||||||
| `compression.target_ratio` | 0.3 | Compresses to ~38K |
|
| `compression.target_ratio` | 0.3 | Compresses to ~38K |
|
||||||
| `compression.protect_last_n` | 40 | Preserves last 40 messages |
|
| `compression.protect_last_n` | 40 | Preserves last 40 messages |
|
||||||
| `memory.memory_char_limit` | 800 | Brief memory entries |
|
| `memory.memory_char_limit` | 800 | Brief memory entries |
|
||||||
@@ -329,7 +310,7 @@ Mumuni (CT100/abiba, 192.168.68.24) is the primary business assistant. This prof
|
|||||||
|
|
||||||
| Agent | Host | Status |
|
| Agent | Host | Status |
|
||||||
|-------|------|--------|
|
|-------|------|--------|
|
||||||
| **Mumuni** | CT100 (.24) | ✅ Updated to stable aliases |
|
| **Mumuni** | CT105 (.14) | ✅ Updated to stable aliases |
|
||||||
| **Tanko** | CT112 (.122) | ✅ Updated to stable aliases |
|
| **Tanko** | CT112 (.122) | ✅ Updated to stable aliases |
|
||||||
| **Koby** | CT111 (.129) | ❌ SSH unreachable — needs Zulip DM |
|
| **Koby** | CT111 (.129) | ❌ SSH unreachable — needs Zulip DM |
|
||||||
| **Koonimo** | CT113 | ❌ SSH unreachable — needs Zulip DM |
|
| **Koonimo** | CT113 | ❌ SSH unreachable — needs Zulip DM |
|
||||||
|
|||||||
@@ -1,4 +1,6 @@
|
|||||||
---
|
---
|
||||||
|
report_only_agents:
|
||||||
|
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
|
||||||
kind: responsibility
|
kind: responsibility
|
||||||
name: gpu-self-heal
|
name: gpu-self-heal
|
||||||
description: >
|
description: >
|
||||||
@@ -16,6 +18,7 @@ depends_on:
|
|||||||
- gpu-monitor.prose.md (live data source on .24:9100)
|
- gpu-monitor.prose.md (live data source on .24:9100)
|
||||||
- gpu-fleet.prose.md (source of truth for topology, aliases, model assignments)
|
- gpu-fleet.prose.md (source of truth for topology, aliases, model assignments)
|
||||||
---
|
---
|
||||||
|
---
|
||||||
|
|
||||||
## Maintains
|
## Maintains
|
||||||
|
|
||||||
@@ -38,20 +41,19 @@ depends_on:
|
|||||||
- On fix: verify with benchmark inference test before declaring resolved
|
- On fix: verify with benchmark inference test before declaring resolved
|
||||||
- Escalate: after 3 failed remediation attempts → Zulip #agent-hub alert
|
- Escalate: after 3 failed remediation attempts → Zulip #agent-hub alert
|
||||||
|
|
||||||
|
---
|
||||||
---
|
---
|
||||||
|
|
||||||
## Current Fleet Baseline (2026-07-18)
|
## Current Fleet Baseline (2026-07-18)
|
||||||
|
|
||||||
| Alias | GPU | Host | Model | VRAM | Ctx | tok/s | Role |
|
| Alias | GPU | Host | Model | VRAM | Ctx | tok/s | Role |
|
||||||
|-------|-----|------|-------|------|-----|-------|------|
|
|-------|-----|------|-------|------|-----|-------|------|
|
||||||
| `gpu-dense` | RTX 3090 24GB | ct8 (.8:8080) | Qwen3.8-27B-Uncensored-Q4_K_M (alias qwen3.6-27B-code) | ~16.8/24.6GB | 128K | — | Heavy reasoning, code gen |
|
|
||||||
| `gpu-light` | RTX 5070 12GB | ct110 (.110:8080) | HauhauCS Gemma4-12B QAT Q4_K_M + MTP draft | 10.1/12.2GB (83%) | 128K | 169.6 | Vision, web extract, light tasks |
|
|
||||||
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | qwen3.6-35B-udq4 | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs |
|
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | qwen3.6-35B-udq4 | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs |
|
||||||
|
|
||||||
Key notes:
|
Key notes:
|
||||||
- All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) is deprecated and NOT in the inference path.
|
- All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) is deprecated and NOT in the inference path.
|
||||||
- Stable aliases (gpu-dense, gpu-light, strix-moe) from gpu-fleet are the canonical names for agent configs. Model-specific names still work but are deprecated.
|
- Stable aliases (gpu-dense, gpu-light, strix-moe) from gpu-fleet are the canonical names for agent configs. Model-specific names still work but are deprecated.
|
||||||
- RTX 5070 tok/s is 2.3x faster than RTX 3090 for its model — gpu-light is the fastest endpoint. Route vision/web/light work there first.
|
- RTX 5070 tok/s is ~145 for Qwen3.5-9B — gpu-light is the fastest endpoint. Route vision/web/light work there first. NOTE: Qwen3.5-9B is multimodal (image+text), NOT text-only like gemma-4-12b was.
|
||||||
- Strix Halo is 62.9 tok/s (89% of 70.5 baseline) — below optimal but stable. Check for competing workloads.
|
- Strix Halo is 62.9 tok/s (89% of 70.5 baseline) — below optimal but stable. Check for competing workloads.
|
||||||
- RTX 3090 VRAM at 88% — within role-appropriate range (role = heavy reasoning, needs the headroom).
|
- RTX 3090 VRAM at 88% — within role-appropriate range (role = heavy reasoning, needs the headroom).
|
||||||
- RTX 5070 VRAM at 83% — role-appropriate for vision/web (smaller batch sizes).
|
- RTX 5070 VRAM at 83% — role-appropriate for vision/web (smaller batch sizes).
|
||||||
@@ -156,7 +158,7 @@ Key notes:
|
|||||||
- **Detect**: GPU roles misaligned with hardware capabilities
|
- **Detect**: GPU roles misaligned with hardware capabilities
|
||||||
- **Target distribution**:
|
- **Target distribution**:
|
||||||
- RTX 3090 (gpu-dense, 24GB, 74.9 tok/s) → Heavy reasoning, code gen, long conversations (slowest per-token but largest context capacity). Weight: 0.55 (LiteLLM).
|
- RTX 3090 (gpu-dense, 24GB, 74.9 tok/s) → Heavy reasoning, code gen, long conversations (slowest per-token but largest context capacity). Weight: 0.55 (LiteLLM).
|
||||||
- RTX 5070 (gpu-light, 12GB, 169.6 tok/s) → Vision/image, web search, lightweight tasks (2.3x faster than 3090 per token). Weight: 0.15 (LiteLLM).
|
- RTX 5070 (gpu-light, 12GB, ~145 tok/s) → Vision (image+text), web search, lightweight tasks. Weight: 0.15 (LiteLLM).
|
||||||
- Strix Halo (strix-moe, 64GB, 62.9 tok/s) → Context compression, summarization, long docs (MoE model). Weight: 0.30 (LiteLLM).
|
- Strix Halo (strix-moe, 64GB, 62.9 tok/s) → Context compression, summarization, long docs (MoE model). Weight: 0.30 (LiteLLM).
|
||||||
- **Note**: RTX 5070 is the fastest endpoint per token. Route high-volume, low-complexity work there first.
|
- **Note**: RTX 5070 is the fastest endpoint per token. Route high-volume, low-complexity work there first.
|
||||||
- **Fix**:
|
- **Fix**:
|
||||||
@@ -166,6 +168,7 @@ Key notes:
|
|||||||
- **Verify**: Each GPU's request pattern matches its designated role within 24h
|
- **Verify**: Each GPU's request pattern matches its designated role within 24h
|
||||||
- **Escalate**: If role mismatch persists >48h → agent alias audit needed
|
- **Escalate**: If role mismatch persists >48h → agent alias audit needed
|
||||||
|
|
||||||
|
---
|
||||||
---
|
---
|
||||||
|
|
||||||
## Execution
|
## Execution
|
||||||
@@ -247,6 +250,7 @@ call update-gpu-health
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
|
---
|
||||||
---
|
---
|
||||||
|
|
||||||
## Reporting
|
## Reporting
|
||||||
@@ -263,6 +267,7 @@ Pushed to `SyslogSolution/health-logs/gpu/{run_id}.json` — versioned, searchab
|
|||||||
- Per-GPU tok/s trend over 7 days
|
- Per-GPU tok/s trend over 7 days
|
||||||
- Regression alerts if any GPU degrades >10% week-over-week
|
- Regression alerts if any GPU degrades >10% week-over-week
|
||||||
|
|
||||||
|
---
|
||||||
---
|
---
|
||||||
|
|
||||||
## Design Decisions (Verified 2026-07-12, Reaffirmed 2026-07-18)
|
## Design Decisions (Verified 2026-07-12, Reaffirmed 2026-07-18)
|
||||||
|
|||||||
@@ -24,8 +24,6 @@ done
|
|||||||
|
|
||||||
| Agent | CT | Node | IP | LiteLLM Alias | Key Source | Platform |
|
| Agent | CT | Node | IP | LiteLLM Alias | Key Source | Platform |
|
||||||
|-------|-----|------|-----|---------------|------------|----------|
|
|-------|-----|------|-----|---------------|------------|----------|
|
||||||
| Tanko | 112 | amdpve | .122 | `tanko` | Infisical vault | Hermes |
|
|
||||||
| Mumuni | 100 | minipve | .24 | `mumuni` | Infisical vault | Hermes |
|
|
||||||
| Koby | 111 | amdpve | .129 | `koby` | Infisical vault | **Hermes** |
|
| Koby | 111 | amdpve | .129 | `koby` | Infisical vault | **Hermes** |
|
||||||
| Koonimo | 113 | amdpve | .114 | `koonimo` | Infisical vault | Hermes |
|
| Koonimo | 113 | amdpve | .114 | `koonimo` | Infisical vault | Hermes |
|
||||||
| Shumba | — | 192.168.68.119 | N/A | N/A (DeepSeek) | Hermes (RETIRED — CT119 now Infisical vault) |
|
| Shumba | — | 192.168.68.119 | N/A | N/A (DeepSeek) | Hermes (RETIRED — CT119 now Infisical vault) |
|
||||||
@@ -53,9 +51,10 @@ Agent (systemd) → LITELLM_API_KEY → LiteLLM (:116/v1) → Router (:9000) →
|
|||||||
|
|
||||||
## Config Pattern — Mandatory Fields
|
## Config Pattern — Mandatory Fields
|
||||||
|
|
||||||
### For Hermes Agents (Tanko, Mumuni, Koonimo)
|
### For Hermes Agents (Mumuni, Koonimo)
|
||||||
|
|
||||||
Every agent's `/root/.hermes/config.yaml` (or `/home/jerome/.hermes/config.yaml`) MUST have:
|
Every Hermes agent's `/root/.hermes/config.yaml` (or `/home/jerome/.hermes/config.yaml`) MUST have:
|
||||||
|
(Tanko is excluded — migrated to DSH/DeepSeek Harness on 2026-08-27, no longer uses Hermes config.)
|
||||||
|
|
||||||
### 1. Main Model
|
### 1. Main Model
|
||||||
```yaml
|
```yaml
|
||||||
@@ -71,7 +70,7 @@ model:
|
|||||||
custom_providers:
|
custom_providers:
|
||||||
- name: harness
|
- name: harness
|
||||||
model: syslog-auto
|
model: syslog-auto
|
||||||
base_url: http://192.168.68.116/v1
|
base_url: http://192.168.68.116/litellm/v1
|
||||||
api_key_env: LITELLM_API_KEY
|
api_key_env: LITELLM_API_KEY
|
||||||
api_mode: chat_completions
|
api_mode: chat_completions
|
||||||
```
|
```
|
||||||
@@ -82,7 +81,7 @@ auxiliary:
|
|||||||
vision:
|
vision:
|
||||||
provider: harness
|
provider: harness
|
||||||
model: gemma-4-12b # or syslog-auto
|
model: gemma-4-12b # or syslog-auto
|
||||||
base_url: http://192.168.68.116/v1
|
base_url: http://192.168.68.116/litellm/v1
|
||||||
api_key_env: LITELLM_API_KEY
|
api_key_env: LITELLM_API_KEY
|
||||||
api_key: <value from: infisical secrets get LITELLM_API_KEY --project=agents --env=production> # ← MANDATORY workaround
|
api_key: <value from: infisical secrets get LITELLM_API_KEY --project=agents --env=production> # ← MANDATORY workaround
|
||||||
timeout: 60
|
timeout: 60
|
||||||
@@ -97,7 +96,7 @@ auxiliary:
|
|||||||
target_ratio: 0.3
|
target_ratio: 0.3
|
||||||
provider: harness
|
provider: harness
|
||||||
model: syslog-auto # or gemma-4-12b
|
model: syslog-auto # or gemma-4-12b
|
||||||
base_url: http://192.168.68.116/v1
|
base_url: http://192.168.68.116/litellm/v1
|
||||||
api_key_env: LITELLM_API_KEY
|
api_key_env: LITELLM_API_KEY
|
||||||
api_key: <value from: infisical secrets get LITELLM_API_KEY --project=agents --env=production> # ← MANDATORY workaround
|
api_key: <value from: infisical secrets get LITELLM_API_KEY --project=agents --env=production> # ← MANDATORY workaround
|
||||||
timeout: 120
|
timeout: 120
|
||||||
@@ -169,11 +168,15 @@ pct-run <CT> grep -A8 "vision:" /root/.hermes/config.yaml | grep api_key
|
|||||||
# Must show both api_key: sk-... and api_key_env: LITELLM_API_KEY
|
# Must show both api_key: sk-... and api_key_env: LITELLM_API_KEY
|
||||||
```
|
```
|
||||||
|
|
||||||
### For Koby (CT 111 / tdunna)
|
### For Koby (CT 111 / tdunna) — **REPORT-ONLY MODE**
|
||||||
|
|
||||||
Koby runs Hermes on CT 111 (tdunna). Config files at `/root/.hermes/config.yaml`.
|
Koby runs Hermes on CT 111 (tdunna). Config files at `/root/.hermes/config.yaml`.
|
||||||
Same Hermes pattern as Tanko/Mumuni/Koonimo — see config sections above.
|
Same Hermes pattern as Tanko/Mumuni/Koonimo — see config sections above.
|
||||||
|
|
||||||
|
**⛔ KOBY IS NEVER REPAIRED (2026-08-17, Captain)**: Diagnostic only — detect and report, never fix on .129.
|
||||||
|
No heal step, no restart, no key rotation, no config edit, no memory rewrite, no disk GC, no service touch, no process kill — ever.
|
||||||
|
If a health check shows Koby degraded, **DO NOT** execute any repair action. Instead, report to Zulip and let Theo fix it.
|
||||||
|
|
||||||
**LiteLLM key**: alias `koby` in LiteLLM DB, injected via `infisical run --` wrapper.
|
**LiteLLM key**: alias `koby` in LiteLLM DB, injected via `infisical run --` wrapper.
|
||||||
|
|
||||||
### For pi Agents (Abiba)
|
### For pi Agents (Abiba)
|
||||||
|
|||||||
@@ -15,7 +15,7 @@ description: >
|
|||||||
|
|
||||||
- template_version: "2.1.0"
|
- template_version: "2.1.0"
|
||||||
- last_applied: timestamp
|
- last_applied: timestamp
|
||||||
- agents_configured: ["tanko", "mumuni", "abiba", "koby", "koonimo", "kagenz0"]
|
- agents_configured: ["mumuni", "abiba", "koby", "koonimo", "kagenz0"] # tanko removed 2026-08-27 (now on DSH/DeepSeek Harness)
|
||||||
- agent_keys: map (see Agent Keys section)
|
- agent_keys: map (see Agent Keys section)
|
||||||
- infra_endpoints_verified: array
|
- infra_endpoints_verified: array
|
||||||
|
|
||||||
@@ -30,8 +30,6 @@ Sub-agent profiles inherit auth from the main config — no separate keys needed
|
|||||||
|
|
||||||
| Agent | Key Alias | Host | SSH | Sub-Agents |
|
| Agent | Key Alias | Host | SSH | Sub-Agents |
|
||||||
|-------|-----------|------|-----|-----------|
|
|-------|-----------|------|-----|-----------|
|
||||||
| Tanko | `tanko-*` | 192.168.68.122 | jerome@.122 | — |
|
|
||||||
| Mumuni | `mumuni` | 192.168.68.24 | root@.24 | 6 profiles ✱ |
|
|
||||||
| Abiba | `abiba-pi` | 192.168.68.24 | local | — |
|
| Abiba | `abiba-pi` | 192.168.68.24 | local | — |
|
||||||
| Koby | `koby` | CT 111 (tdunna) | Zulip | — |
|
| Koby | `koby` | CT 111 (tdunna) | Zulip | — |
|
||||||
| Koonimo | `koonimo` | CT 113 (baggy) | SSH root | — |
|
| Koonimo | `koonimo` | CT 113 (baggy) | SSH root | — |
|
||||||
@@ -348,6 +346,17 @@ curl -s -o /dev/null -w 'key_health: %{http_code}\n' -H "Authorization: Bearer $
|
|||||||
|
|
||||||
### Rule 13: API Key Injection — Two Patterns (UPDATED 2026-07-16, WAL #1300)
|
### Rule 13: API Key Injection — Two Patterns (UPDATED 2026-07-16, WAL #1300)
|
||||||
|
|
||||||
|
### Rule 14: Hermes Context Detection Uses `max_model_tokens`, NOT `max_input_tokens`
|
||||||
|
|
||||||
|
**CRITICAL**: Hermes context detection reads `max_model_tokens` (128K), NOT `max_input_tokens` (64K cap).
|
||||||
|
|
||||||
|
- **Abiba and Hermes agents**: `max_model_tokens: 131072` (128K) — unlimited context
|
||||||
|
- **Crewmates (ops, tune, verify, auth-keys, build)**: `max_input_tokens: 64000` (64K) — capped
|
||||||
|
- If you see `max_input_tokens: 64000` in an Abiba/Hermes config, that's a mistake
|
||||||
|
- Using `max_input_tokens` for Hermes agents causes premature context loss
|
||||||
|
- Check: `grep -n 'max_model_tokens\|max_input_tokens' ~/.hermes/config.yaml`
|
||||||
|
- Expected output: `max_model_tokens: 131072` (not max_input_tokens)
|
||||||
|
|
||||||
Agents inject `LITELLM_API_KEY` via ONE of two mechanisms. Both are valid; the contract
|
Agents inject `LITELLM_API_KEY` via ONE of two mechanisms. Both are valid; the contract
|
||||||
requirement is that the key is a **valid LiteLLM virtual key** (HTTP 200 on /v1/models).
|
requirement is that the key is a **valid LiteLLM virtual key** (HTTP 200 on /v1/models).
|
||||||
|
|
||||||
@@ -424,13 +433,3 @@ curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $K" http://192.
|
|||||||
5. **Set model choice** — Per agent's workload
|
5. **Set model choice** — Per agent's workload
|
||||||
6. **Verify** — curl all shared endpoints, test the model with the new key
|
6. **Verify** — curl all shared endpoints, test the model with the new key
|
||||||
7. **Report** — What was changed, preserved, custom
|
7. **Report** — What was changed, preserved, custom
|
||||||
### Rule 16: Koby Configuration (DeepSeek-primary)
|
|
||||||
(Ref: See Rule 10 for default model behavior, with Koby exception)
|
|
||||||
|
|
||||||
Koby uses a split-model architecture:
|
|
||||||
|
|
||||||
- Primary Model: `deepseek-v4-flash` via `api.deepseek.com` (for reasoning)
|
|
||||||
- Auxiliary Models: `gpu-light` (vision/web_extract) and `syslog-auto` (compression)
|
|
||||||
- Key Hygiene: `api_key_env` is strictly `LITELLM_API_KEY` or `DEEPSEEK_API_KEY`
|
|
||||||
- Constraint: Do NOT touch Koby's primary model/provider/compression settings unless explicitly ruled by the captain.
|
|
||||||
|
|
||||||
|
|||||||
@@ -190,7 +190,7 @@ litellm_settings:
|
|||||||
| Agent | CT | IP | LiteLLM Alias | Key Source | Status | Gateway Wrapper | Last Verified |
|
| Agent | CT | IP | LiteLLM Alias | Key Source | Status | Gateway Wrapper | Last Verified |
|
||||||
|-------|-----|-----|---------------|------------|--------|-----------------|---------------|
|
|-------|-----|-----|---------------|------------|--------|-----------------|---------------|
|
||||||
| Tanko | 112 | .122 | `tanko` | Infisical vault | ✅ Fixed | `infisical run` | 20:17 UTC Jul 5 |
|
| Tanko | 112 | .122 | `tanko` | Infisical vault | ✅ Fixed | `infisical run` | 20:17 UTC Jul 5 |
|
||||||
| Mumuni | 100 (abiba) | .24 | `mumuni` | Infisical vault | ✅ Fixed | Pi Hermes gateway | 2026-07-27 |
|
| Mumuni | 105 (kagentz) | .14 | `mumuni` | Infisical vault | ✅ Fixed | systemd Hermes gateway | 2026-08-29 |
|
||||||
| Koby | 111 | .129 | `koby` | Infisical vault | ✅ Fixed (DeepSeek-primary) | `infisical run` | 23:30 UTC Jul 5 |
|
| Koby | 111 | .129 | `koby` | Infisical vault | ✅ Fixed (DeepSeek-primary) | `infisical run` | 23:30 UTC Jul 5 |
|
||||||
| Koonimo | 113 | .114 | `koonimo` | Infisical vault | ✅ Fixed | `infisical run` (migrated 2026-07-11) | 2026-08-09 |
|
| Koonimo | 113 | .114 | `koonimo` | Infisical vault | ✅ Fixed | `infisical run` (migrated 2026-07-11) | 2026-08-09 |
|
||||||
| Abiba | 100 | .65 | `abiba-pi` | Infisical vault | ✅ N/A (pi native) | — | 19:44 UTC Jul 5 |
|
| Abiba | 100 | .65 | `abiba-pi` | Infisical vault | ✅ N/A (pi native) | — | 19:44 UTC Jul 5 |
|
||||||
@@ -285,13 +285,13 @@ auxiliary:
|
|||||||
vision:
|
vision:
|
||||||
api_key: sk-<agent-key-from-vault> # ← workaround (get via: infisical secrets get LITELLM_API_KEY --project=agents --env=production --plain)
|
api_key: sk-<agent-key-from-vault> # ← workaround (get via: infisical secrets get LITELLM_API_KEY --project=agents --env=production --plain)
|
||||||
api_key_env: LITELLM_API_KEY
|
api_key_env: LITELLM_API_KEY
|
||||||
base_url: http://192.168.68.116/v1
|
base_url: http://192.168.68.116/litellm/v1
|
||||||
model: gemma-4-12b
|
model: gemma-4-12b
|
||||||
provider: harness
|
provider: harness
|
||||||
compression:
|
compression:
|
||||||
api_key: sk-<agent-key-from-vault> # ← workaround (same as above)
|
api_key: sk-<agent-key-from-vault> # ← workaround (same as above)
|
||||||
api_key_env: LITELLM_API_KEY
|
api_key_env: LITELLM_API_KEY
|
||||||
base_url: http://192.168.68.116/v1
|
base_url: http://192.168.68.116/litellm/v1
|
||||||
model: gemma-4-12b
|
model: gemma-4-12b
|
||||||
provider: harness
|
provider: harness
|
||||||
```
|
```
|
||||||
|
|||||||
@@ -28,7 +28,7 @@ connectivity recovery including end-to-end DM validation.
|
|||||||
|
|
||||||
| Param | Type | Required | Default | Description |
|
| Param | Type | Required | Default | Description |
|
||||||
|-------|------|----------|---------|-------------|
|
|-------|------|----------|---------|-------------|
|
||||||
| `target` | string | yes | — | Agent name: `mumuni`, `tanko`, `koby`, or `shumba` |
|
| `target` | string | yes | — | Agent name: `mumuni`, `koby`, or `shumba` (Tanko excluded — on DSH since 2026-08-27, no Hermes plugin) |
|
||||||
| `branch` | string | no | `master` | Git branch to pull (overridable for pinning) |
|
| `branch` | string | no | `master` | Git branch to pull (overridable for pinning) |
|
||||||
|
|
||||||
## Maintains
|
## Maintains
|
||||||
@@ -55,8 +55,7 @@ connectivity recovery including end-to-end DM validation.
|
|||||||
|
|
||||||
| Host | CT | Proxmox | IP (direct) | Hermes Home | User |
|
| Host | CT | Proxmox | IP (direct) | Hermes Home | User |
|
||||||
|------|-----|---------|-------------|-------------|------|
|
|------|-----|---------|-------------|-------------|------|
|
||||||
| Mumuni | CT100 | — | 192.168.68.24 | /root/.hermes | root |
|
| Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome | *(DSH since 2026-08-27 — historical, plugin retired on this host)* |
|
||||||
| Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome |
|
|
||||||
| Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root |
|
| Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root |
|
||||||
| Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky |
|
| Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky |
|
||||||
|
|
||||||
@@ -121,7 +120,8 @@ cp plugins/platforms/zulip/adapter.py \
|
|||||||
plugins/platforms/zulip/plugin.yaml \
|
plugins/platforms/zulip/plugin.yaml \
|
||||||
{{hermes_home}}/hermes-agent/plugins/platforms/zulip/
|
{{hermes_home}}/hermes-agent/plugins/platforms/zulip/
|
||||||
|
|
||||||
# Fix ownership (Tanko only — runs as jerome user)
|
# Fix ownership (was Tanko-only, runs as jerome user)
|
||||||
|
# RETIRED 2026-08-27: tanko no longer uses the Hermes Zulip plugin (DSH).
|
||||||
[ "{{target}}" = "tanko" ] && chown -R jerome:jerome \
|
[ "{{target}}" = "tanko" ] && chown -R jerome:jerome \
|
||||||
{{hermes_home}}/hermes-agent/plugins/platforms/zulip/
|
{{hermes_home}}/hermes-agent/plugins/platforms/zulip/
|
||||||
|
|
||||||
|
|||||||
@@ -1,9 +1,9 @@
|
|||||||
---
|
---
|
||||||
|
report_only_agents:
|
||||||
|
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
|
||||||
kind: function
|
kind: function
|
||||||
name: hermes-zulip-restore
|
name: hermes-zulip-restore
|
||||||
description: >
|
description: >
|
||||||
Restores Zulip connectivity for any Hermes agent (Mumuni CT100, Tanko CT112,
|
|
||||||
Koby CT111, Shumba on Lucky's mini PC). Deploys the zulip-platform adapter to the correct bundled plugin
|
|
||||||
path, verifies env credentials, restarts the gateway, and confirms Zulip
|
path, verifies env credentials, restarts the gateway, and confirms Zulip
|
||||||
connects. Run this whenever a Hermes agent stops responding on Zulip or after
|
connects. Run this whenever a Hermes agent stops responding on Zulip or after
|
||||||
a fresh agent deployment.
|
a fresh agent deployment.
|
||||||
@@ -12,6 +12,7 @@ version: 1.0.0
|
|||||||
status: active
|
status: active
|
||||||
runtime_contract: 2
|
runtime_contract: 2
|
||||||
---
|
---
|
||||||
|
---
|
||||||
|
|
||||||
# Hermes Zulip Restore — Bring Any Agent Back to Good State
|
# Hermes Zulip Restore — Bring Any Agent Back to Good State
|
||||||
|
|
||||||
@@ -23,7 +24,7 @@ gateway restart, and connection validation.
|
|||||||
|
|
||||||
| Param | Type | Required | Default | Description |
|
| Param | Type | Required | Default | Description |
|
||||||
|-------|------|----------|---------|-------------|
|
|-------|------|----------|---------|-------------|
|
||||||
| `target` | string | yes | — | Agent name: `mumuni`, `tanko`, `koby`, or `shumba` |
|
| `target` | string | yes | — | Agent name: `mumuni`, `koby`, or `shumba` (Tanko excluded — DSH since 2026-08-27) |
|
||||||
|
|
||||||
## Maintains
|
## Maintains
|
||||||
|
|
||||||
@@ -36,7 +37,7 @@ gateway restart, and connection validation.
|
|||||||
|
|
||||||
- `_strip_html` function present in `<hermes-agent>/plugins/platforms/zulip/adapter.py`
|
- `_strip_html` function present in `<hermes-agent>/plugins/platforms/zulip/adapter.py`
|
||||||
- All three adapter files (__init__.py, adapter.py, plugin.yaml) present at bundled path
|
- All three adapter files (__init__.py, adapter.py, plugin.yaml) present at bundled path
|
||||||
- Zulip env vars set in `~/.hermes/.env` (or `/home/jerome/.hermes/.env` for Tanko)
|
- Zulip env vars set in `~/.hermes/.env` (or `/home/jerome/.hermes/.env` for Tanko, historical — DSH since 2026-08-27)
|
||||||
- Gateway restarted and zulip platform reports state `connected`
|
- Gateway restarted and zulip platform reports state `connected`
|
||||||
- HTML stripping enabled for `/approve` and `/deny` slash command support
|
- HTML stripping enabled for `/approve` and `/deny` slash command support
|
||||||
|
|
||||||
@@ -51,8 +52,6 @@ gateway restart, and connection validation.
|
|||||||
|
|
||||||
| Host | CT | Proxmox | IP (direct) | Hermes Home | User |
|
| Host | CT | Proxmox | IP (direct) | Hermes Home | User |
|
||||||
|------|-----|---------|-------------|-------------|------|
|
|------|-----|---------|-------------|-------------|------|
|
||||||
| Mumuni | CT100 (abiba) | minipve | 192.168.68.24 | /root/.hermes | root |
|
|
||||||
| Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome |
|
|
||||||
| Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root |
|
| Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root |
|
||||||
| Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky |
|
| Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky |
|
||||||
|
|
||||||
@@ -94,8 +93,8 @@ cp zulip-platform-plugins/plugins/platforms/zulip/adapter.py \
|
|||||||
zulip-platform-plugins/plugins/platforms/zulip/plugin.yaml \
|
zulip-platform-plugins/plugins/platforms/zulip/plugin.yaml \
|
||||||
<HERMES_HOME>/hermes-agent/plugins/platforms/zulip/
|
<HERMES_HOME>/hermes-agent/plugins/platforms/zulip/
|
||||||
|
|
||||||
# Fix ownership (Tanko only)
|
# Fix ownership (was Tanko-only; RETIRED 2026-08-27 — tanko on DSH, no Hermes plugin)
|
||||||
chown -R jerome:jerome <HERMES_HOME>/hermes-agent/plugins/platforms/zulip/ # Tanko only
|
chown -R jerome:jerome <HERMES_HOME>/hermes-agent/plugins/platforms/zulip/ # Tanko only (historical)
|
||||||
|
|
||||||
# Clean up
|
# Clean up
|
||||||
rm -rf /tmp/zulip-deploy
|
rm -rf /tmp/zulip-deploy
|
||||||
@@ -184,6 +183,7 @@ https://git.sysloggh.net/SyslogSolution/zulip-platform-plugins/src/branch/feat/z
|
|||||||
Commit `55ca15d` — `fix(zulip): add _strip_html for slash command matching`
|
Commit `55ca15d` — `fix(zulip): add _strip_html for slash command matching`
|
||||||
Pull request #33 is the primary integration branch.
|
Pull request #33 is the primary integration branch.
|
||||||
|
|
||||||
|
---
|
||||||
---
|
---
|
||||||
|
|
||||||
**Last verified good state**: 2026-07-08 — Mumuni, Tanko, Koby all connected with `_strip_html` applied.
|
**Last verified good state**: 2026-07-08 — Mumuni, Tanko, Koby all connected with `_strip_html` applied.
|
||||||
|
|||||||
@@ -87,8 +87,8 @@ call apply-liteLLM-routing
|
|||||||
|
|
||||||
call apply-agent-compression
|
call apply-agent-compression
|
||||||
agent: mumuni
|
agent: mumuni
|
||||||
host: 192.168.68.24
|
host: 192.168.68.14
|
||||||
config_path: /root/.hermes/config.yaml
|
config_path: /home/hermes/.hermes/config.yaml
|
||||||
|
|
||||||
-- Phase 4: Enable llama.cpp prompt caching on GPU hosts
|
-- Phase 4: Enable llama.cpp prompt caching on GPU hosts
|
||||||
|
|
||||||
|
|||||||
@@ -613,7 +613,7 @@ ssh root@192.168.68.110 "systemctl restart llama-server"
|
|||||||
| 109 | docker-vm | storepve | .7 | Docker host | ❌ |
|
| 109 | docker-vm | storepve | .7 | Docker host | ❌ |
|
||||||
| 110 | gitea | minipve | **.17** | Git | ❌ |
|
| 110 | gitea | minipve | **.17** | Git | ❌ |
|
||||||
| 111 | tdunna | amdpve | .129 | Hermes agent | ✅ |
|
| 111 | tdunna | amdpve | .129 | Hermes agent | ✅ |
|
||||||
| 112 | tanko | amdpve | .122 | Hermes agent | ✅ |
|
| 112 | tanko | amdpve | .122 | DSH (DeepSeek Harness) agent | ✅ |
|
||||||
| 113 | baggy | amdpve | .114 | Hermes agent | ✅ |
|
| 113 | baggy | amdpve | .114 | Hermes agent | ✅ |
|
||||||
| 115 | scottdenya | amdpve | .75 | Denya OneCare | ❌ |
|
| 115 | scottdenya | amdpve | .75 | Denya OneCare | ❌ |
|
||||||
| 116 | syslog-api | minipve | .116 | LiteLLM + Grafana | ❌ |
|
| 116 | syslog-api | minipve | .116 | LiteLLM + Grafana | ❌ |
|
||||||
|
|||||||
@@ -140,7 +140,6 @@ After ALL updates (apt + images + restarts), verify every critical service is ba
|
|||||||
| Zulip | `curl -sf https://chat.sysloggh.net/api/v1/server_settings` | 200 OK |
|
| Zulip | `curl -sf https://chat.sysloggh.net/api/v1/server_settings` | 200 OK |
|
||||||
| Gitea | `curl -sf https://git.sysloggh.net/api/v1/version` | 200 OK |
|
| Gitea | `curl -sf https://git.sysloggh.net/api/v1/version` | 200 OK |
|
||||||
| PM2 processes | `pm2 jlist` (CT 100) | all pi-agent processes `online` |
|
| PM2 processes | `pm2 jlist` (CT 100) | all pi-agent processes `online` |
|
||||||
| Hermes gateways | SSH to Mumuni CT 100, Tanko CT 112; `systemctl is-active hermes-gateway` | `active` for each |
|
|
||||||
|
|
||||||
Regression check: every service that was GREEN in `health-baseline` must still be GREEN. A service that was already RED (and caused a preflight abort) is excluded — but Phase 0 should have aborted before we got here.
|
Regression check: every service that was GREEN in `health-baseline` must still be GREEN. A service that was already RED (and caused a preflight abort) is excluded — but Phase 0 should have aborted before we got here.
|
||||||
|
|
||||||
|
|||||||
@@ -102,6 +102,39 @@ GPU .8 (RTX 3090) GPU .110 (RTX 5070) GPU .15 (Strix Halo)
|
|||||||
- Stack persists across reboots (systemd for exporters, Docker restart policy)
|
- Stack persists across reboots (systemd for exporters, Docker restart policy)
|
||||||
|
|
||||||
## Execution
|
## Execution
|
||||||
|
### check-health
|
||||||
|
|
||||||
|
**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Zulip API health (POST ping)
|
||||||
|
curl -s -o /dev/null -w '%{http_code}' -X POST https://chat.sysloggh.net/api/v1/messages -u 'abiba-bot@chat.sysloggh.net:KEY'
|
||||||
|
# Expected: 200 (HTTP 000 = unreachable/cache)
|
||||||
|
|
||||||
|
# PM2 process health
|
||||||
|
pm2 jlist
|
||||||
|
# Expected: 5/5 online (abiba-telegram, abiba-zulip, zulip-watchdog, gitea-runner, spoton-service)
|
||||||
|
|
||||||
|
# GPU exporters (may be down per DEPLOYMENT STATUS)
|
||||||
|
curl -s http://192.168.68.8:9400/metrics && echo " - OK" || echo " - FAIL"
|
||||||
|
curl -s http://192.168.68.110:9400/metrics && echo " - OK" || echo " - FAIL"
|
||||||
|
curl -s http://192.168.68.15:9400/metrics && echo " - OK" || echo " - FAIL"
|
||||||
|
|
||||||
|
# Prometheus targets
|
||||||
|
curl -s http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets'
|
||||||
|
# Expected: All targets UP (may show some down if exporters not deployed)
|
||||||
|
|
||||||
|
# Grafana health
|
||||||
|
curl -s http://192.168.68.116:3001/api/health | jq '{status, version}'
|
||||||
|
# Expected: {"status":"ok","version":"..."}
|
||||||
|
|
||||||
|
# LiteLLM metrics
|
||||||
|
curl -s http://192.168.68.116:4001/metrics | head -20
|
||||||
|
# Expected: Prometheus-formatted metrics output
|
||||||
|
```
|
||||||
|
|
||||||
|
**Report format**: Summarize actual results from each probe. If any probe returns non-200 or empty output, flag as alert.
|
||||||
|
|
||||||
|
|
||||||
### Phase 1: GPU Exporters
|
### Phase 1: GPU Exporters
|
||||||
|
|
||||||
|
|||||||
@@ -59,7 +59,7 @@ Before ANY update wave:
|
|||||||
| CT 100 (.24) | Abiba (pi) | `apt update && apt upgrade -y` | 3 min |
|
| CT 100 (.24) | Abiba (pi) | `apt update && apt upgrade -y` | 3 min |
|
||||||
| CT 116 (.116) | syslog-api (LiteLLM host) | `apt update && apt upgrade -y` | 3 min |
|
| CT 116 (.116) | syslog-api (LiteLLM host) | `apt update && apt upgrade -y` | 3 min |
|
||||||
| CT 112 (tanko, amdpve) | Tanko | `apt update && apt upgrade -y` | 3 min |
|
| CT 112 (tanko, amdpve) | Tanko | `apt update && apt upgrade -y` | 3 min |
|
||||||
| CT 100 (mumuni/abiba, minipve) | Mumuni | `apt update && apt upgrade -y` | 3 min |
|
| CT 105 (kagentz, minipve) | Mumuni | `apt update && apt upgrade -y` | 3 min |
|
||||||
| VM 101 (.8) | llm-gpu (RTX 3090) | `apt update && apt upgrade -y` | 3 min |
|
| VM 101 (.8) | llm-gpu (RTX 3090) | `apt update && apt upgrade -y` | 3 min |
|
||||||
| VM 103 (.110) | ocu-llm (RTX 5070) | `apt update && apt upgrade -y` | 3 min |
|
| VM 103 (.110) | ocu-llm (RTX 5070) | `apt update && apt upgrade -y` | 3 min |
|
||||||
|
|
||||||
@@ -162,9 +162,9 @@ Before Wave 1, snapshot these files:
|
|||||||
/etc/systemd/system/strix-server.service (amdpve .15 — strix-moe)
|
/etc/systemd/system/strix-server.service (amdpve .15 — strix-moe)
|
||||||
/etc/systemd/system/llama-server.service (VM 101 .8, VM 103 .110)
|
/etc/systemd/system/llama-server.service (VM 101 .8, VM 103 .110)
|
||||||
# Hermes agent configs (key enforcement — 2026-07-10)
|
# Hermes agent configs (key enforcement — 2026-07-10)
|
||||||
/root/.hermes/config.yaml (Mumuni inside CT 100, Tanko CT 112, etc.)
|
/home/hermes/.hermes/config.yaml (Mumuni kagentz CT105; Tanko CT112 uses /home/jerome/.hermes)
|
||||||
/root/.config/systemd/user/hermes-gateway.service (Mumuni inside CT 100 — EnvironmentFile fixed)
|
/etc/systemd/system/hermes-gateway.service (Mumuni kagentz CT105 — system unit, User=hermes)
|
||||||
/etc/environment (Mumuni inside CT 100 — LITELLM_API_KEY)
|
/etc/environment (LITELLM_API_KEY — legacy path, Mumuni now keys via Infisical)
|
||||||
```
|
```
|
||||||
|
|
||||||
## MCP Gateway (2026-07-10)
|
## MCP Gateway (2026-07-10)
|
||||||
|
|||||||
@@ -178,7 +178,7 @@ through its agent wrapper.
|
|||||||
| Agent | Host | Pattern | Keys | Status |
|
| Agent | Host | Pattern | Keys | Status |
|
||||||
|-------|------|---------|------|--------|
|
|-------|------|---------|------|--------|
|
||||||
| abiba | .24 | pi agent wrapper | ABIBA_LITELLM_API_KEY + ABIBA_ZULIP_API_KEY | ✅ vault-backed |
|
| abiba | .24 | pi agent wrapper | ABIBA_LITELLM_API_KEY + ABIBA_ZULIP_API_KEY | ✅ vault-backed |
|
||||||
| mumuni | .24 (CT100 abiba) | Pi Hermes gateway (no systemd) | MUMUNI_LITELLM_API_KEY + MUMUNI_ZULIP_API_KEY | ✅ vault-backed + .env fallback |
|
| mumuni | .14 (kagentz CT105) | systemd unit hermes-gateway.service (user hermes) | MUMUNI_LITELLM_API_KEY + MUMUNI_ZULIP_API_KEY | ✅ vault-backed + .env fallback |
|
||||||
| tanko | .122 | systemd drop-in + while-true wrapper + st.8e848433 (user jerome) | TANKO_LITELLM_API_KEY + TANKO_ZULIP_API_KEY | ✅ vault-backed + .env fallback |
|
| tanko | .122 | systemd drop-in + while-true wrapper + st.8e848433 (user jerome) | TANKO_LITELLM_API_KEY + TANKO_ZULIP_API_KEY | ✅ vault-backed + .env fallback |
|
||||||
| koby | .129 | systemd drop-in + while-true wrapper + st.8e848433 | KOBY_LITELLM_API_KEY, shares TANKO_ZULIP_API_KEY (tanko-bot) | ✅ vault-backed |
|
| koby | .129 | systemd drop-in + while-true wrapper + st.8e848433 | KOBY_LITELLM_API_KEY, shares TANKO_ZULIP_API_KEY (tanko-bot) | ✅ vault-backed |
|
||||||
| koonimo | .114 | systemd drop-in + while-true wrapper + st.8e848433 | KOONIMO_LITELLM_API_KEY + KOONIMO_ZULIP_API_KEY | ✅ vault-backed |
|
| koonimo | .114 | systemd drop-in + while-true wrapper + st.8e848433 | KOONIMO_LITELLM_API_KEY + KOONIMO_ZULIP_API_KEY | ✅ vault-backed |
|
||||||
@@ -257,9 +257,51 @@ reads use per-agent identities. This eliminates the single shared token risk.
|
|||||||
| Agent | .env Keys |
|
| Agent | .env Keys |
|
||||||
|-------|-----------|
|
|-------|-----------|
|
||||||
| Mumuni | MUMUNI_LITELLM_API_KEY, MUMUNI_ZULIP_API_KEY |
|
| Mumuni | MUMUNI_LITELLM_API_KEY, MUMUNI_ZULIP_API_KEY |
|
||||||
| Tanko | TANKO_LITELLM_API_KEY, TANKO_ZULIP_API_KEY |
|
|| Tanko | TANKO_LITELLM_API_KEY, TANKO_ZULIP_API_KEY |
|
||||||
| Koby | (wrapper injects from vault — .env has Telegram token) |
|
|| Koby | (wrapper injects from vault — .env has Telegram token) |
|
||||||
| Koonimo | KOONIMO_LITELLM_API_KEY, KOONIMO_ZULIP_API_KEY |
|
|| Koonimo | KOONIMO_LITELLM_API_KEY, KOONIMO_ZULIP_API_KEY |
|
||||||
|
|| Agent Zero (kagentz .14) | OPENROUTER_API_KEY (direct OpenRouter access) |
|
||||||
|
|
||||||
|
### Agent Zero (kagentz .14) — OpenRouter Integration (2026-09-01)
|
||||||
|
|
||||||
|
Agent Zero runs in Docker on kagentz (CT105) and uses **direct OpenRouter API access**,
|
||||||
|
not via the LiteLLM proxy. This is because Agent Zero's workflow (self-update manager,
|
||||||
|
UI bootstrap, model selection) is built around OpenRouter's native authentication.
|
||||||
|
|
||||||
|
**Key Storage:**
|
||||||
|
- **Container**: `/a0/usr/.env` (line ~72: `API_KEY_OPENROUTER=sk-or-v1-…`)
|
||||||
|
- **Vault**: Infisical secret `OPENROUTER_API_KEY` (project=agents, env=production)
|
||||||
|
- **Fallback**: The container's .env is the primary source; vault sync is optional
|
||||||
|
(unlike fleet agents which require vault injection)
|
||||||
|
|
||||||
|
**Current Key (2026-09-01):**
|
||||||
|
- **Prefix**: `sk-or-v1-0af3f3…`
|
||||||
|
- **User**: `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`
|
||||||
|
- **Plan**: Paid (not free tier)
|
||||||
|
- **Usage**: 0 (as of 2026-09-01)
|
||||||
|
|
||||||
|
**Model Configuration:**
|
||||||
|
- **Preset**: "Cost Efficient" (`/a0/usr/plugins/_model_config/presets.yaml`)
|
||||||
|
- **Model**: `openrouter/moonshotai/kimi-k3`
|
||||||
|
- **API Base**: (empty — uses OpenRouter default)
|
||||||
|
|
||||||
|
**Why not LiteLLM proxy?**
|
||||||
|
Agent Zero's architecture was designed before the fleet adopted the LiteLLM proxy
|
||||||
|
standard. The container runs `/exe/self_update_manager.py` and `/a0/run_ui.py` which
|
||||||
|
directly call OpenRouter via Python's requests library. Converting would require:
|
||||||
|
1. Refactoring all LLM calls to use `litellm` library
|
||||||
|
2. Adding vault wrapper injection
|
||||||
|
3. Updating self_update_manager to use proxy-aware key handling
|
||||||
|
|
||||||
|
**Rotation Procedure:**
|
||||||
|
1. Generate new key in OpenRouter UI
|
||||||
|
2. Update container: `sed -i 's/^API_KEY_OPENROUTER=.*/API_KEY_OPENROUTER=<new_key>/' /a0/usr/.env`
|
||||||
|
3. Update vault: `infisical secrets set OPENROUTER_API_KEY=<new_key> --projectId=agents --env=production`
|
||||||
|
4. Restart container: `sudo docker exec agent-zero supervisorctl restart run_ui`
|
||||||
|
5. Verify: `curl -s https://openrouter.ai/api/v1/auth/key -H "Authorization: Bearer <new_key>"`
|
||||||
|
|
||||||
|
**Related Contract:**
|
||||||
|
- `agent-zero-openrouter-key.prose.md` — Full agent-zero key management contract
|
||||||
|
|
||||||
## Key Rotation Log
|
## Key Rotation Log
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,117 @@
|
|||||||
|
---
|
||||||
|
kind: pattern
|
||||||
|
name: litellm-client-timeouts
|
||||||
|
description: >
|
||||||
|
Standard client timeout and retry policy for ALL agents calling LiteLLM
|
||||||
|
(CT 116, http://192.168.68.116). Created 2026-09-08 after the Sep 6 incident:
|
||||||
|
a backend stall 04:00-06:30 EDT produced 157 client-abandoned 408 failures
|
||||||
|
(82% from abiba-pi, 13 from mumuni) against a backend that was actually
|
||||||
|
succeeding at 20-70s per call once clients stopped giving up. Grounded in
|
||||||
|
measured data: syslog-auto (LiteLLM virtual model, no dedicated GPU — routes
|
||||||
|
to backends) averages 28.8s/request with 25.4s TTFT over 888 calls/24h;
|
||||||
|
nginx already allows 600s (Rule 5, verified 2026-08-09); the gap is entirely
|
||||||
|
client-side. Blast radius if wrong: agents fall back to DeepSeek silently
|
||||||
|
(key/timeout failures present as model degradation, not errors) or abandon
|
||||||
|
healthy-but-slow reasoning calls, fragmenting long tasks.
|
||||||
|
---
|
||||||
|
|
||||||
|
## Maintains
|
||||||
|
|
||||||
|
- client_timeout_standard: "litellm-client-timeouts v1.0 (2026-09-08)"
|
||||||
|
- applies_to: ALL agents and scripts calling http://192.168.68.116 (main model path, auxiliary tasks, health probes, benchmark jobs)
|
||||||
|
- verified_against: Prometheus litellm_* metrics 24h window ending 2026-09-08 ~10:00 EDT; live probe (syslog-auto tiny call 0.56s TTFB, 200 OK); nginx 600s proxy_read_timeout (Rule 5)
|
||||||
|
|
||||||
|
## The measured numbers these values come from
|
||||||
|
|
||||||
|
| Model | avg latency | avg TTFT | p-profile (24h) |
|
||||||
|
|---|---|---|---|
|
||||||
|
| syslog-auto | 28.8s | 25.4s | 68 calls took 30-120s; tail to ~300s under load |
|
||||||
|
| qwen3.6-27B-code | 23.0s | — | same backend class as syslog-auto |
|
||||||
|
| strix-moe | 7.5s | — | Strix Halo, healthy |
|
||||||
|
| gemma-4-12b | 2.6s | — | RTX 5070, healthy |
|
||||||
|
|
||||||
|
Sep 6 incident timeline: failures 04:00-07:00 EDT (0% GPU util = wedged
|
||||||
|
backend), full recovery 07:00-08:00 with ZERO client failures once requests
|
||||||
|
tolerated 20-70s — 392 successful slow calls in the three hours after recovery.
|
||||||
|
LiteLLM's internal queue time is ~0s; the latency is model inference, not
|
||||||
|
proxy queuing.
|
||||||
|
|
||||||
|
## Parameters
|
||||||
|
|
||||||
|
### 1. Primary model path (model.default / custom_providers) — NO client timeout below 300s
|
||||||
|
|
||||||
|
- The default Hermes HTTP timeout (~60s) is TOO SHORT for syslog-auto's healthy
|
||||||
|
28.8s average + 120-300s tail. Every 408 in the incident was a client
|
||||||
|
abandoning a request the backend would have answered.
|
||||||
|
- If the transport exposes a timeout setting for the main model, set it to
|
||||||
|
**300s or more**. If it does not (current Hermes custom-provider path has no
|
||||||
|
timeout knob), that is acceptable ONLY because nginx holds the request for
|
||||||
|
600s — but any wrapper, script, or direct API call you write MUST set its own
|
||||||
|
timeout >= 300s for syslog-auto/qwen-class calls.
|
||||||
|
- Never hardcode a shorter timeout "to fail fast" on this path — failing fast
|
||||||
|
here is what caused the incident.
|
||||||
|
|
||||||
|
### 2. Auxiliary tasks — keep template timeouts, one correction
|
||||||
|
|
||||||
|
- vision: 60s (keep), web_extract: 30s (keep) — gemma-4-12b averages 2.6s;
|
||||||
|
these are fine.
|
||||||
|
- compression: 300s (keep — this was already raised from 60 per gpu-fleet).
|
||||||
|
- **gpu-dense delegation/x_search: set timeout >= 120s.** The RTX 3090
|
||||||
|
(qwen3.6-27B-code backend, 23.0s avg) is the same speed class as
|
||||||
|
syslog-auto; delegation defaults that assume fast responses will 408 the
|
||||||
|
same way.
|
||||||
|
|
||||||
|
### 3. Retry policy — backoff, not repetition
|
||||||
|
|
||||||
|
- On timeout (408) or 5xx: retry up to **2 times** with exponential backoff
|
||||||
|
(**15s, 45s**) before giving up.
|
||||||
|
- Do NOT retry in a tight loop. The Sep 6 spike shape (65 failures in one hour
|
||||||
|
from one key) was a batch job retrying without backoff while the backend was
|
||||||
|
down — it multiplied load during recovery.
|
||||||
|
- On 401/403: do NOT retry — that is a key/permission problem (see
|
||||||
|
litellm-api-keys.prose.md and Rule 11); retrying just spams the log.
|
||||||
|
- On 429: honor the retry-after header if present, else back off 60s.
|
||||||
|
|
||||||
|
### 4. Health probes — identify yourself and time out sanely
|
||||||
|
|
||||||
|
- Probes MUST NOT appear as keyless, model-less failures in the metrics (12
|
||||||
|
such orphans appeared in the incident window and cost investigation time).
|
||||||
|
Send a real model name and use a real (probe-designated) key.
|
||||||
|
- Probe timeout: 30s. A probe that takes longer than 30s IS the alert —
|
||||||
|
report "backend slow (>30s)" rather than hanging.
|
||||||
|
- Probe cadence: at most hourly. The 6h litellm-health cron cadence is the
|
||||||
|
standard; sub-hourly synthetic traffic distorts latency baselines.
|
||||||
|
|
||||||
|
### 5. Batch/benchmark jobs — schedule away from 04:00-07:00 EDT and chunk
|
||||||
|
|
||||||
|
- The incident window showed bulk clients amplifying a backend stall 5:1.
|
||||||
|
- Batch jobs that can tolerate delay: schedule 09:00-17:00 EDT.
|
||||||
|
- Any batch loop over N requests MUST sleep >= 5s between requests and honor
|
||||||
|
the retry policy in section 3.
|
||||||
|
|
||||||
|
## Returns
|
||||||
|
|
||||||
|
- A single standard any agent or script can cite: timeouts >= 300s on the
|
||||||
|
syslog-auto path, >= 120s on gpu-dense delegation, backoff retries (2x,
|
||||||
|
15s/45s), identified probes at 30s/hourly, batch jobs chunked and
|
||||||
|
day-scheduled.
|
||||||
|
- Failure signature recognition: bulk 408s from multiple keys in one window =
|
||||||
|
backend event (check gpu_utilization_percent: 0% = wedged, ~100% = saturated);
|
||||||
|
single-key 408s = that client's timeout is too short.
|
||||||
|
- Cross-references: hermes-config-template.prose.md (Rule 5 nginx 600s,
|
||||||
|
auxiliary timeouts), gpu-fleet.prose.md (compression 300s precedent,
|
||||||
|
stable aliases), litellm-api-keys.prose.md (key/permission failures).
|
||||||
|
|
||||||
|
## Intentionally NOT changed
|
||||||
|
|
||||||
|
- No server-side LiteLLM timeout/cooldown changes proposed — the incident
|
||||||
|
self-recovered and the server is healthy (0.56s live probe); changing
|
||||||
|
server behavior without process-level root cause (CT116 requires root;
|
||||||
|
not reachable from kagentz) would be guessing.
|
||||||
|
- No change to the template's vision/web_extract/compression timeouts —
|
||||||
|
measured data says they are correct.
|
||||||
|
- No per-agent key permission changes — those are litellm-api-keys.prose.md
|
||||||
|
territory (and the open gpu-vision/gemma 403 items are already filed with
|
||||||
|
the key owners).
|
||||||
|
- No model routing changes — syslog-auto's weighted pool behaved correctly
|
||||||
|
throughout the incident.
|
||||||
@@ -1,4 +1,6 @@
|
|||||||
---
|
---
|
||||||
|
report_only_agents:
|
||||||
|
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
|
||||||
kind: responsibility
|
kind: responsibility
|
||||||
name: litellm-self-heal
|
name: litellm-self-heal
|
||||||
status: deployed
|
status: deployed
|
||||||
@@ -22,6 +24,7 @@ description: >
|
|||||||
inference, and agent keys. Applies remediation rules for common failures.
|
inference, and agent keys. Applies remediation rules for common failures.
|
||||||
Reports every action via Zulip DM and Gitea (SyslogSolution/health-logs).
|
Reports every action via Zulip DM and Gitea (SyslogSolution/health-logs).
|
||||||
---
|
---
|
||||||
|
---
|
||||||
|
|
||||||
# LiteLLM Operations — Health Check + Self-Heal
|
# LiteLLM Operations — Health Check + Self-Heal
|
||||||
|
|
||||||
@@ -65,7 +68,15 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
|||||||
|
|
||||||
## LiteLLM Model Surface (ground truth — `/opt/inference-harness/litellm_config.yaml` on CT 116)
|
## LiteLLM Model Surface (ground truth — `/opt/inference-harness/litellm_config.yaml` on CT 116)
|
||||||
|
|
||||||
`model_name`s served: `qwen3.6-27B-code`, `gemma-4-12b`, `qwen3.6-35B-udq4`, `strix-moe`, `gpu-dense`, `gpu-light`, `syslog-auto`.
|
`model_name`s served: `qwen3.6-27B-code`, `gemma-4-12b`, `qwen3.6-35B-udq4`, `strix-moe`, `gpu-dense`, `gpu-light`, `syslog-auto`, `crew-auto` (new 2026-08-20).
|
||||||
|
|
||||||
|
### Context Cap Split (2026-08-20)
|
||||||
|
|
||||||
|
- **Abiba (firstmate)**: 128K uncapped — unlimited context for primary workloads
|
||||||
|
- **Hermes agents** (mumuni, tanko, koby, koonimo): 128K uncapped
|
||||||
|
- **Crewmates** (ops, tune, verify, auth-keys, build): 64K capped — alias `crew-auto` enforces 64K limit
|
||||||
|
|
||||||
|
Preferred implementation: uncap shared pool, add capped alias for crew-only.
|
||||||
|
|
||||||
- `syslog-auto` is a weighted router model: qwen3.6-27B-code (0.55, rpm 500) + qwen3.6-35B-udq4 (0.30, rpm 60) + gemma-4-12b (0.15, rpm 200).
|
- `syslog-auto` is a weighted router model: qwen3.6-27B-code (0.55, rpm 500) + qwen3.6-35B-udq4 (0.30, rpm 60) + gemma-4-12b (0.15, rpm 200).
|
||||||
- `gpu-dense` / `gpu-light` are high-rpm aliases (rpm 500) onto qwen3.6-27B-code / gemma-4-12b respectively.
|
- `gpu-dense` / `gpu-light` are high-rpm aliases (rpm 500) onto qwen3.6-27B-code / gemma-4-12b respectively.
|
||||||
@@ -120,6 +131,7 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
|||||||
- Also wakes on user request
|
- Also wakes on user request
|
||||||
- On failure: re-check after 30s, escalate after 3 consecutive failures
|
- On failure: re-check after 30s, escalate after 3 consecutive failures
|
||||||
|
|
||||||
|
---
|
||||||
---
|
---
|
||||||
|
|
||||||
## Health Check
|
## Health Check
|
||||||
@@ -162,6 +174,7 @@ Determine overall_status from individual check results:
|
|||||||
- "degraded" — 1-2 non-critical checks fail
|
- "degraded" — 1-2 non-critical checks fail
|
||||||
- "down" — critical checks fail
|
- "down" — critical checks fail
|
||||||
|
|
||||||
|
---
|
||||||
---
|
---
|
||||||
|
|
||||||
## Remediation Rules
|
## Remediation Rules
|
||||||
@@ -205,6 +218,7 @@ Escalate → if SSH access unavailable, send Zulip DM
|
|||||||
Router no longer in path so Redis active counters are unused. Rule retained
|
Router no longer in path so Redis active counters are unused. Rule retained
|
||||||
for reference but inactive. If Redis issues occur, check harness-redis container.
|
for reference but inactive. If Redis issues occur, check harness-redis container.
|
||||||
|
|
||||||
|
---
|
||||||
---
|
---
|
||||||
|
|
||||||
## Reporting
|
## Reporting
|
||||||
@@ -227,6 +241,7 @@ top actions, uptime.
|
|||||||
If a fix requires another agent (e.g., Authentik restart), relay sent
|
If a fix requires another agent (e.g., Authentik restart), relay sent
|
||||||
to responsible agent with full context.
|
to responsible agent with full context.
|
||||||
|
|
||||||
|
---
|
||||||
---
|
---
|
||||||
|
|
||||||
## Execution
|
## Execution
|
||||||
|
|||||||
@@ -1,9 +1,12 @@
|
|||||||
---
|
---
|
||||||
|
report_only_agents:
|
||||||
|
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
|
||||||
name: memory-audit-maintenance
|
name: memory-audit-maintenance
|
||||||
kind: responsibility
|
kind: responsibility
|
||||||
description: Shared memory audit and maintenance contract for all Hermes agents (Mumuni, Tanko, Koby, Koonimo). Each agent runs it against its own isolated memory files — no cross-agent access, no shared state. Detects staleness, enforces writer registry, and rotates canary tokens.
|
description: Shared memory audit and maintenance contract for Hermes agents (Mumuni, Koby, Koonimo). Tanko is no longer a Hermes agent (now on DSH/DeepSeek Harness since 2026-08-27) and uses DSH-native memory, so it is excluded from this Hermes roster. Each agent runs it against its own isolated memory files — no cross-agent access, no shared state. Detects staleness, enforces writer registry, and rotates canary tokens.
|
||||||
id: 067NC4KG01RG50R40M30E20918
|
id: 067NC4KG01RG50R40M30E20918
|
||||||
---
|
---
|
||||||
|
---
|
||||||
|
|
||||||
### Goal
|
### Goal
|
||||||
|
|
||||||
@@ -15,11 +18,10 @@ Autonomously audit and reorganize an agent's native memory (MEMORY.md, USER.md,
|
|||||||
|
|
||||||
### Scope
|
### Scope
|
||||||
|
|
||||||
This contract is the **Hermes Agent standard** for memory maintenance. It is shared across all Hermes agents (Mumuni, Tanko, Koby, Koonimo). Each agent runs it against its own memory files only no cross-agent access, no shared state, no shared ledger, no shared canary. The contract is the standard; each agent enforces it independently with fully isolated data.
|
This contract is the **Hermes Agent standard** for memory maintenance. It is shared across Hermes agents (Mumuni, Koby, Koonimo). **Tanko is excluded — it migrated to DSH (DeepSeek Harness) on 2026-08-27 and now uses DSH-native memory (mnemon), not `~/.hermes/memories/`.** Each agent runs it against its own memory files only no cross-agent access, no shared state, no shared ledger, no shared canary. The contract is the standard; each agent enforces it independently with fully isolated data.
|
||||||
|
|
||||||
**Agent Roster:**
|
**Agent Roster (Hermes):**
|
||||||
- Mumuni
|
- Mumuni
|
||||||
- Tanko
|
|
||||||
- Koby (CT 111 / tdunna)
|
- Koby (CT 111 / tdunna)
|
||||||
- Koonimo (CT 113 / baggy)
|
- Koonimo (CT 113 / baggy)
|
||||||
|
|
||||||
@@ -343,4 +345,3 @@ return {
|
|||||||
|
|
||||||
### Per-Agent Notes
|
### Per-Agent Notes
|
||||||
|
|
||||||
Each Hermes agent (Mumuni, Tanko, Tdunna, Baggy) runs this contract against its own `~/.hermes/memories/` directory. The contract is identical across agents, but all data is fully isolated: separate ledgers, separate writer registries, separate canaries. If a new agent is added to the roster, it must be listed in `### Scope` above and given its own isolated memory directory.
|
|
||||||
@@ -1,4 +1,6 @@
|
|||||||
---
|
---
|
||||||
|
report_only_agents:
|
||||||
|
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
|
||||||
kind: pattern
|
kind: pattern
|
||||||
name: memory-fixer
|
name: memory-fixer
|
||||||
description: >
|
description: >
|
||||||
@@ -6,6 +8,7 @@ description: >
|
|||||||
Escalate anything that needs Kwame's input. Executes confirmed Kwame decisions to completion (state + updated_at).
|
Escalate anything that needs Kwame's input. Executes confirmed Kwame decisions to completion (state + updated_at).
|
||||||
version: 2.0.0
|
version: 2.0.0
|
||||||
---
|
---
|
||||||
|
---
|
||||||
|
|
||||||
# Memory Fixer
|
# Memory Fixer
|
||||||
|
|
||||||
|
|||||||
@@ -6,7 +6,7 @@ description: >
|
|||||||
delegation, verification, and delivery. Defines when to delegate, which
|
delegation, verification, and delivery. Defines when to delegate, which
|
||||||
worker to use for what, how to handle failures, and the kanban board
|
worker to use for what, how to handle failures, and the kanban board
|
||||||
protocol. Enforces context-window discipline and separation of concerns.
|
protocol. Enforces context-window discipline and separation of concerns.
|
||||||
Runs on Mumuni (inside Abiba CT100, minipve, .24) via Hermes agent (Pi + Hermes Zulip gateway).
|
Runs on Mumuni (kagentz CT105, minipve, .14) via Hermes agent (Zulip gateway via systemd).
|
||||||
version: 1.0.0
|
version: 1.0.0
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -20,7 +20,7 @@ version: 1.0.0
|
|||||||
## Topology
|
## Topology
|
||||||
|
|
||||||
**Cluster:** 5 Proxmox nodes (ocupve, acerpve, minipve, amdpve, storepve)
|
**Cluster:** 5 Proxmox nodes (ocupve, acerpve, minipve, amdpve, storepve)
|
||||||
**Manager:** Mumuni (inside Abiba CT100, minipve, .24) via Hermes agent
|
**Manager:** Mumuni (kagentz CT105, minipve, .14) via Hermes agent
|
||||||
**Workers:** 6 profiles, all running on the same agent — no separate hosts needed
|
**Workers:** 6 profiles, all running on the same agent — no separate hosts needed
|
||||||
|
|
||||||
This contract is infrastructure-agnostic in terms of which nodes are used.
|
This contract is infrastructure-agnostic in terms of which nodes are used.
|
||||||
|
|||||||
@@ -54,7 +54,7 @@ garbled input. When a user types these commands to Abiba:
|
|||||||
|
|
||||||
- **`/approve`** → "Pi doesn't have pending approvals. Commands execute immediately."
|
- **`/approve`** → "Pi doesn't have pending approvals. Commands execute immediately."
|
||||||
- **`/approve session`** → Same response
|
- **`/approve session`** → Same response
|
||||||
- **`/deny`** → Same response + "For Hermes agents (Tanko, Mumuni), these work with their built-in approval system."
|
- **`/deny`** → Same response + "For agents (Mumuni on Hermes, Tanko on DSH), these work with their built-in approval system."
|
||||||
|
|
||||||
This keeps the UX consistent across agents — users can type `/approve` anywhere
|
This keeps the UX consistent across agents — users can type `/approve` anywhere
|
||||||
without getting confused by LLM responses.
|
without getting confused by LLM responses.
|
||||||
|
|||||||
+2
-16
@@ -2,17 +2,6 @@
|
|||||||
kind: responsibility
|
kind: responsibility
|
||||||
name: pm2-self-heal
|
name: pm2-self-heal
|
||||||
description: >
|
description: >
|
||||||
Monitors critical PM2 processes (abiba-zulip, abiba-telegram, gitea-runner,
|
|
||||||
spoton-service, zulip-watchdog) and auto-restarts any that are stopped or
|
|
||||||
errored. Logs every action to Gitea (SyslogSolution/health-logs — not
|
|
||||||
knowledge graph, hard rule) and alerts the owner via
|
|
||||||
Zulip DM on failures.
|
|
||||||
CRITICAL: Never restart abiba-zulip — it runs this contract.
|
|
||||||
AS-BUILT 2026-08-09 (captain ruling, ecosystem is authoritative):
|
|
||||||
gpu-monitor is systemd-managed (gpu-monitor.service) — NOT PM2;
|
|
||||||
gpu-watchdog decommissioned (function folded into gpu-monitor.service);
|
|
||||||
gitea-runner KEPT (online in PM2); abiba-zulip KEPT (online 4d+, the
|
|
||||||
2026-07-04 'removed/decommissioned' note was stale and is removed).
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## Maintains
|
## Maintains
|
||||||
@@ -24,9 +13,6 @@ description: >
|
|||||||
- zulip-watchdog: { status: "online", uptime: string, restarts: number }
|
- zulip-watchdog: { status: "online", uptime: string, restarts: number }
|
||||||
- last_check: timestamp
|
- last_check: timestamp
|
||||||
|
|
||||||
> **Note (2026-07-04, SUPERSEDED 2026-08-09):** `abiba-zulip` remains ONLINE and
|
|
||||||
> is monitored — the decommission note was stale (process re-added; do not treat
|
|
||||||
> it as removed).
|
|
||||||
|
|
||||||
## Continuity
|
## Continuity
|
||||||
|
|
||||||
@@ -63,9 +49,9 @@ description: >
|
|||||||
- If status is "online" → pass
|
- If status is "online" → pass
|
||||||
- If status is "stopped" or "errored" → apply Rule 1
|
- If status is "stopped" or "errored" → apply Rule 1
|
||||||
- If restarts > 5 → alert owner
|
- If restarts > 5 → alert owner
|
||||||
3. **Check abiba-zulip** (self-process, read-only):
|
3. **Check abiba-zulip** (live Zulip bridge, heartbeating):
|
||||||
- If status is "online" → pass, log restarts count
|
- If status is "online" → pass, log restarts count
|
||||||
- If status is "stopped" or "errored" → **DO NOT RESTART** — alert owner immediately
|
- If status is "stopped" or "errored" → restart (`pm2 restart abiba-zulip` — fully restored)
|
||||||
- If restarts > 5 in last hour → alert owner with full diagnostics
|
- If restarts > 5 in last hour → alert owner with full diagnostics
|
||||||
4. **Log results** — Append to `SyslogSolution/health-logs/pm2/{timestamp}.md` in Gitea (not knowledge graph — hard rule)
|
4. **Log results** — Append to `SyslogSolution/health-logs/pm2/{timestamp}.md` in Gitea (not knowledge graph — hard rule)
|
||||||
5. **Alert** — Send Zulip DM to owner if escalation needed (do NOT run pm2 commands during alerting)
|
5. **Alert** — Send Zulip DM to owner if escalation needed (do NOT run pm2 commands during alerting)
|
||||||
|
|||||||
@@ -39,7 +39,7 @@ PVE_NODES = {
|
|||||||
|
|
||||||
# Agent definitions: ct, host, user, pve_node, vault_key_name
|
# Agent definitions: ct, host, user, pve_node, vault_key_name
|
||||||
AGENTS = {
|
AGENTS = {
|
||||||
"tanko": {"ct": 112, "host": "192.168.68.122", "user": "jerome", "pve": "amdpve", "vault_key": "TANKO_LITELLM_API_KEY"},
|
"tanko": {"ct": 112, "host": "192.168.68.122", "user": "jerome", "pve": "amdpve", "vault_key": "TANKO_LITELLM_API_KEY", "runtime": "dsh"},
|
||||||
"abiba": {"ct": 100, "host": "192.168.68.24", "user": "root", "pve": "minipve", "vault_key": None}, # Pi agent + Mumuni Zulip, no vault key
|
"abiba": {"ct": 100, "host": "192.168.68.24", "user": "root", "pve": "minipve", "vault_key": None}, # Pi agent + Mumuni Zulip, no vault key
|
||||||
"koby": {"ct": 111, "host": "192.168.68.129", "user": "root", "pve": "amdpve", "vault_key": "KOBY_LITELLM_API_KEY"},
|
"koby": {"ct": 111, "host": "192.168.68.129", "user": "root", "pve": "amdpve", "vault_key": "KOBY_LITELLM_API_KEY"},
|
||||||
"koonimo": {"ct": 113, "host": "192.168.68.114", "user": "root", "pve": "amdpve", "vault_key": "KOONIMO_LITELLM_API_KEY"},
|
"koonimo": {"ct": 113, "host": "192.168.68.114", "user": "root", "pve": "amdpve", "vault_key": "KOONIMO_LITELLM_API_KEY"},
|
||||||
@@ -238,20 +238,36 @@ def check_agents():
|
|||||||
host = agent.get("host")
|
host = agent.get("host")
|
||||||
user = agent.get("user")
|
user = agent.get("user")
|
||||||
ct = agent["ct"]
|
ct = agent["ct"]
|
||||||
|
report_only = agent.get("report_only", False)
|
||||||
|
|
||||||
|
# Tanko runs on DSH (DeepSeek Harness) since 2026-08-27 — it no longer runs a
|
||||||
|
# Hermes gateway, so skip the Hermes gateway/state/streaming/journal checks.
|
||||||
|
if agent.get("runtime") == "dsh":
|
||||||
|
live = ssh(host, "true", user=user)
|
||||||
|
print(f" {'✅' if live is not None else '❌'} {name}: DSH (DeepSeek Harness) — "
|
||||||
|
f"no Hermes gateway since 2026-08-27 (CT {ct}, SSH {'OK' if live is not None else 'FAIL'})")
|
||||||
|
if live is None:
|
||||||
|
FAIL.append(f"unreachable:{name}")
|
||||||
|
continue
|
||||||
|
|
||||||
if not host or not user:
|
if not host or not user:
|
||||||
print(f" ⬜ {name} (CT {ct}): cannot SSH — skip liveness check")
|
print(f" ⬜ {name} (CT {ct}): cannot SSH — skip liveness check")
|
||||||
continue
|
continue
|
||||||
|
|
||||||
# Gateway process
|
# ⛔ KOBY IS NEVER REPAIRED — diagnostic only
|
||||||
pid = ssh(host, "pgrep -f 'hermes_cli.main gateway run' | grep -v infisical | head -1", user=user)
|
if report_only:
|
||||||
if not pid:
|
print(f" 🔍 {name}: REPORT-ONLY mode (diagnostic only, no repairs on .129)")
|
||||||
# Try alternate binary name
|
# Still check gateway status for reporting purposes
|
||||||
pid = ssh(host, "pgrep -f 'hermes.*gateway' | grep -v infisical | grep -v bash | head -1", user=user)
|
pid = ssh(host, "pgrep -f 'hermes_cli.main gateway run' | grep -v infisical | head -1", user=user)
|
||||||
if not pid:
|
if not pid:
|
||||||
print(f" ❌ {name}: GATEWAY NOT RUNNING")
|
pid = ssh(host, "pgrep -f 'hermes.*gateway' | grep -v infisical | grep -v bash | head -1", user=user)
|
||||||
FAIL.append(f"gateway-down:{name}")
|
if not pid:
|
||||||
continue
|
print(f" ⚠️ {name}: GATEWAY NOT RUNNING (reported only)")
|
||||||
|
FAIL.append(f"gateway-down:{name}")
|
||||||
|
continue
|
||||||
|
else:
|
||||||
|
print(f" ✅ {name}: gateway running (pid={pid}, report-only mode)")
|
||||||
|
continue # Skip the rest of the check for Koby
|
||||||
|
|
||||||
# Gateway state file
|
# Gateway state file
|
||||||
state = ssh(host, "cat ~/.hermes/gateway_state.json 2>/dev/null", user=user)
|
state = ssh(host, "cat ~/.hermes/gateway_state.json 2>/dev/null", user=user)
|
||||||
@@ -327,6 +343,10 @@ def check_ct_liveness():
|
|||||||
def check_config_integrity():
|
def check_config_integrity():
|
||||||
"""Verify agent config.yaml parses as valid YAML."""
|
"""Verify agent config.yaml parses as valid YAML."""
|
||||||
for name, agent in AGENTS.items():
|
for name, agent in AGENTS.items():
|
||||||
|
# Tanko runs on DSH (DeepSeek Harness) since 2026-08-27 — no Hermes config.yaml.
|
||||||
|
if agent.get("runtime") == "dsh":
|
||||||
|
print(f" ⏭️ {name}: DSH — no Hermes config.yaml since 2026-08-27")
|
||||||
|
continue
|
||||||
host = agent.get("host")
|
host = agent.get("host")
|
||||||
user = agent.get("user")
|
user = agent.get("user")
|
||||||
if not host or not user:
|
if not host or not user:
|
||||||
@@ -356,6 +376,10 @@ def check_config_integrity():
|
|||||||
def check_wrapper_integrity():
|
def check_wrapper_integrity():
|
||||||
"""Verify the hermes CLI wrapper exists and can reach hermes-real."""
|
"""Verify the hermes CLI wrapper exists and can reach hermes-real."""
|
||||||
for name, agent in AGENTS.items():
|
for name, agent in AGENTS.items():
|
||||||
|
# Tanko runs on DSH (DeepSeek Harness) since 2026-08-27 — no hermes CLI wrapper.
|
||||||
|
if agent.get("runtime") == "dsh":
|
||||||
|
print(f" ⏭️ {name}: DSH — no hermes CLI wrapper since 2026-08-27")
|
||||||
|
continue
|
||||||
host = agent.get("host")
|
host = agent.get("host")
|
||||||
user = agent.get("user")
|
user = agent.get("user")
|
||||||
if not host or not user:
|
if not host or not user:
|
||||||
|
|||||||
@@ -296,21 +296,16 @@ def collect():
|
|||||||
"pm2_uptime": pm2.get("uptime", "?"),
|
"pm2_uptime": pm2.get("uptime", "?"),
|
||||||
}
|
}
|
||||||
|
|
||||||
# Tanko (CT 122)
|
# Tanko (CT 112, IP 192.168.68.122) — DSH (DeepSeek Harness), no Hermes gateway
|
||||||
tanko_state = ssh_jerome("192.168.68.122", "cat ~/.hermes/gateway_state.json 2>/dev/null")
|
# since 2026-08-27. There is no ~/.hermes/gateway_state.json on CT 112 anymore;
|
||||||
tanko_data = {}
|
# Zulip/Telegram connectivity is managed by the DSH harness, not the Hermes gateway.
|
||||||
try:
|
|
||||||
tanko_data = json.loads(tanko_state) if tanko_state else {}
|
|
||||||
except:
|
|
||||||
tanko_data = {}
|
|
||||||
platforms = tanko_data.get("platforms", {})
|
|
||||||
report["agents"]["tanko"] = {
|
report["agents"]["tanko"] = {
|
||||||
"platform": "hermes", "ct": 112, "ip": "192.168.68.122",
|
"platform": "dsh", "ct": 112, "ip": "192.168.68.122",
|
||||||
"gateway_state": tanko_data.get("gateway_state", "unknown"),
|
"gateway_state": "n/a (DSH)",
|
||||||
"zulip_state": platforms.get("zulip", {}).get("state", "unknown"),
|
"zulip_state": "unknown",
|
||||||
"telegram_state": platforms.get("telegram", {}).get("state", "unknown"),
|
"telegram_state": "unknown",
|
||||||
"gateway_pid": tanko_data.get("pid"),
|
"gateway_pid": None,
|
||||||
"updated_at": tanko_data.get("updated_at"),
|
"updated_at": "",
|
||||||
}
|
}
|
||||||
|
|
||||||
# Mumuni (CT 100, IP 192.168.68.24)
|
# Mumuni (CT 100, IP 192.168.68.24)
|
||||||
@@ -543,7 +538,7 @@ th {{ color: #8b949e; font-weight: normal; }}
|
|||||||
elif name == "tanko":
|
elif name == "tanko":
|
||||||
zulip_state = "✅" if agent.get("zulip_state") == "connected" else ("❌" if agent.get("zulip_state") == "disconnected" else "⬜")
|
zulip_state = "✅" if agent.get("zulip_state") == "connected" else ("❌" if agent.get("zulip_state") == "disconnected" else "⬜")
|
||||||
gateway = agent.get("gateway_state", "?")
|
gateway = agent.get("gateway_state", "?")
|
||||||
processed = agent.get("updated_at", "")[:10]
|
processed = "DSH"
|
||||||
elif name == "mumuni":
|
elif name == "mumuni":
|
||||||
zulip_state = "⬜" if agent.get("zulip_state") == "not_installed" else ("✅" if agent.get("zulip_state") == "connected" else "⬜")
|
zulip_state = "⬜" if agent.get("zulip_state") == "not_installed" else ("✅" if agent.get("zulip_state") == "connected" else "⬜")
|
||||||
gateway = agent.get("gateway_state", "?")
|
gateway = agent.get("gateway_state", "?")
|
||||||
|
|||||||
@@ -12,10 +12,10 @@ echo ""
|
|||||||
# Authorized agents for restricted contracts
|
# Authorized agents for restricted contracts
|
||||||
# Format: contract_pattern|authorized_agents (comma-separated)
|
# Format: contract_pattern|authorized_agents (comma-separated)
|
||||||
declare -A RESTRICTED
|
declare -A RESTRICTED
|
||||||
RESTRICTED["infrastructure-control.prose.md"]="abiba"
|
RESTRICTED["infrastructure-control.prose.md"]="abiba,abiba-bot,tanko,tanko-bot,mumuni,mumuni-bot"
|
||||||
RESTRICTED["proxmox-monitor.prose.md"]="abiba"
|
RESTRICTED["proxmox-monitor.prose.md"]="abiba,abiba-bot"
|
||||||
RESTRICTED["hermes-config-template.prose.md"]="abiba,mumuni,tanko"
|
RESTRICTED["hermes-config-template.prose.md"]="abiba,abiba-bot,mumuni,mumuni-bot,tanko,tanko-bot"
|
||||||
RESTRICTED["zulip-health.prose.md"]="abiba,mumuni"
|
RESTRICTED["zulip-health.prose.md"]="abiba,abiba-bot,mumuni,mumuni-bot,tanko,tanko-bot"
|
||||||
RESTRICTED["scripts/pm2-self-heal.sh"]="abiba"
|
RESTRICTED["scripts/pm2-self-heal.sh"]="abiba"
|
||||||
RESTRICTED["scripts/prose-lint.sh"]="abiba"
|
RESTRICTED["scripts/prose-lint.sh"]="abiba"
|
||||||
RESTRICTED["scripts/prose-ai-review.sh"]="abiba"
|
RESTRICTED["scripts/prose-ai-review.sh"]="abiba"
|
||||||
|
|||||||
+11
-26
@@ -9,10 +9,11 @@ ZULIP_EMAIL="abiba-bot@chat.sysloggh.net"
|
|||||||
ZULIP_KEY="cKTDMZAPW08dk3zl05sStzO7HRztzyn8"
|
ZULIP_KEY="cKTDMZAPW08dk3zl05sStzO7HRztzyn8"
|
||||||
OWNER_ZULIP_ID="9"
|
OWNER_ZULIP_ID="9"
|
||||||
|
|
||||||
# Email config
|
|
||||||
GMAIL_USER="jtabiri@gmail.com"
|
LOG="/root/zulip-health-monitor.log"
|
||||||
GMAIL_PASS="rgbuomwcydxwbszd"
|
TIMESTAMP=$(date -u '+%Y-%m-%d %H:%M UTC')
|
||||||
EMAIL_TO="jerome@sysloggh.com"
|
ISSUES=0
|
||||||
|
echo "=== Zulip Health Check — $TIMESTAMP ===" >> "$LOG"
|
||||||
|
|
||||||
notify() {
|
notify() {
|
||||||
local severity="$1" msg="$2"
|
local severity="$1" msg="$2"
|
||||||
@@ -24,30 +25,14 @@ notify() {
|
|||||||
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
|
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
|
||||||
-u "${ZULIP_EMAIL}:${ZULIP_KEY}" \
|
-u "${ZULIP_EMAIL}:${ZULIP_KEY}" \
|
||||||
-d "${form}" > /dev/null 2>&1 || true
|
-d "${form}" > /dev/null 2>&1 || true
|
||||||
|
# Zulip stream post to #agent-hub on topic 'zulip-health'
|
||||||
# Email alert
|
local stream_content="${severity} Zulip Monitor: ${msg}"
|
||||||
local subject="${severity} Zulip Monitor Alert"
|
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
|
||||||
python3 -c "
|
-u "${ZULIP_EMAIL}:${ZULIP_KEY}" \
|
||||||
import smtplib
|
-d "type=stream\&to=%5B7%5D\&topic=zulip-health\&content=$(python3 -c "import urllib.parse; print(urllib.parse.quote(str()))")" \
|
||||||
from email.mime.text import MIMEText
|
> /dev/null 2>&1 || true
|
||||||
m = MIMEText('''${msg}''')
|
|
||||||
m['From'] = 'abiba@sysloggh.com'
|
|
||||||
m['To'] = '${EMAIL_TO}'
|
|
||||||
m['Subject'] = '${subject}'
|
|
||||||
s = smtplib.SMTP('smtp.gmail.com', 587)
|
|
||||||
s.starttls()
|
|
||||||
s.login('${GMAIL_USER}', '${GMAIL_PASS}')
|
|
||||||
s.sendmail('abiba@sysloggh.com', ['${EMAIL_TO}'], m.as_string())
|
|
||||||
s.quit()
|
|
||||||
" 2>/dev/null || true
|
|
||||||
}
|
}
|
||||||
|
|
||||||
TIMESTAMP=$(date -u '+%Y-%m-%d %H:%M UTC')
|
|
||||||
ISSUES=0
|
|
||||||
LOG="/root/zulip-health-monitor.log"
|
|
||||||
|
|
||||||
echo "=== Zulip Health Check — $TIMESTAMP ===" >> "$LOG"
|
|
||||||
|
|
||||||
# ── Global: Zulip Server ──
|
# ── Global: Zulip Server ──
|
||||||
SERVER_CODE=$(curl -s -o /dev/null -w "%{http_code}" --connect-timeout 10 \
|
SERVER_CODE=$(curl -s -o /dev/null -w "%{http_code}" --connect-timeout 10 \
|
||||||
https://chat.sysloggh.net/api/v1/server_settings \
|
https://chat.sysloggh.net/api/v1/server_settings \
|
||||||
|
|||||||
+15
-16
@@ -1,22 +1,24 @@
|
|||||||
---
|
---
|
||||||
kind: responsibility
|
kind: responsibility
|
||||||
name: zulip-health
|
name: zulip-health
|
||||||
description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (Agent Zero Docker), Platform B (Hermes agents Tanko/Mumuni), and the Zulip bridge. Verifies bot registration, DM delivery, and cross-platform connectivity.
|
description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (Agent Zero Docker), Platform B (Tanko on DSH / Mumuni on Hermes), and the Zulip bridge. Verifies bot registration, DM delivery, and cross-platform connectivity.
|
||||||
title: Zulip Mesh Health Monitor — Multi-Platform
|
title: Zulip Mesh Health Monitor — Multi-Platform
|
||||||
version: 3.0.0
|
version: 3.0.0
|
||||||
runtime_contract: 2
|
runtime_contract: 2
|
||||||
agent: abiba
|
agent: abiba
|
||||||
|
report_only_agents:
|
||||||
|
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
|
||||||
---
|
---
|
||||||
|
|
||||||
# Zulip Mesh Health Monitor
|
# Zulip Mesh Health Monitor
|
||||||
|
|
||||||
Monitors ALL Zulip-connected agents across three platforms (pi, Hermes, Agent Zero).
|
Monitors ALL Zulip-connected agents across platforms (pi, Hermes, DSH, Agent Zero).
|
||||||
Runs every 15 minutes in the background. Also triggers on session start.
|
Runs every 15 minutes in the background. Also triggers on session start.
|
||||||
|
|
||||||
## Requires
|
## Requires
|
||||||
|
|
||||||
- **Zulip API key** for `abiba-bot@chat.sysloggh.net` in `$ZULIP_API_KEY`
|
- **Zulip API key** for `abiba-bot@chat.sysloggh.net` in `$ZULIP_API_KEY`
|
||||||
- **SSH access** to Tanko (192.168.68.122), Mumuni (192.168.68.24, inside Abiba CT100 on minipve), and Agent Zero Docker host (192.168.68.14)
|
- **SSH access** to Tanko (192.168.68.122), Mumuni (192.168.68.14, kagentz CT105 on minipve), and Agent Zero Docker host (192.168.68.14)
|
||||||
- **PM2** on localhost for pi process management
|
- **PM2** on localhost for pi process management
|
||||||
- **Network access** to `chat.sysloggh.net`, `localhost:9200`
|
- **Network access** to `chat.sysloggh.net`, `localhost:9200`
|
||||||
- **Write access** to `/root/zulip-health-monitor.log` and `/tmp/zulip-monitor-debounce`
|
- **Write access** to `/root/zulip-health-monitor.log` and `/tmp/zulip-monitor-debounce`
|
||||||
@@ -45,7 +47,7 @@ Runs every 15 minutes in the background. Also triggers on session start.
|
|||||||
"severity": "healthy"
|
"severity": "healthy"
|
||||||
},
|
},
|
||||||
"tanko": {
|
"tanko": {
|
||||||
"platform": "hermes",
|
"platform": "dsh",
|
||||||
"zulip_state": "connected",
|
"zulip_state": "connected",
|
||||||
"heartbeat_age_seconds": 45,
|
"heartbeat_age_seconds": 45,
|
||||||
"gateway_pid": 1234,
|
"gateway_pid": 1234,
|
||||||
@@ -81,13 +83,13 @@ Log as "unreachable" — don't treat as critical unless it persists for 3+ conse
|
|||||||
## Streaming Support (2026-07-05)
|
## Streaming Support (2026-07-05)
|
||||||
|
|
||||||
Zulip agents now support progressive message editing during agent generation.
|
Zulip agents now support progressive message editing during agent generation.
|
||||||
When a Hermes agent (Tanko, Mumuni) processes a message, the response is
|
When a Zulip agent (Tanko on DSH, Mumuni on Hermes) processes a message, the response is
|
||||||
streamed in real-time via Zulip's `PATCH /api/v1/messages/{id}` API:
|
streamed in real-time via Zulip's `PATCH /api/v1/messages/{id}` API:
|
||||||
|
|
||||||
- Adapter implements `edit_message()` using `_api_patch()` helper
|
- Adapter implements `edit_message()` using `_api_patch()` helper
|
||||||
- Gateway stream consumer progressively edits the Zulip message
|
- Gateway stream consumer progressively edits the Zulip message
|
||||||
- User sees real-time agent thinking instead of waiting for full response
|
- User sees real-time agent thinking instead of waiting for full response
|
||||||
- Verified: Tanko (CT 112) and Mumuni (inside Abiba CT 100) both have streaming active
|
- Verified: Tanko (CT 112) and Mumuni (kagentz CT 105) both have streaming active
|
||||||
|
|
||||||
### Verification
|
### Verification
|
||||||
```bash
|
```bash
|
||||||
@@ -182,15 +184,14 @@ grep -a "Finalized\|Failed to finalize" /root/.pm2/logs/abiba-zulip-out.log | ta
|
|||||||
| `last_error` set | Log and monitor |
|
| `last_error` set | Log and monitor |
|
||||||
| Crash loop >10/h | Alert user |
|
| Crash loop >10/h | Alert user |
|
||||||
|
|
||||||
### Step 3: Platform B — Hermes (Tanko .122, Mumuni .24)
|
|
||||||
|
|
||||||
**B1: Gateway State**
|
**B1: Gateway State**
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
ssh root@192.168.68.122 "cat ~/.hermes/gateway_state.json"
|
|
||||||
ssh root@192.168.68.24 "cat ~/.hermes/gateway_state.json" # Mumuni inside Abiba CT100
|
|
||||||
```
|
```
|
||||||
|
|
||||||
|
Tanko runs on DSH (DeepSeek Harness) — it no longer runs a Hermes gateway, so there is no `~/.hermes/gateway_state.json` on CT 112 (.122). Verify Tanko's Zulip connectivity via the DSH harness bot status instead.
|
||||||
|
|
||||||
Check `platforms.zulip.state`: `connected` ✅ | `disconnected` ❌ | `error` ❌ | missing → not installed.
|
Check `platforms.zulip.state`: `connected` ✅ | `disconnected` ❌ | `error` ❌ | missing → not installed.
|
||||||
|
|
||||||
**B2: Agent Process**
|
**B2: Agent Process**
|
||||||
@@ -203,24 +204,22 @@ Gateway PID should exist with uptime > 60s. **Dual-gateway detection**: if more
|
|||||||
|
|
||||||
| Agent | Restart command | Notes |
|
| Agent | Restart command | Notes |
|
||||||
|-------|-----------------|-------|
|
|-------|-----------------|-------|
|
||||||
| Mumuni (.24) | `pm2 restart abiba-zulip` | Hermes gateway runs under PM2 as `abiba-zulip` |
|
|
||||||
| Tanko (.122) | `bash /opt/hermes-zulip-plugin/run.sh` (or the agent's systemd/user unit) | Tanko does NOT use PM2 — never run `pm2 restart mumuni-zulip` for Tanko (process does not exist) |
|
|
||||||
|
|
||||||
Check gateway log for "Gateway running with 2 platform(s)" (not 1) to confirm Zulip reloaded.
|
Check gateway log for "Gateway running with 2 platform(s)" (not 1) to confirm Zulip reloaded.
|
||||||
|
|
||||||
**B3: Heartbeat Verification**
|
**B3: Heartbeat Verification** (Hermes agent Mumuni only — Tanko has no Hermes gateway)
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
ssh root@<CT> "grep Heartbeat ~/.hermes/logs/agent.log | tail -3"
|
ssh root@192.168.68.24 "grep Heartbeat ~/.hermes/logs/agent.log | tail -3"
|
||||||
```
|
```
|
||||||
|
|
||||||
Expected: recent heartbeat (within 5 min), `polls=N` incrementing.
|
Expected: recent heartbeat (within 5 min), `polls=N` incrementing.
|
||||||
Silence > 300s → warning. Silence > 600s → critical.
|
Silence > 300s → warning. Silence > 600s → critical.
|
||||||
|
|
||||||
**B4: Response Delivery**
|
**B4: Response Delivery** (Hermes agent Mumuni only)
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
ssh root@<CT> "grep -E 'Finalized|Failed to finalize|Replied to' ~/.hermes/logs/agent.log | tail -10"
|
ssh root@192.168.68.24 "grep -E 'Finalized|Failed to finalize|Replied to' ~/.hermes/logs/agent.log | tail -10"
|
||||||
```
|
```
|
||||||
|
|
||||||
> 50% fail rate → critical.
|
> 50% fail rate → critical.
|
||||||
@@ -229,7 +228,7 @@ ssh root@<CT> "grep -E 'Finalized|Failed to finalize|Replied to' ~/.hermes/logs/
|
|||||||
|
|
||||||
| Condition | Action |
|
| Condition | Action |
|
||||||
|-----------|--------|
|
|-----------|--------|
|
||||||
| `zulip.state != "connected"` | `ssh root@<CT> "pkill -f 'gateway run'; sleep 2; hermes gateway restart"` |
|
| `zulip.state != "connected"` | `ssh root@<CT> "pkill -f 'gateway run'; sleep 2; hermes gateway restart"` (Mumuni) / restart Tanko via DSH service |
|
||||||
| No heartbeat in 10min | Same as above |
|
| No heartbeat in 10min | Same as above |
|
||||||
| `Failed to finalize` > 50% | Check PATCH API, Zulip server |
|
| `Failed to finalize` > 50% | Check PATCH API, Zulip server |
|
||||||
| Response empty/short | Check A2A endpoint / LiteLLM model |
|
| Response empty/short | Check A2A endpoint / LiteLLM model |
|
||||||
|
|||||||
@@ -10,7 +10,7 @@ description: >
|
|||||||
|
|
||||||
> **⚠️ RETIRED** — The pi Zulip extension (`~/.pi/agent/extensions/zulip/`) and
|
> **⚠️ RETIRED** — The pi Zulip extension (`~/.pi/agent/extensions/zulip/`) and
|
||||||
> PM2 process (`abiba-zulip`) have been decommissioned. All mention/reliability
|
> PM2 process (`abiba-zulip`) have been decommissioned. All mention/reliability
|
||||||
> monitoring now happens through Telegram. Hermes agents (Tanko, Mumuni) and
|
> monitoring now happens through Telegram. Agents (Mumuni on Hermes, Tanko on DSH) and
|
||||||
> Agent Zero (kagentz) continue to use Zulip.
|
> Agent Zero (kagentz) continue to use Zulip.
|
||||||
|
|
||||||
## Maintains
|
## Maintains
|
||||||
|
|||||||
@@ -420,7 +420,7 @@ Backup v2 before starting: `cp index.js index.js.v2-backup-$(date +%Y%m%d-%H%M%S
|
|||||||
|-------|----------|-------------|--------------|-------------|
|
|-------|----------|-------------|--------------|-------------|
|
||||||
| **Abiba** | pi (CT 100) | ✅ Connected | API key missing from Infisical injection; poll timeout noise | Added .env fallback; AbortError treated as empty poll (no retry); poll timeout 65s→90s |
|
| **Abiba** | pi (CT 100) | ✅ Connected | API key missing from Infisical injection; poll timeout noise | Added .env fallback; AbortError treated as empty poll (no retry); poll timeout 65s→90s |
|
||||||
| **Tanko** | Hermes (CT 112) | ✅ Connected | Gateway disconnected since Jul 11; watchdog restart didn't re-establish Zulip | Full gateway restart (kill wrapper, let infisical-gateway.sh respawn) |
|
| **Tanko** | Hermes (CT 112) | ✅ Connected | Gateway disconnected since Jul 11; watchdog restart didn't re-establish Zulip | Full gateway restart (kill wrapper, let infisical-gateway.sh respawn) |
|
||||||
| **Mumuni** | Hermes (inside Abiba CT 100) | ✅ Connected | No issues found | None needed |
|
| **Mumuni** | Hermes (kagentz CT 105, migrated 2026-08-29) | ✅ Connected | No issues found | None needed |
|
||||||
|
|
||||||
### Key Fixes Applied
|
### Key Fixes Applied
|
||||||
|
|
||||||
|
|||||||
@@ -13,7 +13,8 @@ triggers:
|
|||||||
|
|
||||||
> **⚠️ RETIRED** — This contract was embedded in the pi Zulip extension code
|
> **⚠️ RETIRED** — This contract was embedded in the pi Zulip extension code
|
||||||
> (`performHealthCheck()`). That code has been removed. Zulip self-healing for
|
> (`performHealthCheck()`). That code has been removed. Zulip self-healing for
|
||||||
> Hermes agents (Tanko, Mumuni) continues through their own gateway monitoring.
|
> agents (Mumuni on Hermes, Tanko on DSH) continues through their own platform
|
||||||
|
> monitoring.
|
||||||
|
|
||||||
## Maintains
|
## Maintains
|
||||||
|
|
||||||
@@ -65,8 +66,7 @@ triggers:
|
|||||||
|------|----|------|---------|
|
|------|----|------|---------|
|
||||||
| Zulip server | 192.168.68.19 | root | Docker: `zulip-zulip-1` |
|
| Zulip server | 192.168.68.19 | root | Docker: `zulip-zulip-1` |
|
||||||
| Abiba (pi) | localhost | root | PM2: `abiba-zulip` |
|
| Abiba (pi) | localhost | root | PM2: `abiba-zulip` |
|
||||||
| Mumuni | 192.168.68.24 (CT100 abiba) | root | `hermes gateway restart` |
|
| Tanko | 192.168.68.122 (CT 112) | jerome | DSH (DeepSeek Harness) — restart via DSH service, not `hermes gateway restart` (no longer a Hermes agent since 2026-08-27) |
|
||||||
| Tanko | 192.168.68.122 | jerome | `PATH=$PATH:/home/jerome/.hermes/hermes-agent hermes gateway restart` |
|
|
||||||
|
|
||||||
## Debounce
|
## Debounce
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user