Compare commits
90
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
767bd22d9c | ||
|
|
e97145c88f | ||
|
|
55f1208eb8 | ||
|
|
b9adf353ee | ||
|
|
aeb66ea22d | ||
|
|
288f74cf84 | ||
|
|
c66671dbee | ||
|
|
b80d3142aa | ||
|
|
266fa1f835 | ||
|
|
85ea1f4f3d | ||
|
|
b2a259fa23 | ||
|
|
fb185ed90a | ||
|
|
b001b657d4 | ||
|
|
a1ffeaad34 | ||
|
|
287657a77a | ||
|
|
21f9073e0b | ||
|
|
32fe7c0652 | ||
|
|
25cf2f5eef | ||
|
|
26f2301188 | ||
|
|
a6a459acc0 | ||
|
|
14bed6e916 | ||
|
|
2e70c834cb | ||
|
|
4f59b82404 | ||
|
|
8a4ee4cf5a | ||
|
|
952fca9c92 | ||
|
|
b1462f3e79 | ||
|
|
19ed186d0a | ||
|
|
88f27e75ed | ||
|
|
bbdf6c1249 | ||
|
|
1974959cc9 | ||
|
|
c59c9fb174 | ||
|
|
194e256ac5 | ||
|
|
c66d9c1e20 | ||
|
|
532250b017 | ||
|
|
f57923b4fa | ||
|
|
7ed2e4e923 | ||
|
|
91d16d2693 | ||
|
|
7ff7ce5b33 | ||
|
|
36f218e255 | ||
|
|
947e8b24e0 | ||
|
|
95b4a0e6b0 | ||
|
|
028f276be4 | ||
|
|
17751e24d1 | ||
|
|
79eeb457fc | ||
|
|
1b8186f6b9 | ||
|
|
9dd0cb18d5 | ||
|
|
4ac60e14a3 | ||
|
|
2e0b737f2d | ||
|
|
bbe9ee533f | ||
|
|
4684ee64e0 | ||
|
|
af3d364242 | ||
|
|
2716e55c16 | ||
|
|
5313e6b9ba | ||
|
|
bfdff13ae7 | ||
|
|
0e4eda0abb | ||
|
|
801de0a25c | ||
|
|
8bf32f6f0f | ||
|
|
36ae464f59 | ||
|
|
a3e97ce72b | ||
|
|
cc5fe0991c | ||
|
|
dc78604360 | ||
|
|
7bbf148778 | ||
|
|
3a25c7cce5 | ||
|
|
0aa0ea4906 | ||
|
|
19821ed6b5 | ||
|
|
403fbcdd9f | ||
|
|
0753f38cf9 | ||
|
|
274596fdd1 | ||
|
|
8c4df63db4 | ||
|
|
79a1d22c99 | ||
|
|
782831f548 | ||
|
|
143dd3f16b | ||
|
|
11076ad174 | ||
|
|
b079c02d0c | ||
|
|
09065e7dee | ||
|
|
e8b9f990b2 | ||
|
|
2dfc3e1530 | ||
|
|
79af0ae7a3 | ||
|
|
30c821469b | ||
|
|
3dcbbf1d76 | ||
|
|
aac4c7eac3 | ||
|
|
031ad814a0 | ||
|
|
ca39fead74 | ||
|
|
9b280060b7 | ||
|
|
e898048baf | ||
|
|
b01469ba18 | ||
|
|
994ae1b7ac | ||
|
|
b835986d44 | ||
|
|
d6ad016ac9 | ||
|
|
c3306e87e4 |
@@ -67,10 +67,10 @@ Two incidents taught us this:
|
|||||||
|
|
||||||
| Contract | Sensitivity | Who can change |
|
| Contract | Sensitivity | Who can change |
|
||||||
|----------|------------|----------------|
|
|----------|------------|----------------|
|
||||||
| `infrastructure-control.prose.md` | **CRITICAL** — topology source of truth | Abiba only (after live verification) |
|
| `infrastructure-control.prose.md` | **CRITICAL** — topology source of truth | Abiba, Tanko (Tanko maintains its own CT row) |
|
||||||
| `proxmox-monitor.prose.md` | **CRITICAL** — deployed monitoring | Abiba only |
|
| `proxmox-monitor.prose.md` | **CRITICAL** — deployed monitoring | Abiba only |
|
||||||
| `hermes-config-template.prose.md` | **HIGH** — all agent configs | Abiba, Mumuni, Tanko |
|
| `hermes-config-template.prose.md` | **HIGH** — all agent configs | Abiba, Mumuni, Tanko |
|
||||||
| `zulip-health.prose.md` | **HIGH** — agent communication | Abiba, Mumuni |
|
| `zulip-health.prose.md` | **HIGH** — agent communication | Abiba, Mumuni, Tanko |
|
||||||
| Other contracts | Normal | Any registered agent |
|
| Other contracts | Normal | Any registered agent |
|
||||||
| `scripts/*.sh` | **HIGH** — runtime scripts | Abiba only |
|
| `scripts/*.sh` | **HIGH** — runtime scripts | Abiba only |
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,2 @@
|
|||||||
|
<!-- Points Claude at AGENTS.md via import; edit AGENTS.md, not this file. -->
|
||||||
|
@AGENTS.md
|
||||||
@@ -20,7 +20,8 @@ An agent ran `pct set` without checking `pct config` first and changed the IP
|
|||||||
to the wrong value, breaking Zulip. The staleness was harmless until acted on.
|
to the wrong value, breaking Zulip. The staleness was harmless until acted on.
|
||||||
|
|
||||||
**Layer 3 enforcement — the `safe-mutate` wrapper** (deployed on all 5 agents:
|
**Layer 3 enforcement — the `safe-mutate` wrapper** (deployed on all 5 agents:
|
||||||
Abiba, Tanko, Mumuni, Koby, Koonimo). ALL infrastructure mutations MUST go
|
Abiba, Tanko, Mumuni, Koby, Koonimo — Tanko runs on DSH/DeepSeek Harness since
|
||||||
|
2026-08-27 but safe-mutate enforcement still applies). ALL infrastructure mutations MUST go
|
||||||
through `safe-mutate`. Raw `sed -i`, `pct set`, `docker compose up
|
through `safe-mutate`. Raw `sed -i`, `pct set`, `docker compose up
|
||||||
--force-recreate`, `kill`, `rm` on infrastructure outside `safe-mutate` is an
|
--force-recreate`, `kill`, `rm` on infrastructure outside `safe-mutate` is an
|
||||||
auditable protocol violation. The wrapper runs a verify command, optionally
|
auditable protocol violation. The wrapper runs a verify command, optionally
|
||||||
@@ -161,7 +162,7 @@ Companion shell scripts that contracts delegate to.
|
|||||||
|---|---|
|
|---|---|
|
||||||
| `daily-infra-report.py` | Generates the daily infrastructure dashboard (HTML email to jerome@sysloggh.com). |
|
| `daily-infra-report.py` | Generates the daily infrastructure dashboard (HTML email to jerome@sysloggh.com). |
|
||||||
| `pm2-self-heal.sh` | Shell companion to the pm2-self-heal contract — restarts crashed PM2 processes (runs every 5 min). |
|
| `pm2-self-heal.sh` | Shell companion to the pm2-self-heal contract — restarts crashed PM2 processes (runs every 5 min). |
|
||||||
| `agent-health-check.py` | Consolidated agent health: LiteLLM key validation + GPU port conflict + streaming checks (every 10 min). Replaced zulip-monitor.sh. |
|
| `agent-health-check.py` | Consolidated agent health: LiteLLM key validation + GPU port conflict + streaming checks (every 10 min). |
|
||||||
|
|
||||||
## Contract Structure
|
## Contract Structure
|
||||||
|
|
||||||
|
|||||||
@@ -1,4 +1,6 @@
|
|||||||
---
|
---
|
||||||
|
report_only_agents:
|
||||||
|
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
|
||||||
kind: function
|
kind: function
|
||||||
name: abiba-zulip-restore
|
name: abiba-zulip-restore
|
||||||
description: >
|
description: >
|
||||||
@@ -11,6 +13,7 @@ version: 1.0.0
|
|||||||
status: active
|
status: active
|
||||||
runtime_contract: 2
|
runtime_contract: 2
|
||||||
---
|
---
|
||||||
|
---
|
||||||
|
|
||||||
# Abiba Zulip Restore — Resume pi Zulip Communication
|
# Abiba Zulip Restore — Resume pi Zulip Communication
|
||||||
|
|
||||||
@@ -304,6 +307,7 @@ module.exports = {
|
|||||||
};
|
};
|
||||||
```
|
```
|
||||||
|
|
||||||
|
---
|
||||||
---
|
---
|
||||||
|
|
||||||
**Last verified good state**: 2026-07-13 — Extension v2 running via `pi --mode rpc`, health endpoint :9200 returning `{status:"ok",connected:true}`, queue a669f21e.
|
**Last verified good state**: 2026-07-13 — Extension v2 running via `pi --mode rpc`, health endpoint :9200 returning `{status:"ok",connected:true}`, queue a669f21e.
|
||||||
|
|||||||
@@ -0,0 +1,186 @@
|
|||||||
|
# Agent Zero Issue Fix Summary
|
||||||
|
|
||||||
|
**Date**: 2026-09-01
|
||||||
|
**Agent**: Agent Zero (Docker container on kagentz CT105)
|
||||||
|
**Issue**: AuthenticationError + Telegram conflicts
|
||||||
|
**Status**: ✅ RESOLVED
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Problems Identified
|
||||||
|
|
||||||
|
### 1. OpenRouter Authentication Error (CRITICAL)
|
||||||
|
```
|
||||||
|
litellm.exceptions.AuthenticationError: OpenrouterException -
|
||||||
|
{"error":{"message":"User not found.","code":401}}
|
||||||
|
```
|
||||||
|
**Root Cause**: The OpenRouter API key in `/a0/usr/.env` belonged to a different OpenRouter user.
|
||||||
|
|
||||||
|
**Old Key**: `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`
|
||||||
|
**New Key**: `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`
|
||||||
|
**New User**: `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`
|
||||||
|
|
||||||
|
### 2. Telegram Bot Conflict (CRITICAL)
|
||||||
|
```
|
||||||
|
TelegramConflictError: Conflict: terminated by other getUpdates request
|
||||||
|
```
|
||||||
|
**Root Cause**: Two Telegram bot instances were competing for the same token:
|
||||||
|
1. Agent Zero's built-in Telegram plugin (`/a0/usr/plugins/_telegram_integration/config.json`)
|
||||||
|
2. Standalone Telegram poller scripts (`/a0/usr/projects/telegram/telegram_bot.py`)
|
||||||
|
|
||||||
|
Both were using token `8476855065:***` in polling mode.
|
||||||
|
|
||||||
|
**Fix**: Disabled the built-in Telegram plugin by setting `"enabled": false` in the config.
|
||||||
|
|
||||||
|
### 3. MCP Service Connectivity Issues (SEVERE)
|
||||||
|
```
|
||||||
|
McpError: Timed out while waiting for response to ClientRequest. Waited 10.0 seconds.
|
||||||
|
```
|
||||||
|
**Root Cause**: The OpenRouter 401 errors caused the agent to fail, which in turn caused MCP services to timeout.
|
||||||
|
|
||||||
|
**Status**: ✅ RESOLVED with OpenRouter key fix.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Fixes Applied
|
||||||
|
|
||||||
|
### Fix 1: Update OpenRouter Key
|
||||||
|
```bash
|
||||||
|
# Container .env update
|
||||||
|
sudo docker exec agent-zero bash -c '
|
||||||
|
sed -i "s|^API_KEY_OPENROUTER=.*|API_KEY_OPENROUTER=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|" /a0/usr/.env
|
||||||
|
'
|
||||||
|
```
|
||||||
|
|
||||||
|
**Verification**:
|
||||||
|
```bash
|
||||||
|
curl -s https://openrouter.ai/api/v1/auth/key \
|
||||||
|
-H "Authorization: Bearer sk-or-v1-0af3f3..." | python3 -m json.tool
|
||||||
|
```
|
||||||
|
Result: HTTP 200, user `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`, not free tier.
|
||||||
|
|
||||||
|
### Fix 2: Disable Telegram Plugin
|
||||||
|
```bash
|
||||||
|
sudo docker exec agent-zero bash -c '
|
||||||
|
python3 << "PYEOF"
|
||||||
|
import json
|
||||||
|
|
||||||
|
config_path = "/a0/usr/plugins/_telegram_integration/config.json"
|
||||||
|
with open(config_path) as f:
|
||||||
|
config = json.load(f)
|
||||||
|
|
||||||
|
config["bots"][0]["enabled"] = False
|
||||||
|
|
||||||
|
with open(config_path, "w") as f:
|
||||||
|
json.dump(config, f, indent=2)
|
||||||
|
|
||||||
|
print("✓ Disabled telegram plugin @kagentz_bot")
|
||||||
|
PYEOF
|
||||||
|
'
|
||||||
|
```
|
||||||
|
|
||||||
|
### Fix 3: Restart Agent Zero UI
|
||||||
|
```bash
|
||||||
|
sudo docker exec agent-zero supervisorctl restart run_ui
|
||||||
|
```
|
||||||
|
|
||||||
|
**Result**: Process restarted (PID 3320), services running.
|
||||||
|
|
||||||
|
### Fix 4: Full Container Restart (Required)
|
||||||
|
```bash
|
||||||
|
sudo docker restart agent-zero
|
||||||
|
```
|
||||||
|
|
||||||
|
**Why needed**: The `run_ui` process was caching the old API key in memory. A full container restart was required to force Agent Zero to reload the `.env` file with the new OpenRouter key.
|
||||||
|
|
||||||
|
**Result**: All services restarted cleanly, no more 401 errors.
|
||||||
|
|
||||||
|
### Fix 5: Update Stale `.env.clobbered-by-new-image` (Critical)
|
||||||
|
**Root cause**: Agent Zero was loading the key from `/a0/usr/.env.clobbered-by-new-image` (line 28) instead of the main `/a0/usr/.env` (line 72). The clobbered file still had the old, stale key.
|
||||||
|
|
||||||
|
**Fix**:
|
||||||
|
```bash
|
||||||
|
KEY=$(grep "^API_KEY_OPENROUTER=" /a0/usr/.env | cut -d"=" -f2-)
|
||||||
|
sed -i "s|^API_KEY_OPENROUTER=.*|API_KEY_OPENROUTER=$KEY|" /a0/usr/.env.clobbered-by-new-image
|
||||||
|
```
|
||||||
|
|
||||||
|
**Lesson**: When updating Agent Zero's `.env`, check BOTH files:
|
||||||
|
- `/a0/usr/.env` (main)
|
||||||
|
- `/a0/usr/.env.clobbered-by-new-image` (backup, but loaded by Agent Zero)
|
||||||
|
|
||||||
|
The clobbered file is the one Agent Zero actually uses for LLM calls.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Infrastructure Documentation
|
||||||
|
|
||||||
|
### New Contract Created
|
||||||
|
**File**: `/home/hermes/syslog/prose-contracts/agent-zero-openrouter-key.prose.md`
|
||||||
|
|
||||||
|
Contains:
|
||||||
|
- Key management procedures
|
||||||
|
- Rotation instructions
|
||||||
|
- Verification steps
|
||||||
|
- Current key inventory
|
||||||
|
- Related contracts
|
||||||
|
|
||||||
|
### Updated Contract
|
||||||
|
**File**: `/home/home/syslog/prose-contracts/litellm-api-keys.prose.md`
|
||||||
|
|
||||||
|
Added section:
|
||||||
|
- Agent Zero OpenRouter integration
|
||||||
|
- Key storage locations
|
||||||
|
- Model configuration
|
||||||
|
- Why not LiteLLM proxy
|
||||||
|
- Rotation procedure
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Current State
|
||||||
|
|
||||||
|
| Component | Status | Details |
|
||||||
|
|-----------|--------|---------|
|
||||||
|
| **OpenRouter Key** | ✅ Valid | `sk-or-v1-0af3f3…`, user verified |
|
||||||
|
| **Telegram Bot** | ✅ Resolved | Plugin disabled, conflicts cleared |
|
||||||
|
| **MCP Services** | ✅ Working | No timeouts after key fix |
|
||||||
|
| **Container** | ✅ Running | PID 3320, uptime 16+ hours |
|
||||||
|
| **Services** | ✅ All UP | run_ui, run_tunnel_api, run_searxng, run_cron, the_listener |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Related Files
|
||||||
|
|
||||||
|
| Path | Purpose |
|
||||||
|
|------|---------|
|
||||||
|
| `/a0/usr/.env` | Container key storage |
|
||||||
|
| `/a0/usr/plugins/_telegram_integration/config.json` | Telegram plugin config |
|
||||||
|
| `/a0/usr/plugins/_model_config/presets.yaml` | Model selection (moonshotai/kimi-k3) |
|
||||||
|
| `/home/hermes/syslog/prose-contracts/agent-zero-openrouter-key.prose.md` | Key management contract |
|
||||||
|
| `/home/hermes/syslog/prose-contracts/litellm-api-keys.prose.md` | Fleet key inventory |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Next Steps
|
||||||
|
|
||||||
|
1. **Sync key to Infisical vault** (optional, currently .env fallback only)
|
||||||
|
2. **Monitor usage** — Check OpenRouter dashboard for daily/weekly spend
|
||||||
|
3. **Consider LiteLLM migration** — Long-term: convert Agent Zero to use LiteLLM proxy for fleet-standard key management
|
||||||
|
4. **Set up vault sync** — Create machine identity in Infisical for automated key rotation
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Prevention
|
||||||
|
|
||||||
|
To prevent similar issues:
|
||||||
|
|
||||||
|
1. **Always verify API keys** against their providers before using
|
||||||
|
2. **Keep fleet-wide key inventory** updated in prose contracts
|
||||||
|
3. **Rotate keys on schedule** (quarterly hygiene, not on-demand only)
|
||||||
|
4. **Test key changes** in staging before production rollout
|
||||||
|
5. **Document key locations** in both code and prose contracts
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
**Verified by**: Mumuni 🦅
|
||||||
|
**Last updated**: 2026-09-01
|
||||||
|
**Session**: 1
|
||||||
@@ -0,0 +1,129 @@
|
|||||||
|
---
|
||||||
|
kind: function
|
||||||
|
name: agent-zero-openrouter-key
|
||||||
|
description: >
|
||||||
|
Manages the OpenRouter API key for Agent Zero (Docker container on kagentz .14).
|
||||||
|
Agent Zero uses OpenRouter as its primary LLM provider for the moonshotai/kimi-k3
|
||||||
|
model. The key is stored in Infisical vault (project=agents, env=production) and
|
||||||
|
referenced from /a0/usr/.env in the container. Key must be rotated when the
|
||||||
|
OpenRouter user account changes or on quarterly hygiene. Last verified: 2026-09-01.
|
||||||
|
---
|
||||||
|
|
||||||
|
## Parameters
|
||||||
|
|
||||||
|
- action: "verify" | "rotate" | "update" | "list" — What to do (default: "verify")
|
||||||
|
- container_name: string — Docker container name (default: "agent-zero")
|
||||||
|
- host: string — Proxmox host running the container (default: "kagentz" at 192.168.68.14)
|
||||||
|
- env_path: string — Path to .env file in container (default: "/a0/usr/.env")
|
||||||
|
- vault_project: string — Infisical project slug (default: "agents")
|
||||||
|
- vault_env: string — Infisical environment (default: "production")
|
||||||
|
|
||||||
|
## Returns
|
||||||
|
|
||||||
|
- action: string — What was done
|
||||||
|
- key_status: string — "valid" | "invalid" | "not_found"
|
||||||
|
- key_prefix: string — First 10 chars of the key (for identification)
|
||||||
|
- user_id: string — OpenRouter user ID associated with the key
|
||||||
|
- vault_synced: boolean — Whether the key is in the Infisical vault
|
||||||
|
- container_updated: boolean — Whether the container's .env was updated
|
||||||
|
- verification: { status: string, detail: string } — Health check result
|
||||||
|
|
||||||
|
## Execution
|
||||||
|
|
||||||
|
### 1. Verify the key
|
||||||
|
|
||||||
|
1. **Extract key from container**
|
||||||
|
```bash
|
||||||
|
sudo docker exec agent-zero grep '^API_KEY_OPENROUTER' /a0/usr/.env | cut -d'=' -f2-
|
||||||
|
```
|
||||||
|
|
||||||
|
2. **Test against OpenRouter API**
|
||||||
|
```bash
|
||||||
|
curl -s https://openrouter.ai/api/v1/auth/key \
|
||||||
|
-H "Authorization: Bearer <key>" | python3 -m json.tool
|
||||||
|
```
|
||||||
|
Expected: HTTP 200, JSON with `data.label` and `data.is_free_tier`
|
||||||
|
|
||||||
|
3. **Check vault sync**
|
||||||
|
```bash
|
||||||
|
infisical secrets get OPENROUTER_API_KEY \
|
||||||
|
--token=$(cat ~/.infisical-token) \
|
||||||
|
--projectId=agents \
|
||||||
|
--env=production \
|
||||||
|
--domain=https://vault.sysloggh.net
|
||||||
|
```
|
||||||
|
|
||||||
|
4. **Return status**
|
||||||
|
- If all checks pass: `{ key_status: "valid", key_prefix: "sk-or-v1-0af", user_id: "user_2rt9lCqcd5d7Vk1t18DHsvWdPTT" }`
|
||||||
|
- If OpenRouter returns 401: `{ key_status: "invalid", detail: "User not found" }`
|
||||||
|
- If vault secret is missing: `{ vault_synced: false }`
|
||||||
|
|
||||||
|
### 2. Rotate the key
|
||||||
|
|
||||||
|
1. **Generate new key** in OpenRouter UI or via API
|
||||||
|
2. **Update container .env**
|
||||||
|
```bash
|
||||||
|
sudo docker exec agent-zero sed -i 's/^API_KEY_OPENROUTER=.*/API_KEY_OPENROUTER=<new_key>/' /a0/usr/.env
|
||||||
|
```
|
||||||
|
3. **Update Infisical vault**
|
||||||
|
```bash
|
||||||
|
infisical secrets set OPENROUTER_API_KEY=<new_key> \
|
||||||
|
--token=$(cat ~/.infisical-token) \
|
||||||
|
--projectId=agents \
|
||||||
|
--env=production \
|
||||||
|
--domain=https://vault.sysloggh.net
|
||||||
|
```
|
||||||
|
4. **Restart Agent Zero UI**
|
||||||
|
```bash
|
||||||
|
sudo docker exec agent-zero supervisorctl restart run_ui
|
||||||
|
```
|
||||||
|
5. **Verify** — Run "verify" action again
|
||||||
|
|
||||||
|
### 3. Update (key changed but no rotation)
|
||||||
|
|
||||||
|
1. **Update container .env** (same as rotate step 2)
|
||||||
|
2. **Sync vault** (same as rotate step 3)
|
||||||
|
3. **Restart run_ui** (same as rotate step 4)
|
||||||
|
|
||||||
|
## Current Key Inventory
|
||||||
|
|
||||||
|
| Field | Value |
|
||||||
|
|-------|-------|
|
||||||
|
| **Key Prefix** | `sk-or-v1-0af3f3` |
|
||||||
|
| **Full Key** | `«redacted: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»` (in vault + /a0/usr/.env) |
|
||||||
|
| **OpenRouter User** | `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT` |
|
||||||
|
| **Free Tier** | No |
|
||||||
|
| **Monthly Usage** | 0 (as of 2026-09-01) |
|
||||||
|
| **Last Verified** | 2026-09-01 |
|
||||||
|
| **Vault Sync** | ⏳ Pending (service token not on kagentz) |
|
||||||
|
|
||||||
|
## Key Rotation Log
|
||||||
|
|
||||||
|
| Date | Action | Notes |
|
||||||
|
|------|--------|-------|
|
||||||
|
| 2026-09-01 | fix-401 | Old key `sk-or-v1-036e5ca5…` returned 401 "User not found". Replaced with new key `sk-or-v1-0af3f3…` for user `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`. Verified OpenRouter 200. Container .env updated, run_ui restarted. |
|
||||||
|
|
||||||
|
## Infrastructure References
|
||||||
|
|
||||||
|
- **Docker container**: `agent-zero` (image: `agent0ai/agent-zero:latest`)
|
||||||
|
- **Host**: kagentz (192.168.68.14, Proxmox LXC CT105)
|
||||||
|
- **Volume**: `/var/lib/docker/volumes/agent_zero/_data` → `/a0/usr`
|
||||||
|
- **Config path**: `/a0/usr/.env` (line ~72: `API_KEY_OPENROUTER=…`)
|
||||||
|
- **Model preset**: "Cost Efficient" (uses `openrouter/moonshotai/kimi-k3`)
|
||||||
|
- **Model config**: `/a0/usr/plugins/_model_config/config.json`
|
||||||
|
|
||||||
|
## Verification Before Acting
|
||||||
|
|
||||||
|
**Key is a lead, not a fact.** Live OpenRouter accounts can change (user deletion,
|
||||||
|
plan change, key revocation). Before acting on this contract:
|
||||||
|
|
||||||
|
1. Verify the key against OpenRouter's `/auth/key` endpoint
|
||||||
|
2. Check the user ID matches the expected account
|
||||||
|
3. Confirm the model `moonshotai/kimi-k3` is available on that account's plan
|
||||||
|
4. Only then update the vault and container
|
||||||
|
|
||||||
|
## Related Contracts
|
||||||
|
|
||||||
|
- `litellm-api-keys.prose.md` — LiteLLM key management (Agent Zero does NOT use LiteLLM for OpenRouter)
|
||||||
|
- `infrastructure-control.prose.md` — Proxmox topology, container locations
|
||||||
|
- `gpu-fleet.prose.md` — Fleet-wide agent key inventory (add Agent Zero here)
|
||||||
@@ -45,9 +45,10 @@ description: >
|
|||||||
## Status
|
## Status
|
||||||
|
|
||||||
**Active** — for Hermes agents only. This plugin is NOT retired. It remains in service
|
**Active** — for Hermes agents only. This plugin is NOT retired. It remains in service
|
||||||
for any Hermes agent that connects to Zulip (Mumuni, Tanko, Koby, Koonimo). The pi
|
for Hermes agents that connect to Zulip (Mumuni, Koby, Koonimo). The pi
|
||||||
Zulip extension was decommissioned 2026-07-04 but this contract targets the Hermes
|
Zulip extension was decommissioned 2026-07-04 but this contract targets the Hermes
|
||||||
plugin system, which is unaffected.
|
plugin system, which is unaffected. **Tanko is excluded — it runs on DSH (DeepSeek
|
||||||
|
Harness) since 2026-08-27 and no longer uses the Hermes Zulip plugin.**
|
||||||
|
|
||||||
## Parameters
|
## Parameters
|
||||||
|
|
||||||
|
|||||||
+74
-4
@@ -628,18 +628,18 @@ contracts:
|
|||||||
sensitivity: high
|
sensitivity: high
|
||||||
status: active
|
status: active
|
||||||
owner: abiba
|
owner: abiba
|
||||||
version: 3.0.0
|
version: 3.2.0
|
||||||
trigger:
|
trigger:
|
||||||
type: scheduled
|
type: scheduled
|
||||||
cadence: '*/15 * * * *'
|
cadence: '*/15 * * * *'
|
||||||
description: "Every 15 minutes \u2014 monitors all Zulip-connected agents"
|
description: "Every 15 minutes \u2014 monitors the Zulip-connected agents under this host's control (pi, DSH, Agent Zero)"
|
||||||
cron_job_id: null
|
cron_job_id: null
|
||||||
execution:
|
execution:
|
||||||
agent: abiba
|
agent: abiba
|
||||||
timeout: 120
|
timeout: 120
|
||||||
requires:
|
requires:
|
||||||
- Zulip API key for abiba-bot@chat.sysloggh.net
|
- Zulip API key for abiba-bot@chat.sysloggh.net
|
||||||
- SSH access to all Hermes agents
|
- SSH access to amdpve (192.168.68.15) for Tanko (CT 112) and the Agent Zero Docker host (.14)
|
||||||
verification:
|
verification:
|
||||||
postconditions:
|
postconditions:
|
||||||
- check: bot registration active
|
- check: bot registration active
|
||||||
@@ -1155,7 +1155,7 @@ contracts:
|
|||||||
type: scheduled
|
type: scheduled
|
||||||
cadence: 0 3 * * *
|
cadence: 0 3 * * *
|
||||||
description: Daily at 3am ET
|
description: Daily at 3am ET
|
||||||
cron_job_id: null
|
cron_job_id: b59f3cc21f4c # provisioned on kagentz 2026-09-08 (okyeame-memory-audit, glm-5.3-flash)
|
||||||
execution:
|
execution:
|
||||||
agent: mumuni
|
agent: mumuni
|
||||||
timeout: 600
|
timeout: 600
|
||||||
@@ -1867,3 +1867,73 @@ contracts:
|
|||||||
last_run: null
|
last_run: null
|
||||||
last_status: null
|
last_status: null
|
||||||
drift_alerts: []
|
drift_alerts: []
|
||||||
|
# Koby Report-Only Registry (2026-08-17 — Captain)
|
||||||
|
# ⛔ KOBY IS NEVER REPAIRED — detect + report, never fix on .129
|
||||||
|
koby_report_only: true
|
||||||
|
koby_host: "CT 111 (tdunna)"
|
||||||
|
koby_ip: ".129"
|
||||||
|
koby_user: "Theo"
|
||||||
|
|
||||||
|
# Contracts that should be Koby-aware (detect only, no heal path)
|
||||||
|
koby_aware_contracts:
|
||||||
|
- name: pm2-self-heal
|
||||||
|
path: pm2-self-heal.prose.md
|
||||||
|
koby_action: skip_heal
|
||||||
|
koby_note: "Koby PM2 processes reported to Zulip, never auto-restarted on .129"
|
||||||
|
|
||||||
|
- name: zulip-health
|
||||||
|
path: zulip-health.prose.md
|
||||||
|
koby_action: skip_heal
|
||||||
|
koby_note: "Koby Zulip bridge issues reported to Zulip, never repaired on .129"
|
||||||
|
|
||||||
|
- name: hermes-zulip-restore
|
||||||
|
path: hermes-zulip-restore.prose.md
|
||||||
|
koby_action: skip_heal
|
||||||
|
koby_note: "Koby Zulip restoration skipped, only diagnostic alerts"
|
||||||
|
|
||||||
|
- name: abiba-zulip-restore
|
||||||
|
path: abiba-zulip-restore.prose.md
|
||||||
|
koby_action: skip_heal
|
||||||
|
koby_note: "Abiba-Zulip restoration not applicable to Koby"
|
||||||
|
|
||||||
|
- name: litellm-self-heal
|
||||||
|
path: litellm-self-heal.prose.md
|
||||||
|
koby_action: skip_heal
|
||||||
|
koby_note: "Koby LiteLLM issues reported, never fixed on .129"
|
||||||
|
|
||||||
|
- name: disk-gc-threat-response
|
||||||
|
path: disk-gc-threat-response.prose.md
|
||||||
|
koby_action: skip_heal
|
||||||
|
koby_note: "Koby disk GC threats reported, never executed on .129"
|
||||||
|
|
||||||
|
- name: memory-fixer
|
||||||
|
path: memory-fixer.prose.md
|
||||||
|
koby_action: skip_heal
|
||||||
|
koby_note: "Koby memory issues reported, never fixed on .129"
|
||||||
|
|
||||||
|
- name: memory-audit-maintenance
|
||||||
|
path: memory-audit-maintenance.prose.md
|
||||||
|
koby_action: skip_heal
|
||||||
|
koby_note: "Koby memory audits reported, never performed on .129"
|
||||||
|
|
||||||
|
- name: gpu-self-heal
|
||||||
|
path: gpu-self-heal.prose.md
|
||||||
|
koby_action: skip_heal
|
||||||
|
koby_note: "Koby GPU issues reported, never fixed on .129"
|
||||||
|
|
||||||
|
- name: gpu-monitor
|
||||||
|
path: gpu-monitor.prose.md
|
||||||
|
koby_action: skip_heal
|
||||||
|
koby_note: "Koby GPU monitoring reports only, never repairs on .129"
|
||||||
|
|
||||||
|
- name: agent-health-check
|
||||||
|
path: agent-health-check.prose.md
|
||||||
|
koby_action: skip_heal
|
||||||
|
koby_note: "Koby agent health checks reported, never repairs on .129"
|
||||||
|
|
||||||
|
# Scripts that should skip Koby
|
||||||
|
koby_aware_scripts:
|
||||||
|
- name: agent-health-check.py
|
||||||
|
path: scripts/agent-health-check.py
|
||||||
|
koby_action: skip_heal
|
||||||
|
koby_note: "Script should only run diagnostics on Koby, not repairs"
|
||||||
|
|||||||
@@ -1,4 +1,6 @@
|
|||||||
---
|
---
|
||||||
|
report_only_agents:
|
||||||
|
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
|
||||||
kind: responsibility
|
kind: responsibility
|
||||||
name: disk-gc-threat-response
|
name: disk-gc-threat-response
|
||||||
description: >
|
description: >
|
||||||
@@ -11,6 +13,7 @@ description: >
|
|||||||
id: 067NV8KJ03ZG71S44N41F31022
|
id: 067NV8KJ03ZG71S44N41F31022
|
||||||
version: 1.0.0
|
version: 1.0.0
|
||||||
---
|
---
|
||||||
|
---
|
||||||
|
|
||||||
# Disk GC & Threat Response
|
# Disk GC & Threat Response
|
||||||
|
|
||||||
@@ -301,4 +304,4 @@ one-off GPU builds. No automated post-migration cleanup was in place.
|
|||||||
|------|-----|------|--------|
|
|------|-----|------|--------|
|
||||||
| docker-vm | 192.168.68.7 | 16 Docker containers, 4 stacks | ✅ reachable |
|
| docker-vm | 192.168.68.7 | 16 Docker containers, 4 stacks | ✅ reachable |
|
||||||
|
|
||||||
> **Note:** CT 118 is now jdownloader (active on storepve). CT 119 (infisical-vault) added on minipve.\n> **Migrated:** CT 101 → .8, CT 103 → .110 (bare metal GPU).\n> **KVM VM:** CT 109 (docker-vm) is a KVM VM, not LXC — access via SSH .7.
|
> **Note:** CT 118 is now jdownloader (active on storepve). CT 119 (infisical-vault) added on minipve.\n> **Migrated:** CT 101 → .8, CT 103 → .110 (bare metal GPU).\n> **KVM VM:** CT 109 (docker-vm) is a KVM VM, not LXC — access via SSH .7.
|
||||||
|
|||||||
@@ -211,6 +211,40 @@ Errors tell the operator what went wrong and what to do about it. Be specific:
|
|||||||
Check router /health/unified at http://192.168.68.116/health/unified instead."
|
Check router /health/unified at http://192.168.68.116/health/unified instead."
|
||||||
```
|
```
|
||||||
|
|
||||||
|
### Report provenance
|
||||||
|
|
||||||
|
Every report a contract produces must lead with the **absolute path the probe
|
||||||
|
executed from** — `pwd -P`, or the running script's absolute path. A live-state
|
||||||
|
report without provenance is unactionable: a report from a stale copy (a worktree
|
||||||
|
clone, a retired cron entry, a diverged consumer) looks identical to a live
|
||||||
|
fault, and the team burns rounds repairing healthy infrastructure. This is not
|
||||||
|
optional. The 2026-09-09 probe-drift rounds cost three false `DEGRADED` reports
|
||||||
|
because a stale consumer probed the wrong port and nothing in the report said
|
||||||
|
where it ran.
|
||||||
|
|
||||||
|
Pair it with the **scoped any-HTTP-response liveness rule**: for unauthenticated
|
||||||
|
or auth-gated endpoints — where any HTTP answer proves a listener is up (the
|
||||||
|
PVE API's `401`, LiteLLM health's `301` redirect) — a probe is ALIVE on ANY HTTP
|
||||||
|
status, including `301` redirects and `401`/`403` auth challenges. **DOWN =
|
||||||
|
connection refused (`000`) or timeout only.**
|
||||||
|
|
||||||
|
Probes whose success condition is specifically a bare `200` are NOT covered by
|
||||||
|
the any-HTTP rule. On those — authenticated probes such as the Zulip message
|
||||||
|
POST and the router `/health` — an unexpected status (`401`/`403` from a bad or
|
||||||
|
missing credential, `5xx`, or anything other than the expected `200`) is an
|
||||||
|
**ALERT**, not "alive".
|
||||||
|
|
||||||
|
```markdown
|
||||||
|
**Report format**: Begin every report with the absolute execution path
|
||||||
|
(`pwd -P` / script path). On auth-gated endpoints, alive = ANY HTTP status and
|
||||||
|
DOWN = `000`/timeout only; on probes whose expected result is `200`, any other
|
||||||
|
status is an alert.
|
||||||
|
```
|
||||||
|
|
||||||
|
The lint pipeline enforces the provenance clause: any contract with a
|
||||||
|
`**Report format**` line must state an absolute path (`pwd -P`, `absolute path`,
|
||||||
|
or `executed from`).
|
||||||
|
|
||||||
### Comments
|
### Comments
|
||||||
|
|
||||||
Comments in contracts explain WHY, not WHAT. The execution steps say what to
|
Comments in contracts explain WHY, not WHAT. The execution steps say what to
|
||||||
|
|||||||
@@ -0,0 +1,278 @@
|
|||||||
|
# Probe-drift round 2 — per-leg before/after evidence
|
||||||
|
|
||||||
|
**Date:** 2026-09-10
|
||||||
|
**Worktree (absolute execution path):** `/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts`
|
||||||
|
**Branch:** `fm/probe-drift-round2-20260909`
|
||||||
|
|
||||||
|
Every command below was run from the absolute path above; output is pasted
|
||||||
|
verbatim. This is the evidence trail for the four scoped corrections; it is not
|
||||||
|
a contract (never `prose run` it).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Leg 1 — agent-health-check (item 1)
|
||||||
|
|
||||||
|
**Before** — from `/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts`,
|
||||||
|
`python3 scripts/agent-health-check.py --no-deploy` (v2, base of this branch):
|
||||||
|
|
||||||
|
```
|
||||||
|
🏥 Agent Health Check v2 — 2026-09-10 01:16 UTC
|
||||||
|
|
||||||
|
🔑 LiteLLM Keys:
|
||||||
|
✅ tanko: key valid → syslog-auto
|
||||||
|
✅ abiba: key valid → syslog-auto
|
||||||
|
✅ koby: key valid → syslog-auto
|
||||||
|
✅ koonimo: key valid → syslog-auto
|
||||||
|
|
||||||
|
🎮 GPU Port Health:
|
||||||
|
✅ gpu-rtx3090 (.8): healthy (pid=472206)
|
||||||
|
✅ gpu-rtx5070 (.110): healthy (pid=207601)
|
||||||
|
✅ gpu-strixhalo (.15): healthy (pid=4098872)
|
||||||
|
|
||||||
|
🤖 Agent Gateways:
|
||||||
|
✅ tanko: DSH (DeepSeek Harness) — no Hermes gateway since 2026-08-27 (CT 112, SSH OK)
|
||||||
|
⚠️ abiba: gw=no-state-file zulip=? streaming=no errors_10m=0 pid=?
|
||||||
|
✅ koby: gw=running zulip=connected streaming=no errors_10m=0 pid=360900
|
||||||
|
✅ koonimo: gw=running zulip=connected streaming=no errors_10m=0 pid=155125
|
||||||
|
|
||||||
|
🖥️ CT Liveness:
|
||||||
|
✅ tanko (CT 112 on amdpve): running
|
||||||
|
✅ abiba (CT 100 on minipve): running
|
||||||
|
❌ koby (CT 111 on amdpve): PVE UNREACHABLE
|
||||||
|
✅ koonimo (CT 113 on amdpve): running
|
||||||
|
|
||||||
|
📝 Config Integrity:
|
||||||
|
⏭️ tanko: DSH — no Hermes config.yaml since 2026-08-27
|
||||||
|
✅ abiba: config.yaml valid YAML
|
||||||
|
✅ koby: config.yaml valid YAML
|
||||||
|
✅ koonimo: config.yaml valid YAML
|
||||||
|
|
||||||
|
🔌 Wrapper/CLI Integrity:
|
||||||
|
⏭️ tanko: DSH — no hermes CLI wrapper since 2026-08-27
|
||||||
|
⚠️ abiba: wrapper infisical path may be wrong (infisical at /usr/bin/infisical)
|
||||||
|
❌ abiba: hermes-real NOT FOUND (wrapper broken)
|
||||||
|
⚠️ abiba: .env may be missing LITELLM_API_KEY entry
|
||||||
|
⚠️ koby: wrapper infisical path may be wrong (infisical at /usr/bin/infisical)
|
||||||
|
❌ koby: hermes-real NOT FOUND (wrapper broken)
|
||||||
|
✅ koby: wrapper + .env key present
|
||||||
|
⚠️ koonimo: wrapper infisical path may be wrong (infisical at /usr/bin/infisical)
|
||||||
|
✅ koonimo: wrapper + .env key present
|
||||||
|
|
||||||
|
🔐 Vault Secrets:
|
||||||
|
✅ tanko: vault TANKO_LITELLM_API_KEY=sk-...x6uw
|
||||||
|
✅ koby: vault KOBY_LITELLM_API_KEY=sk-...jxlg
|
||||||
|
✅ koonimo: vault KOONIMO_LITELLM_API_KEY=sk-...Y0KQ
|
||||||
|
|
||||||
|
❌ 6 FAILURE(S): ct-unreachable:koby:192.168.68.15 | wrapper-infisical-path:abiba | wrapper-no-hermes-real:abiba | wrapper-infisical-path:koby | wrapper-no-hermes-real:koby | wrapper-infisical-path:koonimo
|
||||||
|
```
|
||||||
|
|
||||||
|
Root causes (all stale expectations; no live fault):
|
||||||
|
|
||||||
|
| Failure | Why it was stale |
|
||||||
|
|---|---|
|
||||||
|
| `ct-unreachable:koby:192.168.68.15` | CT 111 (tdunna/koby) runs on **storepve (.6)**, not amdpve (.15). |
|
||||||
|
| `wrapper-*:abiba` | Abiba is pi-only since the harness purge. `/root/.local/bin/hermes` is a dangling symlink; no `hermes-real`, no `~/.hermes/.env`. |
|
||||||
|
| `wrapper-*:koby` | Koby is **report-only** (captain ruling 2026-08-17, Rule 17): detect and report, never repair — its legs must not count as fleet failures. Koby's wrapper is also the genuine no-infisical case: it sources `~/.hermes/.env` rather than `/usr/bin/infisical`, which the check now accepts. |
|
||||||
|
| `wrapper-infisical-path:koonimo` | Koonimo's wrapper **does** reference `/usr/bin/infisical` — but past the old check's `head -20` window, so the check looked for the path in the wrong slice and false-failed. The fix that mattered was reading the full wrapper body (and then verifying any absolute infisical path it finds actually exists). |
|
||||||
|
|
||||||
|
**After** — same absolute path, `python3 scripts/agent-health-check.py --no-deploy` (v4):
|
||||||
|
|
||||||
|
```
|
||||||
|
$ pwd -P
|
||||||
|
/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts
|
||||||
|
$ python3 scripts/agent-health-check.py --no-deploy
|
||||||
|
🏥 Agent Health Check v4 — 2026-09-10 01:22 UTC
|
||||||
|
📍 executed from: script=/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts/scripts/agent-health-check.py cwd=/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts
|
||||||
|
|
||||||
|
🔑 LiteLLM Keys:
|
||||||
|
✅ tanko: key valid → syslog-auto
|
||||||
|
✅ abiba: key valid → syslog-auto
|
||||||
|
✅ koby: key valid → syslog-auto
|
||||||
|
✅ koonimo: key valid → syslog-auto
|
||||||
|
|
||||||
|
🎮 GPU Port Health:
|
||||||
|
✅ gpu-rtx3090 (.8): healthy (pid=472206)
|
||||||
|
✅ gpu-rtx5070 (.110): healthy (pid=207601)
|
||||||
|
✅ gpu-strixhalo (.15): healthy (pid=4098872)
|
||||||
|
|
||||||
|
🤖 Agent Gateways:
|
||||||
|
✅ tanko: DSH (DeepSeek Harness) — no Hermes gateway since 2026-08-27 (CT 112, SSH OK)
|
||||||
|
✅ abiba: pi-only runtime — no Hermes gateway since the harness purge (CT 100, SSH OK)
|
||||||
|
🔍 koby: REPORT-ONLY mode (diagnostic only, no repairs on .129)
|
||||||
|
✅ koby: gateway running (pid=360900, report-only mode)
|
||||||
|
✅ koonimo: gw=running zulip=connected streaming=no errors_10m=0 pid=155125
|
||||||
|
|
||||||
|
🖥️ CT Liveness:
|
||||||
|
✅ tanko (CT 112 on amdpve): running
|
||||||
|
✅ abiba (CT 100 on minipve): running
|
||||||
|
✅ koby (CT 111 on storepve): running
|
||||||
|
✅ koonimo (CT 113 on amdpve): running
|
||||||
|
|
||||||
|
📝 Config Integrity:
|
||||||
|
⏭️ tanko: DSH — no Hermes config.yaml since 2026-08-27
|
||||||
|
⏭️ abiba: pi-only runtime — no Hermes config.yaml since the harness purge
|
||||||
|
✅ koby: config.yaml valid YAML
|
||||||
|
✅ koonimo: config.yaml valid YAML
|
||||||
|
|
||||||
|
🔌 Wrapper/CLI Integrity:
|
||||||
|
⏭️ tanko: DSH — no hermes CLI wrapper since 2026-08-27
|
||||||
|
⏭️ abiba: pi-only runtime — no hermes CLI wrapper since the harness purge
|
||||||
|
ℹ️ koby: wrapper resolves creds without infisical (e.g. ~/.hermes/.env) — OK
|
||||||
|
❌ koby: hermes-real NOT FOUND (wrapper broken)
|
||||||
|
🔍 report-only (koby): wrapper-no-hermes-real:koby — reported, not counted/repaired
|
||||||
|
✅ koby: wrapper + .env key present
|
||||||
|
✅ koonimo: wrapper infisical path OK
|
||||||
|
✅ koonimo: wrapper + .env key present
|
||||||
|
|
||||||
|
🔐 Vault Secrets:
|
||||||
|
✅ tanko: vault TANKO_LITELLM_API_KEY=sk-...x6uw
|
||||||
|
✅ koby: vault KOBY_LITELLM_API_KEY=sk-...jxlg
|
||||||
|
✅ koonimo: vault KOONIMO_LITELLM_API_KEY=sk-...Y0KQ
|
||||||
|
|
||||||
|
✅ All checks passed
|
||||||
|
exit=0
|
||||||
|
```
|
||||||
|
|
||||||
|
**Live vantage proof** (same worktree):
|
||||||
|
|
||||||
|
```
|
||||||
|
$ ssh root@192.168.68.15 "pct status 111"
|
||||||
|
Configuration file 'nodes/amdpve/lxc/111.conf' does not exist
|
||||||
|
$ ssh root@192.168.68.6 "pct status 111; pct list | grep '^ *111'"
|
||||||
|
status: running
|
||||||
|
111 running tdunna
|
||||||
|
$ ssh root@192.168.68.129 "hostname"
|
||||||
|
tdunna
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Leg 2 — infrastructure-monitoring PVE API (item 2)
|
||||||
|
|
||||||
|
**Before** — the contract's probe, aimed at the monitoring host CT 116:
|
||||||
|
|
||||||
|
```
|
||||||
|
$ curl -s -o /dev/null -w '%{http_code}' https://192.168.68.116:8006/api2/json
|
||||||
|
000
|
||||||
|
```
|
||||||
|
|
||||||
|
CT 116 runs no `pveproxy`, so it never answers on `:8006`. The probe target was
|
||||||
|
wrong, which is what read as PVE-API `000`.
|
||||||
|
|
||||||
|
**After** — probing the five real cluster nodes (`:8006/api2/json/version`),
|
||||||
|
alive under the any-HTTP-response rule (`401` = up, unauthenticated):
|
||||||
|
|
||||||
|
```
|
||||||
|
$ for node in 192.168.68.9 192.168.68.5 192.168.68.15 192.168.68.6 192.168.68.12; do
|
||||||
|
printf '%s:8006 -> %s\n' "$node" "$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 5 "https://$node:8006/api2/json/version")"
|
||||||
|
done
|
||||||
|
192.168.68.9:8006 -> 401
|
||||||
|
192.168.68.5:8006 -> 401
|
||||||
|
192.168.68.15:8006 -> 401
|
||||||
|
192.168.68.6:8006 -> 401
|
||||||
|
192.168.68.12:8006 -> 401
|
||||||
|
```
|
||||||
|
|
||||||
|
`401` on every node = alive by design. `DOWN` is `000`/timeout only. (The
|
||||||
|
contract's LiteLLM probe was the same class: `/litellm/health` answers `301` →
|
||||||
|
`/litellm/health/liveliness`, so it is now specified as any-HTTP too.)
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Leg 3 — gpu-monitor GPU probes (item 3)
|
||||||
|
|
||||||
|
**Before** — the false alarm came from probing bare port 80 on GPU hosts:
|
||||||
|
|
||||||
|
```
|
||||||
|
$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.8/health
|
||||||
|
000
|
||||||
|
$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.110/health
|
||||||
|
000
|
||||||
|
```
|
||||||
|
|
||||||
|
Nothing listens on GPU port 80, so the monitor reported
|
||||||
|
`DEGRADED — GPU-rtx3090 000, GPU-rtx5070 000` three times on 2026-09-09.
|
||||||
|
|
||||||
|
**After** — the real endpoints answer:
|
||||||
|
|
||||||
|
```
|
||||||
|
$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.8:8080/health
|
||||||
|
200
|
||||||
|
$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.110:8080/health
|
||||||
|
200
|
||||||
|
$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.15:8080/health
|
||||||
|
200
|
||||||
|
$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health/unified
|
||||||
|
301 # Location: http://192.168.68.116/gpu/gpu-data — the same payload
|
||||||
|
$ curl -s -o /dev/null -w '%{http_code}' -L http://192.168.68.116/health/unified
|
||||||
|
200
|
||||||
|
```
|
||||||
|
|
||||||
|
`301` is healthy under the any-HTTP-response rule. The contract now requires GPU
|
||||||
|
health on `:8080` (or router `/health/unified`) and forbids bare port 80 on a
|
||||||
|
GPU host.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Leg 4 — report provenance (item 4)
|
||||||
|
|
||||||
|
Every contract report must now lead with the absolute path it executed from.
|
||||||
|
`docs/AUTHORING-GUIDE.md` documents the rule and `scripts/prose-lint.sh`
|
||||||
|
enforces it:
|
||||||
|
|
||||||
|
```
|
||||||
|
$ pwd -P
|
||||||
|
/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts
|
||||||
|
$ bash scripts/prose-lint.sh
|
||||||
|
✅ Report provenance present in all report-format contracts
|
||||||
|
...
|
||||||
|
✅ LINT PASSED (12 warning(s))
|
||||||
|
```
|
||||||
|
|
||||||
|
The health script prints `📍 executed from: script=… cwd=…` and includes
|
||||||
|
`execution_path`/`cwd` in `--json` output.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Full suite
|
||||||
|
|
||||||
|
```
|
||||||
|
$ pwd -P
|
||||||
|
/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts
|
||||||
|
$ python3 -m pytest -q
|
||||||
|
24 passed
|
||||||
|
$ shellcheck scripts/prose-lint.sh
|
||||||
|
(clean)
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Follow-up findings (observed, intentionally NOT changed here)
|
||||||
|
|
||||||
|
These are adjacent stale expectations discovered while verifying the four
|
||||||
|
scoped legs. Each touches a CRITICAL/HIGH-sensitivity artifact or an unrelated
|
||||||
|
script, so it is recorded for the captain/verify mate rather than silently
|
||||||
|
repaired.
|
||||||
|
|
||||||
|
1. **`infrastructure-control.prose.md` (CRITICAL) CT 111 node assignment.**
|
||||||
|
Lines ~109 and ~615 place `tdunna` (CT 111, koby) on **amdpve**. Live
|
||||||
|
verification on 2026-09-10 shows `pct status 111` = `running` on
|
||||||
|
**storepve (.6)** and `Configuration file 'nodes/amdpve/lxc/111.conf' does
|
||||||
|
not exist` on .15. `agent-health-check.py` now carries the live-verified
|
||||||
|
`storepve` mapping (the script is not the topology source of truth); the
|
||||||
|
CRITICAL contract itself needs an authorized correction.
|
||||||
|
2. **Strix Halo `:8080` firewall claim is stale.** `prose-ai-review.sh`
|
||||||
|
ground-truth rule #4 and `gpu-monitor.prose.md` say `:8080` is firewalled to
|
||||||
|
`.116` only and `.24` cannot probe it. Live on .15:
|
||||||
|
`-A INPUT -s 192.168.68.24/32 -p tcp --dport 8080 -j ACCEPT`, and a probe
|
||||||
|
from .24 returns `200`. The contract keeps routing Strix via the router
|
||||||
|
(safe), but the claim no longer matches iptables.
|
||||||
|
3. **`contract-registry.yaml` references `agent-health-check.prose.md`**, which
|
||||||
|
does not exist in the repo. The registry entry (with `koby_action: skip_heal`)
|
||||||
|
is aspirational/stale.
|
||||||
|
4. **Pre-existing script defects, untouched:** `scripts/pm2-self-heal.sh` has a
|
||||||
|
bash syntax error at lines 19–20 (`bash -n` fails), and `shellcheck` fails on
|
||||||
|
five untouched scripts (`netbird-add-domain.sh`, `pct-run.sh`,
|
||||||
|
`pm2-self-heal.sh`, `prose-ai-review.sh`, `swap-gpu-dense-model.sh`).
|
||||||
|
`scripts/prose-lint.sh` — the one shell file touched here — is now
|
||||||
|
shellcheck-clean.
|
||||||
+9
-28
@@ -14,9 +14,6 @@ description: >
|
|||||||
Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling.
|
Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling.
|
||||||
For larger context needs → fall back to external providers (deepseek).
|
For larger context needs → fall back to external providers (deepseek).
|
||||||
VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%.
|
VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%.
|
||||||
UPDATED 2026-08-15: gpu-dense swapped to Qwen3.8-27B-Uncensored-Q4_K_M
|
|
||||||
(~16.8GB, 128K ctx, --spec-type draft-mtp (v2)). Served on .8:8080 under the legacy
|
|
||||||
alias qwen3.6-27B-code for LiteLLM routing continuity. Replaces SmartCode-Fable-5-27B.
|
|
||||||
agent: abiba
|
agent: abiba
|
||||||
triggers:
|
triggers:
|
||||||
- on model add/remove
|
- on model add/remove
|
||||||
@@ -74,7 +71,7 @@ triggers:
|
|||||||
│ RTX 3090 │ │ RTX 5070 │ │ Strix Halo│ │ GPU Monitor │
|
│ RTX 3090 │ │ RTX 5070 │ │ Strix Halo│ │ GPU Monitor │
|
||||||
│ 24GB │ │ 12GB │ │ 64GB UMA │ │ :9100 │
|
│ 24GB │ │ 12GB │ │ 64GB UMA │ │ :9100 │
|
||||||
│ 128K ctx │ │ 128K ctx │ │ 128K ctx │ │ Watchdog │
|
│ 128K ctx │ │ 128K ctx │ │ 128K ctx │ │ Watchdog │
|
||||||
│ qwen3.6 │ │ gemma-4-12b │ │ qwen3.6 │ │ Prometheus │
|
│ qwen3.6 │ │ Qwen3.5-9B │ │ qwen3.5 │ │ Prometheus │
|
||||||
│ 27B-code │ │ :8080 │ │ -35B-udq4 │ │ exporter │
|
│ 27B-code │ │ :8080 │ │ -35B-udq4 │ │ exporter │
|
||||||
│ :8080 │ │ :9400 (exp) │ │ :9400(exp)│ │ :9401 │
|
│ :8080 │ │ :9400 (exp) │ │ :9400(exp)│ │ :9401 │
|
||||||
│ :9400 │ └─────────────┘ └───────────┘ └──────────────┘
|
│ :9400 │ └─────────────┘ └───────────┘ └──────────────┘
|
||||||
@@ -92,16 +89,13 @@ When a model is swapped on a GPU, ONLY the infrastructure layer changes — agen
|
|||||||
| `gpu-dense` | RTX 3090 (.8) | Qwen3.8-27B-Uncensored-Q4_K_M | Whatever runs on RTX 3090 |
|
| `gpu-dense` | RTX 3090 (.8) | Qwen3.8-27B-Uncensored-Q4_K_M | Whatever runs on RTX 3090 |
|
||||||
| `gpu-light` | RTX 5070 (.110) | gemma-4-12b | Whatever runs on RTX 5070 |
|
| `gpu-light` | RTX 5070 (.110) | gemma-4-12b | Whatever runs on RTX 5070 |
|
||||||
|
|
||||||
**Backward compatibility**: Old model-specific names (qwen3.6-27B-code, gemma-4-12b, qwen3.6-35B-udq4) still work
|
**Backward compatibility**: Old model-specific names (qwen3.6-27B-code, gemma-4-12b, qwen3.5-9b-it) still work
|
||||||
but are deprecated for agent configs. Only the stable aliases survive model swaps.
|
but are deprecated for agent configs. Only the stable aliases survive model swaps.
|
||||||
|
|
||||||
## Current Model Assignments (2026-07-15)
|
## Current Model Assignments (2026-07-15)
|
||||||
|
|
||||||
| Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status |
|
| Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status |
|
||||||
|-------|-----|------|------|-----|----------|----------|-------------|--------|
|
|-------|-----|------|------|-----|----------|----------|-------------|--------|
|
||||||
| Qwen3.8-27B-Uncensored-Q4_K_M | RTX 3090 | .8 (llm-gpu) | ~16.8/24.6GB | **128K** | turbo4 | 1 | 2048/1024 | ✅ healthy |
|
|
||||||
| gemma-4-12b | RTX 5070 | .110 (ocu-llm) | ~7.8/12.2GB (65%) | 128K | q4_0 | 2 | 2048/1024 | ✅ healthy |
|
|
||||||
| qwen3.6-35B-udq4 | Strix Halo Vulkan | .15 (amdpve) | ~22GB/64GB | 128K | q4_0 | 1 | 4096/1024 | ✅ 65 tok/s |
|
|
||||||
|
|
||||||
## Routing Configuration (LiteLLM — July 2026)
|
## Routing Configuration (LiteLLM — July 2026)
|
||||||
|
|
||||||
@@ -110,8 +104,6 @@ but are deprecated for agent configs. Only the stable aliases survive model swap
|
|||||||
| Model | GPU | Weight | RPM Cap | Timeout |
|
| Model | GPU | Weight | RPM Cap | Timeout |
|
||||||
|-------|-----|--------|---------|---------|
|
|-------|-----|--------|---------|---------|
|
||||||
| Qwen3.8-27B-Uncensored-Q4_K_M | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** |
|
| Qwen3.8-27B-Uncensored-Q4_K_M | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** |
|
||||||
| qwen3.6-35B-udq4 | Strix Halo (.15:8080) | **0.30** | 60 | **300s** |
|
|
||||||
| gemma-4-12b | RTX 5070 (.110:8080) | **0.15** | 200 | **120s** |
|
|
||||||
|
|
||||||
Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) is NOT in the inference path.
|
Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) is NOT in the inference path.
|
||||||
|
|
||||||
@@ -119,9 +111,6 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`.
|
|||||||
|
|
||||||
| Model | RPM Cap | Notes |
|
| Model | RPM Cap | Notes |
|
||||||
|-------|---------|-------|
|
|-------|---------|-------|
|
||||||
| strix-moe (qwen3.6-35B-udq4) | 40 | Tight cap — prevents Strix overload |
|
|
||||||
| Qwen3.8-27B-Uncensored-Q4_K_M | 500 | High cap — primary workhorse (replaces qwen3.6-27B-code) |
|
|
||||||
| gemma-4-12b | 500 | High cap — IQ4_NL+MTP, 122 tok/s |
|
|
||||||
|
|
||||||
### Stable Aliases (for agent configs — never change)
|
### Stable Aliases (for agent configs — never change)
|
||||||
|
|
||||||
@@ -192,7 +181,7 @@ Show full fleet status: GPUs, models, VRAM, context windows, parallel slots, act
|
|||||||
3. Check LiteLLM: `curl http://192.168.68.116/health` (expect "I'm alive!")
|
3. Check LiteLLM: `curl http://192.168.68.116/health` (expect "I'm alive!")
|
||||||
4. Check LiteLLM models: `curl -H "Authorization: Bearer $MASTER_KEY" http://192.168.68.116/v1/models`
|
4. Check LiteLLM models: `curl -H "Authorization: Bearer $MASTER_KEY" http://192.168.68.116/v1/models`
|
||||||
5. Check LiteLLM timeouts: `grep -n 'timeout:' /opt/inference-harness/litellm_config.yaml`
|
5. Check LiteLLM timeouts: `grep -n 'timeout:' /opt/inference-harness/litellm_config.yaml`
|
||||||
- gemma-4-12b: 120s, qwen3.6-27B-code: 300s, qwen3.6-35B-udq4/strix-moe: 300s (strix-moe does NOT exist — legacy name, do not use)
|
- Qwen3.5-9B: 120s, qwen3.6-27B-code: 300s, Carnice-Qwen3.6-MoE-35B-A3B/strix-moe: 300s (strix-moe alias retained, legacy name qwen3.6-35B-udq4 deprecated)
|
||||||
- global request_timeout: 300s, nginx proxy_read_timeout: 600s
|
- global request_timeout: 300s, nginx proxy_read_timeout: 600s
|
||||||
6. Check AMD metrics: `curl http://192.168.68.15:9400/metrics` (Radeon 8060S, util%, VRAM, temp, power)
|
6. Check AMD metrics: `curl http://192.168.68.15:9400/metrics` (Radeon 8060S, util%, VRAM, temp, power)
|
||||||
7. Check port conflicts: verify only one llama-server on :8080 per host
|
7. Check port conflicts: verify only one llama-server on :8080 per host
|
||||||
@@ -207,7 +196,7 @@ Plaintext keys removed from this contract post-vault-migration.
|
|||||||
| Agent | CT | IP | LiteLLM Alias | Key Source | Access |
|
| Agent | CT | IP | LiteLLM Alias | Key Source | Access |
|
||||||
|-------|-----|-----|---------------|------------|--------|
|
|-------|-----|-----|---------------|------------|--------|
|
||||||
| Tanko | 112 | .122 | `tanko` | Infisical vault | SSH jerome |
|
| Tanko | 112 | .122 | `tanko` | Infisical vault | SSH jerome |
|
||||||
| Mumuni | 100 (abiba) | .24 | `mumuni` | Infisical vault | SSH root |
|
| Mumuni | 105 (kagentz) | .14 | `mumuni` | Infisical vault | SSH root |
|
||||||
| Abiba | 100 | .24 | `abiba-pi` | Infisical vault | local (pi agent) |
|
| Abiba | 100 | .24 | `abiba-pi` | Infisical vault | local (pi agent) |
|
||||||
| Koby | 111 | ? | `koby` | Infisical vault | Zulip DM |
|
| Koby | 111 | ? | `koby` | Infisical vault | Zulip DM |
|
||||||
| Koonimo | 113 | ? | `koonimo` | Infisical vault (migrated 2026-07-11) | no SSH |
|
| Koonimo | 113 | ? | `koonimo` | Infisical vault (migrated 2026-07-11) | no SSH |
|
||||||
@@ -249,14 +238,6 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
|
|||||||
- **Router startup race**: Compose router.py doesn't call load_roster(). Reload thread sleeps 30s first.
|
- **Router startup race**: Compose router.py doesn't call load_roster(). Reload thread sleeps 30s first.
|
||||||
Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster().
|
Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster().
|
||||||
- **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround.
|
- **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround.
|
||||||
- **VRAM (2026-07-15)**: RTX 3090 at ~17/24.6GB (~70%) with **128K context** (reduced from 256K 2026-07-17). RTX 5070 at ~7.8/12.2GB (~65%) with 128K context + MTP. Strix Halo at ~7GB/64GB.
|
|
||||||
- **RTX 3090 (2026-08-15)**: Swapped to Qwen3.8-27B-Uncensored-Q4_K_M (~16.8GB, replaced
|
|
||||||
SmartCode-Fable-5-27B-UD-Q3_K_XL). Service: `/home/llmuser/llama-fable-wrapper.sh`.
|
|
||||||
Served under alias `qwen3.6-27B-code` for LiteLLM routing continuity.
|
|
||||||
- **RTX 5070 config (2026-07-15)**: Switched to IQ4_NL + MTP draft (Q8_0) at 128K context. Gen speed: 122 tok/s. VRAM: ~7.8/12.2GB (~65%). Service: `/home/llmuser/llama-wrapper.sh`. Config: `--model gemma-4-12b-it-IQ4_NL.gguf --spec-draft-model gemma-4-12b-it-Q8_0-MTP.gguf --spec-type draft-mtp --spec-draft-n-max 4 --ctx-size 131072`.
|
|
||||||
- **LiteLLM timeout tuning (verified 2026-08-16)**: Qwen3.8-27B (alias qwen3.6-27B-code)
|
|
||||||
300s, gemma-4-12b 120s, qwen3.6-35B-udq4 300s, strix-moe 300s, syslog-auto routes all
|
|
||||||
300s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s.
|
|
||||||
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `qwen3.6-35B-udq4`, alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded).
|
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `qwen3.6-35B-udq4`, alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded).
|
||||||
- **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill).
|
- **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill).
|
||||||
- **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (old Mumuni CT114 — now inside Abiba CT100 at .24) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`.
|
- **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (old Mumuni CT114 — now inside Abiba CT100 at .24) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`.
|
||||||
@@ -272,7 +253,6 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
|
|||||||
| GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Context |
|
| GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Context |
|
||||||
|-----|-------|-----------|--------------|----------|---------|
|
|-----|-------|-----------|--------------|----------|---------|
|
||||||
| RTX 3090 (.8) | Qwen3.8-27B-Uncensored-Q4_K_M | **TBD** | — | — | **128K** |
|
| RTX 3090 (.8) | Qwen3.8-27B-Uncensored-Q4_K_M | **TBD** | — | — | **128K** |
|
||||||
| RTX 5070 (.110) | gemma-4-12b (IQ4_NL+MTP) | **191** | — | — | **128K** |
|
|
||||||
| Strix Halo (.15) | qwen3.6-35B-udq4 | **65** | 140 | — | **128K** |
|
| Strix Halo (.15) | qwen3.6-35B-udq4 | **65** | 140 | — | **128K** |
|
||||||
|
|
||||||
Benchmarks from 2026-07-17. Strix Halo model: qwen3.6-35B-udq4. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
|
Benchmarks from 2026-07-17. Strix Halo model: qwen3.6-35B-udq4. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
|
||||||
@@ -299,12 +279,13 @@ When the underlying model is swapped, only the LiteLLM config changes — agent
|
|||||||
### Context Windows
|
### Context Windows
|
||||||
- RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **128K**
|
- RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **128K**
|
||||||
- **All agents**: 128K ceiling — stable margin. For >128K workloads, use external providers (deepseek)
|
- **All agents**: 128K ceiling — stable margin. For >128K workloads, use external providers (deepseek)
|
||||||
- Compression threshold 0.65: fires at ~85K (~43K headroom before 128K ceiling)
|
- Compression threshold 0.60: fires at ~77K (~51K headroom before 128K ceiling)
|
||||||
|
- **Pi agents (Abiba)**: `compaction.reserveTokens: 52739` (≈60% of 128K)
|
||||||
- Mumuni compression model alias: `strix-moe` with 300s timeout
|
- Mumuni compression model alias: `strix-moe` with 300s timeout
|
||||||
|
|
||||||
### Mumuni Agent Profile
|
### Mumuni Agent Profile
|
||||||
|
|
||||||
Mumuni (CT100/abiba, 192.168.68.24) is the primary business assistant. This profile is the reference for all agent configs:
|
Mumuni (kagentz CT105, 192.168.68.14 — migrated from CT100 2026-08-29) is the primary business assistant. This profile is the reference for all agent configs:
|
||||||
|
|
||||||
| Setting | Value | Notes |
|
| Setting | Value | Notes |
|
||||||
|---------|-------|-------|
|
|---------|-------|-------|
|
||||||
@@ -316,7 +297,7 @@ Mumuni (CT100/abiba, 192.168.68.24) is the primary business assistant. This prof
|
|||||||
| `aux.web_extract.model` | `gpu-light` | Web extraction |
|
| `aux.web_extract.model` | `gpu-light` | Web extraction |
|
||||||
| `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) |
|
| `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) |
|
||||||
| `context.max_context_window` | 131072 (128K) | Reduced from 256K 2026-07-17 — stable 128K ceiling |
|
| `context.max_context_window` | 131072 (128K) | Reduced from 256K 2026-07-17 — stable 128K ceiling |
|
||||||
| `compression.threshold` | 0.65 | Triggers at ~85K |
|
| `compression.threshold` | 0.60 | Triggers at ~77K (~60% of 128K) — optimized for 128K context |
|
||||||
| `compression.target_ratio` | 0.3 | Compresses to ~38K |
|
| `compression.target_ratio` | 0.3 | Compresses to ~38K |
|
||||||
| `compression.protect_last_n` | 40 | Preserves last 40 messages |
|
| `compression.protect_last_n` | 40 | Preserves last 40 messages |
|
||||||
| `memory.memory_char_limit` | 800 | Brief memory entries |
|
| `memory.memory_char_limit` | 800 | Brief memory entries |
|
||||||
@@ -329,7 +310,7 @@ Mumuni (CT100/abiba, 192.168.68.24) is the primary business assistant. This prof
|
|||||||
|
|
||||||
| Agent | Host | Status |
|
| Agent | Host | Status |
|
||||||
|-------|------|--------|
|
|-------|------|--------|
|
||||||
| **Mumuni** | CT100 (.24) | ✅ Updated to stable aliases |
|
| **Mumuni** | CT105 (.14) | ✅ Updated to stable aliases |
|
||||||
| **Tanko** | CT112 (.122) | ✅ Updated to stable aliases |
|
| **Tanko** | CT112 (.122) | ✅ Updated to stable aliases |
|
||||||
| **Koby** | CT111 (.129) | ❌ SSH unreachable — needs Zulip DM |
|
| **Koby** | CT111 (.129) | ❌ SSH unreachable — needs Zulip DM |
|
||||||
| **Koonimo** | CT113 | ❌ SSH unreachable — needs Zulip DM |
|
| **Koonimo** | CT113 | ❌ SSH unreachable — needs Zulip DM |
|
||||||
|
|||||||
+90
-16
@@ -31,20 +31,31 @@ agent: abiba
|
|||||||
└──────┘ │qwen27B│ │LiteLLM │
|
└──────┘ │qwen27B│ │LiteLLM │
|
||||||
└──────┘ │dashboard│
|
└──────┘ │dashboard│
|
||||||
└────────┘
|
└────────┘
|
||||||
|
```
|
||||||
|
|
||||||
Note: JSON sidecar exporters at :8090 were never deployed on any
|
Note: JSON sidecar exporters at :8090 were never deployed on any
|
||||||
GPU host. Router falls back to GPU /health direct probe. Monitor
|
GPU host. Router falls back to GPU /health direct probe. Monitor
|
||||||
should use router /health/unified as source of truth for GPU status.
|
should use router /health/unified as source of truth for GPU status.
|
||||||
Strix Halo :8080 is firewalled to .116 only — monitor on .24 cannot
|
Strix Halo :8080 is firewalled to .116 only — monitor on .24 cannot
|
||||||
poll .15:8080 directly; must go through router on .116.
|
poll .15:8080 directly; must go through router on .116.
|
||||||
```
|
|
||||||
|
**PORT RULE (verified 2026-09-10):** GPU per-host health lives on **:8080**
|
||||||
|
(`http://<gpu-host>:8080/health`); Prometheus GPU exporters live on **:9400**.
|
||||||
|
There is NO listener on bare port 80 for any GPU host — `http://192.168.68.8/health`
|
||||||
|
and `http://192.168.68.110/health` answer `000`. Never use a bare-port-80 probe
|
||||||
|
as a GPU liveness signal: on 2026-09-09 that produced three false
|
||||||
|
`DEGRADED — GPU-rtx3090 000, GPU-rtx5070 000` rounds while
|
||||||
|
`http://192.168.68.8:8080/health` and `http://192.168.68.110:8080/health`
|
||||||
|
answered `200`. Port 80 is valid only on the router (.116), never on a GPU host.
|
||||||
|
|
||||||
### Subsystems Polled
|
### Subsystems Polled
|
||||||
|
|
||||||
| Subsystem | Endpoint | Frequency | Metrics |
|
| Subsystem | Endpoint | Frequency | Metrics |
|
||||||
|-----------|----------|-----------|---------|
|
|-----------|----------|-----------|---------|
|
||||||
| GPU Status (all, via router) | `http://192.168.68.116/health/unified` | 15s | models, CB, scores, GPU status (router probes each GPU /health directly) |
|
| GPU Status (all, via router) | `http://192.168.68.116/health/unified` | 15s | models, CB, scores, GPU status (router probes each GPU /health directly) — `301` → `/gpu/gpu-data` is **alive** |
|
||||||
| Router (unified) | `http://192.168.68.116/health/unified` | 15s | models, CB, scores, GPU status |
|
| GPU .8 (RTX 3090) health | `http://192.168.68.8:8080/health` | 15s | direct liveness fallback — **:8080 ONLY, never bare port 80** |
|
||||||
|
| GPU .110 (RTX 5070) health | `http://192.168.68.110:8080/health` | 15s | direct liveness fallback — **:8080 ONLY, never bare port 80** |
|
||||||
|
| Router (unified) | `http://192.168.68.116/health/unified` | 15s | models, CB, scores, GPU status (`301` → `/gpu/gpu-data` = alive) |
|
||||||
| Router (basic) | `http://192.168.68.116/health` | 15s | basic aliveness |
|
| Router (basic) | `http://192.168.68.116/health` | 15s | basic aliveness |
|
||||||
| LiteLLM | `http://192.168.68.116/litellm/health` | 15s | proxy health, model count |
|
| LiteLLM | `http://192.168.68.116/litellm/health` | 15s | proxy health, model count |
|
||||||
| Strix Halo | `http://192.168.68.116/health/unified` (router) | 15s | Strix Halo status via router — cannot poll .15:8080 directly (firewalled to .116 only) |
|
| Strix Halo | `http://192.168.68.116/health/unified` (router) | 15s | Strix Halo status via router — cannot poll .15:8080 directly (firewalled to .116 only) |
|
||||||
@@ -61,6 +72,22 @@ This replaces the previous DM-only delivery. All agents on the mesh can see and
|
|||||||
|
|
||||||
## Alert Thresholds
|
## Alert Thresholds
|
||||||
|
|
||||||
|
### Liveness rule (scoped)
|
||||||
|
|
||||||
|
The any-HTTP-response rule applies ONLY to redirect/auth-gated liveness
|
||||||
|
endpoints, where any HTTP answer proves a listener is up. Applied here: the
|
||||||
|
router's `/health/unified` answers `301 Moved Permanently` → `/gpu/gpu-data`
|
||||||
|
(the same payload) and LiteLLM's `/litellm/health` answers `301` →
|
||||||
|
`/litellm/health/liveliness`. For those endpoints a probe is **ALIVE** on
|
||||||
|
**ANY** HTTP status — `3xx` redirects and `401`/`403` auth challenges included —
|
||||||
|
and **DOWN = connection refused (`000`) or timeout only**. Same scoped rule as
|
||||||
|
zulip-health (Tanko) and infrastructure-monitoring.
|
||||||
|
|
||||||
|
Probes whose success condition is specifically a bare `200` are NOT covered by
|
||||||
|
the any-HTTP rule. On those — the GPU `:8080/health` endpoints, the router
|
||||||
|
`/health`, and the dashboard — an unexpected status (`401`/`403`, `5xx`, or
|
||||||
|
anything other than the expected `200`) is an **ALERT**, not "alive".
|
||||||
|
|
||||||
| Metric | Warning | Critical |
|
| Metric | Warning | Critical |
|
||||||
|--------|---------|----------|
|
|--------|---------|----------|
|
||||||
| GPU Temp | >80°C | >90°C |
|
| GPU Temp | >80°C | >90°C |
|
||||||
@@ -97,7 +124,47 @@ This replaces the previous DM-only delivery. All agents on the mesh can see and
|
|||||||
`curl http://localhost:9100/gpu-data | jq` — Full fleet status
|
`curl http://localhost:9100/gpu-data | jq` — Full fleet status
|
||||||
|
|
||||||
### check-health
|
### check-health
|
||||||
`curl http://localhost:9100/health` — Monitor self-check
|
|
||||||
|
**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Provenance — run first; paste the absolute path into the report
|
||||||
|
pwd -P
|
||||||
|
|
||||||
|
# GPU Monitor health
|
||||||
|
curl http://localhost:9100/health | jq
|
||||||
|
# Expected: 200 with {"status": "healthy", "cache_age_seconds": <n>}
|
||||||
|
|
||||||
|
# GPU host health — DIRECT on :8080. NEVER probe bare port 80 on a GPU host:
|
||||||
|
# http://192.168.68.8/health has no listener and returns 000 → false DEGRADED.
|
||||||
|
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.8:8080/health
|
||||||
|
# Expected: 200 (bare-200 probe: any other status is an alert; 000/timeout = DOWN)
|
||||||
|
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.110:8080/health
|
||||||
|
# Expected: 200 (bare-200 probe: any other status is an alert; 000/timeout = DOWN)
|
||||||
|
|
||||||
|
# Router unified health (source of truth; 301 → /gpu/gpu-data is HEALTHY)
|
||||||
|
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health/unified
|
||||||
|
# Expected: 301 (or 200 after following the redirect) — any HTTP status = alive
|
||||||
|
|
||||||
|
# Router basic health (via nginx on port 80 — router .116 only, never a GPU host)
|
||||||
|
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health
|
||||||
|
# Expected: 200 (bare-200 probe: any other status is an alert; 000/timeout = DOWN)
|
||||||
|
|
||||||
|
# LiteLLM health (via nginx on port 80)
|
||||||
|
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/litellm/health
|
||||||
|
# Expected: 301 → /litellm/health/liveliness (200 after redirect) — any HTTP status = alive
|
||||||
|
|
||||||
|
# Dashboard
|
||||||
|
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/dashboard/
|
||||||
|
# Expected: 200 (bare-200 probe: any other status is an alert; 000/timeout = DOWN)
|
||||||
|
```
|
||||||
|
|
||||||
|
**Report format**: Begin every report with the **absolute path the probe executed
|
||||||
|
from** (`pwd -P`, or the monitor script's absolute path) so a stale-consumer
|
||||||
|
report is distinguishable from a real fault at read time. Summarize actual
|
||||||
|
results from each probe. Apply the scoped liveness rule above: on auth-gated
|
||||||
|
endpoints only connection-refused (`000`) or timeout is DOWN; on bare-200 probes
|
||||||
|
any other status is an alert. Never probe a GPU host on bare port 80.
|
||||||
|
|
||||||
### view-dashboard
|
### view-dashboard
|
||||||
Open `http://localhost:9100/` in browser — Live HTML dashboard
|
Open `http://localhost:9100/` in browser — Live HTML dashboard
|
||||||
@@ -110,9 +177,11 @@ python3 /root/scripts/gpu-monitor-server.py &
|
|||||||
Or via PM2: `pm2 restart gpu-monitor`
|
Or via PM2: `pm2 restart gpu-monitor`
|
||||||
|
|
||||||
### check-router
|
### check-router
|
||||||
The router health is accessed through nginx on port 80 (NOT port 9000 directly).
|
The router health is accessed through nginx on port 80 on the **router**
|
||||||
`curl http://192.168.68.116/health/unified` — Router unified health via nginx proxy
|
(.116) — NOT port 9000 directly, and NOT bare port 80 on a GPU host.
|
||||||
|
`curl http://192.168.68.116/health/unified` — Router unified health via nginx proxy; answers `301` → `/gpu/gpu-data` (same payload) = alive
|
||||||
`curl http://192.168.68.116:9000/health/unified` — ❌ WILL FAIL (port bound to 127.0.0.1 only)
|
`curl http://192.168.68.116:9000/health/unified` — ❌ WILL FAIL (port bound to 127.0.0.1 only)
|
||||||
|
`curl http://192.168.68.8/health` — ❌ NEVER USE (GPU host, no port-80 listener → false `000`/DEGRADED)
|
||||||
|
|
||||||
## Configuration Files
|
## Configuration Files
|
||||||
|
|
||||||
@@ -124,13 +193,18 @@ The router health is accessed through nginx on port 80 (NOT port 9000 directly).
|
|||||||
|
|
||||||
## Execution
|
## Execution
|
||||||
|
|
||||||
1. **Poll router** (every 15s): GET .116/health/unified — single source of truth for all GPU status (router probes each GPU /health directly via sidecar fallback)
|
**Port discipline:** probe GPU hosts on `:8080` (or the router's
|
||||||
2. **Poll router** (every 15s): GET .116/health via nginx:80
|
`/health/unified`); probe port 80 only on the router (.116). Never bare port 80
|
||||||
3. **Poll LiteLLM** (every 15s): GET .116/litellm/health via nginx:80
|
on a GPU host.
|
||||||
4. **Poll Strix** (every 15s): via router /health/unified (cannot poll .15:8080 directly — firewalled to .116 only)
|
|
||||||
5. **Poll dashboard** (every 15s): GET .116/dashboard/
|
1. **Poll router** (every 15s): GET .116/health/unified — single source of truth for all GPU status (router probes each GPU /health directly via sidecar fallback). `301` → `/gpu/gpu-data` counts as alive.
|
||||||
6. **Check alerts**: Compare metrics against thresholds
|
2. **Fallback direct GPU probe** (only if router /health/unified is DOWN): GET `http://192.168.68.8:8080/health` and `http://192.168.68.110:8080/health` — **:8080 only, never bare port 80**.
|
||||||
7. **Compute summary**: Fleet-wide health aggregation
|
3. **Poll router** (every 15s): GET .116/health via nginx:80
|
||||||
8. **Render dashboard**: Generate HTML at /root/dashboard/gpu-fleet.html
|
4. **Poll LiteLLM** (every 15s): GET .116/litellm/health via nginx:80
|
||||||
9. **Serve API**: HTTP server on port 9100
|
5. **Poll Strix** (every 15s): via router /health/unified (cannot poll .15:8080 directly — firewalled to .116 only)
|
||||||
10. **Repeat** every 15 seconds
|
6. **Poll dashboard** (every 15s): GET .116/dashboard/
|
||||||
|
7. **Check alerts**: Compare metrics against thresholds
|
||||||
|
8. **Compute summary**: Fleet-wide health aggregation
|
||||||
|
9. **Render dashboard**: Generate HTML at /root/dashboard/gpu-fleet.html
|
||||||
|
10. **Serve API**: HTTP server on port 9100
|
||||||
|
11. **Repeat** every 15 seconds
|
||||||
|
|||||||
@@ -1,4 +1,6 @@
|
|||||||
---
|
---
|
||||||
|
report_only_agents:
|
||||||
|
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
|
||||||
kind: responsibility
|
kind: responsibility
|
||||||
name: gpu-self-heal
|
name: gpu-self-heal
|
||||||
description: >
|
description: >
|
||||||
@@ -16,6 +18,7 @@ depends_on:
|
|||||||
- gpu-monitor.prose.md (live data source on .24:9100)
|
- gpu-monitor.prose.md (live data source on .24:9100)
|
||||||
- gpu-fleet.prose.md (source of truth for topology, aliases, model assignments)
|
- gpu-fleet.prose.md (source of truth for topology, aliases, model assignments)
|
||||||
---
|
---
|
||||||
|
---
|
||||||
|
|
||||||
## Maintains
|
## Maintains
|
||||||
|
|
||||||
@@ -38,20 +41,19 @@ depends_on:
|
|||||||
- On fix: verify with benchmark inference test before declaring resolved
|
- On fix: verify with benchmark inference test before declaring resolved
|
||||||
- Escalate: after 3 failed remediation attempts → Zulip #agent-hub alert
|
- Escalate: after 3 failed remediation attempts → Zulip #agent-hub alert
|
||||||
|
|
||||||
|
---
|
||||||
---
|
---
|
||||||
|
|
||||||
## Current Fleet Baseline (2026-07-18)
|
## Current Fleet Baseline (2026-07-18)
|
||||||
|
|
||||||
| Alias | GPU | Host | Model | VRAM | Ctx | tok/s | Role |
|
| Alias | GPU | Host | Model | VRAM | Ctx | tok/s | Role |
|
||||||
|-------|-----|------|-------|------|-----|-------|------|
|
|-------|-----|------|-------|------|-----|-------|------|
|
||||||
| `gpu-dense` | RTX 3090 24GB | ct8 (.8:8080) | Qwen3.8-27B-Uncensored-Q4_K_M (alias qwen3.6-27B-code) | ~16.8/24.6GB | 128K | — | Heavy reasoning, code gen |
|
|
||||||
| `gpu-light` | RTX 5070 12GB | ct110 (.110:8080) | HauhauCS Gemma4-12B QAT Q4_K_M + MTP draft | 10.1/12.2GB (83%) | 128K | 169.6 | Vision, web extract, light tasks |
|
|
||||||
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | qwen3.6-35B-udq4 | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs |
|
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | qwen3.6-35B-udq4 | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs |
|
||||||
|
|
||||||
Key notes:
|
Key notes:
|
||||||
- All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) is deprecated and NOT in the inference path.
|
- All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) is deprecated and NOT in the inference path.
|
||||||
- Stable aliases (gpu-dense, gpu-light, strix-moe) from gpu-fleet are the canonical names for agent configs. Model-specific names still work but are deprecated.
|
- Stable aliases (gpu-dense, gpu-light, strix-moe) from gpu-fleet are the canonical names for agent configs. Model-specific names still work but are deprecated.
|
||||||
- RTX 5070 tok/s is 2.3x faster than RTX 3090 for its model — gpu-light is the fastest endpoint. Route vision/web/light work there first.
|
- RTX 5070 tok/s is ~145 for Qwen3.5-9B — gpu-light is the fastest endpoint. Route vision/web/light work there first. NOTE: Qwen3.5-9B is multimodal (image+text), NOT text-only like gemma-4-12b was.
|
||||||
- Strix Halo is 62.9 tok/s (89% of 70.5 baseline) — below optimal but stable. Check for competing workloads.
|
- Strix Halo is 62.9 tok/s (89% of 70.5 baseline) — below optimal but stable. Check for competing workloads.
|
||||||
- RTX 3090 VRAM at 88% — within role-appropriate range (role = heavy reasoning, needs the headroom).
|
- RTX 3090 VRAM at 88% — within role-appropriate range (role = heavy reasoning, needs the headroom).
|
||||||
- RTX 5070 VRAM at 83% — role-appropriate for vision/web (smaller batch sizes).
|
- RTX 5070 VRAM at 83% — role-appropriate for vision/web (smaller batch sizes).
|
||||||
@@ -156,7 +158,7 @@ Key notes:
|
|||||||
- **Detect**: GPU roles misaligned with hardware capabilities
|
- **Detect**: GPU roles misaligned with hardware capabilities
|
||||||
- **Target distribution**:
|
- **Target distribution**:
|
||||||
- RTX 3090 (gpu-dense, 24GB, 74.9 tok/s) → Heavy reasoning, code gen, long conversations (slowest per-token but largest context capacity). Weight: 0.55 (LiteLLM).
|
- RTX 3090 (gpu-dense, 24GB, 74.9 tok/s) → Heavy reasoning, code gen, long conversations (slowest per-token but largest context capacity). Weight: 0.55 (LiteLLM).
|
||||||
- RTX 5070 (gpu-light, 12GB, 169.6 tok/s) → Vision/image, web search, lightweight tasks (2.3x faster than 3090 per token). Weight: 0.15 (LiteLLM).
|
- RTX 5070 (gpu-light, 12GB, ~145 tok/s) → Vision (image+text), web search, lightweight tasks. Weight: 0.15 (LiteLLM).
|
||||||
- Strix Halo (strix-moe, 64GB, 62.9 tok/s) → Context compression, summarization, long docs (MoE model). Weight: 0.30 (LiteLLM).
|
- Strix Halo (strix-moe, 64GB, 62.9 tok/s) → Context compression, summarization, long docs (MoE model). Weight: 0.30 (LiteLLM).
|
||||||
- **Note**: RTX 5070 is the fastest endpoint per token. Route high-volume, low-complexity work there first.
|
- **Note**: RTX 5070 is the fastest endpoint per token. Route high-volume, low-complexity work there first.
|
||||||
- **Fix**:
|
- **Fix**:
|
||||||
@@ -166,6 +168,7 @@ Key notes:
|
|||||||
- **Verify**: Each GPU's request pattern matches its designated role within 24h
|
- **Verify**: Each GPU's request pattern matches its designated role within 24h
|
||||||
- **Escalate**: If role mismatch persists >48h → agent alias audit needed
|
- **Escalate**: If role mismatch persists >48h → agent alias audit needed
|
||||||
|
|
||||||
|
---
|
||||||
---
|
---
|
||||||
|
|
||||||
## Execution
|
## Execution
|
||||||
@@ -247,6 +250,7 @@ call update-gpu-health
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
|
---
|
||||||
---
|
---
|
||||||
|
|
||||||
## Reporting
|
## Reporting
|
||||||
@@ -263,6 +267,7 @@ Pushed to `SyslogSolution/health-logs/gpu/{run_id}.json` — versioned, searchab
|
|||||||
- Per-GPU tok/s trend over 7 days
|
- Per-GPU tok/s trend over 7 days
|
||||||
- Regression alerts if any GPU degrades >10% week-over-week
|
- Regression alerts if any GPU degrades >10% week-over-week
|
||||||
|
|
||||||
|
---
|
||||||
---
|
---
|
||||||
|
|
||||||
## Design Decisions (Verified 2026-07-12, Reaffirmed 2026-07-18)
|
## Design Decisions (Verified 2026-07-12, Reaffirmed 2026-07-18)
|
||||||
|
|||||||
@@ -24,8 +24,6 @@ done
|
|||||||
|
|
||||||
| Agent | CT | Node | IP | LiteLLM Alias | Key Source | Platform |
|
| Agent | CT | Node | IP | LiteLLM Alias | Key Source | Platform |
|
||||||
|-------|-----|------|-----|---------------|------------|----------|
|
|-------|-----|------|-----|---------------|------------|----------|
|
||||||
| Tanko | 112 | amdpve | .122 | `tanko` | Infisical vault | Hermes |
|
|
||||||
| Mumuni | 100 | minipve | .24 | `mumuni` | Infisical vault | Hermes |
|
|
||||||
| Koby | 111 | amdpve | .129 | `koby` | Infisical vault | **Hermes** |
|
| Koby | 111 | amdpve | .129 | `koby` | Infisical vault | **Hermes** |
|
||||||
| Koonimo | 113 | amdpve | .114 | `koonimo` | Infisical vault | Hermes |
|
| Koonimo | 113 | amdpve | .114 | `koonimo` | Infisical vault | Hermes |
|
||||||
| Shumba | — | 192.168.68.119 | N/A | N/A (DeepSeek) | Hermes (RETIRED — CT119 now Infisical vault) |
|
| Shumba | — | 192.168.68.119 | N/A | N/A (DeepSeek) | Hermes (RETIRED — CT119 now Infisical vault) |
|
||||||
@@ -53,9 +51,10 @@ Agent (systemd) → LITELLM_API_KEY → LiteLLM (:116/v1) → Router (:9000) →
|
|||||||
|
|
||||||
## Config Pattern — Mandatory Fields
|
## Config Pattern — Mandatory Fields
|
||||||
|
|
||||||
### For Hermes Agents (Tanko, Mumuni, Koonimo)
|
### For Hermes Agents (Mumuni, Koonimo)
|
||||||
|
|
||||||
Every agent's `/root/.hermes/config.yaml` (or `/home/jerome/.hermes/config.yaml`) MUST have:
|
Every Hermes agent's `/root/.hermes/config.yaml` (or `/home/jerome/.hermes/config.yaml`) MUST have:
|
||||||
|
(Tanko is excluded — migrated to DSH/DeepSeek Harness on 2026-08-27, no longer uses Hermes config.)
|
||||||
|
|
||||||
### 1. Main Model
|
### 1. Main Model
|
||||||
```yaml
|
```yaml
|
||||||
@@ -71,7 +70,7 @@ model:
|
|||||||
custom_providers:
|
custom_providers:
|
||||||
- name: harness
|
- name: harness
|
||||||
model: syslog-auto
|
model: syslog-auto
|
||||||
base_url: http://192.168.68.116/v1
|
base_url: http://192.168.68.116/litellm/v1
|
||||||
api_key_env: LITELLM_API_KEY
|
api_key_env: LITELLM_API_KEY
|
||||||
api_mode: chat_completions
|
api_mode: chat_completions
|
||||||
```
|
```
|
||||||
@@ -82,7 +81,7 @@ auxiliary:
|
|||||||
vision:
|
vision:
|
||||||
provider: harness
|
provider: harness
|
||||||
model: gemma-4-12b # or syslog-auto
|
model: gemma-4-12b # or syslog-auto
|
||||||
base_url: http://192.168.68.116/v1
|
base_url: http://192.168.68.116/litellm/v1
|
||||||
api_key_env: LITELLM_API_KEY
|
api_key_env: LITELLM_API_KEY
|
||||||
api_key: <value from: infisical secrets get LITELLM_API_KEY --project=agents --env=production> # ← MANDATORY workaround
|
api_key: <value from: infisical secrets get LITELLM_API_KEY --project=agents --env=production> # ← MANDATORY workaround
|
||||||
timeout: 60
|
timeout: 60
|
||||||
@@ -97,7 +96,7 @@ auxiliary:
|
|||||||
target_ratio: 0.3
|
target_ratio: 0.3
|
||||||
provider: harness
|
provider: harness
|
||||||
model: syslog-auto # or gemma-4-12b
|
model: syslog-auto # or gemma-4-12b
|
||||||
base_url: http://192.168.68.116/v1
|
base_url: http://192.168.68.116/litellm/v1
|
||||||
api_key_env: LITELLM_API_KEY
|
api_key_env: LITELLM_API_KEY
|
||||||
api_key: <value from: infisical secrets get LITELLM_API_KEY --project=agents --env=production> # ← MANDATORY workaround
|
api_key: <value from: infisical secrets get LITELLM_API_KEY --project=agents --env=production> # ← MANDATORY workaround
|
||||||
timeout: 120
|
timeout: 120
|
||||||
@@ -169,11 +168,15 @@ pct-run <CT> grep -A8 "vision:" /root/.hermes/config.yaml | grep api_key
|
|||||||
# Must show both api_key: sk-... and api_key_env: LITELLM_API_KEY
|
# Must show both api_key: sk-... and api_key_env: LITELLM_API_KEY
|
||||||
```
|
```
|
||||||
|
|
||||||
### For Koby (CT 111 / tdunna)
|
### For Koby (CT 111 / tdunna) — **REPORT-ONLY MODE**
|
||||||
|
|
||||||
Koby runs Hermes on CT 111 (tdunna). Config files at `/root/.hermes/config.yaml`.
|
Koby runs Hermes on CT 111 (tdunna). Config files at `/root/.hermes/config.yaml`.
|
||||||
Same Hermes pattern as Tanko/Mumuni/Koonimo — see config sections above.
|
Same Hermes pattern as Tanko/Mumuni/Koonimo — see config sections above.
|
||||||
|
|
||||||
|
**⛔ KOBY IS NEVER REPAIRED (2026-08-17, Captain)**: Diagnostic only — detect and report, never fix on .129.
|
||||||
|
No heal step, no restart, no key rotation, no config edit, no memory rewrite, no disk GC, no service touch, no process kill — ever.
|
||||||
|
If a health check shows Koby degraded, **DO NOT** execute any repair action. Instead, report to Zulip and let Theo fix it.
|
||||||
|
|
||||||
**LiteLLM key**: alias `koby` in LiteLLM DB, injected via `infisical run --` wrapper.
|
**LiteLLM key**: alias `koby` in LiteLLM DB, injected via `infisical run --` wrapper.
|
||||||
|
|
||||||
### For pi Agents (Abiba)
|
### For pi Agents (Abiba)
|
||||||
|
|||||||
@@ -15,7 +15,7 @@ description: >
|
|||||||
|
|
||||||
- template_version: "2.1.0"
|
- template_version: "2.1.0"
|
||||||
- last_applied: timestamp
|
- last_applied: timestamp
|
||||||
- agents_configured: ["tanko", "mumuni", "abiba", "koby", "koonimo", "kagenz0"]
|
- agents_configured: ["mumuni", "abiba", "koby", "koonimo", "kagenz0"] # tanko removed 2026-08-27 (now on DSH/DeepSeek Harness)
|
||||||
- agent_keys: map (see Agent Keys section)
|
- agent_keys: map (see Agent Keys section)
|
||||||
- infra_endpoints_verified: array
|
- infra_endpoints_verified: array
|
||||||
|
|
||||||
@@ -30,8 +30,6 @@ Sub-agent profiles inherit auth from the main config — no separate keys needed
|
|||||||
|
|
||||||
| Agent | Key Alias | Host | SSH | Sub-Agents |
|
| Agent | Key Alias | Host | SSH | Sub-Agents |
|
||||||
|-------|-----------|------|-----|-----------|
|
|-------|-----------|------|-----|-----------|
|
||||||
| Tanko | `tanko-*` | 192.168.68.122 | jerome@.122 | — |
|
|
||||||
| Mumuni | `mumuni` | 192.168.68.24 | root@.24 | 6 profiles ✱ |
|
|
||||||
| Abiba | `abiba-pi` | 192.168.68.24 | local | — |
|
| Abiba | `abiba-pi` | 192.168.68.24 | local | — |
|
||||||
| Koby | `koby` | CT 111 (tdunna) | Zulip | — |
|
| Koby | `koby` | CT 111 (tdunna) | Zulip | — |
|
||||||
| Koonimo | `koonimo` | CT 113 (baggy) | SSH root | — |
|
| Koonimo | `koonimo` | CT 113 (baggy) | SSH root | — |
|
||||||
@@ -348,6 +346,17 @@ curl -s -o /dev/null -w 'key_health: %{http_code}\n' -H "Authorization: Bearer $
|
|||||||
|
|
||||||
### Rule 13: API Key Injection — Two Patterns (UPDATED 2026-07-16, WAL #1300)
|
### Rule 13: API Key Injection — Two Patterns (UPDATED 2026-07-16, WAL #1300)
|
||||||
|
|
||||||
|
### Rule 14: Hermes Context Detection Uses `max_model_tokens`, NOT `max_input_tokens`
|
||||||
|
|
||||||
|
**CRITICAL**: Hermes context detection reads `max_model_tokens` (128K), NOT `max_input_tokens` (64K cap).
|
||||||
|
|
||||||
|
- **Abiba and Hermes agents**: `max_model_tokens: 131072` (128K) — unlimited context
|
||||||
|
- **Crewmates (ops, tune, verify, auth-keys, build)**: `max_input_tokens: 64000` (64K) — capped
|
||||||
|
- If you see `max_input_tokens: 64000` in an Abiba/Hermes config, that's a mistake
|
||||||
|
- Using `max_input_tokens` for Hermes agents causes premature context loss
|
||||||
|
- Check: `grep -n 'max_model_tokens\|max_input_tokens' ~/.hermes/config.yaml`
|
||||||
|
- Expected output: `max_model_tokens: 131072` (not max_input_tokens)
|
||||||
|
|
||||||
Agents inject `LITELLM_API_KEY` via ONE of two mechanisms. Both are valid; the contract
|
Agents inject `LITELLM_API_KEY` via ONE of two mechanisms. Both are valid; the contract
|
||||||
requirement is that the key is a **valid LiteLLM virtual key** (HTTP 200 on /v1/models).
|
requirement is that the key is a **valid LiteLLM virtual key** (HTTP 200 on /v1/models).
|
||||||
|
|
||||||
@@ -424,13 +433,3 @@ curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $K" http://192.
|
|||||||
5. **Set model choice** — Per agent's workload
|
5. **Set model choice** — Per agent's workload
|
||||||
6. **Verify** — curl all shared endpoints, test the model with the new key
|
6. **Verify** — curl all shared endpoints, test the model with the new key
|
||||||
7. **Report** — What was changed, preserved, custom
|
7. **Report** — What was changed, preserved, custom
|
||||||
### Rule 16: Koby Configuration (DeepSeek-primary)
|
|
||||||
(Ref: See Rule 10 for default model behavior, with Koby exception)
|
|
||||||
|
|
||||||
Koby uses a split-model architecture:
|
|
||||||
|
|
||||||
- Primary Model: `deepseek-v4-flash` via `api.deepseek.com` (for reasoning)
|
|
||||||
- Auxiliary Models: `gpu-light` (vision/web_extract) and `syslog-auto` (compression)
|
|
||||||
- Key Hygiene: `api_key_env` is strictly `LITELLM_API_KEY` or `DEEPSEEK_API_KEY`
|
|
||||||
- Constraint: Do NOT touch Koby's primary model/provider/compression settings unless explicitly ruled by the captain.
|
|
||||||
|
|
||||||
|
|||||||
@@ -190,7 +190,7 @@ litellm_settings:
|
|||||||
| Agent | CT | IP | LiteLLM Alias | Key Source | Status | Gateway Wrapper | Last Verified |
|
| Agent | CT | IP | LiteLLM Alias | Key Source | Status | Gateway Wrapper | Last Verified |
|
||||||
|-------|-----|-----|---------------|------------|--------|-----------------|---------------|
|
|-------|-----|-----|---------------|------------|--------|-----------------|---------------|
|
||||||
| Tanko | 112 | .122 | `tanko` | Infisical vault | ✅ Fixed | `infisical run` | 20:17 UTC Jul 5 |
|
| Tanko | 112 | .122 | `tanko` | Infisical vault | ✅ Fixed | `infisical run` | 20:17 UTC Jul 5 |
|
||||||
| Mumuni | 100 (abiba) | .24 | `mumuni` | Infisical vault | ✅ Fixed | Pi Hermes gateway | 2026-07-27 |
|
| Mumuni | 105 (kagentz) | .14 | `mumuni` | Infisical vault | ✅ Fixed | systemd Hermes gateway | 2026-08-29 |
|
||||||
| Koby | 111 | .129 | `koby` | Infisical vault | ✅ Fixed (DeepSeek-primary) | `infisical run` | 23:30 UTC Jul 5 |
|
| Koby | 111 | .129 | `koby` | Infisical vault | ✅ Fixed (DeepSeek-primary) | `infisical run` | 23:30 UTC Jul 5 |
|
||||||
| Koonimo | 113 | .114 | `koonimo` | Infisical vault | ✅ Fixed | `infisical run` (migrated 2026-07-11) | 2026-08-09 |
|
| Koonimo | 113 | .114 | `koonimo` | Infisical vault | ✅ Fixed | `infisical run` (migrated 2026-07-11) | 2026-08-09 |
|
||||||
| Abiba | 100 | .65 | `abiba-pi` | Infisical vault | ✅ N/A (pi native) | — | 19:44 UTC Jul 5 |
|
| Abiba | 100 | .65 | `abiba-pi` | Infisical vault | ✅ N/A (pi native) | — | 19:44 UTC Jul 5 |
|
||||||
@@ -285,13 +285,13 @@ auxiliary:
|
|||||||
vision:
|
vision:
|
||||||
api_key: sk-<agent-key-from-vault> # ← workaround (get via: infisical secrets get LITELLM_API_KEY --project=agents --env=production --plain)
|
api_key: sk-<agent-key-from-vault> # ← workaround (get via: infisical secrets get LITELLM_API_KEY --project=agents --env=production --plain)
|
||||||
api_key_env: LITELLM_API_KEY
|
api_key_env: LITELLM_API_KEY
|
||||||
base_url: http://192.168.68.116/v1
|
base_url: http://192.168.68.116/litellm/v1
|
||||||
model: gemma-4-12b
|
model: gemma-4-12b
|
||||||
provider: harness
|
provider: harness
|
||||||
compression:
|
compression:
|
||||||
api_key: sk-<agent-key-from-vault> # ← workaround (same as above)
|
api_key: sk-<agent-key-from-vault> # ← workaround (same as above)
|
||||||
api_key_env: LITELLM_API_KEY
|
api_key_env: LITELLM_API_KEY
|
||||||
base_url: http://192.168.68.116/v1
|
base_url: http://192.168.68.116/litellm/v1
|
||||||
model: gemma-4-12b
|
model: gemma-4-12b
|
||||||
provider: harness
|
provider: harness
|
||||||
```
|
```
|
||||||
|
|||||||
@@ -28,7 +28,7 @@ connectivity recovery including end-to-end DM validation.
|
|||||||
|
|
||||||
| Param | Type | Required | Default | Description |
|
| Param | Type | Required | Default | Description |
|
||||||
|-------|------|----------|---------|-------------|
|
|-------|------|----------|---------|-------------|
|
||||||
| `target` | string | yes | — | Agent name: `mumuni`, `tanko`, `koby`, or `shumba` |
|
| `target` | string | yes | — | Agent name: `mumuni`, `koby`, or `shumba` (Tanko excluded — on DSH since 2026-08-27, no Hermes plugin) |
|
||||||
| `branch` | string | no | `master` | Git branch to pull (overridable for pinning) |
|
| `branch` | string | no | `master` | Git branch to pull (overridable for pinning) |
|
||||||
|
|
||||||
## Maintains
|
## Maintains
|
||||||
@@ -55,8 +55,7 @@ connectivity recovery including end-to-end DM validation.
|
|||||||
|
|
||||||
| Host | CT | Proxmox | IP (direct) | Hermes Home | User |
|
| Host | CT | Proxmox | IP (direct) | Hermes Home | User |
|
||||||
|------|-----|---------|-------------|-------------|------|
|
|------|-----|---------|-------------|-------------|------|
|
||||||
| Mumuni | CT100 | — | 192.168.68.24 | /root/.hermes | root |
|
| Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome | *(DSH since 2026-08-27 — historical, plugin retired on this host)* |
|
||||||
| Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome |
|
|
||||||
| Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root |
|
| Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root |
|
||||||
| Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky |
|
| Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky |
|
||||||
|
|
||||||
@@ -121,7 +120,8 @@ cp plugins/platforms/zulip/adapter.py \
|
|||||||
plugins/platforms/zulip/plugin.yaml \
|
plugins/platforms/zulip/plugin.yaml \
|
||||||
{{hermes_home}}/hermes-agent/plugins/platforms/zulip/
|
{{hermes_home}}/hermes-agent/plugins/platforms/zulip/
|
||||||
|
|
||||||
# Fix ownership (Tanko only — runs as jerome user)
|
# Fix ownership (was Tanko-only, runs as jerome user)
|
||||||
|
# RETIRED 2026-08-27: tanko no longer uses the Hermes Zulip plugin (DSH).
|
||||||
[ "{{target}}" = "tanko" ] && chown -R jerome:jerome \
|
[ "{{target}}" = "tanko" ] && chown -R jerome:jerome \
|
||||||
{{hermes_home}}/hermes-agent/plugins/platforms/zulip/
|
{{hermes_home}}/hermes-agent/plugins/platforms/zulip/
|
||||||
|
|
||||||
|
|||||||
@@ -1,9 +1,9 @@
|
|||||||
---
|
---
|
||||||
|
report_only_agents:
|
||||||
|
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
|
||||||
kind: function
|
kind: function
|
||||||
name: hermes-zulip-restore
|
name: hermes-zulip-restore
|
||||||
description: >
|
description: >
|
||||||
Restores Zulip connectivity for any Hermes agent (Mumuni CT100, Tanko CT112,
|
|
||||||
Koby CT111, Shumba on Lucky's mini PC). Deploys the zulip-platform adapter to the correct bundled plugin
|
|
||||||
path, verifies env credentials, restarts the gateway, and confirms Zulip
|
path, verifies env credentials, restarts the gateway, and confirms Zulip
|
||||||
connects. Run this whenever a Hermes agent stops responding on Zulip or after
|
connects. Run this whenever a Hermes agent stops responding on Zulip or after
|
||||||
a fresh agent deployment.
|
a fresh agent deployment.
|
||||||
@@ -12,6 +12,7 @@ version: 1.0.0
|
|||||||
status: active
|
status: active
|
||||||
runtime_contract: 2
|
runtime_contract: 2
|
||||||
---
|
---
|
||||||
|
---
|
||||||
|
|
||||||
# Hermes Zulip Restore — Bring Any Agent Back to Good State
|
# Hermes Zulip Restore — Bring Any Agent Back to Good State
|
||||||
|
|
||||||
@@ -23,7 +24,7 @@ gateway restart, and connection validation.
|
|||||||
|
|
||||||
| Param | Type | Required | Default | Description |
|
| Param | Type | Required | Default | Description |
|
||||||
|-------|------|----------|---------|-------------|
|
|-------|------|----------|---------|-------------|
|
||||||
| `target` | string | yes | — | Agent name: `mumuni`, `tanko`, `koby`, or `shumba` |
|
| `target` | string | yes | — | Agent name: `mumuni`, `koby`, or `shumba` (Tanko excluded — DSH since 2026-08-27) |
|
||||||
|
|
||||||
## Maintains
|
## Maintains
|
||||||
|
|
||||||
@@ -36,7 +37,7 @@ gateway restart, and connection validation.
|
|||||||
|
|
||||||
- `_strip_html` function present in `<hermes-agent>/plugins/platforms/zulip/adapter.py`
|
- `_strip_html` function present in `<hermes-agent>/plugins/platforms/zulip/adapter.py`
|
||||||
- All three adapter files (__init__.py, adapter.py, plugin.yaml) present at bundled path
|
- All three adapter files (__init__.py, adapter.py, plugin.yaml) present at bundled path
|
||||||
- Zulip env vars set in `~/.hermes/.env` (or `/home/jerome/.hermes/.env` for Tanko)
|
- Zulip env vars set in `~/.hermes/.env` (or `/home/jerome/.hermes/.env` for Tanko, historical — DSH since 2026-08-27)
|
||||||
- Gateway restarted and zulip platform reports state `connected`
|
- Gateway restarted and zulip platform reports state `connected`
|
||||||
- HTML stripping enabled for `/approve` and `/deny` slash command support
|
- HTML stripping enabled for `/approve` and `/deny` slash command support
|
||||||
|
|
||||||
@@ -51,8 +52,6 @@ gateway restart, and connection validation.
|
|||||||
|
|
||||||
| Host | CT | Proxmox | IP (direct) | Hermes Home | User |
|
| Host | CT | Proxmox | IP (direct) | Hermes Home | User |
|
||||||
|------|-----|---------|-------------|-------------|------|
|
|------|-----|---------|-------------|-------------|------|
|
||||||
| Mumuni | CT100 (abiba) | minipve | 192.168.68.24 | /root/.hermes | root |
|
|
||||||
| Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome |
|
|
||||||
| Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root |
|
| Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root |
|
||||||
| Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky |
|
| Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky |
|
||||||
|
|
||||||
@@ -94,8 +93,8 @@ cp zulip-platform-plugins/plugins/platforms/zulip/adapter.py \
|
|||||||
zulip-platform-plugins/plugins/platforms/zulip/plugin.yaml \
|
zulip-platform-plugins/plugins/platforms/zulip/plugin.yaml \
|
||||||
<HERMES_HOME>/hermes-agent/plugins/platforms/zulip/
|
<HERMES_HOME>/hermes-agent/plugins/platforms/zulip/
|
||||||
|
|
||||||
# Fix ownership (Tanko only)
|
# Fix ownership (was Tanko-only; RETIRED 2026-08-27 — tanko on DSH, no Hermes plugin)
|
||||||
chown -R jerome:jerome <HERMES_HOME>/hermes-agent/plugins/platforms/zulip/ # Tanko only
|
chown -R jerome:jerome <HERMES_HOME>/hermes-agent/plugins/platforms/zulip/ # Tanko only (historical)
|
||||||
|
|
||||||
# Clean up
|
# Clean up
|
||||||
rm -rf /tmp/zulip-deploy
|
rm -rf /tmp/zulip-deploy
|
||||||
@@ -184,6 +183,7 @@ https://git.sysloggh.net/SyslogSolution/zulip-platform-plugins/src/branch/feat/z
|
|||||||
Commit `55ca15d` — `fix(zulip): add _strip_html for slash command matching`
|
Commit `55ca15d` — `fix(zulip): add _strip_html for slash command matching`
|
||||||
Pull request #33 is the primary integration branch.
|
Pull request #33 is the primary integration branch.
|
||||||
|
|
||||||
|
---
|
||||||
---
|
---
|
||||||
|
|
||||||
**Last verified good state**: 2026-07-08 — Mumuni, Tanko, Koby all connected with `_strip_html` applied.
|
**Last verified good state**: 2026-07-08 — Mumuni, Tanko, Koby all connected with `_strip_html` applied.
|
||||||
|
|||||||
@@ -87,8 +87,8 @@ call apply-liteLLM-routing
|
|||||||
|
|
||||||
call apply-agent-compression
|
call apply-agent-compression
|
||||||
agent: mumuni
|
agent: mumuni
|
||||||
host: 192.168.68.24
|
host: 192.168.68.14
|
||||||
config_path: /root/.hermes/config.yaml
|
config_path: /home/hermes/.hermes/config.yaml
|
||||||
|
|
||||||
-- Phase 4: Enable llama.cpp prompt caching on GPU hosts
|
-- Phase 4: Enable llama.cpp prompt caching on GPU hosts
|
||||||
|
|
||||||
|
|||||||
@@ -613,7 +613,7 @@ ssh root@192.168.68.110 "systemctl restart llama-server"
|
|||||||
| 109 | docker-vm | storepve | .7 | Docker host | ❌ |
|
| 109 | docker-vm | storepve | .7 | Docker host | ❌ |
|
||||||
| 110 | gitea | minipve | **.17** | Git | ❌ |
|
| 110 | gitea | minipve | **.17** | Git | ❌ |
|
||||||
| 111 | tdunna | amdpve | .129 | Hermes agent | ✅ |
|
| 111 | tdunna | amdpve | .129 | Hermes agent | ✅ |
|
||||||
| 112 | tanko | amdpve | .122 | Hermes agent | ✅ |
|
| 112 | tanko | amdpve | .122 | DSH (DeepSeek Harness) agent | ✅ |
|
||||||
| 113 | baggy | amdpve | .114 | Hermes agent | ✅ |
|
| 113 | baggy | amdpve | .114 | Hermes agent | ✅ |
|
||||||
| 115 | scottdenya | amdpve | .75 | Denya OneCare | ❌ |
|
| 115 | scottdenya | amdpve | .75 | Denya OneCare | ❌ |
|
||||||
| 116 | syslog-api | minipve | .116 | LiteLLM + Grafana | ❌ |
|
| 116 | syslog-api | minipve | .116 | LiteLLM + Grafana | ❌ |
|
||||||
|
|||||||
@@ -140,7 +140,6 @@ After ALL updates (apt + images + restarts), verify every critical service is ba
|
|||||||
| Zulip | `curl -sf https://chat.sysloggh.net/api/v1/server_settings` | 200 OK |
|
| Zulip | `curl -sf https://chat.sysloggh.net/api/v1/server_settings` | 200 OK |
|
||||||
| Gitea | `curl -sf https://git.sysloggh.net/api/v1/version` | 200 OK |
|
| Gitea | `curl -sf https://git.sysloggh.net/api/v1/version` | 200 OK |
|
||||||
| PM2 processes | `pm2 jlist` (CT 100) | all pi-agent processes `online` |
|
| PM2 processes | `pm2 jlist` (CT 100) | all pi-agent processes `online` |
|
||||||
| Hermes gateways | SSH to Mumuni CT 100, Tanko CT 112; `systemctl is-active hermes-gateway` | `active` for each |
|
|
||||||
|
|
||||||
Regression check: every service that was GREEN in `health-baseline` must still be GREEN. A service that was already RED (and caused a preflight abort) is excluded — but Phase 0 should have aborted before we got here.
|
Regression check: every service that was GREEN in `health-baseline` must still be GREEN. A service that was already RED (and caused a preflight abort) is excluded — but Phase 0 should have aborted before we got here.
|
||||||
|
|
||||||
|
|||||||
@@ -103,6 +103,92 @@ GPU .8 (RTX 3090) GPU .110 (RTX 5070) GPU .15 (Strix Halo)
|
|||||||
|
|
||||||
## Execution
|
## Execution
|
||||||
|
|
||||||
|
### Liveness rule (scoped)
|
||||||
|
|
||||||
|
The any-HTTP-response rule applies ONLY to unauthenticated/auth-gated endpoints,
|
||||||
|
where any HTTP answer proves a listener is up: the PVE API
|
||||||
|
(`https://<node>:8006/api2/json/version`) and LiteLLM health
|
||||||
|
(`/litellm/health`, `301` → `/litellm/health/liveliness`). For those endpoints a
|
||||||
|
probe is **ALIVE** on **ANY** HTTP status — `401`/`403` auth challenges and `3xx`
|
||||||
|
redirects included — and **DOWN = connection refused (`000`) or timeout only**.
|
||||||
|
The PVE API legitimately answers `401` to an unauthenticated probe — that is the
|
||||||
|
healthy signal, not a failure. Same scoped rule as zulip-health (Tanko) and
|
||||||
|
gpu-monitor.
|
||||||
|
|
||||||
|
Probes whose success condition is specifically a bare `200` are NOT covered by
|
||||||
|
the any-HTTP rule. On those — the authenticated Zulip POST and the router
|
||||||
|
`/health` — an unexpected status (`401`/`403` from a bad or missing credential,
|
||||||
|
`5xx`, or anything other than the expected `200`) is an **ALERT**, not "alive".
|
||||||
|
|
||||||
|
### check-health
|
||||||
|
|
||||||
|
**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Provenance — run first; paste the absolute path into the report
|
||||||
|
pwd -P
|
||||||
|
|
||||||
|
# Zulip API health (POST ping)
|
||||||
|
source /etc/litellm-monitor.env
|
||||||
|
ZULIP_USER="abiba-bot@chat.sysloggh.net"
|
||||||
|
curl -s -o /dev/null -w '%{http_code}' -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${ZULIP_BOT_KEY}"
|
||||||
|
# Expected: 200 (HTTP 000 = unreachable/cache)
|
||||||
|
|
||||||
|
# PM2 process health
|
||||||
|
pm2 jlist
|
||||||
|
# Expected: 5/5 online (abiba-telegram, abiba-zulip, zulip-watchdog, gitea-runner, spoton-service)
|
||||||
|
|
||||||
|
# GPU exporters (may be down per DEPLOYMENT STATUS)
|
||||||
|
curl -s http://192.168.68.8:9400/metrics && echo " - OK" || echo " - FAIL"
|
||||||
|
curl -s http://192.168.68.110:9400/metrics && echo " - OK" || echo " - FAIL"
|
||||||
|
curl -s http://192.168.68.15:9400/metrics && echo " - OK" || echo " - FAIL"
|
||||||
|
|
||||||
|
# Router health (via nginx on port 80)
|
||||||
|
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health
|
||||||
|
# Expected: 200 (Router is up and responding)
|
||||||
|
|
||||||
|
# LiteLLM health (via nginx on port 80)
|
||||||
|
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/litellm/health
|
||||||
|
# Expected: 301 → /litellm/health/liveliness (200 after redirect) — any HTTP status = alive
|
||||||
|
|
||||||
|
# PVE API liveness — probe the REAL PVE nodes on :8006, never the monitoring
|
||||||
|
# host CT 116. CT 116 runs no pveproxy, so probing it on :8006 returns 000 —
|
||||||
|
# that was the stale-vantage bug this replaces (CT 116 is the monitoring host,
|
||||||
|
# not a cluster node). Unauthenticated GET answers 401 while the API is ALIVE
|
||||||
|
# by design. Alive = ANY HTTP status (401 is the EXPECTED healthy response);
|
||||||
|
# DOWN = connection refused (000) or timeout only.
|
||||||
|
for node in 192.168.68.9 192.168.68.5 192.168.68.15 192.168.68.6 192.168.68.12; do
|
||||||
|
printf '%s:8006 -> %s\n' "$node" \
|
||||||
|
"$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 5 "https://$node:8006/api2/json/version")"
|
||||||
|
done
|
||||||
|
# Expected: 401 on every node (acerpve .9, ocupve .5, amdpve .15, storepve .6, minipve .12)
|
||||||
|
# A node answering 000/timeout is DOWN — flag that node. 401 is NOT a fault.
|
||||||
|
|
||||||
|
# Prometheus targets
|
||||||
|
curl -s http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets'
|
||||||
|
# Expected: All targets UP (may show some down if exporters not deployed)
|
||||||
|
|
||||||
|
# Grafana health
|
||||||
|
curl -s http://192.168.68.116:3001/api/health | jq '{status, version}'
|
||||||
|
# Expected: {"status":"ok","version":"..."}
|
||||||
|
|
||||||
|
# LiteLLM metrics (Prometheus endpoint)
|
||||||
|
curl -s http://192.168.68.116:4001/metrics | head -20
|
||||||
|
# Expected: Prometheus-formatted metrics output
|
||||||
|
```
|
||||||
|
|
||||||
|
**Report format**: Begin every report with the **absolute path the probe executed
|
||||||
|
from** (`pwd -P`, or the script's absolute path) so a stale-consumer report is
|
||||||
|
distinguishable from a real fault at read time. Summarize actual results from
|
||||||
|
each probe. Apply the any-HTTP-response liveness rule ONLY to the auth-gated PVE
|
||||||
|
API and LiteLLM endpoints above: only connection-refused (`000`) or timeout is
|
||||||
|
DOWN; empty output is a warning. For probes whose expected result is a bare `200`
|
||||||
|
(the authenticated Zulip POST, router `/health`), flag an alert on any unexpected
|
||||||
|
status (`401`/`403`/`5xx`) — do not summarize it as alive. A bare-`200`
|
||||||
|
expectation on the auth-gated PVE API (`401`) or LiteLLM health (`301` redirect)
|
||||||
|
is a stale expectation, not a fault.
|
||||||
|
|
||||||
|
|
||||||
### Phase 1: GPU Exporters
|
### Phase 1: GPU Exporters
|
||||||
|
|
||||||
**NVIDIA (.8 and .110)**:
|
**NVIDIA (.8 and .110)**:
|
||||||
|
|||||||
@@ -3,15 +3,16 @@ kind: responsibility
|
|||||||
name: infrastructure-update
|
name: infrastructure-update
|
||||||
description: >
|
description: >
|
||||||
Autonomous system-wide update contract covering all 5 Proxmox nodes,
|
Autonomous system-wide update contract covering all 5 Proxmox nodes,
|
||||||
15+ containers/VMs, and 4 Docker ecosystems. Updates apt packages,
|
15+ containers/VMs, and 5 Docker ecosystems (docker-vm .7, CT 116 .116,
|
||||||
|
CT 117, hwpve .11, NetBird VPS 72.61.0.17). Updates apt packages,
|
||||||
Docker images, and container stacks in safe waves with health checks
|
Docker images, and container stacks in safe waves with health checks
|
||||||
and automatic rollback on failure.
|
and automatic rollback on failure.
|
||||||
agent: abiba
|
agent: abiba
|
||||||
triggers:
|
triggers:
|
||||||
- on "infra update" command
|
- on "infra update" command
|
||||||
- weekly (Sunday 03:00 EDT) via cron
|
- weekly (Sunday 03:00 America/New_York) via Agent Zero scheduler task "weekly-fleet-docker-update" (qSOOVzsU) — implemented 2026-09-08
|
||||||
- on security advisory relay from Mumuni
|
- on security advisory relay from Mumuni
|
||||||
version: 1.2.0
|
version: 1.3.0
|
||||||
---
|
---
|
||||||
|
|
||||||
## Maintains
|
## Maintains
|
||||||
@@ -59,7 +60,7 @@ Before ANY update wave:
|
|||||||
| CT 100 (.24) | Abiba (pi) | `apt update && apt upgrade -y` | 3 min |
|
| CT 100 (.24) | Abiba (pi) | `apt update && apt upgrade -y` | 3 min |
|
||||||
| CT 116 (.116) | syslog-api (LiteLLM host) | `apt update && apt upgrade -y` | 3 min |
|
| CT 116 (.116) | syslog-api (LiteLLM host) | `apt update && apt upgrade -y` | 3 min |
|
||||||
| CT 112 (tanko, amdpve) | Tanko | `apt update && apt upgrade -y` | 3 min |
|
| CT 112 (tanko, amdpve) | Tanko | `apt update && apt upgrade -y` | 3 min |
|
||||||
| CT 100 (mumuni/abiba, minipve) | Mumuni | `apt update && apt upgrade -y` | 3 min |
|
| CT 105 (kagentz, minipve) | Mumuni | `apt update && apt upgrade -y` | 3 min |
|
||||||
| VM 101 (.8) | llm-gpu (RTX 3090) | `apt update && apt upgrade -y` | 3 min |
|
| VM 101 (.8) | llm-gpu (RTX 3090) | `apt update && apt upgrade -y` | 3 min |
|
||||||
| VM 103 (.110) | ocu-llm (RTX 5070) | `apt update && apt upgrade -y` | 3 min |
|
| VM 103 (.110) | ocu-llm (RTX 5070) | `apt update && apt upgrade -y` | 3 min |
|
||||||
|
|
||||||
@@ -80,7 +81,13 @@ Before ANY update wave:
|
|||||||
| VM 109 (.7) | Home stack (Pulse, Stirling PDF) — JDownloader moved to CT 118 LXC 2026-08-01 | `cd /opt/home_stack && docker compose pull && docker compose up -d` |
|
| VM 109 (.7) | Home stack (Pulse, Stirling PDF) — JDownloader moved to CT 118 LXC 2026-08-01 | `cd /opt/home_stack && docker compose pull && docker compose up -d` |
|
||||||
| VM 109 (.7) | Audiobookshelf | `cd /opt/audiobookshelf && docker compose pull && docker compose up -d` |
|
| VM 109 (.7) | Audiobookshelf | `cd /opt/audiobookshelf && docker compose pull && docker compose up -d` |
|
||||||
| CT 116 (.116) | Inference Harness (LiteLLM, Prometheus, Grafana) | `cd /opt/inference-harness && docker compose pull && docker compose up -d` |
|
| CT 116 (.116) | Inference Harness (LiteLLM, Prometheus, Grafana) | `cd /opt/inference-harness && docker compose pull && docker compose up -d` |
|
||||||
| CT 117 (zulip, storepve) | Zulip | `docker pull zulip/docker-zulip:latest && docker restart zulip-zulip-1` |
|
| CT 117 (storepve) | Zulip | `pct exec 117 -- bash -c 'cd /opt/zulip && docker compose pull && docker compose up -d'` (from storepve; compose recreates on zulip_default network) |
|
||||||
|
| CT 117 (storepve) | Jitsi | `pct exec 117 -- bash -c 'cd /opt/jitsi && docker compose pull && docker compose up -d'` (from storepve) |
|
||||||
|
| hwpve (.11) | Authentik (server, worker, postgres) | `ssh root@192.168.68.11 'cd /root && docker compose pull && docker compose up -d'` |
|
||||||
|
| NetBird VPS (72.61.0.17) | NetBird (server, dashboard, proxy, traefik, crowdsec) | `ssh root@72.61.0.17 'cd /root && docker compose pull && docker compose up -d'` |
|
||||||
|
| VM 109 (.7) | Trove test | `cd /opt/trove-test && docker compose pull && docker compose up -d` |
|
||||||
|
| VM 109 (.7) | docker-stats | `cd /opt/docker-stats && docker compose pull && docker compose up -d` |
|
||||||
|
| CT 116 (.116) | Monitoring (Grafana, Prometheus, Alertmanager, PVE exporter) | `cd /opt/monitoring && docker compose pull && docker compose up -d` |
|
||||||
|
|
||||||
**Verify after Wave 3:**
|
**Verify after Wave 3:**
|
||||||
- All containers healthy: `docker ps` on each host
|
- All containers healthy: `docker ps` on each host
|
||||||
@@ -88,8 +95,12 @@ Before ANY update wave:
|
|||||||
- MCP integration test: `curl localhost:4000/mcp-rest/tools/list -H "Authorization: Bearer $MASTER_KEY"` → 90 tools (23 RA-H OS + 67 GitHub)
|
- MCP integration test: `curl localhost:4000/mcp-rest/tools/list -H "Authorization: Bearer $MASTER_KEY"` → 90 tools (23 RA-H OS + 67 GitHub)
|
||||||
- Zulip test: send test message to #agent-hub
|
- Zulip test: send test message to #agent-hub
|
||||||
- Dashboard loading: `curl localhost:3001/` (via CT 116)
|
- Dashboard loading: `curl localhost:3001/` (via CT 116)
|
||||||
- Firecrawl test: `curl :3002/`
|
- Firecrawl test: `curl -X POST http://192.168.68.7:3002/v1/search -H 'Content-Type: application/json' -d '{"query":"health","limit":1}'` → `"success":true` (GET `/` returns 200)
|
||||||
|
- Authentik test: `curl http://192.168.68.11:9000/` → 302 redirect to login
|
||||||
|
- NetBird test: `curl -s -o /dev/null -w '%{http_code}' https://netbird.sysloggh.net/` → 200
|
||||||
|
- harness-litellm cold start: allow 3-5 min after recreate — reports unhealthy and :4000 refuses connections while loading config/DB, then recovers to 200 on its own (verified 2026-09-08)
|
||||||
- SearXNG test: `curl :8888`
|
- SearXNG test: `curl :8888`
|
||||||
|
- Digest-pin sweep: `grep -rn '@sha256:' /opt/*/docker-compose.y*` on every host — digest-pinned images are INVISIBLE to `docker compose pull` (the pin re-pulls the same digest forever, so new releases never appear). Flag every pin in the run report and propose un-pinning to a floating tag with user approval before editing. Found 2026-09-10: audiobookshelf was digest-pinned at 2.34.0 (container created 2026-07-18) and silently missed by every sweep; dockhand stack was also pinned (stack removed 2026-09-10, unused). After un-pinning audiobookshelf to :latest it updated to 2.36.0 and verified HTTP 200.
|
||||||
|
|
||||||
## Wave 4: Proxmox Kernel Reboot
|
## Wave 4: Proxmox Kernel Reboot
|
||||||
|
|
||||||
@@ -158,13 +169,15 @@ Before Wave 1, snapshot these files:
|
|||||||
/opt/search-stack/searxng/docker-compose.yml (VM 109 .7)
|
/opt/search-stack/searxng/docker-compose.yml (VM 109 .7)
|
||||||
/opt/home_stack/docker-compose.yml (VM 109 .7)
|
/opt/home_stack/docker-compose.yml (VM 109 .7)
|
||||||
/opt/audiobookshelf/docker-compose.yml (VM 109 .7)
|
/opt/audiobookshelf/docker-compose.yml (VM 109 .7)
|
||||||
|
/root/compose.yml (hwpve .11 — Authentik server/worker/postgres)
|
||||||
|
/root/docker-compose.yml (NetBird VPS — netbird server/dashboard/proxy, traefik, crowdsec)
|
||||||
/root/.pi/agent/extensions/config.yaml (CT 100 .24)
|
/root/.pi/agent/extensions/config.yaml (CT 100 .24)
|
||||||
/etc/systemd/system/strix-server.service (amdpve .15 — strix-moe)
|
/etc/systemd/system/strix-server.service (amdpve .15 — strix-moe)
|
||||||
/etc/systemd/system/llama-server.service (VM 101 .8, VM 103 .110)
|
/etc/systemd/system/llama-server.service (VM 101 .8, VM 103 .110)
|
||||||
# Hermes agent configs (key enforcement — 2026-07-10)
|
# Hermes agent configs (key enforcement — 2026-07-10)
|
||||||
/root/.hermes/config.yaml (Mumuni inside CT 100, Tanko CT 112, etc.)
|
/home/hermes/.hermes/config.yaml (Mumuni kagentz CT105; Tanko CT112 uses /home/jerome/.hermes)
|
||||||
/root/.config/systemd/user/hermes-gateway.service (Mumuni inside CT 100 — EnvironmentFile fixed)
|
/etc/systemd/system/hermes-gateway.service (Mumuni kagentz CT105 — system unit, User=hermes)
|
||||||
/etc/environment (Mumuni inside CT 100 — LITELLM_API_KEY)
|
/etc/environment (LITELLM_API_KEY — legacy path, Mumuni now keys via Infisical)
|
||||||
```
|
```
|
||||||
|
|
||||||
## MCP Gateway (2026-07-10)
|
## MCP Gateway (2026-07-10)
|
||||||
@@ -220,7 +233,7 @@ When LiteLLM is upgraded to a version supporting per-key MCP grants:
|
|||||||
|
|
||||||
- [ ] All 5 PVE nodes updated, no reboot-loop
|
- [ ] All 5 PVE nodes updated, no reboot-loop
|
||||||
- [ ] All VMs/CTs running post-update
|
- [ ] All VMs/CTs running post-update
|
||||||
- [ ] All Docker containers healthy (VM 109 + CT 116 + CT 117)
|
- [ ] All Docker containers healthy (VM 109 + CT 116 + CT 117 + hwpve .11 + NetBird VPS)
|
||||||
- [ ] LiteLLM inference passing (syslog-auto test)
|
- [ ] LiteLLM inference passing (syslog-auto test)
|
||||||
- [ ] Zulip server + all 3 agents connected
|
- [ ] Zulip server + all 3 agents connected
|
||||||
- [ ] GPU fleet at full capacity (3/3)
|
- [ ] GPU fleet at full capacity (3/3)
|
||||||
|
|||||||
@@ -178,7 +178,7 @@ through its agent wrapper.
|
|||||||
| Agent | Host | Pattern | Keys | Status |
|
| Agent | Host | Pattern | Keys | Status |
|
||||||
|-------|------|---------|------|--------|
|
|-------|------|---------|------|--------|
|
||||||
| abiba | .24 | pi agent wrapper | ABIBA_LITELLM_API_KEY + ABIBA_ZULIP_API_KEY | ✅ vault-backed |
|
| abiba | .24 | pi agent wrapper | ABIBA_LITELLM_API_KEY + ABIBA_ZULIP_API_KEY | ✅ vault-backed |
|
||||||
| mumuni | .24 (CT100 abiba) | Pi Hermes gateway (no systemd) | MUMUNI_LITELLM_API_KEY + MUMUNI_ZULIP_API_KEY | ✅ vault-backed + .env fallback |
|
| mumuni | .14 (kagentz CT105) | systemd unit hermes-gateway.service (user hermes) | MUMUNI_LITELLM_API_KEY + MUMUNI_ZULIP_API_KEY | ✅ vault-backed + .env fallback |
|
||||||
| tanko | .122 | systemd drop-in + while-true wrapper + st.8e848433 (user jerome) | TANKO_LITELLM_API_KEY + TANKO_ZULIP_API_KEY | ✅ vault-backed + .env fallback |
|
| tanko | .122 | systemd drop-in + while-true wrapper + st.8e848433 (user jerome) | TANKO_LITELLM_API_KEY + TANKO_ZULIP_API_KEY | ✅ vault-backed + .env fallback |
|
||||||
| koby | .129 | systemd drop-in + while-true wrapper + st.8e848433 | KOBY_LITELLM_API_KEY, shares TANKO_ZULIP_API_KEY (tanko-bot) | ✅ vault-backed |
|
| koby | .129 | systemd drop-in + while-true wrapper + st.8e848433 | KOBY_LITELLM_API_KEY, shares TANKO_ZULIP_API_KEY (tanko-bot) | ✅ vault-backed |
|
||||||
| koonimo | .114 | systemd drop-in + while-true wrapper + st.8e848433 | KOONIMO_LITELLM_API_KEY + KOONIMO_ZULIP_API_KEY | ✅ vault-backed |
|
| koonimo | .114 | systemd drop-in + while-true wrapper + st.8e848433 | KOONIMO_LITELLM_API_KEY + KOONIMO_ZULIP_API_KEY | ✅ vault-backed |
|
||||||
@@ -257,9 +257,51 @@ reads use per-agent identities. This eliminates the single shared token risk.
|
|||||||
| Agent | .env Keys |
|
| Agent | .env Keys |
|
||||||
|-------|-----------|
|
|-------|-----------|
|
||||||
| Mumuni | MUMUNI_LITELLM_API_KEY, MUMUNI_ZULIP_API_KEY |
|
| Mumuni | MUMUNI_LITELLM_API_KEY, MUMUNI_ZULIP_API_KEY |
|
||||||
| Tanko | TANKO_LITELLM_API_KEY, TANKO_ZULIP_API_KEY |
|
|| Tanko | TANKO_LITELLM_API_KEY, TANKO_ZULIP_API_KEY |
|
||||||
| Koby | (wrapper injects from vault — .env has Telegram token) |
|
|| Koby | (wrapper injects from vault — .env has Telegram token) |
|
||||||
| Koonimo | KOONIMO_LITELLM_API_KEY, KOONIMO_ZULIP_API_KEY |
|
|| Koonimo | KOONIMO_LITELLM_API_KEY, KOONIMO_ZULIP_API_KEY |
|
||||||
|
|| Agent Zero (kagentz .14) | OPENROUTER_API_KEY (direct OpenRouter access) |
|
||||||
|
|
||||||
|
### Agent Zero (kagentz .14) — OpenRouter Integration (2026-09-01)
|
||||||
|
|
||||||
|
Agent Zero runs in Docker on kagentz (CT105) and uses **direct OpenRouter API access**,
|
||||||
|
not via the LiteLLM proxy. This is because Agent Zero's workflow (self-update manager,
|
||||||
|
UI bootstrap, model selection) is built around OpenRouter's native authentication.
|
||||||
|
|
||||||
|
**Key Storage:**
|
||||||
|
- **Container**: `/a0/usr/.env` (line ~72: `API_KEY_OPENROUTER=sk-or-v1-…`)
|
||||||
|
- **Vault**: Infisical secret `OPENROUTER_API_KEY` (project=agents, env=production)
|
||||||
|
- **Fallback**: The container's .env is the primary source; vault sync is optional
|
||||||
|
(unlike fleet agents which require vault injection)
|
||||||
|
|
||||||
|
**Current Key (2026-09-01):**
|
||||||
|
- **Prefix**: `sk-or-v1-0af3f3…`
|
||||||
|
- **User**: `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`
|
||||||
|
- **Plan**: Paid (not free tier)
|
||||||
|
- **Usage**: 0 (as of 2026-09-01)
|
||||||
|
|
||||||
|
**Model Configuration:**
|
||||||
|
- **Preset**: "Cost Efficient" (`/a0/usr/plugins/_model_config/presets.yaml`)
|
||||||
|
- **Model**: `openrouter/moonshotai/kimi-k3`
|
||||||
|
- **API Base**: (empty — uses OpenRouter default)
|
||||||
|
|
||||||
|
**Why not LiteLLM proxy?**
|
||||||
|
Agent Zero's architecture was designed before the fleet adopted the LiteLLM proxy
|
||||||
|
standard. The container runs `/exe/self_update_manager.py` and `/a0/run_ui.py` which
|
||||||
|
directly call OpenRouter via Python's requests library. Converting would require:
|
||||||
|
1. Refactoring all LLM calls to use `litellm` library
|
||||||
|
2. Adding vault wrapper injection
|
||||||
|
3. Updating self_update_manager to use proxy-aware key handling
|
||||||
|
|
||||||
|
**Rotation Procedure:**
|
||||||
|
1. Generate new key in OpenRouter UI
|
||||||
|
2. Update container: `sed -i 's/^API_KEY_OPENROUTER=.*/API_KEY_OPENROUTER=<new_key>/' /a0/usr/.env`
|
||||||
|
3. Update vault: `infisical secrets set OPENROUTER_API_KEY=<new_key> --projectId=agents --env=production`
|
||||||
|
4. Restart container: `sudo docker exec agent-zero supervisorctl restart run_ui`
|
||||||
|
5. Verify: `curl -s https://openrouter.ai/api/v1/auth/key -H "Authorization: Bearer <new_key>"`
|
||||||
|
|
||||||
|
**Related Contract:**
|
||||||
|
- `agent-zero-openrouter-key.prose.md` — Full agent-zero key management contract
|
||||||
|
|
||||||
## Key Rotation Log
|
## Key Rotation Log
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,117 @@
|
|||||||
|
---
|
||||||
|
kind: pattern
|
||||||
|
name: litellm-client-timeouts
|
||||||
|
description: >
|
||||||
|
Standard client timeout and retry policy for ALL agents calling LiteLLM
|
||||||
|
(CT 116, http://192.168.68.116). Created 2026-09-08 after the Sep 6 incident:
|
||||||
|
a backend stall 04:00-06:30 EDT produced 157 client-abandoned 408 failures
|
||||||
|
(82% from abiba-pi, 13 from mumuni) against a backend that was actually
|
||||||
|
succeeding at 20-70s per call once clients stopped giving up. Grounded in
|
||||||
|
measured data: syslog-auto (LiteLLM virtual model, no dedicated GPU — routes
|
||||||
|
to backends) averages 28.8s/request with 25.4s TTFT over 888 calls/24h;
|
||||||
|
nginx already allows 600s (Rule 5, verified 2026-08-09); the gap is entirely
|
||||||
|
client-side. Blast radius if wrong: agents fall back to DeepSeek silently
|
||||||
|
(key/timeout failures present as model degradation, not errors) or abandon
|
||||||
|
healthy-but-slow reasoning calls, fragmenting long tasks.
|
||||||
|
---
|
||||||
|
|
||||||
|
## Maintains
|
||||||
|
|
||||||
|
- client_timeout_standard: "litellm-client-timeouts v1.0 (2026-09-08)"
|
||||||
|
- applies_to: ALL agents and scripts calling http://192.168.68.116 (main model path, auxiliary tasks, health probes, benchmark jobs)
|
||||||
|
- verified_against: Prometheus litellm_* metrics 24h window ending 2026-09-08 ~10:00 EDT; live probe (syslog-auto tiny call 0.56s TTFB, 200 OK); nginx 600s proxy_read_timeout (Rule 5)
|
||||||
|
|
||||||
|
## The measured numbers these values come from
|
||||||
|
|
||||||
|
| Model | avg latency | avg TTFT | p-profile (24h) |
|
||||||
|
|---|---|---|---|
|
||||||
|
| syslog-auto | 28.8s | 25.4s | 68 calls took 30-120s; tail to ~300s under load |
|
||||||
|
| qwen3.6-27B-code | 23.0s | — | same backend class as syslog-auto |
|
||||||
|
| strix-moe | 7.5s | — | Strix Halo, healthy |
|
||||||
|
| gemma-4-12b | 2.6s | — | RTX 5070, healthy |
|
||||||
|
|
||||||
|
Sep 6 incident timeline: failures 04:00-07:00 EDT (0% GPU util = wedged
|
||||||
|
backend), full recovery 07:00-08:00 with ZERO client failures once requests
|
||||||
|
tolerated 20-70s — 392 successful slow calls in the three hours after recovery.
|
||||||
|
LiteLLM's internal queue time is ~0s; the latency is model inference, not
|
||||||
|
proxy queuing.
|
||||||
|
|
||||||
|
## Parameters
|
||||||
|
|
||||||
|
### 1. Primary model path (model.default / custom_providers) — NO client timeout below 300s
|
||||||
|
|
||||||
|
- The default Hermes HTTP timeout (~60s) is TOO SHORT for syslog-auto's healthy
|
||||||
|
28.8s average + 120-300s tail. Every 408 in the incident was a client
|
||||||
|
abandoning a request the backend would have answered.
|
||||||
|
- If the transport exposes a timeout setting for the main model, set it to
|
||||||
|
**300s or more**. If it does not (current Hermes custom-provider path has no
|
||||||
|
timeout knob), that is acceptable ONLY because nginx holds the request for
|
||||||
|
600s — but any wrapper, script, or direct API call you write MUST set its own
|
||||||
|
timeout >= 300s for syslog-auto/qwen-class calls.
|
||||||
|
- Never hardcode a shorter timeout "to fail fast" on this path — failing fast
|
||||||
|
here is what caused the incident.
|
||||||
|
|
||||||
|
### 2. Auxiliary tasks — keep template timeouts, one correction
|
||||||
|
|
||||||
|
- vision: 60s (keep), web_extract: 30s (keep) — gemma-4-12b averages 2.6s;
|
||||||
|
these are fine.
|
||||||
|
- compression: 300s (keep — this was already raised from 60 per gpu-fleet).
|
||||||
|
- **gpu-dense delegation/x_search: set timeout >= 120s.** The RTX 3090
|
||||||
|
(qwen3.6-27B-code backend, 23.0s avg) is the same speed class as
|
||||||
|
syslog-auto; delegation defaults that assume fast responses will 408 the
|
||||||
|
same way.
|
||||||
|
|
||||||
|
### 3. Retry policy — backoff, not repetition
|
||||||
|
|
||||||
|
- On timeout (408) or 5xx: retry up to **2 times** with exponential backoff
|
||||||
|
(**15s, 45s**) before giving up.
|
||||||
|
- Do NOT retry in a tight loop. The Sep 6 spike shape (65 failures in one hour
|
||||||
|
from one key) was a batch job retrying without backoff while the backend was
|
||||||
|
down — it multiplied load during recovery.
|
||||||
|
- On 401/403: do NOT retry — that is a key/permission problem (see
|
||||||
|
litellm-api-keys.prose.md and Rule 11); retrying just spams the log.
|
||||||
|
- On 429: honor the retry-after header if present, else back off 60s.
|
||||||
|
|
||||||
|
### 4. Health probes — identify yourself and time out sanely
|
||||||
|
|
||||||
|
- Probes MUST NOT appear as keyless, model-less failures in the metrics (12
|
||||||
|
such orphans appeared in the incident window and cost investigation time).
|
||||||
|
Send a real model name and use a real (probe-designated) key.
|
||||||
|
- Probe timeout: 30s. A probe that takes longer than 30s IS the alert —
|
||||||
|
report "backend slow (>30s)" rather than hanging.
|
||||||
|
- Probe cadence: at most hourly. The 6h litellm-health cron cadence is the
|
||||||
|
standard; sub-hourly synthetic traffic distorts latency baselines.
|
||||||
|
|
||||||
|
### 5. Batch/benchmark jobs — schedule away from 04:00-07:00 EDT and chunk
|
||||||
|
|
||||||
|
- The incident window showed bulk clients amplifying a backend stall 5:1.
|
||||||
|
- Batch jobs that can tolerate delay: schedule 09:00-17:00 EDT.
|
||||||
|
- Any batch loop over N requests MUST sleep >= 5s between requests and honor
|
||||||
|
the retry policy in section 3.
|
||||||
|
|
||||||
|
## Returns
|
||||||
|
|
||||||
|
- A single standard any agent or script can cite: timeouts >= 300s on the
|
||||||
|
syslog-auto path, >= 120s on gpu-dense delegation, backoff retries (2x,
|
||||||
|
15s/45s), identified probes at 30s/hourly, batch jobs chunked and
|
||||||
|
day-scheduled.
|
||||||
|
- Failure signature recognition: bulk 408s from multiple keys in one window =
|
||||||
|
backend event (check gpu_utilization_percent: 0% = wedged, ~100% = saturated);
|
||||||
|
single-key 408s = that client's timeout is too short.
|
||||||
|
- Cross-references: hermes-config-template.prose.md (Rule 5 nginx 600s,
|
||||||
|
auxiliary timeouts), gpu-fleet.prose.md (compression 300s precedent,
|
||||||
|
stable aliases), litellm-api-keys.prose.md (key/permission failures).
|
||||||
|
|
||||||
|
## Intentionally NOT changed
|
||||||
|
|
||||||
|
- No server-side LiteLLM timeout/cooldown changes proposed — the incident
|
||||||
|
self-recovered and the server is healthy (0.56s live probe); changing
|
||||||
|
server behavior without process-level root cause (CT116 requires root;
|
||||||
|
not reachable from kagentz) would be guessing.
|
||||||
|
- No change to the template's vision/web_extract/compression timeouts —
|
||||||
|
measured data says they are correct.
|
||||||
|
- No per-agent key permission changes — those are litellm-api-keys.prose.md
|
||||||
|
territory (and the open gpu-vision/gemma 403 items are already filed with
|
||||||
|
the key owners).
|
||||||
|
- No model routing changes — syslog-auto's weighted pool behaved correctly
|
||||||
|
throughout the incident.
|
||||||
@@ -1,4 +1,6 @@
|
|||||||
---
|
---
|
||||||
|
report_only_agents:
|
||||||
|
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
|
||||||
kind: responsibility
|
kind: responsibility
|
||||||
name: litellm-self-heal
|
name: litellm-self-heal
|
||||||
status: deployed
|
status: deployed
|
||||||
@@ -22,6 +24,7 @@ description: >
|
|||||||
inference, and agent keys. Applies remediation rules for common failures.
|
inference, and agent keys. Applies remediation rules for common failures.
|
||||||
Reports every action via Zulip DM and Gitea (SyslogSolution/health-logs).
|
Reports every action via Zulip DM and Gitea (SyslogSolution/health-logs).
|
||||||
---
|
---
|
||||||
|
---
|
||||||
|
|
||||||
# LiteLLM Operations — Health Check + Self-Heal
|
# LiteLLM Operations — Health Check + Self-Heal
|
||||||
|
|
||||||
@@ -65,7 +68,15 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
|||||||
|
|
||||||
## LiteLLM Model Surface (ground truth — `/opt/inference-harness/litellm_config.yaml` on CT 116)
|
## LiteLLM Model Surface (ground truth — `/opt/inference-harness/litellm_config.yaml` on CT 116)
|
||||||
|
|
||||||
`model_name`s served: `qwen3.6-27B-code`, `gemma-4-12b`, `qwen3.6-35B-udq4`, `strix-moe`, `gpu-dense`, `gpu-light`, `syslog-auto`.
|
`model_name`s served: `qwen3.6-27B-code`, `gemma-4-12b`, `qwen3.6-35B-udq4`, `strix-moe`, `gpu-dense`, `gpu-light`, `syslog-auto`, `crew-auto` (new 2026-08-20).
|
||||||
|
|
||||||
|
### Context Cap Split (2026-08-20)
|
||||||
|
|
||||||
|
- **Abiba (firstmate)**: 128K uncapped — unlimited context for primary workloads
|
||||||
|
- **Hermes agents** (mumuni, tanko, koby, koonimo): 128K uncapped
|
||||||
|
- **Crewmates** (ops, tune, verify, auth-keys, build): 64K capped — alias `crew-auto` enforces 64K limit
|
||||||
|
|
||||||
|
Preferred implementation: uncap shared pool, add capped alias for crew-only.
|
||||||
|
|
||||||
- `syslog-auto` is a weighted router model: qwen3.6-27B-code (0.55, rpm 500) + qwen3.6-35B-udq4 (0.30, rpm 60) + gemma-4-12b (0.15, rpm 200).
|
- `syslog-auto` is a weighted router model: qwen3.6-27B-code (0.55, rpm 500) + qwen3.6-35B-udq4 (0.30, rpm 60) + gemma-4-12b (0.15, rpm 200).
|
||||||
- `gpu-dense` / `gpu-light` are high-rpm aliases (rpm 500) onto qwen3.6-27B-code / gemma-4-12b respectively.
|
- `gpu-dense` / `gpu-light` are high-rpm aliases (rpm 500) onto qwen3.6-27B-code / gemma-4-12b respectively.
|
||||||
@@ -102,7 +113,7 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
|||||||
|
|
||||||
- **Health-check script** (`/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116): `gpu-fleet` check fails only on **critical** alerts (warnings are informational). Tests `strix-moe` (not `ornith-1.0-35b`).
|
- **Health-check script** (`/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116): `gpu-fleet` check fails only on **critical** alerts (warnings are informational). Tests `strix-moe` (not `ornith-1.0-35b`).
|
||||||
- **GPU monitor** (`/root/scripts/gpu-monitor-server.py` on pi .24): runs as **systemd unit `gpu-monitor.service`** (was bare `&` process). `gpu_count` includes Strix Halo (was 2, now 3). VRAM alert thresholds: warning 93%, critical 97% (raised from 90/95 — 128K context steady-state is ~70% on RTX 3090, not a fault).
|
- **GPU monitor** (`/root/scripts/gpu-monitor-server.py` on pi .24): runs as **systemd unit `gpu-monitor.service`** (was bare `&` process). `gpu_count` includes Strix Halo (was 2, now 3). VRAM alert thresholds: warning 93%, critical 97% (raised from 90/95 — 128K context steady-state is ~70% on RTX 3090, not a fault).
|
||||||
- **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): v2 (2026-07-26) — reads each agent's **agent-specific** `{NAME}_LITELLM_API_KEY` from Infisical vault (not the shared master key). Covers: LiteLLM keys, GPU ports, agent gateways (all 5 agents now SSHa ble), CT liveness (pct status on PVE nodes), config.yaml YAML integrity, wrapper/CLI integrity, vault secret non-emptiness checks. Fleet roster: tanko (.122), mumuni (.24, inside abiba CT100), koby (.129), koonimo (.114), abiba (.24). Legacy `tdunna`/`baggy` replaced with canonical agent hostnames.
|
- **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): v4 (2026-09-10) — vault-backed agents (tanko/koby/koonimo) read their **agent-specific** `{NAME}_LITELLM_API_KEY` from Infisical vault (not the shared master key); abiba (pi agent) reads `LITELLM_API_KEY` from its local `/root/.pi/agent/env.sh` (#735 — moved out of shared `/root/.bashrc`), not from the vault. Abiba is pi-only since the harness purge, so its Hermes config/wrapper/gateway legs are skipped rather than reported as faults; koby is **report-only** (captain's 2026-08-17 ruling) — its findings go to the `--json` `report_only` array and are never counted as fleet failures or repaired, and its CT 111 liveness is probed on storepve (.6). Covers: LiteLLM keys, GPU ports, agent gateways, CT liveness (pct status on PVE nodes), config.yaml YAML integrity, wrapper/CLI integrity, vault secret non-emptiness checks. Every run/report carries the absolute execution path (`script=` + `cwd=`). The current fleet roster is owned by the script changelog (`scripts/agent-health-check.py`); mumuni is no longer probed from this host. Legacy `tdunna`/`baggy` replaced with canonical agent hostnames.
|
||||||
- **Stale keys cleaned**: `daily-infra-report.py` SYNTHETIC_API_KEY was stale (`sk-U_ydi3B` → 401); now reads `LITELLM_MASTER_KEY` from env. Deprecated scripts (`router-original.py`, `router-phase0-backup.py`, `apply-fixes.py`) still reference `sk-syslog-local-master-key` but do not actively poll LiteLLM.
|
- **Stale keys cleaned**: `daily-infra-report.py` SYNTHETIC_API_KEY was stale (`sk-U_ydi3B` → 401); now reads `LITELLM_MASTER_KEY` from env. Deprecated scripts (`router-original.py`, `router-phase0-backup.py`, `apply-fixes.py`) still reference `sk-syslog-local-master-key` but do not actively poll LiteLLM.
|
||||||
|
|
||||||
## Maintains
|
## Maintains
|
||||||
@@ -120,6 +131,7 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
|||||||
- Also wakes on user request
|
- Also wakes on user request
|
||||||
- On failure: re-check after 30s, escalate after 3 consecutive failures
|
- On failure: re-check after 30s, escalate after 3 consecutive failures
|
||||||
|
|
||||||
|
---
|
||||||
---
|
---
|
||||||
|
|
||||||
## Health Check
|
## Health Check
|
||||||
@@ -162,6 +174,7 @@ Determine overall_status from individual check results:
|
|||||||
- "degraded" — 1-2 non-critical checks fail
|
- "degraded" — 1-2 non-critical checks fail
|
||||||
- "down" — critical checks fail
|
- "down" — critical checks fail
|
||||||
|
|
||||||
|
---
|
||||||
---
|
---
|
||||||
|
|
||||||
## Remediation Rules
|
## Remediation Rules
|
||||||
@@ -205,6 +218,7 @@ Escalate → if SSH access unavailable, send Zulip DM
|
|||||||
Router no longer in path so Redis active counters are unused. Rule retained
|
Router no longer in path so Redis active counters are unused. Rule retained
|
||||||
for reference but inactive. If Redis issues occur, check harness-redis container.
|
for reference but inactive. If Redis issues occur, check harness-redis container.
|
||||||
|
|
||||||
|
---
|
||||||
---
|
---
|
||||||
|
|
||||||
## Reporting
|
## Reporting
|
||||||
@@ -227,6 +241,7 @@ top actions, uptime.
|
|||||||
If a fix requires another agent (e.g., Authentik restart), relay sent
|
If a fix requires another agent (e.g., Authentik restart), relay sent
|
||||||
to responsible agent with full context.
|
to responsible agent with full context.
|
||||||
|
|
||||||
|
---
|
||||||
---
|
---
|
||||||
|
|
||||||
## Execution
|
## Execution
|
||||||
|
|||||||
@@ -1,9 +1,12 @@
|
|||||||
---
|
---
|
||||||
|
report_only_agents:
|
||||||
|
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
|
||||||
name: memory-audit-maintenance
|
name: memory-audit-maintenance
|
||||||
kind: responsibility
|
kind: responsibility
|
||||||
description: Shared memory audit and maintenance contract for all Hermes agents (Mumuni, Tanko, Koby, Koonimo). Each agent runs it against its own isolated memory files — no cross-agent access, no shared state. Detects staleness, enforces writer registry, and rotates canary tokens.
|
description: Shared memory audit and maintenance contract for Hermes agents (Mumuni, Koby, Koonimo). Tanko is no longer a Hermes agent (now on DSH/DeepSeek Harness since 2026-08-27) and uses DSH-native memory, so it is excluded from this Hermes roster. Each agent runs it against its own isolated memory files — no cross-agent access, no shared state. Detects staleness, enforces writer registry, and rotates canary tokens.
|
||||||
id: 067NC4KG01RG50R40M30E20918
|
id: 067NC4KG01RG50R40M30E20918
|
||||||
---
|
---
|
||||||
|
---
|
||||||
|
|
||||||
### Goal
|
### Goal
|
||||||
|
|
||||||
@@ -15,11 +18,10 @@ Autonomously audit and reorganize an agent's native memory (MEMORY.md, USER.md,
|
|||||||
|
|
||||||
### Scope
|
### Scope
|
||||||
|
|
||||||
This contract is the **Hermes Agent standard** for memory maintenance. It is shared across all Hermes agents (Mumuni, Tanko, Koby, Koonimo). Each agent runs it against its own memory files only no cross-agent access, no shared state, no shared ledger, no shared canary. The contract is the standard; each agent enforces it independently with fully isolated data.
|
This contract is the **Hermes Agent standard** for memory maintenance. It is shared across Hermes agents (Mumuni, Koby, Koonimo). **Tanko is excluded — it migrated to DSH (DeepSeek Harness) on 2026-08-27 and now uses DSH-native memory (mnemon), not `~/.hermes/memories/`.** Each agent runs it against its own memory files only no cross-agent access, no shared state, no shared ledger, no shared canary. The contract is the standard; each agent enforces it independently with fully isolated data.
|
||||||
|
|
||||||
**Agent Roster:**
|
**Agent Roster (Hermes):**
|
||||||
- Mumuni
|
- Mumuni
|
||||||
- Tanko
|
|
||||||
- Koby (CT 111 / tdunna)
|
- Koby (CT 111 / tdunna)
|
||||||
- Koonimo (CT 113 / baggy)
|
- Koonimo (CT 113 / baggy)
|
||||||
|
|
||||||
@@ -343,4 +345,3 @@ return {
|
|||||||
|
|
||||||
### Per-Agent Notes
|
### Per-Agent Notes
|
||||||
|
|
||||||
Each Hermes agent (Mumuni, Tanko, Tdunna, Baggy) runs this contract against its own `~/.hermes/memories/` directory. The contract is identical across agents, but all data is fully isolated: separate ledgers, separate writer registries, separate canaries. If a new agent is added to the roster, it must be listed in `### Scope` above and given its own isolated memory directory.
|
|
||||||
+31
-6
@@ -1,10 +1,13 @@
|
|||||||
---
|
---
|
||||||
|
report_only_agents:
|
||||||
|
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
|
||||||
kind: pattern
|
kind: pattern
|
||||||
name: memory-fixer
|
name: memory-fixer
|
||||||
description: >
|
description: >
|
||||||
Auto-fix low-hanging fruit in the RA-H OS knowledge graph. No judgment calls — only deterministic Level 1 operations.
|
Auto-fix low-hanging fruit in the RA-H OS knowledge graph. No judgment calls — only deterministic Level 1 operations.
|
||||||
Escalate anything that needs Kwame's input. Executes confirmed Kwame decisions to completion (state + updated_at).
|
Escalate anything that needs Kwame's input. Executes confirmed Kwame decisions to completion (state + updated_at).
|
||||||
version: 2.0.0
|
version: 2.1.0
|
||||||
|
---
|
||||||
---
|
---
|
||||||
|
|
||||||
# Memory Fixer
|
# Memory Fixer
|
||||||
@@ -66,13 +69,15 @@ FROM nodes
|
|||||||
WHERE json_extract(metadata, '$.namespace') IS NULL;
|
WHERE json_extract(metadata, '$.namespace') IS NULL;
|
||||||
```
|
```
|
||||||
|
|
||||||
### 3. Staleness Review Tagging
|
### 3. Staleness Review Tagging (refresh-suggested nodes only)
|
||||||
|
|
||||||
Using the type-based windows from the memory-monitor contract, tag nodes stale beyond their window. **Only process a maximum of 10 nodes per run** to avoid overwhelming Kwame. Prioritize infrastructure first, then dynamic, then ephemeral.
|
Using the type-based windows from the memory-monitor contract, tag nodes stale beyond their window. **Only process a maximum of 10 nodes per run** to avoid overwhelming Kwame. Prioritize infrastructure first, then dynamic, then ephemeral.
|
||||||
|
|
||||||
|
**Archive-suggested nodes are NO LONGER tagged — they are archived outright (see Level 1 fix 4).** Tagging with `[REVIEW: refresh]` applies only to living nodes (infrastructure, deployment, system, system-health, business, philosophy, research, learning, investigation, analysis, project, agent, registry, policy).
|
||||||
|
|
||||||
**Exclusion Rules:**
|
**Exclusion Rules:**
|
||||||
- Nodes with `state` = `review_pending`, `deprecated`, `archived`, or `not_processed` are NOT processed
|
- Nodes with `state` = `review_pending`, `deprecated`, `archived`, or `not_processed` are NOT processed
|
||||||
- Nodes whose `description` already starts with `[REVIEW:` are NOT re-processed
|
- Nodes whose `description` already starts with `[REVIEW:` or `[ARCHIVED]` are NOT re-processed
|
||||||
|
|
||||||
```sql
|
```sql
|
||||||
SELECT id, title, json_extract(metadata, '$.type') as node_type,
|
SELECT id, title, json_extract(metadata, '$.type') as node_type,
|
||||||
@@ -101,9 +106,28 @@ LIMIT 10;
|
|||||||
|
|
||||||
For each identified node, call `updateNode(id, { description: "[REVIEW: action] " + originalDescription })`.
|
For each identified node, call `updateNode(id, { description: "[REVIEW: action] " + originalDescription })`.
|
||||||
|
|
||||||
|
### 4. Stale-Node Archiving (Level 1 — standing Kwame directive, 2026-09-11)
|
||||||
|
|
||||||
|
**Kwame's standing directive: stale nodes CAN be archived by the fixer. No per-batch escalation, no `[REVIEW: archive]` tagging — archive them.**
|
||||||
|
|
||||||
|
For every node whose suggested action is `archive` (i.e. its type is NOT one of the living types in fix 3), archive it in a **single** `updateNode` call:
|
||||||
|
|
||||||
|
```python
|
||||||
|
updateNode(id, {
|
||||||
|
"description": "[ARCHIVED] " + originalDescriptionWithoutReviewTag,
|
||||||
|
"metadata": {"state": "archived"}
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
- `state` transitions **DO work through `updateNode`** (`archived`, and back to `active`). The former "state only accepts processed/not_processed, use SSH" claim was wrong — verified 2026-09-11 by archiving 7 nodes (#61, #373, #388, #465, #475, #526, #1476) over the bridge with `updated_at` auto-bumping. **SSH to the bridge host is a fallback, not a requirement**, and it is blocked from kagentz anyway.
|
||||||
|
- Pass `description` and `metadata` in the **same** call, and always keep the `updates` object nested: `{"id": N, "updates": {…}}`.
|
||||||
|
- Archiving is non-destructive: the node stays in the graph, marked `state: archived` + `[ARCHIVED] ` prefix. **Living nodes (refresh-suggested) are NEVER archived** without a specific Kwame decision — they are the cluster/agent/business canon.
|
||||||
|
|
||||||
|
**Archive candidates are identified by the fix 3 query's `suggested_action = 'archive'` branch** (the `ELSE 'archive'` case: anything not an infrastructure/skill/documentation/strategic/audit type).
|
||||||
|
|
||||||
## Level 2 Escalations (Kwame Decision Required)
|
## Level 2 Escalations (Kwame Decision Required)
|
||||||
|
|
||||||
1. **Stale nodes** flagged with `[REVIEW: …]` — Archive, refresh, or keep?
|
1. **Refresh-suggested stale nodes** flagged with `[REVIEW: refresh]` — refresh or keep? (Archive-suggested nodes are auto-archived under fix 4 and are not escalated.)
|
||||||
2. **Duplicate Nodes** (same title or >70% title overlap) — Merge or keep?
|
2. **Duplicate Nodes** (same title or >70% title overlap) — Merge or keep?
|
||||||
3. **Orphan Nodes >90 days old** — Archive or connect?
|
3. **Orphan Nodes >90 days old** — Archive or connect?
|
||||||
|
|
||||||
@@ -141,7 +165,7 @@ Reply with:
|
|||||||
|
|
||||||
The fixer reads Kwame's previous response and **executes the decision to completion** — it must not leave a node in review-pending forever. Tagging alone is NOT enough; each confirmed decision must also update `state` and `updated_at` so the node drops out of the stale window on the next run.
|
The fixer reads Kwame's previous response and **executes the decision to completion** — it must not leave a node in review-pending forever. Tagging alone is NOT enough; each confirmed decision must also update `state` and `updated_at` so the node drops out of the stale window on the next run.
|
||||||
|
|
||||||
> ⚠️ `updateNode` cannot set `state` to non-standard values (restricted to `processed`/`not_processed`) and cannot add metadata keys. For state transitions and `updated_at` bumps, use **direct SSH + SQLite** on the bridge host:
|
> ⚠️ **Corrected 2026-09-11:** `updateNode` DOES accept `state` changes — `{"updates": {"description": …, "metadata": {"state": "archived"}}}` works over the bridge, and `updated_at` bumps automatically. The old "use direct SSH + SQLite for state transitions" instruction was based on a wrong assumption; SSH is a fallback only (and is blocked from kagentz). Use one `updateNode` call for both the tag and the state.
|
||||||
> ```bash
|
> ```bash
|
||||||
> ssh root@192.168.68.65 "sqlite3 /root/.local/share/RA-H/db/rah.sqlite \"UPDATE nodes SET metadata = json_set(metadata, '$.state', '<state>'), updated_at = datetime('now') WHERE id = <id>;\""
|
> ssh root@192.168.68.65 "sqlite3 /root/.local/share/RA-H/db/rah.sqlite \"UPDATE nodes SET metadata = json_set(metadata, '$.state', '<state>'), updated_at = datetime('now') WHERE id = <id>;\""
|
||||||
> ```
|
> ```
|
||||||
@@ -169,8 +193,9 @@ The result must be 0 rows when all decisions are executed. Report what was done.
|
|||||||
## Checks
|
## Checks
|
||||||
|
|
||||||
- **State integrity:** archived nodes have `state: archived` + `[ARCHIVED]` prefix; kept nodes are `state: active` without a `[REVIEW:]` tag.
|
- **State integrity:** archived nodes have `state: archived` + `[ARCHIVED]` prefix; kept nodes are `state: active` without a `[REVIEW:]` tag.
|
||||||
|
- **Auto-archive applied:** no node should ever be left tagged `[REVIEW: archive]` — that tag is retired. Any `[REVIEW: archive]` found means fix 4 was skipped; archive it and report.
|
||||||
- **No review-pending forever:** after executing Kwame's decisions, `[REVIEW:%` node count must be 0.
|
- **No review-pending forever:** after executing Kwame's decisions, `[REVIEW:%` node count must be 0.
|
||||||
- **Timestamps:** every executed decision bumps `updated_at`, so the node exits the stale window on the next run.
|
- **Timestamps:** every executed decision (and every auto-archive) bumps `updated_at`, so the node exits the stale window on the next run.
|
||||||
|
|
||||||
## Logging
|
## Logging
|
||||||
Every Level 1 fix logged to `~/.hermes/logs/memory-fixer/YYYY-MM-DD.md`
|
Every Level 1 fix logged to `~/.hermes/logs/memory-fixer/YYYY-MM-DD.md`
|
||||||
|
|||||||
@@ -6,7 +6,7 @@ description: >
|
|||||||
delegation, verification, and delivery. Defines when to delegate, which
|
delegation, verification, and delivery. Defines when to delegate, which
|
||||||
worker to use for what, how to handle failures, and the kanban board
|
worker to use for what, how to handle failures, and the kanban board
|
||||||
protocol. Enforces context-window discipline and separation of concerns.
|
protocol. Enforces context-window discipline and separation of concerns.
|
||||||
Runs on Mumuni (inside Abiba CT100, minipve, .24) via Hermes agent (Pi + Hermes Zulip gateway).
|
Runs on Mumuni (kagentz CT105, minipve, .14) via Hermes agent (Zulip gateway via systemd).
|
||||||
version: 1.0.0
|
version: 1.0.0
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -20,7 +20,7 @@ version: 1.0.0
|
|||||||
## Topology
|
## Topology
|
||||||
|
|
||||||
**Cluster:** 5 Proxmox nodes (ocupve, acerpve, minipve, amdpve, storepve)
|
**Cluster:** 5 Proxmox nodes (ocupve, acerpve, minipve, amdpve, storepve)
|
||||||
**Manager:** Mumuni (inside Abiba CT100, minipve, .24) via Hermes agent
|
**Manager:** Mumuni (kagentz CT105, minipve, .14) via Hermes agent
|
||||||
**Workers:** 6 profiles, all running on the same agent — no separate hosts needed
|
**Workers:** 6 profiles, all running on the same agent — no separate hosts needed
|
||||||
|
|
||||||
This contract is infrastructure-agnostic in terms of which nodes are used.
|
This contract is infrastructure-agnostic in terms of which nodes are used.
|
||||||
|
|||||||
@@ -54,7 +54,7 @@ garbled input. When a user types these commands to Abiba:
|
|||||||
|
|
||||||
- **`/approve`** → "Pi doesn't have pending approvals. Commands execute immediately."
|
- **`/approve`** → "Pi doesn't have pending approvals. Commands execute immediately."
|
||||||
- **`/approve session`** → Same response
|
- **`/approve session`** → Same response
|
||||||
- **`/deny`** → Same response + "For Hermes agents (Tanko, Mumuni), these work with their built-in approval system."
|
- **`/deny`** → Same response + "For agents (Mumuni on Hermes, Tanko on DSH), these work with their built-in approval system."
|
||||||
|
|
||||||
This keeps the UX consistent across agents — users can type `/approve` anywhere
|
This keeps the UX consistent across agents — users can type `/approve` anywhere
|
||||||
without getting confused by LLM responses.
|
without getting confused by LLM responses.
|
||||||
|
|||||||
+2
-16
@@ -2,17 +2,6 @@
|
|||||||
kind: responsibility
|
kind: responsibility
|
||||||
name: pm2-self-heal
|
name: pm2-self-heal
|
||||||
description: >
|
description: >
|
||||||
Monitors critical PM2 processes (abiba-zulip, abiba-telegram, gitea-runner,
|
|
||||||
spoton-service, zulip-watchdog) and auto-restarts any that are stopped or
|
|
||||||
errored. Logs every action to Gitea (SyslogSolution/health-logs — not
|
|
||||||
knowledge graph, hard rule) and alerts the owner via
|
|
||||||
Zulip DM on failures.
|
|
||||||
CRITICAL: Never restart abiba-zulip — it runs this contract.
|
|
||||||
AS-BUILT 2026-08-09 (captain ruling, ecosystem is authoritative):
|
|
||||||
gpu-monitor is systemd-managed (gpu-monitor.service) — NOT PM2;
|
|
||||||
gpu-watchdog decommissioned (function folded into gpu-monitor.service);
|
|
||||||
gitea-runner KEPT (online in PM2); abiba-zulip KEPT (online 4d+, the
|
|
||||||
2026-07-04 'removed/decommissioned' note was stale and is removed).
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## Maintains
|
## Maintains
|
||||||
@@ -24,9 +13,6 @@ description: >
|
|||||||
- zulip-watchdog: { status: "online", uptime: string, restarts: number }
|
- zulip-watchdog: { status: "online", uptime: string, restarts: number }
|
||||||
- last_check: timestamp
|
- last_check: timestamp
|
||||||
|
|
||||||
> **Note (2026-07-04, SUPERSEDED 2026-08-09):** `abiba-zulip` remains ONLINE and
|
|
||||||
> is monitored — the decommission note was stale (process re-added; do not treat
|
|
||||||
> it as removed).
|
|
||||||
|
|
||||||
## Continuity
|
## Continuity
|
||||||
|
|
||||||
@@ -63,9 +49,9 @@ description: >
|
|||||||
- If status is "online" → pass
|
- If status is "online" → pass
|
||||||
- If status is "stopped" or "errored" → apply Rule 1
|
- If status is "stopped" or "errored" → apply Rule 1
|
||||||
- If restarts > 5 → alert owner
|
- If restarts > 5 → alert owner
|
||||||
3. **Check abiba-zulip** (self-process, read-only):
|
3. **Check abiba-zulip** (live Zulip bridge, heartbeating):
|
||||||
- If status is "online" → pass, log restarts count
|
- If status is "online" → pass, log restarts count
|
||||||
- If status is "stopped" or "errored" → **DO NOT RESTART** — alert owner immediately
|
- If status is "stopped" or "errored" → restart (`pm2 restart abiba-zulip` — fully restored)
|
||||||
- If restarts > 5 in last hour → alert owner with full diagnostics
|
- If restarts > 5 in last hour → alert owner with full diagnostics
|
||||||
4. **Log results** — Append to `SyslogSolution/health-logs/pm2/{timestamp}.md` in Gitea (not knowledge graph — hard rule)
|
4. **Log results** — Append to `SyslogSolution/health-logs/pm2/{timestamp}.md` in Gitea (not knowledge graph — hard rule)
|
||||||
5. **Alert** — Send Zulip DM to owner if escalation needed (do NOT run pm2 commands during alerting)
|
5. **Alert** — Send Zulip DM to owner if escalation needed (do NOT run pm2 commands during alerting)
|
||||||
|
|||||||
@@ -100,6 +100,35 @@ Edit `build-dashboards.py`, run it, `docker restart harness-grafana`
|
|||||||
### check-targets
|
### check-targets
|
||||||
`curl http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets[] | {job:.labels.job,health}'`
|
`curl http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets[] | {job:.labels.job,health}'`
|
||||||
|
|
||||||
|
### check-health
|
||||||
|
|
||||||
|
**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Prometheus health (bound to 0.0.0.0:9090 on .116)
|
||||||
|
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116:9090/-/healthy
|
||||||
|
# Expected: 200 (Prometheus is up and healthy)
|
||||||
|
|
||||||
|
# Grafana health (bound to 0.0.0.0:3001 on .116)
|
||||||
|
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116:3001/api/health
|
||||||
|
# Expected: 200 (Grafana is up and healthy)
|
||||||
|
|
||||||
|
# Docker Stats exporter (bound to 127.0.0.1:9324 on .116 — must probe from .116 localhost)
|
||||||
|
ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' http://127.0.0.1:9324/metrics"
|
||||||
|
# Expected: 200 (docker-stats-exporter is up and responding)
|
||||||
|
|
||||||
|
# PVE exporter (bound to 127.0.0.1:9221 on .116 — must probe from .116 localhost)
|
||||||
|
ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' http://127.0.0.1:9221/metrics"
|
||||||
|
# Expected: 200 (pve-exporter is up and responding)
|
||||||
|
```
|
||||||
|
|
||||||
|
**Report format**: Begin every report with the **absolute path the probe executed
|
||||||
|
from** (`pwd -P`, or the script's absolute path) so a stale-consumer report is
|
||||||
|
distinguishable from a real fault at read time. Summarize actual results from
|
||||||
|
each probe. If any probe returns non-200, flag as alert.
|
||||||
|
|
||||||
|
**Note**: Docker Stats and PVE Exporter are bound to 127.0.0.1 (localhost-only) so they must be probed from .116 via SSH. Prometheus and Grafana are bound to 0.0.0.0 so they can be probed from the LAN.
|
||||||
|
|
||||||
### restart-exporter
|
### restart-exporter
|
||||||
`cd /opt/monitoring && docker compose restart pve-exporter docker-stats`
|
`cd /opt/monitoring && docker compose restart pve-exporter docker-stats`
|
||||||
|
|
||||||
|
|||||||
+285
-61
@@ -1,6 +1,6 @@
|
|||||||
#!/usr/bin/env python3
|
#!/usr/bin/env python3
|
||||||
"""
|
"""
|
||||||
/root/scripts/agent-health-check.py — Consolidated Agent Health Verification v2
|
/root/scripts/agent-health-check.py — Consolidated Agent Health Verification v4
|
||||||
|
|
||||||
Verifies: LiteLLM keys (agent-specific), GPU port conflicts, agent Zulip streaming,
|
Verifies: LiteLLM keys (agent-specific), GPU port conflicts, agent Zulip streaming,
|
||||||
gateway liveness, gateway log health, CT liveness, config YAML integrity,
|
gateway liveness, gateway log health, CT liveness, config YAML integrity,
|
||||||
@@ -17,11 +17,42 @@ Changelog:
|
|||||||
v2 (2026-07-26): Added CT liveness, config validation, wrapper integrity,
|
v2 (2026-07-26): Added CT liveness, config validation, wrapper integrity,
|
||||||
vault secret emptiness check. Fixed Koby/Koonimo SSH hosts and agent key
|
vault secret emptiness check. Fixed Koby/Koonimo SSH hosts and agent key
|
||||||
name format ({NAME}_LITELLM_API_KEY not LITELLM_API_KEY_{NAME}).
|
name format ({NAME}_LITELLM_API_KEY not LITELLM_API_KEY_{NAME}).
|
||||||
Fleet roster: tanko (.122), mumuni (.24, inside abiba CT100), koby (.129), koonimo (.114),
|
Fleet roster: tanko (.122), koby (.129), koonimo (.114), abiba (.24).
|
||||||
abiba (.24).
|
(v2 also carried a mumuni probe; see v5 — mumuni is no longer probed: she
|
||||||
|
moved to her own container, kagentz CT 105 / .14, and is monitored there.)
|
||||||
|
v3 (2026-09-08): GPU unit repoint verified live (.8 llama-chat-api.service,
|
||||||
|
.110 llama-server.service, .15 strix-server.service) — .8 was probing a stale
|
||||||
|
llama-server unit that reads inactive, producing false UNREACHABLE legs.
|
||||||
|
systemctl is-active no longer swallows non-zero exit as SSH failure.
|
||||||
|
Fixed UnboundLocalError on the abiba/koonimo gateway leg (pid unbound in the
|
||||||
|
summary f-string). Abiba's LiteLLM key now comes from /root/.pi/agent/env.sh
|
||||||
|
(#735 agent separation; creds moved out of shared /root/.bashrc).
|
||||||
|
v4 (2026-09-10): probe-drift round 2 (prose-contracts follow-up to #65/#66/#68).
|
||||||
|
abiba declared pi-only runtime — Hermes-era config/wrapper/gateway checks are
|
||||||
|
skipped (harness purge). koby declared report_only per the captain's
|
||||||
|
2026-08-17 ruling: every koby leg is detected and reported, never counted as a
|
||||||
|
fleet failure and never repaired. koby's PVE mapping corrected to storepve
|
||||||
|
(CT 111 tdunna lives on .6 — the old amdpve mapping produced a false
|
||||||
|
ct-unreachable). The wrapper infisical-path check had two stale-expectation
|
||||||
|
bugs: it read only the first 20 lines of the wrapper, so koonimo (whose
|
||||||
|
wrapper does reference /usr/bin/infisical, just past line 20) was falsely
|
||||||
|
FAILed as "path may be wrong"; and it treated the absence of any infisical
|
||||||
|
reference as a fault, though koby's wrapper sources the key from
|
||||||
|
~/.hermes/.env and never invokes infisical. The check now reads the full
|
||||||
|
wrapper body, accepts a no-infisical wrapper, and verifies that any absolute
|
||||||
|
infisical path the wrapper references actually exists. Report-only findings
|
||||||
|
are surfaced in a machine-readable `report_only` array in --json output,
|
||||||
|
separate from `failures`. Every run prints absolute execution provenance
|
||||||
|
(script + cwd) in the header, in the cron ALERT line, and in --json output so
|
||||||
|
a stale-consumer report is distinguishable from a fault at read time.
|
||||||
|
v5 (2026-09-10): roster correction only, no behavior change. mumuni was removed
|
||||||
|
from the AGENTS dict when she moved off this host onto her own container
|
||||||
|
(kagentz CT 105 on minipve, .14, dedicated `hermes` user) and is monitored
|
||||||
|
from her side. This script must not probe mumuni or .24 — the v2 changelog
|
||||||
|
roster line was the last reference still placing her at .24 / CT100.
|
||||||
"""
|
"""
|
||||||
|
|
||||||
import subprocess, json, sys, os, time
|
import subprocess, json, sys, os, time, re, io, contextlib
|
||||||
from datetime import datetime
|
from datetime import datetime
|
||||||
|
|
||||||
LITELLM = "http://192.168.68.116:80"
|
LITELLM = "http://192.168.68.116:80"
|
||||||
@@ -39,19 +70,56 @@ PVE_NODES = {
|
|||||||
|
|
||||||
# Agent definitions: ct, host, user, pve_node, vault_key_name
|
# Agent definitions: ct, host, user, pve_node, vault_key_name
|
||||||
AGENTS = {
|
AGENTS = {
|
||||||
"tanko": {"ct": 112, "host": "192.168.68.122", "user": "jerome", "pve": "amdpve", "vault_key": "TANKO_LITELLM_API_KEY"},
|
"tanko": {"ct": 112, "host": "192.168.68.122", "user": "jerome", "pve": "amdpve", "vault_key": "TANKO_LITELLM_API_KEY", "runtime": "dsh"},
|
||||||
"abiba": {"ct": 100, "host": "192.168.68.24", "user": "root", "pve": "minipve", "vault_key": None}, # Pi agent + Mumuni Zulip, no vault key
|
# abiba = pi agent (.24) — no vault key; its LiteLLM key is read from its
|
||||||
"koby": {"ct": 111, "host": "192.168.68.129", "user": "root", "pve": "amdpve", "vault_key": "KOBY_LITELLM_API_KEY"},
|
# local env file (key_env below), not from the shared vault or .bashrc.
|
||||||
|
# runtime=pi: abiba has run pi-only since the harness purge. There is no
|
||||||
|
# Hermes gateway, no ~/.hermes/config.yaml and no hermes CLI wrapper on .24
|
||||||
|
# (the /root/.local/bin/hermes symlink is dangling), so the Hermes-era
|
||||||
|
# config/wrapper/gateway legs are skipped rather than reported as faults.
|
||||||
|
"abiba": {"ct": 100, "host": "192.168.68.24", "user": "root", "pve": "minipve",
|
||||||
|
"vault_key": None, "runtime": "pi",
|
||||||
|
"key_env": {"file": "/root/.pi/agent/env.sh", "var": "LITELLM_API_KEY"}},
|
||||||
|
# koby = report-only (captain's 2026-08-17 ruling, Rule 17): detect and
|
||||||
|
# report, NEVER repair, and never count against fleet failures. CT 111
|
||||||
|
# (tdunna) lives on storepve (.6) — verified live 2026-09-10; the previous
|
||||||
|
# amdpve mapping made `pct status 111` fail and read as ct-unreachable.
|
||||||
|
"koby": {"ct": 111, "host": "192.168.68.129", "user": "root", "pve": "storepve", "vault_key": "KOBY_LITELLM_API_KEY", "report_only": True},
|
||||||
"koonimo": {"ct": 113, "host": "192.168.68.114", "user": "root", "pve": "amdpve", "vault_key": "KOONIMO_LITELLM_API_KEY"},
|
"koonimo": {"ct": 113, "host": "192.168.68.114", "user": "root", "pve": "amdpve", "vault_key": "KOONIMO_LITELLM_API_KEY"},
|
||||||
}
|
}
|
||||||
|
|
||||||
|
# Systemd units verified live 2026-09-08 (systemctl list-units on each host):
|
||||||
|
# .8 rtx3090 (gpu-dense) -> llama-chat-api.service (active; the old
|
||||||
|
# llama-server.service unit file is stale/inactive — probing it read as
|
||||||
|
# UNREACHABLE for a healthy process)
|
||||||
|
# .110 rtx5070 (ocu-llm VM) -> llama-server.service (active)
|
||||||
|
# .15 strixhalo (amdpve) -> strix-server.service (active)
|
||||||
GPU_HOSTS = {
|
GPU_HOSTS = {
|
||||||
"gpu-rtx3090 (.8)": {"host": "192.168.68.8", "port": 8080, "service": "llama-server"},
|
"gpu-rtx3090 (.8)": {"host": "192.168.68.8", "port": 8080, "service": "llama-chat-api.service"},
|
||||||
"gpu-rtx5070 (.110)": {"host": "192.168.68.110", "port": 8080, "service": "llama-server"},
|
"gpu-rtx5070 (.110)": {"host": "192.168.68.110", "port": 8080, "service": "llama-server.service"},
|
||||||
"gpu-strixhalo (.15)": {"host": "192.168.68.15", "port": 8080, "service": "strix-server"},
|
"gpu-strixhalo (.15)": {"host": "192.168.68.15", "port": 8080, "service": "strix-server.service"},
|
||||||
}
|
}
|
||||||
|
|
||||||
FAIL = []
|
FAIL = []
|
||||||
|
REPORT_ONLY = []
|
||||||
|
|
||||||
|
|
||||||
|
def _fail(key, agent_name=None):
|
||||||
|
"""Record a failure, except for report-only agents.
|
||||||
|
|
||||||
|
Koby is report-only per the captain's 2026-08-17 ruling (Rule 17): its legs
|
||||||
|
are detected and reported, never repaired and never counted as fleet
|
||||||
|
failures. A red fleet alert on a known report-only leg is a false alarm.
|
||||||
|
Report-only findings are tracked separately so --json consumers can still
|
||||||
|
see them without them counting as fleet failures. Any non-report-only agent
|
||||||
|
(or a leg with no agent, e.g. GPU hosts) records normally.
|
||||||
|
"""
|
||||||
|
if agent_name and AGENTS.get(agent_name, {}).get("report_only"):
|
||||||
|
REPORT_ONLY.append(key)
|
||||||
|
print(f" 🔍 report-only ({agent_name}): {key} — reported, not counted/repaired")
|
||||||
|
return
|
||||||
|
FAIL.append(key)
|
||||||
|
|
||||||
|
|
||||||
INFISICAL_TOKEN = os.environ.get("INFISICAL_TOKEN")
|
INFISICAL_TOKEN = os.environ.get("INFISICAL_TOKEN")
|
||||||
INFISICAL_API_URL = os.environ.get("INFISICAL_API_URL", "https://vault.sysloggh.net")
|
INFISICAL_API_URL = os.environ.get("INFISICAL_API_URL", "https://vault.sysloggh.net")
|
||||||
@@ -161,11 +229,46 @@ def _get_agent_key(agent_name, vault_key_name):
|
|||||||
|
|
||||||
return None
|
return None
|
||||||
|
|
||||||
# Inject keys from vault for each agent
|
|
||||||
for agent_name in AGENTS:
|
def _read_env_export(path, var):
|
||||||
info = AGENTS[agent_name]
|
"""Parse `export VAR=value` (or `VAR=value`) out of a local env file.
|
||||||
key = _get_agent_key(agent_name, info.get("vault_key"))
|
|
||||||
AGENTS[agent_name]["key"] = key
|
#735 agent separation (2026-09-06): agent creds moved out of the shared
|
||||||
|
/root/.bashrc into per-agent env files under /root/.pi/agent/ (bashrc's
|
||||||
|
source line keeps abiba shells resolving them, but the file of record is
|
||||||
|
env.sh). Do NOT fall back to /root/.bashrc here: desktop (.200) SSH
|
||||||
|
sessions override LITELLM_API_KEY with mumuni's key, so sourcing bashrc
|
||||||
|
would validate the wrong identity.
|
||||||
|
"""
|
||||||
|
try:
|
||||||
|
with open(os.path.expanduser(path)) as _f:
|
||||||
|
for line in _f:
|
||||||
|
line = line.strip()
|
||||||
|
if not (line.startswith("export " + var + "=") or line.startswith(var + "=")):
|
||||||
|
continue
|
||||||
|
value = line.split("=", 1)[1].strip().strip('"').strip("'")
|
||||||
|
if value:
|
||||||
|
return value
|
||||||
|
except (OSError, UnicodeDecodeError):
|
||||||
|
pass
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def load_agent_keys():
|
||||||
|
"""Populate AGENTS[*]["key"] from the vault or the agent's local env file.
|
||||||
|
|
||||||
|
Called from main(), not at import: keeping this out of module scope lets the
|
||||||
|
module be imported (and unit tested) without live vault/SSH access. Vault
|
||||||
|
format is {NAME}_LITELLM_API_KEY (project 322fceab-39da-4854-a55a-568e76c0f13f,
|
||||||
|
env prod); abiba has no vault key and reads LITELLM_API_KEY from its local
|
||||||
|
/root/.pi/agent/env.sh (moved there from /root/.bashrc in #735).
|
||||||
|
"""
|
||||||
|
for agent_name in AGENTS:
|
||||||
|
info = AGENTS[agent_name]
|
||||||
|
key = _get_agent_key(agent_name, info.get("vault_key"))
|
||||||
|
if not key and info.get("key_env"):
|
||||||
|
key = _read_env_export(info["key_env"]["file"], info["key_env"]["var"])
|
||||||
|
AGENTS[agent_name]["key"] = key
|
||||||
|
|
||||||
|
|
||||||
# ═══════════════════════════════════════════════════════════════════
|
# ═══════════════════════════════════════════════════════════════════
|
||||||
@@ -176,8 +279,8 @@ def check_keys():
|
|||||||
for name, agent in AGENTS.items():
|
for name, agent in AGENTS.items():
|
||||||
key = agent.get("key")
|
key = agent.get("key")
|
||||||
if not key:
|
if not key:
|
||||||
print(f" ❌ {name}: NO KEY FOUND (vault empty or unreachable)")
|
print(f" ❌ {name}: NO KEY FOUND (vault/env empty or unreachable)")
|
||||||
FAIL.append(f"key:{name}:no-key")
|
_fail(f"key:{name}:no-key", name)
|
||||||
continue
|
continue
|
||||||
data = http_json(f"{LITELLM}/v1/models",
|
data = http_json(f"{LITELLM}/v1/models",
|
||||||
headers={"Authorization": f"Bearer {key}"})
|
headers={"Authorization": f"Bearer {key}"})
|
||||||
@@ -186,11 +289,11 @@ def check_keys():
|
|||||||
print(f" ✅ {name}: key valid → {model}")
|
print(f" ✅ {name}: key valid → {model}")
|
||||||
else:
|
else:
|
||||||
print(f" ❌ {name}: KEY FAILURE — auth rejected or unreachable")
|
print(f" ❌ {name}: KEY FAILURE — auth rejected or unreachable")
|
||||||
FAIL.append(f"key:{name}")
|
_fail(f"key:{name}", name)
|
||||||
|
|
||||||
|
|
||||||
# ═══════════════════════════════════════════════════════════════════
|
# ═══════════════════════════════════════════════════════════════════
|
||||||
# CHECK 2: GPU Port Conflict Detection (unchanged)
|
# CHECK 2: GPU Port Conflict Detection (unit names verified live 2026-09-08)
|
||||||
# ═══════════════════════════════════════════════════════════════════
|
# ═══════════════════════════════════════════════════════════════════
|
||||||
|
|
||||||
def check_gpu_ports():
|
def check_gpu_ports():
|
||||||
@@ -199,7 +302,11 @@ def check_gpu_ports():
|
|||||||
port = gpu["port"]
|
port = gpu["port"]
|
||||||
svc = gpu["service"]
|
svc = gpu["service"]
|
||||||
|
|
||||||
svc_status = ssh(host, f"systemctl is-active {svc}")
|
# `systemctl is-active` exits non-zero when the unit is inactive or
|
||||||
|
# missing, which the ssh() helper would swallow as an SSH failure and
|
||||||
|
# report as UNREACHABLE. `|| true` keeps the real state word so we can
|
||||||
|
# tell "unit inactive" from "host unreachable".
|
||||||
|
svc_status = ssh(host, f"systemctl is-active {svc} || true")
|
||||||
port_owner = ssh(host, f"ss -tlnp 2>/dev/null | grep -Po ':{port}\\s+.*pid=\\K[0-9]+' | head -1")
|
port_owner = ssh(host, f"ss -tlnp 2>/dev/null | grep -Po ':{port}\\s+.*pid=\\K[0-9]+' | head -1")
|
||||||
|
|
||||||
if not svc_status:
|
if not svc_status:
|
||||||
@@ -238,20 +345,49 @@ def check_agents():
|
|||||||
host = agent.get("host")
|
host = agent.get("host")
|
||||||
user = agent.get("user")
|
user = agent.get("user")
|
||||||
ct = agent["ct"]
|
ct = agent["ct"]
|
||||||
|
report_only = agent.get("report_only", False)
|
||||||
|
|
||||||
|
# Tanko runs on DSH (DeepSeek Harness) since 2026-08-27 — it no longer runs a
|
||||||
|
# Hermes gateway, so skip the Hermes gateway/state/streaming/journal checks.
|
||||||
|
# Non-Hermes runtimes have no gateway to probe. dsh = Tanko since
|
||||||
|
# 2026-08-27; pi = abiba since the harness purge (.24 is pi-only).
|
||||||
|
if agent.get("runtime") in ("dsh", "pi"):
|
||||||
|
is_dsh = agent.get("runtime") == "dsh"
|
||||||
|
label = "DSH (DeepSeek Harness)" if is_dsh else "pi-only runtime"
|
||||||
|
since = "since 2026-08-27" if is_dsh else "since the harness purge"
|
||||||
|
live = ssh(host, "true", user=user)
|
||||||
|
print(f" {'✅' if live is not None else '❌'} {name}: {label} — "
|
||||||
|
f"no Hermes gateway {since} (CT {ct}, SSH {'OK' if live is not None else 'FAIL'})")
|
||||||
|
if live is None:
|
||||||
|
_fail(f"unreachable:{name}", name)
|
||||||
|
continue
|
||||||
|
|
||||||
if not host or not user:
|
if not host or not user:
|
||||||
print(f" ⬜ {name} (CT {ct}): cannot SSH — skip liveness check")
|
print(f" ⬜ {name} (CT {ct}): cannot SSH — skip liveness check")
|
||||||
continue
|
continue
|
||||||
|
|
||||||
# Gateway process
|
# Resolve the Hermes gateway PID once, before the report-only branch:
|
||||||
pid = ssh(host, "pgrep -f 'hermes_cli.main gateway run' | grep -v infisical | head -1", user=user)
|
# the summary line below renders `pid`, and it used to be bound only in
|
||||||
|
# the report-only path — leaving it unbound on the abiba/koonimo path
|
||||||
|
# raised UnboundLocalError and crashed the whole check. Agents without
|
||||||
|
# a gateway get pid=?.
|
||||||
|
pid = ssh(host, "pgrep -f '[h]ermes_cli.main gateway run' | grep -v infisical | head -1", user=user)
|
||||||
if not pid:
|
if not pid:
|
||||||
# Try alternate binary name
|
pid = ssh(host, "pgrep -f '[h]ermes.*gateway' | grep -v infisical | grep -v bash | head -1", user=user)
|
||||||
pid = ssh(host, "pgrep -f 'hermes.*gateway' | grep -v infisical | grep -v bash | head -1", user=user)
|
|
||||||
if not pid:
|
if not pid:
|
||||||
print(f" ❌ {name}: GATEWAY NOT RUNNING")
|
pid = "?"
|
||||||
FAIL.append(f"gateway-down:{name}")
|
|
||||||
continue
|
# ⛔ KOBY IS NEVER REPAIRED — diagnostic only
|
||||||
|
if report_only:
|
||||||
|
print(f" 🔍 {name}: REPORT-ONLY mode (diagnostic only, no repairs on .129)")
|
||||||
|
# Still check gateway status for reporting purposes
|
||||||
|
if pid == "?":
|
||||||
|
print(f" ⚠️ {name}: GATEWAY NOT RUNNING (reported only)")
|
||||||
|
_fail(f"gateway-down:{name}", name)
|
||||||
|
continue
|
||||||
|
else:
|
||||||
|
print(f" ✅ {name}: gateway running (pid={pid}, report-only mode)")
|
||||||
|
continue # Skip the rest of the check for Koby
|
||||||
|
|
||||||
# Gateway state file
|
# Gateway state file
|
||||||
state = ssh(host, "cat ~/.hermes/gateway_state.json 2>/dev/null", user=user)
|
state = ssh(host, "cat ~/.hermes/gateway_state.json 2>/dev/null", user=user)
|
||||||
@@ -310,12 +446,12 @@ def check_ct_liveness():
|
|||||||
status = ssh(pve_ip, f"pct status {ct} 2>/dev/null", user="root")
|
status = ssh(pve_ip, f"pct status {ct} 2>/dev/null", user="root")
|
||||||
if not status:
|
if not status:
|
||||||
print(f" ❌ {name} (CT {ct} on {pve_node}): PVE UNREACHABLE")
|
print(f" ❌ {name} (CT {ct} on {pve_node}): PVE UNREACHABLE")
|
||||||
FAIL.append(f"ct-unreachable:{name}:{pve_ip}")
|
_fail(f"ct-unreachable:{name}:{pve_ip}", name)
|
||||||
elif "running" in status:
|
elif "running" in status:
|
||||||
print(f" ✅ {name} (CT {ct} on {pve_node}): running")
|
print(f" ✅ {name} (CT {ct} on {pve_node}): running")
|
||||||
elif "stopped" in status:
|
elif "stopped" in status:
|
||||||
print(f" ❌ {name} (CT {ct} on {pve_node}): STOPPED")
|
print(f" ❌ {name} (CT {ct} on {pve_node}): STOPPED")
|
||||||
FAIL.append(f"ct-stopped:{name}")
|
_fail(f"ct-stopped:{name}", name)
|
||||||
else:
|
else:
|
||||||
print(f" ⚠️ {name} (CT {ct} on {pve_node}): {status.strip()}")
|
print(f" ⚠️ {name} (CT {ct} on {pve_node}): {status.strip()}")
|
||||||
|
|
||||||
@@ -327,6 +463,13 @@ def check_ct_liveness():
|
|||||||
def check_config_integrity():
|
def check_config_integrity():
|
||||||
"""Verify agent config.yaml parses as valid YAML."""
|
"""Verify agent config.yaml parses as valid YAML."""
|
||||||
for name, agent in AGENTS.items():
|
for name, agent in AGENTS.items():
|
||||||
|
# Tanko runs on DSH (DeepSeek Harness) since 2026-08-27 — no Hermes config.yaml.
|
||||||
|
if agent.get("runtime") == "dsh":
|
||||||
|
print(f" ⏭️ {name}: DSH — no Hermes config.yaml since 2026-08-27")
|
||||||
|
continue
|
||||||
|
if agent.get("runtime") == "pi":
|
||||||
|
print(f" ⏭️ {name}: pi-only runtime — no Hermes config.yaml since the harness purge")
|
||||||
|
continue
|
||||||
host = agent.get("host")
|
host = agent.get("host")
|
||||||
user = agent.get("user")
|
user = agent.get("user")
|
||||||
if not host or not user:
|
if not host or not user:
|
||||||
@@ -341,21 +484,48 @@ def check_config_integrity():
|
|||||||
user=user)
|
user=user)
|
||||||
if not yaml_ok:
|
if not yaml_ok:
|
||||||
print(f" ❌ {name}: SSH UNREACHABLE (config check skipped)")
|
print(f" ❌ {name}: SSH UNREACHABLE (config check skipped)")
|
||||||
FAIL.append(f"config-unreachable:{name}")
|
_fail(f"config-unreachable:{name}", name)
|
||||||
elif "OK" in yaml_ok:
|
elif "OK" in yaml_ok:
|
||||||
print(f" ✅ {name}: config.yaml valid YAML")
|
print(f" ✅ {name}: config.yaml valid YAML")
|
||||||
else:
|
else:
|
||||||
print(f" ❌ {name}: config.yaml YAML ERROR — {yaml_ok[:120]}")
|
print(f" ❌ {name}: config.yaml YAML ERROR — {yaml_ok[:120]}")
|
||||||
FAIL.append(f"config-yaml-error:{name}")
|
_fail(f"config-yaml-error:{name}", name)
|
||||||
|
|
||||||
|
|
||||||
# ═══════════════════════════════════════════════════════════════════
|
# ═══════════════════════════════════════════════════════════════════
|
||||||
# CHECK 6: Wrapper/CLI Integrity (NEW)
|
# CHECK 6: Wrapper/CLI Integrity (NEW)
|
||||||
# ═══════════════════════════════════════════════════════════════════
|
# ═══════════════════════════════════════════════════════════════════
|
||||||
|
|
||||||
|
def _infisical_invocation_paths(wrapper_body):
|
||||||
|
"""Absolute infisical paths the wrapper actually invokes.
|
||||||
|
|
||||||
|
Only executed (non-comment) lines count, and only a path followed by a real
|
||||||
|
infisical subcommand (e.g. `/usr/bin/infisical run`) is treated as an
|
||||||
|
invocation. A note such as `# migrated from /usr/local/bin/infisical` is
|
||||||
|
prose, not a call, so it must not manufacture a dangling-path false alarm.
|
||||||
|
"""
|
||||||
|
paths = []
|
||||||
|
for line in wrapper_body.splitlines():
|
||||||
|
code = line.split("#", 1)[0]
|
||||||
|
for _m in re.finditer(
|
||||||
|
r"(/[A-Za-z0-9._/-]*infisical)\s+(?:run|export|secrets|login|logout)\b",
|
||||||
|
code,
|
||||||
|
):
|
||||||
|
if _m.group(1) not in paths:
|
||||||
|
paths.append(_m.group(1))
|
||||||
|
return paths
|
||||||
|
|
||||||
|
|
||||||
def check_wrapper_integrity():
|
def check_wrapper_integrity():
|
||||||
"""Verify the hermes CLI wrapper exists and can reach hermes-real."""
|
"""Verify the hermes CLI wrapper exists and can reach hermes-real."""
|
||||||
for name, agent in AGENTS.items():
|
for name, agent in AGENTS.items():
|
||||||
|
# Tanko runs on DSH (DeepSeek Harness) since 2026-08-27 — no hermes CLI wrapper.
|
||||||
|
if agent.get("runtime") == "dsh":
|
||||||
|
print(f" ⏭️ {name}: DSH — no hermes CLI wrapper since 2026-08-27")
|
||||||
|
continue
|
||||||
|
if agent.get("runtime") == "pi":
|
||||||
|
print(f" ⏭️ {name}: pi-only runtime — no hermes CLI wrapper since the harness purge")
|
||||||
|
continue
|
||||||
host = agent.get("host")
|
host = agent.get("host")
|
||||||
user = agent.get("user")
|
user = agent.get("user")
|
||||||
if not host or not user:
|
if not host or not user:
|
||||||
@@ -369,24 +539,56 @@ def check_wrapper_integrity():
|
|||||||
wrapper = ssh(host, "which hermes 2>/dev/null; command -v hermes 2>/dev/null", user=user)
|
wrapper = ssh(host, "which hermes 2>/dev/null; command -v hermes 2>/dev/null", user=user)
|
||||||
if not wrapper:
|
if not wrapper:
|
||||||
print(f" ❌ {name}: NO HERMES CLI WRAPPER FOUND")
|
print(f" ❌ {name}: NO HERMES CLI WRAPPER FOUND")
|
||||||
FAIL.append(f"wrapper-missing:{name}")
|
_fail(f"wrapper-missing:{name}", name)
|
||||||
continue
|
continue
|
||||||
else:
|
else:
|
||||||
print(f" ⚠️ {name}: hermes at {wrapper.strip()} (not ~/.local/bin/hermes)")
|
print(f" ⚠️ {name}: hermes at {wrapper.strip()} (not ~/.local/bin/hermes)")
|
||||||
|
|
||||||
# Check wrapper has correct infisical path
|
# Credential-injection mechanism. The Hermes-era wrapper injected creds
|
||||||
infisical_path_valid = ssh(host,
|
# with `/usr/bin/infisical run`, but the mechanism is not required to be
|
||||||
"head -20 /root/.local/bin/hermes 2>/dev/null | grep -q '/usr/bin/infisical' && echo OK || echo MISS",
|
# infisical at all: koby's wrapper sources the key from ~/.hermes/.env
|
||||||
user=user)
|
# and never mentions infisical, which is valid. The old check read only
|
||||||
if infisical_path_valid == "MISS":
|
# the first 20 lines, so koonimo's wrapper — which DOES reference
|
||||||
# Check if infisical exists on path
|
# /usr/bin/infisical, just past line 20 — false-failed as "path may be
|
||||||
inf_actual = ssh(host, "command -v infisical 2>/dev/null", user=user)
|
# wrong". Read the full body, accept a no-infisical wrapper, and verify
|
||||||
if not inf_actual:
|
# the absolute infisical path(s) the wrapper actually invokes. Only
|
||||||
print(f" ❌ {name}: INFISICAL NOT INSTALLED (wrapper broken)")
|
# executed (non-comment) lines count: a comment or dead prose mentioning
|
||||||
FAIL.append(f"wrapper-no-infisical:{name}")
|
# a removed path (litellm-api-keys.prose.md documents
|
||||||
|
# `rm -f /usr/local/bin/infisical`) must neither produce a dangling path
|
||||||
|
# nor trigger the PATH check — it is not an invocation.
|
||||||
|
wrapper_body = ssh(host, "cat /root/.local/bin/hermes 2>/dev/null", user=user) or ""
|
||||||
|
wrapper_code = "\n".join(line.split("#", 1)[0] for line in wrapper_body.splitlines())
|
||||||
|
invoked_paths = _infisical_invocation_paths(wrapper_body)
|
||||||
|
if "infisical" in wrapper_code:
|
||||||
|
if invoked_paths:
|
||||||
|
missing = []
|
||||||
|
for _p in invoked_paths:
|
||||||
|
_exists = ssh(host, f"test -x {_p} && echo OK || echo MISS", user=user)
|
||||||
|
if not _exists or _exists.strip().splitlines()[-1] != "OK":
|
||||||
|
missing.append(_p)
|
||||||
|
if len(missing) == len(invoked_paths):
|
||||||
|
inf_actual = ssh(host, "command -v infisical 2>/dev/null", user=user)
|
||||||
|
suffix = f" (infisical at {inf_actual})" if inf_actual else ""
|
||||||
|
print(f" ❌ {name}: wrapper invokes infisical via missing path(s) "
|
||||||
|
f"{', '.join(missing)}{suffix}")
|
||||||
|
_fail(f"wrapper-infisical-path:{name}", name)
|
||||||
|
elif missing:
|
||||||
|
print(f" ⚠️ {name}: wrapper has an unused/missing infisical path "
|
||||||
|
f"({', '.join(missing)}) but a working invocation — informational")
|
||||||
|
elif "/usr/bin/infisical" not in invoked_paths:
|
||||||
|
print(f" ⚠️ {name}: wrapper infisical path differs "
|
||||||
|
f"({', '.join(invoked_paths)}) — informational")
|
||||||
|
else:
|
||||||
|
print(f" ✅ {name}: wrapper infisical path OK")
|
||||||
else:
|
else:
|
||||||
print(f" ⚠️ {name}: wrapper infisical path may be wrong (infisical at {inf_actual})")
|
inf_actual = ssh(host, "command -v infisical 2>/dev/null", user=user)
|
||||||
FAIL.append(f"wrapper-infisical-path:{name}")
|
if not inf_actual:
|
||||||
|
print(f" ❌ {name}: wrapper invokes infisical but the binary is MISSING")
|
||||||
|
_fail(f"wrapper-no-infisical:{name}", name)
|
||||||
|
else:
|
||||||
|
print(f" ✅ {name}: wrapper infisical resolves via PATH ({inf_actual})")
|
||||||
|
else:
|
||||||
|
print(f" ℹ️ {name}: wrapper resolves creds without infisical (e.g. ~/.hermes/.env) — OK")
|
||||||
|
|
||||||
# Check hermes-real exists
|
# Check hermes-real exists
|
||||||
hermes_real = ssh(host,
|
hermes_real = ssh(host,
|
||||||
@@ -399,7 +601,7 @@ def check_wrapper_integrity():
|
|||||||
user=user)
|
user=user)
|
||||||
if not hermes_real or hermes_real.strip() == "MISS":
|
if not hermes_real or hermes_real.strip() == "MISS":
|
||||||
print(f" ❌ {name}: hermes-real NOT FOUND (wrapper broken)")
|
print(f" ❌ {name}: hermes-real NOT FOUND (wrapper broken)")
|
||||||
FAIL.append(f"wrapper-no-hermes-real:{name}")
|
_fail(f"wrapper-no-hermes-real:{name}", name)
|
||||||
else:
|
else:
|
||||||
print(f" ✅ {name}: hermes-real at alt path")
|
print(f" ✅ {name}: hermes-real at alt path")
|
||||||
|
|
||||||
@@ -427,10 +629,10 @@ def check_vault_secrets():
|
|||||||
key = agent.get("key")
|
key = agent.get("key")
|
||||||
if not key:
|
if not key:
|
||||||
print(f" ❌ {name}: vault secret {vault_key_name} MISSING or EMPTY")
|
print(f" ❌ {name}: vault secret {vault_key_name} MISSING or EMPTY")
|
||||||
FAIL.append(f"vault-empty:{name}:{vault_key_name}")
|
_fail(f"vault-empty:{name}:{vault_key_name}", name)
|
||||||
elif not key.startswith("sk-"):
|
elif not key.startswith("sk-"):
|
||||||
print(f" ❌ {name}: vault secret {vault_key_name} WRONG FORMAT (starts '{key[:8]}...')")
|
print(f" ❌ {name}: vault secret {vault_key_name} WRONG FORMAT (starts '{key[:8]}...')")
|
||||||
FAIL.append(f"vault-bad-format:{name}:{vault_key_name}")
|
_fail(f"vault-bad-format:{name}:{vault_key_name}", name)
|
||||||
else:
|
else:
|
||||||
print(f" ✅ {name}: vault {vault_key_name}=sk-...{key[-4:]}")
|
print(f" ✅ {name}: vault {vault_key_name}=sk-...{key[-4:]}")
|
||||||
|
|
||||||
@@ -463,18 +665,7 @@ def deploy_self():
|
|||||||
# MAIN
|
# MAIN
|
||||||
# ═══════════════════════════════════════════════════════════════════
|
# ═══════════════════════════════════════════════════════════════════
|
||||||
|
|
||||||
def main():
|
def _run_checks():
|
||||||
quiet = "--quiet" in sys.argv
|
|
||||||
as_json = "--json" in sys.argv
|
|
||||||
|
|
||||||
# Self-deploy to canonical location
|
|
||||||
if not quiet and "--no-deploy" not in sys.argv:
|
|
||||||
deploy_self()
|
|
||||||
|
|
||||||
if not quiet:
|
|
||||||
print(f"🏥 Agent Health Check v2 — {datetime.now().strftime('%Y-%m-%d %H:%M UTC')}")
|
|
||||||
print()
|
|
||||||
|
|
||||||
print("🔑 LiteLLM Keys:")
|
print("🔑 LiteLLM Keys:")
|
||||||
check_keys()
|
check_keys()
|
||||||
print()
|
print()
|
||||||
@@ -502,16 +693,49 @@ def main():
|
|||||||
print("🔐 Vault Secrets:")
|
print("🔐 Vault Secrets:")
|
||||||
check_vault_secrets()
|
check_vault_secrets()
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
quiet = "--quiet" in sys.argv
|
||||||
|
as_json = "--json" in sys.argv
|
||||||
|
|
||||||
|
# Self-deploy to canonical location
|
||||||
|
if not quiet and "--no-deploy" not in sys.argv:
|
||||||
|
deploy_self()
|
||||||
|
|
||||||
|
# Provenance: a report is only actionable if the reader can tell WHICH copy
|
||||||
|
# of this script produced it. A normal run carries it in the header, --json
|
||||||
|
# carries it for machine consumers, and the cron ALERT line carries it on
|
||||||
|
# failure. --quiet is documented as "only output on failure", so the header
|
||||||
|
# is emitted only when not quiet and a healthy quiet run stays silent.
|
||||||
|
script_path = os.path.abspath(__file__)
|
||||||
|
cwd = os.getcwd()
|
||||||
|
|
||||||
|
if quiet:
|
||||||
|
captured = io.StringIO()
|
||||||
|
with contextlib.redirect_stdout(captured):
|
||||||
|
load_agent_keys()
|
||||||
|
_run_checks()
|
||||||
|
if FAIL:
|
||||||
|
sys.stdout.write(captured.getvalue())
|
||||||
|
else:
|
||||||
|
print(f"🏥 Agent Health Check v4 — {datetime.now().strftime('%Y-%m-%d %H:%M UTC')}")
|
||||||
|
print(f"📍 executed from: script={script_path} cwd={cwd}")
|
||||||
|
print()
|
||||||
|
load_agent_keys()
|
||||||
|
_run_checks()
|
||||||
|
|
||||||
if FAIL:
|
if FAIL:
|
||||||
print(f"\n❌ {len(FAIL)} FAILURE(S): {' | '.join(FAIL)}")
|
print(f"\n❌ {len(FAIL)} FAILURE(S): {' | '.join(FAIL)}")
|
||||||
if quiet:
|
if quiet:
|
||||||
print(f"ALERT agent-health:{','.join(FAIL)}")
|
print(f"ALERT agent-health:{','.join(FAIL)} script={script_path} cwd={cwd}")
|
||||||
elif not quiet:
|
elif not quiet:
|
||||||
print("\n✅ All checks passed")
|
print("\n✅ All checks passed")
|
||||||
|
|
||||||
if as_json:
|
if as_json:
|
||||||
print(json.dumps({"timestamp": datetime.now().isoformat(),
|
print(json.dumps({"timestamp": datetime.now().isoformat(),
|
||||||
"failures": FAIL, "healthy": len(FAIL) == 0}))
|
"execution_path": script_path, "cwd": cwd,
|
||||||
|
"failures": FAIL, "report_only": REPORT_ONLY,
|
||||||
|
"healthy": len(FAIL) == 0}))
|
||||||
|
|
||||||
sys.exit(1 if FAIL else 0)
|
sys.exit(1 if FAIL else 0)
|
||||||
|
|
||||||
|
|||||||
Executable
+227
@@ -0,0 +1,227 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
# capture-dsh-token.sh — refresh the dsh-web login token WITHOUT restarting dsh-web.
|
||||||
|
#
|
||||||
|
# Context (CT 112 / tankodhs.sysloggh.net)
|
||||||
|
# ----------------------------------------
|
||||||
|
# The dsh-web UI (systemd unit `dsh-web.service`, 127.0.0.1:3080) prints a random
|
||||||
|
# launch token to the journal on every start:
|
||||||
|
#
|
||||||
|
# dsh web: http://127.0.0.1:3080/?token=<TOKEN>
|
||||||
|
#
|
||||||
|
# That token is the only way to bootstrap the authority-bound 30-day browser
|
||||||
|
# cookie. It rotates on every dsh-web start, so the Authentik-gated
|
||||||
|
# `location = /dsh-web-login` in /etc/nginx/sites-available/dsh must always
|
||||||
|
# reference the token of the RUNNING process.
|
||||||
|
#
|
||||||
|
# This script:
|
||||||
|
# 1. selects the launch token the RUNNING service actually accepts from the
|
||||||
|
# current systemd invocation — it NEVER stops or starts dsh-web,
|
||||||
|
# 2. records it in /etc/dsh-web/launch-token,
|
||||||
|
# 3. regenerates the nginx include /etc/dsh-web/nginx-login.conf (the
|
||||||
|
# `proxy_pass ...?token=` line consumed by /dsh-web-login),
|
||||||
|
# 4. reloads nginx ONLY when the on-disk include differs from the generated
|
||||||
|
# one or the applied-state stamp does not match the token (the stamp is
|
||||||
|
# written only after a successful reload), rolling the include back on
|
||||||
|
# failure so the next run retries,
|
||||||
|
# 5. removes the legacy unauthenticated :8081 endpoint if it ever reappears.
|
||||||
|
#
|
||||||
|
# Idempotent and safe to run at any time (systemd ExecStartPost or timer).
|
||||||
|
set -euo pipefail
|
||||||
|
umask 077
|
||||||
|
PATH="/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin"
|
||||||
|
|
||||||
|
JOURNAL_UNIT="dsh-web.service"
|
||||||
|
TOKEN_FILE="/etc/dsh-web/launch-token"
|
||||||
|
INCLUDE_FILE="/etc/dsh-web/nginx-login.conf"
|
||||||
|
STAMP_FILE="/etc/dsh-web/nginx-login.conf.applied"
|
||||||
|
PENDING_FILE="/etc/dsh-web/nginx-reload.pending"
|
||||||
|
SITE_ENABLED="/etc/nginx/sites-enabled/dsh"
|
||||||
|
LEGACY_8081="/etc/nginx/sites-enabled/dsh.token"
|
||||||
|
STASH_DIR="/etc/nginx/sites-available"
|
||||||
|
LOCK_FILE="/run/capture-dsh-token.lock"
|
||||||
|
LOGIN_HOST="tankodhs.sysloggh.net"
|
||||||
|
LOGIN_UPSTREAM="http://127.0.0.1:3080"
|
||||||
|
TOKEN_WAIT=120
|
||||||
|
|
||||||
|
log() { printf 'capture-dsh-token: %s\n' "$*" >&2; }
|
||||||
|
die() { printf 'capture-dsh-token: ERROR: %s\n' "$*" >&2; exit 1; }
|
||||||
|
|
||||||
|
[ "$(id -u)" -eq 0 ] || die "must run as root"
|
||||||
|
|
||||||
|
# ── 0. Serialize runs so timer/ExecStartPost/manual runs cannot interleave ──
|
||||||
|
exec 9>"$LOCK_FILE"
|
||||||
|
flock -n 9 || { log "another capture-dsh-token run holds $LOCK_FILE; exiting"; exit 0; }
|
||||||
|
mkdir -p "$(dirname "$PENDING_FILE")"
|
||||||
|
|
||||||
|
# ── 0b. Guarantee the generated include exists before any `nginx -t` ──────
|
||||||
|
# The :80 site includes /etc/dsh-web/nginx-login.conf by literal path, so a
|
||||||
|
# missing include makes every `nginx -t` fail and can wedge recovery. Seed it
|
||||||
|
# from the last known token (or a placeholder); step 4 replaces it.
|
||||||
|
if [ ! -f "$INCLUDE_FILE" ]; then
|
||||||
|
SEED="placeholder"
|
||||||
|
if [ -f "$TOKEN_FILE" ]; then
|
||||||
|
SEED="$(cat "$TOKEN_FILE" 2>/dev/null || true)"
|
||||||
|
[ -n "$SEED" ] || SEED="placeholder"
|
||||||
|
fi
|
||||||
|
printf '%s' "$SEED" | grep -qE '^[A-Za-z0-9._~+/=:@-]+$' || SEED="placeholder"
|
||||||
|
printf 'proxy_pass %s/?token=%s;\n' "$LOGIN_UPSTREAM" "$SEED" > "$INCLUDE_FILE"
|
||||||
|
chmod 600 "$INCLUDE_FILE"
|
||||||
|
log "created missing $INCLUDE_FILE"
|
||||||
|
fi
|
||||||
|
|
||||||
|
# ── 1. Remove the legacy unauthenticated :8081 endpoint, if present ─────────
|
||||||
|
# It bypassed Authentik entirely (listened on 0.0.0.0:8081 with no auth_request)
|
||||||
|
# and must never come back. Stash it rather than delete so it is auditable.
|
||||||
|
if [ -e "$LEGACY_8081" ] || [ -L "$LEGACY_8081" ]; then
|
||||||
|
TS="$(date -u +%Y%m%dT%H%M%SZ)"
|
||||||
|
STASHED="$STASH_DIR/dsh.token.disabled-$TS"
|
||||||
|
mv "$LEGACY_8081" "$STASHED"
|
||||||
|
chmod 600 "$STASHED" 2>/dev/null || true
|
||||||
|
touch "$PENDING_FILE"
|
||||||
|
if ! NGINX_TEST_OUT="$(nginx -t 2>&1)"; then
|
||||||
|
die "nginx config test failed after disabling $LEGACY_8081 (kept disabled at $STASHED): $NGINX_TEST_OUT; a pending reload is recorded so running nginx is reloaded once the config is fixed. The legacy :8081 endpoint will NOT be restored."
|
||||||
|
fi
|
||||||
|
if ! nginx -s reload; then
|
||||||
|
die "nginx reload failed after disabling $LEGACY_8081 (kept disabled at $STASHED); a pending reload is recorded so running nginx is reloaded on the next run. The legacy :8081 endpoint will NOT be restored."
|
||||||
|
fi
|
||||||
|
rm -f "$PENDING_FILE"
|
||||||
|
log "removed legacy :8081 endpoint -> $STASHED"
|
||||||
|
fi
|
||||||
|
|
||||||
|
# ── 1b. Honor a recorded pending reload regardless of token selection ───────
|
||||||
|
# A failed reload leaves PENDING_FILE set so a stashed legacy :8081 file can
|
||||||
|
# never remain loaded in the running nginx while dsh-web is down or not yet
|
||||||
|
# answering. Reconcile it before the token wait.
|
||||||
|
if [ -e "$PENDING_FILE" ]; then
|
||||||
|
if ! NGINX_TEST_OUT="$(nginx -t 2>&1)"; then
|
||||||
|
log "WARNING: pending nginx reload recorded but 'nginx -t' fails: $NGINX_TEST_OUT; continuing so the include can be regenerated; will retry next run"
|
||||||
|
elif ! nginx -s reload; then
|
||||||
|
log "WARNING: pending nginx reload recorded but 'nginx -s reload' failed; will retry next run"
|
||||||
|
else
|
||||||
|
rm -f "$PENDING_FILE"
|
||||||
|
log "completed pending nginx reload"
|
||||||
|
fi
|
||||||
|
fi
|
||||||
|
|
||||||
|
# ── 2. Select the token the RUNNING service actually accepts ────────────────
|
||||||
|
# Re-sample the service's CURRENT systemd invocation on every pass and read
|
||||||
|
# candidates only from it, so a restart that lands during the wait immediately
|
||||||
|
# switches to the new invocation; there is no whole-journal or cross-invocation
|
||||||
|
# fallback, and an empty/unknown invocation just waits. Each candidate is then
|
||||||
|
# functionally verified against the local dsh-web using the public authority,
|
||||||
|
# exactly as the /dsh-web-login proxy does, and the first that answers 303 is
|
||||||
|
# the live token. Candidates are re-probed newest-first on each pass (connection
|
||||||
|
# failures stay eligible) until one is accepted or the wait elapses.
|
||||||
|
journal_tokens() {
|
||||||
|
journalctl -u "$JOURNAL_UNIT" "_SYSTEMD_INVOCATION_ID=$1" --no-pager -o cat 2>/dev/null \
|
||||||
|
| grep -oE 'dsh web: https?://[^[:space:]]+[?&]token=[^[:space:]]+' \
|
||||||
|
| sed -E 's/.*[?&]token=//' \
|
||||||
|
| grep -E '^[A-Za-z0-9._~+/=:@-]+$' \
|
||||||
|
| tac | awk '!seen[$0]++' || true
|
||||||
|
}
|
||||||
|
|
||||||
|
TOKEN=""
|
||||||
|
DEADLINE=$((SECONDS + TOKEN_WAIT))
|
||||||
|
NO_INVOCATION_WARNED=0
|
||||||
|
while [ -z "$TOKEN" ] && [ "$SECONDS" -lt "$DEADLINE" ]; do
|
||||||
|
INVOCATION="$(systemctl show -p InvocationID --value "$JOURNAL_UNIT" 2>/dev/null || true)"
|
||||||
|
if [ -z "$INVOCATION" ] || [ "$INVOCATION" = "n/a" ]; then
|
||||||
|
if [ "$NO_INVOCATION_WARNED" -eq 0 ]; then
|
||||||
|
log "WARNING: no invocation id for $JOURNAL_UNIT; waiting for a live invocation"
|
||||||
|
NO_INVOCATION_WARNED=1
|
||||||
|
fi
|
||||||
|
sleep 2
|
||||||
|
continue
|
||||||
|
fi
|
||||||
|
for cand in $(journal_tokens "$INVOCATION"); do
|
||||||
|
code="$(curl -s -o /dev/null --max-time 5 -w '%{http_code}' \
|
||||||
|
-H "Host: $LOGIN_HOST" "$LOGIN_UPSTREAM/?token=$cand" || true)"
|
||||||
|
if [ "$code" = "303" ]; then
|
||||||
|
TOKEN="$cand"
|
||||||
|
break
|
||||||
|
fi
|
||||||
|
done
|
||||||
|
[ -n "$TOKEN" ] && break
|
||||||
|
sleep 2
|
||||||
|
done
|
||||||
|
|
||||||
|
if [ -z "$TOKEN" ]; then
|
||||||
|
log "no accepted launch token in the current invocation within ${TOKEN_WAIT}s; leaving the include untouched for the next run"
|
||||||
|
[ -e "$PENDING_FILE" ] && die "pending nginx reload could not be completed; will retry next run"
|
||||||
|
exit 0
|
||||||
|
fi
|
||||||
|
|
||||||
|
# ── 3. Record the token (atomic, private) ──────────────────────────────────
|
||||||
|
mkdir -p "$(dirname "$TOKEN_FILE")"
|
||||||
|
if ! printf '%s\n' "$TOKEN" | cmp -s - "$TOKEN_FILE" 2>/dev/null; then
|
||||||
|
printf '%s\n' "$TOKEN" > "$TOKEN_FILE.tmp"
|
||||||
|
chmod 600 "$TOKEN_FILE.tmp"
|
||||||
|
mv "$TOKEN_FILE.tmp" "$TOKEN_FILE"
|
||||||
|
log "recorded live launch token in $TOKEN_FILE"
|
||||||
|
fi
|
||||||
|
chmod 600 "$TOKEN_FILE"
|
||||||
|
|
||||||
|
# ── 4. Regenerate the nginx login include (reload only when it changes) ────
|
||||||
|
NEW_INCLUDE="$(mktemp "$INCLUDE_FILE.XXXXXX")"
|
||||||
|
printf 'proxy_pass %s/?token=%s;\n' "$LOGIN_UPSTREAM" "$TOKEN" > "$NEW_INCLUDE"
|
||||||
|
chmod 600 "$NEW_INCLUDE"
|
||||||
|
|
||||||
|
# The stamp records the token nginx actually loaded. It is written only after a
|
||||||
|
# successful reload, so the early exit is safe only when both the stamp and the
|
||||||
|
# on-disk include agree with the live token; anything else falls through to the
|
||||||
|
# reload path so the include can never silently diverge from what nginx serves.
|
||||||
|
APPLIED=""
|
||||||
|
[ -f "$STAMP_FILE" ] && APPLIED="$(cat "$STAMP_FILE" 2>/dev/null || true)"
|
||||||
|
[ -f "$INCLUDE_FILE" ] && chmod 600 "$INCLUDE_FILE"
|
||||||
|
[ -f "$STAMP_FILE" ] && chmod 600 "$STAMP_FILE"
|
||||||
|
|
||||||
|
if [ "$APPLIED" = "$TOKEN" ] && [ -f "$INCLUDE_FILE" ] && cmp -s "$NEW_INCLUDE" "$INCLUDE_FILE" \
|
||||||
|
&& [ ! -e "$PENDING_FILE" ]; then
|
||||||
|
rm -f "$NEW_INCLUDE"
|
||||||
|
log "token unchanged; nginx not reloaded"
|
||||||
|
exit 0
|
||||||
|
fi
|
||||||
|
|
||||||
|
[ -e "$SITE_ENABLED" ] || { rm -f "$NEW_INCLUDE"; die "$SITE_ENABLED missing; refusing to reload"; }
|
||||||
|
|
||||||
|
RESTORE=""
|
||||||
|
if [ -f "$INCLUDE_FILE" ]; then
|
||||||
|
RESTORE="$(mktemp "$INCLUDE_FILE.bak.XXXXXX")"
|
||||||
|
cp -p "$INCLUDE_FILE" "$RESTORE"
|
||||||
|
chmod 600 "$RESTORE"
|
||||||
|
fi
|
||||||
|
|
||||||
|
mv "$NEW_INCLUDE" "$INCLUDE_FILE"
|
||||||
|
chmod 600 "$INCLUDE_FILE"
|
||||||
|
|
||||||
|
if ! NGINX_TEST_OUT="$(nginx -t 2>&1)"; then
|
||||||
|
if [ -n "$RESTORE" ]; then
|
||||||
|
mv "$RESTORE" "$INCLUDE_FILE"
|
||||||
|
else
|
||||||
|
rm -f "$INCLUDE_FILE"
|
||||||
|
fi
|
||||||
|
die "nginx config test failed: $NGINX_TEST_OUT; previous include restored"
|
||||||
|
fi
|
||||||
|
|
||||||
|
if ! nginx -s reload; then
|
||||||
|
if [ -n "$RESTORE" ]; then
|
||||||
|
mv "$RESTORE" "$INCLUDE_FILE"
|
||||||
|
else
|
||||||
|
rm -f "$INCLUDE_FILE"
|
||||||
|
fi
|
||||||
|
touch "$PENDING_FILE"
|
||||||
|
die "nginx reload failed; previous include restored; a pending reload is recorded so the next run retries"
|
||||||
|
fi
|
||||||
|
|
||||||
|
if [ -n "$RESTORE" ]; then
|
||||||
|
rm -f "$RESTORE"
|
||||||
|
fi
|
||||||
|
|
||||||
|
rm -f "$PENDING_FILE"
|
||||||
|
|
||||||
|
printf '%s\n' "$TOKEN" > "$STAMP_FILE.tmp"
|
||||||
|
chmod 600 "$STAMP_FILE.tmp"
|
||||||
|
mv "$STAMP_FILE.tmp" "$STAMP_FILE"
|
||||||
|
|
||||||
|
log "token changed; nginx reloaded"
|
||||||
|
log "login endpoint: https://$LOGIN_HOST/dsh-web-login (Authentik-gated)"
|
||||||
@@ -15,7 +15,7 @@ import smtplib, json, subprocess, os, sys, datetime, re
|
|||||||
from email.mime.text import MIMEText
|
from email.mime.text import MIMEText
|
||||||
from email.mime.multipart import MIMEMultipart
|
from email.mime.multipart import MIMEMultipart
|
||||||
|
|
||||||
PVE = "https://minipve.sysloggh.net"
|
PVE = "https://192.168.68.12:8006"
|
||||||
AUTH = "Authorization: PVEAPIToken=monitoring@pve!mumuni=eafd56c5-93d4-4d40-a41d-e688be0987f3"
|
AUTH = "Authorization: PVEAPIToken=monitoring@pve!mumuni=eafd56c5-93d4-4d40-a41d-e688be0987f3"
|
||||||
|
|
||||||
# ── Shared credentials —─
|
# ── Shared credentials —─
|
||||||
@@ -38,10 +38,16 @@ TIME_STR = NOW.strftime("%Y-%m-%d %H:%M UTC")
|
|||||||
# ── Helpers ──
|
# ── Helpers ──
|
||||||
|
|
||||||
def pve_get(path):
|
def pve_get(path):
|
||||||
cmd = f'curl -sfk --connect-timeout 10 "{PVE}{path}" -H "{AUTH}"'
|
"""Fetch PVE API data. Returns list on success, None on error (to distinguish from empty list)."""
|
||||||
|
cmd = f'curl -sk --connect-timeout 10 "{PVE}{path}" -H "{AUTH}"'
|
||||||
try:
|
try:
|
||||||
return json.loads(subprocess.check_output(cmd, shell=True))["data"]
|
r = subprocess.run(cmd, shell=True, capture_output=True, text=True, timeout=12)
|
||||||
except: return []
|
if r.returncode != 0:
|
||||||
|
return None
|
||||||
|
data = json.loads(r.stdout)
|
||||||
|
return data.get("data", [])
|
||||||
|
except:
|
||||||
|
return None
|
||||||
|
|
||||||
def ssh(host, cmd):
|
def ssh(host, cmd):
|
||||||
try:
|
try:
|
||||||
@@ -97,21 +103,33 @@ def collect():
|
|||||||
|
|
||||||
# ── Proxmox Nodes ──
|
# ── Proxmox Nodes ──
|
||||||
nodes = pve_get("/api2/json/nodes")
|
nodes = pve_get("/api2/json/nodes")
|
||||||
report["nodes"] = {n["node"]: {
|
if nodes is None:
|
||||||
"cpu_pct": round(n.get('cpu',0)*100, 1),
|
report["nodes"] = {}
|
||||||
"ram": f"{n.get('mem',0)//1024//1024}/{n.get('maxmem',0)//1024//1024}MB",
|
report["node_count"] = 0
|
||||||
"ram_pct": round(n.get('mem',0)/n.get('maxmem',1)*100, 0),
|
report["nodes_online"] = 0
|
||||||
"disk": f"{n.get('disk',0)//1024//1024//1024}/{n.get('maxdisk',0)//1024//1024//1024}GB",
|
report["pve_probe_status"] = "unreachable"
|
||||||
"disk_pct": round(n.get('disk',0)/n.get('maxdisk',1)*100, 0),
|
else:
|
||||||
"uptime_h": n.get('uptime',0)//3600,
|
report["nodes"] = {n["node"]: {
|
||||||
"status": n["status"]
|
"cpu_pct": round(n.get('cpu',0)*100, 1),
|
||||||
} for n in nodes}
|
"ram": f"{n.get('mem',0)//1024//1024}/{n.get('maxmem',0)//1024//1024}MB",
|
||||||
report["node_count"] = len(nodes)
|
"ram_pct": round(n.get('mem',0)/n.get('maxmem',1)*100, 0),
|
||||||
report["nodes_online"] = sum(1 for n in nodes if n["status"] == "online")
|
"disk": f"{n.get('disk',0)//1024//1024//1024}/{n.get('maxdisk',0)//1024//1024//1024}GB",
|
||||||
|
"disk_pct": round(n.get('disk',0)/n.get('maxdisk',1)*100, 0),
|
||||||
|
"uptime_h": n.get('uptime',0)//3600,
|
||||||
|
"status": n["status"]
|
||||||
|
} for n in nodes}
|
||||||
|
report["node_count"] = len(nodes)
|
||||||
|
report["nodes_online"] = sum(1 for n in nodes if n["status"] == "online")
|
||||||
|
report["pve_probe_status"] = "ok"
|
||||||
|
|
||||||
# ── VMs/CTs ──
|
# ── VMs/CTs ──
|
||||||
resources = pve_get("/api2/json/cluster/resources")
|
resources = pve_get("/api2/json/cluster/resources")
|
||||||
vms = [r for r in resources if r.get("type") in ("qemu","lxc")]
|
if resources is None:
|
||||||
|
vms = []
|
||||||
|
report["resources_probe_status"] = "unreachable"
|
||||||
|
else:
|
||||||
|
vms = [r for r in resources if r.get("type") in ("qemu","lxc")]
|
||||||
|
report["resources_probe_status"] = "ok"
|
||||||
report["total_vms"] = len(vms)
|
report["total_vms"] = len(vms)
|
||||||
report["running_vms"] = sum(1 for v in vms if v.get("status") == "running")
|
report["running_vms"] = sum(1 for v in vms if v.get("status") == "running")
|
||||||
stopped = [v for v in vms if v.get("status") != "running"]
|
stopped = [v for v in vms if v.get("status") != "running"]
|
||||||
@@ -183,7 +201,7 @@ def collect():
|
|||||||
("Authentik", "https://auth.sysloggh.net"),
|
("Authentik", "https://auth.sysloggh.net"),
|
||||||
("Zulip", "https://chat.sysloggh.net"),
|
("Zulip", "https://chat.sysloggh.net"),
|
||||||
("Pulse", "https://pulse.sysloggh.net"),
|
("Pulse", "https://pulse.sysloggh.net"),
|
||||||
("Proxmox", "https://minipve.sysloggh.net"),
|
("Proxmox", "https://192.168.68.12:8006"),
|
||||||
("SearXNG", "http://192.168.68.7:8888"),
|
("SearXNG", "http://192.168.68.7:8888"),
|
||||||
("Firecrawl", "http://192.168.68.7:3002/health"),
|
("Firecrawl", "http://192.168.68.7:3002/health"),
|
||||||
]
|
]
|
||||||
@@ -247,11 +265,13 @@ def collect():
|
|||||||
zulip_health = json.loads(health_body) if health_body else {}
|
zulip_health = json.loads(health_body) if health_body else {}
|
||||||
except:
|
except:
|
||||||
zulip_health = {}
|
zulip_health = {}
|
||||||
report["zulip_ext"]["connected"] = zulip_health.get("connected", False)
|
# Live state is nested under 'zulip' key
|
||||||
report["zulip_ext"]["queue_id"] = zulip_health.get("queue_id")
|
zulip_state = zulip_health.get("zulip", {})
|
||||||
report["zulip_ext"]["last_error"] = zulip_health.get("last_error")
|
report["zulip_ext"]["connected"] = zulip_state.get("connected", False)
|
||||||
report["zulip_ext"]["messages_processed"] = zulip_health.get("messages_processed", 0)
|
report["zulip_ext"]["queue_id"] = zulip_state.get("queue_id")
|
||||||
report["zulip_ext"]["retry_count"] = zulip_health.get("retry_count", 0)
|
report["zulip_ext"]["last_error"] = zulip_state.get("last_error")
|
||||||
|
report["zulip_ext"]["messages_processed"] = zulip_state.get("messages_processed", 0)
|
||||||
|
report["zulip_ext"]["skipped"] = zulip_state.get("skipped", 0)
|
||||||
|
|
||||||
# Phase 2: PM2 process check
|
# Phase 2: PM2 process check
|
||||||
pm2_raw = subprocess.check_output(
|
pm2_raw = subprocess.check_output(
|
||||||
@@ -289,50 +309,32 @@ def collect():
|
|||||||
# Abiba (pi)
|
# Abiba (pi)
|
||||||
report["agents"]["abiba"] = {
|
report["agents"]["abiba"] = {
|
||||||
"platform": "pi", "ct": 100, "ip": "192.168.68.24",
|
"platform": "pi", "ct": 100, "ip": "192.168.68.24",
|
||||||
"zulip_connected": zulip_health.get("connected", False),
|
"zulip_connected": zulip_state.get("connected", False),
|
||||||
"zulip_processed": zulip_health.get("messages_processed", 0),
|
"zulip_processed": zulip_state.get("messages_processed", 0),
|
||||||
"pm2_status": pm2.get("status", "unknown"),
|
"pm2_status": pm2.get("status", "unknown"),
|
||||||
"pm2_restarts": pm2.get("restarts", "?"),
|
"pm2_restarts": pm2.get("restarts", "?"),
|
||||||
"pm2_uptime": pm2.get("uptime", "?"),
|
"pm2_uptime": pm2.get("uptime", "?"),
|
||||||
}
|
}
|
||||||
|
|
||||||
# Tanko (CT 122)
|
# Tanko (CT 112, IP 192.168.68.122) — DSH (DeepSeek Harness), no Hermes gateway
|
||||||
tanko_state = ssh_jerome("192.168.68.122", "cat ~/.hermes/gateway_state.json 2>/dev/null")
|
# since 2026-08-27. There is no ~/.hermes/gateway_state.json on CT 112 anymore;
|
||||||
tanko_data = {}
|
# Zulip/Telegram connectivity is managed by the DSH harness, not the Hermes gateway.
|
||||||
try:
|
|
||||||
tanko_data = json.loads(tanko_state) if tanko_state else {}
|
|
||||||
except:
|
|
||||||
tanko_data = {}
|
|
||||||
platforms = tanko_data.get("platforms", {})
|
|
||||||
report["agents"]["tanko"] = {
|
report["agents"]["tanko"] = {
|
||||||
"platform": "hermes", "ct": 112, "ip": "192.168.68.122",
|
"platform": "dsh", "ct": 112, "ip": "192.168.68.122",
|
||||||
"gateway_state": tanko_data.get("gateway_state", "unknown"),
|
"gateway_state": "n/a (DSH)",
|
||||||
"zulip_state": platforms.get("zulip", {}).get("state", "unknown"),
|
"zulip_state": "unknown",
|
||||||
"telegram_state": platforms.get("telegram", {}).get("state", "unknown"),
|
"telegram_state": "unknown",
|
||||||
"gateway_pid": tanko_data.get("pid"),
|
"gateway_pid": None,
|
||||||
"updated_at": tanko_data.get("updated_at"),
|
"updated_at": "",
|
||||||
}
|
}
|
||||||
|
|
||||||
# Mumuni (CT 100, IP 192.168.68.24)
|
# Mumuni is deliberately absent from this digest: captain ruling 2026-09-10.
|
||||||
mumuni_state = ssh("192.168.68.24", "cat ~/.hermes/gateway_state.json 2>/dev/null")
|
# She moved off this host onto her own container (kagentz CT 105 on minipve,
|
||||||
mumuni_data = {}
|
# 192.168.68.14, dedicated `hermes` user) and is monitored from her side. The
|
||||||
try:
|
# former probe ssh'd to 192.168.68.24 for the decommissioned deployment's
|
||||||
mumuni_data = json.loads(mumuni_state) if mumuni_state else {}
|
# ~/.hermes/gateway_state.json, always read "unknown", and published a false
|
||||||
except:
|
# "mumuni:unknown" line in the agent table and the gateway-unknown issue
|
||||||
mumuni_data = {}
|
# count of every digest. Do NOT re-add an .24 / gateway_state probe.
|
||||||
mumuni_platforms = mumuni_data.get("platforms", {})
|
|
||||||
report["agents"]["mumuni"] = {
|
|
||||||
"platform": "hermes", "ct": 100, "ip": "192.168.68.24",
|
|
||||||
"gateway_state": mumuni_data.get("gateway_state", "unknown"),
|
|
||||||
"telegram_state": mumuni_platforms.get("telegram", {}).get("state", "unknown"),
|
|
||||||
"zulip_state": mumuni_platforms.get("zulip", {}).get("state", "not_installed"),
|
|
||||||
"email_state": mumuni_platforms.get("email", {}).get("state", "unknown"),
|
|
||||||
"hermes_version": "",
|
|
||||||
}
|
|
||||||
# Get Hermes version
|
|
||||||
ver = ssh("192.168.68.24", "hermes --version 2>/dev/null | head -1")
|
|
||||||
if ver:
|
|
||||||
report["agents"]["mumuni"]["hermes_version"] = ver.split("·")[0].replace("Hermes Agent ","").strip()
|
|
||||||
|
|
||||||
return report
|
return report
|
||||||
|
|
||||||
@@ -424,7 +426,7 @@ th {{ color: #8b949e; font-weight: normal; }}
|
|||||||
<div class="alert {'good' if not issues else 'bad' if any('🔴' in i for i in issues) else 'warn'}">
|
<div class="alert {'good' if not issues else 'bad' if any('🔴' in i for i in issues) else 'warn'}">
|
||||||
<p style="margin:0;font-size:16px"><b>{status}</b></p>
|
<p style="margin:0;font-size:16px"><b>{status}</b></p>
|
||||||
<p style="margin:4px 0 0 0;font-size:13px">
|
<p style="margin:4px 0 0 0;font-size:13px">
|
||||||
{r['node_count']} PVE nodes · {r['total_vms']} VMs/CTs · {r['running_vms']} running ·
|
Proxmox: {r.get('pve_probe_status', 'ok')} ({r['nodes_online']}/{r['node_count']}) · {r['total_vms']} VMs/CTs · {r['running_vms']} running ·
|
||||||
{r['docker_vm']['total'] + r['docker_syslog']['total'] + r['docker_netbird']['total']} containers ·
|
{r['docker_vm']['total'] + r['docker_syslog']['total'] + r['docker_netbird']['total']} containers ·
|
||||||
{len(r['endpoints'])} endpoints · {len(r.get('agents',{}))} agents
|
{len(r['endpoints'])} endpoints · {len(r.get('agents',{}))} agents
|
||||||
</p>
|
</p>
|
||||||
@@ -440,8 +442,10 @@ th {{ color: #8b949e; font-weight: normal; }}
|
|||||||
|
|
||||||
# ── Quick Stats ──
|
# ── Quick Stats ──
|
||||||
html += '<div class="card"><h2>📊 Quick Stats</h2><div class="grid">'
|
html += '<div class="card"><h2>📊 Quick Stats</h2><div class="grid">'
|
||||||
|
pve_status_label = "unreachable" if r.get('pve_probe_status') == 'unreachable' else f"{r['nodes_online']}/{r['node_count']}"
|
||||||
|
pve_status_color = "red" if r.get('pve_probe_status') == 'unreachable' or r['nodes_online'] != r['node_count'] else "green"
|
||||||
stats = [
|
stats = [
|
||||||
("PVE Nodes", f"{r['nodes_online']}/{r['node_count']}", "green" if r['nodes_online'] == r['node_count'] else "red"),
|
("PVE Nodes", pve_status_label, pve_status_color),
|
||||||
("VMs/CTs", f"{r['running_vms']}/{r['total_vms']}", "green" if r['running_vms'] == r['total_vms'] else "red"),
|
("VMs/CTs", f"{r['running_vms']}/{r['total_vms']}", "green" if r['running_vms'] == r['total_vms'] else "red"),
|
||||||
("Containers", f"{r['docker_vm']['running']}/{r['docker_vm']['total']}", "green" if r['docker_vm']['running'] == r['docker_vm']['total'] else "yellow"),
|
("Containers", f"{r['docker_vm']['running']}/{r['docker_vm']['total']}", "green" if r['docker_vm']['running'] == r['docker_vm']['total'] else "yellow"),
|
||||||
("LiteLLM Ctrs", f"{r['docker_syslog']['running']}/{r['docker_syslog']['total']}", "green" if r['docker_syslog']['running'] == r['docker_syslog']['total'] else "red"),
|
("LiteLLM Ctrs", f"{r['docker_syslog']['running']}/{r['docker_syslog']['total']}", "green" if r['docker_syslog']['running'] == r['docker_syslog']['total'] else "red"),
|
||||||
@@ -543,13 +547,7 @@ th {{ color: #8b949e; font-weight: normal; }}
|
|||||||
elif name == "tanko":
|
elif name == "tanko":
|
||||||
zulip_state = "✅" if agent.get("zulip_state") == "connected" else ("❌" if agent.get("zulip_state") == "disconnected" else "⬜")
|
zulip_state = "✅" if agent.get("zulip_state") == "connected" else ("❌" if agent.get("zulip_state") == "disconnected" else "⬜")
|
||||||
gateway = agent.get("gateway_state", "?")
|
gateway = agent.get("gateway_state", "?")
|
||||||
processed = agent.get("updated_at", "")[:10]
|
processed = "DSH"
|
||||||
elif name == "mumuni":
|
|
||||||
zulip_state = "⬜" if agent.get("zulip_state") == "not_installed" else ("✅" if agent.get("zulip_state") == "connected" else "⬜")
|
|
||||||
gateway = agent.get("gateway_state", "?")
|
|
||||||
tg = "✅" if agent.get("telegram_state") == "connected" else "❌"
|
|
||||||
ver = agent.get("hermes_version", "")
|
|
||||||
processed = f"TG:{tg} v{ver}"
|
|
||||||
else:
|
else:
|
||||||
zulip_state = "⬜"
|
zulip_state = "⬜"
|
||||||
gateway = agent.get("gateway_state", "?")
|
gateway = agent.get("gateway_state", "?")
|
||||||
@@ -703,6 +701,6 @@ if __name__ == "__main__":
|
|||||||
print(f" Zulip Ext: {'✅' if report.get('zulip_ext',{}).get('connected') else '❌'}")
|
print(f" Zulip Ext: {'✅' if report.get('zulip_ext',{}).get('connected') else '❌'}")
|
||||||
print(f" LiteLLM: {sum(1 for c in report.get('litellm',{}).get('checks',[]) if c['status']=='pass')}/{len(report.get('litellm',{}).get('checks',[]))} checks pass")
|
print(f" LiteLLM: {sum(1 for c in report.get('litellm',{}).get('checks',[]) if c['status']=='pass')}/{len(report.get('litellm',{}).get('checks',[]))} checks pass")
|
||||||
agent_parts = []
|
agent_parts = []
|
||||||
for k,v in report.get('agents',{}).items():
|
for k,v in report.get('agents',{}).items():
|
||||||
agent_parts.append(f"{k}:{v.get('gateway_state',v.get('pm2_status','?'))}")
|
agent_parts.append(f"{k}:{v.get('gateway_state',v.get('pm2_status','?'))}")
|
||||||
print(f" Agents: {', '.join(agent_parts)}")
|
print(f" Agents: {', '.join(agent_parts)}")
|
||||||
|
|||||||
@@ -5,6 +5,7 @@
|
|||||||
# Field positions (awk -F'│'): $7=pid $8=uptime $9=restarts $10=status
|
# Field positions (awk -F'│'): $7=pid $8=uptime $9=restarts $10=status
|
||||||
|
|
||||||
TELEGRAM_BOT_TOKEN="$(grep TELEGRAM_BOT_TOKEN /root/.pi/agent/extensions/telegram/.env 2>/dev/null | cut -d= -f2 || echo '')"
|
TELEGRAM_BOT_TOKEN="$(grep TELEGRAM_BOT_TOKEN /root/.pi/agent/extensions/telegram/.env 2>/dev/null | cut -d= -f2 || echo '')"
|
||||||
|
LOG="/root/pm2-self-heal.log"
|
||||||
TELEGRAM_CHAT_ID="5822977936"
|
TELEGRAM_CHAT_ID="5822977936"
|
||||||
|
|
||||||
notify_tg() {
|
notify_tg() {
|
||||||
@@ -16,8 +17,6 @@ notify_tg() {
|
|||||||
-d "text=${msg}" \
|
-d "text=${msg}" \
|
||||||
-d "parse_mode=HTML" > /dev/null 2>&1 || true
|
-d "parse_mode=HTML" > /dev/null 2>&1 || true
|
||||||
}
|
}
|
||||||
ALERTS="${ALERTS}$msg"
|
|
||||||
}
|
|
||||||
|
|
||||||
# Log-only mode: replaced by prose contract pm2-self-heal.prose.md
|
# Log-only mode: replaced by prose contract pm2-self-heal.prose.md
|
||||||
# Only alerts Telegram on actual failure (status != online)
|
# Only alerts Telegram on actual failure (status != online)
|
||||||
|
|||||||
@@ -12,10 +12,10 @@ echo ""
|
|||||||
# Authorized agents for restricted contracts
|
# Authorized agents for restricted contracts
|
||||||
# Format: contract_pattern|authorized_agents (comma-separated)
|
# Format: contract_pattern|authorized_agents (comma-separated)
|
||||||
declare -A RESTRICTED
|
declare -A RESTRICTED
|
||||||
RESTRICTED["infrastructure-control.prose.md"]="abiba"
|
RESTRICTED["infrastructure-control.prose.md"]="abiba,abiba-bot,tanko,tanko-bot,mumuni,mumuni-bot"
|
||||||
RESTRICTED["proxmox-monitor.prose.md"]="abiba"
|
RESTRICTED["proxmox-monitor.prose.md"]="abiba,abiba-bot"
|
||||||
RESTRICTED["hermes-config-template.prose.md"]="abiba,mumuni,tanko"
|
RESTRICTED["hermes-config-template.prose.md"]="abiba,abiba-bot,mumuni,mumuni-bot,tanko,tanko-bot"
|
||||||
RESTRICTED["zulip-health.prose.md"]="abiba,mumuni"
|
RESTRICTED["zulip-health.prose.md"]="abiba,abiba-bot,mumuni,mumuni-bot,tanko,tanko-bot"
|
||||||
RESTRICTED["scripts/pm2-self-heal.sh"]="abiba"
|
RESTRICTED["scripts/pm2-self-heal.sh"]="abiba"
|
||||||
RESTRICTED["scripts/prose-lint.sh"]="abiba"
|
RESTRICTED["scripts/prose-lint.sh"]="abiba"
|
||||||
RESTRICTED["scripts/prose-ai-review.sh"]="abiba"
|
RESTRICTED["scripts/prose-ai-review.sh"]="abiba"
|
||||||
|
|||||||
+31
-2
@@ -57,7 +57,7 @@ echo "── 2. Regression detection ──"
|
|||||||
|
|
||||||
# Grafana /grafana/ as nginx route or URL path (reverted 2026-07-02)
|
# Grafana /grafana/ as nginx route or URL path (reverted 2026-07-02)
|
||||||
# EXCLUDE: filesystem paths (/opt/monitoring/grafana/...), directory creation, revert docs
|
# EXCLUDE: filesystem paths (/opt/monitoring/grafana/...), directory creation, revert docs
|
||||||
GRAFANA_HITS=$(grep -rn '/grafana/' *.prose.md 2>/dev/null \
|
GRAFANA_HITS=$(grep -rn '/grafana/' ./*.prose.md 2>/dev/null \
|
||||||
| grep -v '/opt/monitoring/grafana/' \
|
| grep -v '/opt/monitoring/grafana/' \
|
||||||
| grep -v 'was tried and reverted\|was reverted\|do not re-add\|NOT recommended' \
|
| grep -v 'was tried and reverted\|was reverted\|do not re-add\|NOT recommended' \
|
||||||
| grep -v 'mkdir.*grafana\|Create.*grafana' \
|
| grep -v 'mkdir.*grafana\|Create.*grafana' \
|
||||||
@@ -71,7 +71,7 @@ else
|
|||||||
fi
|
fi
|
||||||
|
|
||||||
# Stale CT IDs (CT 122, CT 123 as CT IDs — not IPs .122, .123)
|
# Stale CT IDs (CT 122, CT 123 as CT IDs — not IPs .122, .123)
|
||||||
CT_STALE=$(grep -rn '\bCT 122\b' *.prose.md 2>/dev/null || true)
|
CT_STALE=$(grep -rn '\bCT 122\b' ./*.prose.md 2>/dev/null || true)
|
||||||
if [ -n "$CT_STALE" ]; then
|
if [ -n "$CT_STALE" ]; then
|
||||||
echo " ❌ REGRESSION: CT 122 used as CT ID — Tanko is CT 112"
|
echo " ❌ REGRESSION: CT 122 used as CT ID — Tanko is CT 112"
|
||||||
echo "$CT_STALE"
|
echo "$CT_STALE"
|
||||||
@@ -84,6 +84,35 @@ fi
|
|||||||
# .122/.123 are correct — verified reachable bridge IPs for Tanko/Mumuni
|
# .122/.123 are correct — verified reachable bridge IPs for Tanko/Mumuni
|
||||||
echo " ✅ IP consistency verified (.19=.122=.123 all reachable)"
|
echo " ✅ IP consistency verified (.19=.122=.123 all reachable)"
|
||||||
|
|
||||||
|
# Report provenance — every contract report must state the absolute path it
|
||||||
|
# executed from, so a stale-consumer report is distinguishable from a real fault
|
||||||
|
# at read time (2026-09-09 probe-drift incident: three false DEGRADED rounds).
|
||||||
|
# Enforced only inside the **Report format** paragraph, and a check-health
|
||||||
|
# contract with no Report format paragraph FAILs rather than being skipped.
|
||||||
|
PROV_FILES=$(grep -rlE '^### check-health|\*\*Report format\*\*' ./*.prose.md 2>/dev/null || true)
|
||||||
|
if [ -z "$PROV_FILES" ]; then
|
||||||
|
echo " ❌ No check-health/report-format contracts found — provenance not enforced"
|
||||||
|
FAILED=1
|
||||||
|
else
|
||||||
|
PROV_BAD=0
|
||||||
|
while IFS= read -r f; do
|
||||||
|
[ -n "$f" ] || continue
|
||||||
|
REPORT_PARA=$(awk '/\*\*Report format\*\*/{found=1} found{print} found && /^[[:space:]]*$/{exit}' "$f")
|
||||||
|
if [ -z "$REPORT_PARA" ]; then
|
||||||
|
echo " ❌ $f: check-health contract has no **Report format** paragraph"
|
||||||
|
PROV_BAD=1
|
||||||
|
elif ! printf '%s\n' "$REPORT_PARA" | grep -qE 'absolute path|pwd -P|executed from'; then
|
||||||
|
echo " ❌ $f: **Report format** lacks execution provenance (absolute path / pwd -P)"
|
||||||
|
PROV_BAD=1
|
||||||
|
fi
|
||||||
|
done <<< "$PROV_FILES"
|
||||||
|
if [ "$PROV_BAD" -eq 1 ]; then
|
||||||
|
FAILED=1
|
||||||
|
else
|
||||||
|
echo " ✅ Report provenance present in all report-format contracts"
|
||||||
|
fi
|
||||||
|
fi
|
||||||
|
|
||||||
# ── 3. Cross-contract consistency ──
|
# ── 3. Cross-contract consistency ──
|
||||||
echo ""
|
echo ""
|
||||||
echo "── 3. Cross-contract consistency ──"
|
echo "── 3. Cross-contract consistency ──"
|
||||||
|
|||||||
+117
-80
@@ -1,7 +1,10 @@
|
|||||||
#!/bin/bash
|
#!/bin/bash
|
||||||
# /root/scripts/zulip-monitor.sh — Zulip Mesh Health Monitor
|
# /root/scripts/zulip-monitor.sh — Zulip Mesh Health Monitor
|
||||||
# Implements zulip-health.prose.md v2
|
# Implements zulip-health.prose.md v3
|
||||||
# Runs every 15 min via cron. Alerts via Telegram.
|
# Runs every 15 min via cron. Alerts: Zulip private DM to the owner plus a stream post to #agent-hub on topic 'zulip-health'.
|
||||||
|
# Legs: global Zulip server, Platform A pi/Abiba (the Zulip bridge), Platform B
|
||||||
|
# Tanko (DSH), Platform C Agent Zero (kagentz). The former Platform B Hermes
|
||||||
|
# agent leg is retired — see the note after the Tanko leg.
|
||||||
set -euo pipefail
|
set -euo pipefail
|
||||||
|
|
||||||
ZULIP_SITE="https://chat.sysloggh.net"
|
ZULIP_SITE="https://chat.sysloggh.net"
|
||||||
@@ -9,10 +12,11 @@ ZULIP_EMAIL="abiba-bot@chat.sysloggh.net"
|
|||||||
ZULIP_KEY="cKTDMZAPW08dk3zl05sStzO7HRztzyn8"
|
ZULIP_KEY="cKTDMZAPW08dk3zl05sStzO7HRztzyn8"
|
||||||
OWNER_ZULIP_ID="9"
|
OWNER_ZULIP_ID="9"
|
||||||
|
|
||||||
# Email config
|
|
||||||
GMAIL_USER="jtabiri@gmail.com"
|
LOG="/root/zulip-health-monitor.log"
|
||||||
GMAIL_PASS="rgbuomwcydxwbszd"
|
TIMESTAMP=$(date -u '+%Y-%m-%d %H:%M UTC')
|
||||||
EMAIL_TO="jerome@sysloggh.com"
|
ISSUES=0
|
||||||
|
echo "=== Zulip Health Check — $TIMESTAMP ===" >> "$LOG"
|
||||||
|
|
||||||
notify() {
|
notify() {
|
||||||
local severity="$1" msg="$2"
|
local severity="$1" msg="$2"
|
||||||
@@ -20,34 +24,20 @@ notify() {
|
|||||||
|
|
||||||
# Zulip DM to owner
|
# Zulip DM to owner
|
||||||
local content="${severity} Zulip Monitor: ${msg}"
|
local content="${severity} Zulip Monitor: ${msg}"
|
||||||
local form="type=private&to=%5B${OWNER_ZULIP_ID}%5D&content=$(python3 -c "import urllib.parse; print(urllib.parse.quote('''${content}'''))")"
|
local form
|
||||||
|
form="type=private&to=%5B${OWNER_ZULIP_ID}%5D&content=$(python3 -c "import urllib.parse; print(urllib.parse.quote('''${content}'''))")"
|
||||||
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
|
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
|
||||||
-u "${ZULIP_EMAIL}:${ZULIP_KEY}" \
|
-u "${ZULIP_EMAIL}:${ZULIP_KEY}" \
|
||||||
-d "${form}" > /dev/null 2>&1 || true
|
-d "${form}" > /dev/null 2>&1 || true
|
||||||
|
# Zulip stream post to #agent-hub on topic 'zulip-health'
|
||||||
# Email alert
|
local stream_content="${severity} Zulip Monitor: ${msg}"
|
||||||
local subject="${severity} Zulip Monitor Alert"
|
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
|
||||||
python3 -c "
|
-u "${ZULIP_EMAIL}:${ZULIP_KEY}" \
|
||||||
import smtplib
|
-d "type=stream&to=%5B7%5D&topic=zulip-health&content=$(printf '%s' "${stream_content}" | python3 -c "import sys,urllib.parse; print(urllib.parse.quote_from_bytes(sys.stdin.buffer.read()))")" \
|
||||||
from email.mime.text import MIMEText
|
> /dev/null 2>&1 \
|
||||||
m = MIMEText('''${msg}''')
|
|| echo " WARN: stream alert to #agent-hub (zulip-health) delivery failed (curl exit $?)" >> "$LOG"
|
||||||
m['From'] = 'abiba@sysloggh.com'
|
|
||||||
m['To'] = '${EMAIL_TO}'
|
|
||||||
m['Subject'] = '${subject}'
|
|
||||||
s = smtplib.SMTP('smtp.gmail.com', 587)
|
|
||||||
s.starttls()
|
|
||||||
s.login('${GMAIL_USER}', '${GMAIL_PASS}')
|
|
||||||
s.sendmail('abiba@sysloggh.com', ['${EMAIL_TO}'], m.as_string())
|
|
||||||
s.quit()
|
|
||||||
" 2>/dev/null || true
|
|
||||||
}
|
}
|
||||||
|
|
||||||
TIMESTAMP=$(date -u '+%Y-%m-%d %H:%M UTC')
|
|
||||||
ISSUES=0
|
|
||||||
LOG="/root/zulip-health-monitor.log"
|
|
||||||
|
|
||||||
echo "=== Zulip Health Check — $TIMESTAMP ===" >> "$LOG"
|
|
||||||
|
|
||||||
# ── Global: Zulip Server ──
|
# ── Global: Zulip Server ──
|
||||||
SERVER_CODE=$(curl -s -o /dev/null -w "%{http_code}" --connect-timeout 10 \
|
SERVER_CODE=$(curl -s -o /dev/null -w "%{http_code}" --connect-timeout 10 \
|
||||||
https://chat.sysloggh.net/api/v1/server_settings \
|
https://chat.sysloggh.net/api/v1/server_settings \
|
||||||
@@ -60,62 +50,109 @@ else
|
|||||||
fi
|
fi
|
||||||
|
|
||||||
# ── Platform A: pi (Abiba) ──
|
# ── Platform A: pi (Abiba) ──
|
||||||
PI_HEALTH=$(curl -sf --connect-timeout 5 http://localhost:9200/health 2>/dev/null || echo "{}")
|
# Probes the pi Zulip extension health endpoint (:9200/health, served by the
|
||||||
PI_CONNECTED=$(echo "$PI_HEALTH" | python3 -c "import sys,json; d=json.load(sys.stdin); print(d.get('connected',False))" 2>/dev/null)
|
# extension's startHealthServer; shape documented in zulip-health.prose.md).
|
||||||
PI_ERROR=$(echo "$PI_HEALTH" | python3 -c "import sys,json; d=json.load(sys.stdin); print(d.get('last_error') or '')" 2>/dev/null)
|
# FAIL-SAFE contract (pinned by tests/zulip-monitor-abiba.sh): connection state
|
||||||
PI_RETRIES=$(echo "$PI_HEALTH" | python3 -c "import sys,json; d=json.load(sys.stdin); print(d.get('retry_count',0))" 2>/dev/null)
|
# lives NESTED at zulip.connected / zulip.last_error — there is no top-level
|
||||||
|
# `connected` and no retry counter in the payload. A fetch error, non-2xx
|
||||||
|
# response, empty/unparseable body, or payload missing a boolean
|
||||||
|
# zulip.connected is a PROBE FAILURE: it alerts and NEVER calls pm2 restart.
|
||||||
|
# pm2 restart runs ONLY on affirmative zulip.connected=false.
|
||||||
|
# -- abiba-leg-start (verbatim-extracted by tests/zulip-monitor-abiba.sh)
|
||||||
|
PI_HTTP=$(curl -s -o /dev/null --connect-timeout 5 --max-time 10 -w '%{http_code}' http://localhost:9200/health 2>/dev/null || echo "000")
|
||||||
|
PI_BODY=$(curl -s --connect-timeout 5 --max-time 10 http://localhost:9200/health 2>/dev/null || true)
|
||||||
|
PI_STATE=$(printf '%s' "$PI_BODY" | python3 -c '
|
||||||
|
import sys, json
|
||||||
|
code = sys.argv[1]
|
||||||
|
body = sys.stdin.read()
|
||||||
|
try:
|
||||||
|
d = json.loads(body)
|
||||||
|
except Exception:
|
||||||
|
sys.stdout.write("probe-failed|unparseable body")
|
||||||
|
sys.exit(0)
|
||||||
|
if not code.startswith("2"):
|
||||||
|
sys.stdout.write("probe-failed|HTTP %s" % code)
|
||||||
|
sys.exit(0)
|
||||||
|
if not isinstance(d, dict) or not isinstance(d.get("zulip"), dict):
|
||||||
|
sys.stdout.write("probe-failed|missing zulip.connected")
|
||||||
|
sys.exit(0)
|
||||||
|
z = d["zulip"]
|
||||||
|
if "connected" not in z or not isinstance(z["connected"], bool):
|
||||||
|
sys.stdout.write("probe-failed|missing or non-boolean zulip.connected")
|
||||||
|
sys.exit(0)
|
||||||
|
err = z.get("last_error") or ""
|
||||||
|
if z["connected"]:
|
||||||
|
if err:
|
||||||
|
sys.stdout.write("degraded|%s" % err)
|
||||||
|
else:
|
||||||
|
sys.stdout.write("healthy|%s" % z.get("messages_processed", 0))
|
||||||
|
else:
|
||||||
|
sys.stdout.write("disconnected|")
|
||||||
|
' "$PI_HTTP" 2>/dev/null) || PI_STATE="probe-failed|python error"
|
||||||
|
PI_VERDICT=${PI_STATE%%|*}
|
||||||
|
PI_DETAIL=${PI_STATE#*|}
|
||||||
|
|
||||||
if [ "$PI_CONNECTED" != "True" ]; then
|
case "$PI_VERDICT" in
|
||||||
notify "🔴" "Abiba pi extension DISCONNECTED — restarting"
|
healthy)
|
||||||
pm2 restart abiba-zulip 2>/dev/null || true
|
echo " Abiba: ✅ Connected (processed=$PI_DETAIL)" >> "$LOG" ;;
|
||||||
|
degraded)
|
||||||
|
notify "🟡" "Abiba pi extension error: ${PI_DETAIL:0:100}"
|
||||||
|
echo " Abiba: 🟡 Error: ${PI_DETAIL:0:100}" >> "$LOG" ;;
|
||||||
|
disconnected)
|
||||||
|
notify "🔴" "Abiba pi extension DISCONNECTED — restarting"
|
||||||
|
pm2 restart abiba-zulip 2>/dev/null || true
|
||||||
|
ISSUES=$((ISSUES + 1))
|
||||||
|
echo " Abiba: ❌ Disconnected — restarted" >> "$LOG" ;;
|
||||||
|
probe-failed)
|
||||||
|
notify "🟠" "Abiba pi extension health probe FAILED (${PI_DETAIL}; HTTP $PI_HTTP) — NOT restarting, manual check needed"
|
||||||
|
ISSUES=$((ISSUES + 1))
|
||||||
|
echo " Abiba: ⚠️ Probe failed (${PI_DETAIL}; HTTP $PI_HTTP) — NOT restarted" >> "$LOG" ;;
|
||||||
|
*)
|
||||||
|
notify "🟠" "Abiba pi extension health probe returned unexpected verdict (${PI_STATE}) — NOT restarting, manual check needed"
|
||||||
|
ISSUES=$((ISSUES + 1))
|
||||||
|
echo " Abiba: ⚠️ Unexpected probe verdict (${PI_STATE}) — NOT restarted" >> "$LOG" ;;
|
||||||
|
esac
|
||||||
|
# -- abiba-leg-end
|
||||||
|
|
||||||
|
# ── Platform B: Tanko (DSH dsh-web on amdpve CT 112) ──
|
||||||
|
# Direct SSH to 192.168.68.122 is not a dependency of this monitor — per-worker
|
||||||
|
# key availability varies — so probes run from the amdpve vantage via `pct exec`.
|
||||||
|
# Tanko's Zulip gateway runs as the dsh-web systemd unit inside CT 112 on amdpve
|
||||||
|
# (192.168.68.15). The gateway binds 127.0.0.1:3080 loopback-only by design — a
|
||||||
|
# remote :3080 probe is refused and is NOT a fault.
|
||||||
|
TANKO_SVC=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.15 \
|
||||||
|
"pct exec 112 -- systemctl is-active dsh-web" 2>/dev/null || true)
|
||||||
|
[ -n "$TANKO_SVC" ] || TANKO_SVC="unknown"
|
||||||
|
TANKO_HTTP=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.15 \
|
||||||
|
"pct exec 112 -- curl -s --connect-timeout 5 --max-time 10 -o /dev/null -w '%{http_code}' http://127.0.0.1:3080/" 2>/dev/null || true)
|
||||||
|
[ -n "$TANKO_HTTP" ] || TANKO_HTTP="000"
|
||||||
|
|
||||||
|
if [ "$TANKO_SVC" != "active" ]; then
|
||||||
|
notify "🔴" "Tanko (DSH dsh-web) service state: $TANKO_SVC — needs restart"
|
||||||
ISSUES=$((ISSUES + 1))
|
ISSUES=$((ISSUES + 1))
|
||||||
echo " Abiba: ❌ Disconnected — restarted" >> "$LOG"
|
echo " Tanko: ❌ service=$TANKO_SVC" >> "$LOG"
|
||||||
elif [ -n "$PI_ERROR" ]; then
|
elif [ "$TANKO_HTTP" = "000" ]; then
|
||||||
notify "🟡" "Abiba pi extension error: ${PI_ERROR:0:100}"
|
notify "🔴" "Tanko (DSH dsh-web) HTTP :3080 connection refused/timeout — needs restart"
|
||||||
echo " Abiba: 🟡 Error: ${PI_ERROR:0:100}" >> "$LOG"
|
ISSUES=$((ISSUES + 1))
|
||||||
elif [ "$PI_RETRIES" -ge 3 ]; then
|
echo " Tanko: ❌ http=000 (refused/timeout)" >> "$LOG"
|
||||||
notify "🟡" "Abiba pi extension: $PI_RETRIES retries — restarting"
|
|
||||||
pm2 restart abiba-zulip 2>/dev/null || true
|
|
||||||
echo " Abiba: 🟡 $PI_RETRIES retries — restarted" >> "$LOG"
|
|
||||||
else
|
else
|
||||||
echo " Abiba: ✅ Connected (processed=$(echo "$PI_HEALTH" | python3 -c "import sys,json; d=json.load(sys.stdin); print(d.get('messages_processed',0))" 2>/dev/null))" >> "$LOG"
|
case "$TANKO_HTTP" in
|
||||||
|
200|301|302|307|308|401|403)
|
||||||
|
echo " Tanko: ✅ service=active http=$TANKO_HTTP" >> "$LOG" ;;
|
||||||
|
*)
|
||||||
|
notify "🟡" "Tanko (DSH dsh-web) HTTP :3080 answered $TANKO_HTTP — running, unexpected status"
|
||||||
|
echo " Tanko: 🟡 service=active http=$TANKO_HTTP (running, warning)" >> "$LOG" ;;
|
||||||
|
esac
|
||||||
fi
|
fi
|
||||||
|
|
||||||
# ── Platform B: Hermes (Tanko) ──
|
# ── Removed: the former "Platform B: Hermes" agent leg ──
|
||||||
TANKO_STATE=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 jerome@192.168.68.122 \
|
# Captain ruling 2026-09-10: that agent moved off this host onto her own
|
||||||
"cat ~/.hermes/gateway_state.json 2>/dev/null" 2>/dev/null || echo "{}")
|
# container (kagentz CT 105 on minipve, dedicated `hermes` user) and is now
|
||||||
TANKO_ZULIP=$(echo "$TANKO_STATE" | python3 -c "
|
# monitored on her side — see the out-of-scope note in zulip-health.prose.md.
|
||||||
import sys,json
|
# The old leg ssh'd to her former CT 100 deployment and read its Hermes gateway
|
||||||
d=json.load(sys.stdin)
|
# state, which reported "unknown" on every run and posted a false 🔴 DM plus an
|
||||||
p=d.get('platforms',{}).get('zulip',{})
|
# #agent-hub stream alert. Do NOT re-add a probe for her: this monitor must
|
||||||
print(p.get('state','unknown'))
|
# never contact her former host.
|
||||||
" 2>/dev/null)
|
|
||||||
|
|
||||||
if [ "$TANKO_ZULIP" != "connected" ]; then
|
|
||||||
notify "🔴" "Tanko (Hermes) Zulip state: $TANKO_ZULIP — needs restart"
|
|
||||||
ISSUES=$((ISSUES + 1))
|
|
||||||
echo " Tanko: ❌ state=$TANKO_ZULIP" >> "$LOG"
|
|
||||||
else
|
|
||||||
echo " Tanko: ✅ Zulip connected" >> "$LOG"
|
|
||||||
fi
|
|
||||||
|
|
||||||
# ── Platform B: Hermes (Mumuni) ──
|
|
||||||
MUMUNI_STATE=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.24 \
|
|
||||||
"cat ~/.hermes/gateway_state.json 2>/dev/null" 2>/dev/null || echo "{}")
|
|
||||||
MUMUNI_ZULIP=$(echo "$MUMUNI_STATE" | python3 -c "
|
|
||||||
import sys,json
|
|
||||||
d=json.load(sys.stdin)
|
|
||||||
p=d.get('platforms',{}).get('zulip',{})
|
|
||||||
print(p.get('state','unknown'))
|
|
||||||
" 2>/dev/null)
|
|
||||||
|
|
||||||
if [ "$MUMUNI_ZULIP" != "connected" ]; then
|
|
||||||
notify "🔴" "Mumuni (Hermes) Zulip state: $MUMUNI_ZULIP"
|
|
||||||
ISSUES=$((ISSUES + 1))
|
|
||||||
echo " Mumuni: ❌ state=$MUMUNI_ZULIP" >> "$LOG"
|
|
||||||
else
|
|
||||||
echo " Mumuni: ✅ Zulip connected" >> "$LOG"
|
|
||||||
fi
|
|
||||||
|
|
||||||
# ── Platform C: Agent Zero (kagentz) ──
|
# ── Platform C: Agent Zero (kagentz) ──
|
||||||
AZ_A2A=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.14 \
|
AZ_A2A=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.14 \
|
||||||
|
|||||||
+25
@@ -0,0 +1,25 @@
|
|||||||
|
{
|
||||||
|
"status": "ok",
|
||||||
|
"platform": "pi",
|
||||||
|
"agent": "abiba",
|
||||||
|
"zulip": {
|
||||||
|
"connected": true,
|
||||||
|
"site": "https://chat.sysloggh.net",
|
||||||
|
"email": "abiba-bot@chat.sysloggh.net",
|
||||||
|
"queue_id": "ee7f8b6d-9d53-48a7-ad58-f6e999771001",
|
||||||
|
"bot_user_id": 21,
|
||||||
|
"messages_processed": 0,
|
||||||
|
"skipped": 0,
|
||||||
|
"last_error": null
|
||||||
|
},
|
||||||
|
"circuit_breaker": {
|
||||||
|
"state": "CLOSED",
|
||||||
|
"failures": 0,
|
||||||
|
"successes": 5,
|
||||||
|
"totalRequests": 5,
|
||||||
|
"failureRate": "0.000",
|
||||||
|
"openedAt": null
|
||||||
|
},
|
||||||
|
"workers": [],
|
||||||
|
"worker_count": 0
|
||||||
|
}
|
||||||
+25
@@ -0,0 +1,25 @@
|
|||||||
|
{
|
||||||
|
"status": "down",
|
||||||
|
"platform": "pi",
|
||||||
|
"agent": "abiba",
|
||||||
|
"zulip": {
|
||||||
|
"connected": false,
|
||||||
|
"site": "https://chat.sysloggh.net",
|
||||||
|
"email": "abiba-bot@chat.sysloggh.net",
|
||||||
|
"queue_id": null,
|
||||||
|
"bot_user_id": null,
|
||||||
|
"messages_processed": 0,
|
||||||
|
"skipped": 0,
|
||||||
|
"last_error": "Zulip API error 401: queue registration failed"
|
||||||
|
},
|
||||||
|
"circuit_breaker": {
|
||||||
|
"state": "CLOSED",
|
||||||
|
"failures": 0,
|
||||||
|
"successes": 0,
|
||||||
|
"totalRequests": 0,
|
||||||
|
"failureRate": "0.000",
|
||||||
|
"openedAt": null
|
||||||
|
},
|
||||||
|
"workers": [],
|
||||||
|
"worker_count": 0
|
||||||
|
}
|
||||||
@@ -0,0 +1,138 @@
|
|||||||
|
"""
|
||||||
|
Regression tests for daily-infra-report.py fixes (PR #64).
|
||||||
|
|
||||||
|
Tests:
|
||||||
|
(a) Asserts the nested zulip read feeds the agent-card fields
|
||||||
|
(b) Asserts an unreachable pve_get renders labelled-unreachable, not "0/0"
|
||||||
|
"""
|
||||||
|
import json
|
||||||
|
import subprocess
|
||||||
|
import sys
|
||||||
|
from pathlib import Path
|
||||||
|
from unittest.mock import patch, MagicMock
|
||||||
|
|
||||||
|
# Add scripts to path
|
||||||
|
sys.path.insert(0, str(Path(__file__).parent.parent / "scripts"))
|
||||||
|
import importlib.util
|
||||||
|
|
||||||
|
def load_script():
|
||||||
|
"""Load the daily-infra-report script as a module."""
|
||||||
|
script_path = Path(__file__).parent.parent / "scripts" / "daily-infra-report.py"
|
||||||
|
spec = importlib.util.spec_from_file_location("daily_infra_report", script_path)
|
||||||
|
module = importlib.util.module_from_spec(spec)
|
||||||
|
spec.loader.exec_module(module)
|
||||||
|
return module
|
||||||
|
|
||||||
|
|
||||||
|
def test_nested_zulip_read_feeds_agent_card():
|
||||||
|
"""Test that Zulip state is read from the nested 'zulip' key and feeds agent-card fields."""
|
||||||
|
# Mock the http_get_body response with nested structure
|
||||||
|
mock_health_response = json.dumps({
|
||||||
|
"status": "ok",
|
||||||
|
"platform": "pi",
|
||||||
|
"agent": "abiba",
|
||||||
|
"zulip": {
|
||||||
|
"connected": True,
|
||||||
|
"queue_id": "test-queue-id",
|
||||||
|
"messages_processed": 42,
|
||||||
|
"skipped": 5,
|
||||||
|
"last_error": None
|
||||||
|
}
|
||||||
|
})
|
||||||
|
|
||||||
|
# Import and patch
|
||||||
|
report_mod = load_script()
|
||||||
|
|
||||||
|
with patch.object(report_mod, 'http_get_body', return_value=mock_health_response):
|
||||||
|
# Simulate the collect() function's Zulip section
|
||||||
|
zulip_health = json.loads(report_mod.http_get_body("http://localhost:9200/health"))
|
||||||
|
zulip_state = zulip_health.get("zulip", {})
|
||||||
|
|
||||||
|
# Assert the nested key is read correctly
|
||||||
|
assert zulip_state.get("connected") == True, "Zulip connected should be True from nested key"
|
||||||
|
assert zulip_state.get("messages_processed") == 42, "messages_processed should be 42 from nested key"
|
||||||
|
assert zulip_state.get("queue_id") == "test-queue-id", "queue_id should be read from nested key"
|
||||||
|
|
||||||
|
# Simulate the agent card field population
|
||||||
|
agent_card = {
|
||||||
|
"zulip_connected": zulip_state.get("connected", False),
|
||||||
|
"zulip_processed": zulip_state.get("messages_processed", 0),
|
||||||
|
}
|
||||||
|
|
||||||
|
assert agent_card["zulip_connected"] == True, "Agent card should show Zulip connected"
|
||||||
|
assert agent_card["zulip_processed"] == 42, "Agent card should show 42 processed messages"
|
||||||
|
|
||||||
|
|
||||||
|
def test_unreachable_pve_get_renders_labelled_unreachable():
|
||||||
|
"""Test that an unreachable PVE API renders 'unreachable' instead of '0/0'."""
|
||||||
|
# Import and patch
|
||||||
|
report_mod = load_script()
|
||||||
|
|
||||||
|
# Test pve_get returns None on error
|
||||||
|
with patch.object(report_mod.subprocess, 'run') as mock_run:
|
||||||
|
mock_run.return_value.returncode = 7 # Connection failure
|
||||||
|
result = report_mod.pve_get("/api2/json/nodes")
|
||||||
|
assert result is None, "pve_get should return None on connection failure"
|
||||||
|
|
||||||
|
# Test the render logic
|
||||||
|
report = {
|
||||||
|
"nodes": {},
|
||||||
|
"node_count": 0,
|
||||||
|
"nodes_online": 0,
|
||||||
|
"pve_probe_status": "unreachable",
|
||||||
|
"total_vms": 0,
|
||||||
|
"running_vms": 0,
|
||||||
|
}
|
||||||
|
|
||||||
|
# The render should show "unreachable" not "0/0"
|
||||||
|
pve_status_label = "unreachable" if report.get('pve_probe_status') == 'unreachable' else f"{report['nodes_online']}/{report['node_count']}"
|
||||||
|
|
||||||
|
assert pve_status_label == "unreachable", "PVE status should show 'unreachable' when probe fails, not '0/0'"
|
||||||
|
|
||||||
|
|
||||||
|
def test_unreachable_resources_renders_labelled_unreachable():
|
||||||
|
"""Test that unreachable resources probe renders 'unreachable' instead of '0/0'."""
|
||||||
|
report_mod = load_script()
|
||||||
|
|
||||||
|
# Test resources probe returns None
|
||||||
|
with patch.object(report_mod.subprocess, 'run') as mock_run:
|
||||||
|
mock_run.return_value.returncode = 7
|
||||||
|
result = report_mod.pve_get("/api2/json/cluster/resources")
|
||||||
|
assert result is None, "pve_get for resources should return None on connection failure"
|
||||||
|
|
||||||
|
# Test the render logic
|
||||||
|
report = {
|
||||||
|
"resources_probe_status": "unreachable",
|
||||||
|
"total_vms": 0,
|
||||||
|
"running_vms": 0,
|
||||||
|
}
|
||||||
|
|
||||||
|
resources_label = "unreachable" if report.get('resources_probe_status') == 'unreachable' else f"{report['running_vms']}/{report['total_vms']}"
|
||||||
|
|
||||||
|
assert resources_label == "unreachable", "Resources status should show 'unreachable' when probe fails, not '0/0'"
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
print("Running tests...")
|
||||||
|
try:
|
||||||
|
test_nested_zulip_read_feeds_agent_card()
|
||||||
|
print("✓ test_nested_zulip_read_feeds_agent_card passed")
|
||||||
|
except AssertionError as e:
|
||||||
|
print(f"✗ test_nested_zulip_read_feeds_agent_card failed: {e}")
|
||||||
|
sys.exit(1)
|
||||||
|
|
||||||
|
try:
|
||||||
|
test_unreachable_pve_get_renders_labelled_unreachable()
|
||||||
|
print("✓ test_unreachable_pve_get_renders_labelled_unreachable passed")
|
||||||
|
except AssertionError as e:
|
||||||
|
print(f"✗ test_unreachable_pve_get_renders_labelled_unreachable failed: {e}")
|
||||||
|
sys.exit(1)
|
||||||
|
|
||||||
|
try:
|
||||||
|
test_unreachable_resources_renders_labelled_unreachable()
|
||||||
|
print("✓ test_unreachable_resources_renders_labelled_unreachable passed")
|
||||||
|
except AssertionError as e:
|
||||||
|
print(f"✗ test_unreachable_resources_renders_labelled_unreachable failed: {e}")
|
||||||
|
sys.exit(1)
|
||||||
|
|
||||||
|
print("All tests passed!")
|
||||||
@@ -0,0 +1,352 @@
|
|||||||
|
"""Regression tests for the 2026-09-10 retirement of the Mumuni monitoring leg.
|
||||||
|
|
||||||
|
WHY THIS FILE EXISTS: captain ruling 2026-09-10 — Mumuni moved off this host
|
||||||
|
onto her own container (kagentz CT 105 on minipve, 192.168.68.14, dedicated
|
||||||
|
`hermes` user) and is monitored from her side. The monitor nevertheless kept
|
||||||
|
ssh'ing to root@192.168.68.24 for `~/.hermes/gateway_state.json` on the
|
||||||
|
decommissioned deployment, read "unknown" on every run, and posted a false 🔴
|
||||||
|
"Mumuni (Hermes) Zulip state: unknown" DM + #agent-hub stream alert to the
|
||||||
|
captain. The daily infra digest published a matching `mumuni:unknown` row.
|
||||||
|
|
||||||
|
CONTRACT UNDER TEST:
|
||||||
|
* `scripts/zulip-monitor.sh` carries NO Mumuni probe and NO 192.168.68.24
|
||||||
|
reference; it never ssh'es .24, and even on a failing run it emits no Mumuni
|
||||||
|
notify (stdout alert, Zulip payload, or log line).
|
||||||
|
* The Abiba (pi — the Zulip bridge), Tanko (DSH) and Agent Zero (kagentz) legs
|
||||||
|
still work: deleting the Mumuni leg must not have gutted the rest.
|
||||||
|
* `scripts/daily-infra-report.py` no longer probes .24 for a Hermes gateway
|
||||||
|
state and no longer emits a `mumuni` agent entry.
|
||||||
|
* `scripts/agent-health-check.py`'s AGENTS roster has no mumuni entry. This is
|
||||||
|
a pin, not a behavior change — verify the probe was already gone.
|
||||||
|
* `zulip-health.prose.md` retires the Mumuni-only steps and says explicitly
|
||||||
|
that Mumuni is not monitored from this host.
|
||||||
|
|
||||||
|
HOW: behavioral execution plus one named deliverable-text contract. The sandbox
|
||||||
|
copies the shipped monitor verbatim and rewrites only its LOG constant, then
|
||||||
|
runs it with stub ssh/curl on PATH; the ssh stub records every host it is asked
|
||||||
|
to reach, so "never probes .24" and "no Mumuni notify" are asserted from
|
||||||
|
observed behavior. The daily digest is pinned by importing it and exercising
|
||||||
|
collect() and build_html() directly. The single source-text assertion is the
|
||||||
|
deliverable-text contract the captain acceptance names for the shipped monitor.
|
||||||
|
|
||||||
|
Usage: python3 -m pytest tests/test_mumuni_monitor_removal.py
|
||||||
|
"""
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import importlib.util
|
||||||
|
import os
|
||||||
|
import pathlib
|
||||||
|
import stat
|
||||||
|
import subprocess
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
ROOT = pathlib.Path(__file__).resolve().parents[1]
|
||||||
|
ZULIP_MONITOR = ROOT / "scripts" / "zulip-monitor.sh"
|
||||||
|
DAILY_REPORT = ROOT / "scripts" / "daily-infra-report.py"
|
||||||
|
AHC = ROOT / "scripts" / "agent-health-check.py"
|
||||||
|
HEALTH_CONTRACT = ROOT / "zulip-health.prose.md"
|
||||||
|
CONNECTED_FIXTURE = ROOT / "tests" / "fixtures" / "zulip-health-connected.json"
|
||||||
|
|
||||||
|
MUMUNI_IP = "192.168.68.24" # Mumuni's old (decommissioned) deployment
|
||||||
|
TANKO_VANTAGE = "192.168.68.15" # amdpve — Tanko CT 112 via pct exec
|
||||||
|
AGENT_ZERO_HOST = "192.168.68.14" # kagentz host, Agent Zero docker
|
||||||
|
|
||||||
|
|
||||||
|
# ── scripts/zulip-monitor.sh: deliverable-text contract ─────────────
|
||||||
|
|
||||||
|
def test_zulip_monitor_deliverable_text_contract():
|
||||||
|
"""Owned deliverable-text contract for scripts/zulip-monitor.sh.
|
||||||
|
|
||||||
|
Captain acceptance requires the shipped monitor to contain no Mumuni probe
|
||||||
|
identifier and no 192.168.68.24 literal. Behavioral proof that the monitor
|
||||||
|
never contacts that host and never emits a Mumuni notify lives in the
|
||||||
|
sandbox tests below; this only pins the named text contract.
|
||||||
|
"""
|
||||||
|
text = ZULIP_MONITOR.read_text()
|
||||||
|
assert "mumuni" not in text.lower()
|
||||||
|
assert MUMUNI_IP not in text
|
||||||
|
|
||||||
|
|
||||||
|
# ── scripts/zulip-monitor.sh: behavioral sandbox ─────────────────────
|
||||||
|
|
||||||
|
SSH_STUB = r"""#!/usr/bin/env bash
|
||||||
|
# Stub ssh: record the target host, then answer by host + remote command.
|
||||||
|
printf '%s\n' "$*" >> "$RECORD_DIR/ssh.calls"
|
||||||
|
host=""
|
||||||
|
for a in "$@"; do
|
||||||
|
case "$a" in
|
||||||
|
*@192.168.*) host="${a##*@}" ;;
|
||||||
|
esac
|
||||||
|
done
|
||||||
|
printf '%s\n' "$host" >> "$RECORD_DIR/ssh.hosts"
|
||||||
|
cmd="${*: -1}"
|
||||||
|
case "$host" in
|
||||||
|
192.168.68.15)
|
||||||
|
case "$cmd" in
|
||||||
|
*"systemctl is-active"*) printf '%s' "$TANKO_SVC" ;;
|
||||||
|
*curl*) printf '%s' "$TANKO_HTTP" ;;
|
||||||
|
esac ;;
|
||||||
|
192.168.68.14)
|
||||||
|
case "$cmd" in
|
||||||
|
*agent.json*) printf '%s' "$AZ_A2A" ;;
|
||||||
|
*"ps aux"*) printf '%s\n' "$AZ_PS" ;;
|
||||||
|
esac ;;
|
||||||
|
*)
|
||||||
|
printf 'UNEXPECTED-SSH-HOST %s\n' "$host" >> "$RECORD_DIR/unexpected-ssh" ;;
|
||||||
|
esac
|
||||||
|
exit 0
|
||||||
|
"""
|
||||||
|
|
||||||
|
CURL_STUB = r"""#!/usr/bin/env bash
|
||||||
|
# Stub curl: serve the Abiba health fixture and the Zulip server 200, and
|
||||||
|
# record every call (including notify) payloads.
|
||||||
|
printf '%s\n' "$*" >> "$RECORD_DIR/curl.calls"
|
||||||
|
case "$*" in
|
||||||
|
*:9200/health*)
|
||||||
|
case " $* " in
|
||||||
|
*" -w "*) printf '%s' "$PI_HTTP" ;; # -w '%{http_code}' probe
|
||||||
|
*) printf '%s' "$PI_BODY" ;; # body probe
|
||||||
|
esac ;;
|
||||||
|
*server_settings*)
|
||||||
|
printf '%s' "$SERVER_HTTP" ;;
|
||||||
|
esac
|
||||||
|
exit 0
|
||||||
|
"""
|
||||||
|
|
||||||
|
|
||||||
|
def _write_exec(path: pathlib.Path, body: str) -> None:
|
||||||
|
path.write_text(body)
|
||||||
|
path.chmod(path.stat().st_mode
|
||||||
|
| stat.S_IXUSR | stat.S_IXGRP | stat.S_IXOTH)
|
||||||
|
|
||||||
|
|
||||||
|
def _run_monitor(tmp_path, *, tanko_svc="active", tanko_http="200",
|
||||||
|
az_a2a='{"name":"kagentz"}',
|
||||||
|
az_ps="root 111 0.1 0.2 /opt/venv-a0/bin/python3 -u adapter.py"):
|
||||||
|
"""Run the shipped monitor in a sandbox; return (proc, record_dir, log_path).
|
||||||
|
|
||||||
|
Only the LOG constant is rewritten (to keep the run inside the worktree).
|
||||||
|
Everything else — legs, labels, notify logic — is the shipped script.
|
||||||
|
"""
|
||||||
|
sandbox = tmp_path / "sandbox"
|
||||||
|
bindir = sandbox / "bin"
|
||||||
|
record = sandbox / "record"
|
||||||
|
bindir.mkdir(parents=True)
|
||||||
|
record.mkdir()
|
||||||
|
|
||||||
|
_write_exec(bindir / "ssh", SSH_STUB)
|
||||||
|
_write_exec(bindir / "curl", CURL_STUB)
|
||||||
|
|
||||||
|
source = ZULIP_MONITOR.read_text()
|
||||||
|
log_line = 'LOG="/root/zulip-health-monitor.log"'
|
||||||
|
assert log_line in source, "LOG constant moved — update the sandbox harness"
|
||||||
|
log_path = sandbox / "zulip-health-monitor.log"
|
||||||
|
script = sandbox / "zulip-monitor.sh"
|
||||||
|
script.write_text(source.replace(log_line, f'LOG="{log_path}"'))
|
||||||
|
|
||||||
|
env = dict(os.environ)
|
||||||
|
env.update({
|
||||||
|
"PATH": f"{bindir}:{env['PATH']}",
|
||||||
|
"RECORD_DIR": str(record),
|
||||||
|
"TANKO_SVC": tanko_svc,
|
||||||
|
"TANKO_HTTP": tanko_http,
|
||||||
|
"AZ_A2A": az_a2a,
|
||||||
|
"AZ_PS": az_ps,
|
||||||
|
"PI_HTTP": "200",
|
||||||
|
"PI_BODY": CONNECTED_FIXTURE.read_text(),
|
||||||
|
"SERVER_HTTP": "200",
|
||||||
|
})
|
||||||
|
proc = subprocess.run(["bash", str(script)], cwd=sandbox, env=env,
|
||||||
|
capture_output=True, text=True)
|
||||||
|
return proc, record, log_path
|
||||||
|
|
||||||
|
|
||||||
|
def test_healthy_run_is_quiet_and_never_reaches_mumuni(tmp_path):
|
||||||
|
proc, record, log_path = _run_monitor(tmp_path)
|
||||||
|
assert proc.returncode == 0, proc.stderr
|
||||||
|
log = log_path.read_text()
|
||||||
|
|
||||||
|
# Every retained leg actually ran and passed.
|
||||||
|
assert "Server: ✅ HTTP 200" in log
|
||||||
|
assert "Abiba: ✅ Connected" in log
|
||||||
|
assert "Tanko: ✅ service=active http=200" in log
|
||||||
|
assert "kagentz: ✅ A2A alive" in log
|
||||||
|
assert "kagentz: ✅ Adapter running" in log
|
||||||
|
assert "Result: ✅ All healthy" in log
|
||||||
|
|
||||||
|
# A healthy run emits no notify at all — and certainly no Mumuni one.
|
||||||
|
assert proc.stdout == ""
|
||||||
|
assert "Mumuni" not in log
|
||||||
|
assert "🔴" not in log
|
||||||
|
|
||||||
|
# Observed behavior: .24 is never resolved, only Tanko's vantage and the
|
||||||
|
# Agent Zero host are contacted.
|
||||||
|
hosts = record.joinpath("ssh.hosts").read_text().split()
|
||||||
|
assert MUMUNI_IP not in hosts
|
||||||
|
assert set(hosts) == {TANKO_VANTAGE, AGENT_ZERO_HOST}
|
||||||
|
assert not record.joinpath("unexpected-ssh").exists()
|
||||||
|
|
||||||
|
|
||||||
|
def test_failing_run_alerts_on_tanko_but_never_on_mumuni(tmp_path):
|
||||||
|
# Failure path: exercises notify() end to end so "no Mumuni notify" is
|
||||||
|
# proven on the alert path, not only on the quiet healthy path.
|
||||||
|
proc, record, log_path = _run_monitor(tmp_path, tanko_svc="inactive",
|
||||||
|
tanko_http="000")
|
||||||
|
assert proc.returncode == 0, proc.stderr
|
||||||
|
|
||||||
|
alerts = proc.stdout
|
||||||
|
assert "Tanko (DSH dsh-web) service state: inactive" in alerts
|
||||||
|
assert "1 issue(s) found" in alerts
|
||||||
|
|
||||||
|
# No Mumuni text in stdout, the log, or any Zulip DM/stream payload.
|
||||||
|
assert "Mumuni" not in alerts
|
||||||
|
assert "Mumuni" not in log_path.read_text()
|
||||||
|
assert MUMUNI_IP not in alerts + log_path.read_text()
|
||||||
|
payloads = record.joinpath("curl.calls").read_text()
|
||||||
|
assert "Mumuni" not in payloads
|
||||||
|
assert MUMUNI_IP not in payloads
|
||||||
|
|
||||||
|
# The rest of the monitor still ran alongside the failing Tanko leg.
|
||||||
|
log = log_path.read_text()
|
||||||
|
assert "Abiba: ✅ Connected" in log
|
||||||
|
assert "kagentz: ✅ A2A alive" in log
|
||||||
|
assert "Result: 🔴 1 issue(s) found" in log
|
||||||
|
|
||||||
|
|
||||||
|
# ── scripts/daily-infra-report.py: behavioral digest checks ──────────
|
||||||
|
|
||||||
|
@pytest.fixture(scope="module")
|
||||||
|
def daily():
|
||||||
|
spec = importlib.util.spec_from_file_location("daily_infra_report", DAILY_REPORT)
|
||||||
|
assert spec and spec.loader
|
||||||
|
module = importlib.util.module_from_spec(spec)
|
||||||
|
spec.loader.exec_module(module)
|
||||||
|
return module
|
||||||
|
|
||||||
|
|
||||||
|
DAILY_AGENTS = {
|
||||||
|
"abiba": {
|
||||||
|
"platform": "pi", "ct": 100, "ip": MUMUNI_IP,
|
||||||
|
"zulip_connected": True, "zulip_processed": 5,
|
||||||
|
"pm2_status": "online", "pm2_restarts": "0", "pm2_uptime": "1h",
|
||||||
|
},
|
||||||
|
"tanko": {
|
||||||
|
"platform": "dsh", "ct": 112, "ip": "192.168.68.122",
|
||||||
|
"gateway_state": "n/a (DSH)", "zulip_state": "connected",
|
||||||
|
"telegram_state": "unknown", "gateway_pid": None, "updated_at": "",
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _fabricated_report(agents):
|
||||||
|
return {
|
||||||
|
"nodes": {},
|
||||||
|
"node_count": 1,
|
||||||
|
"nodes_online": 1,
|
||||||
|
"total_vms": 0,
|
||||||
|
"running_vms": 0,
|
||||||
|
"stopped_vms": [],
|
||||||
|
"vms_by_node": {n: [] for n in
|
||||||
|
["amdpve", "minipve", "storepve", "acerpve", "ocupve"]},
|
||||||
|
"storage": [],
|
||||||
|
"docker_vm": {"total": 0, "running": 0, "unhealthy": [],
|
||||||
|
"containers": [], "reclaimable": "", "disk_used": "1%"},
|
||||||
|
"docker_syslog": {"total": 0, "running": 0, "containers": []},
|
||||||
|
"docker_netbird": {"total": 0, "running": 0, "containers": []},
|
||||||
|
"endpoints": [],
|
||||||
|
"litellm": {"checks": []},
|
||||||
|
"nfs": [],
|
||||||
|
"zulip_ext": {
|
||||||
|
"connected": True, "queue_id": "queue", "last_error": None,
|
||||||
|
"messages_processed": 0, "retry_count": 0, "pm2": {},
|
||||||
|
"pm2_healthy": True, "bot_skipped_15min": 0, "finalized_1h": 0,
|
||||||
|
"failed_finalize_1h": 0, "finalize_fail_pct": 0,
|
||||||
|
"server_status": "200",
|
||||||
|
},
|
||||||
|
"agents": agents,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _agent_status_card(html):
|
||||||
|
start = html.index("🤖 Agent Status")
|
||||||
|
end = html.index("💬 Zulip Extension")
|
||||||
|
return html[start:end]
|
||||||
|
|
||||||
|
|
||||||
|
def test_daily_report_renders_only_abiba_and_tanko_agents(daily):
|
||||||
|
"""build_html() over a Mumuni-free agent set must render no Mumuni row and
|
||||||
|
no Mumuni gateway-unknown issue, while abiba and tanko rows still render."""
|
||||||
|
html = daily.build_html(_fabricated_report(dict(DAILY_AGENTS)))
|
||||||
|
card = _agent_status_card(html)
|
||||||
|
assert "mumuni" not in card.lower()
|
||||||
|
assert "abiba" in card
|
||||||
|
assert "tanko" in card
|
||||||
|
assert "mumuni" not in html.lower()
|
||||||
|
|
||||||
|
|
||||||
|
def test_daily_report_collect_never_probes_mumuni(monkeypatch, daily):
|
||||||
|
"""collect() with ssh stubbed must add no mumuni agent and must never ssh
|
||||||
|
its decommissioned .24 host."""
|
||||||
|
probed = []
|
||||||
|
|
||||||
|
class _NoSubprocess:
|
||||||
|
@staticmethod
|
||||||
|
def check_output(*args, **kwargs):
|
||||||
|
return b""
|
||||||
|
|
||||||
|
def fake_ssh(host, cmd):
|
||||||
|
probed.append(host)
|
||||||
|
return ""
|
||||||
|
|
||||||
|
monkeypatch.setattr(daily, "pve_get", lambda path: [])
|
||||||
|
monkeypatch.setattr(daily, "ssh_jerome", lambda host, cmd: "")
|
||||||
|
monkeypatch.setattr(daily, "ssh", fake_ssh)
|
||||||
|
monkeypatch.setattr(daily, "http_get",
|
||||||
|
lambda url, auth=None, timeout=10: "200")
|
||||||
|
monkeypatch.setattr(daily, "http_get_body",
|
||||||
|
lambda url, auth=None, timeout=10: "")
|
||||||
|
monkeypatch.setattr(daily, "count_in_log", lambda *a, **k: 0)
|
||||||
|
monkeypatch.setattr(daily, "subprocess", _NoSubprocess)
|
||||||
|
|
||||||
|
report = daily.collect()
|
||||||
|
assert "mumuni" not in report["agents"]
|
||||||
|
assert MUMUNI_IP not in probed
|
||||||
|
|
||||||
|
|
||||||
|
# ── scripts/agent-health-check.py: roster pin ───────────────────────
|
||||||
|
|
||||||
|
@pytest.fixture(scope="module")
|
||||||
|
def ahc():
|
||||||
|
spec = importlib.util.spec_from_file_location("agent_health_check_roster", AHC)
|
||||||
|
assert spec and spec.loader
|
||||||
|
module = importlib.util.module_from_spec(spec)
|
||||||
|
spec.loader.exec_module(module)
|
||||||
|
return module
|
||||||
|
|
||||||
|
|
||||||
|
def test_agent_health_roster_has_no_mumuni_entry(ahc):
|
||||||
|
assert "mumuni" not in ahc.AGENTS
|
||||||
|
|
||||||
|
|
||||||
|
# ── zulip-health.prose.md: contract reconciliation ──────────────────
|
||||||
|
|
||||||
|
def test_health_contract_retires_mumuni_only_steps():
|
||||||
|
text = HEALTH_CONTRACT.read_text()
|
||||||
|
assert MUMUNI_IP not in text
|
||||||
|
for step in ("**B4:", "**B5:", "**B6:"):
|
||||||
|
assert step not in text
|
||||||
|
|
||||||
|
|
||||||
|
def test_health_contract_states_mumuni_is_not_monitored_from_this_host():
|
||||||
|
text = HEALTH_CONTRACT.read_text()
|
||||||
|
assert "Mumuni is NOT monitored from this host" in text
|
||||||
|
assert "monitored on her side" in text
|
||||||
|
assert "her own container" in text
|
||||||
|
|
||||||
|
|
||||||
|
def test_health_contract_keeps_tanko_agent_zero_and_bridge_steps():
|
||||||
|
text = HEALTH_CONTRACT.read_text()
|
||||||
|
for marker in ("**B1:", "**B2:", "**B3:", "Step 4: Platform C",
|
||||||
|
"Step 2: Platform A", "Step 1: Zulip Server Liveness"):
|
||||||
|
assert marker in text, marker
|
||||||
@@ -0,0 +1,466 @@
|
|||||||
|
"""Regression tests for the 2026-09-09/10 probe-drift corrections.
|
||||||
|
|
||||||
|
WHY THIS FILE EXISTS: the monitoring contracts kept emitting false alarms from
|
||||||
|
stale expectations rather than live faults.
|
||||||
|
|
||||||
|
* agent-health-check v3 reported 6 failures that were all stale expectations:
|
||||||
|
abiba (pi-only since the harness purge) was tested as a Hermes host, koby
|
||||||
|
(report-only per the captain's 2026-08-17 ruling) was counted as repairable,
|
||||||
|
koby's CT 111 was probed on amdpve where it does not exist (it runs on
|
||||||
|
storepve .6), and the wrapper infisical check had two bugs — it read only
|
||||||
|
the first 20 lines, so koonimo's wrapper (which references /usr/bin/infisical
|
||||||
|
past line 20) false-failed, and it treated koby's genuine no-infisical
|
||||||
|
(~/.hermes/.env) wrapper as broken.
|
||||||
|
* gpu-monitor emitted "DEGRADED — GPU-rtx3090 000, GPU-rtx5070 000" three
|
||||||
|
times from probing bare port 80 on GPU hosts while :8080 answered 200.
|
||||||
|
* infrastructure-monitoring probed CT 116 for the PVE API (no pveproxy ->
|
||||||
|
000) instead of the five real cluster nodes, which answer 401 = alive.
|
||||||
|
|
||||||
|
These tests execute the health script (with SSH/vault stubbed) and the real
|
||||||
|
provenance consumer (scripts/prose-lint.sh), and parse the contracts' executable
|
||||||
|
check-health probe blocks into normalized probe sets. No live network, vault, or
|
||||||
|
SSH access is required.
|
||||||
|
"""
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import importlib.util
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
import pathlib
|
||||||
|
import re
|
||||||
|
import subprocess
|
||||||
|
import sys
|
||||||
|
import textwrap
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
ROOT = pathlib.Path(__file__).resolve().parents[1]
|
||||||
|
AHC = ROOT / "scripts" / "agent-health-check.py"
|
||||||
|
LINT = ROOT / "scripts" / "prose-lint.sh"
|
||||||
|
GPU = ROOT / "gpu-monitor.prose.md"
|
||||||
|
INFRA = ROOT / "infrastructure-monitoring.prose.md"
|
||||||
|
|
||||||
|
PVE_NODE_IPS = {
|
||||||
|
"192.168.68.9",
|
||||||
|
"192.168.68.5",
|
||||||
|
"192.168.68.15",
|
||||||
|
"192.168.68.6",
|
||||||
|
"192.168.68.12",
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.fixture(scope="module")
|
||||||
|
def ahc():
|
||||||
|
"""Import agent-health-check.py without live network/SSH side effects."""
|
||||||
|
spec = importlib.util.spec_from_file_location("agent_health_check", AHC)
|
||||||
|
assert spec and spec.loader
|
||||||
|
module = importlib.util.module_from_spec(spec)
|
||||||
|
spec.loader.exec_module(module)
|
||||||
|
return module
|
||||||
|
|
||||||
|
|
||||||
|
# ── helpers: execute the health script with SSH/vault stubbed ─────────
|
||||||
|
|
||||||
|
def _run_main(ahc, monkeypatch, capsys, argv, ssh_result=None):
|
||||||
|
ahc.FAIL.clear()
|
||||||
|
ahc.REPORT_ONLY.clear()
|
||||||
|
monkeypatch.setattr(ahc, "load_agent_keys", lambda: None)
|
||||||
|
monkeypatch.setattr(ahc, "ssh", lambda *a, **k: ssh_result)
|
||||||
|
monkeypatch.setattr(sys, "argv", ["agent-health-check.py", "--no-deploy", *argv])
|
||||||
|
with pytest.raises(SystemExit) as exc:
|
||||||
|
ahc.main()
|
||||||
|
return exc.value.code, capsys.readouterr().out
|
||||||
|
|
||||||
|
|
||||||
|
def _json_payload(out):
|
||||||
|
for line in reversed(out.splitlines()):
|
||||||
|
if line.startswith('{"timestamp"'):
|
||||||
|
return json.loads(line)
|
||||||
|
raise AssertionError(f"no JSON payload in output:\n{out}")
|
||||||
|
|
||||||
|
|
||||||
|
# ── agent-health-check: stale-expectation legs ───────────────────────
|
||||||
|
|
||||||
|
def test_import_does_not_contact_vault(ahc):
|
||||||
|
# Keys are loaded in main() via load_agent_keys(); importing must stay inert.
|
||||||
|
assert callable(ahc.load_agent_keys)
|
||||||
|
assert all(agent.get("key") is None for agent in ahc.AGENTS.values())
|
||||||
|
|
||||||
|
|
||||||
|
def test_abiba_is_pi_only_runtime(ahc):
|
||||||
|
# .24 has run pi-only since the harness purge: no Hermes gateway, config, or
|
||||||
|
# wrapper. Probing those legs produced false failures.
|
||||||
|
assert ahc.AGENTS["abiba"]["runtime"] == "pi"
|
||||||
|
|
||||||
|
|
||||||
|
def test_koby_is_report_only(ahc):
|
||||||
|
# Captain's 2026-08-17 ruling (Rule 17): detect and report, never repair.
|
||||||
|
assert ahc.AGENTS["koby"]["report_only"] is True
|
||||||
|
|
||||||
|
|
||||||
|
def test_koby_ct111_is_on_storepve(ahc):
|
||||||
|
# Live-verified 2026-09-10: `pct status 111` = running on storepve (.6);
|
||||||
|
# amdpve has no lxc/111.conf, which is what false-failed before.
|
||||||
|
assert ahc.AGENTS["koby"]["pve"] == "storepve"
|
||||||
|
|
||||||
|
|
||||||
|
def test_report_only_legs_never_count_as_failures(ahc):
|
||||||
|
for agent, report_only in (("koby", True), ("koonimo", False), ("tanko", False)):
|
||||||
|
ahc.FAIL.clear()
|
||||||
|
ahc.REPORT_ONLY.clear()
|
||||||
|
ahc._fail(f"probe:{agent}", agent)
|
||||||
|
if report_only:
|
||||||
|
assert ahc.FAIL == []
|
||||||
|
assert ahc.REPORT_ONLY == [f"probe:{agent}"]
|
||||||
|
else:
|
||||||
|
assert ahc.FAIL == [f"probe:{agent}"]
|
||||||
|
assert ahc.REPORT_ONLY == []
|
||||||
|
ahc.FAIL.clear()
|
||||||
|
ahc.REPORT_ONLY.clear()
|
||||||
|
|
||||||
|
|
||||||
|
def test_failure_recording_accepts_agentless_keys(ahc):
|
||||||
|
ahc.FAIL.clear()
|
||||||
|
try:
|
||||||
|
ahc._fail("gpu-no-port:gpu-rtx3090 (.8)")
|
||||||
|
assert ahc.FAIL == ["gpu-no-port:gpu-rtx3090 (.8)"]
|
||||||
|
finally:
|
||||||
|
ahc.FAIL.clear()
|
||||||
|
|
||||||
|
|
||||||
|
def test_json_reports_absolute_execution_provenance(ahc, monkeypatch, capsys):
|
||||||
|
code, out = _run_main(ahc, monkeypatch, capsys, ["--json"])
|
||||||
|
payload = _json_payload(out)
|
||||||
|
assert payload["execution_path"] == os.path.abspath(str(AHC))
|
||||||
|
assert payload["cwd"] == os.getcwd()
|
||||||
|
assert code == 1 # stubbed SSH fails every leg, but provenance is still emitted
|
||||||
|
|
||||||
|
|
||||||
|
def test_quiet_run_still_carries_provenance_on_the_alert_path(ahc, monkeypatch, capsys):
|
||||||
|
# The cron runs --quiet; a failure report must still carry provenance. The
|
||||||
|
# header line is suppressed in quiet mode, so the ALERT line is the carrier.
|
||||||
|
code, out = _run_main(ahc, monkeypatch, capsys, ["--quiet"])
|
||||||
|
assert code == 1
|
||||||
|
assert "📍 executed from:" not in out
|
||||||
|
alerts = [ln for ln in out.splitlines() if ln.startswith("ALERT agent-health:")]
|
||||||
|
assert alerts, out
|
||||||
|
assert f"script={os.path.abspath(str(AHC))}" in alerts[0]
|
||||||
|
assert f"cwd={os.getcwd()}" in alerts[0]
|
||||||
|
|
||||||
|
|
||||||
|
def test_quiet_healthy_run_emits_no_stdout(ahc, monkeypatch, capsys):
|
||||||
|
# --quiet is documented as "only output on failure": a run with no fleet
|
||||||
|
# failures must produce no stdout at all (the production cron runs --quiet).
|
||||||
|
for name in ("check_keys", "check_gpu_ports", "check_agents", "check_ct_liveness",
|
||||||
|
"check_config_integrity", "check_wrapper_integrity", "check_vault_secrets"):
|
||||||
|
monkeypatch.setattr(ahc, name, lambda: None)
|
||||||
|
code, out = _run_main(ahc, monkeypatch, capsys, ["--quiet"])
|
||||||
|
assert code == 0
|
||||||
|
assert out == ""
|
||||||
|
|
||||||
|
|
||||||
|
def test_json_surfaces_report_only_findings_separately(ahc, monkeypatch, capsys):
|
||||||
|
# Koby's down legs are reported but must not count as fleet failures; the
|
||||||
|
# --json payload exposes them in their own array (item 1 + f8).
|
||||||
|
_, out = _run_main(ahc, monkeypatch, capsys, ["--json"])
|
||||||
|
payload = _json_payload(out)
|
||||||
|
assert isinstance(payload["report_only"], list)
|
||||||
|
assert any(key.startswith(("gateway-down:koby", "ct-unreachable:koby"))
|
||||||
|
for key in payload["report_only"])
|
||||||
|
assert not any("koby" in key for key in payload["failures"])
|
||||||
|
|
||||||
|
|
||||||
|
# ── agent-health-check: wrapper infisical behavior (f3) ───────────────
|
||||||
|
|
||||||
|
def _stub_wrapper_ssh(ahc, monkeypatch, wrapper_body, test_x_result="OK", command_v="/usr/local/bin/infisical"):
|
||||||
|
def fake_ssh(host, cmd, user="root"):
|
||||||
|
if cmd.startswith("cat /root/.local/bin/hermes"):
|
||||||
|
return wrapper_body
|
||||||
|
if cmd.startswith("ls -la /root/.local/bin/hermes "):
|
||||||
|
return "-rwxr-xr-x 1 root root 0 Jan 1 00:00 /root/.local/bin/hermes"
|
||||||
|
if cmd.startswith("ls -la /root/.local/bin/hermes-real") or "venv/bin/hermes" in cmd:
|
||||||
|
return "-rwxr-xr-x 1 root root 0 Jan 1 00:00 /root/.local/bin/hermes-real"
|
||||||
|
if cmd.startswith("grep -c 'LITELLM_API_KEY'"):
|
||||||
|
return "1"
|
||||||
|
if cmd.startswith("test -x "):
|
||||||
|
path = cmd[len("test -x "):].split()[0]
|
||||||
|
if isinstance(test_x_result, dict):
|
||||||
|
return test_x_result.get(path, "MISS")
|
||||||
|
return test_x_result
|
||||||
|
if cmd.startswith("command -v infisical"):
|
||||||
|
return command_v
|
||||||
|
return None
|
||||||
|
|
||||||
|
monkeypatch.setattr(ahc, "ssh", fake_ssh)
|
||||||
|
monkeypatch.setattr(ahc, "AGENTS", {"koonimo": dict(ahc.AGENTS["koonimo"])})
|
||||||
|
ahc.FAIL.clear()
|
||||||
|
ahc.REPORT_ONLY.clear()
|
||||||
|
|
||||||
|
|
||||||
|
def test_env_based_wrapper_without_infisical_is_not_failed(ahc, monkeypatch, capsys):
|
||||||
|
_stub_wrapper_ssh(ahc, monkeypatch,
|
||||||
|
"#!/bin/bash\nsource ~/.hermes/.env\nexec hermes-real \"$@\"\n")
|
||||||
|
ahc.check_wrapper_integrity()
|
||||||
|
out = capsys.readouterr().out
|
||||||
|
assert ahc.FAIL == []
|
||||||
|
assert "wrapper resolves creds without infisical" in out
|
||||||
|
|
||||||
|
|
||||||
|
def test_dangling_absolute_infisical_path_is_failed(ahc, monkeypatch, capsys):
|
||||||
|
# Wrapper hardcodes /usr/bin/infisical, which is absent, while PATH resolves
|
||||||
|
# infisical to /usr/local/bin/infisical. The literal path must be verified,
|
||||||
|
# not inferred from PATH resolution.
|
||||||
|
_stub_wrapper_ssh(ahc, monkeypatch,
|
||||||
|
"#!/bin/bash\n/usr/bin/infisical run -- hermes-real \"$@\"\n",
|
||||||
|
test_x_result="MISS", command_v="/usr/local/bin/infisical")
|
||||||
|
ahc.check_wrapper_integrity()
|
||||||
|
assert "wrapper-infisical-path:koonimo" in ahc.FAIL
|
||||||
|
|
||||||
|
|
||||||
|
def test_existing_absolute_infisical_path_passes(ahc, monkeypatch, capsys):
|
||||||
|
_stub_wrapper_ssh(ahc, monkeypatch,
|
||||||
|
"#!/bin/bash\n/usr/bin/infisical run -- hermes-real \"$@\"\n",
|
||||||
|
test_x_result="OK")
|
||||||
|
ahc.check_wrapper_integrity()
|
||||||
|
out = capsys.readouterr().out
|
||||||
|
assert ahc.FAIL == []
|
||||||
|
assert "wrapper infisical path OK" in out
|
||||||
|
|
||||||
|
|
||||||
|
def test_comment_mentioning_removed_infisical_path_is_not_failed(ahc, monkeypatch, capsys):
|
||||||
|
# litellm-api-keys.prose.md documents `rm -f /usr/local/bin/infisical`; a
|
||||||
|
# wrapper comment about that migration must not manufacture a dangling path
|
||||||
|
# when the real invocation (/usr/bin/infisical) is present and executable.
|
||||||
|
_stub_wrapper_ssh(ahc, monkeypatch,
|
||||||
|
"#!/bin/bash\n# migrated from /usr/local/bin/infisical\n"
|
||||||
|
"exec /usr/bin/infisical run -- hermes-real \"$@\"\n",
|
||||||
|
test_x_result={"/usr/bin/infisical": "OK",
|
||||||
|
"/usr/local/bin/infisical": "MISS"})
|
||||||
|
ahc.check_wrapper_integrity()
|
||||||
|
out = capsys.readouterr().out
|
||||||
|
assert ahc.FAIL == []
|
||||||
|
assert "wrapper infisical path OK" in out
|
||||||
|
|
||||||
|
|
||||||
|
def test_comment_only_infisical_mention_does_not_reach_path_check(ahc, monkeypatch, capsys):
|
||||||
|
# A comment-only mention of a removed infisical path on a healthy .env-based
|
||||||
|
# wrapper is not an invocation: it must not fall through to the `command -v`
|
||||||
|
# PATH check and false-FAIL `wrapper-no-infisical`.
|
||||||
|
_stub_wrapper_ssh(ahc, monkeypatch,
|
||||||
|
"#!/bin/bash\n# migrated from /usr/local/bin/infisical\n"
|
||||||
|
"source ~/.hermes/.env\nexec hermes-real \"$@\"\n",
|
||||||
|
test_x_result="MISS", command_v=None)
|
||||||
|
ahc.check_wrapper_integrity()
|
||||||
|
out = capsys.readouterr().out
|
||||||
|
assert ahc.FAIL == []
|
||||||
|
assert "wrapper resolves creds without infisical" in out
|
||||||
|
|
||||||
|
|
||||||
|
# ── item 4: prose-lint enforces report provenance (real consumer) ─────
|
||||||
|
|
||||||
|
GOOD_CONTRACT = textwrap.dedent("""\
|
||||||
|
---
|
||||||
|
kind: function
|
||||||
|
name: good
|
||||||
|
description: fixture with provenance
|
||||||
|
---
|
||||||
|
|
||||||
|
## Parameters
|
||||||
|
|
||||||
|
- x: y
|
||||||
|
|
||||||
|
## Returns
|
||||||
|
|
||||||
|
ok
|
||||||
|
|
||||||
|
### check-health
|
||||||
|
|
||||||
|
```bash
|
||||||
|
pwd -P
|
||||||
|
```
|
||||||
|
|
||||||
|
**Report format**: Begin with the absolute path the probe executed from.
|
||||||
|
""")
|
||||||
|
|
||||||
|
DECOY_CONTRACT = textwrap.dedent("""\
|
||||||
|
---
|
||||||
|
kind: function
|
||||||
|
name: decoy
|
||||||
|
description: fixture with provenance only outside the report format
|
||||||
|
---
|
||||||
|
|
||||||
|
## Parameters
|
||||||
|
|
||||||
|
- x: y
|
||||||
|
|
||||||
|
## Returns
|
||||||
|
|
||||||
|
ok
|
||||||
|
|
||||||
|
The absolute path of the config is /etc/foo.
|
||||||
|
|
||||||
|
### check-health
|
||||||
|
|
||||||
|
```bash
|
||||||
|
true
|
||||||
|
```
|
||||||
|
|
||||||
|
**Report format**: Summarize actual results from each probe.
|
||||||
|
""")
|
||||||
|
|
||||||
|
MISSING_CONTRACT = textwrap.dedent("""\
|
||||||
|
---
|
||||||
|
kind: function
|
||||||
|
name: missing
|
||||||
|
description: check-health contract with no report format
|
||||||
|
---
|
||||||
|
|
||||||
|
## Parameters
|
||||||
|
|
||||||
|
- x: y
|
||||||
|
|
||||||
|
## Returns
|
||||||
|
|
||||||
|
ok
|
||||||
|
|
||||||
|
### check-health
|
||||||
|
|
||||||
|
```bash
|
||||||
|
pwd -P
|
||||||
|
```
|
||||||
|
""")
|
||||||
|
|
||||||
|
|
||||||
|
def _run_lint(tmp_path, text, name):
|
||||||
|
(tmp_path / name).write_text(text)
|
||||||
|
return subprocess.run(["bash", str(LINT)], cwd=tmp_path,
|
||||||
|
capture_output=True, text=True)
|
||||||
|
|
||||||
|
|
||||||
|
def test_prose_lint_accepts_report_format_with_provenance(tmp_path):
|
||||||
|
result = _run_lint(tmp_path, GOOD_CONTRACT, "good.prose.md")
|
||||||
|
assert result.returncode == 0, result.stdout + result.stderr
|
||||||
|
|
||||||
|
|
||||||
|
def test_prose_lint_rejects_report_format_without_provenance(tmp_path):
|
||||||
|
result = _run_lint(tmp_path, DECOY_CONTRACT, "decoy.prose.md")
|
||||||
|
assert result.returncode == 1, result.stdout
|
||||||
|
assert "lacks execution provenance" in result.stdout
|
||||||
|
|
||||||
|
|
||||||
|
def test_prose_lint_requires_report_format_on_check_health_contract(tmp_path):
|
||||||
|
result = _run_lint(tmp_path, MISSING_CONTRACT, "missing.prose.md")
|
||||||
|
assert result.returncode == 1, result.stdout
|
||||||
|
assert "no **Report format** paragraph" in result.stdout
|
||||||
|
|
||||||
|
|
||||||
|
# ── contracts: parse the executable check-health probe block ─────────
|
||||||
|
|
||||||
|
def _check_health_block(contract):
|
||||||
|
"""Extract the bash probe block under ### check-health (the probe interface)."""
|
||||||
|
text = contract.read_text()
|
||||||
|
marker = "### check-health"
|
||||||
|
assert marker in text, f"{contract.name} has no {marker}"
|
||||||
|
after = text.split(marker, 1)[1]
|
||||||
|
match = re.search(r"```bash\n(.*?)```", after, re.S)
|
||||||
|
assert match, f"{contract.name} check-health has no bash probe block"
|
||||||
|
return match.group(1)
|
||||||
|
|
||||||
|
|
||||||
|
def _loop_nodes(block):
|
||||||
|
nodes = []
|
||||||
|
for line in block.splitlines():
|
||||||
|
match = re.match(r"\s*for\s+\w+\s+in\s+(.+?);?\s*do\b", line)
|
||||||
|
if match:
|
||||||
|
nodes = match.group(1).split()
|
||||||
|
return nodes
|
||||||
|
|
||||||
|
|
||||||
|
def _record(url):
|
||||||
|
"""Normalize a URL into a probe record: host, port, path, expected status."""
|
||||||
|
match = re.match(r"https?://([^/\s\"')]+)(/[^\s\"')]*)?", url)
|
||||||
|
assert match, f"unparseable probe URL: {url}"
|
||||||
|
hostport = match.group(1)
|
||||||
|
if "@" in hostport:
|
||||||
|
hostport = hostport.split("@", 1)[1]
|
||||||
|
if hostport.startswith("["):
|
||||||
|
host, port = hostport[1:hostport.index("]")], None
|
||||||
|
elif ":" in hostport:
|
||||||
|
host, raw_port = hostport.rsplit(":", 1)
|
||||||
|
port = int(raw_port) if raw_port.isdigit() else None
|
||||||
|
else:
|
||||||
|
host, port = hostport, None
|
||||||
|
return {"host": host, "port": port, "path": match.group(2) or "/",
|
||||||
|
"expected": None}
|
||||||
|
|
||||||
|
|
||||||
|
def _probes(block):
|
||||||
|
"""Parse the executable check-health bash block into a normalized probe model.
|
||||||
|
|
||||||
|
Comments are not probes; an `# Expected: <status>` comment annotates the
|
||||||
|
preceding probe. URLs using the block's shell-loop variable `$node` are
|
||||||
|
expanded over the loop's node list.
|
||||||
|
"""
|
||||||
|
loop_nodes = _loop_nodes(block)
|
||||||
|
probes = []
|
||||||
|
last = None
|
||||||
|
for raw in block.splitlines():
|
||||||
|
stripped = raw.strip()
|
||||||
|
if stripped.startswith("#"):
|
||||||
|
expected = re.search(r"Expected:\s*(\d{3})", stripped, re.I)
|
||||||
|
if expected and last is not None:
|
||||||
|
last["expected"] = int(expected.group(1))
|
||||||
|
continue
|
||||||
|
for url in re.findall(r"https?://[^\s\"')]+", raw):
|
||||||
|
hosts = loop_nodes if "$node" in url else [None]
|
||||||
|
for node in hosts:
|
||||||
|
record = _record(url.replace("$node", node) if node else url)
|
||||||
|
probes.append(record)
|
||||||
|
last = record
|
||||||
|
return probes
|
||||||
|
|
||||||
|
|
||||||
|
def test_gpu_monitor_probes_every_gpu_health_on_8080():
|
||||||
|
probes = _probes(_check_health_block(GPU))
|
||||||
|
targets = {(p["host"], p["port"], p["path"]) for p in probes}
|
||||||
|
assert ("192.168.68.8", 8080, "/health") in targets
|
||||||
|
assert ("192.168.68.110", 8080, "/health") in targets
|
||||||
|
|
||||||
|
|
||||||
|
def test_gpu_monitor_never_probes_bare_port_80_on_gpu_hosts():
|
||||||
|
probes = _probes(_check_health_block(GPU))
|
||||||
|
gpu_hosts = {"192.168.68.8", "192.168.68.110", "192.168.68.15"}
|
||||||
|
offenders = [p for p in probes
|
||||||
|
if p["host"] in gpu_hosts and p["port"] in (None, 80)]
|
||||||
|
assert offenders == []
|
||||||
|
|
||||||
|
|
||||||
|
def test_probe_model_flags_explicit_port_80_on_gpu_host():
|
||||||
|
# Regression: a bare-port probe may be spelled with an explicit :80.
|
||||||
|
block = ("curl -s -o /dev/null -w '%{http_code}' "
|
||||||
|
"http://192.168.68.8:80/health\n")
|
||||||
|
gpu_hosts = {"192.168.68.8", "192.168.68.110", "192.168.68.15"}
|
||||||
|
offenders = [p for p in _probes(block)
|
||||||
|
if p["host"] in gpu_hosts and p["port"] in (None, 80)]
|
||||||
|
assert offenders and offenders[0]["port"] == 80
|
||||||
|
|
||||||
|
|
||||||
|
def test_gpu_monitor_treats_router_301_as_alive():
|
||||||
|
probes = _probes(_check_health_block(GPU))
|
||||||
|
unified = [p for p in probes
|
||||||
|
if p["host"] == "192.168.68.116" and p["path"] == "/health/unified"]
|
||||||
|
assert unified, "router /health/unified probe missing"
|
||||||
|
assert unified[0]["expected"] == 301
|
||||||
|
|
||||||
|
|
||||||
|
def test_infra_monitoring_probes_every_real_pve_node():
|
||||||
|
probes = _probes(_check_health_block(INFRA))
|
||||||
|
pve = {(p["host"], p["port"], p["path"]) for p in probes if p["port"] == 8006}
|
||||||
|
assert {host for host, _, _ in pve} == PVE_NODE_IPS
|
||||||
|
assert {path for _, _, path in pve} == {"/api2/json/version"}
|
||||||
|
|
||||||
|
|
||||||
|
def test_infra_monitoring_does_not_probe_ct116_for_pve_api():
|
||||||
|
probes = _probes(_check_health_block(INFRA))
|
||||||
|
assert not any(p["host"] == "192.168.68.116" and p["port"] == 8006
|
||||||
|
for p in probes)
|
||||||
Executable
+211
@@ -0,0 +1,211 @@
|
|||||||
|
#!/bin/bash
|
||||||
|
# tests/zulip-monitor-abiba.sh — regression test pinning the producer→consumer
|
||||||
|
# contract between the pi Zulip extension's :9200/health payload and the Abiba
|
||||||
|
# leg of scripts/zulip-monitor.sh.
|
||||||
|
#
|
||||||
|
# WHY THIS TEST EXISTS: 2026-09-09 live incident. The monitor parsed the health
|
||||||
|
# payload at the WRONG nesting level (d.get('connected') at top level, while the
|
||||||
|
# extension serves zulip.connected) so PI_CONNECTED was always False and every
|
||||||
|
# monitor run restarted a healthy bot: pm2 showed restarts=8 with the process
|
||||||
|
# created 2026-09-09T09:35:09Z, the monitor log recorded four ❌ Abiba verdicts
|
||||||
|
# (04:23, 05:35, 06:55, 09:35 UTC) and zero ✅, while the Zulip server answered
|
||||||
|
# HTTP 200 and the bot logged a clean connect plus continuing heartbeats. The
|
||||||
|
# watchdog was the fault, not the connection. This test makes that class of
|
||||||
|
# regression fail loudly instead of silently restarting healthy services.
|
||||||
|
#
|
||||||
|
# CONTRACT UNDER TEST (must hold for scripts/zulip-monitor.sh):
|
||||||
|
# * Connection state is NESTED: zulip.connected (boolean) and zulip.last_error
|
||||||
|
# live inside the `zulip` object. There is NO top-level `connected` and NO
|
||||||
|
# retry counter anywhere in the payload (verified against the extension's
|
||||||
|
# startHealthServer handler) — the old retry_count branch was dropped.
|
||||||
|
# * zulip.connected=true -> log "✅ Connected", NO pm2 restart.
|
||||||
|
# * zulip.connected=false -> alert, pm2 restart abiba-zulip.
|
||||||
|
# * fetch error / non-2xx / empty body / unparseable body / missing or
|
||||||
|
# non-boolean zulip.connected -> "⚠️ Probe failed" alert with a
|
||||||
|
# "NOT restarting" label, NO pm2 restart. A parse miss must never kill a
|
||||||
|
# healthy service.
|
||||||
|
# * zulip.connected=true with last_error -> degraded 🟡 warning, no restart.
|
||||||
|
#
|
||||||
|
# HOW: the Abiba leg of the shipped script sits between the
|
||||||
|
# `# -- abiba-leg-start` / `# -- abiba-leg-end` marker comments. This runner
|
||||||
|
# extracts that block verbatim and executes it with a stubbed curl (fixture body
|
||||||
|
# + HTTP code), recorded notify()/pm2 shims, and a temp $LOG. If the markers
|
||||||
|
# disappear (fix reverted or renamed) extraction yields nothing and the suite
|
||||||
|
# fails — the bug cannot return silently.
|
||||||
|
#
|
||||||
|
# Usage: bash tests/zulip-monitor-abiba.sh [path/to/zulip-monitor.sh]
|
||||||
|
# Exit 0 iff every check passes.
|
||||||
|
#
|
||||||
|
# shellcheck disable=SC2034,SC2329,SC1090
|
||||||
|
# LOG/ISSUES and the notify/pm2/curl stubs below are consumed at runtime by
|
||||||
|
# the leg extracted between the marker comments and `source`d in each case;
|
||||||
|
# the static analyzer cannot see across that dynamic source, so it flags them.
|
||||||
|
set -uo pipefail
|
||||||
|
|
||||||
|
ROOT=$(cd "$(dirname "$0")/.." && pwd)
|
||||||
|
SCRIPT=${1:-"$ROOT/scripts/zulip-monitor.sh"}
|
||||||
|
FIXTURES="$ROOT/tests/fixtures"
|
||||||
|
TMP=$(mktemp -d)
|
||||||
|
trap 'rm -rf "$TMP"' EXIT
|
||||||
|
|
||||||
|
PASS=0
|
||||||
|
FAIL=0
|
||||||
|
ok() { PASS=$((PASS + 1)); printf ' \033[32m✔\033[0m %s\n' "$1"; }
|
||||||
|
bad() { FAIL=$((FAIL + 1)); printf ' \033[31m✘\033[0m %s\n' "$1"; }
|
||||||
|
|
||||||
|
echo "== tests/zulip-monitor-abiba.sh — Abiba leg vs :9200/health producer contract =="
|
||||||
|
echo "target script: $SCRIPT"
|
||||||
|
|
||||||
|
# --- structural guards -------------------------------------------------------
|
||||||
|
if ! grep -q '^# -- abiba-leg-start' "$SCRIPT"; then
|
||||||
|
echo "✘ FATAL: $SCRIPT has no '# -- abiba-leg-start' marker — the fix has been reverted or renamed."
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
if ! grep -q '^# -- abiba-leg-end' "$SCRIPT"; then
|
||||||
|
echo "✘ FATAL: $SCRIPT has no '# -- abiba-leg-end' marker."
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
|
||||||
|
LEG="$TMP/leg.sh"
|
||||||
|
awk '/^# -- abiba-leg-start/{f=1; next}
|
||||||
|
/^# -- abiba-leg-end/{f=0; next}
|
||||||
|
f' "$SCRIPT" > "$LEG"
|
||||||
|
if [ ! -s "$LEG" ]; then
|
||||||
|
echo "✘ FATAL: extracted Abiba leg is empty."
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
echo "== structural =="
|
||||||
|
if bash -n "$SCRIPT"; then ok "syntax: bash -n $SCRIPT"; else bad "syntax: bash -n $SCRIPT failed"; fi
|
||||||
|
if bash -n "$LEG"; then ok "syntax: extracted leg parses (bash -n)"; else bad "syntax: extracted leg fails bash -n"; fi
|
||||||
|
|
||||||
|
# --- per-case harness ---------------------------------------------------------
|
||||||
|
CURRENT_NAME=""
|
||||||
|
CURRENT_DIR=""
|
||||||
|
|
||||||
|
# $1 case name, $2 http-code, $3 body (file path or literal)
|
||||||
|
run_case() {
|
||||||
|
local name="$1" http="$2" body_src="$3" body
|
||||||
|
CURRENT_NAME="$name"
|
||||||
|
CURRENT_DIR=$(mktemp -d "$TMP/case.XXXXXX")
|
||||||
|
if [ -f "$body_src" ]; then
|
||||||
|
body=$(cat "$body_src")
|
||||||
|
else
|
||||||
|
body="$body_src"
|
||||||
|
fi
|
||||||
|
(
|
||||||
|
LOG="$CURRENT_DIR/log"; ISSUES=0
|
||||||
|
notify() { printf 'ALERT [%s] %s\n' "$1" "$2" >> "$CURRENT_DIR/alerts"; }
|
||||||
|
pm2() { printf 'PM2 %s\n' "$*" >> "$CURRENT_DIR/pm2"; }
|
||||||
|
curl() {
|
||||||
|
local url=""
|
||||||
|
for a in "$@"; do case "$a" in http*) url="$a";; esac; done
|
||||||
|
case "$url" in
|
||||||
|
*:9200/health*)
|
||||||
|
case " $* " in
|
||||||
|
*"-w"*) printf '%s' "$http" ;; # -w '%{http_code}' code probe
|
||||||
|
*) printf '%s' "$body" ;; # body probe
|
||||||
|
esac ;;
|
||||||
|
*)
|
||||||
|
printf 'UNEXPECTED-CURL %s\n' "$*" >> "$CURRENT_DIR/unexpected-curl"
|
||||||
|
return 7 ;;
|
||||||
|
esac
|
||||||
|
return 0
|
||||||
|
}
|
||||||
|
source "$LEG"
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
assert_log_has() {
|
||||||
|
if grep -qF -- "$1" "$CURRENT_DIR/log"; then ok "$CURRENT_NAME — log has: $1"; else bad "$CURRENT_NAME — log MISSING: $1"; fi
|
||||||
|
}
|
||||||
|
assert_log_lacks() {
|
||||||
|
if grep -qF -- "$1" "$CURRENT_DIR/log"; then bad "$CURRENT_NAME — log must NOT contain: $1"; else ok "$CURRENT_NAME — log correctly lacks: $1"; fi
|
||||||
|
}
|
||||||
|
assert_alert_has() {
|
||||||
|
if grep -qF -- "$1" "$CURRENT_DIR/alerts"; then ok "$CURRENT_NAME — alert sent: $1"; else bad "$CURRENT_NAME — alert MISSING: $1"; fi
|
||||||
|
}
|
||||||
|
assert_alert_empty() {
|
||||||
|
if [ ! -s "$CURRENT_DIR/alerts" ]; then ok "$CURRENT_NAME — no alert sent (quiet healthy path)"; else bad "$CURRENT_NAME — unexpected alert: $(cat "$CURRENT_DIR/alerts")"; fi
|
||||||
|
}
|
||||||
|
assert_pm2_restarted() {
|
||||||
|
if grep -qF "PM2 restart abiba-zulip" "$CURRENT_DIR/pm2"; then ok "$CURRENT_NAME — pm2 restart abiba-zulip was called"; else bad "$CURRENT_NAME — expected pm2 restart abiba-zulip, pm2 log: $(cat "$CURRENT_DIR/pm2" 2>/dev/null)"; fi
|
||||||
|
}
|
||||||
|
assert_no_restart() {
|
||||||
|
if [ ! -s "$CURRENT_DIR/pm2" ]; then ok "$CURRENT_NAME — NO pm2 restart (fail-safe holds)"; else bad "$CURRENT_NAME — pm2 was called but must NOT be: $(cat "$CURRENT_DIR/pm2")"; fi
|
||||||
|
}
|
||||||
|
assert_no_unexpected_curl() {
|
||||||
|
if [ ! -s "$CURRENT_DIR/unexpected-curl" ]; then ok "$CURRENT_NAME — only :9200/health was probed"; else bad "$CURRENT_NAME — unexpected curl: $(cat "$CURRENT_DIR/unexpected-curl")"; fi
|
||||||
|
}
|
||||||
|
|
||||||
|
# --- case 1: real payload shape, zulip.connected=true -> healthy, no restart --
|
||||||
|
echo "== case 1: connected (real producer payload: nested zulip.connected=true) =="
|
||||||
|
run_case "connected" 200 "$FIXTURES/zulip-health-connected.json"
|
||||||
|
assert_log_has "Abiba: ✅ Connected (processed=0)"
|
||||||
|
assert_log_lacks "Disconnected"
|
||||||
|
assert_alert_empty
|
||||||
|
assert_no_restart
|
||||||
|
assert_no_unexpected_curl
|
||||||
|
|
||||||
|
# --- case 2: zulip.connected=false -> disconnected, restart -------------------
|
||||||
|
echo "== case 2: disconnected (nested zulip.connected=false triggers restart) =="
|
||||||
|
run_case "disconnected" 200 "$FIXTURES/zulip-health-disconnected.json"
|
||||||
|
assert_log_has "Abiba: ❌ Disconnected — restarted"
|
||||||
|
assert_alert_has "DISCONNECTED — restarting"
|
||||||
|
assert_pm2_restarted
|
||||||
|
assert_no_unexpected_curl
|
||||||
|
|
||||||
|
# --- cases 3-9: probe failures must alert and MUST NOT restart ----------------
|
||||||
|
echo "== probe-failure cases: alert 'NOT restarting', zero pm2 restarts =="
|
||||||
|
|
||||||
|
run_case "empty body" 200 ""
|
||||||
|
assert_log_has "Abiba: ⚠️ Probe failed"
|
||||||
|
assert_log_lacks "❌ Disconnected"
|
||||||
|
assert_alert_has "NOT restarting"
|
||||||
|
assert_no_restart
|
||||||
|
|
||||||
|
run_case "garbage body" 200 '{not valid json!!'
|
||||||
|
assert_log_has "Abiba: ⚠️ Probe failed"
|
||||||
|
assert_alert_has "NOT restarting"
|
||||||
|
assert_no_restart
|
||||||
|
|
||||||
|
run_case "missing zulip key" 200 '{"status":"ok","platform":"pi","agent":"abiba"}'
|
||||||
|
assert_log_has "Probe failed"
|
||||||
|
assert_alert_has "NOT restarting"
|
||||||
|
assert_no_restart
|
||||||
|
|
||||||
|
run_case "zulip without connected" 200 '{"status":"ok","zulip":{"last_error":null}}'
|
||||||
|
assert_log_has "Probe failed"
|
||||||
|
assert_alert_has "NOT restarting"
|
||||||
|
assert_no_restart
|
||||||
|
|
||||||
|
run_case "non-boolean connected" 200 '{"status":"ok","zulip":{"connected":"true"}}'
|
||||||
|
assert_log_has "Probe failed"
|
||||||
|
assert_alert_has "NOT restarting"
|
||||||
|
assert_no_restart
|
||||||
|
|
||||||
|
run_case "fetch failure http 000" 000 ""
|
||||||
|
assert_log_has "Probe failed"
|
||||||
|
assert_alert_has "NOT restarting"
|
||||||
|
assert_no_restart
|
||||||
|
|
||||||
|
run_case "non-2xx http 500" 500 '{"error":"boom"}'
|
||||||
|
assert_log_has "Probe failed"
|
||||||
|
assert_alert_has "NOT restarting"
|
||||||
|
assert_no_restart
|
||||||
|
|
||||||
|
# --- case 10: connected but last_error set -> degraded 🟡, no restart ---------
|
||||||
|
echo "== case 10: degraded (connected=true but last_error set) warns, no restart =="
|
||||||
|
run_case "degraded" 200 '{"status":"ok","zulip":{"connected":true,"last_error":"transient queue hiccup","messages_processed":3}}'
|
||||||
|
assert_log_has "Abiba: 🟡 Error: transient queue hiccup"
|
||||||
|
assert_log_lacks "❌ Disconnected"
|
||||||
|
assert_no_restart
|
||||||
|
|
||||||
|
# --- summary -------------------------------------------------------------------
|
||||||
|
echo ""
|
||||||
|
if [ "$FAIL" -eq 0 ]; then
|
||||||
|
echo "✅ ALL CHECKS PASSED ($PASS/$PASS) — tests/zulip-monitor-abiba.sh"
|
||||||
|
exit 0
|
||||||
|
else
|
||||||
|
echo "❌ $FAIL CHECK(S) FAILED ($PASS passed) — tests/zulip-monitor-abiba.sh"
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
+245
-43
@@ -1,22 +1,34 @@
|
|||||||
---
|
---
|
||||||
kind: responsibility
|
kind: responsibility
|
||||||
name: zulip-health
|
name: zulip-health
|
||||||
description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (Agent Zero Docker), Platform B (Hermes agents Tanko/Mumuni), and the Zulip bridge. Verifies bot registration, DM delivery, and cross-platform connectivity.
|
description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (pi/Abiba Zulip bridge), Platform B (Tanko on DSH), and Platform C (Agent Zero Docker). Verifies bot registration, DM delivery, and cross-platform connectivity. Mumuni is no longer monitored from this host — she runs on her own container (kagentz CT 105 on minipve, .14) and is monitored on her side.
|
||||||
title: Zulip Mesh Health Monitor — Multi-Platform
|
title: Zulip Mesh Health Monitor — Multi-Platform
|
||||||
version: 3.0.0
|
version: 3.2.0
|
||||||
runtime_contract: 2
|
runtime_contract: 2
|
||||||
agent: abiba
|
agent: abiba
|
||||||
|
report_only_agents:
|
||||||
|
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
|
||||||
---
|
---
|
||||||
|
|
||||||
# Zulip Mesh Health Monitor
|
# Zulip Mesh Health Monitor
|
||||||
|
|
||||||
Monitors ALL Zulip-connected agents across three platforms (pi, Hermes, Agent Zero).
|
Monitors the Zulip-connected agents under this host's operational control (pi,
|
||||||
Runs every 15 minutes in the background. Also triggers on session start.
|
DSH, Agent Zero). Runs every 15 minutes in the background. Also triggers on
|
||||||
|
session start.
|
||||||
|
|
||||||
|
> **Mumuni is NOT monitored from this host (captain ruling 2026-09-10).** She
|
||||||
|
> moved off this host onto her own container — kagentz CT 105 on minipve
|
||||||
|
> (192.168.68.14), running a dedicated `hermes` user — and is monitored on her
|
||||||
|
> side. No step in this contract, and no leg of `scripts/zulip-monitor.sh`, may
|
||||||
|
> ssh to her old deployment, read her `~/.hermes/gateway_state.json`, or alert on
|
||||||
|
> her state. The former Platform-B-for-Mumuni steps (gateway process, heartbeat,
|
||||||
|
> response delivery) are retired: they always read "unknown" against the
|
||||||
|
> decommissioned deployment and produced a false 🔴 alert on every run.
|
||||||
|
|
||||||
## Requires
|
## Requires
|
||||||
|
|
||||||
- **Zulip API key** for `abiba-bot@chat.sysloggh.net` in `$ZULIP_API_KEY`
|
- **Zulip API key** for `abiba-bot@chat.sysloggh.net` in `$ZULIP_API_KEY`
|
||||||
- **SSH access** to Tanko (192.168.68.122), Mumuni (192.168.68.24, inside Abiba CT100 on minipve), and Agent Zero Docker host (192.168.68.14)
|
- **SSH access** to amdpve (192.168.68.15) for Tanko — CT 112 reached via `pct exec` (direct SSH to .122 is not a dependency of this contract: per-worker key availability varies); and the Agent Zero Docker host (192.168.68.14)
|
||||||
- **PM2** on localhost for pi process management
|
- **PM2** on localhost for pi process management
|
||||||
- **Network access** to `chat.sysloggh.net`, `localhost:9200`
|
- **Network access** to `chat.sysloggh.net`, `localhost:9200`
|
||||||
- **Write access** to `/root/zulip-health-monitor.log` and `/tmp/zulip-monitor-debounce`
|
- **Write access** to `/root/zulip-health-monitor.log` and `/tmp/zulip-monitor-debounce`
|
||||||
@@ -45,11 +57,9 @@ Runs every 15 minutes in the background. Also triggers on session start.
|
|||||||
"severity": "healthy"
|
"severity": "healthy"
|
||||||
},
|
},
|
||||||
"tanko": {
|
"tanko": {
|
||||||
"platform": "hermes",
|
"platform": "dsh",
|
||||||
"zulip_state": "connected",
|
"service_state": "active",
|
||||||
"heartbeat_age_seconds": 45,
|
"http_status": 200,
|
||||||
"gateway_pid": 1234,
|
|
||||||
"edit_fail_rate_pct": 0,
|
|
||||||
"severity": "healthy"
|
"severity": "healthy"
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -81,13 +91,13 @@ Log as "unreachable" — don't treat as critical unless it persists for 3+ conse
|
|||||||
## Streaming Support (2026-07-05)
|
## Streaming Support (2026-07-05)
|
||||||
|
|
||||||
Zulip agents now support progressive message editing during agent generation.
|
Zulip agents now support progressive message editing during agent generation.
|
||||||
When a Hermes agent (Tanko, Mumuni) processes a message, the response is
|
When a Zulip agent under this monitor's scope (Tanko on DSH) processes a message, the response is
|
||||||
streamed in real-time via Zulip's `PATCH /api/v1/messages/{id}` API:
|
streamed in real-time via Zulip's `PATCH /api/v1/messages/{id}` API:
|
||||||
|
|
||||||
- Adapter implements `edit_message()` using `_api_patch()` helper
|
- Adapter implements `edit_message()` using `_api_patch()` helper
|
||||||
- Gateway stream consumer progressively edits the Zulip message
|
- Gateway stream consumer progressively edits the Zulip message
|
||||||
- User sees real-time agent thinking instead of waiting for full response
|
- User sees real-time agent thinking instead of waiting for full response
|
||||||
- Verified: Tanko (CT 112) and Mumuni (inside Abiba CT 100) both have streaming active
|
- Verified: Tanko (CT 112) has streaming active; Mumuni's (kagentz CT 105) is verified on her own host, not from here
|
||||||
|
|
||||||
### Verification
|
### Verification
|
||||||
```bash
|
```bash
|
||||||
@@ -182,67 +192,257 @@ grep -a "Finalized\|Failed to finalize" /root/.pm2/logs/abiba-zulip-out.log | ta
|
|||||||
| `last_error` set | Log and monitor |
|
| `last_error` set | Log and monitor |
|
||||||
| Crash loop >10/h | Alert user |
|
| Crash loop >10/h | Alert user |
|
||||||
|
|
||||||
### Step 3: Platform B — Hermes (Tanko .122, Mumuni .24)
|
|
||||||
|
|
||||||
**B1: Gateway State**
|
### Step 3: Platform B — Tanko (DSH on amdpve CT 112)
|
||||||
|
|
||||||
|
Mumuni is out of scope for this host (see the note above): she runs on her own
|
||||||
|
container and is monitored on her side.
|
||||||
|
|
||||||
|
Tanko runs on DSH (DeepSeek Harness) — it no longer runs a Hermes gateway, so
|
||||||
|
there is no `~/.hermes/gateway_state.json` on CT 112. Tanko's Zulip gateway runs
|
||||||
|
as the `dsh-web` systemd unit inside **CT 112**, which resides on the **amdpve**
|
||||||
|
PVE host (**192.168.68.15**). Direct SSH to 192.168.68.122 is not a dependency
|
||||||
|
of this contract — per-worker key availability varies — so CT 112 probes run
|
||||||
|
from the amdpve vantage via `pct exec`:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
ssh root@192.168.68.122 "cat ~/.hermes/gateway_state.json"
|
ssh root@192.168.68.15 "pct exec 112 -- <command>"
|
||||||
ssh root@192.168.68.24 "cat ~/.hermes/gateway_state.json" # Mumuni inside Abiba CT100
|
|
||||||
```
|
```
|
||||||
|
|
||||||
Check `platforms.zulip.state`: `connected` ✅ | `disconnected` ❌ | `error` ❌ | missing → not installed.
|
> **By design (verified 2026-09-08):** the `dsh-web` gateway binds
|
||||||
|
> `127.0.0.1:3080` **loopback-only**. A remote probe against
|
||||||
|
> `192.168.68.122:3080` gets connection-refused — that is EXPECTED, NOT a fault,
|
||||||
|
> and must never be raised as Tanko down. Only loopback probes from inside
|
||||||
|
> CT 112 (or the public-URL fallback below) are valid health signals.
|
||||||
|
|
||||||
**B2: Agent Process**
|
**B1: Gateway Service State (Tanko)**
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
ssh root@<CT> "ps aux | grep 'gateway run' | grep -v grep"
|
ssh root@192.168.68.15 "pct exec 112 -- systemctl is-active dsh-web"
|
||||||
```
|
```
|
||||||
|
|
||||||
Gateway PID should exist with uptime > 60s. **Dual-gateway detection**: if more than one `gateway run` process is found, the gateway has a collision (typically one `--force` and one `--replace` process). Kill the newer/duplicate process, then restart the remaining gateway per-agent (parameterized 2026-08-09, captain ruling):
|
Expected: `active`. Anything else → gateway service down → apply the Tanko heal
|
||||||
|
(restart via DSH service, Platform B Actions table below).
|
||||||
|
|
||||||
| Agent | Restart command | Notes |
|
**B2: Gateway HTTP Liveness (Tanko — loopback-only :3080)**
|
||||||
|-------|-----------------|-------|
|
|
||||||
| Mumuni (.24) | `pm2 restart abiba-zulip` | Hermes gateway runs under PM2 as `abiba-zulip` |
|
|
||||||
| Tanko (.122) | `bash /opt/hermes-zulip-plugin/run.sh` (or the agent's systemd/user unit) | Tanko does NOT use PM2 — never run `pm2 restart mumuni-zulip` for Tanko (process does not exist) |
|
|
||||||
|
|
||||||
Check gateway log for "Gateway running with 2 platform(s)" (not 1) to confirm Zulip reloaded.
|
|
||||||
|
|
||||||
**B3: Heartbeat Verification**
|
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
ssh root@<CT> "grep Heartbeat ~/.hermes/logs/agent.log | tail -3"
|
ssh root@192.168.68.15 "pct exec 112 -- curl -s --connect-timeout 5 --max-time 10 -o /dev/null -w '%{http_code}' http://127.0.0.1:3080/"
|
||||||
```
|
```
|
||||||
|
|
||||||
Expected: recent heartbeat (within 5 min), `polls=N` incrementing.
|
Alive = **ANY** HTTP status response from the endpoint — the expected set is
|
||||||
Silence > 300s → warning. Silence > 600s → critical.
|
`200`/`301`/`302`/`307`/`308`/`401`/`403` (the gateway UI is token-gated and
|
||||||
|
legitimately answers with redirects/auth-challenges, so never require a bare
|
||||||
|
`200`), and any other status, including `404`/`5xx`, also counts alive: a
|
||||||
|
process answering `503` is running and self-heal must NOT restart-loop it.
|
||||||
|
Down = connection refused (`000`) or timeout only. Statuses outside the
|
||||||
|
expected set are logged/reported as a warning — reported, never healed on.
|
||||||
|
|
||||||
**B4: Response Delivery**
|
**B3: Public-URL Fallback Probe (Tanko — for nodes without pct/ssh access to amdpve)**
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
ssh root@<CT> "grep -E 'Finalized|Failed to finalize|Replied to' ~/.hermes/logs/agent.log | tail -10"
|
curl -s --connect-timeout 10 --max-time 15 -o /dev/null -w '%{http_code}' https://tankodhs.sysloggh.net/
|
||||||
```
|
```
|
||||||
|
|
||||||
> 50% fail rate → critical.
|
Fallback only — used when the monitoring node has no pct/SSH path to amdpve.
|
||||||
|
Alive = **ANY** HTTP status response from the endpoint — healthy signals are
|
||||||
|
`302` (authentik proxy-auth redirect) and `401` (auth-gated), and any other
|
||||||
|
status, including `404`/`5xx`, also counts alive: the endpoint is up and
|
||||||
|
answering and must NOT be restart-looped. Down = connection refused (`000`) or
|
||||||
|
timeout only. Never expect a bare `200` — the public URL terminates in the
|
||||||
|
token-gated authentik chain. Statuses outside the healthy set are
|
||||||
|
logged/reported as a warning — reported, never healed on.
|
||||||
|
|
||||||
**Platform B Actions**
|
**Platform B Actions**
|
||||||
|
|
||||||
| Condition | Action |
|
| Condition | Action |
|
||||||
|-----------|--------|
|
|-----------|--------|
|
||||||
| `zulip.state != "connected"` | `ssh root@<CT> "pkill -f 'gateway run'; sleep 2; hermes gateway restart"` |
|
| `dsh-web` service not `active` | Restart Tanko via DSH service |
|
||||||
| No heartbeat in 10min | Same as above |
|
| HTTP `:3080` connection refused/timeout (`000`) | Same as above |
|
||||||
| `Failed to finalize` > 50% | Check PATCH API, Zulip server |
|
| HTTP status outside the expected set | Log/report as a warning — reported, never healed on |
|
||||||
| Response empty/short | Check A2A endpoint / LiteLLM model |
|
|
||||||
|
**B4: dsh-web Authentication (Tanko — restart-persistent login)**
|
||||||
|
|
||||||
|
The dsh-web UI is token-gated. On every start the process prints a random
|
||||||
|
launch token to the journal:
|
||||||
|
|
||||||
|
```
|
||||||
|
dsh web: http://127.0.0.1:3080/?token=<TOKEN>
|
||||||
|
```
|
||||||
|
|
||||||
|
The token only bootstraps an authority-bound, HMAC-signed browser cookie with a
|
||||||
|
30-day lifetime. The signing secret is durable in
|
||||||
|
`/root/.dsh/.credentials.yaml` (key `client-connection/browser-session`), so a
|
||||||
|
cookie minted once keeps working across `dsh-web` restarts; the launch token
|
||||||
|
itself rotates on every restart.
|
||||||
|
|
||||||
|
**Login endpoint (public, Authentik-gated):**
|
||||||
|
`https://tankodhs.sysloggh.net/dsh-web-login`
|
||||||
|
|
||||||
|
It lives inside the Authentik-gated `:80` server block
|
||||||
|
(`/etc/nginx/sites-available/dsh`, symlinked from
|
||||||
|
`/etc/nginx/sites-enabled/dsh`) as `location = /dsh-web-login`, guarded by
|
||||||
|
`auth_request /outpost.goauthentik.io/auth/nginx`. It proxies to dsh-web with
|
||||||
|
`Host: tankodhs.sysloggh.net`, so the minted cookie is bound to the public
|
||||||
|
authority — never to `127.0.0.1:3080`. The token-dependent line is isolated in
|
||||||
|
the generated include `/etc/dsh-web/nginx-login.conf`:
|
||||||
|
|
||||||
|
```
|
||||||
|
proxy_pass http://127.0.0.1:3080/?token=<TOKEN>;
|
||||||
|
```
|
||||||
|
|
||||||
|
**Token refresh (non-disruptive):**
|
||||||
|
`/opt/deepseek-harness/capture-dsh-token.sh` (source:
|
||||||
|
`scripts/capture-dsh-token.sh`) reads candidate launch tokens from the journal
|
||||||
|
**scoped to the service's current systemd invocation**
|
||||||
|
(`systemctl show -p InvocationID` + `_SYSTEMD_INVOCATION_ID=`), re-sampling the
|
||||||
|
invocation on every pass so a restart that lands during the wait switches to the
|
||||||
|
new invocation; a restarted process's stale token is never considered while its
|
||||||
|
new startup banner is still pending and there is no whole-journal or
|
||||||
|
cross-invocation fallback. Each candidate
|
||||||
|
is then functionally verified against dsh-web with `Host: tankodhs.sysloggh.net`,
|
||||||
|
using the first the running process accepts with `303`. It waits up to 120s for
|
||||||
|
a restarted process to accept a token and re-probes every current-invocation
|
||||||
|
candidate on each pass, so a token that briefly returns `000` while the service
|
||||||
|
is still starting is not disqualified. If none is accepted it leaves the include
|
||||||
|
untouched and exits so the timer retries (exiting non-zero when a pending reload
|
||||||
|
is still outstanding). It writes
|
||||||
|
`/etc/dsh-web/launch-token` and regenerates `/etc/dsh-web/nginx-login.conf`,
|
||||||
|
reloading nginx only when the on-disk include differs from the generated one or
|
||||||
|
the applied-state stamp does not match the token (`nginx -t` guards the reload,
|
||||||
|
and the stamp is written only after a successful `nginx -s reload`, so a failed
|
||||||
|
or interrupted reload is retried on the next run). Any failed reload records a
|
||||||
|
pending-reload marker under `/etc/dsh-web/`; the next run attempts the reload
|
||||||
|
before the token wait, independent of token state, and clears the marker only
|
||||||
|
once the reload succeeds, so a disabled legacy `:8081` file can never leave the
|
||||||
|
running nginx unreloaded. The generated include is recreated before any
|
||||||
|
`nginx -t` if it is missing, so a failed run cannot wedge recovery.
|
||||||
|
Runs are serialized with `flock` on `/run/capture-dsh-token.lock`. It **never
|
||||||
|
stops or starts `dsh-web`**.
|
||||||
|
It is triggered by the `dsh-web.service` drop-in
|
||||||
|
`/etc/systemd/system/dsh-web.service.d/20-token-refresh.conf`
|
||||||
|
(`ExecStartPost=/bin/systemctl --no-block start dsh-web-token.service`) and by
|
||||||
|
`dsh-web-token.timer` every 2 minutes for reconciliation.
|
||||||
|
|
||||||
|
<details><summary>Installed systemd wiring (CT 112)</summary>
|
||||||
|
|
||||||
|
```ini
|
||||||
|
# /etc/systemd/system/dsh-web-token.service
|
||||||
|
[Unit]
|
||||||
|
Description=Refresh the dsh-web launch token for the nginx login endpoint
|
||||||
|
After=dsh-web.service
|
||||||
|
[Service]
|
||||||
|
Type=oneshot
|
||||||
|
TimeoutStartSec=180
|
||||||
|
ExecStart=/opt/deepseek-harness/capture-dsh-token.sh
|
||||||
|
|
||||||
|
# /etc/systemd/system/dsh-web-token.timer
|
||||||
|
[Unit]
|
||||||
|
Description=Periodically refresh the dsh-web login token
|
||||||
|
[Timer]
|
||||||
|
OnBootSec=90s
|
||||||
|
OnUnitActiveSec=120s
|
||||||
|
AccuracySec=10s
|
||||||
|
Persistent=true
|
||||||
|
[Install]
|
||||||
|
WantedBy=timers.target
|
||||||
|
|
||||||
|
# /etc/systemd/system/dsh-web.service.d/20-token-refresh.conf
|
||||||
|
[Service]
|
||||||
|
ExecStartPost=/bin/systemctl --no-block start dsh-web-token.service
|
||||||
|
```
|
||||||
|
|
||||||
|
</details>
|
||||||
|
|
||||||
|
> **Do NOT reintroduce the `:8081` endpoint.** It listened on `0.0.0.0:8081`
|
||||||
|
> with no `auth_request` and was a full Authentik bypass for anyone on the LAN.
|
||||||
|
> The script now removes `/etc/nginx/sites-enabled/dsh.token` automatically if
|
||||||
|
> it ever reappears.
|
||||||
|
|
||||||
|
**Authentication flow:**
|
||||||
|
1. `GET https://tankodhs.sysloggh.net/dsh-web-login`
|
||||||
|
2. Unauthenticated → Authentik sign-in; once authenticated the request reaches
|
||||||
|
dsh-web with `Host: tankodhs.sysloggh.net`.
|
||||||
|
3. dsh-web accepts the launch token on `GET /`, writes the
|
||||||
|
`dsh-auth-<authority-hash>` cookie (30 days, `HttpOnly`, `SameSite=Strict`)
|
||||||
|
and returns `303` to `/`.
|
||||||
|
4. Every later request through `/` presents that cookie; the token is not needed
|
||||||
|
again until the cookie expires or a new browser is used.
|
||||||
|
|
||||||
|
**Verification** (amdpve vantage):
|
||||||
|
```bash
|
||||||
|
# 1. Login endpoint is Authentik-gated: unauthenticated -> 302 (not 200/303).
|
||||||
|
ssh root@192.168.68.15 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}\n' \
|
||||||
|
-H 'Host: tankodhs.sysloggh.net' http://127.0.0.1/dsh-web-login"
|
||||||
|
# Expected: 302
|
||||||
|
|
||||||
|
# 2. Legacy :8081 endpoint is gone (connection refused -> 000).
|
||||||
|
ssh root@192.168.68.15 "pct exec 112 -- curl -s --max-time 3 -o /dev/null \
|
||||||
|
-w '%{http_code}\n' http://192.168.68.122:8081/"
|
||||||
|
# Expected: 000
|
||||||
|
|
||||||
|
# 3. Backend cookie mint + reuse (exactly what /dsh-web-login proxies to).
|
||||||
|
TOKEN=$(ssh root@192.168.68.15 "pct exec 112 -- cat /etc/dsh-web/launch-token")
|
||||||
|
ssh root@192.168.68.15 "pct exec 112 -- curl -s -c /tmp/dsh.jar -o /dev/null \
|
||||||
|
-H 'Host: tankodhs.sysloggh.net' 'http://127.0.0.1:3080/?token=$TOKEN'"
|
||||||
|
ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh.jar -o /dev/null \
|
||||||
|
-w '%{http_code}\n' -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/"
|
||||||
|
# Expected: 200 — the minted dsh-auth-... cookie (authority
|
||||||
|
# tankodhs.sysloggh.net) is replayed on the next request and accepted.
|
||||||
|
|
||||||
|
# 4. Token refresh is non-disruptive and idempotent.
|
||||||
|
ssh root@192.168.68.15 "pct exec 112 -- /opt/deepseek-harness/capture-dsh-token.sh"
|
||||||
|
# Expected: "token unchanged; nginx not reloaded" when nothing changed
|
||||||
|
```
|
||||||
|
|
||||||
|
**Restart durability (acceptance):** after `systemctl restart dsh-web`, (a) the
|
||||||
|
cookie minted before the restart still returns `200` on `/`, and (b) the
|
||||||
|
refreshed `/etc/dsh-web/nginx-login.conf` carries the new token and mints a
|
||||||
|
fresh cookie. Both verified live 2026-09-11.
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# 5. Cookie survives a dsh-web restart, and the new token mints a new cookie.
|
||||||
|
ssh root@192.168.68.15 "pct exec 112 -- systemctl restart dsh-web"
|
||||||
|
# dsh-web is Type=simple: restart returns before :3080 is listening. Bounded-poll
|
||||||
|
# until the socket answers (any status but 000) before asserting the cookie.
|
||||||
|
for i in $(seq 1 60); do
|
||||||
|
UP=$(ssh root@192.168.68.15 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}' \
|
||||||
|
-H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/")
|
||||||
|
[ "$UP" != "000" ] && break
|
||||||
|
sleep 2
|
||||||
|
done
|
||||||
|
ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh.jar -o /dev/null \
|
||||||
|
-w '%{http_code}\n' -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/"
|
||||||
|
# Expected: 200 — the pre-restart cookie is still accepted.
|
||||||
|
# The restart's ExecStartPost (or the 2-minute timer) refreshes the include. A
|
||||||
|
# manual run may no-op on the flock, so poll until the include carries a token
|
||||||
|
# the running process accepts (bounded wait) before the mint+reuse check.
|
||||||
|
for i in $(seq 1 60); do
|
||||||
|
TOKEN=$(ssh root@192.168.68.15 "pct exec 112 -- sed -n 's/.*token=//p' /etc/dsh-web/nginx-login.conf | tr -d ';\n'")
|
||||||
|
CODE=$(ssh root@192.168.68.15 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}' \
|
||||||
|
-H 'Host: tankodhs.sysloggh.net' 'http://127.0.0.1:3080/?token=$TOKEN'")
|
||||||
|
[ "$CODE" = "303" ] && break
|
||||||
|
sleep 2
|
||||||
|
done
|
||||||
|
# Expected: 303 — the include now holds the token the running process accepts.
|
||||||
|
ssh root@192.168.68.15 "pct exec 112 -- curl -s -c /tmp/dsh-new.jar -o /dev/null \
|
||||||
|
-H 'Host: tankodhs.sysloggh.net' 'http://127.0.0.1:3080/?token=$TOKEN'"
|
||||||
|
ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh-new.jar -o /dev/null \
|
||||||
|
-w '%{http_code}\n' -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/"
|
||||||
|
# Expected: 200 — the refreshed token minted a fresh cookie.
|
||||||
|
```
|
||||||
|
|
||||||
|
|
||||||
### Step 4: Platform C — Agent Zero (kagentz, CT 105 via Docker host .14)
|
### Step 4: Platform C — Agent Zero (kagentz, CT 105 via Docker host .14)
|
||||||
|
|
||||||
**C1: A2A Server Health**
|
**C1: A2A Server Health**
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
ssh root@192.168.68.14 "docker exec agent-zero curl -s --connect-timeout 5 http://127.0.0.1:8001/.well-known/agent.json"
|
# A2A server is on :50080 (not :8001) and is auth-gated (401 expected for unauthenticated)
|
||||||
|
ssh root@192.168.68.14 "curl -s --connect-timeout 5 -o /dev/null -w '%{http_code}' http://127.0.0.1:50080/a2a/"
|
||||||
```
|
```
|
||||||
|
|
||||||
Expected: `{"name":"kagentz",...}`. Connection refused → A2A server down.
|
Expected: `401` (auth-gated, A2A server is up and responding) or `200` (if no auth required). Connection refused (000) → A2A server down.
|
||||||
|
|
||||||
**C2: Adapter Process**
|
**C2: Adapter Process**
|
||||||
|
|
||||||
@@ -263,12 +463,14 @@ Check: `processed=N` incrementing, `silence < 600s`, `reconnects` ≈ 0.
|
|||||||
**C4: A2A Response Verification**
|
**C4: A2A Response Verification**
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
ssh root@192.168.68.14 "docker exec agent-zero curl -s -X POST http://127.0.0.1:8001/a2a \
|
# A2A server is on :50080 (not :8001) and is auth-gated (401 expected for unauthenticated)
|
||||||
|
ssh root@192.168.68.14 "curl -s -X POST http://127.0.0.1:50080/a2a \
|
||||||
-H 'Content-Type: application/json' \
|
-H 'Content-Type: application/json' \
|
||||||
|
-H 'Authorization: Bearer $LITELLM_KEY' \
|
||||||
-d '{\"jsonrpc\":\"2.0\",\"method\":\"tasks/send\",\"params\":{\"message\":{\"role\":\"user\",\"parts\":[{\"text\":\"ping\"}]}},\"id\":1}'"
|
-d '{\"jsonrpc\":\"2.0\",\"method\":\"tasks/send\",\"params\":{\"message\":{\"role\":\"user\",\"parts\":[{\"text\":\"ping\"}]}},\"id\":1}'"
|
||||||
```
|
```
|
||||||
|
|
||||||
Expected: task ID with "working" status. Poll for completion with `tasks/get`.
|
Expected: task ID with "working" status. Poll for completion with `tasks/get`. If 401, check LITELLM_KEY is set.
|
||||||
|
|
||||||
**Platform C Actions**
|
**Platform C Actions**
|
||||||
|
|
||||||
@@ -285,7 +487,7 @@ Expected: task ID with "working" status. Poll for completion with `tasks/get`.
|
|||||||
|
|
||||||
Check each agent's log for excessive bot-to-bot chatter:
|
Check each agent's log for excessive bot-to-bot chatter:
|
||||||
- Abiba: `Skipped.*bot msgs` count
|
- Abiba: `Skipped.*bot msgs` count
|
||||||
- Tanko/Mumuni: Repeated DM exchanges between bots
|
- Tanko: Repeated DM exchanges between bots
|
||||||
- kagentz: Adapter log for bot DMs being processed
|
- kagentz: Adapter log for bot DMs being processed
|
||||||
|
|
||||||
If any bot processes >50 bot-originated messages in 15min → warning.
|
If any bot processes >50 bot-originated messages in 15min → warning.
|
||||||
|
|||||||
@@ -10,7 +10,7 @@ description: >
|
|||||||
|
|
||||||
> **⚠️ RETIRED** — The pi Zulip extension (`~/.pi/agent/extensions/zulip/`) and
|
> **⚠️ RETIRED** — The pi Zulip extension (`~/.pi/agent/extensions/zulip/`) and
|
||||||
> PM2 process (`abiba-zulip`) have been decommissioned. All mention/reliability
|
> PM2 process (`abiba-zulip`) have been decommissioned. All mention/reliability
|
||||||
> monitoring now happens through Telegram. Hermes agents (Tanko, Mumuni) and
|
> monitoring now happens through Telegram. Agents (Mumuni on Hermes, Tanko on DSH) and
|
||||||
> Agent Zero (kagentz) continue to use Zulip.
|
> Agent Zero (kagentz) continue to use Zulip.
|
||||||
|
|
||||||
## Maintains
|
## Maintains
|
||||||
|
|||||||
@@ -420,7 +420,7 @@ Backup v2 before starting: `cp index.js index.js.v2-backup-$(date +%Y%m%d-%H%M%S
|
|||||||
|-------|----------|-------------|--------------|-------------|
|
|-------|----------|-------------|--------------|-------------|
|
||||||
| **Abiba** | pi (CT 100) | ✅ Connected | API key missing from Infisical injection; poll timeout noise | Added .env fallback; AbortError treated as empty poll (no retry); poll timeout 65s→90s |
|
| **Abiba** | pi (CT 100) | ✅ Connected | API key missing from Infisical injection; poll timeout noise | Added .env fallback; AbortError treated as empty poll (no retry); poll timeout 65s→90s |
|
||||||
| **Tanko** | Hermes (CT 112) | ✅ Connected | Gateway disconnected since Jul 11; watchdog restart didn't re-establish Zulip | Full gateway restart (kill wrapper, let infisical-gateway.sh respawn) |
|
| **Tanko** | Hermes (CT 112) | ✅ Connected | Gateway disconnected since Jul 11; watchdog restart didn't re-establish Zulip | Full gateway restart (kill wrapper, let infisical-gateway.sh respawn) |
|
||||||
| **Mumuni** | Hermes (inside Abiba CT 100) | ✅ Connected | No issues found | None needed |
|
| **Mumuni** | Hermes (kagentz CT 105, migrated 2026-08-29) | ✅ Connected | No issues found | None needed |
|
||||||
|
|
||||||
### Key Fixes Applied
|
### Key Fixes Applied
|
||||||
|
|
||||||
|
|||||||
@@ -13,7 +13,8 @@ triggers:
|
|||||||
|
|
||||||
> **⚠️ RETIRED** — This contract was embedded in the pi Zulip extension code
|
> **⚠️ RETIRED** — This contract was embedded in the pi Zulip extension code
|
||||||
> (`performHealthCheck()`). That code has been removed. Zulip self-healing for
|
> (`performHealthCheck()`). That code has been removed. Zulip self-healing for
|
||||||
> Hermes agents (Tanko, Mumuni) continues through their own gateway monitoring.
|
> agents (Mumuni on Hermes, Tanko on DSH) continues through their own platform
|
||||||
|
> monitoring.
|
||||||
|
|
||||||
## Maintains
|
## Maintains
|
||||||
|
|
||||||
@@ -65,8 +66,7 @@ triggers:
|
|||||||
|------|----|------|---------|
|
|------|----|------|---------|
|
||||||
| Zulip server | 192.168.68.19 | root | Docker: `zulip-zulip-1` |
|
| Zulip server | 192.168.68.19 | root | Docker: `zulip-zulip-1` |
|
||||||
| Abiba (pi) | localhost | root | PM2: `abiba-zulip` |
|
| Abiba (pi) | localhost | root | PM2: `abiba-zulip` |
|
||||||
| Mumuni | 192.168.68.24 (CT100 abiba) | root | `hermes gateway restart` |
|
| Tanko | 192.168.68.122 (CT 112) | jerome | DSH (DeepSeek Harness) — restart via DSH service, not `hermes gateway restart` (no longer a Hermes agent since 2026-08-27) |
|
||||||
| Tanko | 192.168.68.122 | jerome | `PATH=$PATH:/home/jerome/.hermes/hermes-agent hermes gateway restart` |
|
|
||||||
|
|
||||||
## Debounce
|
## Debounce
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user