Compare commits

..
Author SHA1 Message Date
abiba-bot 6b2ba1bba5 Merge pull request 'fix(alignment): clarify tanko live LiteLLM proxy status in health-check description' (#145) from fix/contract-alignment-f2-hermes-baseline-20260928 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 27s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 33s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 14s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 4s
2026-09-28 12:59:50 +00:00
root a48b947242 fix(alignment): correct tanko status - hybrid (DSH + Hermes), not DSH-only
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 21s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 21s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 20s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
PR #145 originally claimed 'the other 3 CTs are DSH-only' but all 4 agents
run a Hermes gateway. Only tanko additionally runs DSH, making it the hybrid
one. Fixed line 300 to say what the contract actually checks for each agent,
and fixed line 76 to stop asserting tanko's Hermes config is gone.
2026-09-28 12:57:00 +00:00
abiba-bot ba9d29b4b9 Merge pull request 'fix(alignment): recognize tanko as hybrid (DSH + Hermes) in agent-health-check' (#146) from fix/contract-alignment-f3-tanko-hybrid-20260928 into master 2026-09-28 12:54:42 +00:00
abiba-bot d6376e5142 Merge pull request 'fix(alignment): use llmuser with sudo for swap-gpu-dense-model.sh' (#144) from fix/contract-alignment-f1-swap-gpu-user-20260928 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 18s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 24s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 15s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 2s
2026-09-28 12:53:53 +00:00
root 308265e7ce fix(alignment): recognize tanko as hybrid (DSH + Hermes) in agent-health-check
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 17s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 32s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 17s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 3s
Tanko runs both DSH (pnpm dsh web) and Hermes (hermes gateway) concurrently.
The false premise was that tanko is DSH-only with no Hermes gateway, which
caused the gateway liveness, config.yaml, and wrapper integrity checks to be
skipped entirely. Now tanko gets the full Hermes-era checks like koonimo and
koby, while dsh/pi-only agents still skip those legs correctly.
2026-09-28 12:44:25 +00:00
root 88e9243ce4 fix(alignment): clarify tanko's live LiteLLM proxy status in health-check description 2026-09-28 12:38:42 +00:00
root 13ac189365 fix(alignment): use llmuser with sudo for swap-gpu-dense-model.sh (root SSH to .8 denied)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 25s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 16s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 42s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 23s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-09-28 12:34:32 +00:00
root 2851a0cfc8 Merge PR #143: fix(agent-health): use llmuser for .8 GPU health probe instead of root
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 17s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 14s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 22s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 3s
2026-09-28 07:27:50 +00:00
root e0c9852de8 no-mistakes(document): Document llmuser SSH user for .8 GPU health probe
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 12s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 24s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 23s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
2026-09-28 07:22:13 +00:00
root bf3a1ba523 fix(agent-health): use llmuser for .8 GPU health probe instead of root
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Failing after 28s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
The rebuilt VM 101 (192.168.68.8) kept only one SSH key in root's
authorized_keys, so the health check's root probe returns Permission
denied and misreports the healthy host as UNREACHABLE. The llama-server
runs as llmuser, so that user can see the :8080 pid via ss -tlnp.

Add per-host user to GPU_HOSTS (default root, llmuser for .8) and pass
it through check_gpu_ports() into all ssh() calls.

Closes the gpu-unreachable:192.168.68.8 leg while leaving the root SSH
security decision for the captain.
2026-09-28 07:12:57 +00:00
root f230812e3a Merge PR #142: sync tanko CT 112 PVE mapping to minipve across consumers
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 18s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-28 07:01:19 +00:00
root 96769a103f no-mistakes(review): sync tanko CT 112 mapping to minipve across consumers
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 12s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 17s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 3s
2026-09-28 06:47:46 +00:00
root d697baa7b6 fix(agent-health): update tanko PVE mapping from amdpve to minipve (CT 112 migrated 2026-09-27)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 13s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 24s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 26s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-09-28 06:35:34 +00:00
abiba-bot fb7f351a2b Merge pull request 'fix(audit-hermes): handle fallback_providers as list or dict' (#141) from fix/audit-hermes-fallback-list-20260927 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-27 11:35:07 +00:00
18 changed files with 124 additions and 65 deletions
+1 -1
View File
@@ -24,7 +24,7 @@ Runs every 4 hours (2, 6, 10, 14, 18, 22 UTC at :35) via cron (`35 2,6,10,14,18,
## Requires
- **LiteLLM admin key** for key validation (retrieved from `/root/.pi/agent/env.sh`)
- **SSH access** to GPU hosts (.8, .110, .15) and agent CTs (.122, .129, .114, .24)
- **SSH access** to GPU hosts — `llmuser` on .8 (owns `llama-server`), `root` on .110 and .15 — and agent CTs (.122, .129, .114, .24)
- **Python 3** for script execution
- **Network access** to LiteLLM (:4000), GPU exporters (:9400), and gateway endpoints
+1 -1
View File
@@ -643,7 +643,7 @@ contracts:
timeout: 120
requires:
- Zulip API key for abiba-bot@chat.sysloggh.net
- SSH access to amdpve (192.168.68.15) for Tanko (CT 112) and the Agent Zero Docker host (.14)
- SSH access to minipve (192.168.68.12) for Tanko (CT 112) and the Agent Zero Docker host (.14)
verification:
postconditions:
- check: bot registration active
+1 -1
View File
@@ -414,7 +414,7 @@ one-off GPU builds. No automated post-migration cleanup was in place.
| 108 | media | storepve | lxc | ✅ reachable |
| 110 | gitea | minipve | lxc | ✅ reachable |
| 111 | tdunna | **storepve** | lxc | ⛔ **REPORT-ONLY** (192.168.68.129, Theo's box — no GC at any level) |
| 112 | tanko | amdpve | lxc | ✅ reachable |
| 112 | tanko | minipve | lxc | ✅ reachable |
| 113 | baggy | amdpve | lxc | ✅ reachable |
| 115 | scottdenya | amdpve | lxc | ✅ reachable |
| 116 | syslog-api | minipve | lxc | ✅ reachable |
+3 -3
View File
@@ -70,10 +70,10 @@ Agent (systemd) → LITELLM_API_KEY → LiteLLM (:116/v1) → GPU (llama-server)
## Config Pattern — Mandatory Fields
### For Hermes Agents (Mumuni, Koonimo)
### For Hermes Agents (Mumuni, Koonimo, Tanko-hybrid)
Every Hermes agent's `/root/.hermes/config.yaml` (or `/home/jerome/.hermes/config.yaml`) MUST have:
(Tanko is excluded — migrated to DSH/DeepSeek Harness on 2026-08-27, no longer uses Hermes config.)
(Tanko is hybrid — runs both DSH and Hermes since 2026-08-27, so its Hermes config is also checked.)
### 1. Main Model
```yaml
@@ -297,7 +297,7 @@ Run the consolidated health check:
```bash
python3 /root/scripts/agent-health-check.py
```
This validates all 4 LiteLLM keys, detects GPU port conflicts (ghost processes),
This validates each agent's live LiteLLM key against the gateway, including tanko, which runs HYBRID (DSH + Hermes) since 2026-08-27; detects GPU port conflicts (ghost processes),
verifies gateway liveness, confirms Zulip streaming (`edit_message` present),
and counts recent errors. Non-disruptive — never restarts anything.
+2 -2
View File
@@ -55,7 +55,7 @@ connectivity recovery including end-to-end DM validation.
| Host | CT | Proxmox | IP (direct) | Hermes Home | User |
|------|-----|---------|-------------|-------------|------|
| Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome | *(DSH since 2026-08-27 — historical, plugin retired on this host)* |
| Tanko | CT112 | minipve | 192.168.68.122 | /home/jerome/.hermes | jerome | *(DSH since 2026-08-27 — historical, plugin retired on this host)* |
| Koby | CT111 | storepve | 192.168.68.129 | /root/.hermes | root |
| Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky |
@@ -72,7 +72,7 @@ connectivity recovery including end-to-end DM validation.
### Step 1: Resolve Target
Map `target` to host, CT ID, hermes_home, and user from the live-state table.
For CT112 route through `ssh root@amdpve`; for CT111 route through `ssh root@storepve` — then `pct exec <id>`.
For CT112 route through `ssh root@minipve`; for CT111 route through `ssh root@storepve` — then `pct exec <id>`.
### Step 2: Pull Latest Plugin Source
+1 -1
View File
@@ -67,7 +67,7 @@ gateway restart, and connection validation.
### Step 1: Locate Target
Map `target` to connectivity parameters from the live-state table above.
For CT112 route through `ssh root@amdpve`; for CT111 route through `ssh root@storepve` — then `pct exec <id>`.
For CT112 route through `ssh root@minipve`; for CT111 route through `ssh root@storepve` — then `pct exec <id>`.
### Step 2: Deploy Zulip Adapter
+4 -4
View File
@@ -105,8 +105,8 @@ description: >
| Node | IP | CPU | RAM | VMs/CTs | Role |
|------|----|-----|-----|---------|------|
| minipve | .12 | 16C | 30GB | abiba, authentik, gitea, syslog-api, infisical-vault, jitsi | Auth, git, messaging |
| amdpve | .15 | 32C | 62GB | kagentz, tanko, baggy, scottdenya, adguard2 | Agents, compute |
| minipve | .12 | 16C | 30GB | abiba, tanko, authentik, gitea, syslog-api, infisical-vault, jitsi | Auth, git, messaging |
| amdpve | .15 | 32C | 62GB | kagentz, baggy, scottdenya, adguard2 | Agents, compute |
| storepve | .6 | 28C | 31GB | docker-vm, ra-h-os, PBS, media, jdownloader, zulip, tdunna | Docker, storage, chat |
| acerpve | .9 | 28C | 31GB | llm-gpu | GPU VMs |
| ocupve | .5 | 12C | 14GB | ocu-llm | GPU VMs |
@@ -682,7 +682,7 @@ ssh root@192.168.68.110 "systemctl restart llama-server"
| 109 | docker-vm | storepve | .7 | Docker host | ❌ |
| 110 | gitea | minipve | **.17** | Git | ❌ |
| 111 | tdunna | storepve | .129 | Hermes agent — ⛔ REPORT-ONLY (Theo's box, no GC) | ✅ |
| 112 | tanko | amdpve | .122 | DSH (DeepSeek Harness) agent | ✅ |
| 112 | tanko | minipve | .122 | DSH (DeepSeek Harness) agent | ✅ |
| 113 | baggy | amdpve | .114 | Hermes agent | ✅ |
| 115 | scottdenya | amdpve | .75 | Denya OneCare | ❌ |
| 116 | syslog-api | minipve | .116 | LiteLLM + Grafana | ❌ |
@@ -712,7 +712,7 @@ Source of truth: `/root/scripts/pct-run.sh` or `prose-contracts/scripts/pct-run.
| 100 | abiba | minipve | `pct-run 100` |
| 105 | kagentz | amdpve | `pct-run 105` |
| 111 | tdunna | storepve | `pct-run 111` (⛔ report-only — no GC) |
| 112 | tanko | amdpve | `pct-run 112` |
| 112 | tanko | minipve | `pct-run 112` |
| 113 | baggy | amdpve | `pct-run 113` |
| 115 | scottdenya | amdpve | `pct-run 115` |
| 104 | authentik | minipve | `pct-run 104` |
+1 -1
View File
@@ -59,7 +59,7 @@ Before ANY update wave:
| ocupve (.5) | Proxmox node | `apt update && apt upgrade -y` | 5 min |
| CT 100 (.24) | Abiba (pi) | `apt update && apt upgrade -y` | 3 min |
| CT 116 (.116) | syslog-api (LiteLLM host) | `apt update && apt upgrade -y` | 3 min |
| CT 112 (tanko, amdpve) | Tanko | `apt update && apt upgrade -y` | 3 min |
| CT 112 (tanko, minipve) | Tanko | `apt update && apt upgrade -y` | 3 min |
| CT 105 (kagentz, minipve) | Mumuni | `apt update && apt upgrade -y` | 3 min |
| VM 101 (.8) | llm-gpu (RTX 3090) | `apt update && apt upgrade -y` | 3 min |
| VM 103 (.110) | ocu-llm (RTX 5070) | `apt update && apt upgrade -y` | 3 min |
+17 -11
View File
@@ -50,6 +50,11 @@ Changelog:
(kagentz CT 105 on minipve, .14, dedicated `hermes` user) and is monitored
from her side. This script must not probe mumuni or .24 — the v2 changelog
roster line was the last reference still placing her at .24 / CT100.
v6 (2026-09-28): .8 GPU health probe now runs as `llmuser` instead of `root`.
Root SSH to .8 was lost when the guest was rebuilt, so every .8 leg read as
UNREACHABLE for a healthy host. llmuser owns llama-server and can read
`systemctl is-active`, `systemctl show -p MainPID`, and the :8080 pid.
.110 and .15 keep the default `root` user.
"""
import subprocess, json, sys, os, time, re, io, contextlib
@@ -70,7 +75,7 @@ PVE_NODES = {
# Agent definitions: ct, host, user, pve_node, vault_key_name
AGENTS = {
"tanko": {"ct": 112, "host": "192.168.68.122", "user": "jerome", "pve": "amdpve", "vault_key": "TANKO_LITELLM_API_KEY", "runtime": "dsh"},
"tanko": {"ct": 112, "host": "192.168.68.122", "user": "jerome", "pve": "minipve", "vault_key": "TANKO_LITELLM_API_KEY", "runtime": "hybrid"},
# abiba = pi agent (.24) — no vault key; its LiteLLM key is read from its
# local env file (key_env below), not from the shared vault or .bashrc.
# runtime=pi: abiba has run pi-only since the harness purge. There is no
@@ -95,7 +100,7 @@ AGENTS = {
# .110 rtx5070 (ocu-llm VM) -> llama-server.service (active)
# .15 strixhalo (amdpve) -> strix-server.service (active)
GPU_HOSTS = {
"gpu-rtx3090 (.8)": {"host": "192.168.68.8", "port": 8080, "service": "llama-chat-api.service"},
"gpu-rtx3090 (.8)": {"host": "192.168.68.8", "port": 8080, "service": "llama-chat-api.service", "user": "llmuser"},
"gpu-rtx5070 (.110)": {"host": "192.168.68.110", "port": 8080, "service": "llama-server.service"},
"gpu-strixhalo (.15)": {"host": "192.168.68.15", "port": 8080, "service": "strix-server.service"},
}
@@ -300,13 +305,14 @@ def check_gpu_ports():
host = gpu["host"]
port = gpu["port"]
svc = gpu["service"]
user = gpu.get("user", "root") # default root, overridden per-host where needed
# `systemctl is-active` exits non-zero when the unit is inactive or
# missing, which the ssh() helper would swallow as an SSH failure and
# report as UNREACHABLE. `|| true` keeps the real state word so we can
# tell "unit inactive" from "host unreachable".
svc_status = ssh(host, f"systemctl is-active {svc} || true")
port_owner = ssh(host, f"ss -tlnp 2>/dev/null | grep -Po ':{port}\\s+.*pid=\\K[0-9]+' | head -1")
svc_status = ssh(host, f"systemctl is-active {svc} || true", user=user)
port_owner = ssh(host, f"ss -tlnp 2>/dev/null | grep -Po ':{port}\\s+.*pid=\\K[0-9]+' | head -1", user=user)
if not svc_status:
print(f" ❌ {label}: UNREACHABLE")
@@ -317,14 +323,14 @@ def check_gpu_ports():
print(f" ❌ {label}: PORT {port} NOT LISTENING (svc={svc_status})")
FAIL.append(f"gpu-no-port:{label}")
elif svc_status != "active":
svc_pid = ssh(host, f"systemctl show {svc} -p MainPID 2>/dev/null | cut -d= -f2")
svc_pid = ssh(host, f"systemctl show {svc} -p MainPID 2>/dev/null | cut -d= -f2", user=user)
if svc_pid and port_owner != svc_pid:
print(f" ❌ {label}: GHOST PROCESS — port owned by pid {port_owner}, svc pid {svc_pid} (svc={svc_status})")
FAIL.append(f"gpu-ghost:{label}:{port_owner}")
else:
print(f" ⚠️ {label}: svc={svc_status}, port owned by {port_owner}")
else:
health = ssh(host, f"curl -s --max-time 5 http://localhost:{port}/health")
health = ssh(host, f"curl -s --max-time 5 http://localhost:{port}/health", user=user)
if health and '"status":"ok"' in health:
print(f" ✅ {label}: healthy (pid={port_owner})")
elif health and '"status":"no slot available"' in health:
@@ -376,10 +382,10 @@ def check_agents():
ct = agent["ct"]
report_only = agent.get("report_only", False)
# Tanko runs on DSH (DeepSeek Harness) since 2026-08-27 — it no longer runs a
# Tanko runs hybrid (DSH + Hermes) since 2026-08-27 — it runs both DSH and Hermes gateway.
# Hermes gateway, so skip the Hermes gateway/state/streaming/journal checks.
# Non-Hermes runtimes have no gateway to probe. dsh = Tanko since
# 2026-08-27; pi = abiba since the harness purge (.24 is pi-only).
# Non-Hermes runtimes have no gateway to probe. dsh/pi-only skip the check;
# hybrid runs both DSH and Hermes and is checked normally.
if agent.get("runtime") in ("dsh", "pi"):
is_dsh = agent.get("runtime") == "dsh"
label = "DSH (DeepSeek Harness)" if is_dsh else "pi-only runtime"
@@ -500,7 +506,7 @@ def check_ct_liveness():
def check_config_integrity():
"""Verify agent config.yaml parses as valid YAML."""
for name, agent in AGENTS.items():
# Tanko runs on DSH (DeepSeek Harness) since 2026-08-27 — no Hermes config.yaml.
# DSH/pi-only runtimes have no Hermes config.yaml; hybrid has both.
if agent.get("runtime") == "dsh":
print(f" ⏭️ {name}: DSH — no Hermes config.yaml since 2026-08-27")
continue
@@ -556,7 +562,7 @@ def _infisical_invocation_paths(wrapper_body):
def check_wrapper_integrity():
"""Verify the hermes CLI wrapper exists and can reach hermes-real."""
for name, agent in AGENTS.items():
# Tanko runs on DSH (DeepSeek Harness) since 2026-08-27 — no hermes CLI wrapper.
# DSH/pi-only runtimes have no hermes CLI wrapper; hybrid has both.
if agent.get("runtime") == "dsh":
print(f" ⏭️ {name}: DSH — no hermes CLI wrapper since 2026-08-27")
continue
+2 -2
View File
@@ -101,8 +101,6 @@ GUESTS: list[Guest] = [
# amdpve (192.168.68.15)
Guest(ct_id="105", hostname="kagentz", ip="192.168.68.105", node="amdpve",
access_method="ssh-host", probe_target="kagentz (CT 105, amdpve)"),
Guest(ct_id="112", hostname="tanko", ip="192.168.68.112", node="amdpve",
access_method="pct-run", probe_target="tanko (CT 112, amdpve)"),
Guest(ct_id="113", hostname="baggy", ip="192.168.68.113", node="amdpve",
access_method="pct-run", probe_target="baggy (CT 113, amdpve)"),
Guest(ct_id="115", hostname="scottdenya", ip="192.168.68.115", node="amdpve",
@@ -110,6 +108,8 @@ GUESTS: list[Guest] = [
Guest(ct_id="120", hostname="adguard2", ip="192.168.68.120", node="amdpve",
access_method="pct-run", probe_target="adguard2 (CT 120, amdpve)"),
# minipve (192.168.68.12)
Guest(ct_id="112", hostname="tanko", ip="192.168.68.112", node="minipve",
access_method="pct-run", probe_target="tanko (CT 112, minipve)"),
Guest(ct_id="100", hostname="abiba", ip="192.168.68.100", node="minipve",
access_method="pct-run", probe_target="abiba (CT 100, minipve)"),
Guest(ct_id="102", hostname="adguard", ip="192.168.68.102", node="minipve",
+1 -1
View File
@@ -12,7 +12,6 @@ set -euo pipefail
declare -A CT_NODES=(
# amdpve (192.168.68.15)
[105]=amdpve # kagentz (was hwepve — corrected 2026-09-12; live per pvesh)
[112]=amdpve # tanko
[113]=amdpve # baggy
[115]=amdpve # scottdenya
[120]=amdpve # adguard2 (added 2026-09-12)
@@ -21,6 +20,7 @@ declare -A CT_NODES=(
[102]=minipve # adguard (was acerpve)
[104]=minipve # authentik
[110]=minipve # gitea
[112]=minipve # tanko (was amdpve — migrated 2026-09-27; live per pvesh)
[116]=minipve # syslog-api
[119]=minipve # infisical-vault
# storepve (192.168.68.6)
+2 -2
View File
@@ -47,8 +47,8 @@ You are a code reviewer for OpenProse infrastructure contracts in the Syslog Sol
The infrastructure-control.prose.md contract is the canonical reference for the cluster topology:
**Proxmox Cluster "Tabiri" (5 nodes):**
- amdpve (192.168.68.15): kagentz, tanko, baggy, scottdenya, adguard2
- minipve (192.168.68.12): abiba, adguard, authentik, gitea, syslog-api, infisical-vault
- amdpve (192.168.68.15): kagentz, baggy, scottdenya, adguard2
- minipve (192.168.68.12): abiba, tanko, adguard, authentik, gitea, syslog-api, infisical-vault
- storepve (192.168.68.6): docker-vm, ra-h-os, PBS, media, jdownloader, zulip, tdunna
- acerpve (192.168.68.9): llm-gpu
- ocupve (192.168.68.5): ocu-llm
+1 -1
View File
@@ -1,6 +1,6 @@
#!/bin/bash
# swap-gpu-dense-model.sh — Swap RTX 3090 from qwen3.6-27B-code to SmartCode-Fable-5
# Run when download completes: ssh root@192.168.68.8 'bash -s' < this script
# Run when download completes: ssh llmuser@192.168.68.8 'sudo bash -s' < this script
#
# Usage: bash swap-gpu-dense-model.sh
# Requires: new model at /home/llmuser/models/SmartCode-Fable-5-27B-UD-Q4_K_XL.gguf
+6 -6
View File
@@ -135,16 +135,16 @@ case "$PI_VERDICT" in
esac
# -- abiba-leg-end
# ── Platform B: Tanko (DSH dsh-web on amdpve CT 112) ──
# ── Platform B: Tanko (DSH dsh-web on minipve CT 112) ──
# Direct SSH to 192.168.68.122 is not a dependency of this monitor — per-worker
# key availability varies — so probes run from the amdpve vantage via `pct exec`.
# Tanko's Zulip gateway runs as the dsh-web systemd unit inside CT 112 on amdpve
# (192.168.68.15). The gateway binds 127.0.0.1:3080 loopback-only by design — a
# key availability varies — so probes run from the minipve vantage via `pct exec`.
# Tanko's Zulip gateway runs as the dsh-web systemd unit inside CT 112 on minipve
# (192.168.68.12). The gateway binds 127.0.0.1:3080 loopback-only by design — a
# remote :3080 probe is refused and is NOT a fault.
TANKO_SVC=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.15 \
TANKO_SVC=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.12 \
"pct exec 112 -- systemctl is-active dsh-web" 2>/dev/null || true)
[ -n "$TANKO_SVC" ] || TANKO_SVC="unknown"
TANKO_HTTP=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.15 \
TANKO_HTTP=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.12 \
"pct exec 112 -- curl -s --connect-timeout 5 --max-time 10 -o /dev/null -w '%{http_code}' http://127.0.0.1:3080/" 2>/dev/null || true)
[ -n "$TANKO_HTTP" ] || TANKO_HTTP="000"
+2 -2
View File
@@ -49,7 +49,7 @@ HEALTH_CONTRACT = ROOT / "zulip-health.prose.md"
CONNECTED_FIXTURE = ROOT / "tests" / "fixtures" / "zulip-health-connected.json"
MUMUNI_IP = "192.168.68.24" # Mumuni's old (decommissioned) deployment
TANKO_VANTAGE = "192.168.68.15" # amdpve — Tanko CT 112 via pct exec
TANKO_VANTAGE = "192.168.68.12" # minipve — Tanko CT 112 via pct exec
AGENT_ZERO_HOST = "192.168.68.14" # kagentz host, Agent Zero docker
@@ -82,7 +82,7 @@ done
printf '%s\n' "$host" >> "$RECORD_DIR/ssh.hosts"
cmd="${*: -1}"
case "$host" in
192.168.68.15)
192.168.68.12)
case "$cmd" in
*"systemctl is-active"*) printf '%s' "$TANKO_SVC" ;;
*curl*) printf '%s' "$TANKO_HTTP" ;;
+53
View File
@@ -104,6 +104,59 @@ def test_koby_ct111_is_on_storepve(ahc):
assert ahc.AGENTS["koby"]["pve"] == "storepve"
def test_tanko_ct112_is_probed_on_minipve(ahc, monkeypatch, capsys):
# CT 112 (tanko) was live-migrated to minipve (.12) on 2026-09-27; the
# amdpve mapping made `pct status 112` fail and read as ct-unreachable.
# Execute the probe and assert the host the script actually contacts.
probes = []
monkeypatch.setattr(
ahc, "ssh",
lambda host, cmd, user="root": probes.append((host, cmd)) or "status: running",
)
ahc.FAIL.clear()
ahc.REPORT_ONLY.clear()
try:
ahc.check_ct_liveness()
tanko_hosts = [h for h, cmd in probes if cmd == "pct status 112 2>/dev/null"]
assert tanko_hosts == ["192.168.68.12"]
finally:
ahc.FAIL.clear()
ahc.REPORT_ONLY.clear()
def test_gpu_rtx3090_probe_uses_llmuser_not_root(ahc, monkeypatch, capsys):
# 2026-09-28: root SSH to .8 was lost when the guest was rebuilt; llmuser
# owns llama-server and can read systemctl status and the :8080 pid. A root
# probe reads as UNREACHABLE for a healthy host (the reported bug). Execute
# check_gpu_ports() against an SSH boundary that only accepts llmuser@.8 and
# assert the .8 leg does not produce the false UNREACHABLE failure.
seen = []
def fake_ssh(host, cmd, user="root"):
seen.append((host, user))
if host == "192.168.68.8" and user != "llmuser":
return None # root SSH denied -> baseline false UNREACHABLE
if cmd.startswith("systemctl is-active"):
return "active"
if cmd.startswith("ss -tlnp"):
return "48351"
if cmd.startswith("curl"):
return '{"status":"ok"}'
return None
monkeypatch.setattr(ahc, "ssh", fake_ssh)
ahc.FAIL.clear()
try:
ahc.check_gpu_ports()
out = capsys.readouterr().out
assert "gpu-unreachable:192.168.68.8" not in ahc.FAIL
assert "\u2705 gpu-rtx3090 (.8): healthy" in out
assert ("192.168.68.8", "llmuser") in seen
assert not any(host == "192.168.68.8" and user == "root" for host, user in seen)
finally:
ahc.FAIL.clear()
def test_report_only_legs_never_count_as_failures(ahc):
for agent, report_only in (("koby", True), ("koonimo", False), ("tanko", False)):
ahc.FAIL.clear()
+2 -2
View File
@@ -34,7 +34,7 @@ ROOT = pathlib.Path(__file__).resolve().parents[1]
ZULIP_MONITOR = ROOT / "scripts" / "zulip-monitor.sh"
CONNECTED_FIXTURE = ROOT / "tests" / "fixtures" / "zulip-health-connected.json"
TANKO_VANTAGE = "192.168.68.15" # amdpve — Tanko CT 112 via pct exec
TANKO_VANTAGE = "192.168.68.12" # minipve — Tanko CT 112 via pct exec
AGENT_ZERO_HOST = "192.168.68.14" # kagentz host, Agent Zero docker
@@ -52,7 +52,7 @@ done
printf '%s\n' "$host" >> "$RECORD_DIR/ssh.hosts"
cmd="${*: -1}"
case "$host" in
192.168.68.15)
192.168.68.12)
case "$cmd" in
*"systemctl is-active"*) printf '%s' "$TANKO_SVC" ;;
*curl*) printf '%s' "$TANKO_HTTP" ;;
+24 -24
View File
@@ -28,7 +28,7 @@ session start.
## Requires
- **Zulip API key** for `abiba-bot@chat.sysloggh.net` in `$ZULIP_API_KEY`
- **SSH access** to amdpve (192.168.68.15) for Tanko — CT 112 reached via `pct exec` (direct SSH to .122 is not a dependency of this contract: per-worker key availability varies); and the Agent Zero Docker host (192.168.68.14)
- **SSH access** to minipve (192.168.68.12) for Tanko — CT 112 reached via `pct exec` (direct SSH to .122 is not a dependency of this contract: per-worker key availability varies); and the Agent Zero Docker host (192.168.68.14)
- **PM2** on localhost for pi process management
- **Network access** to `chat.sysloggh.net`, `kagentz.sysloggh.net` (C3 public path), `localhost:9200`
- **Write access** to `/root/zulip-health-monitor.log` and `/tmp/zulip-monitor-debounce`
@@ -228,20 +228,20 @@ grep -a "Finalized\|Failed to finalize" /root/.pm2/logs/abiba-zulip-out.log | ta
| Crash loop >10/h | Alert user |
### Step 3: Platform B — Tanko (DSH on amdpve CT 112)
### Step 3: Platform B — Tanko (DSH on minipve CT 112)
Mumuni is out of scope for this host (see the note above): she runs on her own
container and is monitored on her side.
Tanko runs on DSH (DeepSeek Harness) — it no longer runs a Hermes gateway, so
there is no `~/.hermes/gateway_state.json` on CT 112. Tanko's Zulip gateway runs
as the `dsh-web` systemd unit inside **CT 112**, which resides on the **amdpve**
PVE host (**192.168.68.15**). Direct SSH to 192.168.68.122 is not a dependency
as the `dsh-web` systemd unit inside **CT 112**, which resides on the **minipve**
PVE host (**192.168.68.12**). Direct SSH to 192.168.68.122 is not a dependency
of this contract — per-worker key availability varies — so CT 112 probes run
from the amdpve vantage via `pct exec`:
from the minipve vantage via `pct exec`:
```bash
ssh root@192.168.68.15 "pct exec 112 -- <command>"
ssh root@192.168.68.12 "pct exec 112 -- <command>"
```
> **By design (verified 2026-09-08):** the `dsh-web` gateway binds
@@ -253,7 +253,7 @@ ssh root@192.168.68.15 "pct exec 112 -- <command>"
**B1: Gateway Service State (Tanko)**
```bash
ssh root@192.168.68.15 "pct exec 112 -- systemctl is-active dsh-web"
ssh root@192.168.68.12 "pct exec 112 -- systemctl is-active dsh-web"
```
Expected: `active`. Anything else → gateway service down → apply the Tanko heal
@@ -262,7 +262,7 @@ Expected: `active`. Anything else → gateway service down → apply the Tanko h
**B2: Gateway HTTP Liveness (Tanko — loopback-only :3080)**
```bash
ssh root@192.168.68.15 "pct exec 112 -- curl -s --connect-timeout 5 --max-time 10 -o /dev/null -w '%{http_code}' http://127.0.0.1:3080/"
ssh root@192.168.68.12 "pct exec 112 -- curl -s --connect-timeout 5 --max-time 10 -o /dev/null -w '%{http_code}' http://127.0.0.1:3080/"
```
Alive = **ANY** HTTP status response from the endpoint — the expected set is
@@ -273,13 +273,13 @@ process answering `503` is running and self-heal must NOT restart-loop it.
Down = connection refused (`000`) or timeout only. Statuses outside the
expected set are logged/reported as a warning — reported, never healed on.
**B3: Public-URL Fallback Probe (Tanko — for nodes without pct/ssh access to amdpve)**
**B3: Public-URL Fallback Probe (Tanko — for nodes without pct/ssh access to minipve)**
```bash
curl -s --connect-timeout 10 --max-time 15 -o /dev/null -w '%{http_code}' https://tankodhs.sysloggh.net/
```
Fallback only — used when the monitoring node has no pct/SSH path to amdpve.
Fallback only — used when the monitoring node has no pct/SSH path to minipve.
Alive = **ANY** HTTP status response from the endpoint — healthy signals are
`302` (authentik proxy-auth redirect) and `401` (auth-gated), and any other
status, including `404`/`5xx`, also counts alive: the endpoint is up and
@@ -404,29 +404,29 @@ ExecStartPost=/bin/systemctl --no-block start dsh-web-token.service
4. Every later request through `/` presents that cookie; the token is not needed
again until the cookie expires or a new browser is used.
**Verification** (amdpve vantage):
**Verification** (minipve vantage):
```bash
# 1. Login endpoint is Authentik-gated: unauthenticated -> 302 (not 200/303).
ssh root@192.168.68.15 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}\n' \
ssh root@192.168.68.12 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}\n' \
-H 'Host: tankodhs.sysloggh.net' http://127.0.0.1/dsh-web-login"
# Expected: 302
# 2. Legacy :8081 endpoint is gone (connection refused -> 000).
ssh root@192.168.68.15 "pct exec 112 -- curl -s --max-time 3 -o /dev/null \
ssh root@192.168.68.12 "pct exec 112 -- curl -s --max-time 3 -o /dev/null \
-w '%{http_code}\n' http://192.168.68.122:8081/"
# Expected: 000
# 3. Backend cookie mint + reuse (exactly what /dsh-web-login proxies to).
TOKEN=$(ssh root@192.168.68.15 "pct exec 112 -- cat /etc/dsh-web/launch-token")
ssh root@192.168.68.15 "pct exec 112 -- curl -s -c /tmp/dsh.jar -o /dev/null \
TOKEN=$(ssh root@192.168.68.12 "pct exec 112 -- cat /etc/dsh-web/launch-token")
ssh root@192.168.68.12 "pct exec 112 -- curl -s -c /tmp/dsh.jar -o /dev/null \
-H 'Host: tankodhs.sysloggh.net' 'http://127.0.0.1:3080/?token=$TOKEN'"
ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh.jar -o /dev/null \
ssh root@192.168.68.12 "pct exec 112 -- curl -s -b /tmp/dsh.jar -o /dev/null \
-w '%{http_code}\n' -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/"
# Expected: 200 — the minted dsh-auth-... cookie (authority
# tankodhs.sysloggh.net) is replayed on the next request and accepted.
# 4. Token refresh is non-disruptive and idempotent.
ssh root@192.168.68.15 "pct exec 112 -- /opt/deepseek-harness/capture-dsh-token.sh"
ssh root@192.168.68.12 "pct exec 112 -- /opt/deepseek-harness/capture-dsh-token.sh"
# Expected: "token unchanged; nginx not reloaded" when nothing changed
```
@@ -437,32 +437,32 @@ fresh cookie. Both verified live 2026-09-11.
```bash
# 5. Cookie survives a dsh-web restart, and the new token mints a new cookie.
ssh root@192.168.68.15 "pct exec 112 -- systemctl restart dsh-web"
ssh root@192.168.68.12 "pct exec 112 -- systemctl restart dsh-web"
# dsh-web is Type=simple: restart returns before :3080 is listening. Bounded-poll
# until the socket answers (any status but 000) before asserting the cookie.
for i in $(seq 1 60); do
UP=$(ssh root@192.168.68.15 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}' \
UP=$(ssh root@192.168.68.12 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}' \
-H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/")
[ "$UP" != "000" ] && break
sleep 2
done
ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh.jar -o /dev/null \
ssh root@192.168.68.12 "pct exec 112 -- curl -s -b /tmp/dsh.jar -o /dev/null \
-w '%{http_code}\n' -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/"
# Expected: 200 — the pre-restart cookie is still accepted.
# The restart's ExecStartPost (or the 2-minute timer) refreshes the include. A
# manual run may no-op on the flock, so poll until the include carries a token
# the running process accepts (bounded wait) before the mint+reuse check.
for i in $(seq 1 60); do
TOKEN=$(ssh root@192.168.68.15 "pct exec 112 -- sed -n 's/.*token=//p' /etc/dsh-web/nginx-login.conf | tr -d ';\n'")
CODE=$(ssh root@192.168.68.15 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}' \
TOKEN=$(ssh root@192.168.68.12 "pct exec 112 -- sed -n 's/.*token=//p' /etc/dsh-web/nginx-login.conf | tr -d ';\n'")
CODE=$(ssh root@192.168.68.12 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}' \
-H 'Host: tankodhs.sysloggh.net' 'http://127.0.0.1:3080/?token=$TOKEN'")
[ "$CODE" = "303" ] && break
sleep 2
done
# Expected: 303 — the include now holds the token the running process accepts.
ssh root@192.168.68.15 "pct exec 112 -- curl -s -c /tmp/dsh-new.jar -o /dev/null \
ssh root@192.168.68.12 "pct exec 112 -- curl -s -c /tmp/dsh-new.jar -o /dev/null \
-H 'Host: tankodhs.sysloggh.net' 'http://127.0.0.1:3080/?token=$TOKEN'"
ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh-new.jar -o /dev/null \
ssh root@192.168.68.12 "pct exec 112 -- curl -s -b /tmp/dsh-new.jar -o /dev/null \
-w '%{http_code}\n' -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/"
# Expected: 200 — the refreshed token minted a fresh cookie.
```