Compare commits

...
Author SHA1 Message Date
root a48b947242 fix(alignment): correct tanko status - hybrid (DSH + Hermes), not DSH-only
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 21s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 21s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 20s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
PR #145 originally claimed 'the other 3 CTs are DSH-only' but all 4 agents
run a Hermes gateway. Only tanko additionally runs DSH, making it the hybrid
one. Fixed line 300 to say what the contract actually checks for each agent,
and fixed line 76 to stop asserting tanko's Hermes config is gone.
2026-09-28 12:57:00 +00:00
root 88e9243ce4 fix(alignment): clarify tanko's live LiteLLM proxy status in health-check description 2026-09-28 12:38:42 +00:00
root 2851a0cfc8 Merge PR #143: fix(agent-health): use llmuser for .8 GPU health probe instead of root
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 17s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 14s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 22s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 3s
2026-09-28 07:27:50 +00:00
root e0c9852de8 no-mistakes(document): Document llmuser SSH user for .8 GPU health probe
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 12s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 24s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 23s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
2026-09-28 07:22:13 +00:00
root bf3a1ba523 fix(agent-health): use llmuser for .8 GPU health probe instead of root
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Failing after 28s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
The rebuilt VM 101 (192.168.68.8) kept only one SSH key in root's
authorized_keys, so the health check's root probe returns Permission
denied and misreports the healthy host as UNREACHABLE. The llama-server
runs as llmuser, so that user can see the :8080 pid via ss -tlnp.

Add per-host user to GPU_HOSTS (default root, llmuser for .8) and pass
it through check_gpu_ports() into all ssh() calls.

Closes the gpu-unreachable:192.168.68.8 leg while leaving the root SSH
security decision for the captain.
2026-09-28 07:12:57 +00:00
4 changed files with 48 additions and 9 deletions
+1 -1
View File
@@ -24,7 +24,7 @@ Runs every 4 hours (2, 6, 10, 14, 18, 22 UTC at :35) via cron (`35 2,6,10,14,18,
## Requires
- **LiteLLM admin key** for key validation (retrieved from `/root/.pi/agent/env.sh`)
- **SSH access** to GPU hosts (.8, .110, .15) and agent CTs (.122, .129, .114, .24)
- **SSH access** to GPU hosts — `llmuser` on .8 (owns `llama-server`), `root` on .110 and .15 — and agent CTs (.122, .129, .114, .24)
- **Python 3** for script execution
- **Network access** to LiteLLM (:4000), GPU exporters (:9400), and gateway endpoints
+3 -3
View File
@@ -70,10 +70,10 @@ Agent (systemd) → LITELLM_API_KEY → LiteLLM (:116/v1) → GPU (llama-server)
## Config Pattern — Mandatory Fields
### For Hermes Agents (Mumuni, Koonimo)
### For Hermes Agents (Mumuni, Koonimo, Tanko-hybrid)
Every Hermes agent's `/root/.hermes/config.yaml` (or `/home/jerome/.hermes/config.yaml`) MUST have:
(Tanko is excluded — migrated to DSH/DeepSeek Harness on 2026-08-27, no longer uses Hermes config.)
(Tanko is hybrid — runs both DSH and Hermes since 2026-08-27, so its Hermes config is also checked.)
### 1. Main Model
```yaml
@@ -297,7 +297,7 @@ Run the consolidated health check:
```bash
python3 /root/scripts/agent-health-check.py
```
This validates all 4 LiteLLM keys, detects GPU port conflicts (ghost processes),
This validates each agent's live LiteLLM key against the gateway, including tanko, which runs HYBRID (DSH + Hermes) since 2026-08-27; detects GPU port conflicts (ghost processes),
verifies gateway liveness, confirms Zulip streaming (`edit_message` present),
and counts recent errors. Non-disruptive — never restarts anything.
+11 -5
View File
@@ -50,6 +50,11 @@ Changelog:
(kagentz CT 105 on minipve, .14, dedicated `hermes` user) and is monitored
from her side. This script must not probe mumuni or .24 — the v2 changelog
roster line was the last reference still placing her at .24 / CT100.
v6 (2026-09-28): .8 GPU health probe now runs as `llmuser` instead of `root`.
Root SSH to .8 was lost when the guest was rebuilt, so every .8 leg read as
UNREACHABLE for a healthy host. llmuser owns llama-server and can read
`systemctl is-active`, `systemctl show -p MainPID`, and the :8080 pid.
.110 and .15 keep the default `root` user.
"""
import subprocess, json, sys, os, time, re, io, contextlib
@@ -95,7 +100,7 @@ AGENTS = {
# .110 rtx5070 (ocu-llm VM) -> llama-server.service (active)
# .15 strixhalo (amdpve) -> strix-server.service (active)
GPU_HOSTS = {
"gpu-rtx3090 (.8)": {"host": "192.168.68.8", "port": 8080, "service": "llama-chat-api.service"},
"gpu-rtx3090 (.8)": {"host": "192.168.68.8", "port": 8080, "service": "llama-chat-api.service", "user": "llmuser"},
"gpu-rtx5070 (.110)": {"host": "192.168.68.110", "port": 8080, "service": "llama-server.service"},
"gpu-strixhalo (.15)": {"host": "192.168.68.15", "port": 8080, "service": "strix-server.service"},
}
@@ -300,13 +305,14 @@ def check_gpu_ports():
host = gpu["host"]
port = gpu["port"]
svc = gpu["service"]
user = gpu.get("user", "root") # default root, overridden per-host where needed
# `systemctl is-active` exits non-zero when the unit is inactive or
# missing, which the ssh() helper would swallow as an SSH failure and
# report as UNREACHABLE. `|| true` keeps the real state word so we can
# tell "unit inactive" from "host unreachable".
svc_status = ssh(host, f"systemctl is-active {svc} || true")
port_owner = ssh(host, f"ss -tlnp 2>/dev/null | grep -Po ':{port}\\s+.*pid=\\K[0-9]+' | head -1")
svc_status = ssh(host, f"systemctl is-active {svc} || true", user=user)
port_owner = ssh(host, f"ss -tlnp 2>/dev/null | grep -Po ':{port}\\s+.*pid=\\K[0-9]+' | head -1", user=user)
if not svc_status:
print(f" ❌ {label}: UNREACHABLE")
@@ -317,14 +323,14 @@ def check_gpu_ports():
print(f" ❌ {label}: PORT {port} NOT LISTENING (svc={svc_status})")
FAIL.append(f"gpu-no-port:{label}")
elif svc_status != "active":
svc_pid = ssh(host, f"systemctl show {svc} -p MainPID 2>/dev/null | cut -d= -f2")
svc_pid = ssh(host, f"systemctl show {svc} -p MainPID 2>/dev/null | cut -d= -f2", user=user)
if svc_pid and port_owner != svc_pid:
print(f" ❌ {label}: GHOST PROCESS — port owned by pid {port_owner}, svc pid {svc_pid} (svc={svc_status})")
FAIL.append(f"gpu-ghost:{label}:{port_owner}")
else:
print(f" ⚠️ {label}: svc={svc_status}, port owned by {port_owner}")
else:
health = ssh(host, f"curl -s --max-time 5 http://localhost:{port}/health")
health = ssh(host, f"curl -s --max-time 5 http://localhost:{port}/health", user=user)
if health and '"status":"ok"' in health:
print(f" ✅ {label}: healthy (pid={port_owner})")
elif health and '"status":"no slot available"' in health:
+33
View File
@@ -124,6 +124,39 @@ def test_tanko_ct112_is_probed_on_minipve(ahc, monkeypatch, capsys):
ahc.REPORT_ONLY.clear()
def test_gpu_rtx3090_probe_uses_llmuser_not_root(ahc, monkeypatch, capsys):
# 2026-09-28: root SSH to .8 was lost when the guest was rebuilt; llmuser
# owns llama-server and can read systemctl status and the :8080 pid. A root
# probe reads as UNREACHABLE for a healthy host (the reported bug). Execute
# check_gpu_ports() against an SSH boundary that only accepts llmuser@.8 and
# assert the .8 leg does not produce the false UNREACHABLE failure.
seen = []
def fake_ssh(host, cmd, user="root"):
seen.append((host, user))
if host == "192.168.68.8" and user != "llmuser":
return None # root SSH denied -> baseline false UNREACHABLE
if cmd.startswith("systemctl is-active"):
return "active"
if cmd.startswith("ss -tlnp"):
return "48351"
if cmd.startswith("curl"):
return '{"status":"ok"}'
return None
monkeypatch.setattr(ahc, "ssh", fake_ssh)
ahc.FAIL.clear()
try:
ahc.check_gpu_ports()
out = capsys.readouterr().out
assert "gpu-unreachable:192.168.68.8" not in ahc.FAIL
assert "\u2705 gpu-rtx3090 (.8): healthy" in out
assert ("192.168.68.8", "llmuser") in seen
assert not any(host == "192.168.68.8" and user == "root" for host, user in seen)
finally:
ahc.FAIL.clear()
def test_report_only_legs_never_count_as_failures(ahc):
for agent, report_only in (("koby", True), ("koonimo", False), ("tanko", False)):
ahc.FAIL.clear()