Compare commits

...
Author SHA1 Message Date
root 8a4dd08b05 fix: capture 401/403 response body and key alias for credential faults
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
- get_response_body() returns first 200 chars of response body (single line)
- On 401/403 model probe: report code + body + key_alias
- Monitor key alias: monitor-20260813 (from /etc/litellm-monitor.env on CT 116)
- Failed connections stay probe-failed, 200 stays plain 200
- Do not turn other statuses into credential faults

Signed-off-by: Abiba
2026-09-15 05:11:07 +00:00
root 88b6decb31 fix: remove false Infisical claim - master key NOT in infrastructure project
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
- Replace Infisical retrieval path with proven docker exec + .env note
- State explicitly that master key is NOT in Infisical project=infrastructure
- Keep the live-key check and never-trust-a-literal instruction
- All other corrections from PR #95 preserved

Signed-off-by: Abiba
2026-09-15 04:42:31 +00:00
root 7bf9f78fc6 fix: gpu-dense probe timeout handling - report probe-failed with kind, not service verdict
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
- probe_http now returns (code, failure_kind) tuple
- Model probes report 'probe-failed: <model> <kind> (Ns timeout)' on 000
- Do not assert a service verdict from a failed probe
- 30s timeout for single-host aliases (RTX 3090 needs long warmup/prefill)
- 60s timeout for syslog-auto pool alias with retry on 000

Signed-off-by: Abiba
2026-09-15 04:22:46 +00:00
root dd9e68329e fix: correct infisical --plain flag and use verified localhost:4000 endpoint
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
- Replace broken --plain flag (prints nothing on CLI 0.43.110) with awk parsing
- Note that --plain is broken so nobody fixes it back
- Replace unproven nginx path with verified direct endpoint http://127.0.0.1:4000/key/list

Signed-off-by: Abiba
2026-09-15 03:05:30 +00:00
root cd9ec6a0df fix: correct key-lifecycle contracts to match measured reality
- hermes-key-enforcement.prose.md:
  - State that expiry must be set EXPLICITLY at creation with duration
  - Record that config default is NOT honoured by LiteLLM 1.99.1
  - Describe daily audit as AUDIT-ONLY (reports non-expiring and soon-to-expire)
  - State that renewal is NOT implemented
  - Document exclusions: abiba-pi and all crewmate keys stay WITHOUT expiry
  - koby is report-only

- litellm-api-keys.prose.md:
  - Replace literal master key with retrieval path (docker exec + infisical)
  - State that literal values must never be trusted again (key rotates)
  - Add live-key check (200 from /key/list)

Signed-off-by: Abiba
2026-09-15 02:59:19 +00:00
root 2f961d7e7a fix: agent-health-check gateway leg — deterministic probe-failed reporting
Per defect report 1154.msg (agent-health-gateway-leg-flap-20260913):
1. Add _ssh_retry() helper with one retry at longer timeout (25s)
2. Name probe target explicitly: ssh {user}@{host} <command>
3. Print probe-failed when first attempt fails, then retry
4. Only declare gateway-down after retry fails
5. Make Koby report-only explicit in output

Changes:
- check_agents(): all gateway probes now use _ssh_retry()
- All output lines name the probe target (ssh host:port)
- Koby's report-only status is explicit in output
- Never print bare "gateway down" — always name target and failure kind

Verified: koonimo shows "probe-failed" on first attempt (transient SSH),
retries at 25s, succeeds, reports ✅ koonimo: gw=running
2026-09-14 15:21:36 +00:00
abiba-bot 4fe4f3621d Merge pull request 'fix(monitoring): probe precision - any HTTP status means alive, failed probes never become service verdicts' (#92) from fix/probe-precision-20260914 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-14 15:07:35 +00:00
abiba-bot 6a1c4967db Merge pull request 'fix(disk-gc): deterministic reachability verdict and correctly labelled disk figures' (#91) from fix/disk-gc-probe-deterministic-20260914 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-14 14:43:43 +00:00
4 changed files with 172 additions and 58 deletions
+4 -2
View File
@@ -207,9 +207,11 @@ The agent picks up the new key via `infisical run --` at gateway startup.
**Keys are permanent and use bare agent name aliases.**
- **Duration**: `null` — keys never expire. NOT enforced today: CT 116 `litellm_config.yaml` has no `default_key_generate_params` block, and a key generated with no explicit models comes back with an empty models list. OPEN policy question: should agent keys expire by default? (captain security-policy decision, raised separately.)
- **Duration**: `null` — keys never expire by default. **Expiry must be set EXPLICITLY at creation** with the `duration` parameter (e.g., `90d` for 90 days). The 90-day default is the standard; however, the config default is **NOT honoured** by LiteLLM 1.99.1 (verified on CT 116: a key generated with no explicit duration returns `expires=null`). This has been recorded in `/opt/inference-harness/litellm_config.yaml` to prevent re-filing as a bug.
- **Daily Audit**: A daily audit job runs at 00:00 UTC (`/usr/local/bin/litellm-key-renewal-ct116.sh`, cron 00:00). It is **AUDIT-ONLY** and does not perform renewal. It lists every key, reports those with no expiry and those inside a 14-day warning window, explicitly EXCLUDES `abiba-pi` and `koby` (report-only, and .129 must never be touched), and logs `RENEWAL-REQUIRED-BUT-NOT-PERFORMED + NO KEY WAS CHANGED` when renewal is skipped. **Renewal is NOT implemented** — keys must not be rotated until delivery (vault injection + consumer verification) exists and is proven end-to-end.
- **Exclusions**: `abiba-pi` and every firstmate/secondmate/crewmate key stay **WITHOUT an expiry** until a proven renewal path exists. `koby` is **report-only** (never touched). These exclusions are enforced by the audit job.
- **Alias convention**: bare agent name only (e.g., `tanko`, `mumuni`, `koby`, `koonimo`). No dates, no versions. The alias IS the identity.
- **Rotation triggers**: compromise, personnel departure, or quarterly security hygiene. NOT calendar-driven.
- **Rotation triggers**: compromise, personnel departure, or quarterly security hygiene. NOT calendar-driven. Manual rotation is permitted only when the renewal delivery path is proven and verified on a throwaway consumer before production use.
- **Max budget**: $100 per key (config default).
```yaml
+9 -1
View File
@@ -320,7 +320,15 @@ directly call OpenRouter via Python's requests library. Converting would require
## LiteLLM Master Key (use sparingly — agents should NOT use it directly)
- Master key: `sk-litellm-7f96080dd99b15c36bd4b333b58a6796` (in /opt/inference-harness/.env on CT116, Infisical project=infrastructure env=production secret=LITELLM_MASTER_KEY)
- Master key: **Retrieval path (do not trust a literal value in this file — the key rotates)**:
```bash
# PRIMARY (proven, runs on CT 116 with no extra tooling):
docker exec harness-litellm printenv LITELLM_MASTER_KEY
# Note: the same value is stored in /opt/inference-harness/.env on CT 116 (verified matching)
# The master key is NOT in the Infisical vault (project=infrastructure env=production does not contain it)
# Prove a key is live with a 200 from /key/list on the CT 116 host (the container has no curl):
curl -s -H "Authorization: Bearer <key>" http://127.0.0.1:4000/key/list | jq length
```
- Used for /key/generate, /key/delete, /key/list (GET), DB queries
- **Known violation (RESOLVED 2026-07-16):** Abiba's LITELLM_API_KEY was previously the master key.
It is now a dedicated agent key `sk-sxbphLvk1OU…` (vault secret `ABIBA_LITELLM_API_KEY`, alias `abiba-pi`).
+64 -26
View File
@@ -340,6 +340,36 @@ def check_gpu_ports():
# CHECK 3: Agent Gateway Liveness + Streaming (now covers all agents)
# ═══════════════════════════════════════════════════════════════════
def _ssh_retry(host, cmd, user="root", timeout=15, retry_timeout=25, label=""):
"""SSH with one retry at a longer timeout.
Returns (stdout_or_None, probe_failed_bool, fail_kind).
When probe_failed is True, fail_kind is one of: timeout, ssh-failed.
"""
import subprocess as _sp
def _attempt(tmo, conn_tmo):
try:
r = _sp.run(
["ssh", "-o", "StrictHostKeyChecking=no", "-o", f"ConnectTimeout={conn_tmo}",
f"{user}@{host}", cmd],
capture_output=True, text=True, timeout=tmo)
return r.stdout.strip() if r.returncode == 0 else None
except _sp.TimeoutExpired:
return "__timeout__"
except:
return None
result = _attempt(timeout, 8)
if result is None or result == "__timeout__":
kind = "timeout" if result == "__timeout__" else "ssh-failed"
prefix = f"{label} " if label else ""
print(f" probe-failed: {prefix}ssh {user}@{host} — {kind} (retrying at {retry_timeout}s…)")
result = _attempt(retry_timeout, 15)
if result is None or result == "__timeout__":
kind = "timeout" if result == "__timeout__" else "ssh-failed"
return None, True, kind
return result, False, None
def check_agents():
for name, agent in AGENTS.items():
host = agent.get("host")
@@ -355,42 +385,49 @@ def check_agents():
is_dsh = agent.get("runtime") == "dsh"
label = "DSH (DeepSeek Harness)" if is_dsh else "pi-only runtime"
since = "since 2026-08-27" if is_dsh else "since the harness purge"
live = ssh(host, "true", user=user)
print(f" {'✅' if live is not None else '❌'} {name}: {label} — "
f"no Hermes gateway {since} (CT {ct}, SSH {'OK' if live is not None else 'FAIL'})")
if live is None:
_fail(f"unreachable:{name}", name)
live, probe_failed, fail_kind = _ssh_retry(host, "true", user=user)
if probe_failed:
print(f" ❌ {name}: {label} — probe-failed: ssh {user}@{host} {fail_kind} "
f"(retried at 25s: also {fail_kind}) [CT {ct}]")
_fail(f"probe-failed:{name}:{fail_kind}", name)
else:
print(f" ✅ {name}: {label} — no Hermes gateway {since} "
f"(ssh {user}@{host} OK, CT {ct})")
continue
if not host or not user:
print(f" ⬜ {name} (CT {ct}): cannot SSH — skip liveness check")
continue
# Resolve the Hermes gateway PID once, before the report-only branch:
# the summary line below renders `pid`, and it used to be bound only in
# the report-only path — leaving it unbound on the abiba/koonimo path
# raised UnboundLocalError and crashed the whole check. Agents without
# a gateway get pid=?.
pid = ssh(host, "pgrep -f '[h]ermes_cli.main gateway run' | grep -v infisical | head -1", user=user)
if not pid:
pid = ssh(host, "pgrep -f '[h]ermes.*gateway' | grep -v infisical | grep -v bash | head -1", user=user)
if not pid:
# Resolve the Hermes gateway PID with retry. The probe target is
# explicit: ssh {user}@{host} pgrep -f hermes gateway.
pid, probe_failed, fail_kind = _ssh_retry(
host, "pgrep -f '[h]ermes_cli.main gateway run' | grep -v infisical | head -1", user=user)
if not pid and not probe_failed:
pid, probe_failed, fail_kind = _ssh_retry(
host, "pgrep -f '[h]ermes.*gateway' | grep -v infisical | grep -v bash | head -1", user=user)
if not pid and not probe_failed:
pid = "?"
# ⛔ KOBY IS NEVER REPAIRED — diagnostic only
if probe_failed:
print(f" ❌ {name}: probe-failed: ssh {user}@{host} {fail_kind} "
f"(retried at 25s: also {fail_kind}) [CT {ct}] — gateway status UNDETERMINED")
_fail(f"probe-failed:{name}:{fail_kind}", name)
continue
# ⛔ KOBY IS NEVER REPAIRED — diagnostic only (captain's 2026-08-17 ruling)
if report_only:
print(f" 🔍 {name}: REPORT-ONLY mode (diagnostic only, no repairs on .129)")
# Still check gateway status for reporting purposes
if pid == "?":
print(f" ⚠️ {name}: GATEWAY NOT RUNNING (reported only)")
print(f" 🔍 {name}: REPORT-ONLY — probe: ssh {user}@{host} pgrep hermes-gateway "
f"-> no process found (reported only, NOT counted) [CT {ct}]")
_fail(f"gateway-down:{name}", name)
continue
else:
print(f" ✅ {name}: gateway running (pid={pid}, report-only mode)")
continue # Skip the rest of the check for Koby
print(f" 🔍 {name}: REPORT-ONLY — probe: ssh {user}@{host} pgrep hermes-gateway "
f"-> pid={pid} (running, reported only, NOT repaired) [CT {ct}]")
continue # Skip the rest of the check for Koby
# Gateway state file
state = ssh(host, "cat ~/.hermes/gateway_state.json 2>/dev/null", user=user)
state, _, _ = _ssh_retry(host, "cat ~/.hermes/gateway_state.json 2>/dev/null", user=user)
if state:
try:
st = json.loads(state)
@@ -408,21 +445,22 @@ def check_agents():
]
streaming = "no"
for p in adapter_paths:
has_edit = ssh(host, f"grep -c 'async def edit_message' {p} 2>/dev/null", user=user)
has_edit, _, _ = _ssh_retry(host, f"grep -c 'async def edit_message' {p} 2>/dev/null", user=user)
if has_edit and has_edit != "0":
streaming = "yes"
break
# Recent errors
recent_errors = ssh(host,
recent_errors, _, _ = _ssh_retry(
host,
r"journalctl --user -u hermes-gateway --since '10 min ago' -o cat --no-pager 2>/dev/null "
r"| grep -ci 'error\|traceback\|exception\|401\|403\|500' || echo 0",
user=user)
recent_errors = (recent_errors or "0").strip().split("\n")[-1]
print(f" {'✅' if gw_state == 'running' and zulip == 'connected' else '⚠️'} "
f"{name}: gw={gw_state} zulip={zulip} streaming={streaming} "
f"errors_10m={recent_errors.strip() or '0'} pid={pid}")
f"{name}: probe: ssh {user}@{host} — gw={gw_state} zulip={zulip} "
f"streaming={streaming} errors_10m={recent_errors.strip() or '0'} pid={pid} [CT {ct}]")
# ═══════════════════════════════════════════════════════════════════
+95 -29
View File
@@ -35,7 +35,12 @@ def run_command(cmd, timeout=15):
return 1, "", str(e)
def probe_http(url, method="GET", bearer_token=None, data=None, timeout=10, follow_redirects=False):
"""Probe HTTP endpoint and return status code"""
"""Probe HTTP endpoint and return (status_code, failure_kind)
Returns:
(code, None) if successful or HTTP response received
(000, kind) if connection failed, where kind is 'timeout', 'refused', 'dns', etc.
"""
cmd = "curl -s -o /dev/null -w '%{http_code}' -m " + str(timeout)
if method == "POST":
cmd += " -X POST"
@@ -47,15 +52,47 @@ def probe_http(url, method="GET", bearer_token=None, data=None, timeout=10, foll
cmd += " -L"
cmd += " '" + url + "'"
rc, stdout, stderr = run_command(cmd, timeout)
if rc != 0 and "TIMEOUT" not in stderr:
return 000 # Connection failed
try:
rc, stdout, stderr = run_command(cmd, timeout)
if rc != 0:
# Determine failure kind from curl exit code
# curl exit codes: 28=timeout, 7=refused, 6=dns, 35=ssl, 52=empty
if rc == 28:
return (000, "timeout after " + str(timeout) + "s")
elif rc == 7:
return (000, "connection refused")
elif rc == 6:
return (000, "dns failure")
elif rc == 35:
return (000, "ssl error")
elif rc == 52:
return (000, "empty response")
else:
return (000, "curl exit " + str(rc))
return (int(stdout), None) if stdout.isdigit() else (000, "unparseable response")
except subprocess.TimeoutExpired:
return (000, "timeout after " + str(timeout) + "s")
def get_response_body(url, method="POST", bearer_token=None, data=None, timeout=30):
"""Get response body for 401/403 credential faults (truncated to 200 chars)"""
cmd = "curl -s -m " + str(timeout)
if method == "POST":
cmd += " -X POST"
if bearer_token:
cmd += " -H 'Authorization: Bearer " + bearer_token + "'"
if data:
cmd += " -H 'Content-Type: application/json' -d '" + data + "'"
cmd += " '" + url + "'"
return int(stdout) if stdout.isdigit() else 000
rc, stdout, stderr = run_command(cmd, timeout)
# Return first 200 chars, single line
body = stdout.replace('\n', ' ').replace('\t', ' ')[:200] if stdout else ""
return body
def check_liveliness():
"""Step 1: Liveliness probe"""
code = probe_http("http://" + BACKEND_HOST + "/litellm/health/liveliness")
code, _ = probe_http("http://" + BACKEND_HOST + "/litellm/health/liveliness")
return "Liveliness", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/health/liveliness)"
def check_containers():
@@ -79,32 +116,61 @@ def check_model_probes():
results = []
for model in ["gpu-dense", "gpu-vision", "strix-moe"]:
# Single-host aliases: 30s timeout
code = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
method="POST",
bearer_token=monitor_key,
data='{"model":"' + model + '","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
timeout=30)
# Single-host aliases: 30s timeout each
# gpu-dense (RTX 3090) may need long warmup/prefill - timeout is acceptable on cold-start
code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
method="POST",
bearer_token=monitor_key,
data='{"model":"' + model + '","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
timeout=30)
results.append((model, code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")"))
if code == 000 and failure_kind:
# Report probe failure with kind, do not assert a service verdict
results.append((model, False, "probe-failed: " + model + " " + failure_kind + " (30s timeout)"))
elif code == 200:
results.append((model, True, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")"))
elif code in (401, 403):
# Credential fault - capture body and key alias
body = get_response_body("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
method="POST",
bearer_token=monitor_key,
data='{"model":"' + model + '","messages":[{"role":"user","content":"health"}],"max_tokens":4}',
timeout=10)
# Resolve key alias
alias = "monitor-20260813" # Known from /etc/litellm-monitor.env on CT 116
results.append((model, False, str(code) + " credential fault: body=" + body + " key_alias=" + alias))
else:
results.append((model, False, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")"))
# Pool alias (syslog-auto): 60s timeout, retry once on 000
code = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
method="POST",
bearer_token=monitor_key,
data='{"model":"syslog-auto","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
timeout=60)
code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
method="POST",
bearer_token=monitor_key,
data='{"model":"syslog-auto","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
timeout=60)
if code == 000:
if code == 000 and failure_kind:
# Retry once with same timeout
time.sleep(1)
code = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
method="POST",
bearer_token=monitor_key,
data='{"model":"syslog-auto","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
timeout=60)
results.append(("syslog-auto", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=syslog-auto)"))
code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
method="POST",
bearer_token=monitor_key,
data='{"model":"syslog-auto","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
timeout=60)
if code == 000 and failure_kind:
results.append(("syslog-auto", False, "probe-failed: syslog-auto " + failure_kind + " (60s timeout, retry)"))
elif code in (401, 403):
body = get_response_body("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
method="POST",
bearer_token=monitor_key,
data='{"model":"syslog-auto","messages":[{"role":"user","content":"health"}],"max_tokens":4}',
timeout=10)
alias = "monitor-20260813"
results.append(("syslog-auto", False, str(code) + " credential fault: body=" + body + " key_alias=" + alias))
else:
results.append(("syslog-auto", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=syslog-auto)"))
else:
results.append(("syslog-auto", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=syslog-auto)"))
return results
@@ -146,18 +212,18 @@ def check_admin_key_list():
def check_github_status():
"""Step 3: GitHub status - 301 redirect is acceptable for status page"""
code = probe_http("https://status.github.com/api/status.json", timeout=15)
code, _ = probe_http("https://status.github.com/api/status.json", timeout=15)
# GitHub status API returns 301 redirect, which is expected behavior
return "GitHub Status", code == 301, str(code)
def check_prometheus():
"""Step 4: Prometheus health"""
code = probe_http("http://" + BACKEND_HOST + ":9090/-/healthy")
code, _ = probe_http("http://" + BACKEND_HOST + ":9090/-/healthy")
return "Prometheus", code == 200, str(code) + " (target: " + BACKEND_HOST + ":9090/-/healthy)"
def check_grafana():
"""Step 9: Grafana health"""
code = probe_http("http://" + BACKEND_HOST + ":3001/api/health")
code, _ = probe_http("http://" + BACKEND_HOST + ":3001/api/health")
return "Grafana", code == 200, str(code) + " (target: " + BACKEND_HOST + ":3001/api/health)"
def check_docker_stats():