Commit Graph
136 Commits
Author SHA1 Message Date
root 5ed6f8179c fix: add os import for HELPER_PCT_RUN.readable() check
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 3s
2026-09-14 14:14:15 +00:00
root b9322973ce fix disk-gc-scan: CWD independence (FIX A) and df column parsing (FIX B) 2026-09-14 14:05:06 +00:00
root d28df4f4de add disk-gc-scan.py: deterministic fleet disk probe with per-guest access methods 2026-09-14 13:14:10 +00:00
root 4e34b7a2a2 fix(hermes): wire all three contracts to reachability helper
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
The helper was correct but dead code - nothing called it. This commit:
1. Makes the helper runnable standalone: scripts/hermes-reachability-check.sh <host> <pattern> <path>
2. Wires all THREE contracts to it:
   - hermes-key-enforcement.prose.md
   - hermes-config-template.prose.md
   - hermes-agent-baseline.prose.md
3. Each contract now explicitly instructs to run the helper and interpret the three outcomes
4. States that the bug this replaces was deriving reachability from the remote grep's exit code

Files changed (4):
- scripts/hermes-reachability-check.sh (standalone mode added)
- hermes-key-enforcement.prose.md (reachability section added)
- hermes-config-template.prose.md (reachability section added)
- hermes-agent-baseline.prose.md (reachability section added)
2026-09-14 12:01:32 +00:00
root f9f6661dd5 fix(hermes): add shared reachability check helper
Bug: The reachability check used 'ssh ... grep ... || echo unreachable',
which conflated grep's 'no matches found' (exit 1) with SSH failure.
This caused clean hosts to be reported as unreachable.

Fix: Add scripts/hermes-reachability-check.sh with the pattern:
  out=$(ssh -o BatchMode=yes root@HOST "grep ... 2>/dev/null; true")
  if [ $? -ne 0 ]; then verdict="unreachable"
  elif [ -n "$out" ]; then verdict="violation: $out"
  else verdict="compliant"
  fi

This correctly distinguishes:
- SSH failure (connection/auth/route) -> unreachable
- SSH success + grep found matches -> violation
- SSH success + grep found nothing -> compliant

Evidence: 3 of 4 hosts (Tanko, Mumuni, Koonimo) were reported as
'unreachable' when they were actually compliant. Only Koby (.129)
has a real finding (plaintext key in state snapshot).

Used by: hermes-key-enforcement, hermes-config-template,
hermes-agent-baseline contracts.
2026-09-14 11:55:41 +00:00
root 05366bd58d Fix litellm-health-check.py robustness defects
1. Timeout fix for pool alias (syslog-auto):
   - Single-host aliases (gpu-dense, gpu-vision, strix-moe): 30s timeout
   - Pool alias (syslog-auto): 60s timeout, retry once on 000 before failing
   - Cold first request to pool alias can take ~13s; 10s was too short

2. Remove DEBUG prints from output:
   - Removed 'DEBUG: keylen=...' and 'DEBUG: response=...' lines
   - These leaked key inventory to status logs
   - Success output now shows only counts (e.g., 'Admin Key List: 10 keys')
2026-09-14 03:24:30 +00:00
root e94fadfa60 Add litellm-health-check.py: standardized health check script for LiteLLM fleet monitoring
Implements all 11 checks from litellm-health contract:
- Liveliness, Containers, Prometheus, Grafana health probes
- Model probes for gpu-dense, gpu-vision, strix-moe, syslog-auto
- Admin API key list (10 keys), GitHub status, Docker Stats metrics

Fixed quoting for SSH commands and response parsing (dict with 'keys' field).
Backend edge uses internal IP 192.168.68.116, not public URL.
Docker Stats fetched from CT 116 host itself (127.0.0.1:9324/metrics).

All 11 checks passing consistently.
2026-09-14 03:15:27 +00:00
root 6272d29978 no-mistakes(review): Fail closed on empty report-only exclusion list 2026-09-12 18:55:34 +00:00
root 4bd6132cf5 no-mistakes(review): Reject report-only exclusion entries lacking identity keys 2026-09-12 18:51:42 +00:00
root 6d65cba064 no-mistakes(review): Canonicalize guest identities in GC report-only gate 2026-09-12 18:47:17 +00:00
root de1428b4ae no-mistakes(review): Harden GC gate key aliases, fail closed, fix baseline 2026-09-12 18:42:36 +00:00
root 4ea2d0309f no-mistakes(review): Fix report-only gate, dedupe list, correct fleet map 2026-09-12 18:38:19 +00:00
abiba 7b8cc5f9ac fix(disk-gc): hard guest-level report-only gate for CT 111/.129; correct stale fleet map
CT 111 (tdunna, 192.168.68.129) is Theo's box and is report-only per the captain
(2026-08-17, re-confirmed 2026-09-10). The contract defined AMBER as 'GC scheduled
for next run' and its Execution loop called gc-executor for EVERY threat with no
guest-level exclusion - so a single AMBER reading there would have scheduled apt
clean / journal vacuum / log+tmp deletion against someone else's box. The only
marker was frontmatter report_only_agents, which names an AGENT while the scan unit
is a GUEST.

- Gate the Execution loop on a guest/host-keyed report_only_guests block (guest id,
  hostname and IP all match); an excluded guest is alerted and skipped, so no
  gc-executor call is constructed for it at any level.
- Carry the ruling in the contract body next to the loop, not only in frontmatter.
- scripts/disk-gc-plan.py: executable planner that reads the contract's
  authoritative exclusion block and emits the action plan; tests/ covers it.
- Correct the stale fleet map against pvesh /cluster/resources: CT 105 -> amdpve
  (was minipve), CT 111 -> storepve (was amdpve), add guests 118/119/120, and fix
  the '15 CTs' counts (20 guests: 17 LXC + 3 QEMU VMs).
- scripts/pct-run.sh: same stale map (105/111 wrong node, 120 missing) - this is
  why pct-run 111/105 failed.
- Fold in disk-gc-ct100-probe-gap-20260911: CT 100 verifiably works through
  pct-run now that the map is correct; documented that the scanner must probe it
  like any other guest, never via a local-only path (the scanner runs inside CT 100).
2026-09-12 18:33:02 +00:00
root 5f582e2c9c no-mistakes(review): fix duplicate-000 probe capture at assignment boundary 2026-09-12 14:23:35 +00:00
root b1b3b4c010 no-mistakes(review): fix kagentz A2A status logging, tests, and contract retirement 2026-09-12 14:19:13 +00:00
root 5d9b9847bc fix: remove kagentz Zulip adapter leg (code no longer exists)
- Remove adapter process check and restart logic
- Keep A2A probe (port 80, HTTP code check)
- The adapter code at /a0/usr/kagentz-zulip/ no longer exists
- Captain's ruling: Zulip communication with agent zero is not priority
2026-09-12 14:09:01 +00:00
root d29da3cc69 Make missing key file fail loudly instead of using dead key fallback
- Drop dead key literal from except block
- Report 'no-key-file' as check status when key file missing/unreadable
2026-09-12 12:49:15 +00:00
root 2ba1016ca8 Fix LiteLLM API key source and auth header
- Load API key from durable file /root/.abiba-workspace/secrets/litellm-key.txt (works in cron)
- Fix http_get to use Bearer token instead of Basic Auth for API endpoint check
- All 6 LiteLLM checks now pass (was 5/6)
2026-09-12 12:45:50 +00:00
abiba-bot 9c8637bcaf Merge pull request 'fix: correct Agent Zero A2A probe port in zulip-monitor.sh' (#76) from fix/a2a-port-20260911 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-11 22:17:37 +00:00
root aec62f7e77 fix: correct Agent Zero A2A probe from stale port 8001 to correct port 80
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
- Changed A2A probe from http://127.0.0.1:8001/.well-known/agent.json to
  http://127.0.0.1:80/a2a/ inside agent-zero container
- Port 8001 does not exist inside container (nothing listens there)
- Port 80 maps to external port 50080; returns 401 (auth-gated, alive by design)
- Updated A2A_URL in adapter.py restart command to use port 80 instead of 8001
- Verified: probe now returns 401 (auth-gated) instead of 000 (connection refused)
- Before: kagentz A2A reported DOWN on every run (false positive due to stale port)
- After: kagentz A2A reports ✅ A2A alive (auth-gated 401 = healthy)
2026-09-11 21:54:13 +00:00
kagentz-bot af27530edc Merge pull request 'fix(router): decommission legacy GPU router across contracts and fleet scripts' (#74) from fix/decommission-router-20260911 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 3s
2026-09-11 19:19:13 +00:00
agent-zero c26255f5ff fix(router): decommission legacy GPU router across contracts and fleet scripts
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
Router container, image, and config are purged on CT 116 (verified: no
container, no image, inference-harness-router:latest removed, :9000 free,
11 containers healthy, 7 models, live syslog-auto completion OK). Updates:

- gpu-fleet.prose.md: topology diagram rebuilt without the router tier
- gpu-monitor / gpu-self-heal / litellm-health / litellm-self-heal: router steps dropped
- infrastructure-control.prose.md: container inventory + litellm row corrected
- scripts/daily-infra-report.py, scripts/prose-ai-review.sh: host-scoped checks

Verified: 40/40 anchors matched uniquely, 0 U+FFFD across changed files,
daily-infra-report.py compiles, no stray harness-router references remain.
2026-09-11 14:20:18 -04:00
abiba-bot e97145c88f no-mistakes(review): re-sample systemd invocation each pass during token wait 2026-09-11 17:34:53 +00:00
abiba-bot 55f1208eb8 no-mistakes(review): scope dsh token to invocation; log nginx diagnostics 2026-09-11 17:30:33 +00:00
abiba-bot b9adf353ee no-mistakes(review): guard missing include, non-fatal pending reload, chmod stash 2026-09-11 17:24:12 +00:00
abiba-bot aeb66ea22d no-mistakes(review): fix pending-reload path, token modes, restart check 2026-09-11 17:17:26 +00:00
abiba-bot 288f74cf84 no-mistakes(review): reprobe dsh tokens; persist pending nginx reload on failure 2026-09-11 17:12:11 +00:00
abiba-bot c66671dbee no-mistakes(review): simplify dsh token selection and reload state machine 2026-09-11 17:07:00 +00:00
abiba-bot b80d3142aa no-mistakes(review): harden dsh token reload retry, legacy bypass, cookie verification 2026-09-11 16:59:55 +00:00
abiba-bot 266fa1f835 fix(dsh-web-auth): Authentik-gated :80 login + non-disruptive token capture
Correct the dsh-web authentication fix after parent correction:

- Remove the unauthenticated :8081 endpoint (0.0.0.0 bind with no
  auth_request = full Authentik bypass for the LAN). The script now removes
  /etc/nginx/sites-enabled/dsh.token automatically if it reappears.
- Put the login path inside the Authentik-gated :80 server block as
  location = /dsh-web-login; proxy to dsh-web with Host =
  tankodhs.sysloggh.net so the 30-day cookie is bound to the public
  authority, never to 127.0.0.1:3080.
- Isolate the rotating token in a generated include
  /etc/dsh-web/nginx-login.conf; reload nginx only when it changes.
- Replace the disruptive capture (systemctl stop/start dsh-web) with a
  non-disruptive read of the running service's journal, scoped to the
  current systemd invocation so a restarted process's stale token is never
  reused while the new banner is still pending.
- Keep x-dsh-task-board-proxy-token and Host $ak_origin_host intact in '/'.
- Document the corrected design (B4) in zulip-health.prose.md, v3.2.0.

Live-verified 2026-09-11: no auth bypass (302), :8081 refused (000), a
cookie minted before two dsh-web restarts still returns 200, the refreshed
token mints a fresh cookie, and the systemd ExecStartPost/timer refreshes
the token automatically without touching dsh-web.
2026-09-11 16:52:45 +00:00
abiba-bot 85ea1f4f3d no-mistakes(review): harden dsh token capture: loopback bind, atomic nginx config 2026-09-11 15:22:41 +00:00
abiba-bot b2a259fa23 fix: add dsh-web restart-persistent authentication
- Add capture-dsh-token.sh script that captures the dsh-web launch token
- Add login endpoint (/dsh-web-login on :8081) that mints 30-day auth cookie
- Document the authentication flow in zulip-health.prose.md (Platform B4)
- Cookie is authority-bound to 127.0.0.1:3080 with 30-day expiry
- After first login, subsequent requests use the cookie — no token required
2026-09-11 14:56:57 +00:00
root 287657a77a Merge remote-tracking branch 'origin/master' into update/docker-ecosystems-20260908
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-09-10 23:01:32 +00:00
root 21f9073e0b fix(daily-infra-report): use pve_probe_status in render + add regression tests (PR #64 round 2) 2026-09-10 23:01:18 +00:00
root 32fe7c0652 fix(daily-infra-report): complete Zulip nested-key fix + Proxmox port/auth fix (PR #64 review fix 1-2)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-09-10 22:03:48 +00:00
root 25cf2f5eef fix(daily-infra-report): read Zulip state from nested 'zulip' key (PR #64 fix 3a)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-09-10 20:55:17 +00:00
root 26f2301188 fix(pm2-self-heal): remove stray fragment so script parses (PR #64 fix 2)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-09-10 20:43:01 +00:00
abiba-bot 8a4ee4cf5a no-mistakes(review): Pin Mumuni removal behaviorally and fix digest import bug 2026-09-10 10:18:38 +00:00
abiba 952fca9c92 fix(monitoring): retire the Mumuni leg — stop probing her decommissioned .24 deployment
Captain ruling 2026-09-10: Mumuni moved off this host onto her own container
(kagentz CT 105 on minipve, 192.168.68.14, dedicated `hermes` user) and is
monitored from her side. This host must not monitor anything Mumuni.

The stale probes fired false alerts repeatedly:
  * scripts/zulip-monitor.sh ssh'd to root@192.168.68.24 for the
    decommissioned deployment's ~/.hermes/gateway_state.json, read "unknown"
    on every run, and posted a 🔴 "Mumuni (Hermes) Zulip state: unknown" DM +
    #agent-hub stream alert each cycle.
  * scripts/daily-infra-report.py published a matching "mumuni:unknown" row in
    every digest.

Changes:
  * zulip-monitor.sh: delete the "Platform B: Hermes (Mumuni)" leg and its
    notify; keep the Zulip-server, Platform A pi/Abiba (bridge), Platform B
    Tanko and Platform C Agent Zero legs. A comment records why the leg is
    retired so it is not re-added. Also fixes SC2155 so shellcheck is clean.
  * daily-infra-report.py: delete the .24 ~/.hermes/gateway_state.json agent
    probe, its hermes --version probe, and the now-dead mumuni render branch.
    Abiba (CT 100, its own .24 address) and Tanko legs unchanged; the PVE API
    token name is untouched.
  * zulip-health.prose.md (v3.1.0): drop the Mumuni-only B4 gateway-process,
    B5 heartbeat and B6 response-delivery steps and the stale 192.168.68.24
    references; state explicitly that Mumuni is not monitored from this host.
    Tanko/Agent-Zero/bridge steps retained.
  * agent-health-check.py: correct the v2 changelog roster comment that still
    placed mumuni at .24/CT100. No behavior change — the mumuni probe was
    already absent from the AGENTS dict; v5 changelog notes the correction.

Tests: tests/test_mumuni_monitor_removal.py pins the removal structurally and
behaviorally — the shipped zulip-monitor.sh is run in a sandbox (only its LOG
constant rewritten) with stub ssh/curl on PATH; the ssh stub records every
target host, so "never reaches .24" and "no Mumuni notify even on the alert
path" are asserted from observed behavior. A mutation check (re-inject the old
leg) fails the suite, so the guarantee is not vacuous.
2026-09-10 10:11:19 +00:00
abiba-bot 88f27e75ed no-mistakes(document): Refresh stale health-check version, provenance, and lint evidence 2026-09-10 02:01:20 +00:00
abiba-bot bbdf6c1249 no-mistakes(review): Restore quiet-mode silence; keep provenance on failure alert 2026-09-10 01:52:38 +00:00
abiba-bot 1974959cc9 no-mistakes(review): Gate wrapper checks on executed infisical; scope liveness guide 2026-09-10 01:46:06 +00:00
abiba-bot c59c9fb174 no-mistakes(review): Ignore commented infisical paths; normalize probe-model tests 2026-09-10 01:41:27 +00:00
abiba-bot 194e256ac5 no-mistakes(review): Harden health-check provenance, infisical verification, report-only JSON, tests 2026-09-10 01:35:44 +00:00
abiba-bot c66d9c1e20 fix: probe-drift round 2 — stale monitoring expectations (health check, PVE API, GPU port 80, report provenance)
Second probe-drift correction pass after #65/#66/#68. All four legs were stale
consumer expectations, not live faults.

1. scripts/agent-health-check.py (v4)
   - abiba declared pi-only runtime (harness purge): Hermes-era gateway, config
     and wrapper legs are skipped instead of failing.
   - koby declared report_only (captain ruling 2026-08-17, Rule 17): every koby
     leg is detected and reported, never counted as a fleet failure or repaired.
   - koby's CT 111 mapping corrected to storepve (.6); the old amdpve mapping
     made `pct status 111` fail and read as ct-unreachable.
   - wrapper check no longer FAILs .env-based wrappers that legitimately never
     invoke infisical (koonimo).
   - keys load in main() (load_agent_keys) so the module is importable/testable.
   - every run prints absolute execution provenance (script + cwd), in the header
     and in --json.
   Before: 6 FAILURE(S). After: 0 failures, koby reported read-only.

2. infrastructure-monitoring.prose.md
   - PVE API probe repointed from CT 116 (no pveproxy, 000) to the five real
     nodes on https://<node>:8006/api2/json/version, all 401 = alive.
   - any-HTTP-response liveness rule added (401/3xx alive; 000/timeout = DOWN).
   - LiteLLM health documented as 301 -> /litellm/health/liveliness, not bare 200.

3. gpu-monitor.prose.md
   - GPU health probes on :8080 (or router /health/unified); bare port 80 on a
     GPU host is forbidden (no listener -> false DEGRADED).
   - router /health/unified 301 -> /gpu/gpu-data documented as alive.
   - port-discipline + liveness rule + direct-fallback execution step.

4. Report provenance (all contracts)
   - docs/AUTHORING-GUIDE.md documents the rule; scripts/prose-lint.sh enforces
     that any **Report format** contract states an absolute path (pwd -P).
   - provenance added to gpu-monitor, infrastructure-monitoring, proxmox-monitor.

Tests: tests/test_probe_drift.py (23 passed); prose-lint.sh clean + shellcheck
clean; CI frontmatter validation passes. Evidence with per-leg before/after and
absolute paths: docs/probe-drift-round2-evidence.md.
2026-09-10 01:23:54 +00:00
root 7ed2e4e923 fix(zulip-monitor): Abiba leg reads nested zulip.connected; probe failures never restart
The monitor's Abiba leg parsed the :9200/health payload at the top level
(d.get('connected',False)) while the pi Zulip extension serves connection
state NESTED at zulip.connected / zulip.last_error (verified against the
extension's startHealthServer handler and the live endpoint). PI_CONNECTED
was therefore always False and every run took the DISCONNECTED path,
restarting a healthy bot: pm2 restarts=8 with the process created
2026-09-09T09:35:09Z, four ❌ Abiba verdicts today (04:23/05:35/06:55/09:35
UTC) and zero ✅, while the Zulip server answered HTTP 200 and the bot kept
heartbeating. The watchdog was the fault, not the connection.

Changes (Abiba leg only; every other leg byte-identical):
- Read the real nested shape: zulip.connected, zulip.last_error and
  zulip.messages_processed. The old retry_count branch is DROPPED — the
  payload exposes no retry counter (the extension keeps retryCount internal
  and never serialises it), so the branch is fabricated and cannot stay.
- Fail-safe restart decision: a fetch error (HTTP 000), non-2xx response,
  empty/unparseable body, or payload missing a boolean zulip.connected is a
  clearly-labelled PROBE FAILURE (🟠 alert + ⚠️ log line naming reason and
  HTTP code) and NEVER calls pm2 restart. pm2 restart runs only on
  affirmative zulip.connected=false (🔴/❌ path unchanged in wording).
  connected=true with last_error keeps the degraded 🟡 warn-no-restart path.

Regression test (new tests/zulip-monitor-abiba.sh, 36 checks): extracts the
real Abiba leg from between the # -- abiba-leg-start/-end markers in the
shipped script and executes it verbatim with stubbed curl/notify/pm2 against
fixtures of the real payload shape — asserts connected=true → no restart,
connected=false → restart, and empty/garbage/missing-key/non-boolean/HTTP
000/HTTP 500 bodies → probe-failure alert with zero restarts. The suite
fails loudly on the pre-fix base and on a marker-intact top-level-parse
variant, so this class of bug cannot return silently.
2026-09-09 09:51:46 +00:00
root 7ff7ce5b33 no-mistakes(document): docs: fix stale replaced-claim and Telegram header comment
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
2026-09-09 03:27:09 +00:00
root 947e8b24e0 fix(zulip-monitor): stream alert body now carries real content, plain & params, WARN on delivery failure
The #agent-hub / zulip-health stream post in notify() encoded content with
python quote(str()) (always empty), so every stream alert posted empty
content and Zulip rejected it silently behind '|| true'. The merged body
also used backslash-escaped ampersands inside double quotes, which curl
transmits literally (type=stream\ -> Zulip 400 'Invalid type').

Stream post now percent-encodes the real message text by piping it through
the encoder (locale-proof: quote_from_bytes on stdin.buffer), uses plain &
separators, and on failure appends one WARN line to the monitor log with
the curl exit status instead of silently swallowing it. Delivery failure
stays non-fatal and unretried. DM path, probes, alive rule, exit codes and
log format unchanged.
2026-09-09 02:47:44 +00:00
abiba-bot 95b4a0e6b0 Merge pull request 'fix(agent-health-check): correct unit names, abiba key path, pid crash, uptime math (script-only)' (#65) from fm/agent-health-probe-repoint-20260907 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-09 01:39:29 +00:00
root 1b8186f6b9 no-mistakes(document): docs: reword Tanko .122 SSH dependency claims in zulip-health
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-09-08 13:00:05 +00:00