Commit Graph
354 Commits
Author SHA1 Message Date
abiba-bot 4f59b82404 no-mistakes(document): Refresh stale Mumuni roster docs and registry metadata
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-09-10 10:26:00 +00:00
abiba-bot 8a4ee4cf5a no-mistakes(review): Pin Mumuni removal behaviorally and fix digest import bug 2026-09-10 10:18:38 +00:00
abiba 952fca9c92 fix(monitoring): retire the Mumuni leg — stop probing her decommissioned .24 deployment
Captain ruling 2026-09-10: Mumuni moved off this host onto her own container
(kagentz CT 105 on minipve, 192.168.68.14, dedicated `hermes` user) and is
monitored from her side. This host must not monitor anything Mumuni.

The stale probes fired false alerts repeatedly:
  * scripts/zulip-monitor.sh ssh'd to root@192.168.68.24 for the
    decommissioned deployment's ~/.hermes/gateway_state.json, read "unknown"
    on every run, and posted a 🔴 "Mumuni (Hermes) Zulip state: unknown" DM +
    #agent-hub stream alert each cycle.
  * scripts/daily-infra-report.py published a matching "mumuni:unknown" row in
    every digest.

Changes:
  * zulip-monitor.sh: delete the "Platform B: Hermes (Mumuni)" leg and its
    notify; keep the Zulip-server, Platform A pi/Abiba (bridge), Platform B
    Tanko and Platform C Agent Zero legs. A comment records why the leg is
    retired so it is not re-added. Also fixes SC2155 so shellcheck is clean.
  * daily-infra-report.py: delete the .24 ~/.hermes/gateway_state.json agent
    probe, its hermes --version probe, and the now-dead mumuni render branch.
    Abiba (CT 100, its own .24 address) and Tanko legs unchanged; the PVE API
    token name is untouched.
  * zulip-health.prose.md (v3.1.0): drop the Mumuni-only B4 gateway-process,
    B5 heartbeat and B6 response-delivery steps and the stale 192.168.68.24
    references; state explicitly that Mumuni is not monitored from this host.
    Tanko/Agent-Zero/bridge steps retained.
  * agent-health-check.py: correct the v2 changelog roster comment that still
    placed mumuni at .24/CT100. No behavior change — the mumuni probe was
    already absent from the AGENTS dict; v5 changelog notes the correction.

Tests: tests/test_mumuni_monitor_removal.py pins the removal structurally and
behaviorally — the shipped zulip-monitor.sh is run in a sandbox (only its LOG
constant rewritten) with stub ssh/curl on PATH; the ssh stub records every
target host, so "never reaches .24" and "no Mumuni notify even on the alert
path" are asserted from observed behavior. A mutation check (re-inject the old
leg) fails the suite, so the guarantee is not vacuous.
2026-09-10 10:11:19 +00:00
abiba-bot b1462f3e79 Merge pull request 'fix(monitoring): probe-drift round 2 — gpu port, PVE liveness, report-only visibility, report provenance' (#70) from fm/probe-drift-round2-20260909 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Failing after 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Skipped
2026-09-10 02:17:01 +00:00
abiba-bot 19ed186d0a no-mistakes(document): Scope gpu-monitor liveness rule to bare-200 probes
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
2026-09-10 02:04:07 +00:00
abiba-bot 88f27e75ed no-mistakes(document): Refresh stale health-check version, provenance, and lint evidence 2026-09-10 02:01:20 +00:00
abiba-bot bbdf6c1249 no-mistakes(review): Restore quiet-mode silence; keep provenance on failure alert 2026-09-10 01:52:38 +00:00
abiba-bot 1974959cc9 no-mistakes(review): Gate wrapper checks on executed infisical; scope liveness guide 2026-09-10 01:46:06 +00:00
abiba-bot c59c9fb174 no-mistakes(review): Ignore commented infisical paths; normalize probe-model tests 2026-09-10 01:41:27 +00:00
abiba-bot 194e256ac5 no-mistakes(review): Harden health-check provenance, infisical verification, report-only JSON, tests 2026-09-10 01:35:44 +00:00
abiba-bot c66d9c1e20 fix: probe-drift round 2 — stale monitoring expectations (health check, PVE API, GPU port 80, report provenance)
Second probe-drift correction pass after #65/#66/#68. All four legs were stale
consumer expectations, not live faults.

1. scripts/agent-health-check.py (v4)
   - abiba declared pi-only runtime (harness purge): Hermes-era gateway, config
     and wrapper legs are skipped instead of failing.
   - koby declared report_only (captain ruling 2026-08-17, Rule 17): every koby
     leg is detected and reported, never counted as a fleet failure or repaired.
   - koby's CT 111 mapping corrected to storepve (.6); the old amdpve mapping
     made `pct status 111` fail and read as ct-unreachable.
   - wrapper check no longer FAILs .env-based wrappers that legitimately never
     invoke infisical (koonimo).
   - keys load in main() (load_agent_keys) so the module is importable/testable.
   - every run prints absolute execution provenance (script + cwd), in the header
     and in --json.
   Before: 6 FAILURE(S). After: 0 failures, koby reported read-only.

2. infrastructure-monitoring.prose.md
   - PVE API probe repointed from CT 116 (no pveproxy, 000) to the five real
     nodes on https://<node>:8006/api2/json/version, all 401 = alive.
   - any-HTTP-response liveness rule added (401/3xx alive; 000/timeout = DOWN).
   - LiteLLM health documented as 301 -> /litellm/health/liveliness, not bare 200.

3. gpu-monitor.prose.md
   - GPU health probes on :8080 (or router /health/unified); bare port 80 on a
     GPU host is forbidden (no listener -> false DEGRADED).
   - router /health/unified 301 -> /gpu/gpu-data documented as alive.
   - port-discipline + liveness rule + direct-fallback execution step.

4. Report provenance (all contracts)
   - docs/AUTHORING-GUIDE.md documents the rule; scripts/prose-lint.sh enforces
     that any **Report format** contract states an absolute path (pwd -P).
   - provenance added to gpu-monitor, infrastructure-monitoring, proxmox-monitor.

Tests: tests/test_probe_drift.py (23 passed); prose-lint.sh clean + shellcheck
clean; CI frontmatter validation passes. Evidence with per-leg before/after and
absolute paths: docs/probe-drift-round2-evidence.md.
2026-09-10 01:23:54 +00:00
abiba-bot 532250b017 Merge pull request 'fix(zulip-monitor): read nested zulip.connected; probe failures alert and never restart' (#69) from fm/zulip-monitor-false-selfheal-20260909 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
2026-09-09 11:09:14 +00:00
root f57923b4fa no-mistakes(document): test shellcheck hygiene; flagged monitor contract doc drift
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Failing after 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
2026-09-09 10:26:35 +00:00
root 7ed2e4e923 fix(zulip-monitor): Abiba leg reads nested zulip.connected; probe failures never restart
The monitor's Abiba leg parsed the :9200/health payload at the top level
(d.get('connected',False)) while the pi Zulip extension serves connection
state NESTED at zulip.connected / zulip.last_error (verified against the
extension's startHealthServer handler and the live endpoint). PI_CONNECTED
was therefore always False and every run took the DISCONNECTED path,
restarting a healthy bot: pm2 restarts=8 with the process created
2026-09-09T09:35:09Z, four ❌ Abiba verdicts today (04:23/05:35/06:55/09:35
UTC) and zero ✅, while the Zulip server answered HTTP 200 and the bot kept
heartbeating. The watchdog was the fault, not the connection.

Changes (Abiba leg only; every other leg byte-identical):
- Read the real nested shape: zulip.connected, zulip.last_error and
  zulip.messages_processed. The old retry_count branch is DROPPED — the
  payload exposes no retry counter (the extension keeps retryCount internal
  and never serialises it), so the branch is fabricated and cannot stay.
- Fail-safe restart decision: a fetch error (HTTP 000), non-2xx response,
  empty/unparseable body, or payload missing a boolean zulip.connected is a
  clearly-labelled PROBE FAILURE (🟠 alert + ⚠️ log line naming reason and
  HTTP code) and NEVER calls pm2 restart. pm2 restart runs only on
  affirmative zulip.connected=false (🔴/❌ path unchanged in wording).
  connected=true with last_error keeps the degraded 🟡 warn-no-restart path.

Regression test (new tests/zulip-monitor-abiba.sh, 36 checks): extracts the
real Abiba leg from between the # -- abiba-leg-start/-end markers in the
shipped script and executes it verbatim with stubbed curl/notify/pm2 against
fixtures of the real payload shape — asserts connected=true → no restart,
connected=false → restart, and empty/garbage/missing-key/non-boolean/HTTP
000/HTTP 500 bodies → probe-failure alert with zero restarts. The suite
fails loudly on the pre-fix base and on a marker-intact top-level-parse
variant, so this class of bug cannot return silently.
2026-09-09 09:51:46 +00:00
abiba-bot 91d16d2693 Merge pull request 'fix(zulip-monitor): stream alert body carries real content; WARN on delivery failure' (#68) from fm/zulip-monitor-stream-body-20260909 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-09 04:21:37 +00:00
root 7ff7ce5b33 no-mistakes(document): docs: fix stale replaced-claim and Telegram header comment
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
2026-09-09 03:27:09 +00:00
root 36f218e255 chore(agents): replace CLAUDE.md symlink with canonical @AGENTS.md pointer file
fm-ensure-agents-md.sh converted the tracked AGENTS.md symlink into the
real-file pointer form, removing the dangling-symlink hazard.
2026-09-09 02:48:17 +00:00
root 947e8b24e0 fix(zulip-monitor): stream alert body now carries real content, plain & params, WARN on delivery failure
The #agent-hub / zulip-health stream post in notify() encoded content with
python quote(str()) (always empty), so every stream alert posted empty
content and Zulip rejected it silently behind '|| true'. The merged body
also used backslash-escaped ampersands inside double quotes, which curl
transmits literally (type=stream\ -> Zulip 400 'Invalid type').

Stream post now percent-encodes the real message text by piping it through
the encoder (locale-proof: quote_from_bytes on stdin.buffer), uses plain &
separators, and on failure appends one WARN line to the monitor log with
the curl exit status instead of silently swallowing it. Delivery failure
stays non-fatal and unretried. DM path, probes, alive rule, exit codes and
log format unchanged.
2026-09-09 02:47:44 +00:00
abiba-bot 95b4a0e6b0 Merge pull request 'fix(agent-health-check): correct unit names, abiba key path, pid crash, uptime math (script-only)' (#65) from fm/agent-health-probe-repoint-20260907 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-09 01:39:29 +00:00
abiba-bot 028f276be4 Merge pull request 'fix(zulip-health): tanko probe via amdpve pct + loopback-any-HTTP alive rule + zulip-monitor.sh stale-path fix' (#66) from fm/zulip-health-contract-tanko-probe-via-am-65 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-09 01:39:16 +00:00
mumuni-bot 17751e24d1 Merge pull request 'provision: register okyeame-memory-audit cron job id' (#67) from provision/okyeame-cron-jobs into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
2026-09-08 16:29:14 +00:00
mumuni-bot 79eeb457fc provision: register okyeame-memory-audit cron job id (kagentz 2026-09-08)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
2026-09-08 16:22:51 +00:00
root 1b8186f6b9 no-mistakes(document): docs: reword Tanko .122 SSH dependency claims in zulip-health
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-09-08 13:00:05 +00:00
root 9dd0cb18d5 no-mistakes(document): docs: align zulip-health tanko snapshot schema to dsh-web probes 2026-09-08 12:49:34 +00:00
root 4ac60e14a3 no-mistakes(review): fix zulip-monitor Tanko probe fallback-echo contamination in down detection 2026-09-08 12:38:45 +00:00
root 2e0b737f2d revert: drop co-gated prose edits (gpu-fleet, infrastructure-control) — script-only scope per captain decision
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
The no-mistakes document step had auto-updated the .8 systemd unit name
(llama-server -> llama-chat-api.service) in gpu-fleet.prose.md and
infrastructure-control.prose.md. Those contracts are co-gated with another
agent and cannot be amended inside a script-scope task (firstmate decision
2026-09-08, key prose-doc-step-scope); the unit-name doc sync is separate
follow-up work. This restores both files to their origin/master content.
The litellm-self-heal.prose.md agent-health-check v3/env.sh description
update (same document commit) is intentional and kept.
2026-09-08 12:30:13 +00:00
root bbe9ee533f no-mistakes(document): Docs updated for .8 unit repoint and abiba env.sh sourcing 2026-09-08 12:18:14 +00:00
root 4684ee64e0 no-mistakes(review): Align zulip-monitor Tanko probe and B2/B3 criteria to dsh-web 2026-09-08 12:10:58 +00:00
root af3d364242 no-mistakes(review): Bracket pgrep patterns to stop ssh wrapper self-match 2026-09-08 12:04:01 +00:00
root 2716e55c16 fix(agent-health): repoint .8 GPU unit, fix pid UnboundLocalError, source abiba key from env.sh
- GPU unit repoint verified live 2026-09-08: .8 rtx3090 probes
  llama-chat-api.service (stale llama-server unit read inactive -> false
  UNREACHABLE for a healthy process); .110 keeps llama-server.service
  (ocu-llm VM), .15 keeps strix-server.service. is-active no longer
  swallowed as SSH failure (|| true).
- check_agents: bind pid before the summary f-string so the non-report-only
  path (abiba/koonimo) no longer raises UnboundLocalError (line 305 crash).
- abiba key leg: read LITELLM_API_KEY from /root/.pi/agent/env.sh (#735
  moved creds out of shared /root/.bashrc); 'abiba NO KEY' gone on healthy
  setup.
2026-09-08 11:08:28 +00:00
root 5313e6b9ba fix: tanko zulip-health probes via amdpve pct exec (dsh-web loopback :3080 + public-URL fallback) 2026-09-08 10:26:56 +00:00
agent-zero bfdff13ae7 docs(infrastructure-update): point weekly trigger at scheduler task qSOOVzsU
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Failing after 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
Weekly Sunday 03:00 America/New_York run is now implemented as Agent Zero
scheduled task 'weekly-fleet-docker-update' (qSOOVzsU) with built-in
post-update service verification and old-image pruning.
2026-09-08 05:27:38 -04:00
abiba-bot 0e4eda0abb Merge pull request #63: probe alignment across monitoring contracts (proxmox/docker-stats 9324, pve-exporter 9221 localhost-only; gpu-monitor, zulip-health)
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-08 08:54:30 +00:00
root 801de0a25c fix: correct proxmox-monitor probe endpoints to use localhost for exporters
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
Docker Stats and PVE Exporter are bound to 127.0.0.1 (localhost-only), not 0.0.0.0.
They must be probed from .116 via SSH. Prometheus and Grafana remain 0.0.0.0 (LAN-reachable).

Fixes false alarm from remote probe showing connection-refused (by design for localhost binds).
2026-09-08 08:41:21 +00:00
agent-zero 8bf32f6f0f update(infrastructure-update): v1.3.0 — full Docker ecosystem coverage from 2026-09-08 live run
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 3s
- Add 5 Docker ecosystems: docker-vm .7, CT 116 .116, CT 117 (zulip+jitsi),
  hwpve .11 (Authentik), NetBird VPS 72.61.0.17
- Wave 3: add trove-test, docker-stats, monitoring stack, Zulip (pct exec + compose
  recreate preserves zulip_default network), Jitsi, Authentik, NetBird
- Firecrawl test is now POST /v1/search (GET / returns 404 by design)
- Add Authentik 302 + NetBird dashboard 200 checks
- Document harness-litellm 3-5 min cold start after recreate (verified 2026-09-08)
- Add hwpve + VPS compose files to config backup list; success criteria covers 5 hosts
2026-09-08 03:23:03 -04:00
root 36ae464f59 fix: probe alignment across multiple monitoring contracts
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
- infra-monitoring: Added Router+LiteLLM+PVE API probes with expected HTTP codes (200/401/500)
- gpu-monitor: Added explicit check-health section with Router+LiteLLM probe commands
- proxmox-monitor: Added check-health section for Prometheus/Grafana/exporters (all bound to 0.0.0.0)
- zulip-health: Fixed A2A port from :8001 to :50080 and added auth-gated expectation (401)

Closes: infra-monitoring-probe-alignment, gpu-monitor-probe-alignment, proxmox-monitor-probe-alignment, zulip-health-probe-alignment
2026-09-08 07:20:41 +00:00
root a3e97ce72b fix: add Router+LiteLLM+PVE API probes to infrastructure-monitoring check-health
- Added Router health probe (http://192.168.68.116/health) — expected 200
- Added LiteLLM health probe (http://192.168.68.116/litellm/health) — expected 200
- Added PVE API probe (https://192.168.68.116:8006/api2/json) — 401 expected for unauthenticated
- Clarified that 401 means API is up, 000 means unreachable, 500 means API down

Closes: infra-monitoring-probe-alignment
2026-09-08 07:20:41 +00:00
root cc5fe0991c Move Zulip bot creds from inline to env/vault reference
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
The abiba-bot@chat.sysloggh.net:KEY is now sourced from
/etc/litellm-monitor.env as ZULIP_USER and ZULIP_BOT_KEY.
2026-09-08 06:18:21 +00:00
mumuni-bot dc78604360 Merge pull request 'feat: litellm-client-timeouts — standard client timeout/retry policy for all agents' (#61) from fix/litellm-client-timeouts into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Failing after 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Skipped
2026-09-08 06:14:03 +00:00
mumuni-bot 7bbf148778 feat: litellm-client-timeouts contract — standard client timeout/retry policy
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 10s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
Root cause of the 2026-09-06 incident: 157 client-abandoned 408s against a
backend that was succeeding at 20-70s/call once clients stopped giving up.
Measured grounding: syslog-auto 28.8s avg / 25.4s TTFT over 888 calls;
nginx already allows 600s (Rule 5); the gap was entirely client-side.

Values: >=300s primary path, >=120s gpu-dense delegation, 2 retries with
15s/45s backoff, identified probes at 30s/hourly, batch jobs chunked +
day-scheduled. Verified live: Prometheus metrics, 0.56s syslog-auto probe,
gpu-fleet /health/unified topology.
2026-09-08 06:07:43 +00:00
mumuni-bot 3a25c7cce5 Merge pull request 'fix: restore agent-health-check AGENTS roster + land missed Tanko-DSH rows' (#60) from fix/tanko-dsh-reland into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
2026-09-07 07:34:34 +00:00
mumuni-bot 0aa0ea4906 fix: restore agent-health-check AGENTS roster + land missed Tanko-DSH rows
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2dfc3e1 (PR #50's out-of-band squash) overwrote agent-health-check.py with
a version whose AGENTS dict was empty — the health check silently skipped
every agent since. Restore the pre-stomp roster verbatim: tanko (runtime:
dsh), abiba, koby, koonimo.

Also land two Tanko-DSH rows from PR #52 that the earlier merge missed:
- hermes-zulip-plugin live-state table (plugin retired on CT112)
- zulip-self-heal restart-services table (restart via DSH service, not
  hermes gateway restart)

Mumuni rows from the branch were NOT restored: they carry pre-migration
CT100/.24 data, superseded by PRs #54-56 (Mumuni now on kagentz CT105/.14).

Closes the reland of PR #52's substance; the stale branch head stays closed.
2026-09-07 07:31:34 +00:00
mumuni-bot 19821ed6b5 Merge pull request 'fix: remove duplicated ## Execution block in infrastructure-monitoring' (#59) from fix/infra-monitoring-dedup into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 10s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 14s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-07 07:12:19 +00:00
mumuni-bot 403fbcdd9f fix: remove duplicated ## Execution block in infrastructure-monitoring
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 15s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
PR #50's content was squash-merged out-of-band (2dfc3e1, 2026-08-30) while
the PR itself stayed open, double-applying the check-health/Execution
section. This drops the second copy; keeps one canonical
Execution -> check-health -> Phases 1-4 -> Verification Commands flow.
2026-09-07 07:08:49 +00:00
mumuni-bot 0753f38cf9 Merge pull request 'fix: rename agent-zero fix-summary out of contract scan path' (#58) from fix/agent-zero-summary-not-a-contract into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-03 05:15:58 +00:00
mumuni-bot 274596fdd1 fix: rename agent-zero fix-summary out of contract scan path
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
agent-zero-fix-summary.prose.md has no YAML frontmatter (kind/name/description),
so CI validate fails on every master push that touches it (runs 226-228).
It is a dated session fix-log, not a contract - the durable knowledge already
lives in agent-zero-openrouter-key.prose.md (kind: function, referenced from
the summary). Rename to .md so validate/lint (which scan only *.prose.md)
stop rejecting master; matches repo root docs like cron-prompts-review.md.

Verified live: local validate repro PASS (36 files), prose-lint PASS at
baseline 15 warnings, no new warnings. Intentionally NOT changed: no content
edits, no other files, no contract frontmatter added to non-contract docs.
2026-09-03 05:02:41 +00:00
mumuni-bot 8c4df63db4 docs: add critical fix for .env.clobbered-by-new-image
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Failing after 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Skipped
- Root cause: Agent Zero loaded key from .env.clobbered-by-new-image
- Not main .env file!
- Both files must be updated when changing OpenRouter key
- Added to prose contract for future reference
2026-09-01 21:24:06 +00:00
mumuni-bot 79a1d22c99 docs: note that full container restart was required for key change
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Failing after 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Skipped
- run_ui process caches API keys in memory
- supervisorctl restart run_ui was insufficient
- Full container restart (docker restart agent-zero) required
2026-09-01 20:37:11 +00:00
mumuni-bot 782831f548 docs: add vault sync status to agent-zero key contract
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Failing after 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Skipped
- Note that vault sync is pending (service token not on kagentz)
- Key is stored in /home/hermes/syslog/agent-zero-keys.env as fallback
2026-09-01 20:33:05 +00:00
mumuni-bot 143dd3f16b docs: add Agent Zero fix summary
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Failing after 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Skipped
- Documents root cause analysis (OpenRouter 401 + Telegram conflicts)
- Details fixes applied (key update, telegram plugin disable, restart)
- Includes current state verification
- Provides next steps and prevention measures
2026-09-01 18:01:11 +00:00