Folds into the same branch as the Zulip delivery change, as instructed.
FINDING 1 - the Firecrawl probe path was wrong; the service is fine.
scripts/daily-infra-report.py probed http://192.168.68.7:3002/health, which
Firecrawl does not serve - it 404s. The root answers 200 with
{"message":"Firecrawl API",...}. Live 2026-09-26:
Firecrawl(/) -> 200
Firecrawl(/health) -> 404 <- what the report was showing
The probe is now the root, which is its liveness endpoint.
FINDING 2 - the Network Endpoints classification was wrong twice over.
It read: color = green if code in (200,302,401) else (yellow if code >= 400
else red). Two defects:
(a) it ignored the fleet's own probe policy, codified 2026-09-14 in the
monitoring contracts: ANY HTTP status proves the service answered, so the
service is ALIVE, and only a failed CONNECTION is a failed probe. A 404
from a wrong path is not a service fault.
(b) was a STRING comparison. Reproduced: 301 -> red (a live
redirect rendered as a failure), 404 -> yellow, 500 -> yellow (a real
server error softened to a warning).
Replaced with classify_endpoint(), which returns:
any 2xx/3xx/4xx -> green 'alive' (code still shown)
5xx -> yellow 'server error' (kept distinct from 4xx, as asked)
000/no answer -> red 'no connection'
Verified against the live endpoints after the change:
Gitea 200, Authentik 302, Zulip 302, Pulse 200, Proxmox 200, SearXNG 200,
Firecrawl 200 - all green/alive; the only red state is a genuine no-connection.
ALSO CHECKED, as asked: scripts/search-stack-check.py does NOT depend on the
wrong route. It POSTs to {FIRECRAWL_URL}/v1/scrape with formats=[markdown], and
that path really works - live POST returned HTTP 200 and 180 chars of markdown
for https://example.com. It was never using /health.
prose-lint: PASSED.
Captain's decision 2026-09-26, clarified the same day: the digest is delivered to
his Zulip DM (user id 9) from abiba-bot as an HTML FILE - an attachment, not HTML
rendered in the message body and not a Markdown translation of it. Closes
daily-digest-mail-transport-20260921; the Google dependency is gone (no SMTP, no
EMAIL_PASSWORD, no app password, nothing to rotate).
WHAT CHANGES
* scripts/daily-infra-report.py: send_email() is replaced by send_zulip(), which
writes the styled dashboard to /var/log/daily-infra-report/infra-report-<ts>.html,
uploads it via POST /api/v1/user_uploads, then posts a SHORT Markdown pointer to
user 9. The message body carries subject, top-line status and the attachment
link; it does not reproduce the report.
* the 10,000-character cap is irrelevant here - it bounds message TEXT only, and
the report travels as a file, so nothing is shrunk to fit.
* the key is abiba-bot's, already on the execution host at
/root/.pi/agent/extensions/zulip/.env (mode 600). No vault entry was added:
under the auth-keys charter that is a captain decision.
* daily-health-digest.prose.md -> v2.0.0 and contract-registry.yaml updated:
transport, healthy/degraded definitions, and exit codes now match observed
behaviour. There is NO degraded delivery leg any more - delivery is the only
output path, so a missing or rejected key is a real failure (exit 1).
* queued defect folded in: a failed delivery used to print only the transport
error while the report body never surfaced. Now the HTML is printed to stdout
AND persisted on every failure, and the message names which step failed.
EVIDENCE (all against the live stack)
* real send: message id 86221 to user 9, attachment 16208 bytes at
/user_uploads/2/45/m1cQesBFV78BGeNY2lN8xkN5/infra-report-20260926-153406.html
* the message is type=private, sender abiba-bot@chat.sysloggh.net, recipients
[9, 21], body carries the top-line status and the attachment link, and does NOT
contain a <table> - i.e. it does not reproduce the report
* the attachment fetches HTTP 200, 16208 bytes, content-type text/html, starts
with <!DOCTYPE html>, and contains <style>, <table> and 16 class="card" blocks -
it opens as a standalone styled document
* failure path: a bad key gives 'Delivery FAILED at upload: Malformed API key',
EXIT=1, the HTML is printed to stdout and persisted to disk
* scheduled path: the run's own output is pasted in the PR
prose-lint: PASSED (19 warnings); secret scan clean.
Master's lint is RED right now, and it is my doing.
at 483a66b (before #133): prose-lint -> LINT PASSED
at 9d64b0b (after #133): prose-lint -> LINT FAILED, 4 credential-shaped strings
The regression came from #133's new pve_auth() work. The secret scanner flags
the literal header shape PVEAPIToken=<value> (rule proxmox-token), and three
occurrences landed in the tree:
scripts/daily-infra-report.py the f-string building the auth header
tests/test_daily_infra_report.py a docstring quoting the old placeholder
tests/test_daily_infra_report.py a synthetic token in an assertion
Fixed without allowlisting anything, because none of these is a credential:
* the header prefix becomes PVE_AUTH_HEADER = "PVEAPIToken=", a constant ending
at '=' so the scanner's pattern (which needs a character after '=') cannot
match, and the f-string no longer contains the literal;
* the test docstring no longer reproduces the old placeholder verbatim;
* the test builds its expected value from the constant plus a local sample
variable instead of embedding a credential-shaped literal.
Behaviour is unchanged and re-verified: with PVE_TOKEN unset the script still
exits 1 with the probe failure, and with the vault token it still reports
pve_probe_status ok / node_count 5 / nodes_online 5.
before: LINT FAILED — 4 credential-shaped strings
after: LINT PASSED (18 warnings)
The Proxmox leg of the daily digest has been reporting NOTHING while exiting 0.
Root cause: the auth header was a literal placeholder string,
AUTH = "Authorization: PVEAPIToken=«vault: infrastructure/production PVE_API_TOKEN»"
which was sent verbatim. The API rejected it, pve_get() returned None, and the
report rendered node_count=0 / nodes_online=0 with pve_probe_status='unreachable'
while still exiting 0. A monitoring gap that looks like a healthy run.
Also: the vault key is PVE_TOKEN, not PVE_API_TOKEN, so even reading os.environ
by the old name would not have found it.
Fixes:
* pve_auth() resolves the token at call time from PVE_TOKEN (injected by
'infisical run --env=prod'). Nothing is hardcoded; a missing token raises.
* pve_get() builds the command inside its try block, so a missing token degrades
to None instead of escaping as an unhandled exception.
* PROBE_FAILURES records an unreachable node/resources probe. Probe failures are
deliberately separate from DEGRADED_LEGS: a missing credential stays exit 0
(existing intent), but a probe with no data now exits 1 in both the report and
--json paths, so it cannot pass unnoticed.
Measured effect on the live host: pve_probe_status unreachable -> ok,
node_count 0 -> 5, nodes_online 0 -> 5, total_vms 0 -> 22, running_vms 0 -> 22.
Tests: 4 new regression tests; all 4 fail against the pre-fix script and pass
after, and the 3 pre-existing tests still pass (7/7).
Not fixed here (needs the captain): the email leg fails with
'534 5.7.9 Application-specific password required' - EMAIL_PASSWORD in the vault
is not a valid Gmail app password for jtabiri@gmail.com. That is a credential
action, not a code change.
1. Remove vestigial ZULIP_API_KEY requirement:
- /api/v1/server_settings is a PUBLIC endpoint (verified HTTP 200 with or without credential)
- No Zulip API key is required for this call
- If a future leg genuinely needs abiba-bot's key, it must prove it with a 200 from
/api/v1/users/me as abiba-bot and label itself degraded when it cannot
- Never fall back to the vault's shared ZULIP_API_KEY
2. Make failed sends exit non-zero:
- A degraded leg (no credential configured) must stay exit 0
- A failed send (attempted and failed) must exit 1
- This distinguishes 'not configured' from 'attempted and failed'
Test evidence:
- No-credential run: exit 0, digest still produced
- Wrong password: exit 1, labelled SMTP error
- grep -n ZULIP_API_KEY: only comment reference remains
- ZULIP_API_KEY: no longer SystemExit, now reports 'credential-missing: ZULIP_API_KEY'
- EMAIL_PASSWORD: no longer sys.exit(1), now appends to DEGRADED_LEGS and returns success
- PVE API: fixed None check in storage section
- Summary: reports degraded legs before summary
This allows the digest to be produced and emailed even when credentials are missing,
while still explicitly logging which legs are degraded.
Test: empty env produces JSON report + degraded leg labels, no SystemExit.
Scripts that had hardcoded credentials:
- scripts/zulip-monitor.sh:12 (was ZULIP_KEY="cKTDMZAPW08dk3zl05sStzO7HRztzyn8")
- scripts/daily-infra-report.py:25 (was ZULIP_KEY="cKTDMZAPW08dk3zl05sStzO7HRztzyn8")
Both now read from environment variable ZULIP_API_KEY (set by vault-backed start script)
with loud failure if not present.
Other credentials in scripts/:
- capture-dsh-token.sh: uses TOKEN variable with fallback (not a secret)
- pm2-self-heal.sh: reads TELEGRAM_BOT_TOKEN from /root/.pi/agent/extensions/telegram/.env (acceptable)
- prose-ai-review.sh: uses GITEA_TOKEN from .env file with LITELLM_KEY fallback (not secrets)
No other hardcoded credentials found.
Proof of behavior:
With ZULIP_API_KEY set:
bash scripts/zulip-monitor.sh → Server: HTTP 200 (authenticated)
python3 scripts/daily-infra-report.py --json → Collecting infrastructure data...
Without ZULIP_API_KEY set:
bash scripts/zulip-monitor.sh → "ZULIP_API_KEY not set — refusing to run with no credential"
python3 scripts/daily-infra-report.py --json → "ZULIP_API_KEY not set — refusing to run with no credential"
Cred source: environment variable ZULIP_API_KEY (set by vault-backed start script)
No key rotation (that is a separate decision).
- Load API key from durable file /root/.abiba-workspace/secrets/litellm-key.txt (works in cron)
- Fix http_get to use Bearer token instead of Basic Auth for API endpoint check
- All 6 LiteLLM checks now pass (was 5/6)
Captain ruling 2026-09-10: Mumuni moved off this host onto her own container
(kagentz CT 105 on minipve, 192.168.68.14, dedicated `hermes` user) and is
monitored from her side. This host must not monitor anything Mumuni.
The stale probes fired false alerts repeatedly:
* scripts/zulip-monitor.sh ssh'd to root@192.168.68.24 for the
decommissioned deployment's ~/.hermes/gateway_state.json, read "unknown"
on every run, and posted a 🔴 "Mumuni (Hermes) Zulip state: unknown" DM +
#agent-hub stream alert each cycle.
* scripts/daily-infra-report.py published a matching "mumuni:unknown" row in
every digest.
Changes:
* zulip-monitor.sh: delete the "Platform B: Hermes (Mumuni)" leg and its
notify; keep the Zulip-server, Platform A pi/Abiba (bridge), Platform B
Tanko and Platform C Agent Zero legs. A comment records why the leg is
retired so it is not re-added. Also fixes SC2155 so shellcheck is clean.
* daily-infra-report.py: delete the .24 ~/.hermes/gateway_state.json agent
probe, its hermes --version probe, and the now-dead mumuni render branch.
Abiba (CT 100, its own .24 address) and Tanko legs unchanged; the PVE API
token name is untouched.
* zulip-health.prose.md (v3.1.0): drop the Mumuni-only B4 gateway-process,
B5 heartbeat and B6 response-delivery steps and the stale 192.168.68.24
references; state explicitly that Mumuni is not monitored from this host.
Tanko/Agent-Zero/bridge steps retained.
* agent-health-check.py: correct the v2 changelog roster comment that still
placed mumuni at .24/CT100. No behavior change — the mumuni probe was
already absent from the AGENTS dict; v5 changelog notes the correction.
Tests: tests/test_mumuni_monitor_removal.py pins the removal structurally and
behaviorally — the shipped zulip-monitor.sh is run in a sandbox (only its LOG
constant rewritten) with stub ssh/curl on PATH; the ssh stub records every
target host, so "never reaches .24" and "no Mumuni notify even on the alert
path" are asserted from observed behavior. A mutation check (re-inject the old
leg) fails the suite, so the guarantee is not vacuous.