Commit Graph
19 Commits
Author SHA1 Message Date
root a13457bcd6 fix(infra): Fix Docker Stats (9324) and PVE Exporter (9221) probe ports
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
Previously probed wrong ports:
- Docker Stats was at 9323 (dockerd metrics) but should be 9324
  (harness-docker-stats, docker_container_* metrics)
- PVE Exporter was at 9324 (harness-docker-stats) but should be 9221
  (harness-pve-exporter, 5 pve_* metrics)

Both exporters bind to 127.0.0.1 on CT 116 and must be probed via SSH.

Updated infrastructure-monitoring.prose.md to document the correct ports.
Added test assertions verifying the exact ports are probed.

Branch: fix/infra-monitoring-probe-ports-20260919
2026-09-19 02:32:00 +00:00
root 3c7f5d7d65 fix(infra+zulip): PR #115 round-2 findings F1-F4+F7
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
F1: Make kind classification real — append (<kind>) to every failure
     line, assign kind=tls on curl exit 35/60, fix :112 where
     kind=refused was set on successful retry. Update prose shape.
F3: Move credential placeholder skip from server probe to notify()
     only — server is always probed (200 without auth verified live).
F4: notify() logs ALERT SUPPRESSED when credential unusable so
     alerts from other legs are not silently dropped.
F7: Restore trailing newline in infra-monitoring.sh.

F5 (DO NOT CHANGE): Verified directly — ssh root@192.168.68.6
     'grep -n keep-daily /etc/pve/jobs.cfg' returns five
     prune-backups keep-daily=35 lines. Prose is CORRECT.

Test: 18 passed, 0 failed (bash scripts/test_infra_monitoring.sh)
2026-09-18 18:21:12 +00:00
root 385f7e0623 fix(infra-monitoring): resolve PR #115 review findings (A1-A3, B1-B2, C1-C2)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 10s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
A1: Test -k assertion now checks use_k:+-k syntax (actual bash pattern)
A2: PVE_NODES assertions now count expected nodes and verify exact array size
A3: Test now asserts liveness behavior (PVE_API_LIVENESS=1) not source text
B1: disk-gc GC schedule corrected: cron runs pbs-gc.sh (not proxmox-backup-manager),
    schedule is 20:00 LOCAL (00:00 UTC, not 20:00 UTC), host timezone America/New_York
B2: PROBE SHAPE now documents actual output shape including TLS flag notes
C1: TLS kind is now printed in PVE API failure output
C2: SSH retry logic clarified - retry is in probe_http function (not unreachable)
2026-09-18 05:54:06 +00:00
root 93f15709d1 fix(infra-monitoring): move probes to versioned script with port-drift test
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
The 2026-09-17 false-verdict incident (third recurrence) showed that prose
policy is not a control: the agent probed :9325/:9405 (nonexistent ports),
CT 116 for PVE API (should be real PVE nodes), and rendered TLS failures as
connection-refused. This moves the canonical probe set into
scripts/infra-monitoring.sh (executed verbatim by the contract) and adds
scripts/test_infra_monitoring.sh which asserts every probed port matches the
documented value.

- scripts/infra-monitoring.sh: one script per contract pattern; all targets,
  ports, paths, and expected-status rules in code; -k for PVE self-signed
  certs; non-zero exit naming every failed target; no OK summary on failure
- scripts/test_infra_monitoring.sh: 20 assertions covering port drift,
  monitoring-host-as-PVE-node, and missing -k flag
- infrastructure-monitoring.prose.md: check-health section now references the
  script as executable owner; paste its raw output verbatim

Proof: all 13 legs pass (exit 0); deliberately broken Grafana port (9325)
produces 'probe-failed: 192.168.68.116:9325 (expected 200)' and exit 1.
2026-09-18 05:17:35 +00:00
root efe9381283 fix: probe precision — correct GPU exporter path, Grafana port, add retry + probe-failed reporting
Per defect report 1150.msg:
- GPU exporters: probe /metrics (Prometheus scrape target), not bare /
- Grafana: correct port 3001 (not 3000)
- All probes: print target name + full URL + HTTP code
- All probes: retry once at 25s on 000/timeout
- Apply standing rules: any HTTP status = ALIVE; only 000/timeout/refused = probe-failed

Verified: all probes now return real HTTP codes (GPU 200, Grafana 200, Router 200, LiteLLM 301)
2026-09-14 14:44:37 +00:00
root 6f40a3be60 fix: address credential-sourcing review findings
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
- infrastructure-monitoring: Correct Zulip key path to /etc/litellm-monitor.env
  (not /etc/zulip-bot.env which doesn't exist); add credential-missing check;
  remove unused monitor_key variable

- litellm-health: Clarify that syslog-auto is a fallback pool alias, not a
  step 7 probe; monitor key must be scoped for all four aliases
2026-09-12 23:17:27 +00:00
root d9368467ff fix: credential-sourcing fix for litellm-health and infrastructure-monitoring
- litellm-health.prose.md: Make monitor key retrieval explicit (ssh from CT 100 to CT 116)
  with executable commands; add syslog-auto alias; document credential-missing failure
  condition (not bare 401 or 0 keys)

- infrastructure-monitoring.prose.md: Fix Zulip POST probe to retrieve keys from CT 116
  via ssh instead of sourcing local env file that doesn't exist on executor host

- Verify model inference probes return 200 for gpu-dense, gpu-vision, strix-moe,
  syslog-auto with corrected credential retrieval
2026-09-12 23:13:24 +00:00
agent-zero c26255f5ff fix(router): decommission legacy GPU router across contracts and fleet scripts
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
Router container, image, and config are purged on CT 116 (verified: no
container, no image, inference-harness-router:latest removed, :9000 free,
11 containers healthy, 7 models, live syslog-auto completion OK). Updates:

- gpu-fleet.prose.md: topology diagram rebuilt without the router tier
- gpu-monitor / gpu-self-heal / litellm-health / litellm-self-heal: router steps dropped
- infrastructure-control.prose.md: container inventory + litellm row corrected
- scripts/daily-infra-report.py, scripts/prose-ai-review.sh: host-scoped checks

Verified: 40/40 anchors matched uniquely, 0 U+FFFD across changed files,
daily-infra-report.py compiles, no stray harness-router references remain.
2026-09-11 14:20:18 -04:00
abiba-bot 194e256ac5 no-mistakes(review): Harden health-check provenance, infisical verification, report-only JSON, tests 2026-09-10 01:35:44 +00:00
abiba-bot c66d9c1e20 fix: probe-drift round 2 — stale monitoring expectations (health check, PVE API, GPU port 80, report provenance)
Second probe-drift correction pass after #65/#66/#68. All four legs were stale
consumer expectations, not live faults.

1. scripts/agent-health-check.py (v4)
   - abiba declared pi-only runtime (harness purge): Hermes-era gateway, config
     and wrapper legs are skipped instead of failing.
   - koby declared report_only (captain ruling 2026-08-17, Rule 17): every koby
     leg is detected and reported, never counted as a fleet failure or repaired.
   - koby's CT 111 mapping corrected to storepve (.6); the old amdpve mapping
     made `pct status 111` fail and read as ct-unreachable.
   - wrapper check no longer FAILs .env-based wrappers that legitimately never
     invoke infisical (koonimo).
   - keys load in main() (load_agent_keys) so the module is importable/testable.
   - every run prints absolute execution provenance (script + cwd), in the header
     and in --json.
   Before: 6 FAILURE(S). After: 0 failures, koby reported read-only.

2. infrastructure-monitoring.prose.md
   - PVE API probe repointed from CT 116 (no pveproxy, 000) to the five real
     nodes on https://<node>:8006/api2/json/version, all 401 = alive.
   - any-HTTP-response liveness rule added (401/3xx alive; 000/timeout = DOWN).
   - LiteLLM health documented as 301 -> /litellm/health/liveliness, not bare 200.

3. gpu-monitor.prose.md
   - GPU health probes on :8080 (or router /health/unified); bare port 80 on a
     GPU host is forbidden (no listener -> false DEGRADED).
   - router /health/unified 301 -> /gpu/gpu-data documented as alive.
   - port-discipline + liveness rule + direct-fallback execution step.

4. Report provenance (all contracts)
   - docs/AUTHORING-GUIDE.md documents the rule; scripts/prose-lint.sh enforces
     that any **Report format** contract states an absolute path (pwd -P).
   - provenance added to gpu-monitor, infrastructure-monitoring, proxmox-monitor.

Tests: tests/test_probe_drift.py (23 passed); prose-lint.sh clean + shellcheck
clean; CI frontmatter validation passes. Evidence with per-leg before/after and
absolute paths: docs/probe-drift-round2-evidence.md.
2026-09-10 01:23:54 +00:00
root a3e97ce72b fix: add Router+LiteLLM+PVE API probes to infrastructure-monitoring check-health
- Added Router health probe (http://192.168.68.116/health) — expected 200
- Added LiteLLM health probe (http://192.168.68.116/litellm/health) — expected 200
- Added PVE API probe (https://192.168.68.116:8006/api2/json) — 401 expected for unauthenticated
- Clarified that 401 means API is up, 000 means unreachable, 500 means API down

Closes: infra-monitoring-probe-alignment
2026-09-08 07:20:41 +00:00
root cc5fe0991c Move Zulip bot creds from inline to env/vault reference
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
The abiba-bot@chat.sysloggh.net:KEY is now sourced from
/etc/litellm-monitor.env as ZULIP_USER and ZULIP_BOT_KEY.
2026-09-08 06:18:21 +00:00
mumuni-bot 403fbcdd9f fix: remove duplicated ## Execution block in infrastructure-monitoring
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 15s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
PR #50's content was squash-merged out-of-band (2dfc3e1, 2026-08-30) while
the PR itself stayed open, double-applying the check-health/Execution
section. This drops the second copy; keeps one canonical
Execution -> check-health -> Phases 1-4 -> Verification Commands flow.
2026-09-07 07:08:49 +00:00
root 2dfc3e1530 Merged PR #50: fix/gpu-dense-docs
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
2026-08-30 12:16:39 +00:00
root d6ad016ac9 Add check-health section with live probes to infrastructure-monitoring contract
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 16s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 13s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
Fixes the canned 'API UNREACHABLE (HTTP 000)' reporting. Adds an explicit
check-health execution section (Zulip POST, pm2, GPU exporters, Prometheus,
Grafana, LiteLLM probes) with a hard RUN LIVE, NEVER ECHO rule, mirroring
the gpu-monitor contract pattern. Captain priority 2026-08-22.
2026-08-22 23:59:28 +00:00
root 68f9ebff74 Ship: monitoring contract fixes (2026-08-09)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-08-09 21:57:44 +00:00
root 97977f320c fix: contract improvements from 2026-07-09 run log review
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Failing after 14m43s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Has been skipped
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Has been skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Has been skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Has been skipped
- pct-run.sh: add CT 109 note (KVM VM), document decommissioned/migrated CTs
- disk-gc-threat-response: add docker-vm (SSH .7), add amdpve Docker scope,
  remove jitsi (decommissioned), update access matrix with KVM VM section
- infrastructure-control: docker-vm container count 11→16, add Trove agents stack
- infrastructure-monitoring: deployment status updated (core live, GPU exporters not deployed)
2026-07-09 05:32:00 +00:00
root 0d0215e9dd prose: annotate infrastructure-monitoring with current deployment status
- Core stack (Prometheus + Grafana + pve/node/docker exporters) already
  deployed by proxmox-monitor contract (Jul 2)
- GPU exporters (.8/.110/.15) NOT yet live — gpu-exporter crash-loops,
  sidecars never deployed
- Strike nginx /monitoring/ proxy suggestion (sub-path pattern already
  tried and reverted for /grafana/ per proxmox-monitor)
- Add cross-reference: this contract defines target state, proxmox-monitor
  is the as-built
2026-07-04 21:55:43 +00:00
Abiba bb20637b32 fix: align contract name fields to filenames (litellm-health, infrastructure-monitoring)
- litellm-health: 'check-litellm-health' -> 'litellm-health'
- infrastructure-monitoring: 'deploy-monitoring-stack' -> 'infrastructure-monitoring'
Fixes 'prose run <filename>' mismatch. No external references broken.
2026-07-02 20:35:44 +00:00