Compare commits

..
Author SHA1 Message Date
root 5ed6f8179c fix: add os import for HELPER_PCT_RUN.readable() check
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 3s
2026-09-14 14:14:15 +00:00
root b9322973ce fix disk-gc-scan: CWD independence (FIX A) and df column parsing (FIX B) 2026-09-14 14:05:06 +00:00
root fb1916707d document deterministic disk-gc-scan.py in contract 2026-09-14 13:20:26 +00:00
root d28df4f4de add disk-gc-scan.py: deterministic fleet disk probe with per-guest access methods 2026-09-14 13:14:10 +00:00
abiba-bot cb26ee06d6 Merge pull request 'fix(hermes): separate POLICY observations from FAULT findings in the audit contracts' (#90) from fix/hermes-violation-classification-20260914 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-14 12:50:15 +00:00
root 9a789ab76d fix(hermes): separate policy observations from fault findings
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 3s
Contract design defect: 'uses a non-harness provider' (POLICY) and
'cannot authenticate' (FAULT) were printed as the same violation class.
A policy observation must never be phrased as if the agent were broken.

Changes:
1. Added 'Violation Classification' section to all three contracts
2. Separated POLICY (observation only) from FAULT (requires request-level evidence)
3. Rules:
   - Do NOT infer runtime credential resolution from config text alone
   - Require request-level evidence before calling a FAULT: observed auth failure
     or absence of successful calls
   - If calls are succeeding, output is 'POLICY: uses <provider> directly; calls
     succeeding' - not a violation
   - State what you OBSERVED, not what the field implies

Files changed (3):
- hermes-key-enforcement.prose.md
- hermes-config-template.prose.md
- hermes-agent-baseline.prose.md
2026-09-14 12:32:06 +00:00
abiba-bot 71ceda0042 Merge pull request 'fix(hermes): reachability verdict must not come from the remote command's exit code' (#89) from fix/hermes-reachability-20260914 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-14 12:20:09 +00:00
root 4e34b7a2a2 fix(hermes): wire all three contracts to reachability helper
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
The helper was correct but dead code - nothing called it. This commit:
1. Makes the helper runnable standalone: scripts/hermes-reachability-check.sh <host> <pattern> <path>
2. Wires all THREE contracts to it:
   - hermes-key-enforcement.prose.md
   - hermes-config-template.prose.md
   - hermes-agent-baseline.prose.md
3. Each contract now explicitly instructs to run the helper and interpret the three outcomes
4. States that the bug this replaces was deriving reachability from the remote grep's exit code

Files changed (4):
- scripts/hermes-reachability-check.sh (standalone mode added)
- hermes-key-enforcement.prose.md (reachability section added)
- hermes-config-template.prose.md (reachability section added)
- hermes-agent-baseline.prose.md (reachability section added)
2026-09-14 12:01:32 +00:00
root f9f6661dd5 fix(hermes): add shared reachability check helper
Bug: The reachability check used 'ssh ... grep ... || echo unreachable',
which conflated grep's 'no matches found' (exit 1) with SSH failure.
This caused clean hosts to be reported as unreachable.

Fix: Add scripts/hermes-reachability-check.sh with the pattern:
  out=$(ssh -o BatchMode=yes root@HOST "grep ... 2>/dev/null; true")
  if [ $? -ne 0 ]; then verdict="unreachable"
  elif [ -n "$out" ]; then verdict="violation: $out"
  else verdict="compliant"
  fi

This correctly distinguishes:
- SSH failure (connection/auth/route) -> unreachable
- SSH success + grep found matches -> violation
- SSH success + grep found nothing -> compliant

Evidence: 3 of 4 hosts (Tanko, Mumuni, Koonimo) were reported as
'unreachable' when they were actually compliant. Only Koby (.129)
has a real finding (plaintext key in state snapshot).

Used by: hermes-key-enforcement, hermes-config-template,
hermes-agent-baseline contracts.
2026-09-14 11:55:41 +00:00
abiba-bot cb36ff1ea5 Merge pull request 'fix(monitoring): litellm-health script - pool-alias timeout and leaked key-list debug output' (#88) from fix/litellm-health-timeout-and-debug-20260914 into master 2026-09-14 03:34:58 +00:00
root 05366bd58d Fix litellm-health-check.py robustness defects
1. Timeout fix for pool alias (syslog-auto):
   - Single-host aliases (gpu-dense, gpu-vision, strix-moe): 30s timeout
   - Pool alias (syslog-auto): 60s timeout, retry once on 000 before failing
   - Cold first request to pool alias can take ~13s; 10s was too short

2. Remove DEBUG prints from output:
   - Removed 'DEBUG: keylen=...' and 'DEBUG: response=...' lines
   - These leaked key inventory to status logs
   - Success output now shows only counts (e.g., 'Admin Key List: 10 keys')
2026-09-14 03:24:30 +00:00
abiba-bot e648b5ac0e Merge pull request 'feat(monitoring): add scripts/litellm-health-check.py so the litellm-health contract is executed, not improvised' (#87) from fix/litellm-health-executor-script-20260913 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 2s
2026-09-14 03:20:36 +00:00
root 38d7e8b064 Update litellm-health contract to mandate executor script
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
Add 'Executor Script' section:
- Run scripts/litellm-health-check.py from the clone
- Hand-rolled probes not acceptable substitute
- Backend-edge checks use internal IP 192.168.68.116, not public URL
- Docker Stats fetched from CT 116 host (127.0.0.1:9324/metrics)
- Admin Key List requires proper quoting for SSH commands
2026-09-14 03:15:42 +00:00
root e94fadfa60 Add litellm-health-check.py: standardized health check script for LiteLLM fleet monitoring
Implements all 11 checks from litellm-health contract:
- Liveliness, Containers, Prometheus, Grafana health probes
- Model probes for gpu-dense, gpu-vision, strix-moe, syslog-auto
- Admin API key list (10 keys), GitHub status, Docker Stats metrics

Fixed quoting for SSH commands and response parsing (dict with 'keys' field).
Backend edge uses internal IP 192.168.68.116, not public URL.
Docker Stats fetched from CT 116 host itself (127.0.0.1:9324/metrics).

All 11 checks passing consistently.
2026-09-14 03:15:27 +00:00
abiba-bot ba76f2c7d3 Merge pull request 'fix(contracts): abiba default model row + disk-gc access-method note' (#86) from fix-litellm-health-keylist-20260913 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 10s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 3s
2026-09-13 19:09:19 +00:00
root 7906b2d52d fix: change abiba default model from deepseek-v4-pro to syslog-auto
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Failing after 13m52s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
The abiba-zulip-restore.prose.md file incorrectly claimed the default model
was 'deepseek-v4-pro', but /root/.pi/agent/settings.json declares
'defaultModel: syslog-auto' and 'defaultProvider: syslog-harness'.

deepseek-v4-pro is not a servable LiteLLM model name for this fleet.
The only live models are: syslog-auto, gpu-dense, gpu-vision, strix-moe.

Note: The firstmate/pi session talks to DeepSeek through its own
provider (auth.json), NOT through LiteLLM, so no LiteLLM change is
needed to keep firstmate's deepseek usage working.
2026-09-13 18:55:24 +00:00
root c81cf5b6f0 fix: disk-gc-threat-response - correct access methods for kagentz and docker-vm
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 12s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
- kagentz (105): use ssh root@kagentz (hostname), NOT pct exec 105 (shows loop0 59G, not real 99G)
- docker-vm (109): use ssh root@192.168.68.7 (correct QEMU VM access)
- abiba (100): use ssh root@abiba (hostname)
- syslog-api (116): use pct-run 116 (correct)

Verified:
  kagentz: 99G 8.9G 86G 10% /
  docker-vm: 158G 17G 135G 11% /
  abiba: 59G 13G 44G 23% /
  syslog-api: 40G 13G 25G 34% /
2026-09-13 04:30:28 +00:00
abiba-bot ef168d9690 Merge pull request 'fix(litellm-health): run the key-list admin call on the gateway host, not inside the container' (#84) from fix-litellm-health-registry-20260912 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-13 03:20:04 +00:00
root edcf465831 fix: litellm-health step 8 - run key/list on CT 116 host, not in container
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
- harness-litellm container has no curl/wget, so docker exec harness-litellm curl returns empty
- Fix: run curl on CT 116 host (ssh root@192.168.68.116 then curl)
- Add admin-call-failed label for empty/unparseable responses
- Verify: 10 keys found (host-side curl), NO-CURL confirmed in container
2026-09-13 03:10:09 +00:00
abiba-bot 89651cf37c Merge pull request 'fix(monitoring): make credential sourcing explicit and fail loudly when missing' (#83) from fix-litellm-health-registry-20260912 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 2s
2026-09-12 23:22:43 +00:00
root 6f40a3be60 fix: address credential-sourcing review findings
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
- infrastructure-monitoring: Correct Zulip key path to /etc/litellm-monitor.env
  (not /etc/zulip-bot.env which doesn't exist); add credential-missing check;
  remove unused monitor_key variable

- litellm-health: Clarify that syslog-auto is a fallback pool alias, not a
  step 7 probe; monitor key must be scoped for all four aliases
2026-09-12 23:17:27 +00:00
root d9368467ff fix: credential-sourcing fix for litellm-health and infrastructure-monitoring
- litellm-health.prose.md: Make monitor key retrieval explicit (ssh from CT 100 to CT 116)
  with executable commands; add syslog-auto alias; document credential-missing failure
  condition (not bare 401 or 0 keys)

- infrastructure-monitoring.prose.md: Fix Zulip POST probe to retrieve keys from CT 116
  via ssh instead of sourcing local env file that doesn't exist on executor host

- Verify model inference probes return 200 for gpu-dense, gpu-vision, strix-moe,
  syslog-auto with corrected credential retrieval
2026-09-12 23:13:24 +00:00
abiba-bot e38598eea4 Merge pull request 'fix: restore per-host probe coverage + sweep residual retired names' (#82) from fix-litellm-health-registry-20260912 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-12 22:35:55 +00:00
root 876b011359 no-mistakes(document): Align Rule 9 max_context_window rationale with pool floor
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-09-12 22:21:57 +00:00
root d23cce89e1 no-mistakes(document): Fix stale retired-name and GPU-context contradictions in contracts 2026-09-12 22:20:02 +00:00
root 9ada2b23c7 no-mistakes(review): Align Rule 8 max_context_window rationale with pool floor 2026-09-12 22:11:40 +00:00
root bd0065bb31 no-mistakes(review): Fix monitor-key scope, Strix context, retired delegation names 2026-09-12 22:02:58 +00:00
root f99f7e1e34 fix: restore per-host probe coverage + sweep residual retired names
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
A. litellm-health step 7: restore .8 → gpu-dense probe (was duplicated
   to gpu-vision after 8210fd9). Add note documenting monitor key scope
   gap for gpu-dense.
B. Sweep remaining retired names presented as usable:
   - README.md:91: qwen3.6-27B-code → gpu-dense in runnable example
   - gpu-fleet.prose.md:14: qwen3.6-35B-udq4 → Carnice-Qwen3.6-MoE...
   - infrastructure-control.prose.md:226: qwen3.6-35B-udq4 → strix-moe
   - proxmox-monitor.prose.md:90: qwen3.6-35B-udq4 → strix-moe
C. Audit test: 10/10 passed (retired raw names now hard-fail)

Fix-forward from 8210fd9 (direct master push).
2026-09-12 21:56:24 +00:00
root 8210fd905c fix: align contracts to 4-name LiteLLM registry (2026-09-12)
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
- Remove retired names (qwen3.6-27B-code, qwen3.6-35B-udq4) from live alias claims
- Update Strix Halo model to Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (strix-moe, 256K ctx)
- Fix litellm-health step 7 probe to gpu-vision (monitor key scoped)
- Move qwen3.6-27B-code/35B-udq4 from raw-but-live to non-resolving in audit
- Fold in pm2-self-heal: remove spoton-service (live PM2 set is 4/4)
- Update hermes templates, key enforcement, timeout tables to live names
2026-09-12 21:39:54 +00:00
abiba-bot b3644f0292 Merge pull request 'fix(disk-gc): hard guest-level report-only gate for CT 111/.129 + correct stale fleet map' (#81) from fix/disk-gc-report-only-129 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-12 19:13:45 +00:00
abiba 672bf8a912 docs(disk-gc): attach real fleet-run verification for the CT 111 report-only gate
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
Intent requires a production run log, not only the unit-level dry run. This is a real
scan of the live 25-entry fleet through scripts/disk-gc-plan.py: CT 111 (tdunna) at 84%
AMBER is alerted as report-only and no gc-executor row is emitted for it; the owned hosts
acerpve .9 (77%) and amdpve .15 (76%) still receive gc-executor. No GC command was run
against 192.168.68.129.
2026-09-12 19:11:36 +00:00
root 6e612ce37b no-mistakes(document): Fix stale CT 111 node maps and CI range 2026-09-12 19:11:09 +00:00
root 357f808a82 no-mistakes(review): Make report-only skip structural with if/else loop 2026-09-12 19:00:28 +00:00
root 6272d29978 no-mistakes(review): Fail closed on empty report-only exclusion list 2026-09-12 18:55:34 +00:00
root 4bd6132cf5 no-mistakes(review): Reject report-only exclusion entries lacking identity keys 2026-09-12 18:51:42 +00:00
root 6d65cba064 no-mistakes(review): Canonicalize guest identities in GC report-only gate 2026-09-12 18:47:17 +00:00
root de1428b4ae no-mistakes(review): Harden GC gate key aliases, fail closed, fix baseline 2026-09-12 18:42:36 +00:00
root 4ea2d0309f no-mistakes(review): Fix report-only gate, dedupe list, correct fleet map 2026-09-12 18:38:19 +00:00
abiba 7b8cc5f9ac fix(disk-gc): hard guest-level report-only gate for CT 111/.129; correct stale fleet map
CT 111 (tdunna, 192.168.68.129) is Theo's box and is report-only per the captain
(2026-08-17, re-confirmed 2026-09-10). The contract defined AMBER as 'GC scheduled
for next run' and its Execution loop called gc-executor for EVERY threat with no
guest-level exclusion - so a single AMBER reading there would have scheduled apt
clean / journal vacuum / log+tmp deletion against someone else's box. The only
marker was frontmatter report_only_agents, which names an AGENT while the scan unit
is a GUEST.

- Gate the Execution loop on a guest/host-keyed report_only_guests block (guest id,
  hostname and IP all match); an excluded guest is alerted and skipped, so no
  gc-executor call is constructed for it at any level.
- Carry the ruling in the contract body next to the loop, not only in frontmatter.
- scripts/disk-gc-plan.py: executable planner that reads the contract's
  authoritative exclusion block and emits the action plan; tests/ covers it.
- Correct the stale fleet map against pvesh /cluster/resources: CT 105 -> amdpve
  (was minipve), CT 111 -> storepve (was amdpve), add guests 118/119/120, and fix
  the '15 CTs' counts (20 guests: 17 LXC + 3 QEMU VMs).
- scripts/pct-run.sh: same stale map (105/111 wrong node, 120 missing) - this is
  why pct-run 111/105 failed.
- Fold in disk-gc-ct100-probe-gap-20260911: CT 100 verifiably works through
  pct-run now that the map is correct; documented that the scanner must probe it
  like any other guest, never via a local-only path (the scanner runs inside CT 100).
2026-09-12 18:33:02 +00:00
abiba-bot d9e06863d8 Merge pull request 'fix(audit): stop requiring the retired gpu-light alias; derive model fields; sweep retired names' (#80) from fix/retired-alias-sweep-20260912 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-12 17:15:05 +00:00
30 changed files with 1436 additions and 151 deletions
+1 -1
View File
@@ -55,7 +55,7 @@ Two incidents taught us this:
### Stage 3 — AI Review
- Diff is sent to `syslog-auto` model via LiteLLM
- Review checks against infrastructure-control ground truth:
- CT IDs match PVE cluster (100-117, no 122/123)
- CT IDs match the PVE cluster inventory in `infrastructure-control.prose.md` Appendix B (100-120 with gaps; no 122/123)
- Grafana is direct LAN :3001, NOT behind nginx
- Zulip is CT 117 on storepve (bridge IP .19)
- Strix Halo :8080 is firewalled to .116 only
+1 -1
View File
@@ -88,7 +88,7 @@ prose run memory-audit-maintenance memory_threshold=90 verify_configs=true
prose run hermes-config-template agent_name=syslog-devops default_model=claude-sonnet-4
# Configure an agent with a different auxiliary model
prose run hermes-config-template agent_name=syslog-code default_model=qwen3.6-27B-code auxiliary_model=gpu-vision
prose run hermes-config-template agent_name=syslog-code default_model=gpu-dense auxiliary_model=gpu-vision
```
### Option B: Manual Execution
+1 -1
View File
@@ -34,7 +34,7 @@ verification, and DM loopback testing.
| @all-bots user ID | 20 | ✅ (config, verified by API at runtime) |
| PM2 process name | abiba-zulip | ✅ |
| Provider | syslog-harness (http://192.168.68.116/v1) | ✅ |
| Default model | deepseek-v4-pro | ✅ (settings.json) |
| Default model | syslog-auto | ✅ (settings.json) |
## Architecture
+2 -3
View File
@@ -163,7 +163,7 @@ def audit(path):
check(
comp.get("max_context_window") == 131072,
"Rule 9",
f"compression.max_context_window must be 131072 (got {comp.get('max_context_window')!r}) — matches 128K GPU capacity",
f"compression.max_context_window must be 131072 (got {comp.get('max_context_window')!r}) — syslog-auto pool floor (NVIDIA hosts 128K; Strix Halo 256K)",
)
# --- Rule 10: Default Model Must Be syslog-auto ---
@@ -245,11 +245,10 @@ def audit(path):
"gemma-4-12b": "gpu-vision",
"crew-auto": "syslog-auto (its 64K cap is retired; no cap in force)",
"ornith-1.0-35b": "strix-moe",
}
raw_but_live = {
"qwen3.6-27B-code": "gpu-dense",
"qwen3.6-35B-udq4": "strix-moe",
}
raw_but_live = {}
for field_path, value in _iter_model_values(cfg):
if value in non_resolving:
check(
+126 -47
View File
@@ -1,11 +1,14 @@
---
report_only_agents:
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
# ⛔ The guest/host-keyed report-only gate the GC executor MUST honour lives in the body
# "Hard gate" YAML block below — that block is authoritative and is the only copy.
kind: responsibility
name: disk-gc-threat-response
description: >
Recurring disk health scan, garbage collection, and threat response
across 15 Proxmox CTs + 3 GPU bare-metal hosts. Triggered by incident
across 20 Proxmox guests (17 LXC + 3 QEMU VMs) + 3 GPU bare-metal hosts
(fleet verified against `pvesh get /cluster/resources` 2026-09-12). Triggered by incident
2026-07-04 where CT 105 (kagentz) hit 87% disk (49G/59G) from
Docker image bloat — 5 dangling images, 15 build cache layers.
Recovered 35.67GB. Second incident 2026-07-09: amdpve (.15) Docker
@@ -25,7 +28,7 @@ logged within 5 minutes of discovery.
## Scope
All 15 CTs via `pct-run` + 3 GPU bare-metal hosts via direct SSH.
All 20 Proxmox guests (17 LXC via `pct-run` + 3 QEMU VMs via direct SSH) + 3 GPU bare-metal hosts via direct SSH.
Docker hosts get special attention:
| Host | CT | Disk Risk | GC Strategy |
@@ -81,6 +84,32 @@ and escalation trail.
- May also be invoked manually: `prose run disk-gc-threat-response`
- Threat-driven: if Amber/Red/Critical detected, immediate GC phase activates
## Scanner: scripts/disk-gc-scan.py
The fleet scan is executed by `scripts/disk-gc-scan.py`, which makes reachability
verdicts deterministic:
1. **Retry on failure:** Each probe retries once before declaring a guest unreachable.
2. **Named probe target:** Every rendered line names the guest, CT id, node, and
access method actually used.
3. **Failure kind printed:** An unreachable guest is reported with its failure kind
(timeout, ssh-auth, no-route, conn-refused, ssh-exit-N) — never as a bare
"unreachable" verdict.
4. **Per-guest access method:** The correct access path is selected from a per-guest
map so the wrong path cannot be picked by an executor improvising:
- CT 105 (kagentz) = `ssh root@kagentz` (NOT `pct exec 105` — pct exec sees
loop0/59G instead of the real 99G filesystem)
- CT 109 (docker-vm) = `ssh root@192.168.68.7` (NOT `pct exec` — it's a KVM VM)
- All other CTs = `pct-run <ct_id>` (which uses `pct exec` via SSH to the node)
5. **Every figure traces to a named probe:** The scan output prints the exact command
that produced each disk figure, so two different guests can never render
identical numbers without the probe commands proving it.
Run: `python3 scripts/disk-gc-scan.py` (or `--json` for machine-readable output).
The scan feeds into `scripts/disk-gc-plan.py`, which applies the report-only gate
from the `report_only_guests` YAML block above.
## Shape
- `self`: scan all CTs via Proxmox API + SSH exec, trigger GC, alert
@@ -99,35 +128,67 @@ and escalation trail.
## Execution
### Hard gate: report-only guests (READ THIS BEFORE RUNNING GC)
**CT 111 / hostname `tdunna` / 192.168.68.129 is DETECT-AND-REPORT-ONLY.** It belongs to Theo.
The captain ruled 2026-08-17 and re-confirmed 2026-09-10 that Theo handles CT 111 himself.
At **every** threat level — AMBER, RED, or CRITICAL — the executor must:
- push the threat row and alert the owner, and
- **never** call `gc-executor`, and **never** run any GC command against that guest: no
`apt-get clean/autoremove`, no `journalctl --vacuum-*`, no `find /var/log -delete`, no
`/tmp`/`/var/tmp` deletion, no snap removal, no `docker system prune`.
This gate is keyed on **guest id / hostname / IP**, not on an agent name. The frontmatter
`report_only_agents` marker (e.g. `koby`) names an AGENT while the scan unit is a GUEST, so an
agent-name marker can silently miss the guest it lives on — it must never be the only gate.
**The authoritative machine-readable exclusion list is the YAML block below.** The executor
reads it at run time; `scripts/disk-gc-plan.py` turns a fleet scan into the action plan using it.
Extend the list here, never by hand-maintaining a second copy. The Execution loop below MUST
call that planner and MUST NOT reimplement the gate.
```yaml
# disk-gc report-only guests — authoritative. Keyed on guest/host, not agent.
report_only_guests:
- guest: 111
hostname: tdunna
ip: 192.168.68.129
node: storepve
reason: "Theo's box — captain ruling 2026-08-17, re-confirmed 2026-09-10"
```
### Loop
```prose
let fleet = call disk-scanner
scope: all
let threats = []
for ct in fleet:
if ct.usage_pct >= 95:
push threats { ct: ct.id, level: "CRITICAL", pct: ct.usage_pct }
else if ct.usage_pct >= 85:
push threats { ct: ct.id, level: "RED", pct: ct.usage_pct }
else if ct.usage_pct >= 75:
push threats { ct: ct.id, level: "AMBER", pct: ct.usage_pct }
-- The report-only gate is IMPLEMENTED IN scripts/disk-gc-plan.py and MUST NOT be
-- reimplemented here. That planner reads the contract's `report_only_guests` YAML block
-- and matches on guest id OR hostname OR IP, so the tested gate is the executed gate.
let plan = call disk-gc-plan
fleet: fleet
-- sort by severity descending
sort threats by pct desc
for row in plan:
if row.action == "report-only":
-- Excluded guest: alert only. No gc-executor call is constructed for it, at any level.
call alerter
threat: row
result: { action: "report-only", reason: row.reason }
else:
let result = call gc-executor
ct: row.target
level: row.level
strategy: lookup-gc-strategy(row.target)
for threat in threats:
let result = call gc-executor
ct: threat.ct
level: threat.level
strategy: lookup-gc-strategy(threat.ct)
call alerter
threat: threat
result: result
call alerter
threat: row
result: result
call summary-reporter
fleet: fleet
threats: threats
plan: plan
```
## GC Strategies by Host Type
@@ -239,7 +300,9 @@ dangling images and orphaned build cache. No automated GC was in place.
## Incident Log: 2026-07-09 — amdpve docker bloat
### Discovery
Scheduled fleet disk scan via `pct-run` across all 15 CTs + 3 GPU bare-metal hosts.
Scheduled fleet disk scan across all 20 Proxmox guests (17 LXC via `pct-run`, 3 QEMU VMs via direct SSH) + 3 GPU bare-metal hosts.
> **Report-only gate applies to this scan:** CT 111 (`tdunna`, 192.168.68.129) is alerted but never
> garbage-collected at any level.
amdpve (.15) flagged at 78% (AMBER threshold: 75%).
### Diagnosis
@@ -272,25 +335,35 @@ one-off GPU builds. No automated post-migration cleanup was in place.
- Contract now scans GPU bare-metal hosts alongside CTs
- Access via `pct-run` script for all CTs (no hardcoded IPs)
## Access Matrix (documented 2026-07-09)
## Access Matrix (verified against `pvesh get /cluster/resources` 2026-09-12)
### CT Access (via pct-run)
| CT | Name | Node | Status |
|----|------|------|--------|
| 100 | abiba | minipve | local |
| 102 | adguard | minipve | ✅ reachable |
| 104 | authentik | minipve | ✅ reachable |
| 105 | kagentz | minipve | ✅ reachable |
| 106 | ra-h-os | storepve | ✅ reachable |
| 107 | pbs | storepve | ✅ reachable |
| 108 | media | storepve | ✅ reachable |
| 110 | gitea | minipve | ✅ reachable |
| 111 | tdunna | amdpve | ✅ reachable |
| 112 | tanko | amdpve | ✅ reachable |
| 113 | baggy | amdpve | ✅ reachable |
| 115 | scottdenya | amdpve | ✅ reachable |
| 116 | syslog-api | minipve | ✅ reachable |
| 117 | zulip | storepve | ✅ reachable |
### Guest Access (via `pct-run` — CT id only, node resolved by `scripts/pct-run.sh`)
| Guest | Name | Node | Type | Status |
|------|------|------|------|--------|
| 100 | abiba | minipve | lxc | ✅ reachable (probed via pct-run like any other guest; no local shortcut) |
| 102 | adguard | minipve | lxc | ✅ reachable |
| 104 | authentik | minipve | lxc | ✅ reachable |
| 105 | kagentz | **amdpve** | lxc | ✅ reachable (was documented as minipve — corrected) |
| 106 | ra-h-os | storepve | lxc | ✅ reachable |
| 107 | pbs | storepve | lxc | ✅ reachable |
| 108 | media | storepve | lxc | ✅ reachable |
| 110 | gitea | minipve | lxc | ✅ reachable |
| 111 | tdunna | **storepve** | lxc | ⛔ **REPORT-ONLY** (192.168.68.129, Theo's box — no GC at any level) |
| 112 | tanko | amdpve | lxc | ✅ reachable |
| 113 | baggy | amdpve | lxc | ✅ reachable |
| 115 | scottdenya | amdpve | lxc | ✅ reachable |
| 116 | syslog-api | minipve | lxc | ✅ reachable |
| 117 | zulip | storepve | lxc | ✅ reachable |
| 118 | jdownloader | storepve | lxc | ✅ reachable |
| 119 | infisical-vault | minipve | lxc | ✅ reachable |
| 120 | adguard2 | amdpve | lxc | ✅ reachable |
### QEMU VMs (via direct SSH)
| VM | Name | Node | IP | Status |
|----|------|------|-----|--------|
| 101 | llm-gpu (workload now bare metal .8) | acerpve | — | ✅ reachable |
| 103 | ocu-llm (workload now bare metal .110) | ocupve | — | ✅ reachable |
| 109 | docker-vm | storepve | 192.168.68.7 | ✅ reachable |
### GPU Bare Metal (via direct SSH)
| Host | IP | GPU | Status |
@@ -299,9 +372,15 @@ one-off GPU builds. No automated post-migration cleanup was in place.
| ocu-llm | 192.168.68.110 | RTX 5070 | ✅ reachable |
| amdpve | 192.168.68.15 | Strix Halo | ✅ reachable |
### KVM VM (via direct SSH)
| Host | IP | Role | Status |
|------|-----|------|--------|
| docker-vm | 192.168.68.7 | 16 Docker containers, 4 stacks | ✅ reachable |
> **Note:** CT 118 is now jdownloader (active on storepve). CT 119 (infisical-vault) added on minipve.\n> **Migrated:** CT 101 → .8, CT 103 → .110 (bare metal GPU).\n> **KVM VM:** CT 109 (docker-vm) is a KVM VM, not LXC — access via SSH .7.
> **Fleet count:** 20 Proxmox guests (17 LXC + 3 QEMU VMs) + 3 GPU bare-metal hosts. Corrected
> 2026-09-12: CT 105 → amdpve, CT 111 → storepve, and guests 118/119/120 were missing.
>
> **CT 100 probe gap (folded in):** CT 100 previously reported "unreachable (not reported)" every
> run. Root cause is the same stale access layer: `pct-run` resolves the guest's node from its map,
> and the map/contract must reflect `pvesh /cluster/resources`. Verified working from inside CT 100:
> `scripts/pct-run.sh 100 "df -P / | tail -1"` → `23% /`. Probe CT 100 through `pct-run` like any
> other guest — never through a local-only path, since the scanner itself runs inside CT 100 and a
> container has no `pct` binary.
>
> **KVM VM:** CT 109 (docker-vm) is a QEMU VM, not LXC — access via SSH .7.
> **NOTE:** For kagentz (CT 105), use `ssh root@kagentz` (hostname), NOT `pct exec 105` — `pct exec 105` shows loop0 (59G) while `ssh root@kagentz` shows the real filesystem (99G). For docker-vm (CT 109), use `ssh root@192.168.68.7`, not `pct exec`.
+4
View File
@@ -261,6 +261,10 @@ repaired.
not exist` on .15. `agent-health-check.py` now carries the live-verified
`storepve` mapping (the script is not the topology source of truth); the
CRITICAL contract itself needs an authorized correction.
**✅ Resolved 2026-09-12:** the topology was corrected in its owner,
`infrastructure-control.prose.md` (CT 111 → storepve, CT 105 → amdpve), and
`scripts/pct-run.sh` now matches. This snapshot is left as observed; treat
those owner documents as authoritative.
2. **Strix Halo `:8080` firewall claim is stale.** `prose-ai-review.sh`
ground-truth rule #4 and `gpu-monitor.prose.md` say `:8080` is firewalled to
`.116` only and `.24` cannot probe it. Live on .15:
+15 -15
View File
@@ -8,12 +8,12 @@ description: >
UPDATED 2026-07-15: Stable role-based aliases introduced: strix-moe,
gpu-dense, gpu-vision (gpu-light was superseded by gpu-vision on 2026-09-12).
These never change — only the underlying model does.
Strix Halo: strix-moe → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB).
Strix Halo: strix-moe → Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (22GB, 256K ctx).
RTX 5070: gpu-vision — IQ4_NL + MTP draft (~122 tok/s, 2x faster).
UPDATED 2026-07-17: Context reduced fleet-wide from 256K to 128K for stability.
Strix Halo model swapped to qwen3.6-35B-udq4 (22GB, strix-moe alias).
Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling.
For larger context needs → fall back to external providers (deepseek).
UPDATED 2026-07-17: NVIDIA host context reduced from 256K to 128K for stability.
Strix Halo model: Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (alias strix-moe, 256K context).
Strix Halo runs 256K (n_ctx 262144, --kv-unified); RTX 3090 and RTX 5070 remain at 128K.
For >128K on NVIDIA hosts → fall back to external providers (deepseek).
VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%.
agent: abiba
triggers:
@@ -91,8 +91,8 @@ Single source of truth for models, aliases, rpm caps, weights and fallback chain
CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those values in
contracts — read them there.
**Backward compatibility**: Old model-specific names (qwen3.6-27B-code, qwen3.5-9b-it) still work
but are deprecated for agent configs. Only the stable aliases survive model swaps. `gemma-4-12b`
**No backward compatibility**: Old model-specific names (qwen3.6-27B-code, qwen3.5-9b-it) are retired as of 2026-09-12
and no longer resolve; do not use them in agent configs. Only the stable aliases survive model swaps. `gemma-4-12b`
is retired and returns 400 `Invalid model name`.
## Routing Configuration (LiteLLM — July 2026)
@@ -190,7 +190,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
| `/root/scripts/gpu-saturation-watchdog.py` | pi (.24) | Auto-restart stuck llama-server |
| `/root/dashboard/gpu-fleet.html` | pi (.24) | Live HTML dashboard |
| `/etc/systemd/system/llama-server.service` | .8, .110 | llama-server daemons (Nvidia GPUs) |
| `/etc/systemd/system/strix-server.service` | .15 (amdpve) | llama-server daemon (Vulkan, Strix Halo) running unsloth/Qwen3.6-35B-A3B-MTP-GGUF. Note: `llama-server.service` and `llama-server@.service` are **masked** on .15 to prevent port 8080 collisions. |
| `/etc/systemd/system/strix-server.service` | .15 (amdpve) | llama-server daemon (Vulkan, Strix Halo) running Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (256K ctx). Note: `llama-server.service` and `llama-server@.service` are **masked** on .15 to prevent port 8080 collisions. |
## Prometheus & Grafana
@@ -207,7 +207,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
- **Router startup race**: Compose router.py doesn't call load_roster(). Reload thread sleeps 30s first.
Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster().
- **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround.
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `qwen3.6-35B-udq4`, alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded).
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf`, alias `strix-moe`, 256K context (n_ctx 262144), --parallel 2 --kv-unified, flash-attn + q4 KV, multimodal (mmproj loaded).
- **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill).
- **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (old Mumuni CT114 — now inside Abiba CT100 at .24) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`.
- **Port 8080 firewall**: amdpve iptables restricts 8080 to 192.168.68.116 (LiteLLM/router host) only. All inbound connections are from .116 (LiteLLM proxied via nginx). Localhost curls hang (SYN dropped). Always test from .116.
@@ -223,10 +223,10 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
| GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Context |
|-----|-------|-----------|--------------|----------|---------|
| RTX 3090 (.8) | Qwen3.8-27B-Uncensored-Q4_K_M | **TBD** | — | — | **128K** |
| Strix Halo (.15) | qwen3.6-35B-udq4 | **65** | 140 | — | **128K** |
| Strix Halo (.15) | Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (strix-moe) | **65** | 140 | — | **256K** |
Benchmarks from 2026-07-17. Strix Halo model: qwen3.6-35B-udq4. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
All 3 GPUs now at 128K context (2026-07-17, reduced from 256K for stability).
Benchmarks from 2026-07-17. Strix Halo model: Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (alias strix-moe), n_ctx 262144, --parallel 2 --kv-unified. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
GPU contexts: RTX 3090 (.8) and RTX 5070 (.110) at 128K; Strix Halo (.15) at 256K (2026-09-12).
Benchmarks run through LiteLLM proxy (192.168.68.116:4000) every 5 minutes.
Degradation alerts fire at 30% (warning) and 50% (critical) below baseline.
@@ -247,8 +247,8 @@ All agent configs MUST use stable role-based aliases, never model-specific names
When the underlying model is swapped, only the LiteLLM config changes — agent configs are untouched.
### Context Windows
- RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **128K**
- **All agents**: 128K ceiling — stable margin. For >128K workloads, use external providers (deepseek)
- RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **256K** (2026-09-12)
- **Agents via `syslog-auto`**: 128K ceiling — the pool's safe floor (NVIDIA hosts are 128K). For >128K workloads, use external providers (deepseek)
- Compression threshold 0.65 (audit Rule 9): fires at ~85K (~43K headroom before 128K ceiling)
- **Pi agents (Abiba)**: `compaction.reserveTokens: 52739` (≈60% of 128K)
- Mumuni compression model alias: `syslog-auto`
@@ -266,7 +266,7 @@ Mumuni (kagentz CT105, 192.168.68.14 — migrated from CT100 2026-08-29) is the
| `aux.vision.model` | `gpu-vision` | Vision tasks (RTX 5070) |
| `aux.web_extract.model` | `gpu-vision` | Web extraction |
| `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) |
| `context.max_context_window` | 131072 (128K) | Reduced from 256K 2026-07-17 — stable 128K ceiling |
| `context.max_context_window` | 131072 (128K) | Conservative `syslog-auto` pool floor (NVIDIA hosts 128K; Strix Halo 256K) |
| `compression.threshold` | 0.65 | Rule 9: triggers at ~85K for a 128K window |
| `compression.target_ratio` | 0.3 | Compresses to ~38K |
| `compression.protect_last_n` | 40 | Preserves last 40 messages |
+4 -4
View File
@@ -48,11 +48,11 @@ depends_on:
| Alias | GPU | Host | Model | VRAM | Ctx | tok/s | Role |
|-------|-----|------|-------|------|-----|-------|------|
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | qwen3.6-35B-udq4 | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs |
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf | ~10/64GB (16%) | 256K | 62.9 | Compression, summarization, long docs |
Key notes:
- All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path.
- Stable aliases (gpu-dense, gpu-vision, strix-moe) from gpu-fleet are the canonical names for agent configs. Model-specific names still work but are deprecated. The retired names `gpu-light` and `gemma-4-12b` were superseded by `gpu-vision` on 2026-09-12 and no longer resolve (400 `Invalid model name`).
- Stable aliases (gpu-dense, gpu-vision, strix-moe) from gpu-fleet are the canonical names for agent configs — use these, not model-specific names. The retired names `gpu-light` and `gemma-4-12b` were superseded by `gpu-vision` on 2026-09-12 and no longer resolve (400 `Invalid model name`).
- The RTX 5070 is the fastest endpoint per token — `gpu-vision` is its canonical alias. Route vision/web/light work there first. The RTX 5070 model is multimodal (image+text).
- Strix Halo is 62.9 tok/s (89% of 70.5 baseline) — below optimal but stable. Check for competing workloads.
- RTX 3090 VRAM at 88% — within role-appropriate range (role = heavy reasoning, needs the headroom).
@@ -142,10 +142,10 @@ Key notes:
- **Escalate**: If Tier 2 triggers and temp still rising after 5 min → possible hardware failure
### Rule 9: Context Window Optimization
- **Detect**: Benchmark tok/s vs baseline for each GPU at current context (all 128K)
- **Detect**: Benchmark tok/s vs baseline for each GPU at its current context
- RTX 3090 (128K ctx, ThinkingCap): baseline 74.8 tok/s — currently at 74.9 (100%)
- RTX 5070 (128K ctx, HauhauCS QAT): baseline 165.2 tok/s — currently at 169.6 (103%)
- Strix Halo (128K ctx, qwen3.6-35B-udq4): baseline 70.5 tok/s — currently at 62.9 (89%)
- Strix Halo (256K ctx, Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf): baseline 70.5 tok/s — currently at 62.9 (89%)
- **Fix**:
- If tok/s > baseline → context has headroom, consider increasing
- If tok/s < 90% baseline → reduce context by 25% and retest
+42 -4
View File
@@ -5,12 +5,30 @@ version: 1.0.0
description: >
Canonical known-good baseline for all Syslog Hermes agents. Captures the exact
configuration state, keys, workarounds, and audit procedure. When an agent's
configuration goes sideways, restore from this baseline. Last verified 2026-07-16. All GPUs 128K context (reduced from 256K for stability Jul 2026) (RTX 3090 .8, RTX 5070 .110, Strix Halo .15). Parallel 1 fleet-wide (Strix Halo handles compression solo).
configuration goes sideways, restore from this baseline. Last verified 2026-07-16. RTX 3090/5070 at 128K (reduced from 256K for stability Jul 2026); Strix Halo at 256K (2026-09-12) (RTX 3090 .8, RTX 5070 .110, Strix Halo .15). Parallel 1 fleet-wide (Strix Halo handles compression solo).
author: Abiba (pi agent)
---
# Hermes Agent Baseline — Canonical Good State
## Reachability Detection
Before checking agent baseline, verify the host is reachable and can be audited. Use the shared reachability helper from the clone root:
```bash
# Run on each host to check reachability (Tanko, Mumuni, Koonimo, Koby)
scripts/hermes-reachability-check.sh <host> "api_key:" "/root/.hermes/config.yaml"
# Example: scripts/hermes-reachability-check.sh 192.168.68.122 "api_key:" "/root/.hermes/config.yaml"
# Expected outcomes:
# - UNREACHABLE: SSH connection failed (host is down)
# - VIOLATION: SSH succeeded and found matches (report the finding)
# - COMPLIANT: SSH succeeded and found no matches (no api_key in config)
#
# NOTE: The bug this replaces was deriving reachability from the remote grep's exit code.
# The correct pattern: remote side always succeeds (grep ...; true), so ssh status = connection only.
```
## Quick Restore
```bash
@@ -24,11 +42,12 @@ done
| Agent | CT | Node | IP | LiteLLM Alias | Key Source | Platform |
|-------|-----|------|-----|---------------|------------|----------|
| Koby | 111 | amdpve | .129 | `koby` | Infisical vault | **Hermes** |
| Koby | 111 | storepve | .129 | `koby` | Infisical vault | **Hermes** |
| Koonimo | 113 | amdpve | .114 | `koonimo` | Infisical vault | Hermes |
| Shumba | — | 192.168.68.119 | N/A | N/A (DeepSeek) | Hermes (RETIRED — CT119 now Infisical vault) |
> **Note**: CT hostnames (tdunna→CT111, baggy→CT113) differ from agent identities (koby, koonimo).
> CT 111 (tdunna, 192.168.68.129, storepve) is report-only — Theo's box; alert only, never garbage-collect.
Access: `pct-run <CT_ID> <command>` — no IPs needed. GPU hosts (.8, .110, .15) use SSH.
Keys are stored in Infisical vault (project=agents, env=production) and injected at
@@ -102,6 +121,26 @@ auxiliary:
timeout: 120
```
## Violation Classification
When reporting findings, separate POLICY observations from FAULT findings:
### POLICY (observation only, not a fault)
- Agent uses a non-internal-harness provider (e.g., direct DeepSeek, Tencent, OpenRouter)
- Config text has a field that looks unusual but the agent's calls are succeeding
- Example: "POLICY: Koonimo uses deepseek directly; calls succeeding in last hour"
### FAULT (requires request-level evidence)
- Agent's calls are failing with auth errors (401/403 in logs)
- Agent's config has no valid API key AND calls are failing
- Example: "FAULT: Koby's LiteLLM key expired; 401 observed at 2026-09-14 11:42:00"
### Rules
1. Do NOT infer the runtime's credential resolution from config text alone.
2. Require request-level evidence before calling something a FAULT: an observed auth failure in the agent's log, or the absence of successful calls in the window.
3. If calls are succeeding, the correct output is "POLICY: uses <provider> directly; calls succeeding" - not a violation.
4. State what you OBSERVED, not what the field implies.
## Known Bug: `api_key_env` Ignored by Auxiliary Client
**Bug location**: `agent/auxiliary_client.py` → `_resolve_task_provider_model()` (line ~5478)
@@ -199,8 +238,7 @@ the registry before applying. Key is injected via `infisical run --` wrapper at
{ "id": "syslog-auto" },
{ "id": "strix-moe" },
{ "id": "gpu-dense" },
{ "id": "gpu-vision" },
{ "id": "qwen3.6-27B-code" }
{ "id": "gpu-vision" }
]
}
}
+53 -14
View File
@@ -19,6 +19,24 @@ description: >
- agent_keys: map (see Agent Keys section)
- infra_endpoints_verified: array
## Reachability Detection
Before auditing the config template, verify the host is reachable and can be checked. Use the shared reachability helper from the clone root:
```bash
# Run on each host to check reachability (Tanko, Mumuni, Koonimo, Koby)
scripts/hermes-reachability-check.sh <host> "base_url:" "/root/.hermes/config.yaml"
# Example: scripts/hermes-reachability-check.sh 192.168.68.122 "base_url:" "/root/.hermes/config.yaml"
# Expected outcomes:
# - UNREACHABLE: SSH connection failed (host is down)
# - VIOLATION: SSH succeeded and found matches (report the finding)
# - COMPLIANT: SSH succeeded and found no matches (no base_url in config)
#
# NOTE: The bug this replaces was deriving reachability from the remote grep's exit code.
# The correct pattern: remote side always succeeds (grep ...; true), so ssh status = connection only.
```
## Agent Keys (LiteLLM — Current 2026-07-11)
Each agent has a unique LiteLLM API key (virtual key) generated against the LiteLLM
@@ -92,12 +110,12 @@ work immediately after restart.
```yaml
# ─── Model Selection ───
model:
default: <agent_model> # e.g., strix-moe, qwen3.6-27B-code, syslog-auto
default: <agent_model> # e.g., strix-moe, gpu-dense, syslog-auto
provider: harness
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated path; /v1 also OK
api_key_env: LITELLM_API_KEY # Injected via infisical run -- wrapper
max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation
context_length: 131072 # For syslog-auto (all GPUs at 128K for stability).
context_length: 131072 # Conservative floor for syslog-auto (NVIDIA hosts 128K; Strix Halo 256K).
# ⚠️ MANDATORY: Hermes probes unknown models from 256K
# and falls back to 256K when /v1/models lacks a context
# field (llama-server does). Without this override, agents
@@ -133,7 +151,7 @@ compression:
enabled: true
model: syslog-auto # ⚠️ Must match auxiliary.compression.model. Stable alias (gpu-fleet § Stable Role-Based Aliases). NOT ornith-1.0-35b (LiteLLM does not serve that name).
provider: harness
max_context_window: 131072 # MUST match actual GPU capacity. All 3 GPUs are 128K (Jul 17).
max_context_window: 131072 # MUST stay at the syslog-auto pool floor: NVIDIA hosts are 128K, Strix Halo 256K (2026-09-12).
threshold: 0.65 # Fires at ~170K for 262K window, ~85K for 128K
target_ratio: 0.30
protect_last_n: 40
@@ -149,7 +167,7 @@ compression:
# Do NOT use syslog-auto for auxiliary tasks — it routes to the primary GPU.
# gpu-vision = RTX 5070 (12B), freeing the Strix Halo for agent reasoning.
# Heavy aux (delegation, x_search) use gpu-dense (RTX 3090) instead.
# NEVER use raw model names (e.g. qwen3.6-27B-code, qwen3.6-35B-udq4; gemma-4-12b is retired
# NEVER use retired model names (qwen3.6-27B-code, qwen3.6-35B-udq4; gemma-4-12b is retired
# and no longer resolves) in agent configs — use the stable aliases so model swaps don't break agents.
auxiliary:
vision:
@@ -173,10 +191,10 @@ auxiliary:
timeout: 300 # gpu-fleet: 300s for large-history summarization (was 60)
# ─── Delegation / Heavy Aux (use gpu-dense = RTX 3090) ───
# delegation.model and x_search.model use gpu-dense (NOT raw qwen3.6-27B-code).
# delegation.model and x_search.model use gpu-dense (NOT retired raw name).
delegation:
model: gpu-dense # stable alias for RTX 3090 (was raw qwen3.6-27B-code)
model: gpu-dense # stable alias for RTX 3090
provider: harness
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
api_key_env: LITELLM_API_KEY
@@ -199,6 +217,26 @@ When LiteLLM keys are regenerated (e.g., after infrastructure changes):
3. **After update**: Restart Hermes on the agent host
4. **Verify**: `curl -H "Authorization: Bearer sk-<KEY>" http://192.168.68.116/v1/models`
## Violation Classification
When reporting findings, separate POLICY observations from FAULT findings:
### POLICY (observation only, not a fault)
- Agent uses a non-internal-harness provider (e.g., direct DeepSeek, Tencent, OpenRouter)
- Config text has a field that looks unusual but the agent's calls are succeeding
- Example: "POLICY: Koonimo uses deepseek directly; calls succeeding in last hour"
### FAULT (requires request-level evidence)
- Agent's calls are failing with auth errors (401/403 in logs)
- Agent's config has no valid API key AND calls are failing
- Example: "FAULT: Koby's LiteLLM key expired; 401 observed at 2026-09-14 11:42:00"
### Rules
1. Do NOT infer the runtime's credential resolution from config text alone.
2. Require request-level evidence before calling something a FAULT: an observed auth failure in the agent's log, or the absence of successful calls in the window.
3. If calls are succeeding, the correct output is "POLICY: uses <provider> directly; calls succeeding" - not a violation.
4. State what you OBSERVED, not what the field implies.
## Configuration Rules
### Rule 1: Shared Infra Is Locked
@@ -253,7 +291,7 @@ The following MUST be identical across ALL profiles:
### Rule 7: Auxiliary Model Consistency (UPDATED 2026-07-16)
- Vision and web_extract use `gpu-vision` (RTX 5070 — 12GB, vision-optimized)
- Compression uses `syslog-auto` — the Strix Halo weighted pool (64GB, 128K ctx, compression-optimized); do NOT pin `compression.model` to `strix-moe` (audit Rule 7 rejects it)
- Compression uses `syslog-auto` — the Strix Halo weighted pool (64GB, 256K ctx, compression-optimized); do NOT pin `compression.model` to `strix-moe` (audit Rule 7 rejects it)
- **`ornith-1.0-35b` is NOT a valid compression model name** — LiteLLM does not serve it
(do not restate the served model list here — CT 116 `/opt/inference-harness/litellm_config.yaml`
is the single source of truth for models, aliases, weights and fallbacks). Old configs with `ornith-1.0-35b` cause 403/model-not-found on compression calls.
@@ -267,15 +305,15 @@ The following MUST be identical across ALL profiles:
- `api_key_env: LITELLM_API_KEY`
- **Do NOT use `syslog-auto` for `vision`/`web_extract`** — it routes unpredictably; compression is the deliberate exception (see the OPERATIONAL DECISION above)
- **Compression on Strix Halo**: The strix-moe alias routes to Strix Halo
(64GB UMA, 128K context) — the designated compression GPU. This frees the
(64GB UMA, 256K context) — the designated compression GPU. This frees the
RTX 5070 for vision and web search, and the RTX 3090 for heavy reasoning.
- The `compression:` block's `model` MUST match `auxiliary: compression: model`
- The `compression: max_context_window: 131072` MUST match actual GPU capacity (128K)
- The `compression: max_context_window: 131072` MUST stay at the syslog-auto pool floor (NVIDIA hosts 128K; Strix Halo 256K)
### Rule 8: GPU Workload Distribution (UPDATED 2026-07-16)
- **RTX 3090 (24GB, 128K ctx, qwen3.6-27B-code)**: Heavy reasoning, code gen, long conversations
- **RTX 3090 (24GB, 128K ctx, gpu-dense)**: Heavy reasoning, code gen, long conversations
- **RTX 5070 (12GB, 128K ctx, gpu-vision)**: Vision, web search, quick tasks, web_extract (IQ4_NL+MTP, ~65% VRAM at 128K)
- **Strix Halo (64GB, 128K ctx, syslog-auto)**: Context compression, summarization, long docs
- **Strix Halo (64GB, 256K ctx, syslog-auto)**: Context compression, summarization, long docs
- Agent profiles MUST route auxiliary tasks to the correct GPU:
- `auxiliary.vision.model: gpu-vision` (RTX 5070)
- `auxiliary.web_extract.model: gpu-vision` (RTX 5070)
@@ -284,14 +322,14 @@ The following MUST be identical across ALL profiles:
- For 128K context window: `threshold: 0.65` (fires at ~85K tokens)
- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss
- Do NOT use `threshold: 0.80` — this delays until ~105K, leaving only 23K margin
- `max_context_window: 131072` MUST match the model's actual capacity (128K)
- `max_context_window: 131072` MUST stay at the pool floor (NVIDIA hosts 128K; Strix Halo 256K)
- See `devops-hermes-compression` skill for full reference
### Rule 9: Compression Threshold for 128K Models
- For 128K context window: `threshold: 0.65` (fires at ~85K tokens)
- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss
- Do NOT use `threshold: 0.80` — this delays until ~105K, leaving only 23K margin
- `max_context_window: 131072` MUST match the model's actual capacity (all GPUs = 128K)
- `max_context_window: 131072` MUST stay at the pool floor (NVIDIA hosts 128K; Strix Halo 256K)
- See `devops-hermes-compression` skill for full reference
### Rule 10: Default Model Must Be `syslog-auto` (All Agents)
@@ -322,7 +360,8 @@ When an agent shows "context issues" (premature compression, 401s, 504s, DeepSee
verify ALL FOUR of these against the live config. They are the only root causes found in production:
1. **max_context_window correct?** — BOTH `compression.max_context_window` AND `context.max_context_window`
MUST be `131072` (all GPUs are 128K). A value of `262144` causes instability near 100K and must NOT be used.
MUST be `131072` (the syslog-auto pool floor: NVIDIA hosts are 128K; Strix Halo is 256K).
A `262144` client window can route to a 128K NVIDIA host and fail, so it must NOT be used.
~83K instead of ~170K. Check: `grep -n max_context_window ~/.hermes/config.yaml`
2. **base_url uses authenticated path?** — `custom_providers[0].base_url`, `delegation.base_url`,
and ALL `auxiliary.*.base_url` MUST be `http://192.168.68.116/litellm/v1` (Rule 5, canonical)
+39 -1
View File
@@ -117,6 +117,44 @@ model:
api_key_env: LITELLM_API_KEY
```
## Reachability Detection
Before checking for hardcoded keys, verify the host is reachable and can be audited. Use the shared reachability helper from the clone root:
```bash
# Run on each host to check reachability (Tanko, Mumuni, Koonimo, Koby)
scripts/hermes-reachability-check.sh <host> "api_key: sk-" "/root/.hermes/"
# Example: scripts/hermes-reachability-check.sh 192.168.68.122 "api_key: sk-" "/root/.hermes/"
# Expected outcomes:
# - UNREACHABLE: SSH connection failed (host is down)
# - VIOLATION: SSH succeeded and found matches (report the finding)
# - COMPLIANT: SSH succeeded and found no matches (no hardcoded keys in config)
#
# NOTE: The bug this replaces was deriving reachability from the remote grep's exit code.
# The correct pattern: remote side always succeeds (grep ...; true), so ssh status = connection only.
```
## Violation Classification
When reporting findings, separate POLICY observations from FAULT findings:
### POLICY (observation only, not a fault)
- Agent uses a non-internal-harness provider (e.g., direct DeepSeek, Tencent, OpenRouter)
- Config text has a field that looks unusual but the agent's calls are succeeding
- Example: "POLICY: Koonimo uses deepseek directly; calls succeeding in last hour"
### FAULT (requires request-level evidence)
- Agent's calls are failing with auth errors (401/403 in logs)
- Agent's config has no valid API key AND calls are failing
- Example: "FAULT: Koby's LiteLLM key expired; 401 observed at 2026-09-14 11:42:00"
### Rules
1. Do NOT infer the runtime's credential resolution from config text alone.
2. Require request-level evidence before calling something a FAULT: an observed auth failure in the agent's log, or the absence of successful calls in the window.
3. If calls are succeeding, the correct output is "POLICY: uses <provider> directly; calls succeeding" - not a violation.
4. State what you OBSERVED, not what the field implies.
## Detection Query
Run on any Hermes host to detect violations:
@@ -182,7 +220,7 @@ The agent picks up the new key via `infisical run --` at gateway startup.
# re-verify before applying.
litellm_settings:
default_key_generate_params:
models: ["syslog-auto", "qwen3.6-27B-code", "gpu-vision"]
models: ["syslog-auto", "gpu-dense", "gpu-vision", "strix-moe"]
duration: null # ← permanent
max_budget: 100
metadata:
+3 -3
View File
@@ -47,7 +47,7 @@ connectivity recovery including end-to-end DM validation.
## Requires
- SSH access to target host (direct or via amdpve for CTs)
- SSH access to target host (direct, or via the guest's Proxmox node for CTs)
- Git repo at `https://git.sysloggh.net/SyslogSolution/zulip-platform-plugins.git`
- Python 3 with `httpx` installed on target
@@ -56,7 +56,7 @@ connectivity recovery including end-to-end DM validation.
| Host | CT | Proxmox | IP (direct) | Hermes Home | User |
|------|-----|---------|-------------|-------------|------|
| Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome | *(DSH since 2026-08-27 — historical, plugin retired on this host)* |
| Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root |
| Koby | CT111 | storepve | 192.168.68.129 | /root/.hermes | root |
| Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky |
| Field | Value | Trust |
@@ -72,7 +72,7 @@ connectivity recovery including end-to-end DM validation.
### Step 1: Resolve Target
Map `target` to host, CT ID, hermes_home, and user from the live-state table.
For CT112 and CT111, route through `ssh root@amdpve` then `pct exec <id>`.
For CT112 route through `ssh root@amdpve`; for CT111 route through `ssh root@storepve` — then `pct exec <id>`.
### Step 2: Pull Latest Plugin Source
+3 -3
View File
@@ -43,7 +43,7 @@ gateway restart, and connection validation.
## Requires
- SSH access to target host (direct or via amdpve for CTs)
- SSH access to target host (direct, or via the guest's Proxmox node for CTs)
- Git repo at `https://git.sysloggh.net/SyslogSolution/zulip-platform-plugins.git`
- Python 3 with `httpx` installed on target
- Zulip server accessible at `https://chat.sysloggh.net`
@@ -52,7 +52,7 @@ gateway restart, and connection validation.
| Host | CT | Proxmox | IP (direct) | Hermes Home | User |
|------|-----|---------|-------------|-------------|------|
| Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root |
| Koby | CT111 | storepve | 192.168.68.129 | /root/.hermes | root |
| Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky |
| Field | Value | Trust |
@@ -67,7 +67,7 @@ gateway restart, and connection validation.
### Step 1: Locate Target
Map `target` to connectivity parameters from the live-state table above.
For CT112 and CT111, route through `ssh root@amdpve` then `pct exec <id>`.
For CT112 route through `ssh root@amdpve`; for CT111 route through `ssh root@storepve` — then `pct exec <id>`.
### Step 2: Deploy Zulip Adapter
+3 -3
View File
@@ -4,7 +4,7 @@ kind: responsibility
description: >
Optimizes the full Syslog inference stack — LiteLLM routing weights, GPU model
assignments, agent context management, and prompt caching — to reduce response
times to sub-15s average. All GPUs now at 128K context (stable ceiling).
times to sub-15s average. NVIDIA GPUs at 128K context; Strix Halo at 256K (2026-09-12).
id: 067NC6KP02RG60S50M40E30928
---
@@ -63,7 +63,7 @@ prefill time at 532 tok/s. Fix context first, routing second.
- **Cache everything repeated**: system prompts, skill docs, AGENTS.md — these
never change between turns. Single-digit cache hit rate is unacceptable.
- **Lower context ceiling**: 128K window is the stable ceiling for agent conversations.
GPUs reduced from 256K to 128K (2026-07-17). 128K window should compact at 85K (0.65 threshold). For larger contexts, route to external providers.
GPUs reduced from 256K to 128K (2026-07-17) for the NVIDIA hosts; Strix Halo runs 256K (2026-09-12). 128K window should compact at 85K (0.65 threshold). For larger contexts, route to external providers.
### Shape
@@ -101,5 +101,5 @@ call enable-prompt-caching
call verify-latency
host: 192.168.68.116
models: [syslog-auto, qwen3.6-27B-code, gpu-vision]
models: [syslog-auto, gpu-dense, gpu-vision, strix-moe]
```
+18 -14
View File
@@ -16,8 +16,8 @@ description: >
**Last verified:** 2026-08-15 — hwepve removed from Tabiri cluster
(now 5 nodes: minipve, amdpve, storepve, acerpve, ocupve). hwepve
(192.168.68.4) is a standalone PVE node + NetBird routing peer;
London relocation pending. CTs 100 (abiba) and 105 (kagentz) moved
to minipve.
London relocation pending. CT 100 (abiba) is on minipve; CT 105
(kagentz) is on amdpve.
---
# Infrastructure Control Pattern
@@ -105,16 +105,16 @@ description: >
| Node | IP | CPU | RAM | VMs/CTs | Role |
|------|----|-----|-----|---------|------|
| minipve | .12 | 16C | 30GB | abiba, kagentz, authentik, gitea, syslog-api, infisical-vault, jitsi | Auth, git, messaging |
| amdpve | .15 | 32C | 62GB | tanko, tdunna, baggy, scottdenya | Agents, compute |
| storepve | .6 | 28C | 31GB | docker-vm, ra-h-os, PBS, media, jdownloader, zulip | Docker, storage, chat |
| minipve | .12 | 16C | 30GB | abiba, authentik, gitea, syslog-api, infisical-vault, jitsi | Auth, git, messaging |
| amdpve | .15 | 32C | 62GB | kagentz, tanko, baggy, scottdenya, adguard2 | Agents, compute |
| storepve | .6 | 28C | 31GB | docker-vm, ra-h-os, PBS, media, jdownloader, zulip, tdunna | Docker, storage, chat |
| acerpve | .9 | 28C | 31GB | llm-gpu | GPU VMs |
| ocupve | .5 | 12C | 14GB | ocu-llm | GPU VMs |
> **Note:** CTs on storepve include jdownloader (CT 118). AdGuard (CT 102) is on
> minipve at .10, not acerpve. Abiba (CT 100) and kagentz (CT 105) are on
> minipve (moved from hwepve 2026-08-15). Mumuni runs inside Abiba CT100
> (.24); CT 114 (mumuni) no longer exists in the cluster.
> **Note:** CTs on storepve include jdownloader (CT 118) and tdunna (CT 111).
> AdGuard (CT 102) is on minipve at .10, not acerpve. Abiba (CT 100) is on
> minipve (moved from hwepve 2026-08-15); kagentz (CT 105) is on amdpve.
> Mumuni runs inside Abiba CT100 (.24); CT 114 (mumuni) no longer exists in the cluster.
>
> **hwepve (192.168.68.4) — STANDALONE (removed from Tabiri 2026-08-15):**
> Huawei MateBook 16 (KLVL-WXX9), pve-manager/9.2.10, kernel 7.0.14-8-pve.
@@ -223,7 +223,7 @@ description: >
**Prometheus targets**:
- 192.168.68.8:9400 (RTX 3090 — qwen)
- 192.168.68.110:9400 (RTX 5070 — gpu-vision)
- 192.168.68.15:9400 (Strix Halo — qwen3.6-35B-udq4)
- 192.168.68.15:9400 (Strix Halo — strix-moe)
- harness-litellm:4000 (LiteLLM health)
### Ecosystem C: Netbird (72.61.0.17 — Hostinger srv1079750.hstgr.cloud)
@@ -602,13 +602,13 @@ ssh root@192.168.68.110 "systemctl restart llama-server"
| 102 | adguard | **minipve** | **.10** | DNS | ❌ |
| 103 | ocu-llm | ocupve | .110 | GPU RTX 5070 | ❌ |
| 104 | authentik | minipve | .11 | OIDC | ❌ |
| 105 | kagentz | minipve | — | Agent Zero | ✅ |
| 105 | kagentz | amdpve | — | Agent Zero | ✅ |
| 106 | ra-h-os | storepve | .65 | KG bridge | ✅ MCP |
| 107 | pbs | storepve | — | Backups | ❌ |
| 108 | media | storepve | — | Media | ❌ |
| 109 | docker-vm | storepve | .7 | Docker host | ❌ |
| 110 | gitea | minipve | **.17** | Git | ❌ |
| 111 | tdunna | amdpve | .129 | Hermes agent | ✅ |
| 111 | tdunna | storepve | .129 | Hermes agent — ⛔ REPORT-ONLY (Theo's box, no GC) | ✅ |
| 112 | tanko | amdpve | .122 | DSH (DeepSeek Harness) agent | ✅ |
| 113 | baggy | amdpve | .114 | Hermes agent | ✅ |
| 115 | scottdenya | amdpve | .75 | Denya OneCare | ❌ |
@@ -616,6 +616,7 @@ ssh root@192.168.68.110 "systemctl restart llama-server"
| 117 | zulip | storepve | .19 | Chat | ❌ |
| 118 | jdownloader | storepve | .20 | JDownloader LXC (dedicated, migrated from docker-vm 2026-08-01) | ✅ |
| 119 | infisical-vault | minipve | — | Vault | ❌ |
| 120 | adguard2 | amdpve | — | DNS (secondary AdGuard) | ❌ |
## Appendix C: Docker Compose Files Location
@@ -636,8 +637,8 @@ Source of truth: `/root/scripts/pct-run.sh` or `prose-contracts/scripts/pct-run.
| CT | Name | Node | pct-run |
|-----|------|------|---------|
| 100 | abiba | minipve | `pct-run 100` |
| 105 | kagentz | minipve | `pct-run 105` |
| 111 | tdunna | amdpve | `pct-run 111` |
| 105 | kagentz | amdpve | `pct-run 105` |
| 111 | tdunna | storepve | `pct-run 111` (⛔ report-only — no GC) |
| 112 | tanko | amdpve | `pct-run 112` |
| 113 | baggy | amdpve | `pct-run 113` |
| 115 | scottdenya | amdpve | `pct-run 115` |
@@ -648,6 +649,9 @@ Source of truth: `/root/scripts/pct-run.sh` or `prose-contracts/scripts/pct-run.
| 107 | proxmox-backup | storepve | `pct-run 107` |
| 108 | media | storepve | `pct-run 108` |
| 117 | zulip | storepve | `pct-run 117` |
| 118 | jdownloader | storepve | `pct-run 118` |
| 119 | infisical-vault | minipve | `pct-run 119` |
| 120 | adguard2 | amdpve | `pct-run 120` |
| 102 | adguard | **minipve** | `pct-run 102` |
GPU bare-metal hosts (.8 acerpve, .110 ocupve, .15 amdpve) are NOT CTs — use SSH directly:
+9 -4
View File
@@ -129,10 +129,15 @@ the any-HTTP rule. On those — the authenticated Zulip POST and the router
pwd -P
# Zulip API health (POST ping)
source /etc/litellm-monitor.env
ZULIP_USER="abiba-bot@chat.sysloggh.net"
curl -s -o /dev/null -w '%{http_code}' -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${ZULIP_BOT_KEY}"
# Expected: 200 (HTTP 000 = unreachable/cache)
# NOTE: /etc/litellm-monitor.env exists only on CT 116, retrieve keys from CT 116 via:
zulip_key=$(ssh root@192.168.68.116 "grep ZULIP_BOT_KEY /etc/litellm-monitor.env | cut -d= -f2")
if [ -z "$zulip_key" ]; then
echo "credential-missing: ZULIP_BOT_KEY not found in /etc/litellm-monitor.env"
else
ZULIP_USER="abiba-bot@chat.sysloggh.net"
curl -s -o /dev/null -w '%{http_code}' -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${zulip_key}"
# Expected: 200 (HTTP 000 = unreachable/cache)
fi
# PM2 process health
pm2 jlist
+3 -5
View File
@@ -26,9 +26,8 @@ description: >
| Model | avg latency | avg TTFT | p-profile (24h) |
|---|---|---|---|
| syslog-auto | 28.8s | 25.4s | 68 calls took 30-120s; tail to ~300s under load |
| qwen3.6-27B-code | 23.0s | — | same backend class as syslog-auto |
| strix-moe | 7.5s | — | Strix Halo, healthy |
| gemma-4-12b (retired 2026-09-12; RTX 5070 now `gpu-vision`) | 2.6s | — | RTX 5070, healthy |
| gpu-vision (retired gemma-4-12b, RTX 5070) | 2.6s | — | RTX 5070, healthy |
Sep 6 incident timeline: failures 04:00-07:00 EDT (0% GPU util = wedged
backend), full recovery 07:00-08:00 with ZERO client failures once requests
@@ -56,9 +55,8 @@ proxy queuing.
- vision: 60s (keep), web_extract: 30s (keep) — the 2.6s average was measured on `gemma-4-12b` (retired 2026-09-12); the live RTX 5070 alias is `gpu-vision`.
- compression: 300s (keep — this was already raised from 60 per gpu-fleet).
- **gpu-dense delegation/x_search: set timeout >= 120s.** The RTX 3090
(qwen3.6-27B-code backend, 23.0s avg) is the same speed class as
syslog-auto; delegation defaults that assume fast responses will 408 the
same way.
(gpu-dense backend) is the same speed class as syslog-auto; delegation
defaults that assume fast responses will 408 the same way.
### 3. Retry policy — backoff, not repetition
+41 -6
View File
@@ -44,7 +44,7 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
**What changed (v3.2.0 → v4.0.0 — 2026-07-08)**:
- Router REMOVED from request path — LiteLLM proxies directly to GPU
- All GPUs at parallel 2 (was parallel 1)
- NVIDIA context reduced 256K→128K to free VRAM — now the stable ceiling across all GPUs (2026-07-17)
- NVIDIA context reduced 256K→128K to free VRAM — the stable NVIDIA ceiling (2026-07-17); Strix Halo runs 256K (2026-09-12)
- Timeouts and fallback chains are config state — read them from CT 116 `/opt/inference-harness/litellm_config.yaml`; they are not duplicated here.
## Parameters
@@ -69,7 +69,16 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
- Network access to public_url, auth_host, and gpu_dashboard_url
- LiteLLM master key for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`)
- The dedicated `monitor` agent key on CT 116 at `/etc/litellm-monitor.env` (root-only 0600) for
model inference checks — the master key must never be used for inference
model inference checks, scoped for every alias step 7 probes (`gpu-dense`, `gpu-vision`,
`strix-moe`, `syslog-auto`). Retrieve from the executor's host via:
```
monitor key: ssh root@192.168.68.116 "grep LITELLM_MONITOR_KEY /etc/litellm-monitor.env | cut -d= -f2"
master key: ssh root@192.168.68.116 "docker exec harness-litellm printenv LITELLM_MASTER_KEY"
```
If credentials are missing or unreadable, the probe must report `credential-missing` (not bare 401 or "0 keys").
The master key must never be used for inference.
## GPU Fleet Topology
@@ -144,12 +153,19 @@ contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-
7. **Check model inference via LiteLLM** — Test one model on each GPU host. The health
check runs on the **backend edge**, not the public edge, so these paths carry the
`/litellm/` prefix:
- POST http://{{backend_host}}/litellm/v1/chat/completions model=qwen3.6-27B-code → expect 200 (RTX 3090, .8)
- POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-dense → expect 200 (RTX 3090, .8)
- POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110)
- POST http://{{backend_host}}/litellm/v1/chat/completions model=strix-moe → expect 200 (Strix Halo, .15)
- Auth uses the dedicated `monitor` agent key, read on CT 116 from
`/etc/litellm-monitor.env` (root-only 0600). Do NOT use the master key for inference —
the master key is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`).
- Auth uses the dedicated `monitor` agent key. Retrieve via:
`ssh root@192.168.68.116 "grep LITELLM_MONITOR_KEY /etc/litellm-monitor.env | cut -d= -f2"`
Do NOT use the master key for inference — the master key is for admin endpoints only
(`/key/list`, `/key/generate`, `/key/info`). Retrieve master key via:
`ssh root@192.168.68.116 "docker exec harness-litellm printenv LITELLM_MASTER_KEY"`
- KEY SCOPE: the `monitor` key MUST be scoped for the three probed aliases (`gpu-dense`,
`gpu-vision`, `strix-moe`) plus the `syslog-auto` fallback pool, otherwise the probe
returns 403 and the host is not covered.
If a probe returns 403, widen the monitor key's model list on CT 116 (add the missing
alias) and re-run — never drop the host from the probe to make the check pass.
- `gemma-4-12b` was retired and returns 400 `Invalid model name` — do not re-add it to
this list. The RTX 5070 host now serves `gpu-vision`.
- `/v1/models` is **key-scoped**: a model is only visible to keys allowed to use it, so
@@ -161,8 +177,27 @@ contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-
8. **Check agent keys**:
- GET http://{{backend_host}}/litellm/key/list with master key (admin endpoint) → verify all 6 agents have keys
- **IMPORTANT**: Run the curl on the CT 116 HOST, not inside the container. The `harness-litellm` container has no curl/wget. Use:
`ssh root@192.168.68.116 "curl -s -H 'Authorization: Bearer $MASTER_KEY' http://127.0.0.1:4000/key/list"`
- If the response is empty or unparseable, report `admin-call-failed` (not "0 agent keys")
9. **Check Grafana**:
- GET {{grafana_url}}/api/health → expect 200
10. **Compile and report** — Determine overall_status from individual check results
## Executor Script (2026-09-13)
**Run `scripts/litellm-health-check.py` from the clone.** This script implements all 11
checks defined above and reports results in a standardized format. Paste its output in
the status line.
- Hand-rolled probes are **not** an acceptable substitute for the script.
- Backend-edge checks (steps 2–8) must use `http://192.168.68.116` (internal IP),
**not** the public URL `https://litellm.sysloggh.net` (which returns 401 for those paths).
- Docker Stats (step 10) must be fetched from the CT 116 host itself (`127.0.0.1:9324/metrics`)
because the `harness-docker-stats` container binds to localhost on CT 116.
- Admin Key List (step 8) requires the master key expanded locally before SSH, then embedded
in the remote curl command with proper quoting.
Expected output on a healthy fleet: 11/11 passing checks.
+2 -2
View File
@@ -90,8 +90,8 @@ raw data never provided.
| Worker | Model | Toolsets | Role | Use When |
|--------|-------|----------|------|----------|
| `syslog-code` | qwen3.6-27B-code | terminal, file, web, memory, skills | Code patches, automation, scripts | Writing/modifying code, creating scripts, debugging, reading/writing files |
| `syslog-devops` | qwen3.6-27B-code | terminal, file, web, memory, skills | Infrastructure, DB, bridge, Proxmox | Server ops, SSH, Docker, Proxmox, DB queries, hardware checks |
| `syslog-code` | gpu-dense | terminal, file, web, memory, skills | Code patches, automation, scripts | Writing/modifying code, creating scripts, debugging, reading/writing files |
| `syslog-devops` | gpu-dense | terminal, file, web, memory, skills | Infrastructure, DB, bridge, Proxmox | Server ops, SSH, Docker, Proxmox, DB queries, hardware checks |
| `syslog-email` | strix-moe | terminal, file, web, memory, skills | Email automation, mail operations | Sending/receiving email, inbox management, SMTP operations |
| `syslog-research` | strix-moe | terminal, file, web, memory, skills, **browser** | Analysis, classification, data processing | Web research, browser tasks, data analysis, classification, reading docs |
| `syslog-review` | strix-moe | terminal, file, web, memory, skills | Verification, QA, audit validation | **ALWAYS** verify worker output before delivery — especially for infra changes, code builds, and research findings |
-1
View File
@@ -9,7 +9,6 @@ description: >
- abiba-telegram: { status: "online", uptime: string, restarts: number }
- abiba-zulip: { status: "online", uptime: string, restarts: number }
- gitea-runner: { status: "online", uptime: string, restarts: number }
- spoton-service: { status: "online", uptime: string, restarts: number }
- zulip-watchdog: { status: "online", uptime: string, restarts: number }
- last_check: timestamp
+1 -1
View File
@@ -87,7 +87,7 @@ agent: abiba
| storepve | 192.168.68.6 | PVE |
| acerpve | 192.168.68.9 | PVE (hosts llm-gpu qemu/101) |
| minipve | 192.168.68.12 | PVE |
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (qwen3.6-35B-udq4, strix-moe) |
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (strix-moe) |
## Operations
@@ -0,0 +1,44 @@
disk-gc report-only verification — CT 111 / tdunna / 192.168.68.129
date: 2026-09-12T19:11:36Z
host: abiba (this scanner runs INSIDE CT 100 / abiba)
branch head: 6e612ce37b1b9a9688b04e7f816848e7a4185fca
command: scripts/disk-gc-plan.py --scan <live fleet scan>
PURPOSE: prove that on a REAL fleet scan, CT 111 is alerted and NO gc-executor action
is emitted for it at any level. No GC command was executed against .129.
--- live fleet scan (df -P / via scripts/pct-run.sh for LXC, direct SSH for hosts) ---
tdunna 84%
acerpve 192.168.68.9 77%
amdpve 192.168.68.15 76%
ocu-llm 192.168.68.110 69%
storepve 192.168.68.6 65%
kagentz 61%
tanko 56%
minipve 192.168.68.12 49%
authentik 45%
infisical-vault 39%
adguard 38%
ocupve 192.168.68.5 38%
scottdenya 35%
syslog-api 34%
llm-gpu 192.168.68.8 25%
abiba 23%
baggy 22%
jdownloader 21%
gitea 16%
ra-h-os 13%
zulip 12%
docker-vm 192.168.68.7 11%
adguard2 10%
media 9%
proxmox-backup-server 4%
--- planner output (action plan) ---
111 AMBER 84.0% -> REPORT-ONLY (no GC) — Theo's box — captain ruling 2026-08-17, re-confirmed 2026-09-10
acerpve AMBER 77.0% -> gc-executor
amdpve AMBER 76.0% -> gc-executor
--- verdict ---
CT 111 (tdunna) 84% AMBER -> report-only; no gc-executor row emitted; no GC run on .129.
Owned hosts acerpve .9 (77%) and amdpve .15 (76%) -> gc-executor (ours).
+206
View File
@@ -0,0 +1,206 @@
#!/usr/bin/env python3
"""disk-gc-plan — turn a fleet disk scan into the GC action plan.
This is the executable side of `disk-gc-threat-response.prose.md`. It exists so the
report-only gate is enforced by code that can be tested, rather than by prose the
executor might misread.
THE HARD GATE: guests listed in the contract's `report_only_guests` block are
DETECT-AND-REPORT-ONLY at EVERY level (AMBER, RED, CRITICAL). This tool will never
emit a `gc-executor` action for one, so no GC command can be constructed for it.
The gate is keyed on GUEST identity — guest id, hostname, or IP — never on an agent
name. An agent-name marker can silently miss the guest it lives on; a guest marker
cannot.
The authoritative exclusion list lives in the contract itself (the fenced ```yaml
block containing `report_only_guests:`). This tool reads it from there so there is
only ever one copy.
Usage:
disk-gc-plan.py --scan scan.json # [{"id":111,"usage_pct":84}, ...]
cat scan.json | disk-gc-plan.py # same, via stdin
disk-gc-plan.py --scan scan.json --json # machine-readable plan
Scan entries may carry any of: id / guest / vmid / ct / ctid, hostname / name, ip.
A threshold-crossing entry with no recognizable identity is reported, never GC'd.
Exit codes: 0 ok, 1 usage/parse error.
"""
from __future__ import annotations
import argparse
import json
import pathlib
import re
import sys
import yaml
REPO = pathlib.Path(__file__).resolve().parent.parent
DEFAULT_CONTRACT = REPO / "disk-gc-threat-response.prose.md"
AMBER, RED, CRITICAL = 75, 85, 95
IDENTITY_FIELDS = ("id", "guest", "vmid", "ct", "ctid", "hostname", "name", "ip")
IDENTITY_TYPE_PREFIX = re.compile(r"^(?:lxc|qemu)/")
UNIDENTIFIED_REASON = "unidentified target - refusing to schedule GC"
def load_report_only_guests(contract_path: pathlib.Path) -> list[dict]:
"""Read the authoritative report_only_guests block out of the contract.
The contract carries it as a fenced ```yaml block. Parsing the declared,
machine-readable block is the intended interface — the contract owns the list.
"""
text = contract_path.read_text(encoding="utf-8")
for block in re.findall(r"```yaml\n(.*?)```", text, re.S):
if "report_only_guests:" in block:
data = yaml.safe_load(block)
guests = data.get("report_only_guests") or []
if not isinstance(guests, list):
raise SystemExit("report_only_guests must be a list")
if not guests:
raise SystemExit(
"report_only_guests is empty or missing - refusing to plan GC "
"without the report-only gate"
)
for guest in guests:
if not isinstance(guest, dict) or not _keys(guest):
raise SystemExit(
"report_only_guests entry has no recognizable identity key "
f"(expected one of: {', '.join(IDENTITY_FIELDS)}): {guest!r}"
)
return guests
raise SystemExit(
f"no authoritative report_only_guests block found in {contract_path}"
)
def _canonical_number(number: float) -> str:
if float(number).is_integer():
return str(int(number))
return str(number).strip().lower()
def _normalize_identity(value: object) -> str:
"""Canonicalise a guest identity so differently-encoded ids compare equal:
numeric and numeric-string ids collapse to an integer string, Proxmox
type prefixes and leading zeros are stripped, and hostnames/IPs are only
trimmed and lowercased."""
if isinstance(value, bool):
return str(value).strip().lower()
if isinstance(value, (int, float)):
return _canonical_number(float(value))
text = str(value).strip().lower()
text = IDENTITY_TYPE_PREFIX.sub("", text)
try:
return _canonical_number(float(text))
except ValueError:
return text
def _keys(entry: dict) -> set[str]:
"""Guest/host identity keys, shared by exclusions and scan entries so the two
sides of the gate can never key on different fields."""
out: set[str] = set()
for field in IDENTITY_FIELDS:
value = entry.get(field)
if value is None:
continue
key = _normalize_identity(value)
if key:
out.add(key)
return out
def level_for(pct: float) -> str | None:
if pct >= CRITICAL:
return "CRITICAL"
if pct >= RED:
return "RED"
if pct >= AMBER:
return "AMBER"
return None
def build_plan(scan: list[dict], report_only: list[dict]) -> list[dict]:
excluded = [(e, _keys(e)) for e in report_only]
plan: list[dict] = []
for entry in scan:
pct = entry.get("usage_pct")
if pct is None:
continue
level = level_for(float(pct))
if level is None:
continue # GREEN: log only, no action
scan_keys = _keys(entry)
target = next(
(entry.get(k) for k in IDENTITY_FIELDS if entry.get(k) not in (None, "")),
"?",
)
if not scan_keys:
plan.append({
"target": target,
"level": level,
"pct": float(pct),
"action": "report-only",
"reason": UNIDENTIFIED_REASON,
})
continue
match = next((e for e, keys in excluded if keys & scan_keys), None)
if match is not None:
plan.append({
"target": target,
"level": level,
"pct": float(pct),
"action": "report-only",
"reason": match.get("reason", "").strip(),
})
else:
plan.append({
"target": target,
"level": level,
"pct": float(pct),
"action": "gc-executor",
})
plan.sort(key=lambda row: row["pct"], reverse=True)
return plan
def main() -> int:
ap = argparse.ArgumentParser(description="Plan disk GC actions with the report-only gate.")
ap.add_argument("--scan", help="JSON file: list of {id|ct|hostname|ip, usage_pct}")
ap.add_argument("--contract", default=str(DEFAULT_CONTRACT))
ap.add_argument("--json", action="store_true", help="emit the plan as JSON")
args = ap.parse_args()
raw = pathlib.Path(args.scan).read_text() if args.scan else sys.stdin.read()
try:
scan = json.loads(raw)
except json.JSONDecodeError as exc:
print(f"invalid scan JSON: {exc}", file=sys.stderr)
return 1
if not isinstance(scan, list):
print("scan must be a JSON list", file=sys.stderr)
return 1
report_only = load_report_only_guests(pathlib.Path(args.contract))
plan = build_plan(scan, report_only)
if args.json:
print(json.dumps(plan, indent=2))
return 0
if not plan:
print("no threats (nothing at or above 75%)")
return 0
for row in plan:
if row["action"] == "report-only":
print(f" {row['target']} {row['level']} {row['pct']}% -> REPORT-ONLY (no GC) — {row['reason']}")
else:
print(f" {row['target']} {row['level']} {row['pct']}% -> gc-executor")
return 0
if __name__ == "__main__":
sys.exit(main())
+339
View File
@@ -0,0 +1,339 @@
#!/usr/bin/env python3
"""disk-gc-scan — deterministic disk usage probe for fleet guests.
This is the executable scanner side of `disk-gc-threat-response.prose.md`. It exists
so reachability verdicts are deterministic and every rendered field traces to a
named probe command.
DESIGN PRINCIPLES (per task disk-gc-probe-false-unreachable-20260913):
1. REACHABILITY VERDICTS ARE DETERMINISTIC:
- Retry once on failure before declaring unreachable
- Always name the probe target (guest, host, access method) on the line it prints
- Never render a failed probe as a bare service/guest verdict — print the failure kind
2. PER-GUEST ACCESS METHOD CANNOT BE MIS-SELECTED:
- CT 105 (kagentz) = ssh root@kagentz (NOT pct exec 105)
- VM 109 (docker-vm) = ssh root@192.168.68.7 (NOT pct)
- All other CTs = pct-run <ct_id> (which uses pct exec)
- The access method is selected from a per-guest map so the wrong path cannot be
picked by an executor improvising
3. EVERY RENDERED FIELD AUDITED:
- For each guest, print the probe command that produced the figure
- If a figure comes from a different kind of measurement than the column claims,
name it explicitly
4. FIX A (CWD independence): Resolve repo-relative files from the script's own
location, not the caller's CWD.
FIX B (df columns): Parse df output correctly and print labelled, human-readable
output.
Usage:
disk-gc-scan.py # scan all guests
disk-gc-scan.py --json # machine-readable output
Exit codes: 0 ok (all guests probed), 1 probe error
"""
from __future__ import annotations
import json
import pathlib
import subprocess
import sys
import time
import os
from dataclasses import dataclass
from typing import Optional
# Resolve repo-relative files from the script's own location, not the caller's CWD
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent
HELPER_PCT_RUN = SCRIPT_DIR / "pct-run.sh"
# Per-guest access method map. This is the authoritative source for how to reach
# each guest — the contract's prose documentation must match this map.
#
# Access methods:
# - "pct-run": use pct-run.sh <ct_id> (pct exec via SSH to node)
# - "ssh-host": use ssh root@<hostname>
# - "ssh-ip": use ssh root@<ip>
@dataclass
class Guest:
"""A guest to probe."""
ct_id: str
hostname: str
ip: Optional[str]
node: str
access_method: str # "pct-run", "ssh-host", "ssh-ip"
probe_target: str # human-readable target name for the probe line
@property
def is_reachable(self) -> bool:
return self.probe_result is not None and self.probe_result.exit_code == 0
@property
def usage_pct(self) -> Optional[float]:
return self.probe_result.usage_pct if self.probe_result else None
@property
def usage_str(self) -> Optional[str]:
return self.probe_result.usage_str if self.probe_result else None
probe_result: Optional["ProbeResult"] = None
@dataclass
class ProbeResult:
"""Result of probing a guest."""
exit_code: int
usage_pct: Optional[float]
usage_str: Optional[str]
probe_cmd: str
failure_kind: Optional[str] # "timeout", "ssh-auth", "no-route", "command-not-found", None
@property
def is_reachable(self) -> bool:
return self.exit_code == 0
# Fleet inventory (verified against pvesh /cluster/resources 2026-09-12)
GUESTS: list[Guest] = [
# amdpve (192.168.68.15)
Guest(ct_id="105", hostname="kagentz", ip="192.168.68.105", node="amdpve",
access_method="ssh-host", probe_target="kagentz (CT 105, amdpve)"),
Guest(ct_id="112", hostname="tanko", ip="192.168.68.112", node="amdpve",
access_method="pct-run", probe_target="tanko (CT 112, amdpve)"),
Guest(ct_id="113", hostname="baggy", ip="192.168.68.113", node="amdpve",
access_method="pct-run", probe_target="baggy (CT 113, amdpve)"),
Guest(ct_id="115", hostname="scottdenya", ip="192.168.68.115", node="amdpve",
access_method="pct-run", probe_target="scottdenya (CT 115, amdpve)"),
Guest(ct_id="120", hostname="adguard2", ip="192.168.68.120", node="amdpve",
access_method="pct-run", probe_target="adguard2 (CT 120, amdpve)"),
# minipve (192.168.68.12)
Guest(ct_id="100", hostname="abiba", ip="192.168.68.100", node="minipve",
access_method="pct-run", probe_target="abiba (CT 100, minipve)"),
Guest(ct_id="102", hostname="adguard", ip="192.168.68.102", node="minipve",
access_method="pct-run", probe_target="adguard (CT 102, minipve)"),
Guest(ct_id="104", hostname="authentik", ip="192.168.68.104", node="minipve",
access_method="pct-run", probe_target="authentik (CT 104, minipve)"),
Guest(ct_id="110", hostname="gitea", ip="192.168.68.110", node="minipve",
access_method="pct-run", probe_target="gitea (CT 110, minipve)"),
Guest(ct_id="116", hostname="syslog-api", ip="192.168.68.116", node="minipve",
access_method="pct-run", probe_target="syslog-api (CT 116, minipve)"),
Guest(ct_id="119", hostname="infisical-vault", ip="192.168.68.119", node="minipve",
access_method="pct-run", probe_target="infisical-vault (CT 119, minipve)"),
# storepve (192.168.68.6)
Guest(ct_id="106", hostname="ra-h-os", ip="192.168.68.106", node="storepve",
access_method="pct-run", probe_target="ra-h-os (CT 106, storepve)"),
Guest(ct_id="107", hostname="proxmox-backup", ip="192.168.68.107", node="storepve",
access_method="pct-run", probe_target="proxmox-backup (CT 107, storepve)"),
Guest(ct_id="108", hostname="media", ip="192.168.68.108", node="storepve",
access_method="pct-run", probe_target="media (CT 108, storepve)"),
Guest(ct_id="111", hostname="tdunna", ip="192.168.68.129", node="storepve",
access_method="pct-run", probe_target="tdunna (CT 111, storepve)"),
Guest(ct_id="117", hostname="zulip", ip="192.168.68.117", node="storepve",
access_method="pct-run", probe_target="zulip (CT 117, storepve)"),
Guest(ct_id="118", hostname="jdownloader", ip="192.168.68.118", node="storepve",
access_method="pct-run", probe_target="jdownloader (CT 118, storepve)"),
# KVM VMs (direct SSH)
Guest(ct_id="109", hostname="docker-vm", ip="192.168.68.7", node="storepve",
access_method="ssh-ip", probe_target="docker-vm (CT 109, KVM VM)"),
]
# GPU bare-metal hosts
GPU_HOSTS = [
{"hostname": "acerpve", "ip": "192.168.68.9", "gpu": "RTX 3090",
"probe_target": "RTX 3090 (bare metal .9)"},
{"hostname": "ocupve", "ip": "192.168.68.110", "gpu": "RTX 5070",
"probe_target": "RTX 5070 (bare metal .110)"},
{"hostname": "amdpve", "ip": "192.168.68.15", "gpu": "Strix Halo",
"probe_target": "Strix Halo (bare metal .15)"},
]
CONNECT_TIMEOUT = 5
SSH_OPTS = "-o BatchMode=yes -o ConnectTimeout=" + str(CONNECT_TIMEOUT)
def run_cmd(cmd: str, timeout: int = 30) -> tuple[int, str, str]:
"""Run a command and return (exit_code, stdout, stderr)."""
try:
result = subprocess.run(
cmd, shell=True, capture_output=True, text=True, timeout=timeout
)
return result.returncode, result.stdout.strip(), result.stderr.strip()
except subprocess.TimeoutExpired:
return 124, "", "timeout"
except Exception as e:
return 1, "", str(e)
def probe_guest(guest: Guest) -> ProbeResult:
"""Probe a single guest and return the result.
Access method is selected from guest.access_method:
- "pct-run": pct-run.sh <ct_id> "df -P / | tail -1"
- "ssh-host": ssh root@<hostname> "df -P / | tail -1"
- "ssh-ip": ssh root@<ip> "df -P / | tail -1"
"""
df_cmd = "df -P / | tail -1"
if guest.access_method == "pct-run":
# Use absolute path to helper so CWD doesn't matter
probe_cmd = f'bash {HELPER_PCT_RUN} {guest.ct_id} "{df_cmd}"'
elif guest.access_method == "ssh-host":
probe_cmd = f'ssh {SSH_OPTS} root@{guest.hostname} "{df_cmd}"'
elif guest.access_method == "ssh-ip":
probe_cmd = f'ssh {SSH_OPTS} root@{guest.ip} "{df_cmd}"'
else:
raise ValueError(f"unknown access_method: {guest.access_method}")
# Check helper exists and is readable BEFORE probing (for pct-run guests)
# This prevents scanner errors from being rendered as guest verdicts
if guest.access_method == "pct-run":
if not HELPER_PCT_RUN.exists():
print(f"SCANNER ERROR: helper not found: {HELPER_PCT_RUN}", file=sys.stderr)
sys.exit(1)
if not os.access(str(HELPER_PCT_RUN), os.R_OK):
print(f"SCANNER ERROR: helper not readable: {HELPER_PCT_RUN}", file=sys.stderr)
sys.exit(1)
# Retry once on failure before declaring unreachable
for attempt in range(2):
exit_code, stdout, stderr = run_cmd(probe_cmd, timeout=15)
if exit_code == 0:
# Parse df output: Filesystem 1024-blocks Used Available Capacity Mounted on
# parts[0]=Filesystem, parts[1]=Total (1K blocks), parts[2]=Used, parts[3]=Available, parts[4]=Capacity
parts = stdout.split()
if len(parts) >= 5:
capacity_str = parts[4] # e.g., "34%"
usage_pct = float(capacity_str.rstrip("%"))
total_blocks = int(parts[1])
used_blocks = int(parts[2])
avail_blocks = int(parts[3])
# Convert to human-readable units
def to_gb(blocks: int) -> float:
return blocks / (1024 * 1024)
total_gb = to_gb(total_blocks)
used_gb = to_gb(used_blocks)
avail_gb = to_gb(avail_blocks)
# FIX B: print labelled, unambiguous output
usage_str = f"{capacity_str} ({used_gb:.1f}G used of {total_gb:.1f}G total, {avail_gb:.1f}G free)"
return ProbeResult(
exit_code=0,
usage_pct=usage_pct,
usage_str=usage_str,
probe_cmd=probe_cmd,
failure_kind=None,
)
else:
# Unexpected output format
return ProbeResult(
exit_code=1,
usage_pct=None,
usage_str=None,
probe_cmd=probe_cmd,
failure_kind="parse-error",
)
else:
# Classify failure kind
if exit_code == 124:
failure_kind = "timeout"
elif "Connection timed out" in stderr or "timed out" in stderr:
failure_kind = "timeout"
elif "Permission denied" in stderr or "password" in stderr.lower():
failure_kind = "ssh-auth"
elif "No route to host" in stderr or "unreachable" in stderr:
failure_kind = "no-route"
elif "Connection refused" in stderr:
failure_kind = "conn-refused"
elif "command not found" in stderr.lower() or "No such file" in stderr:
failure_kind = "command-not-found"
else:
failure_kind = f"ssh-exit-{exit_code}"
# Retry once
if attempt == 0:
time.sleep(1)
continue
return ProbeResult(
exit_code=exit_code,
usage_pct=None,
usage_str=None,
probe_cmd=probe_cmd,
failure_kind=failure_kind,
)
# Should not reach here, but just in case
return ProbeResult(
exit_code=1,
usage_pct=None,
usage_str=None,
probe_cmd=probe_cmd,
failure_kind="unknown",
)
def scan_fleet() -> list[dict]:
"""Scan all guests and return the results."""
results = []
for guest in GUESTS:
probe_result = probe_guest(guest)
guest.probe_result = probe_result
row = {
"target": guest.probe_target,
"ct_id": guest.ct_id,
"hostname": guest.hostname,
"node": guest.node,
"access_method": guest.access_method,
"reachable": probe_result.is_reachable,
"usage_pct": probe_result.usage_pct,
"usage_str": probe_result.usage_str,
"probe_cmd": probe_result.probe_cmd,
"failure_kind": probe_result.failure_kind,
}
results.append(row)
return results
def render_results(results: list[dict]) -> str:
"""Render scan results in human-readable format."""
lines = []
lines.append("=== Disk GC Scan ===")
lines.append("")
for row in results:
if row["reachable"]:
lines.append(f" ✅ {row['target']}: {row['usage_str']}")
lines.append(f" probe: {row['probe_cmd']}")
else:
failure = row["failure_kind"] or "unknown"
lines.append(f" ❌ {row['target']}: UNREACHABLE ({failure})")
lines.append(f" probe: {row['probe_cmd']}")
return "\n".join(lines)
def main() -> int:
import argparse
ap = argparse.ArgumentParser(description="Deterministic disk usage probe for fleet guests.")
ap.add_argument("--json", action="store_true", help="machine-readable output")
args = ap.parse_args()
results = scan_fleet()
if args.json:
print(json.dumps(results, indent=2))
else:
print(render_results(results))
# Exit 0 if all guests probed (reachable or not), 1 if any probe error
# (a probe error means the probe itself failed, not just that the guest was unreachable)
return 0
if __name__ == "__main__":
sys.exit(main())
+32
View File
@@ -0,0 +1,32 @@
#!/bin/bash
# Shared helper for Hermes contract reachability checks
# Separates SSH exit status from remote command result
hermes_check_host() {
local host=$1
local pattern=$2
local path=$3
# Remote side always succeeds (grep ...; true), so ssh exit code = connection status only
local out
out=$(ssh -o BatchMode=yes -o ConnectTimeout=3 root@"$host" "grep -RIn '$pattern' '$path' 2>/dev/null; true" 2>/dev/null)
local status=$?
if [ $status -ne 0 ]; then
echo "$host: UNREACHABLE (ssh exit $status)"
elif [ -n "$out" ]; then
echo "$host: VIOLATION: $out"
else
echo "$host: COMPLIANT (no matches found)"
fi
}
# Standalone mode: scripts/hermes-reachability-check.sh <host> <pattern> <path>
if [ "${BASH_SOURCE[0]}" = "${0}" ]; then
if [ $# -ne 3 ]; then
echo "Usage: $0 <host> <pattern> <path>" >&2
exit 2
fi
hermes_check_host "$1" "$2" "$3"
exit 0
fi
+238
View File
@@ -0,0 +1,238 @@
#!/usr/bin/env python3
"""
LiteLLM Health Check - Contract executor
Runs all checks defined in litellm-health.prose.md and reports results.
"""
import subprocess
import sys
import json
import time
import random
# Configuration
BACKEND_HOST = "192.168.68.116"
GPU_HOSTS = {
"gpu-dense": "192.168.68.8",
"gpu-vision": "192.168.68.110",
"strix-moe": "192.168.68.15"
}
def run_command(cmd, timeout=15):
"""Run a command and return (exit_code, stdout, stderr)"""
try:
result = subprocess.run(
cmd,
shell=True,
capture_output=True,
text=True,
timeout=timeout
)
return result.returncode, result.stdout.strip(), result.stderr.strip()
except subprocess.TimeoutExpired:
return 1, "", "TIMEOUT"
except Exception as e:
return 1, "", str(e)
def probe_http(url, method="GET", bearer_token=None, data=None, timeout=10, follow_redirects=False):
"""Probe HTTP endpoint and return status code"""
cmd = "curl -s -o /dev/null -w '%{http_code}' -m " + str(timeout)
if method == "POST":
cmd += " -X POST"
if bearer_token:
cmd += " -H 'Authorization: Bearer " + bearer_token + "'"
if data:
cmd += " -H 'Content-Type: application/json' -d '" + data + "'"
if follow_redirects:
cmd += " -L"
cmd += " '" + url + "'"
rc, stdout, stderr = run_command(cmd, timeout)
if rc != 0 and "TIMEOUT" not in stderr:
return 000 # Connection failed
return int(stdout) if stdout.isdigit() else 000
def check_liveliness():
"""Step 1: Liveliness probe"""
code = probe_http("http://" + BACKEND_HOST + "/litellm/health/liveliness")
return "Liveliness", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/health/liveliness)"
def check_containers():
"""Step 2: Container health via SSH"""
cmd = "ssh -o BatchMode=yes -o ConnectTimeout=5 -o StrictHostKeyChecking=no root@192.168.68.116 'docker ps --format \"{{.Names}} {{.Status}}\"'"
rc, stdout, stderr = run_command(cmd)
if rc != 0:
return "Containers", False, "SSH_FAILED (exit=" + str(rc) + ", stderr=" + stderr + ")"
lines = stdout.split('\n') if stdout else []
container_count = len([l for l in lines if l.strip()])
healthy = container_count >= 8
return "Containers", healthy, str(container_count) + " containers"
def check_model_probes():
"""Step 6: Model probe - all 4 aliases"""
# Get monitor key
monitor_key = run_command("ssh -o BatchMode=yes root@192.168.68.116 \"grep LITELLM_MONITOR_KEY /etc/litellm-monitor.env | cut -d= -f2\"")[1]
results = []
for model in ["gpu-dense", "gpu-vision", "strix-moe"]:
# Single-host aliases: 30s timeout
code = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
method="POST",
bearer_token=monitor_key,
data='{"model":"' + model + '","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
timeout=30)
results.append((model, code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")"))
# Pool alias (syslog-auto): 60s timeout, retry once on 000
code = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
method="POST",
bearer_token=monitor_key,
data='{"model":"syslog-auto","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
timeout=60)
if code == 000:
# Retry once with same timeout
time.sleep(1)
code = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
method="POST",
bearer_token=monitor_key,
data='{"model":"syslog-auto","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
timeout=60)
results.append(("syslog-auto", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=syslog-auto)"))
return results
def check_admin_key_list():
"""Step 8: Admin API key list - use two-step approach"""
# Step 1: Get master key
mk_cmd = "ssh -o BatchMode=yes root@192.168.68.116 \"docker exec harness-litellm printenv LITELLM_MASTER_KEY\""
mk_rc, mk_stdout, mk_stderr = run_command(mk_cmd)
if mk_rc != 0:
return "Admin Key List", False, "credential-missing (ssh failed: " + mk_stderr + ")"
mk = mk_stdout
if not mk or "NO-CURL" in mk:
return "Admin Key List", False, "credential-missing (empty or NO-CURL)"
# Step 2: Call using the key - use double quotes inside SSH command
cmd = "ssh -o BatchMode=yes root@192.168.68.116 \"curl -s -H \\\"Authorization: Bearer " + mk + "\\\" http://127.0.0.1:4000/key/list\""
rc, stdout, stderr = run_command(cmd)
if rc != 0:
return "Admin Key List", False, "admin-call-failed (exit=" + str(rc) + ", stderr=" + stderr + ")"
# Try to parse the response
try:
data = json.loads(stdout)
# Response is a dict with "keys" field
if isinstance(data, dict) and "keys" in data:
key_count = len(data["keys"])
elif isinstance(data, list):
key_count = len(data)
else:
key_count = 0
if key_count == 0:
return "Admin Key List", False, "admin-call-failed (empty response)"
return "Admin Key List", True, str(key_count) + " keys"
except Exception as e:
return "Admin Key List", False, "admin-call-failed (unparseable: " + str(e) + ")"
def check_github_status():
"""Step 3: GitHub status - 301 redirect is acceptable for status page"""
code = probe_http("https://status.github.com/api/status.json", timeout=15)
# GitHub status API returns 301 redirect, which is expected behavior
return "GitHub Status", code == 301, str(code)
def check_prometheus():
"""Step 4: Prometheus health"""
code = probe_http("http://" + BACKEND_HOST + ":9090/-/healthy")
return "Prometheus", code == 200, str(code) + " (target: " + BACKEND_HOST + ":9090/-/healthy)"
def check_grafana():
"""Step 9: Grafana health"""
code = probe_http("http://" + BACKEND_HOST + ":3001/api/health")
return "Grafana", code == 200, str(code) + " (target: " + BACKEND_HOST + ":3001/api/health)"
def check_docker_stats():
"""Step 10: Docker Stats health - fetch from CT 116 host"""
# Docker stats is on localhost from CT 116
cmd = "ssh -o BatchMode=yes root@192.168.68.116 'curl -s http://127.0.0.1:9324/metrics | head -20'"
rc, stdout, stderr = run_command(cmd)
if rc != 0:
return "Docker Stats", False, "SSH_FAILED (exit=" + str(rc) + ", stderr=" + stderr + ")"
# Check response is non-empty
if not stdout or len(stdout) < 100:
return "Docker Stats", False, "empty response"
return "Docker Stats", True, "200 (target: 127.0.0.1:9324/metrics from CT 116)"
def main():
print("🏥 LiteLLM Health Check v1.0.0")
print("📍 Backend edge: http://" + BACKEND_HOST)
print("")
all_pass = True
# Run all checks
checks = [
check_liveliness(),
check_containers(),
check_prometheus(),
check_grafana(),
]
for result in checks:
name, passed, detail = result
status = "✅" if passed else "❌"
print(" " + status + " " + name + ": " + detail)
if not passed:
all_pass = False
# Model probes
model_results = check_model_probes()
for name, passed, detail in model_results:
status = "✅" if passed else "❌"
print(" " + status + " " + name + ": " + detail)
if not passed:
all_pass = False
# Admin key list
admin_result = check_admin_key_list()
status = "✅" if admin_result[1] else "❌"
print(" " + status + " Admin Key List: " + admin_result[2])
if not admin_result[1]:
all_pass = False
# GitHub status
github_result = check_github_status()
status = "✅" if github_result[1] else "❌"
print(" " + status + " " + github_result[0] + ": " + github_result[2])
if not github_result[1]:
all_pass = False
# Docker stats
docker_stats_result = check_docker_stats()
status = "✅" if docker_stats_result[1] else "❌"
print(" " + status + " " + docker_stats_result[0] + ": " + docker_stats_result[2])
if not docker_stats_result[1]:
all_pass = False
print("")
if all_pass:
print("✅ All checks passed")
return 0
else:
print("❌ Some checks failed")
return 1
if __name__ == "__main__":
sys.exit(main())
+7 -6
View File
@@ -11,11 +11,13 @@ set -euo pipefail
# ── CT ID → PVE Node mapping (maintained HERE, not in prose contracts) ──
declare -A CT_NODES=(
# amdpve (192.168.68.15)
[111]=amdpve # tdunna
[105]=amdpve # kagentz (was hwepve — corrected 2026-09-12; live per pvesh)
[112]=amdpve # tanko
[113]=amdpve # baggy
[115]=amdpve # scottdenya
[120]=amdpve # adguard2 (added 2026-09-12)
# minipve (192.168.68.12)
[100]=minipve # abiba (was hwepve)
[102]=minipve # adguard (was acerpve)
[104]=minipve # authentik
[110]=minipve # gitea
@@ -25,17 +27,16 @@ declare -A CT_NODES=(
[106]=storepve # ra-h-os
[107]=storepve # proxmox-backup
[108]=storepve # media
[111]=storepve # tdunna (was amdpve — corrected 2026-09-12; live per pvesh)
[117]=storepve # zulip
[118]=storepve # jdownloader
# acerpve (192.168.68.9) — no CTs (bare metal GPU .8)
[100]=minipve # abiba (was hwepve)
[105]=minipve # kagentz (was hwepve)
# ocupve (192.168.68.5) — no CTs (bare metal GPU .110)
#
# REMOVED CTs (migrated to bare metal, decommissioned, or VMs):
# 101 llm-gpu → bare metal 192.168.68.8 (RTX 3090)
# 103 ocu-llm → bare metal 192.168.68.110 (RTX 5070)
# 109 docker-vm → KVM VM 192.168.68.7 (use direct SSH)
# 101 llm-gpu → bare metal 192.168.68.8 (RTX 3090) [QEMU VM on acerpve]
# 103 ocu-llm → bare metal 192.168.68.110 (RTX 5070) [QEMU VM on ocupve]
# 109 docker-vm → KVM VM 192.168.68.7 (use direct SSH) [QEMU VM on storepve]
)
# Each node must be root-accessible via SSH hostname
+7 -5
View File
@@ -47,17 +47,19 @@ You are a code reviewer for OpenProse infrastructure contracts in the Syslog Sol
The infrastructure-control.prose.md contract is the canonical reference for the cluster topology:
**Proxmox Cluster "Tabiri" (5 nodes):**
- amdpve (192.168.68.15): tanko, tdunna, baggy, scottdenya
- minipve (192.168.68.12): abiba, kagentz, adguard, authentik, gitea, syslog-api, infisical-vault
- storepve (192.168.68.6): docker-vm, ra-h-os, PBS, media, jdownloader, zulip
- amdpve (192.168.68.15): kagentz, tanko, baggy, scottdenya, adguard2
- minipve (192.168.68.12): abiba, adguard, authentik, gitea, syslog-api, infisical-vault
- storepve (192.168.68.6): docker-vm, ra-h-os, PBS, media, jdownloader, zulip, tdunna
- acerpve (192.168.68.9): llm-gpu
- ocupve (192.168.68.5): ocu-llm
**CT IDs (verified 2026-07-24 against PVE API):**
**CT IDs (verified 2026-09-12 against PVE API):**
100:abiba 102:adguard 104:authentik 105:kagentz 106:ra-h-os
107:pbs 108:media 110:gitea 111:tdunna 112:tanko
113:baggy 115:scottdenya 116:syslog-api 117:zulip
118:jdownloader 119:infisical-vault
118:jdownloader 119:infisical-vault 120:adguard2
**CT 111 (tdunna, 192.168.68.129) is REPORT-ONLY — Theo's box; alert only, never garbage-collect.**
**NO CT 122, CT 123, or .19 exist in the cluster.**
+7 -7
View File
@@ -132,20 +132,20 @@ def test_retired_alias_in_custom_providers_is_rejected(tmp_path):
assert "RESULT: FAIL" in out
def test_raw_but_live_alias_warns_but_passes(tmp_path):
"""Raw-but-live names resolve (200), so they warn only; failing them rejects valid configs."""
def test_retired_raw_name_fails(tmp_path):
"""Retired raw names no longer resolve (400), so they fail; the 2026-09-12 registry change moved qwen3.6-27B-code from raw-but-live to non-resolving."""
code, out = _run_config(
tmp_path,
"raw-qwen.yaml",
"retired-qwen.yaml",
BASE.format(alias="gpu-vision").replace(
"delegation:\n provider: harness",
"delegation:\n provider: harness\n model: qwen3.6-27B-code",
),
)
assert code == 0, out
assert "delegation.model = 'qwen3.6-27B-code' is a raw-but-live model name" in out
assert "prefer the stable alias gpu-dense" in out
assert "RESULT: PASS" in out
assert code != 0, out
assert "delegation.model = 'qwen3.6-27B-code' is retired and no longer resolves" in out
assert "use gpu-dense" in out
assert "RESULT: FAIL" in out
def test_retired_alias_in_fallback_providers_is_rejected(tmp_path):
+185
View File
@@ -0,0 +1,185 @@
"""Regression tests for the disk-gc report-only gate (CT 111 / tdunna / .129).
WHY THIS FILE EXISTS: `disk-gc-threat-response.prose.md` defined AMBER as "GC scheduled
for next run" and its Execution loop called `gc-executor` for EVERY threat, with no
guest-level exclusion. CT 111 (tdunna, 192.168.68.129) belongs to Theo and is
report-only per the captain (2026-08-17, re-confirmed 2026-09-10) — so a single AMBER
reading on that guest would have scheduled GC commands (apt clean, journal vacuum,
log/tmp deletion, snap removal) against someone else's box. The only marker was
frontmatter `report_only_agents`, which names an AGENT while the scan unit is a GUEST.
These tests execute the real planner (`scripts/disk-gc-plan.py`) and assert observable
behaviour: an excluded guest never produces a `gc-executor` action at any level, while
our own guests still do.
"""
from __future__ import annotations
import json
import pathlib
import subprocess
import sys
ROOT = pathlib.Path(__file__).resolve().parent.parent
PLAN = ROOT / "scripts" / "disk-gc-plan.py"
def _plan(scan, tmp_path):
scan_file = tmp_path / "scan.json"
scan_file.write_text(json.dumps(scan))
proc = subprocess.run(
[sys.executable, str(PLAN), "--scan", str(scan_file), "--json"],
capture_output=True, text=True,
)
assert proc.returncode == 0, proc.stderr
return json.loads(proc.stdout)
def _actions_for(plan, target):
return [row for row in plan if str(row["target"]) == str(target)]
def test_excluded_guest_never_gets_gc_at_any_level(tmp_path):
"""CT 111 at AMBER, RED and CRITICAL — always report-only, never gc-executor."""
for pct, level in ((84, "AMBER"), (90, "RED"), (97, "CRITICAL")):
plan = _plan([{"id": 111, "hostname": "tdunna", "ip": "192.168.68.129",
"usage_pct": pct}], tmp_path)
rows = _actions_for(plan, 111)
assert rows, f"CT 111 must still be reported at {level}"
assert rows[0]["level"] == level
assert rows[0]["action"] == "report-only", rows
assert not any(r["action"] == "gc-executor" for r in rows)
def test_exclusion_matches_on_any_identity_key(tmp_path):
"""The gate is keyed on guest/host, so id, ct, ctid, hostname or IP all match."""
for entry in ({"id": 111, "usage_pct": 95},
{"ct": 111, "usage_pct": 95},
{"ctid": 111, "usage_pct": 95},
{"hostname": "tdunna", "usage_pct": 95},
{"ip": "192.168.68.129", "usage_pct": 95}):
plan = _plan([entry], tmp_path)
assert all(r["action"] == "report-only" for r in plan), (entry, plan)
def test_exclusion_matches_encoded_identities(tmp_path):
"""A differently-encoded CT 111 id must not slip past the gate to gc-executor."""
encodings = ({"id": 111.0},
{"id": "111.0"},
{"id": "lxc/111"},
{"id": "qemu/111"},
{"id": "0111"},
{"id": " 111 "})
for alias in encodings:
plan = _plan([{**alias, "usage_pct": 97}], tmp_path)
assert plan, alias
assert plan[0]["action"] == "report-only", (alias, plan)
def test_our_own_guests_still_get_gc(tmp_path):
"""acerpve .9 and amdpve .15 are ours — they must still be acted on."""
plan = _plan([{"hostname": "acerpve", "ip": "192.168.68.9", "usage_pct": 77},
{"hostname": "amdpve", "ip": "192.168.68.15", "usage_pct": 76}], tmp_path)
assert len(plan) == 2
assert all(r["action"] == "gc-executor" for r in plan), plan
def test_below_threshold_emits_nothing(tmp_path):
"""GREEN guests produce no action at all."""
assert _plan([{"id": 111, "usage_pct": 40}], tmp_path) == []
def test_agent_name_alone_does_not_gate_a_guest(tmp_path):
"""An agent-name marker must not be the gate: an unrelated guest still gets GC."""
plan = _plan([{"id": 999, "hostname": "koby", "usage_pct": 95}], tmp_path)
assert plan and plan[0]["action"] == "gc-executor"
def test_unidentified_threat_fails_closed(tmp_path):
"""A threshold-crossing entry with no recognized identity must not schedule GC."""
plan = _plan([{"usage_pct": 97}], tmp_path)
assert plan, "an unidentified threat must still be reported"
assert plan[0]["action"] == "report-only", plan
assert plan[0]["reason"], plan
def test_exclusion_entry_without_identity_fails_closed(tmp_path):
"""A mis-typed exclusion entry must break the run, never silently disable the gate."""
contract = tmp_path / "broken.prose.md"
contract.write_text(
"```yaml\n"
"report_only_guests:\n"
" - node: storepve\n"
" reason: \"typo - no guest identity\"\n"
"```\n"
)
scan_file = tmp_path / "scan.json"
scan_file.write_text(json.dumps([{"ct": 111, "usage_pct": 97}]))
proc = subprocess.run(
[sys.executable, str(PLAN), "--scan", str(scan_file),
"--contract", str(contract), "--json"],
capture_output=True, text=True,
)
assert proc.returncode != 0, proc.stdout
assert "gc-executor" not in proc.stdout
assert "identity" in proc.stderr.lower(), proc.stderr
def test_empty_exclusion_block_fails_closed(tmp_path):
"""An emptied report_only_guests list must break the run, not disable the gate."""
contract = tmp_path / "empty.prose.md"
contract.write_text("```yaml\nreport_only_guests: []\n```\n")
scan_file = tmp_path / "scan.json"
scan_file.write_text(json.dumps([{"ct": 111, "usage_pct": 97}]))
proc = subprocess.run(
[sys.executable, str(PLAN), "--scan", str(scan_file),
"--contract", str(contract), "--json"],
capture_output=True, text=True,
)
assert proc.returncode != 0, proc.stdout
assert "gc-executor" not in proc.stdout
assert "report-only gate" in proc.stderr.lower(), proc.stderr
def _plan_with_contract(scan, contract_text, tmp_path, name):
contract = tmp_path / name
contract.write_text(contract_text)
scan_file = tmp_path / f"scan-{name}.json"
scan_file.write_text(json.dumps(scan))
proc = subprocess.run(
[sys.executable, str(PLAN), "--scan", str(scan_file),
"--contract", str(contract), "--json"],
capture_output=True, text=True,
)
assert proc.returncode == 0, proc.stderr
return json.loads(proc.stdout)
def test_gate_is_read_from_the_contract_block(tmp_path):
"""The gate is data-driven by the contract block: the planner excludes the guest
when the block names it and acts on it when the block does not. Executes the real
planner interface against both fixtures so the behaviour change is observable."""
scan = [{"id": 111, "hostname": "tdunna", "ip": "192.168.68.129", "usage_pct": 95}]
with_gate = (
"```yaml\n"
"report_only_guests:\n"
" - guest: 111\n"
" hostname: tdunna\n"
" ip: 192.168.68.129\n"
" reason: \"fixture reason\"\n"
"```\n"
)
without_gate = (
"```yaml\n"
"report_only_guests:\n"
" - guest: 999\n"
" reason: \"fixture excludes a different guest\"\n"
"```\n"
)
gated = _plan_with_contract(scan, with_gate, tmp_path, "gated.prose.md")
assert gated[0]["action"] == "report-only", gated
assert gated[0]["reason"] == "fixture reason", gated
ungated = _plan_with_contract(scan, without_gate, tmp_path, "ungated.prose.md")
assert ungated[0]["action"] == "gc-executor", ungated
assert ungated[0]["action"] != gated[0]["action"]