root
96769a103f
no-mistakes(review): sync tanko CT 112 mapping to minipve across consumers
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 12s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 17s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 3s
2026-09-28 06:47:46 +00:00
root
c72436b406
no-mistakes(document): docs: reconcile zulip-monitor status in infrastructure-control
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-09-19 22:40:25 +00:00
root
30b2fe3fdc
Fix PR #112 security review - restore docs, remove live credentials
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 10s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-09-17 06:51:52 +00:00
root
5112c566c8
Remove Stirling PDF credentials (password + API key) from 2 files
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 15s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 12s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-09-17 06:40:59 +00:00
root
b4b5321011
fix: add hostname resolution warning and fix dmsetup field documentation
...
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
Add warning that bare PVE hostnames (acerpve, amdpve, etc.) resolve to VPS
via *.dns.sysloggh.net wildcard, not to actual nodes. List IP addresses:
- acerpve 192.168.68.9
- amdpve 192.168.68.15
- storepve 192.168.68.6
- minipve 192.168.68.12
- ocupve 192.168.68.5
Update acerpve example in backup preflight to include address (192.168.68.9).
Fix dmsetup comment to show full field order:
=start =length =thin-pool =transaction-id
=metadata_used/metadata_total =data_used/data_total
remaining fields are flags
Make it clear lvs command is the primary source for percentages, dmsetup is only for error-state check.
Incidents now include addresses: acerpve (192.168.68.9) and amdpve (192.168.68.15).
2026-09-15 12:09:59 +00:00
root
69940bc9eb
fix: correct backup preflight commands - lvs pve/data and dmsetup field documentation
...
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
Fix two errors in the PREFLIGHT section (measured on acerpve 2026-09-15):
1. lvs -o ... pve/data (not pve-data-tpool) - this is the PRIMARY check that yields percentages directly
- Quote the acerpve example: data 29.95% 1.22% <816.21g
2. dmsetup status pve-data-tpool - document fields correctly:
- = transaction ID (99), NOT data_percent
- = metadata used/total blocks
- = data used/total sectors
- Show how to derive percentages if needed
3. Keep the error-state check (grep -q 'Error|Fail') - this is how the incident presented
Everything else stays: 1777 tmpdir requirement with EACCES symptom, incidents as rationale,
GPU-host fact, --output-format json rule, honest note that metadata/snapshot pressure is unproven.
2026-09-15 11:42:10 +00:00
root
65eaffe1c6
fix: add backup safety preconditions - thin-pool headroom and tmpdir 1777
...
Add documented preflight checks for VM/CT backups on LVM thin-pool hosts:
- dmsetup status pve-data-tpool + lvs to verify data_percent < 90% and metadata_percent < 70%
- Exit 1 if pool shows Error/Fail state (takes down entire VG including host root)
- tmpdir must be mode 1777 (world-traversable) for vzdump archive step
- --output-format json for tasks started from truncating shells
- Document two incidents: acerpve thin-pool VM 101 (twice on 2026-09-13) and amdpve 0700 tmpdir (2026-09-14)
- Note metadata/snapshot-pressure hypothesis is UNPROVEN; preflight is the control
- Document GPU-host fact: VM 101 (llm-gpu) and VM 103 (ocu-llm) have no scheduled backup
2026-09-15 11:37:41 +00:00
root
f99f7e1e34
fix: restore per-host probe coverage + sweep residual retired names
...
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
A. litellm-health step 7: restore .8 → gpu-dense probe (was duplicated
to gpu-vision after 8210fd9 ). Add note documenting monitor key scope
gap for gpu-dense.
B. Sweep remaining retired names presented as usable:
- README.md:91: qwen3.6-27B-code → gpu-dense in runnable example
- gpu-fleet.prose.md:14: qwen3.6-35B-udq4 → Carnice-Qwen3.6-MoE...
- infrastructure-control.prose.md:226: qwen3.6-35B-udq4 → strix-moe
- proxmox-monitor.prose.md:90: qwen3.6-35B-udq4 → strix-moe
C. Audit test: 10/10 passed (retired raw names now hard-fail)
Fix-forward from 8210fd9 (direct master push).
2026-09-12 21:56:24 +00:00
root
4ea2d0309f
no-mistakes(review): Fix report-only gate, dedupe list, correct fleet map
2026-09-12 18:38:19 +00:00
root
78b501798f
no-mistakes(document): Sweep residual gemma labels; align compression rule contradiction
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Failing after 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
2026-09-12 16:54:41 +00:00
agent-zero
7730cc7c03
feat(litellm): 1.99.1 update coverage, trove agent, nginx /ui /docs path fixes
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
2026-09-11 15:32:20 -04:00
agent-zero
832b184af6
fix(litellm): upgrade contract refs 1.90.0-rc.1 -> 1.99.1 and correct redis role
...
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
- harness-litellm pinned tag updated to 1.99.1 in infrastructure-control,
litellm-health, litellm-self-heal container tables
- infrastructure-update MCP per-key limitation note now cites v1.99.1
- harness-redis role corrected: dead 'Router slots, circuit breakers' ->
'LiteLLM cache + rate-limit state' (router decommissioned 2026-09-11)
- litellm-self-heal Rule 9 wording clarified for cache/rate-limit only
Live verification 2026-09-11: harness-litellm healthy, /health 200/200,
7 models, 46 keys, prisma migrations 127 -> 157.
2026-09-11 15:05:00 -04:00
agent-zero
c26255f5ff
fix(router): decommission legacy GPU router across contracts and fleet scripts
...
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
Router container, image, and config are purged on CT 116 (verified: no
container, no image, inference-harness-router:latest removed, :9000 free,
11 containers healthy, 7 models, live syslog-auto completion OK). Updates:
- gpu-fleet.prose.md: topology diagram rebuilt without the router tier
- gpu-monitor / gpu-self-heal / litellm-health / litellm-self-heal: router steps dropped
- infrastructure-control.prose.md: container inventory + litellm row corrected
- scripts/daily-infra-report.py, scripts/prose-ai-review.sh: host-scoped checks
Verified: 40/40 anchors matched uniquely, 0 U+FFFD across changed files,
daily-infra-report.py compiles, no stray harness-router references remain.
2026-09-11 14:20:18 -04:00
root
79af0ae7a3
Merged PR #52 : fix/tanko-runtime
2026-08-30 12:16:28 +00:00
kagentz-bot
44f7008302
docs: reflect hwepve removal from Tabiri cluster (5 nodes + standalone London node)
...
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
- Remove hwepve from cluster member lists/diagrams in infrastructure-control,
proxmox-monitor, infrastructure-update, litellm-health, mumuni-delegation
- CTs 100 (abiba) and 105 (kagentz) moved to minipve; CT 114 (mumuni) no longer
exists (Mumuni runs inside Abiba CT 100)
- Add standalone London role + NetBird routing peer description for hwepve
- Update scripts: agent-health-check.py, pct-run.sh, prose-ai-review.sh
- Update disk-gc CT access table, hermes baselines/restore host references
2026-08-15 20:03:02 -04:00
root
8f719ca7c7
fix(contracts): Pulse port 3001→7655 (verified), jdownloader decommissioned from docker-vm → CT 118 LXC (.20)
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Failing after 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Skipped
2026-08-01 20:28:50 +00:00
root
b3193e5e1b
fix(agent-health): 14 stale-reference fixes (Mumuni CT100/.24, Koby .129, Koonimo CT113, wrapper paths)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-08-01 14:10:52 +00:00
root
87b4d67067
no-mistakes(review): Fix 5 review findings: duplicate table header, hwepve specs/Wave2 target, dangling ref, Authentik port
2026-07-24 21:31:28 +00:00
root
fa26b7a579
no-mistakes(test): Fixed 7 cross-table inconsistencies: kagentz placement, mumuni placement, stale 5-node references
2026-07-24 20:14:49 +00:00
root
b17c60f997
fix: contract accuracy updates post fleet-wide reboot
...
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
infrastructure-control.prose.md:
- Add hwepve as 6th Proxmox node
- Fix Gitea IP: .17 (was .110)
- Fix AdGuard IP: .10 on minipve (was .102 on acerpve)
- Fix Abiba placement: hwepve (was amdpve)
- Fix Mumuni placement: hwepve (was minipve)
- Fix Authentik port: add :9000
- Add CT 118 (jdownloader), CT 119 (infisical-vault)
- Add last-verified date (2026-07-24)
infrastructure-update.prose.md:
- Add post-reboot CT sweep procedure
- Add Zulip Docker network recovery steps
- Add WireGuard tunnel verification
scripts/netbird-add-domain.sh:
- New script to register domains in Netbird proxy store.db
2026-07-24 20:01:43 +00:00
root
1c44bf1259
contract updates: ornith decommissioned, 256K→128K context, Mumuni Discord disabled
...
Change 1: Strix Halo — ornith decommissioned
- gpu-fleet: Genesis Hermes V3 APEX → qwen3.6-35B-udq4 throughout
- inference-optimization: ornith→strix-moe/qwen3.6-35B-udq4
- gpu-monitor: ornith status → Strix Halo status
- infrastructure-control: Strix Halo — ornith → qwen3.6-35B-udq4
- infrastructure-update: ornith→strix-moe via router
- proxmox-monitor: Strix Halo LLM (ornith) → (qwen3.6-35B-udq4, strix-moe)
Change 2: GPU context 256K→128K fleet-wide
- hermes-agent-baseline: frontmatter description updated
- litellm-health: GPU Fleet Topology table 256K→128K
- litellm-self-heal: GPU Fleet Topology, engine flags, VRAM alert
- inference-optimization: compress threshold 256K→128K compact at 85K
- gpu-fleet: instability note updated
Change 3: Mumuni Discord platform disabled
- gpu-fleet: Mumuni platforms: removed discord
2026-07-23 09:00:42 +00:00
root
79d4a73895
docs: lessons learned from 2026-07-12 session
...
- gpu-fleet: Corrected architecture (direct GPU, no router in path).
Updated context values (RTX 3090=256K, not 128K). Added api-key
standardization requirement.
- gpu-self-heal: Added Lessons Learned section with 5 critical findings:
L1: API key standardization (RTX 5070 sk-loc...5678 vs not-needed)
L2: Fallback chain cascading failure loop detection
L3: Verify running state, not documentation
L4: Infisical fallback requirement (.env must have uncommented key)
L5: Zulip event queue can silently die after ~40 reconnects
- litellm-self-heal: Updated status manual-only→deployed, cron schedule
- litellm-api-keys: Added Infisical token expiry warning + .env fallback
- hermes-config-template: Rule 3 updated with .env fallback requirement
2026-07-12 22:49:39 +00:00
root
97977f320c
fix: contract improvements from 2026-07-09 run log review
...
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Failing after 14m43s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Has been skipped
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Has been skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Has been skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Has been skipped
- pct-run.sh: add CT 109 note (KVM VM), document decommissioned/migrated CTs
- disk-gc-threat-response: add docker-vm (SSH .7), add amdpve Docker scope,
remove jitsi (decommissioned), update access matrix with KVM VM section
- infrastructure-control: docker-vm container count 11→16, add Trove agents stack
- infrastructure-monitoring: deployment status updated (core live, GPU exporters not deployed)
2026-07-09 05:32:00 +00:00
root
5aa93117de
docs: capture all session changes — port conflict, health check, streaming
...
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Failing after 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Has been skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Has been skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Has been skipped
Updated contracts to reflect all infrastructure changes made 2026-07-05/06:
gpu-fleet.prose.md:
- Port conflict detection on all 3 GPU wrappers
- Ghost process root cause documented (.8 stale pid 25836)
infrastructure-control.prose.md:
- Section 7: Agent Health Check (consolidated, non-disruptive)
- Disabled scripts documented (zulip-watchdog, zulip-monitor)
hermes-agent-baseline.prose.md:
- Health verification section + GPU port conflict detection
- Change log updated
zulip-health.prose.md:
- Streaming support section (edit_message + PATCH API)
2026-07-06 02:46:41 +00:00
root
5c90eb8f6c
rename: Koonimo→Baggy, Koby→Tdunna — match hostnames
...
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
Standardize agent names across all contracts and pct-run.sh.
No functional changes — same CT IDs, same keys, same nodes.
2026-07-05 20:04:22 +00:00
root
b65116b847
feat: pct-run.sh — run commands in CTs by ID, no IPs needed
...
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Failing after 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Has been skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Has been skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Has been skipped
Adds /scripts/pct-run.sh that maps CT IDs to PVE nodes and runs
pct exec directly. No hardcoded IPs — Proxmox is the source of truth.
Updates infrastructure-control with CT access table using pct-run.
Usage examples:
pct-run 112 cat /etc/hostname # Tanko
pct-run 114 grep LITELLM /etc/environment # Mumuni key check
pct-run 116 docker ps # LiteLLM containers
2026-07-05 20:02:20 +00:00
root
0db95a4d1c
fix: correct IPs, CT IDs, and key status across all contracts
...
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Failing after 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Has been skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Has been skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Has been skipped
hermes-key-enforcement.prose.md:
- Add IP column to verified agents table
- Fix Tanko status: was using master key via systemd drop-in (not compliant)
- Add detection query step 2: check systemd drop-ins for master key leaks
- Add step 3: verify running process env against dedicated key
- Mark Mumuni, Koby, Koonimo as Unverified (last check was surface-level only)
gpu-fleet.prose.md:
- Add CT ID column to agent table for cross-reference
- Fix Abiba IP .24 → .65 (pi agent runs on RA-H OS host)
- Update Abiba LiteLLM key to match actual models.json key
- Set Koonimo IP to unknown (was .114 — incorrect per user)
- Fix pi-specific paths .24 → .65
hermes-config-template.prose.md:
- Set Koonimo IP to unknown
infrastructure-control.prose.md:
- Set Koonimo IP to unknown
- Fix Koonimo CT reference
2026-07-05 19:58:14 +00:00
root
263c4080f3
ci: add automated PR validation pipeline + fix regressions
...
NEW: Gitea Actions CI pipeline (.gitea/workflows/pr-pipeline.yaml)
- Stage 1 (validate): YAML frontmatter validation for all .prose.md files
- Stage 2 (lint): Structural checks + regression detection
* /grafana/ nginx route (reverted 2026-07-02)
* Stale CT IDs (122→112, no 123)
* Verified: .19/.122/.123 = correct bridge IPs
- Stage 3 (ai-review): LiteLLM-powered diff review against ground truth
- Stage 4 (auto-merge): Gate — safe to merge when all green
NEW: scripts/prose-lint.sh — contract structure + regression checker
NEW: scripts/prose-ai-review.sh — AI review via syslog-auto model
FIXES FOUND BY LINT:
- gpu-fleet.prose.md: removed /grafana/ from nginx routing diagram
- gpu-fleet.prose.md: removed /grafana/ from config file description
- infrastructure-control.prose.md: CT 122→112 in topology diagram
- scripts/prose-lint.sh: updated to not flag .19 (verified correct IP)
2026-07-04 22:13:51 +00:00
root
253e19f680
prose: review, fix, and consolidate infrastructure contracts
...
- DELETE infrastructure-monitoring.prose.md (redundant — proxmox-monitor already deployed this)
- FIX infrastructure-control.prose.md: remove /grafana/ nginx route (reverted Jul 2, broke gpu-fleet)
- FIX infrastructure-update.prose.md: wrong CT IDs (122→112 Tanko, 123→114 Mumuni, .19→storepve Zulip)
- FIX infrastructure-update.prose.md: Strix Halo :8080→router health check (firewalled to .116 only)
- FIX proxmox-monitor.prose.md: stale /grafana/ URLs→:3001 (post-revert cleanup)
- ADD disk-gc-threat-response.prose.md: verified run (2026-07-04, reclaimed 35.67GB from kagentz)
- ADD infrastructure-update.prose.md, pi-approval-architecture.prose.md, zulip-approval-fix.prose.md, zulip-self-heal.prose.md
- UPDATE hermes-config, litellm-health/self-heal, pm2-self-heal, zulip-* contracts
- ADD runs/20260704-005804-51f9a9 (disk-gc-threat-response run artifacts)
All contracts verified against live PVE API, nginx config, and firewall state.
2026-07-04 21:47:33 +00:00
Abiba
d31ae0c25a
Verify-before-mutate doctrine: annotate live-state fields, fix stale Zulip IP (.117→.19)
...
- Added FIELD TRUST convention to frontmatter (live-state vs policy fields)
- Marked credentials/IPs with VERIFY-BEFORE-USE labels
- Fixed 3 stale Zulip IP references (.117 → .19) that caused the 2026-07-03 outage
- Added verify-before-mutate skill reference
- See knowledge graph node #604 for post-mortem
2026-07-03 20:10:45 +00:00
Abiba
ae143d7a88
Add Stirling-PDF agent API access contract + skill. OAuth2 configured but disabled (needs license).
2026-07-03 19:10:22 +00:00
Abiba
8dfbad5415
Remove stale Bento-PDF reference from migration note
2026-07-03 18:20:08 +00:00
Abiba
c1bdbbc491
2026-07-03: Hermes key enforcement contract + Stirling-PDF docs + config standardization
...
- hermes-key-enforcement.prose.md: NEW contract enforcing api_key_env for all harness/LiteLLM providers, exempting external providers
- hermes-config-template.prose.md: Updated with full fix log (15 total fixes across Koby/Koonimo/Mumuni/Tanko), model:auto detection rule, key-to-agent mapping
- infrastructure-control.prose.md: Added Stirling-PDF service details (port 8989, credentials, API key, Swagger URL) replacing Bento-PDF
2026-07-03 18:19:41 +00:00
Abiba
b0c0331727
feat: infrastructure-control network-verification + IP-first doctrine, plus 3 supporting contracts
...
- infrastructure-control: add Section 5.3 (Network Verification & Routing) with
dual-path checks, NetBird ingress health, DNS drift detection, and routing
regression alerts. Add Section 7 (Configuration Doctrine — IP-First) mandating
LAN IPs for all configs/agents, URLs reserved for human browser access only.
Also enrich access matrix (Mumuni, Koonimo, LiteLLM Admin, Grafana entries),
reachability matrix, and CT 116 nginx routing + Prometheus targets.
- build-zulip-plugin: fold Success Criteria into Maintains postconditions.
- pi-approval-architecture: new reference doc — pi approval model vs Hermes,
architectural limits, /approve and /deny stub commands.
- zulip-approval-fix: new responsibility — documents the Zulip HTML/prefix bug
that broke /approve and /deny, and the one-line fix applied.
2026-07-02 20:39:40 +00:00
root
14bb145300
feat: expand infrastructure-control pattern — Proxmox 5-node cluster, Docker 3-ecosystem, NFS storage, network services
...
Full discovery and expansion based on PVE API exploration:
Proxmox (5 nodes):
- minipve (.12): authentik, gitea, mumuni, syslog-api, jitsi
- amdpve (.15): abiba, kagentz, tanko, tdunna, baggy, scottdenya
- storepve (.6): docker-vm, ra-h-os, PBS, media, zulip
- acerpve (.9): llm-gpu, adguard
- ocupve (.5): ocu-llm
Docker (3 ecosystems, 22 containers):
- docker-vm (.7): Firecrawl, SearXNG, Home stack, Audiobookshelf
- CT 116: LiteLLM stack (6 containers)
- Netbird (.17): VPN controller + dashboard + crowdsec + traefik
Pattern now includes: access matrix, per-section checks/remediations,
alert routing, escalation chain, quick-start commands, full CT inventory.
2026-06-27 13:11:55 +00:00
root
e090b27e20
feat: expand infrastructure-control pattern — Proxmox, Docker, storage, network
...
Full expansion of the infrastructure-control pattern covering:
Layer 1 — Proxmox Cluster (5 nodes)
- 4 monitoring rules (offline, resource pressure, stopped VMs/CTs, storage)
- Auto-start for 8 critical CTs/VMs on unexpected stop
- Node resource threshold alerting
Layer 2 — Docker Ecosystems (3 hosts, 22 containers)
- 6 monitoring rules across docker-vm, syslog-api, netbird
- Auto-restart for critical containers and stacks
- Disk pressure auto-prune
Layer 3 — Storage Fabric (3.6TB + 7.3TB NFS)
- Mount availability monitoring
- Growth rate tracking for high-use volumes
- PBS backup compliance checking
Layer 4 — Network Services (DNS, auth, git, chat, VPN)
- Service endpoint monitoring
- Escalation to Mumuni for authentik issues
Layer 5 — Agent Health (6 agents across 3 platforms)
- CT auto-start for stopped agents
- Health endpoint verification
Full access matrix and priority-based SLA table.
2026-06-27 13:11:15 +00:00
root
acfe16e5e7
feat: infrastructure-control — reference pattern for environment reach
...
Maps how OpenProse contracts connect to Gitea, containers, remote
agents, and infrastructure. Documents the access matrix, what works
today, and what gaps remain for full enforcement.
2026-06-26 18:52:54 +00:00