Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
fdb22948d9 | ||
|
|
44ba53cf71 | ||
|
|
b96b334283 | ||
|
|
dc829bb09d | ||
|
|
a8354e88bc | ||
|
|
98d8ce6772 | ||
|
|
548e421fa1 | ||
|
|
0134ad8be6 | ||
|
|
546ca79849 | ||
|
|
c31dbad95f | ||
|
|
d829f595b4 | ||
|
|
4164ee5ed0 |
@@ -160,8 +160,8 @@ Companion shell scripts that contracts delegate to.
|
|||||||
| File | Description |
|
| File | Description |
|
||||||
|---|---|
|
|---|---|
|
||||||
| `daily-infra-report.py` | Generates the daily infrastructure dashboard (HTML email to jerome@sysloggh.com). |
|
| `daily-infra-report.py` | Generates the daily infrastructure dashboard (HTML email to jerome@sysloggh.com). |
|
||||||
| `zulip-monitor.sh` | Monitors Zulip bot health and message flow. |
|
| `pm2-self-heal.sh` | Shell companion to the pm2-self-heal contract — restarts crashed PM2 processes (runs every 5 min). |
|
||||||
| `pm2-self-heal.sh` | Shell companion to the pm2-self-heal contract — restarts crashed PM2 processes. |
|
| `agent-health-check.py` | Consolidated agent health: LiteLLM key validation + GPU port conflict + streaming checks (every 10 min). Replaced zulip-monitor.sh. |
|
||||||
|
|
||||||
## Contract Structure
|
## Contract Structure
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,256 @@
|
|||||||
|
---
|
||||||
|
kind: pattern
|
||||||
|
name: delegation-prose-contract
|
||||||
|
description: >
|
||||||
|
Manager (Mumuni) operating doctrine for task decomposition, worker
|
||||||
|
delegation, verification, and delivery. Defines when to delegate, which
|
||||||
|
worker to use for what, how to handle failures, and the kanban board
|
||||||
|
protocol. Enforces context-window discipline and separation of concerns.
|
||||||
|
Runs on Mumuni (CT 118, storepve, .6) via Hermes agent.
|
||||||
|
version: 1.0.0
|
||||||
|
---
|
||||||
|
|
||||||
|
## Maintains
|
||||||
|
|
||||||
|
- Worker roster: 6 profiles (`syslog-code`, `syslog-devops`, `syslog-email`,
|
||||||
|
`syslog-research`, `syslog-review`, `syslog-writer`)
|
||||||
|
- Kanban board state at `~/.hermes/kanban/kanban.json`
|
||||||
|
- Context window budget: ~65K tokens per request (131K total, 60% threshold)
|
||||||
|
|
||||||
|
## Topology
|
||||||
|
|
||||||
|
**Cluster:** 5 Proxmox nodes (ocupve, acerpve, minipve, amdpve, storepve)
|
||||||
|
**Manager:** Mumuni (CT 118, storepve, .6) via Hermes agent
|
||||||
|
**Workers:** 6 profiles, all running on the same agent — no separate hosts needed
|
||||||
|
|
||||||
|
This contract is infrastructure-agnostic in terms of which nodes are used.
|
||||||
|
Workers execute tasks on whatever infrastructure they're given — SSH to .6,
|
||||||
|
.pm, .9, .12, or .15 depending on the task. The contract defines the
|
||||||
|
**who** and **when** — not the **where**.
|
||||||
|
|
||||||
|
## Why This Matters
|
||||||
|
|
||||||
|
Without enforced delegation, the manager consumes the full iteration budget
|
||||||
|
(60 calls) on single-turn tasks — SSH to 5 nodes, check each VM, read logs —
|
||||||
|
leaving no capacity for actual coordination. The result: context overflow
|
||||||
|
(59K tokens in system prompt), iteration exhaustion, and degraded response
|
||||||
|
quality. This contract exists because I blew through my budget checking
|
||||||
|
Proxmox node status instead of delegating to `syslog-devops`.
|
||||||
|
|
||||||
|
## Context Window Discipline
|
||||||
|
|
||||||
|
**The system prompt is ~6.5K tokens (stable: ~4.5K tool schemas + ~2K other guidance).**
|
||||||
|
**Volatile (MEMORY.md + USER.md): ~300 tokens.**
|
||||||
|
**Total base: ~6,800 tokens per request.**
|
||||||
|
|
||||||
|
The remaining budget is the conversation. Every tool call result adds to it.
|
||||||
|
If a single call returns >10K tokens (e.g., `grep` on a large file, SSH output
|
||||||
|
from multiple nodes), the context fills fast. That's why we delegate: workers
|
||||||
|
process in isolation and return compact results.
|
||||||
|
|
||||||
|
## Trigger Conditions
|
||||||
|
|
||||||
|
Delegation is **mandatory** when any of these apply:
|
||||||
|
|
||||||
|
| Condition | Threshold | Example |
|
||||||
|
|-----------|-----------|---------|
|
||||||
|
| Multiple tool calls needed | 2+ calls with intermediate logic | Read file → analyze → write report |
|
||||||
|
| Large data retrieval | Output >5K tokens | `grep -r "pattern" /path` on large dirs |
|
||||||
|
| Cross-domain work | Spans 2+ worker specialties | Infra check + email filter |
|
||||||
|
| Infrastructure changes | Any mutating operation | `qm set`, `systemctl restart`, `git push` |
|
||||||
|
| Research/analysis | Needs browser or deep reading | Web research, code review, data analysis |
|
||||||
|
| Code builds or changes | Writing or modifying code | Scripts, configs, patches |
|
||||||
|
| Sequential dependencies | Worker B needs Worker A's output | Code → Review → Deliver |
|
||||||
|
|
||||||
|
**Single tool calls stay at manager level.** Quick `grep`, `ls`, `cat`,
|
||||||
|
`curl`, `hermes tools list` — these are decision-making tools. The manager
|
||||||
|
reads them directly.
|
||||||
|
|
||||||
|
## Worker Selection Matrix
|
||||||
|
|
||||||
|
| Worker | Model | Toolsets | Role | Use When |
|
||||||
|
|--------|-------|----------|------|----------|
|
||||||
|
| `syslog-code` | qwen3.6-27B-code | terminal, file, web, memory, skills | Code patches, automation, scripts | Writing/modifying code, creating scripts, debugging, reading/writing files |
|
||||||
|
| `syslog-devops` | qwen3.6-27B-code | terminal, file, web, memory, skills | Infrastructure, DB, bridge, Proxmox | Server ops, SSH, Docker, Proxmox, DB queries, hardware checks |
|
||||||
|
| `syslog-email` | ornith-1.0-35b | terminal, file, web, memory, skills | Email automation, mail operations | Sending/receiving email, inbox management, SMTP operations |
|
||||||
|
| `syslog-research` | ornith-1.0-35b | terminal, file, web, memory, skills, **browser** | Analysis, classification, data processing | Web research, browser tasks, data analysis, classification, reading docs |
|
||||||
|
| `syslog-review` | ornith-1.0-35b | terminal, file, web, memory, skills | Verification, QA, audit validation | **ALWAYS** verify worker output before delivery — especially for infra changes, code builds, and research findings |
|
||||||
|
| `syslog-writer` | ornith-1.0-35b | terminal, file, web, memory, skills | Docs, content, branding, reports | Writing docs, reports, proposals, content, markdown formatting |
|
||||||
|
|
||||||
|
### Selection Rules
|
||||||
|
|
||||||
|
1. **Match specialty first.** A code task → `syslog-code`. An infra task →
|
||||||
|
`syslog-devops`. Don't put a `syslog-email` worker on a code review.
|
||||||
|
2. **Research tasks with browser needs → `syslog-research`.** Other workers
|
||||||
|
don't have the browser toolset.
|
||||||
|
3. **Verification → `syslog-review`.** Never deliver raw worker output.
|
||||||
|
4. **Documentation/content → `syslog-writer`.** Let them own the prose.
|
||||||
|
5. **If unsure, delegate to `syslog-research`** — it has the broadest toolset
|
||||||
|
(includes browser) and high reasoning effort.
|
||||||
|
|
||||||
|
## Delegation Protocol
|
||||||
|
|
||||||
|
### Step 1: Decompose
|
||||||
|
|
||||||
|
Break the task into lanes. Each lane does ONE thing. Workers are independent —
|
||||||
|
no lane depends on another's output mid-flight. If lanes depend on each other,
|
||||||
|
dispatch sequentially.
|
||||||
|
|
||||||
|
### Step 2: Dispatch
|
||||||
|
|
||||||
|
Fire workers via `delegate_task`:
|
||||||
|
|
||||||
|
**Parallel (independent lanes):**
|
||||||
|
```
|
||||||
|
delegate_task(
|
||||||
|
tasks=[
|
||||||
|
{"goal": "Check all 5 Proxmox nodes for VM status", "context": "SSH to each node via 192.168.68.x, run 'qm list'"},
|
||||||
|
{"goal": "Check Docker container health on .7/.116/.17", "context": "SSH to each host, check container status"},
|
||||||
|
]
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
**Sequential (dependent lanes):**
|
||||||
|
Dispatch lane 1 → wait for result → dispatch lane 2.
|
||||||
|
|
||||||
|
### Step 3: Verify
|
||||||
|
|
||||||
|
**MANDATORY for:**
|
||||||
|
- Infrastructure changes (any `qm`, `pct`, `systemctl`, `git push`)
|
||||||
|
- Code builds and modifications
|
||||||
|
- Research findings (web data, external sources)
|
||||||
|
- Any output that will reach the user
|
||||||
|
|
||||||
|
**Fire `syslog-review` to verify:**
|
||||||
|
```
|
||||||
|
delegate_task(
|
||||||
|
goal="Review the output of the devops worker. Verify the node status
|
||||||
|
report is accurate, check for inconsistencies, confirm all nodes were
|
||||||
|
reachable.",
|
||||||
|
context="Worker was syslog-devops. Output is at /tmp/node-report.md.
|
||||||
|
Verify against live system."
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
**If verification fails:**
|
||||||
|
1. Send work back to original worker with review feedback
|
||||||
|
2. Re-verify
|
||||||
|
3. Max 2 re-verify cycles before escalating to Kwame
|
||||||
|
|
||||||
|
### Step 4: Deliver
|
||||||
|
|
||||||
|
Only verified results reach Kwame. Format per channel:
|
||||||
|
- Telegram: Use `telegram-formatting` skill
|
||||||
|
- Zulip: Use Zulip Markdown (CommonMark)
|
||||||
|
- Email: Use `syslog-email` skill
|
||||||
|
|
||||||
|
## Kanban Board Protocol
|
||||||
|
|
||||||
|
**File:** `~/.hermes/kanban/kanban.json`
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"task_id": "unique-id",
|
||||||
|
"title": "Task description",
|
||||||
|
"created": "2026-07-09T01:00:00",
|
||||||
|
"status": "backlog|in_progress|review|done",
|
||||||
|
"lanes": [
|
||||||
|
{
|
||||||
|
"lane_id": "devops-check",
|
||||||
|
"worker": "syslog-devops",
|
||||||
|
"goal": "Check all 5 Proxmox nodes",
|
||||||
|
"status": "dispatched|completed|failed",
|
||||||
|
"output_file": "/tmp/node-report.md"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Update the board on every state change.**
|
||||||
|
|
||||||
|
## Failure Handling
|
||||||
|
|
||||||
|
### Worker Timeouts
|
||||||
|
|
||||||
|
- Child timeout: **900 seconds** (15 minutes)
|
||||||
|
- Worker model `syslog-auto` is slow — it can hit the timeout limit with
|
||||||
|
22+ API calls
|
||||||
|
- **If a worker times out:** Re-dispatch with a narrower scope. Break the
|
||||||
|
task into smaller pieces that fit in the timeout window.
|
||||||
|
- **Avoid delegating sequential SSH hops** — each SSH connection adds latency
|
||||||
|
that compounds quickly. Prefer API-based or local approaches when possible.
|
||||||
|
|
||||||
|
### Worker Selection Failures
|
||||||
|
|
||||||
|
- `syslog-devops` is best for infrastructure tasks (SSH, Proxmox, Docker)
|
||||||
|
- `syslog-code` is best for code-level work (reading files, writing scripts)
|
||||||
|
- `syslog-research` has the browser toolset — use for web research
|
||||||
|
- `syslog-review` is the QA gate — always fire before delivery
|
||||||
|
- **Never fire more than 3 parallel workers** (max_concurrent_children: 3)
|
||||||
|
- **Never nest delegation** (max_spawn_depth: 1)
|
||||||
|
|
||||||
|
### Context Overflow
|
||||||
|
|
||||||
|
- If a task requires >10K tokens of output, delegate the processing
|
||||||
|
- Workers return compact summaries, not raw data dumps
|
||||||
|
- Pass file paths and concrete goals — never dump raw data into context
|
||||||
|
|
||||||
|
## Anti-patterns
|
||||||
|
|
||||||
|
- ❌ Reading large files into your own context before deciding → delegate the read
|
||||||
|
- ❌ Carrying SSH/grep/output results in your context → delegate the analysis
|
||||||
|
- ❌ Doing work yourself and then "pretending" to delegate → the user can tell
|
||||||
|
- ❌ Skipping verification → raw worker output never reaches the user
|
||||||
|
- ❌ Delegating single tool calls → keep quick reads/writes at manager level
|
||||||
|
- ❌ Firing more than 3 workers in parallel → hard limit
|
||||||
|
|
||||||
|
## Emergency Exception
|
||||||
|
|
||||||
|
**In an emergency (server down, service must be restored immediately):**
|
||||||
|
- Delegate the diagnosis (find the problem)
|
||||||
|
- Execute the fix yourself (minimize handoff latency)
|
||||||
|
- Verify the fix after delivery
|
||||||
|
- Log the exception in the kanban board
|
||||||
|
|
||||||
|
The emergency exception exists because the user needs the service back NOW,
|
||||||
|
not after three worker round-trips. But it's an exception — not the rule.
|
||||||
|
|
||||||
|
## What This Contract Doesn't Cover
|
||||||
|
|
||||||
|
1. **Worker profile configuration** — covered by `hermes-config-template.prose.md`
|
||||||
|
2. **SSH key management** — covered by existing SSH/Proxmox contracts
|
||||||
|
3. **Git workflow** — covered by `AGENTS.md` in the prose-contracts repo
|
||||||
|
4. **Cron job management** — covered by individual cron contracts
|
||||||
|
5. **Infra verification** — covered by `verify-before-mutate` protocol
|
||||||
|
|
||||||
|
## Verification
|
||||||
|
|
||||||
|
Run `scripts/worker-audit.py` to verify all 6 profiles are aligned:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python3 /root/.hermes/skills/kanban-orchestrator/scripts/worker-audit.py
|
||||||
|
```
|
||||||
|
|
||||||
|
## References
|
||||||
|
|
||||||
|
- `kanban-orchestrator` skill: The operational playbook (detailed execution steps)
|
||||||
|
- `worker-profile-audit.md` (skill reference): Worker configuration audit notes
|
||||||
|
- `delegation-timeout-patterns.md` (skill reference): Timeout handling patterns
|
||||||
|
- `verify-before-mutate` protocol: Infrastructure change verification
|
||||||
|
- `hermes-config-template.prose.md`: Worker profile configuration
|
||||||
|
|
||||||
|
## Success Criteria
|
||||||
|
|
||||||
|
This contract succeeds when:
|
||||||
|
|
||||||
|
1. **No context overflow** — single-turn tasks don't exhaust the iteration budget
|
||||||
|
2. **Workers do the work** — manager coordinates, doesn't execute
|
||||||
|
3. **Verification before delivery** — all output passes through `syslog-review`
|
||||||
|
4. **Kanban board is current** — every task has a lane, every lane has a status
|
||||||
|
5. **User gets verified results** — raw worker output never reaches Kwame
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
**Last updated:** 2026-07-09
|
||||||
|
**Author:** Mumuni (with Kwame's input on triggers and exception criteria)
|
||||||
|
**Status:** Draft — awaiting PR review and merge to prose-contracts main
|
||||||
+5
-5
@@ -142,7 +142,7 @@ Show full fleet status: GPUs, models, VRAM, context windows, parallel slots, act
|
|||||||
|-------|-----|-----|-----|--------|
|
|-------|-----|-----|-----|--------|
|
||||||
| Tanko | 112 | .122 | `sk-CggiHWlamQyShxWC3Hx6uw` | SSH jerome |
|
| Tanko | 112 | .122 | `sk-CggiHWlamQyShxWC3Hx6uw` | SSH jerome |
|
||||||
| Mumuni | 114 | .123 | `sk-VrqCNlwUgzoNGOpikJ7nwQ` | SSH root |
|
| Mumuni | 114 | .123 | `sk-VrqCNlwUgzoNGOpikJ7nwQ` | SSH root |
|
||||||
| Abiba | 100 | .65 | `sk-Qvzi4uYQBhlSK_XstEhcyQ` | local (pi agent) |
|
| Abiba | 100 | .24 | `sk-Qvzi4uYQBhlSK_XstEhcyQ` | local (pi agent) |
|
||||||
| Tdunna | 111 | ? | `sk-6sbCNjz2T6lTVDBdlNHXsA` | Zulip DM |
|
| Tdunna | 111 | ? | `sk-6sbCNjz2T6lTVDBdlNHXsA` | Zulip DM |
|
||||||
| Baggy | 113 | ? | `sk-krnw_zGBwvvL5b7l2t-s-A` | no SSH |
|
| Baggy | 113 | ? | `sk-krnw_zGBwvvL5b7l2t-s-A` | no SSH |
|
||||||
| Kagenz0 | 105 | ? | `sk-Dh4CDkaHebMLEp8qqq20qA` | no SSH |
|
| Kagenz0 | 105 | ? | `sk-Dh4CDkaHebMLEp8qqq20qA` | no SSH |
|
||||||
@@ -159,10 +159,10 @@ If no SSH access, send Zulip DM via abiba-bot.
|
|||||||
| `/opt/inference-harness/router/router.py` | CT 116 | Router source (builds via compose) |
|
| `/opt/inference-harness/router/router.py` | CT 116 | Router source (builds via compose) |
|
||||||
| `/etc/nginx/nginx.conf` | CT 116 (nginx container) | Routes /v1→LiteLLM, /dashboard/, /litellm/, /health |
|
| `/etc/nginx/nginx.conf` | CT 116 (nginx container) | Routes /v1→LiteLLM, /dashboard/, /litellm/, /health |
|
||||||
| `/opt/monitoring/prometheus.yml` | CT 116 | Prometheus scrape config (5 targets) |
|
| `/opt/monitoring/prometheus.yml` | CT 116 | Prometheus scrape config (5 targets) |
|
||||||
| `/root/scripts/gpu-monitor-server.py` | pi (.65) | GPU fleet monitor v2.1.0 (with benchmarks) |
|
| `/root/scripts/gpu-monitor-server.py` | pi (.24) | GPU fleet monitor v2.1.0 (with benchmarks) |
|
||||||
| `/root/scripts/gpu_benchmark.py` | pi (.65) | GPU inference benchmark module (tok/s tracking) |
|
| `/root/scripts/gpu_benchmark.py` | pi (.24) | GPU inference benchmark module (tok/s tracking) |
|
||||||
| `/root/scripts/gpu-saturation-watchdog.py` | pi (.65) | Auto-restart stuck llama-server |
|
| `/root/scripts/gpu-saturation-watchdog.py` | pi (.24) | Auto-restart stuck llama-server |
|
||||||
| `/root/dashboard/gpu-fleet.html` | pi (.65) | Live HTML dashboard |
|
| `/root/dashboard/gpu-fleet.html` | pi (.24) | Live HTML dashboard |
|
||||||
| `/etc/systemd/system/llama-server.service` | .8, .110 | llama-server daemons (Nvidia GPUs) |
|
| `/etc/systemd/system/llama-server.service` | .8, .110 | llama-server daemons (Nvidia GPUs) |
|
||||||
| `/etc/systemd/system/ornith-server.service` | .15 (amdpve) | llama-server daemon (Vulkan, Strix Halo). Note: `llama-server.service` and `llama-server@.service` are **masked** on .15 to prevent port 8080 collisions. |
|
| `/etc/systemd/system/ornith-server.service` | .15 (amdpve) | llama-server daemon (Vulkan, Strix Halo). Note: `llama-server.service` and `llama-server@.service` are **masked** on .15 to prevent port 8080 collisions. |
|
||||||
|
|
||||||
|
|||||||
@@ -19,7 +19,7 @@ agent: abiba
|
|||||||
|
|
||||||
```
|
```
|
||||||
┌─────────────────────────────────────────────────────────┐
|
┌─────────────────────────────────────────────────────────┐
|
||||||
│ GPU Monitor Server (:9100) — 192.168.68.65 │
|
│ GPU Monitor Server (:9100) — 192.168.68.24 │
|
||||||
│ Polls every 15s, serves dashboard + JSON API │
|
│ Polls every 15s, serves dashboard + JSON API │
|
||||||
└──┬──────────┬──────────┬────────────────────────────────┘
|
└──┬──────────┬──────────┬────────────────────────────────┘
|
||||||
│ │ │
|
│ │ │
|
||||||
@@ -35,7 +35,7 @@ agent: abiba
|
|||||||
Note: JSON sidecar exporters at :8090 were never deployed on any
|
Note: JSON sidecar exporters at :8090 were never deployed on any
|
||||||
GPU host. Router falls back to GPU /health direct probe. Monitor
|
GPU host. Router falls back to GPU /health direct probe. Monitor
|
||||||
should use router /health/unified as source of truth for GPU status.
|
should use router /health/unified as source of truth for GPU status.
|
||||||
Strix Halo :8080 is firewalled to .116 only — monitor on .65 cannot
|
Strix Halo :8080 is firewalled to .116 only — monitor on .24 cannot
|
||||||
poll .15:8080 directly; must go through router on .116.
|
poll .15:8080 directly; must go through router on .116.
|
||||||
```
|
```
|
||||||
|
|
||||||
@@ -118,8 +118,8 @@ The router health is accessed through nginx on port 80 (NOT port 9000 directly).
|
|||||||
|
|
||||||
| File | Host | Purpose |
|
| File | Host | Purpose |
|
||||||
|------|------|---------|
|
|------|------|---------|
|
||||||
| `/root/scripts/gpu-monitor-server.py` | pi (.65) | Monitor server v2.0.0 |
|
| `/root/scripts/gpu-monitor-server.py` | pi (.24) | Monitor server v2.0.0 |
|
||||||
| `/root/dashboard/gpu-fleet.html` | pi (.65) | Live HTML dashboard |
|
| `/root/dashboard/gpu-fleet.html` | pi (.24) | Live HTML dashboard |
|
||||||
| `/etc/nginx/nginx.conf` | CT 116 | Routes /health/* → router |
|
| `/etc/nginx/nginx.conf` | CT 116 | Routes /health/* → router |
|
||||||
|
|
||||||
## Execution
|
## Execution
|
||||||
|
|||||||
@@ -24,10 +24,11 @@ version: 1.1.0
|
|||||||
|
|
||||||
Before ANY update wave:
|
Before ANY update wave:
|
||||||
1. ✅ All Proxmox nodes online (`GET /api2/json/nodes`)
|
1. ✅ All Proxmox nodes online (`GET /api2/json/nodes`)
|
||||||
2. ✅ All critical VMs/CTs running (VM 101 llm-gpu, VM 103 ocu-llm, VM 109 docker-vm, CT 116 syslog-api, CT 117 zulip, CT 106 ra-h-os)
|
2. ✅ All critical VMs/CTs running (VM 109 docker-vm, CT 116 syslog-api, CT 117 zulip, CT 106 ra-h-os)
|
||||||
3. ✅ Docker healthy on VM 109 (.7), CT 116 (.116)
|
3. ✅ GPU bare-metal hosts reachable: .8 (RTX 3090), .110 (RTX 5070), .15 (Strix Halo)
|
||||||
4. ✅ LiteLLM health check passing
|
4. ✅ Docker healthy on VM 109 (.7), CT 116 (.116)
|
||||||
5. ✅ Zulip server reachable
|
5. ✅ LiteLLM health check passing
|
||||||
|
6. ✅ Zulip server reachable
|
||||||
6. ✅ GPU fleet healthy (all 3 GPUs: RTX 3090, RTX 5070, RX 7600)
|
6. ✅ GPU fleet healthy (all 3 GPUs: RTX 3090, RTX 5070, RX 7600)
|
||||||
7. ✅ Disk >20% free on all nodes
|
7. ✅ Disk >20% free on all nodes
|
||||||
8. 📋 Snapshot critical configs (LiteLLM, nginx, docker-compose files)
|
8. 📋 Snapshot critical configs (LiteLLM, nginx, docker-compose files)
|
||||||
|
|||||||
@@ -81,6 +81,8 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
|||||||
| harness-dashboard | inference-harness-dashboard | :3000 | /health |
|
| harness-dashboard | inference-harness-dashboard | :3000 | /health |
|
||||||
| harness-grafana | grafana/grafana | :3000→:3001 | /api/health |
|
| harness-grafana | grafana/grafana | :3000→:3001 | /api/health |
|
||||||
| harness-prometheus | prom/prometheus | :9090 | /-/healthy |
|
| harness-prometheus | prom/prometheus | :9090 | /-/healthy |
|
||||||
|
| harness-docker-stats | python:3.12-alpine | — | container stats exporter |
|
||||||
|
| harness-pve-exporter | prompve/prometheus-pve-exporter | — | Proxmox metrics → Prometheus |
|
||||||
|
|
||||||
## Maintains
|
## Maintains
|
||||||
|
|
||||||
@@ -111,9 +113,10 @@ Run this first on every cycle. Results feed into remediation rules below.
|
|||||||
- GET http://{{backend_host}}/litellm/health/liveliness → expect 200
|
- GET http://{{backend_host}}/litellm/health/liveliness → expect 200
|
||||||
|
|
||||||
### 3. Check backend container health
|
### 3. Check backend container health
|
||||||
- SSH to {{backend_host}} → `docker ps` → verify 8 containers healthy
|
- SSH to {{backend_host}} → `docker ps` → verify 10 containers healthy
|
||||||
- Critical: harness-litellm, harness-router, harness-nginx, harness-postgres
|
- Critical: harness-litellm, harness-nginx, harness-postgres
|
||||||
- Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus
|
- Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus, harness-docker-stats, harness-pve-exporter
|
||||||
|
- Deprecated but running: harness-router (not in path, reference only)
|
||||||
|
|
||||||
### 4. Check GPU fleet health (via fleet dashboard)
|
### 4. Check GPU fleet health (via fleet dashboard)
|
||||||
- GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON
|
- GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON
|
||||||
|
|||||||
@@ -0,0 +1,290 @@
|
|||||||
|
---
|
||||||
|
kind: pattern
|
||||||
|
name: mumuni-delegation
|
||||||
|
description: >
|
||||||
|
Mumuni-specific operating doctrine for task decomposition, worker
|
||||||
|
delegation, verification, and delivery. Defines when to delegate, which
|
||||||
|
worker to use for what, how to handle failures, and the kanban board
|
||||||
|
protocol. Enforces context-window discipline and separation of concerns.
|
||||||
|
Runs on Mumuni (CT 118, storepve, .6) via Hermes agent.
|
||||||
|
version: 1.0.0
|
||||||
|
---
|
||||||
|
|
||||||
|
## Maintains
|
||||||
|
|
||||||
|
- Worker roster: 6 profiles (`syslog-code`, `syslog-devops`, `syslog-email`,
|
||||||
|
`syslog-research`, `syslog-review`, `syslog-writer`)
|
||||||
|
- Kanban board state at `~/.hermes/kanban/kanban.json`
|
||||||
|
- Context window budget: ~65K tokens per request (131K total, 60% threshold)
|
||||||
|
|
||||||
|
## Topology
|
||||||
|
|
||||||
|
**Cluster:** 5 Proxmox nodes (ocupve, acerpve, minipve, amdpve, storepve)
|
||||||
|
**Manager:** Mumuni (CT 118, storepve, .6) via Hermes agent
|
||||||
|
**Workers:** 6 profiles, all running on the same agent — no separate hosts needed
|
||||||
|
|
||||||
|
This contract is infrastructure-agnostic in terms of which nodes are used.
|
||||||
|
Workers execute tasks on whatever infrastructure they're given — SSH to .6,
|
||||||
|
.pm, .9, .12, or .15 depending on the task. The contract defines the
|
||||||
|
**who** and **when** — not the **where**.
|
||||||
|
|
||||||
|
## Why This Matters
|
||||||
|
|
||||||
|
Without enforced delegation, the manager consumes the full iteration budget
|
||||||
|
(60 calls) on single-turn tasks — SSH to 5 nodes, check each VM, read logs —
|
||||||
|
leaving no capacity for actual coordination. The result: context overflow
|
||||||
|
(59K tokens in system prompt), iteration exhaustion, and degraded response
|
||||||
|
quality. This contract exists because I blew through my budget checking
|
||||||
|
Proxmox node status instead of delegating to `syslog-devops`.
|
||||||
|
|
||||||
|
## Context Window Discipline
|
||||||
|
|
||||||
|
**The system prompt is ~6.5K tokens (stable: ~4.5K tool schemas + ~2K other guidance).**
|
||||||
|
**Volatile (MEMORY.md + USER.md): ~300 tokens.**
|
||||||
|
**Total base: ~6,800 tokens per request.**
|
||||||
|
|
||||||
|
The remaining budget is the conversation. Every tool call result adds to it.
|
||||||
|
If a single call returns >10K tokens (e.g., `grep` on a large file, SSH output
|
||||||
|
from multiple nodes), the context fills fast. That's why we delegate: workers
|
||||||
|
process in isolation and return compact results.
|
||||||
|
|
||||||
|
## Trigger Conditions
|
||||||
|
|
||||||
|
Delegation is **mandatory** when any of these apply:
|
||||||
|
|
||||||
|
| Condition | Threshold | Example |
|
||||||
|
|-----------|-----------|---------|
|
||||||
|
| Multiple tool calls needed | 2+ calls with intermediate logic | Read file → analyze → write report |
|
||||||
|
| Large data retrieval | Output >5K tokens | `grep -r "pattern" /path` on large dirs |
|
||||||
|
| Cross-domain work | Spans 2+ worker specialties | Infra check + email filter |
|
||||||
|
| Infrastructure changes | Any mutating operation | `qm set`, `systemctl restart`, `git push` |
|
||||||
|
| Research/analysis | Needs browser or deep reading | Web research, code review, data analysis |
|
||||||
|
| Code builds or changes | Writing or modifying code | Scripts, configs, patches |
|
||||||
|
| Sequential dependencies | Worker B needs Worker A's output | Code → Review → Deliver |
|
||||||
|
|
||||||
|
**Single tool calls stay at manager level.** Quick `grep`, `ls`, `cat`,
|
||||||
|
`curl`, `hermes tools list` — these are decision-making tools. The manager
|
||||||
|
reads them directly.
|
||||||
|
|
||||||
|
## Data Source Integrity (CRITICAL)
|
||||||
|
|
||||||
|
**Workers MUST use the data provided in their task context. They MUST NOT
|
||||||
|
fetch their own data from external sources unless explicitly told to.**
|
||||||
|
|
||||||
|
When a task says "Read file X and format it", the worker reads file X. It does
|
||||||
|
not query a separate API, run its own diagnostics, or pull data from a different
|
||||||
|
system. This is the #1 source of cross-worker inconsistency: one worker gathers
|
||||||
|
SSH data, another queries the Proxmox API, and the report merges two incompatible
|
||||||
|
datasets.
|
||||||
|
|
||||||
|
**Rule:** If a worker needs additional data beyond what's in its task description,
|
||||||
|
it asks the manager (via relay) — it doesn't go find it on its own.
|
||||||
|
|
||||||
|
**This is a hard rule, not a recommendation.** Violating it produces the exact
|
||||||
|
type of discrepancy the kanban pipeline exists to prevent: a review worker finds
|
||||||
|
"5 nodes present" in the raw data but "5/5 online" in the report — even though
|
||||||
|
one of those nodes was unreachable. The report lied because it used data the
|
||||||
|
raw data never provided.
|
||||||
|
|
||||||
|
## Worker Selection Matrix
|
||||||
|
|
||||||
|
| Worker | Model | Toolsets | Role | Use When |
|
||||||
|
|--------|-------|----------|------|----------|
|
||||||
|
| `syslog-code` | qwen3.6-27B-code | terminal, file, web, memory, skills | Code patches, automation, scripts | Writing/modifying code, creating scripts, debugging, reading/writing files |
|
||||||
|
| `syslog-devops` | qwen3.6-27B-code | terminal, file, web, memory, skills | Infrastructure, DB, bridge, Proxmox | Server ops, SSH, Docker, Proxmox, DB queries, hardware checks |
|
||||||
|
| `syslog-email` | ornith-1.0-35b | terminal, file, web, memory, skills | Email automation, mail operations | Sending/receiving email, inbox management, SMTP operations |
|
||||||
|
| `syslog-research` | ornith-1.0-35b | terminal, file, web, memory, skills, **browser** | Analysis, classification, data processing | Web research, browser tasks, data analysis, classification, reading docs |
|
||||||
|
| `syslog-review` | ornith-1.0-35b | terminal, file, web, memory, skills | Verification, QA, audit validation | **ALWAYS** verify worker output before delivery — especially for infra changes, code builds, and research findings |
|
||||||
|
| `syslog-writer` | ornith-1.0-35b | terminal, file, web, memory, skills | Docs, content, branding, reports | Writing docs, reports, proposals, content, markdown formatting |
|
||||||
|
|
||||||
|
### Selection Rules
|
||||||
|
|
||||||
|
1. **Match specialty first.** A code task → `syslog-code`. An infra task →
|
||||||
|
`syslog-devops`. Don't put a `syslog-email` worker on a code review.
|
||||||
|
2. **Research tasks with browser needs → `syslog-research`.** Other workers
|
||||||
|
don't have the browser toolset.
|
||||||
|
3. **Verification → `syslog-review`.** Never deliver raw worker output.
|
||||||
|
4. **Documentation/content → `syslog-writer`.** Let them own the prose.
|
||||||
|
5. **If unsure, delegate to `syslog-research`** — it has the broadest toolset
|
||||||
|
(includes browser) and high reasoning effort.
|
||||||
|
|
||||||
|
## Delegation Protocol
|
||||||
|
|
||||||
|
### Step 1: Decompose
|
||||||
|
|
||||||
|
Break the task into lanes. Each lane does ONE thing. Workers are independent —
|
||||||
|
no lane depends on another's output mid-flight. If lanes depend on each other,
|
||||||
|
dispatch sequentially.
|
||||||
|
|
||||||
|
### Step 2: Dispatch
|
||||||
|
|
||||||
|
Fire workers via `delegate_task`:
|
||||||
|
|
||||||
|
**Critical: Pass the data, not just the goal.** When dispatching a worker that
|
||||||
|
processes output from another worker, include the file path AND explicit
|
||||||
|
instructions to use ONLY that source. Example:
|
||||||
|
|
||||||
|
```
|
||||||
|
delegate_task(
|
||||||
|
goal="Format the cluster check into a clean report",
|
||||||
|
context="Source data is at /tmp/proxmox-check-raw.md. Format ONLY the data
|
||||||
|
in that file. Do NOT query the Proxmox API or any other data source. Use the
|
||||||
|
file as your sole source of truth."
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
**Parallel (independent lanes):**
|
||||||
|
```
|
||||||
|
delegate_task(
|
||||||
|
tasks=[
|
||||||
|
{"goal": "Check all 5 Proxmox nodes for VM status", "context": "SSH to each node via 192.168.68.x, run 'qm list'"},
|
||||||
|
{"goal": "Check Docker container health on .7/.116/.17", "context": "SSH to each host, check container status"},
|
||||||
|
]
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
**Sequential (dependent lanes):**
|
||||||
|
Dispatch lane 1 → wait for result → dispatch lane 2.
|
||||||
|
|
||||||
|
### Step 3: Verify
|
||||||
|
|
||||||
|
**MANDATORY for:**
|
||||||
|
- Infrastructure changes (any `qm`, `pct`, `systemctl`, `git push`)
|
||||||
|
- Code builds and modifications
|
||||||
|
- Research findings (web data, external sources)
|
||||||
|
- Any output that will reach the user
|
||||||
|
|
||||||
|
**Fire `syslog-review` to verify:**
|
||||||
|
```
|
||||||
|
delegate_task(
|
||||||
|
goal="Review the output of the devops worker. Verify the node status
|
||||||
|
report is accurate, check for inconsistencies, confirm all nodes were
|
||||||
|
reachable.",
|
||||||
|
context="Worker was syslog-devops. Output is at /tmp/node-report.md.
|
||||||
|
Verify against live system."
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
**If verification fails:**
|
||||||
|
1. Send work back to original worker with review feedback
|
||||||
|
2. Re-verify
|
||||||
|
3. Max 2 re-verify cycles before escalating to Kwame
|
||||||
|
|
||||||
|
### Step 4: Deliver
|
||||||
|
|
||||||
|
Only verified results reach Kwame. Format per channel:
|
||||||
|
- Telegram: Use `telegram-formatting` skill
|
||||||
|
- Zulip: Use Zulip Markdown (CommonMark)
|
||||||
|
- Email: Use `syslog-email` skill
|
||||||
|
|
||||||
|
## Kanban Board Protocol
|
||||||
|
|
||||||
|
**File:** `~/.hermes/kanban/kanban.json`
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"task_id": "unique-id",
|
||||||
|
"title": "Task description",
|
||||||
|
"created": "2026-07-09T01:00:00",
|
||||||
|
"status": "backlog|in_progress|review|done",
|
||||||
|
"lanes": [
|
||||||
|
{
|
||||||
|
"lane_id": "devops-check",
|
||||||
|
"worker": "syslog-devops",
|
||||||
|
"goal": "Check all 5 Proxmox nodes",
|
||||||
|
"status": "dispatched|completed|failed",
|
||||||
|
"output_file": "/tmp/node-report.md"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Update the board on every state change.**
|
||||||
|
|
||||||
|
## Failure Handling
|
||||||
|
|
||||||
|
### Worker Timeouts
|
||||||
|
|
||||||
|
- Child timeout: **900 seconds** (15 minutes)
|
||||||
|
- Worker model `syslog-auto` is slow — it can hit the timeout limit with
|
||||||
|
22+ API calls
|
||||||
|
- **If a worker times out:** Re-dispatch with a narrower scope. Break the
|
||||||
|
task into smaller pieces that fit in the timeout window.
|
||||||
|
- **Avoid delegating sequential SSH hops** — each SSH connection adds latency
|
||||||
|
that compounds quickly. Prefer API-based or local approaches when possible.
|
||||||
|
|
||||||
|
### Worker Selection Failures
|
||||||
|
|
||||||
|
- `syslog-devops` is best for infrastructure tasks (SSH, Proxmox, Docker)
|
||||||
|
- `syslog-code` is best for code-level work (reading files, writing scripts)
|
||||||
|
- `syslog-research` has the browser toolset — use for web research
|
||||||
|
- `syslog-review` is the QA gate — always fire before delivery
|
||||||
|
- **Never fire more than 3 parallel workers** (max_concurrent_children: 3)
|
||||||
|
- **Never nest delegation** (max_spawn_depth: 1)
|
||||||
|
|
||||||
|
### Context Overflow
|
||||||
|
|
||||||
|
- If a task requires >10K tokens of output, delegate the processing
|
||||||
|
- Workers return compact summaries, not raw data dumps
|
||||||
|
- Pass file paths and concrete goals — never dump raw data into context
|
||||||
|
|
||||||
|
## Anti-patterns
|
||||||
|
|
||||||
|
- ❌ Reading large files into your own context before deciding → delegate the read
|
||||||
|
- ❌ Carrying SSH/grep/output results in your context → delegate the analysis
|
||||||
|
- ❌ Doing work yourself and then "pretending" to delegate → the user can tell
|
||||||
|
- ❌ Skipping verification → raw worker output never reaches the user
|
||||||
|
- ❌ Delegating single tool calls → keep quick reads/writes at manager level
|
||||||
|
- ❌ Firing more than 3 workers in parallel → hard limit
|
||||||
|
- ❌ **Workers fetching their own data sources** → a writer worker that queries the Proxmox API when told to "format the raw file" is fabricating data. Use the input given, not external sources
|
||||||
|
|
||||||
|
## Emergency Exception
|
||||||
|
|
||||||
|
**In an emergency (server down, service must be restored immediately):**
|
||||||
|
- Delegate the diagnosis (find the problem)
|
||||||
|
- Execute the fix yourself (minimize handoff latency)
|
||||||
|
- Verify the fix after delivery
|
||||||
|
- Log the exception in the kanban board
|
||||||
|
|
||||||
|
The emergency exception exists because the user needs the service back NOW,
|
||||||
|
not after three worker round-trips. But it's an exception — not the rule.
|
||||||
|
|
||||||
|
## What This Contract Doesn't Cover
|
||||||
|
|
||||||
|
1. **Worker profile configuration** — covered by `hermes-config-template.prose.md`
|
||||||
|
2. **SSH key management** — covered by existing SSH/Proxmox contracts
|
||||||
|
3. **Git workflow** — covered by `AGENTS.md` in the prose-contracts repo
|
||||||
|
4. **Cron job management** — covered by individual cron contracts
|
||||||
|
5. **Infra verification** — covered by `verify-before-mutate` protocol
|
||||||
|
|
||||||
|
## Verification
|
||||||
|
|
||||||
|
Run `scripts/worker-audit.py` to verify all 6 profiles are aligned:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python3 /root/.hermes/skills/kanban-orchestrator/scripts/worker-audit.py
|
||||||
|
```
|
||||||
|
|
||||||
|
## References
|
||||||
|
|
||||||
|
- `kanban-orchestrator` skill: The operational playbook (detailed execution steps)
|
||||||
|
- `worker-profile-audit.md` (skill reference): Worker configuration audit notes
|
||||||
|
- `delegation-timeout-patterns.md` (skill reference): Timeout handling patterns
|
||||||
|
- `verify-before-mutate` protocol: Infrastructure change verification
|
||||||
|
- `hermes-config-template.prose.md`: Worker profile configuration
|
||||||
|
|
||||||
|
## Success Criteria
|
||||||
|
|
||||||
|
This contract succeeds when:
|
||||||
|
|
||||||
|
1. **No context overflow** — single-turn tasks don't exhaust the iteration budget
|
||||||
|
2. **Workers do the work** — manager coordinates, doesn't execute
|
||||||
|
3. **Verification before delivery** — all output passes through `syslog-review`
|
||||||
|
4. **Kanban board is current** — every task has a lane, every lane has a status
|
||||||
|
5. **User gets verified results** — raw worker output never reaches Kwame
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
**Last updated:** 2026-07-09
|
||||||
|
**Author:** Mumuni (with Kwame's input on triggers and exception criteria)
|
||||||
|
**Status:** Draft — awaiting PR review and merge to prose-contracts main
|
||||||
@@ -12,6 +12,7 @@ description: >
|
|||||||
|
|
||||||
- abiba-telegram: { status: "online", uptime: string, restarts: number }
|
- abiba-telegram: { status: "online", uptime: string, restarts: number }
|
||||||
- gpu-watchdog: { status: "online", uptime: string, restarts: number }
|
- gpu-watchdog: { status: "online", uptime: string, restarts: number }
|
||||||
|
- gpu-monitor: { status: "online", uptime: string, restarts: number }
|
||||||
- gitea-runner: { status: "online", uptime: string, restarts: number }
|
- gitea-runner: { status: "online", uptime: string, restarts: number }
|
||||||
- last_check: timestamp
|
- last_check: timestamp
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,210 @@
|
|||||||
|
---
|
||||||
|
kind: pattern
|
||||||
|
name: ra-h-os-custodianship-contract
|
||||||
|
description: >
|
||||||
|
Operational standards for maintaining RA-H OS as a high-functioning shared
|
||||||
|
memory system across all agents. Defines rules for embedding consistency,
|
||||||
|
namespace discipline, staleness management, orphan prevention, agent
|
||||||
|
custodianship, and recovery protocols. Prevents knowledge graph degradation
|
||||||
|
and ensures reliable semantic search capabilities.
|
||||||
|
---
|
||||||
|
|
||||||
|
# RA-H OS Custodianship Contract — Shared Memory Protocol
|
||||||
|
|
||||||
|
## Purpose
|
||||||
|
This contract establishes the operational standards for maintaining RA-H OS as a high-functioning shared memory system across all agents. It defines the rules, responsibilities, and recovery protocols that prevent knowledge graph degradation and ensure consistent, reliable semantic search capabilities.
|
||||||
|
|
||||||
|
## Core Principles
|
||||||
|
|
||||||
|
### 1. Embedding Consistency
|
||||||
|
**Rule:** All nodes MUST use the RA-H OS built-in embedding pipeline via `http://192.168.68.65:8080/v1/embeddings` for consistent vector representation.
|
||||||
|
|
||||||
|
**Requirements:**
|
||||||
|
- The embedding service at `192.168.68.65:8080` is the single source of truth for vector generation
|
||||||
|
- No external embedding APIs (OpenAI, Cohere, etc.) may be used for RA-H OS nodes
|
||||||
|
- All nodes must have `embedding_status: "chunked"` and `chunk_status: "chunked"` upon creation
|
||||||
|
- If embedding fails, the node must be immediately flagged as `state: "unsearchable"` and logged
|
||||||
|
|
||||||
|
### 2. Namespace Discipline
|
||||||
|
**Rule:** All nodes MUST be categorized into one of three namespaces with strict metadata requirements.
|
||||||
|
|
||||||
|
**Namespace Structure:**
|
||||||
|
- `agent-private`: Isolated working notes for individual agents (Mumuni, Tanko, Okyeame, etc.)
|
||||||
|
- `shared`: Policy files, registry nodes, collective knowledge
|
||||||
|
- `syslogsolution`: Business nodes, client data, operational context
|
||||||
|
|
||||||
|
**Metadata Requirements:**
|
||||||
|
Every node MUST have these metadata fields:
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"agent_id": "<agent_name>",
|
||||||
|
"namespace": "<namespace>",
|
||||||
|
"state": "<state>",
|
||||||
|
"type": "<type>",
|
||||||
|
"owner": "<owner>",
|
||||||
|
"tenant": "<tenant>",
|
||||||
|
"visibility": "<visibility>",
|
||||||
|
"source": "<source>",
|
||||||
|
"captured_by": "<captured_by>"
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### 3. Staleness Management
|
||||||
|
**Rule:** Nodes are categorized by type and have specific staleness thresholds.
|
||||||
|
|
||||||
|
**Staleness Thresholds:**
|
||||||
|
- `policy` type: 30 days (e.g., Shared Memory Policy, Agent Registry)
|
||||||
|
- `registry` type: 30 days (e.g., Node State Management, Health Dashboard)
|
||||||
|
- `template` type: 30 days (e.g., Agent SOUL.md)
|
||||||
|
- `note` type: 60 days (e.g., business context, project notes)
|
||||||
|
- `information` type: 90 days (e.g., reference documentation)
|
||||||
|
- `idea` type: 30 days (e.g., brainstorming, experimental notes)
|
||||||
|
|
||||||
|
**Transition Protocol:**
|
||||||
|
- `active` → `stale`: After days since `updated_at` exceeds threshold
|
||||||
|
- `stale` → `review_pending`: Daily cron check flags for human review
|
||||||
|
- `review_pending` → `active`: Human reviews and updates content
|
||||||
|
- `review_pending` → `archived`: After 15 days in review_pending state (automatic purge)
|
||||||
|
|
||||||
|
### 4. Orphan Prevention
|
||||||
|
**Rule:** No node should remain orphaned (0 edges) for more than 7 days.
|
||||||
|
|
||||||
|
**Orphan Detection:**
|
||||||
|
- Daily cron job identifies nodes with 0 incoming/outgoing edges
|
||||||
|
- Orphans are flagged with `state: "review_pending"` and `metadata.orphan_detected: true`
|
||||||
|
- 7-day grace period: Orphans must be either:
|
||||||
|
- Connected to relevant nodes via edges
|
||||||
|
- Merged into existing parent nodes
|
||||||
|
- Archived after 15 days in review_pending state
|
||||||
|
|
||||||
|
### 5. Agent Custodianship
|
||||||
|
**Rule:** Each agent is responsible for maintaining their own created nodes.
|
||||||
|
|
||||||
|
**Responsibilities:**
|
||||||
|
- **Mumuni**: Primary custodian for `syslogsolution` namespace, health dashboard, and all business nodes
|
||||||
|
- **Tanko**: Custodian for `agent-private` namespace, personal notes, and fitness/creative work
|
||||||
|
- **Okyeame**: Custodian for `shared` namespace, policy files, and collective knowledge nodes
|
||||||
|
- **All Agents**: Must verify embedding status before reporting node creation as complete
|
||||||
|
|
||||||
|
### 6. Embedding Failure Recovery
|
||||||
|
**Rule:** When embedding pipeline is unavailable, nodes are blocked from creation or immediately flagged.
|
||||||
|
|
||||||
|
**Recovery Protocol:**
|
||||||
|
1. **Detection**: Daily health check monitors `http://192.168.68.65:8080/health` endpoint
|
||||||
|
2. **Alert**: If embedding service is down, alert is sent to all agents via relay
|
||||||
|
3. **Mitigation**:
|
||||||
|
- New node creation is suspended
|
||||||
|
- Existing nodes with `chunk_status: "not_chunked"` are flagged as `unsearchable`
|
||||||
|
4. **Recovery**: When service returns:
|
||||||
|
- All `not_chunked` nodes are queued for re-embedding
|
||||||
|
- Re-embedding status logged in `embedding_retry_log` table
|
||||||
|
- 5-minute retry window with exponential backoff (1s, 2s, 4s, 8s, 16s)
|
||||||
|
|
||||||
|
### 7. Chunking Verification
|
||||||
|
**Rule:** All node modifications must verify chunking status within 5 minutes.
|
||||||
|
|
||||||
|
**Verification Protocol:**
|
||||||
|
- After `updateNode` or `createNode`, query `chunks` table for `node_id`
|
||||||
|
- If `chunk_status != "chunked"` after 5 minutes, log failure and flag node
|
||||||
|
- Automated cleanup script runs every 6 hours to retry failed chunks
|
||||||
|
- Maximum 3 retry attempts per node before escalation to human
|
||||||
|
|
||||||
|
### 8. Graph Health Monitoring
|
||||||
|
**Rule:** Health dashboard (Node #74) must reflect real-time graph status.
|
||||||
|
|
||||||
|
**Dashboard Requirements:**
|
||||||
|
- Total nodes, edges, orphans, stale nodes, chunk errors
|
||||||
|
- Embedding service health status
|
||||||
|
- Last successful embedding attempt
|
||||||
|
- Alert thresholds:
|
||||||
|
- Orphan rate > 40%: Warning
|
||||||
|
- Stale node rate > 50%: Critical
|
||||||
|
- Chunk error rate > 30%: Critical
|
||||||
|
- Embedding service down: Critical (immediate alert)
|
||||||
|
|
||||||
|
## Operational Procedures
|
||||||
|
|
||||||
|
### Daily Maintenance (Automated)
|
||||||
|
1. **Staleness Check**: Identify nodes past their threshold
|
||||||
|
2. **Orphan Detection**: Flag nodes with 0 edges
|
||||||
|
3. **Chunk Verification**: Retry failed chunks, log errors
|
||||||
|
4. **Health Update**: Refresh Node #74 dashboard
|
||||||
|
|
||||||
|
### Weekly Maintenance (Human Review)
|
||||||
|
1. **Review Pending Nodes**: Examine flagged nodes
|
||||||
|
2. **Edge Optimization**: Connect related orphaned nodes
|
||||||
|
3. **Archive Cleanup**: Remove nodes past 15-day review period
|
||||||
|
4. **Embedding Audit**: Verify all nodes have valid embeddings
|
||||||
|
|
||||||
|
### Emergency Recovery
|
||||||
|
1. **Embedding Service Down**:
|
||||||
|
- Suspend node creation
|
||||||
|
- Alert all agents
|
||||||
|
- Investigate root cause (check `192.168.68.65:8080`)
|
||||||
|
2. **Database Corruption**:
|
||||||
|
- Restore from latest PBS backup
|
||||||
|
- Verify chunk integrity
|
||||||
|
- Re-run embedding for affected nodes
|
||||||
|
3. **Mass Orphan Creation**:
|
||||||
|
- Identify source agent/namespace
|
||||||
|
- Review recent changes
|
||||||
|
- Reconnect or archive affected nodes
|
||||||
|
|
||||||
|
## Compliance & Enforcement
|
||||||
|
|
||||||
|
### Audit Schedule
|
||||||
|
- **Daily**: Automated health checks
|
||||||
|
- **Weekly**: Human review of review_pending nodes
|
||||||
|
- **Monthly**: Comprehensive graph audit (full schema validation)
|
||||||
|
- **Quarterly**: Contract review and threshold adjustment
|
||||||
|
|
||||||
|
### Violation Consequences
|
||||||
|
- **First**: Alert to responsible agent
|
||||||
|
- **Second**: Node placed in `review_pending` state
|
||||||
|
- **Third**: Agent access suspended until remediation
|
||||||
|
- **Fourth**: Escalation to human (Kwame)
|
||||||
|
|
||||||
|
### Metrics for Success
|
||||||
|
- Orphan rate: < 10%
|
||||||
|
- Stale node rate: < 20%
|
||||||
|
- Chunk error rate: < 5%
|
||||||
|
- Embedding service uptime: > 99.9%
|
||||||
|
- Node creation verification: 100%
|
||||||
|
|
||||||
|
## Implementation Notes
|
||||||
|
|
||||||
|
### Database Schema Requirements
|
||||||
|
- `nodes` table: Standard fields plus `embedding_status`, `chunk_status`, `last_embedding_attempt`
|
||||||
|
- `chunks` table: Must track `chunk_status` and `embedding_status`
|
||||||
|
- `embedding_retry_log` table: Track retry attempts for failed chunks
|
||||||
|
- `node_state_history` table: Log state transitions for audit trail
|
||||||
|
|
||||||
|
### Cron Job Requirements
|
||||||
|
- `ra-h-health-check`: Daily (4h interval)
|
||||||
|
- `ra-h-staleness-detection`: Daily
|
||||||
|
- `ra-h-orphan-detection`: Daily
|
||||||
|
- `ra-h-chunk-retry`: Every 6 hours
|
||||||
|
- `ra-h-review-cleanup`: Weekly (15-day archive)
|
||||||
|
|
||||||
|
### Agent Onboarding
|
||||||
|
All new agents must:
|
||||||
|
1. Read this contract
|
||||||
|
2. Understand namespace responsibilities
|
||||||
|
3. Know how to verify embedding status
|
||||||
|
4. Know the alert escalation path
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Approval
|
||||||
|
This contract is effective immediately upon approval by the primary custodian (Mumuni) and system owner (Kwame).
|
||||||
|
|
||||||
|
**Effective Date:** July 10, 2026
|
||||||
|
**Review Date:** October 10, 2026
|
||||||
|
**Primary Custodian:** Mumuni 🦅
|
||||||
|
**System Owner:** Jerome Tabiri
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
*Contract Version: 1.0*
|
||||||
|
*Last Updated: July 10, 2026*
|
||||||
|
*Storage: `/root/.hermes/skills/ra-h-os-custodianship-contract.prose.md`*
|
||||||
Reference in New Issue
Block a user