Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
0c298eb9d9 | ||
|
|
af41f8f57a | ||
|
|
be02b0e843 | ||
|
|
79d4a73895 | ||
|
|
19b6db9891 |
@@ -1,256 +0,0 @@
|
|||||||
---
|
|
||||||
kind: pattern
|
|
||||||
name: delegation-prose-contract
|
|
||||||
description: >
|
|
||||||
Manager (Mumuni) operating doctrine for task decomposition, worker
|
|
||||||
delegation, verification, and delivery. Defines when to delegate, which
|
|
||||||
worker to use for what, how to handle failures, and the kanban board
|
|
||||||
protocol. Enforces context-window discipline and separation of concerns.
|
|
||||||
Runs on Mumuni (CT 118, storepve, .6) via Hermes agent.
|
|
||||||
version: 1.0.0
|
|
||||||
---
|
|
||||||
|
|
||||||
## Maintains
|
|
||||||
|
|
||||||
- Worker roster: 6 profiles (`syslog-code`, `syslog-devops`, `syslog-email`,
|
|
||||||
`syslog-research`, `syslog-review`, `syslog-writer`)
|
|
||||||
- Kanban board state at `~/.hermes/kanban/kanban.json`
|
|
||||||
- Context window budget: ~65K tokens per request (131K total, 60% threshold)
|
|
||||||
|
|
||||||
## Topology
|
|
||||||
|
|
||||||
**Cluster:** 5 Proxmox nodes (ocupve, acerpve, minipve, amdpve, storepve)
|
|
||||||
**Manager:** Mumuni (CT 118, storepve, .6) via Hermes agent
|
|
||||||
**Workers:** 6 profiles, all running on the same agent — no separate hosts needed
|
|
||||||
|
|
||||||
This contract is infrastructure-agnostic in terms of which nodes are used.
|
|
||||||
Workers execute tasks on whatever infrastructure they're given — SSH to .6,
|
|
||||||
.pm, .9, .12, or .15 depending on the task. The contract defines the
|
|
||||||
**who** and **when** — not the **where**.
|
|
||||||
|
|
||||||
## Why This Matters
|
|
||||||
|
|
||||||
Without enforced delegation, the manager consumes the full iteration budget
|
|
||||||
(60 calls) on single-turn tasks — SSH to 5 nodes, check each VM, read logs —
|
|
||||||
leaving no capacity for actual coordination. The result: context overflow
|
|
||||||
(59K tokens in system prompt), iteration exhaustion, and degraded response
|
|
||||||
quality. This contract exists because I blew through my budget checking
|
|
||||||
Proxmox node status instead of delegating to `syslog-devops`.
|
|
||||||
|
|
||||||
## Context Window Discipline
|
|
||||||
|
|
||||||
**The system prompt is ~6.5K tokens (stable: ~4.5K tool schemas + ~2K other guidance).**
|
|
||||||
**Volatile (MEMORY.md + USER.md): ~300 tokens.**
|
|
||||||
**Total base: ~6,800 tokens per request.**
|
|
||||||
|
|
||||||
The remaining budget is the conversation. Every tool call result adds to it.
|
|
||||||
If a single call returns >10K tokens (e.g., `grep` on a large file, SSH output
|
|
||||||
from multiple nodes), the context fills fast. That's why we delegate: workers
|
|
||||||
process in isolation and return compact results.
|
|
||||||
|
|
||||||
## Trigger Conditions
|
|
||||||
|
|
||||||
Delegation is **mandatory** when any of these apply:
|
|
||||||
|
|
||||||
| Condition | Threshold | Example |
|
|
||||||
|-----------|-----------|---------|
|
|
||||||
| Multiple tool calls needed | 2+ calls with intermediate logic | Read file → analyze → write report |
|
|
||||||
| Large data retrieval | Output >5K tokens | `grep -r "pattern" /path` on large dirs |
|
|
||||||
| Cross-domain work | Spans 2+ worker specialties | Infra check + email filter |
|
|
||||||
| Infrastructure changes | Any mutating operation | `qm set`, `systemctl restart`, `git push` |
|
|
||||||
| Research/analysis | Needs browser or deep reading | Web research, code review, data analysis |
|
|
||||||
| Code builds or changes | Writing or modifying code | Scripts, configs, patches |
|
|
||||||
| Sequential dependencies | Worker B needs Worker A's output | Code → Review → Deliver |
|
|
||||||
|
|
||||||
**Single tool calls stay at manager level.** Quick `grep`, `ls`, `cat`,
|
|
||||||
`curl`, `hermes tools list` — these are decision-making tools. The manager
|
|
||||||
reads them directly.
|
|
||||||
|
|
||||||
## Worker Selection Matrix
|
|
||||||
|
|
||||||
| Worker | Model | Toolsets | Role | Use When |
|
|
||||||
|--------|-------|----------|------|----------|
|
|
||||||
| `syslog-code` | qwen3.6-27B-code | terminal, file, web, memory, skills | Code patches, automation, scripts | Writing/modifying code, creating scripts, debugging, reading/writing files |
|
|
||||||
| `syslog-devops` | qwen3.6-27B-code | terminal, file, web, memory, skills | Infrastructure, DB, bridge, Proxmox | Server ops, SSH, Docker, Proxmox, DB queries, hardware checks |
|
|
||||||
| `syslog-email` | ornith-1.0-35b | terminal, file, web, memory, skills | Email automation, mail operations | Sending/receiving email, inbox management, SMTP operations |
|
|
||||||
| `syslog-research` | ornith-1.0-35b | terminal, file, web, memory, skills, **browser** | Analysis, classification, data processing | Web research, browser tasks, data analysis, classification, reading docs |
|
|
||||||
| `syslog-review` | ornith-1.0-35b | terminal, file, web, memory, skills | Verification, QA, audit validation | **ALWAYS** verify worker output before delivery — especially for infra changes, code builds, and research findings |
|
|
||||||
| `syslog-writer` | ornith-1.0-35b | terminal, file, web, memory, skills | Docs, content, branding, reports | Writing docs, reports, proposals, content, markdown formatting |
|
|
||||||
|
|
||||||
### Selection Rules
|
|
||||||
|
|
||||||
1. **Match specialty first.** A code task → `syslog-code`. An infra task →
|
|
||||||
`syslog-devops`. Don't put a `syslog-email` worker on a code review.
|
|
||||||
2. **Research tasks with browser needs → `syslog-research`.** Other workers
|
|
||||||
don't have the browser toolset.
|
|
||||||
3. **Verification → `syslog-review`.** Never deliver raw worker output.
|
|
||||||
4. **Documentation/content → `syslog-writer`.** Let them own the prose.
|
|
||||||
5. **If unsure, delegate to `syslog-research`** — it has the broadest toolset
|
|
||||||
(includes browser) and high reasoning effort.
|
|
||||||
|
|
||||||
## Delegation Protocol
|
|
||||||
|
|
||||||
### Step 1: Decompose
|
|
||||||
|
|
||||||
Break the task into lanes. Each lane does ONE thing. Workers are independent —
|
|
||||||
no lane depends on another's output mid-flight. If lanes depend on each other,
|
|
||||||
dispatch sequentially.
|
|
||||||
|
|
||||||
### Step 2: Dispatch
|
|
||||||
|
|
||||||
Fire workers via `delegate_task`:
|
|
||||||
|
|
||||||
**Parallel (independent lanes):**
|
|
||||||
```
|
|
||||||
delegate_task(
|
|
||||||
tasks=[
|
|
||||||
{"goal": "Check all 5 Proxmox nodes for VM status", "context": "SSH to each node via 192.168.68.x, run 'qm list'"},
|
|
||||||
{"goal": "Check Docker container health on .7/.116/.17", "context": "SSH to each host, check container status"},
|
|
||||||
]
|
|
||||||
)
|
|
||||||
```
|
|
||||||
|
|
||||||
**Sequential (dependent lanes):**
|
|
||||||
Dispatch lane 1 → wait for result → dispatch lane 2.
|
|
||||||
|
|
||||||
### Step 3: Verify
|
|
||||||
|
|
||||||
**MANDATORY for:**
|
|
||||||
- Infrastructure changes (any `qm`, `pct`, `systemctl`, `git push`)
|
|
||||||
- Code builds and modifications
|
|
||||||
- Research findings (web data, external sources)
|
|
||||||
- Any output that will reach the user
|
|
||||||
|
|
||||||
**Fire `syslog-review` to verify:**
|
|
||||||
```
|
|
||||||
delegate_task(
|
|
||||||
goal="Review the output of the devops worker. Verify the node status
|
|
||||||
report is accurate, check for inconsistencies, confirm all nodes were
|
|
||||||
reachable.",
|
|
||||||
context="Worker was syslog-devops. Output is at /tmp/node-report.md.
|
|
||||||
Verify against live system."
|
|
||||||
)
|
|
||||||
```
|
|
||||||
|
|
||||||
**If verification fails:**
|
|
||||||
1. Send work back to original worker with review feedback
|
|
||||||
2. Re-verify
|
|
||||||
3. Max 2 re-verify cycles before escalating to Kwame
|
|
||||||
|
|
||||||
### Step 4: Deliver
|
|
||||||
|
|
||||||
Only verified results reach Kwame. Format per channel:
|
|
||||||
- Telegram: Use `telegram-formatting` skill
|
|
||||||
- Zulip: Use Zulip Markdown (CommonMark)
|
|
||||||
- Email: Use `syslog-email` skill
|
|
||||||
|
|
||||||
## Kanban Board Protocol
|
|
||||||
|
|
||||||
**File:** `~/.hermes/kanban/kanban.json`
|
|
||||||
|
|
||||||
```json
|
|
||||||
{
|
|
||||||
"task_id": "unique-id",
|
|
||||||
"title": "Task description",
|
|
||||||
"created": "2026-07-09T01:00:00",
|
|
||||||
"status": "backlog|in_progress|review|done",
|
|
||||||
"lanes": [
|
|
||||||
{
|
|
||||||
"lane_id": "devops-check",
|
|
||||||
"worker": "syslog-devops",
|
|
||||||
"goal": "Check all 5 Proxmox nodes",
|
|
||||||
"status": "dispatched|completed|failed",
|
|
||||||
"output_file": "/tmp/node-report.md"
|
|
||||||
}
|
|
||||||
]
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
**Update the board on every state change.**
|
|
||||||
|
|
||||||
## Failure Handling
|
|
||||||
|
|
||||||
### Worker Timeouts
|
|
||||||
|
|
||||||
- Child timeout: **900 seconds** (15 minutes)
|
|
||||||
- Worker model `syslog-auto` is slow — it can hit the timeout limit with
|
|
||||||
22+ API calls
|
|
||||||
- **If a worker times out:** Re-dispatch with a narrower scope. Break the
|
|
||||||
task into smaller pieces that fit in the timeout window.
|
|
||||||
- **Avoid delegating sequential SSH hops** — each SSH connection adds latency
|
|
||||||
that compounds quickly. Prefer API-based or local approaches when possible.
|
|
||||||
|
|
||||||
### Worker Selection Failures
|
|
||||||
|
|
||||||
- `syslog-devops` is best for infrastructure tasks (SSH, Proxmox, Docker)
|
|
||||||
- `syslog-code` is best for code-level work (reading files, writing scripts)
|
|
||||||
- `syslog-research` has the browser toolset — use for web research
|
|
||||||
- `syslog-review` is the QA gate — always fire before delivery
|
|
||||||
- **Never fire more than 3 parallel workers** (max_concurrent_children: 3)
|
|
||||||
- **Never nest delegation** (max_spawn_depth: 1)
|
|
||||||
|
|
||||||
### Context Overflow
|
|
||||||
|
|
||||||
- If a task requires >10K tokens of output, delegate the processing
|
|
||||||
- Workers return compact summaries, not raw data dumps
|
|
||||||
- Pass file paths and concrete goals — never dump raw data into context
|
|
||||||
|
|
||||||
## Anti-patterns
|
|
||||||
|
|
||||||
- ❌ Reading large files into your own context before deciding → delegate the read
|
|
||||||
- ❌ Carrying SSH/grep/output results in your context → delegate the analysis
|
|
||||||
- ❌ Doing work yourself and then "pretending" to delegate → the user can tell
|
|
||||||
- ❌ Skipping verification → raw worker output never reaches the user
|
|
||||||
- ❌ Delegating single tool calls → keep quick reads/writes at manager level
|
|
||||||
- ❌ Firing more than 3 workers in parallel → hard limit
|
|
||||||
|
|
||||||
## Emergency Exception
|
|
||||||
|
|
||||||
**In an emergency (server down, service must be restored immediately):**
|
|
||||||
- Delegate the diagnosis (find the problem)
|
|
||||||
- Execute the fix yourself (minimize handoff latency)
|
|
||||||
- Verify the fix after delivery
|
|
||||||
- Log the exception in the kanban board
|
|
||||||
|
|
||||||
The emergency exception exists because the user needs the service back NOW,
|
|
||||||
not after three worker round-trips. But it's an exception — not the rule.
|
|
||||||
|
|
||||||
## What This Contract Doesn't Cover
|
|
||||||
|
|
||||||
1. **Worker profile configuration** — covered by `hermes-config-template.prose.md`
|
|
||||||
2. **SSH key management** — covered by existing SSH/Proxmox contracts
|
|
||||||
3. **Git workflow** — covered by `AGENTS.md` in the prose-contracts repo
|
|
||||||
4. **Cron job management** — covered by individual cron contracts
|
|
||||||
5. **Infra verification** — covered by `verify-before-mutate` protocol
|
|
||||||
|
|
||||||
## Verification
|
|
||||||
|
|
||||||
Run `scripts/worker-audit.py` to verify all 6 profiles are aligned:
|
|
||||||
|
|
||||||
```bash
|
|
||||||
python3 /root/.hermes/skills/kanban-orchestrator/scripts/worker-audit.py
|
|
||||||
```
|
|
||||||
|
|
||||||
## References
|
|
||||||
|
|
||||||
- `kanban-orchestrator` skill: The operational playbook (detailed execution steps)
|
|
||||||
- `worker-profile-audit.md` (skill reference): Worker configuration audit notes
|
|
||||||
- `delegation-timeout-patterns.md` (skill reference): Timeout handling patterns
|
|
||||||
- `verify-before-mutate` protocol: Infrastructure change verification
|
|
||||||
- `hermes-config-template.prose.md`: Worker profile configuration
|
|
||||||
|
|
||||||
## Success Criteria
|
|
||||||
|
|
||||||
This contract succeeds when:
|
|
||||||
|
|
||||||
1. **No context overflow** — single-turn tasks don't exhaust the iteration budget
|
|
||||||
2. **Workers do the work** — manager coordinates, doesn't execute
|
|
||||||
3. **Verification before delivery** — all output passes through `syslog-review`
|
|
||||||
4. **Kanban board is current** — every task has a lane, every lane has a status
|
|
||||||
5. **User gets verified results** — raw worker output never reaches Kwame
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
**Last updated:** 2026-07-09
|
|
||||||
**Author:** Mumuni (with Kwame's input on triggers and exception criteria)
|
|
||||||
**Status:** Draft — awaiting PR review and merge to prose-contracts main
|
|
||||||
+76
-37
@@ -5,9 +5,13 @@ description: >
|
|||||||
Manages the GPU inference fleet across all hosts. Handles model deployment,
|
Manages the GPU inference fleet across all hosts. Handles model deployment,
|
||||||
registration, health checks, LiteLLM sync, agent key management, GPU
|
registration, health checks, LiteLLM sync, agent key management, GPU
|
||||||
saturation watchdog, Prometheus/Grafana monitoring, and self-healing.
|
saturation watchdog, Prometheus/Grafana monitoring, and self-healing.
|
||||||
Current as of 2026-07-08: context reduced to 128K on NVIDIA GPUs, parallel 2
|
UPDATED 2026-07-12: Architecture is DIRECT GPU — LiteLLM routes directly
|
||||||
on all GPUs, LiteLLM timeouts tuned (gemma 25→120s, qwen 40→90s), router fully
|
to llama-server on each GPU host (no router in inference path). Router
|
||||||
deprecated — nginx routes /v1 → LiteLLM directly.
|
(port 9000) is running but NOT in request path. All GPUs standardized on
|
||||||
|
api-key 'not-needed'. RTX 5070 had api-key mismatch (sk-loc...5678) that
|
||||||
|
caused cascading 401→timeout→401 fallback loops — fixed.
|
||||||
|
Context: RTX 3090 verified at 256K (was documented as 128K — WRONG).
|
||||||
|
Workload: Compression moved to Strix Halo, RTX 5070 → vision/web only.
|
||||||
agent: abiba
|
agent: abiba
|
||||||
triggers:
|
triggers:
|
||||||
- on model add/remove
|
- on model add/remove
|
||||||
@@ -30,7 +34,7 @@ triggers:
|
|||||||
- prometheus: { status: "running", targets: 5 } — Scrapes GPU :9400 exporters + LiteLLM
|
- prometheus: { status: "running", targets: 5 } — Scrapes GPU :9400 exporters + LiteLLM
|
||||||
- port_conflict_detection: { status: "active" } — All 3 GPU wrappers detect ghost processes before binding
|
- port_conflict_detection: { status: "active" } — All 3 GPU wrappers detect ghost processes before binding
|
||||||
|
|
||||||
## Fleet Topology (Current — June 2026)
|
## Fleet Topology (Current — July 2026)
|
||||||
|
|
||||||
```
|
```
|
||||||
┌──────────────────────────────────────────────────────────────────┐
|
┌──────────────────────────────────────────────────────────────────┐
|
||||||
@@ -47,9 +51,9 @@ triggers:
|
|||||||
│ Containers: │
|
│ Containers: │
|
||||||
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
|
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
|
||||||
│ │ LiteLLM │ │ Router │ │Dashboard │ │ Grafana │ │
|
│ │ LiteLLM │ │ Router │ │Dashboard │ │ Grafana │ │
|
||||||
│ │ :4000 │─▶│ :9000 │ │ :3000 │ │ :3000 │ │
|
│ │ :4000 │ │ :9000 │ │ :3000 │ │ :3000 │ │
|
||||||
│ │ keys+sync│ │internal │ │ harness │ │ Prometheus│ │
|
│ │ keys+sync│ │deprecated│ │ harness │ │ Prometheus│ │
|
||||||
│ │ fallback │ │only! │ │ UI │ │ data src │ │
|
│ │ fallback │ │not in │ │ UI │ │ data src │ │
|
||||||
│ └──────────┘ └───┬──────┘ └──────────┘ └──────────┘ │
|
│ └──────────┘ └───┬──────┘ └──────────┘ └──────────┘ │
|
||||||
│ │ │
|
│ │ │
|
||||||
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
|
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
|
||||||
@@ -64,7 +68,7 @@ triggers:
|
|||||||
│ CT 8 │ │ CT 110 │ │ CT 15 │ │ pi (.24) │
|
│ CT 8 │ │ CT 110 │ │ CT 15 │ │ pi (.24) │
|
||||||
│ RTX 3090 │ │ RTX 5070 │ │ Strix Halo│ │ GPU Monitor │
|
│ RTX 3090 │ │ RTX 5070 │ │ Strix Halo│ │ GPU Monitor │
|
||||||
│ 24GB │ │ 12GB │ │ 64GB UMA │ │ :9100 │
|
│ 24GB │ │ 12GB │ │ 64GB UMA │ │ :9100 │
|
||||||
│ 128K ctx │ │ 128K ctx │ │ 256K ctx │ │ Watchdog │
|
│ 256K ctx │ │ 131K ctx │ │ 256K ctx │ │ Watchdog │
|
||||||
│ qwen3.6 │ │ gemma-4-12b │ │ ornith35B │ │ Prometheus │
|
│ qwen3.6 │ │ gemma-4-12b │ │ ornith35B │ │ Prometheus │
|
||||||
│ 27B-code │ │ :8080 │ │ :8080 │ │ exporter │
|
│ 27B-code │ │ :8080 │ │ :8080 │ │ exporter │
|
||||||
│ :8080 │ │ :9400 (exp) │ │ :9400(exp)│ │ :9401 │
|
│ :8080 │ │ :9400 (exp) │ │ :9400(exp)│ │ :9401 │
|
||||||
@@ -72,13 +76,42 @@ triggers:
|
|||||||
└──────────┘
|
└──────────┘
|
||||||
```
|
```
|
||||||
|
|
||||||
## Current Model Assignments (2026-07-08)
|
## Current Model Assignments (2026-07-12)
|
||||||
|
|
||||||
| Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status |
|
| Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status |
|
||||||
|-------|-----|------|------|-----|----------|----------|-------------|--------|
|
|-------|-----|------|------|-----|----------|----------|-------------|--------|
|
||||||
| qwen3.6-27B-code | RTX 3090 | .8 (llm-gpu) | 20.3/24GB (83%) | 128K | turbo4 | 2 | 512/512 | ✅ healthy |
|
| qwen3.6-27B-code | RTX 3090 | .8 (llm-gpu) | 20.7/24GB (84%) | **256K** | turbo4 | 1 | default | ✅ healthy |
|
||||||
| gemma-4-12b | RTX 5070 | .110 (ocu-llm) | 9.4/12.2GB (77%) | 128K | q4_0 | 2 | 2048/512 | ✅ healthy |
|
| gemma-4-12b | RTX 5070 | .110 (ocu-llm) | 10.0/12.2GB (82%) | 131K | q4_0 | 2 | 2048/1024 | ✅ healthy |
|
||||||
| ornith-1.0-35b | Strix Halo Vulkan | .15 (amdpve) | 24.4/64GB (35%) | 256K | q8_0 | 2 | 2048/512 | ✅ healthy |
|
| ornith-1.0-35b | Strix Halo Vulkan | .15 (amdpve) | ~9GB/64GB | 256K | q8_0 | 2 | 2048/512 | ✅ healthy |
|
||||||
|
|
||||||
|
## Routing Configuration (LiteLLM — July 2026)
|
||||||
|
|
||||||
|
### syslog-auto Weighted Pool
|
||||||
|
|
||||||
|
| Model | GPU | Weight | RPM Cap | Purpose |
|
||||||
|
|-------|-----|--------|---------|---------|
|
||||||
|
| qwen3.6-27B-code | RTX 3090 | 0.55 | 500 | Heavy reasoning, code, long context |
|
||||||
|
| ornith-1.0-35b | Strix Halo | 0.30 | **60** | Agentic workflows, tool calling |
|
||||||
|
| gemma-4-12b | RTX 5070 | 0.15 | 200 | Overflow + vision |
|
||||||
|
|
||||||
|
### Direct Model Endpoints
|
||||||
|
|
||||||
|
| Model | RPM Cap | Notes |
|
||||||
|
|-------|---------|-------|
|
||||||
|
| ornith-1.0-35b | 40 | Tight cap — prevents Strix overload |
|
||||||
|
| qwen3.6-27B-code | 500 | High cap — primary workhorse |
|
||||||
|
| gemma-4-12b | 500 | High cap — fast 12B |
|
||||||
|
|
||||||
|
### Fallback Chains
|
||||||
|
- gemma → qwen
|
||||||
|
- qwen → gemma
|
||||||
|
- ornith → qwen → gemma
|
||||||
|
- syslog-auto → qwen → gemma → ornith
|
||||||
|
|
||||||
|
### Why ornith RPM Is Capped
|
||||||
|
- Direct: 40 RPM (tight) — Strix Halo is shared with compression tasks
|
||||||
|
- Via syslog-auto: 60 RPM (moderate) — prevents flooding when multiple agents use syslog-auto simultaneously
|
||||||
|
- Combined max: ~100 RPM across both paths — Strix Halo can sustain this at 80°C
|
||||||
|
|
||||||
## Operations
|
## Operations
|
||||||
|
|
||||||
@@ -114,10 +147,10 @@ triggers:
|
|||||||
|
|
||||||
### sync-keys
|
### sync-keys
|
||||||
1. List all agent keys in LiteLLM DB via `GET /key/list`
|
1. List all agent keys in LiteLLM DB via `GET /key/list`
|
||||||
2. Compare against expected agent list: [tanko, mumuni, abiba, tdunna, baggy, kagenz0]
|
2. Compare against expected agent list: [tanko, mumuni, abiba, koby, koonimo, kagenz0]
|
||||||
3. Generate missing keys via `POST /key/generate` with unlimited budget
|
3. Generate missing keys via `POST /key/generate` with unlimited budget
|
||||||
4. Update agent configs — `/etc/environment` LITELLM_API_KEY
|
4. Update Infisical vault: `infisical secrets set LITELLM_API_KEY=<key> --project=agents --env=production`
|
||||||
5. Send Zulip DM to agents that can't be reached via SSH
|
5. Send Zulip DM to agents that can't be reached via SSH (provide vault login instructions)
|
||||||
6. Verify each key with test request through full chain
|
6. Verify each key with test request through full chain
|
||||||
7. Document keys in knowledge graph
|
7. Document keys in knowledge graph
|
||||||
|
|
||||||
@@ -136,19 +169,25 @@ Show full fleet status: GPUs, models, VRAM, context windows, parallel slots, act
|
|||||||
7. Check port conflicts: verify only one llama-server on :8080 per host
|
7. Check port conflicts: verify only one llama-server on :8080 per host
|
||||||
8. Verify agent keys: 9 keys in LiteLLM DB (`GET /key/list`)
|
8. Verify agent keys: 9 keys in LiteLLM DB (`GET /key/list`)
|
||||||
|
|
||||||
## Agent Keys (LiteLLM DB — Current 2026-06-30)
|
## Agent Keys (LiteLLM DB — Current 2026-07-11)
|
||||||
|
|
||||||
| Agent | CT | IP | Key | Access |
|
Keys stored in Infisical vault (project=agents, env=production, secret=LITELLM_API_KEY).
|
||||||
|-------|-----|-----|-----|--------|
|
Agent gateways inject keys at runtime via `infisical run --` wrapper.
|
||||||
| Tanko | 112 | .122 | `sk-CggiHWlamQyShxWC3Hx6uw` | SSH jerome |
|
Plaintext keys removed from this contract post-vault-migration.
|
||||||
| Mumuni | 114 | .123 | `sk-VrqCNlwUgzoNGOpikJ7nwQ` | SSH root |
|
|
||||||
| Abiba | 100 | .24 | `sk-Qvzi4uYQBhlSK_XstEhcyQ` | local (pi agent) |
|
|
||||||
| Tdunna | 111 | ? | `sk-6sbCNjz2T6lTVDBdlNHXsA` | Zulip DM |
|
|
||||||
| Baggy | 113 | ? | `sk-krnw_zGBwvvL5b7l2t-s-A` | no SSH |
|
|
||||||
| Kagenz0 | 105 | ? | `sk-Dh4CDkaHebMLEp8qqq20qA` | no SSH |
|
|
||||||
|
|
||||||
**Key update procedure**: Update `/etc/environment` → `LITELLM_API_KEY=sk-...` → restart Hermes.
|
| Agent | CT | IP | LiteLLM Alias | Key Source | Access |
|
||||||
If no SSH access, send Zulip DM via abiba-bot.
|
|-------|-----|-----|---------------|------------|--------|
|
||||||
|
| Tanko | 112 | .122 | `tanko` | Infisical vault | SSH jerome |
|
||||||
|
| Mumuni | 114 | .123 | `mumuni` | Infisical vault | SSH root |
|
||||||
|
| Abiba | 100 | .24 | `abiba-pi` | Infisical vault | local (pi agent) |
|
||||||
|
| Koby | 111 | ? | `koby` | Infisical vault | Zulip DM |
|
||||||
|
| Koonimo | 113 | ? | `koonimo` | Infisical vault (migrated 2026-07-11) | no SSH |
|
||||||
|
| Kagenz0 | 105 | ? | `kagenz0` | Infisical vault | no SSH |
|
||||||
|
|
||||||
|
> **Note**: CT hostnames differ from agent identities. CT111=tdunna runs koby; CT113=baggy runs koonimo.
|
||||||
|
|
||||||
|
**Key update procedure**: Update Infisical vault → `infisical secrets set LITELLM_API_KEY=sk-... --project=agents --env=production` → restart agent gateway. Agent picks up new key via `infisical run --` wrapper at startup.
|
||||||
|
If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
|
||||||
|
|
||||||
## Configuration Files
|
## Configuration Files
|
||||||
|
|
||||||
@@ -170,7 +209,7 @@ If no SSH access, send Zulip DM via abiba-bot.
|
|||||||
|
|
||||||
| Component | URL | Details |
|
| Component | URL | Details |
|
||||||
|-----------|-----|---------|
|
|-----------|-----|---------|
|
||||||
| Grafana | `http://192.168.68.116:3001/` | admin / syslog-grafana-2026 |
|
| Grafana | `http://192.168.68.116:3001/` | admin / vault (`GRAFANA_ADMIN_PASSWORD`) |
|
||||||
| GPU Dashboard | `http://192.168.68.116:3001/d/gpu-fleet` | Gauges + time series |
|
| GPU Dashboard | `http://192.168.68.116:3001/d/gpu-fleet` | Gauges + time series |
|
||||||
| Prometheus | `http://192.168.68.116:9090/` (internal) | 5 scrape targets |
|
| Prometheus | `http://192.168.68.116:9090/` (internal) | 5 scrape targets |
|
||||||
| GPU Exporters | `:9400/metrics` on .8, .110, .15 | NVIDIA/AMD GPU metrics |
|
| GPU Exporters | `:9400/metrics` on .8, .110, .15 | NVIDIA/AMD GPU metrics |
|
||||||
@@ -181,10 +220,10 @@ If no SSH access, send Zulip DM via abiba-bot.
|
|||||||
- **Router startup race**: Compose router.py doesn't call load_roster(). Reload thread sleeps 30s first.
|
- **Router startup race**: Compose router.py doesn't call load_roster(). Reload thread sleeps 30s first.
|
||||||
Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster().
|
Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster().
|
||||||
- **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround.
|
- **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround.
|
||||||
- **VRAM (2026-07-08)**: RTX 3090 at 20.3/24GB (83%), RTX 5070 at 9.4/12.2GB (77%), Strix Halo at 24.4/64GB (35%). Context reduced from 256K→128K on NVIDIA GPUs freed ~3.3GB (.8) and ~1.5GB (.110).
|
- **VRAM (2026-07-12)**: RTX 3090 at 20.7/24GB (84%), RTX 5070 at 10.0/12.2GB (82%). RTX 3090 context increased to 256K (was incorrectly documented as 128K — verified via /proc/PID/cmdline).
|
||||||
- **All GPUs at `--parallel 2` (2026-07-08)**: Fleet serves 6 concurrent requests (was 3). 2× throughput.
|
- **RTX 3090 runs `--parallel 1`** (verified 2026-07-12). RTX 5070 and Strix at parallel 2.
|
||||||
- **RTX 3090 config**: `-c 131072 -ctk turbo4 -ctv turbo4 --parallel 2`. No explicit batch flags (512/512 default). Service: `/home/llmuser/llama-wrapper.sh`.
|
- **RTX 3090 config**: `-c 262144 -ctk turbo4 -ctv turbo4 --parallel 1 --flash-attn on --cont-batching`. Service: `/home/llmuser/llama-wrapper.sh`.
|
||||||
- **RTX 5070 config**: `--ctx-size 131072 --cache-type-k q4_0 --cache-type-v q4_0 --batch-size 2048 --ubatch-size 512 --parallel 2`. Ubatch fixed 4096→512 (was inverted — ubatch > batch killed prompt throughput). Service: `/home/llmuser/llama-wrapper.sh`.
|
- **RTX 5070 config**: `--ctx-size 131072 --cache-type-k q4_0 --cache-type-v q4_0 --batch-size 2048 --ubatch-size 1024 --parallel 2`. Api-key standardized to `not-needed` (was `sk-loc...5678` causing 401 loops). Service: `/home/llmuser/llama-wrapper.sh`.
|
||||||
- **LiteLLM timeout tuning (2026-07-08)**: gemma-4-12b 25→120s, qwen3.6-27B-code 40→90s, syslog-auto (qwen route) 40→90s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s. Config at `/opt/inference-harness/litellm_config.yaml`.
|
- **LiteLLM timeout tuning (2026-07-08)**: gemma-4-12b 25→120s, qwen3.6-27B-code 40→90s, syslog-auto (qwen route) 40→90s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s. Config at `/opt/inference-harness/litellm_config.yaml`.
|
||||||
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `ornith-server.service` on port 8080, 256K context, flash-attn + q8 KV.
|
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `ornith-server.service` on port 8080, 256K context, flash-attn + q8 KV.
|
||||||
- **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill).
|
- **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill).
|
||||||
@@ -210,12 +249,12 @@ History stored at `/root/data/toks-history.json` with 7-day rolling window.
|
|||||||
|
|
||||||
**Note (2026-07-01)**: Strix Halo prompt tok/s jumped 209→532 after Vulkan rebuild (cooperative-matrix fast path now active on GFX1151). Baseline may need re-calibration.
|
**Note (2026-07-01)**: Strix Halo prompt tok/s jumped 209→532 after Vulkan rebuild (cooperative-matrix fast path now active on GFX1151). Baseline may need re-calibration.
|
||||||
|
|
||||||
## Agent Config Implications (2026-07-08)
|
## Agent Config Implications (2026-07-12)
|
||||||
|
|
||||||
With NVIDIA GPUs at 128K context:
|
With RTX 3090 at 256K context (verified July 2026):
|
||||||
- Agents using `syslog-auto` (50/50 qwen+ornith): keep `context_length: 262144` — ornith supports it, Litellm fallbacks handle qwen overflow
|
- Agents using `syslog-auto` (55/30/15 qwen+ornith+gemma): `context_length: 262144` — ornith and qwen both support it
|
||||||
- Agents using `qwen3.6-27B-code` directly: set `context_length: 131072` and `max_tokens: 4096` per thermal safety rule
|
- Agents using `qwen3.6-27B-code` directly: `context_length: 262144` (256K ctx verified)
|
||||||
- Agents using `gemma-4-12b` directly (auxiliary tasks): set `context_length: 131072`
|
- Agents using `gemma-4-12b` directly (auxiliary tasks): `context_length: 131072`
|
||||||
- Compression threshold at 0.65: fires at ~170K for syslog-auto (262K ctx), ~85K for direct qwen/gemma (128K ctx)
|
- Compression threshold at 0.65: fires at ~170K for 262K context window on Strix Halo
|
||||||
- All Hermes clients MUST set `max_tokens: 4096` — first line of defense before server-side `-n 8192` cap
|
- All Hermes clients MUST set `max_tokens: 4096` — first line of defense before server-side `-n 8192` cap
|
||||||
- Port 8080 is used on all 3 GPU hosts (not 8090 as previously documented)
|
- Port 8080 is used on all 3 GPU hosts (not 8090 as previously documented)
|
||||||
|
|||||||
@@ -0,0 +1,287 @@
|
|||||||
|
---
|
||||||
|
kind: responsibility
|
||||||
|
name: gpu-self-heal
|
||||||
|
description: >
|
||||||
|
GPU fleet self-healing — detects anomalies, applies remediation, tracks
|
||||||
|
benchmarks, and predicts failures before they happen. Extends gpu-monitor
|
||||||
|
(v2.1.0) with active remediation rules, Prometheus metrics consumption,
|
||||||
|
VRAM trend analysis, and predictive alerting.
|
||||||
|
agent: abiba
|
||||||
|
depends_on:
|
||||||
|
- gpu-monitor.prose.md (live data source on .24:9100)
|
||||||
|
- gpu-fleet.prose.md (source of truth for topology)
|
||||||
|
---
|
||||||
|
|
||||||
|
## Maintains
|
||||||
|
|
||||||
|
- gpu-health: { status: "healthy"|"degraded"|"down", issues: array, actions: array }
|
||||||
|
- gpu-self-heal-log: array of { timestamp, gpu, issue, action, result } — audit trail
|
||||||
|
- benchmark-regression: { gpu, baseline_tok_sec, current_tok_sec, trend, alerts }
|
||||||
|
- vram-trend: { gpu, current_mb, rate_mb_per_hour, projected_full_in_hours }
|
||||||
|
- circuit-breaker-status: { gpu, open, auto_reset_attempted, last_reset }
|
||||||
|
|
||||||
|
## Requires
|
||||||
|
|
||||||
|
- gpu-monitor:function — Live fleet data from .24:9100/gpu-data
|
||||||
|
- Prometheus exporters on all 3 GPUs (:9400/metrics)
|
||||||
|
- SSH access to GPU hosts for restart operations
|
||||||
|
|
||||||
|
## Continuity
|
||||||
|
|
||||||
|
- Self-driven: check every 60 seconds against GPU monitor data
|
||||||
|
- Also wakes on gpu-fleet health degradation
|
||||||
|
- On fix: verify with benchmark inference test before declaring resolved
|
||||||
|
- Escalate: after 3 failed remediation attempts → Zulip #agent-hub alert
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Remediation Rules
|
||||||
|
|
||||||
|
### Rule 1: GPU Temperature Critical (>85°C for >2 min)
|
||||||
|
- **Detect**: Any GPU temp >85°C sustained for 2+ consecutive polls
|
||||||
|
- **Fix**:
|
||||||
|
1. Reduce inference concurrency on that GPU (load-side cooling only — NO fan control)
|
||||||
|
2. Redirect new requests to cooler GPUs via LiteLLM fallback chains
|
||||||
|
3. If all GPUs hot, alert about cooling infrastructure
|
||||||
|
- **Verify**: Temp drops below 80°C within 5 minutes
|
||||||
|
- **Escalate after**: 3 verification failures → Zulip alert
|
||||||
|
|
||||||
|
### Rule 2: VRAM Leak Detection (tiered by GPU capacity)
|
||||||
|
- **Detect**: VRAM growing at sustained rate over 6+ hour window
|
||||||
|
- RTX 3090 (24GB): ≥100MB/hour
|
||||||
|
- RTX 5070 (12GB): ≥50MB/hour
|
||||||
|
- Strix Halo (64GB UMA): ≥200MB/hour
|
||||||
|
- **Fix**:
|
||||||
|
1. Log VRAM snapshot with process list (nvidia-smi/rocm-smi + ps aux)
|
||||||
|
2. If llama-server is the growth source → restart with memory cap flag
|
||||||
|
3. If unknown process → kill and alert
|
||||||
|
- **Verify**: VRAM growth rate drops below threshold
|
||||||
|
- **Escalate after**: persistent leak after restart → hardware investigation
|
||||||
|
|
||||||
|
### Rule 3: Model Inference Timeout / GPU Stuck
|
||||||
|
- **Detect**: >50% failure rate over 60s window + 30s grace period (not just single stuck request)
|
||||||
|
- **Fix**:
|
||||||
|
1. Restart llama-server on affected GPU host
|
||||||
|
2. Wait 15s for model to reload
|
||||||
|
3. Run benchmark inference test
|
||||||
|
- **Verify**: Model returns 200 with <30s response, failure rate drops to 0%
|
||||||
|
- **Escalate after**: 3 restarts in 1 hour → GPU hardware check
|
||||||
|
|
||||||
|
### Rule 4: Benchmark Regression (>20% drop)
|
||||||
|
- **Detect**: gen_tok_per_sec drops >20% below baseline over 3+ benchmarks
|
||||||
|
- **Fix**:
|
||||||
|
1. Check GPU utilization — if >90%, other process is competing
|
||||||
|
2. Check power limit — if throttled, restore to max
|
||||||
|
3. Check thermal — if hot, apply Rule 1
|
||||||
|
- **Verify**: Benchmark returns to within 10% of baseline
|
||||||
|
- **Escalate after**: persistent regression → possible hardware degradation
|
||||||
|
|
||||||
|
### Rule 5: Circuit Breaker Stuck Open
|
||||||
|
- **Detect**: Circuit breaker open >10 minutes with GPU reporting healthy
|
||||||
|
- **Fix**:
|
||||||
|
1. Verify GPU /health returns 200
|
||||||
|
2. If GPU healthy, send 1 test inference
|
||||||
|
3. If test succeeds → reset circuit breaker via router API
|
||||||
|
4. 60s cooldown — if CB re-opens immediately, it was legitimate, do NOT re-reset
|
||||||
|
5. Max 1 auto-reset per GPU per hour
|
||||||
|
- **Verify**: CB closes, inference succeeds, CB stays closed for 60s+
|
||||||
|
- **Escalate after**: CB won't close after reset → router issue
|
||||||
|
|
||||||
|
### Rule 6: Strix Halo Unreachable
|
||||||
|
- **Detect**: Strix not responding — probe .15:8080 directly (firewall opened .24→.15)
|
||||||
|
- **Fix**:
|
||||||
|
1. SSH to .15 → check llama-server process
|
||||||
|
2. Restart llama-server if not running
|
||||||
|
3. Verify through both direct probe AND router
|
||||||
|
- **Verify**: Direct health probe returns 200, router reports Strix healthy
|
||||||
|
- **Escalate**: If host .15 itself is unreachable → infrastructure alert
|
||||||
|
|
||||||
|
### Rule 7: Prometheus Exporter Down
|
||||||
|
- **Detect**: Any GPU :9400/metrics unreachable for >2 polls
|
||||||
|
- **Fix**:
|
||||||
|
1. SSH to GPU host → check prometheus-exporter process
|
||||||
|
2. Restart exporter if dead
|
||||||
|
3. While exporter is down, fall back to nvidia-smi/rocm-smi direct probes
|
||||||
|
4. If exporter is running but unreachable → check firewall/host networking
|
||||||
|
- **Verify**: :9400/metrics returns 200
|
||||||
|
- **Escalate after**: 3 failed restarts → networking issue
|
||||||
|
|
||||||
|
### Rule 8: Predictive Thermal Warning (two-tier)
|
||||||
|
- **Detect**:
|
||||||
|
- Tier 1 (warning): temp >70°C AND rising >2°C/min → reduce concurrency, no alert
|
||||||
|
- Tier 2 (critical): temp >80°C AND still rising → full alert + aggressive load shedding
|
||||||
|
- **Fix**:
|
||||||
|
- Tier 1: silently reduce parallel requests to that GPU by 50%
|
||||||
|
- Tier 2: redirect all new requests away, alert #agent-hub, apply Rule 1 logic
|
||||||
|
- **Verify**: Temp rise rate drops below 1°C/min (Tier 1) or temp drops below 80°C (Tier 2)
|
||||||
|
- **Escalate**: If Tier 2 triggers and temp still rising after 5 min → possible hardware failure
|
||||||
|
|
||||||
|
### Rule 9: Context Window Optimization
|
||||||
|
- **Detect**: Benchmark tok/s vs baseline for each GPU at current context
|
||||||
|
- RTX 3090 (256K ctx, qwen3.6-27B-code): target 75+ tok/s — currently at baseline
|
||||||
|
- RTX 5070 (131K ctx, gemma-4-12b): target 76+ tok/s — optimal for vision/web role
|
||||||
|
- Strix Halo (256K ctx, ornith-1.0-35b): target 70+ tok/s — currently above baseline
|
||||||
|
- **Fix**:
|
||||||
|
- If tok/s > baseline → context has headroom, consider increasing
|
||||||
|
- If tok/s < 90% baseline → reduce context by 25% and retest
|
||||||
|
- If tok/s within 10% of baseline → optimal, no change
|
||||||
|
- **Verify**: Re-benchmark after context change, confirm within 10% of target
|
||||||
|
- **Escalate**: If context can't be adjusted without significant perf loss
|
||||||
|
|
||||||
|
### Rule 10: Workload Distribution Optimization
|
||||||
|
- **Detect**: GPU roles misaligned with hardware capabilities
|
||||||
|
- **Target distribution**:
|
||||||
|
- RTX 3090 (24GB, 256K, 75 tok/s) → Heavy reasoning, code gen, long conversations
|
||||||
|
- RTX 5070 (12GB, 131K, 76 tok/s) → Vision/image, web search, quick lightweight tasks
|
||||||
|
- Strix Halo (64GB, 256K, 72 tok/s) → Context compression, summarization, long docs
|
||||||
|
- **Fix**:
|
||||||
|
- Alert if any GPU is handling workload outside its designated role
|
||||||
|
- Recommend Hermes agent profile updates to match workload to GPU
|
||||||
|
- Track per-GPU request distribution via LiteLLM spend logs
|
||||||
|
- **Verify**: Each GPU's request pattern matches its designated role within 24h
|
||||||
|
- **Escalate**: If role mismatch persists >48h → agent profile audit needed
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Execution
|
||||||
|
|
||||||
|
```prose
|
||||||
|
-- Phase 1: Fetch live GPU data
|
||||||
|
let fleet = call gpu-monitor
|
||||||
|
endpoint: "http://192.168.68.24:9100/gpu-data"
|
||||||
|
|
||||||
|
-- Phase 2: Evaluate each GPU against remediation rules
|
||||||
|
let actions = []
|
||||||
|
for gpu in fleet.gpus:
|
||||||
|
-- Rule 1: Thermal critical
|
||||||
|
if gpu.temp_c > 85 and sustained_for(gpu, 120):
|
||||||
|
push actions apply-thermal-fix(gpu)
|
||||||
|
|
||||||
|
-- Rule 2: VRAM leak
|
||||||
|
let vram_rate = calculate-vram-trend(gpu, hours=6)
|
||||||
|
if vram_rate > 50:
|
||||||
|
push actions apply-vram-fix(gpu, vram_rate)
|
||||||
|
|
||||||
|
-- Rule 4: Benchmark regression
|
||||||
|
let bench = fleet.benchmarks[gpu.hostname]
|
||||||
|
if bench.current_tok_sec < bench.baseline_tok_sec * 0.8:
|
||||||
|
push actions apply-benchmark-fix(gpu, bench)
|
||||||
|
|
||||||
|
-- Rule 3: Model stuck
|
||||||
|
for model in fleet.router.available_models:
|
||||||
|
if model.consecutive_timeouts >= 3:
|
||||||
|
push actions apply-model-restart(model)
|
||||||
|
|
||||||
|
-- Rule 5: Circuit breaker
|
||||||
|
for cb in fleet.router.circuit_breaker:
|
||||||
|
if cb.open and cb.open_duration > 600 and gpu_is_healthy(cb.gpu):
|
||||||
|
push actions apply-cb-reset(cb)
|
||||||
|
|
||||||
|
-- Rule 6: Strix Halo
|
||||||
|
if not fleet.strix.running and pingable("192.168.68.15"):
|
||||||
|
push actions apply-strix-restart()
|
||||||
|
|
||||||
|
-- Rule 7: Prometheus exporters
|
||||||
|
for gpu in fleet.gpus:
|
||||||
|
if not prometheus_reachable(gpu.hostname, 9400):
|
||||||
|
push actions apply-exporter-restart(gpu)
|
||||||
|
|
||||||
|
-- Rule 8: Predictive thermal
|
||||||
|
for gpu in fleet.gpus:
|
||||||
|
let rise_rate = calculate-temp-rise(gpu, minutes=5)
|
||||||
|
if rise_rate > 2.0 and gpu.temp_c < 80:
|
||||||
|
push actions apply-proactive-cooling(gpu)
|
||||||
|
|
||||||
|
-- Phase 3: Execute actions, verify, log
|
||||||
|
for action in actions:
|
||||||
|
let result = execute-with-verify(action)
|
||||||
|
log-to-kg(action, result)
|
||||||
|
if result.failed:
|
||||||
|
escalate-if-needed(action)
|
||||||
|
|
||||||
|
-- Phase 4: Update health state
|
||||||
|
call update-gpu-health
|
||||||
|
gpus: fleet.gpus
|
||||||
|
actions: actions
|
||||||
|
status: derive-overall-status(fleet, actions)
|
||||||
|
|
||||||
|
-- Wait 60s and repeat
|
||||||
|
```
|
||||||
|
|
||||||
|
## Audit Trail Format
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"run_id": "gpu-self-heal-20260712-001",
|
||||||
|
"timestamp": "2026-07-12T16:00:00Z",
|
||||||
|
"gpu": "ct8-rtx3090",
|
||||||
|
"issue": "thermal-critical",
|
||||||
|
"detected": { "temp_c": 87, "duration_s": 180 },
|
||||||
|
"action": "set-fan-100pct",
|
||||||
|
"result": "resolved",
|
||||||
|
"verification": { "temp_c": 76, "after_s": 300 },
|
||||||
|
"escalated": false
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Reporting
|
||||||
|
|
||||||
|
### 1. Knowledge Graph
|
||||||
|
Every action logged as `[GPU-SELF-HEAL] <run_id>` node with full audit trail.
|
||||||
|
|
||||||
|
### 2. Zulip Alerts (#agent-hub → alerts-gpu)
|
||||||
|
- `issues_fixed > 0` → "🛠 GPU Self-Heal — <gpu> <issue> resolved"
|
||||||
|
- `issues_escalated > 0` → "⚠ GPU Self-Heal — <gpu> needs attention"
|
||||||
|
- Every 100th clean cycle → "✅ GPU Fleet: All Clear"
|
||||||
|
|
||||||
|
### 3. Prometheus/Grafana Integration
|
||||||
|
- GPU self-heal actions exposed as Prometheus counter metrics
|
||||||
|
- Dashboard panel: "GPU Interventions (24h)" showing count/type/result
|
||||||
|
|
||||||
|
### 4. Weekly Benchmark Report
|
||||||
|
- Per-GPU tok/s trend over 7 days
|
||||||
|
- Regression alerts if any GPU degrades >10% week-over-week
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Design Decisions (Grilled & Confirmed — 2026-07-12)
|
||||||
|
|
||||||
|
1. **Fan control**: ❌ NO auto fan control. Load-side cooling only (reduce concurrency, redirect).
|
||||||
|
2. **Model restart**: ✅ Only if >50% failure rate over 60s + 30s grace period. Not on single stuck request.
|
||||||
|
3. **Strix direct access**: ✅ Open firewall .15:8080 → .24 for direct health probe + restart.
|
||||||
|
4. **VRAM thresholds**: Tiered — 100MB/h (RTX 3090), 50MB/h (RTX 5070), 200MB/h (Strix).
|
||||||
|
5. **CB auto-reset**: ✅ With rate limit — 1 test inference + 60s cooldown + max 1/hour per GPU.
|
||||||
|
6. **Benchmark baseline**: Rolling 30-day average, recalculated weekly. Original baseline kept in Grafana.
|
||||||
|
7. **Predictive alerts**: Two-tier — warn at >70°C+rising (>2°C/min), critical at >80°C+rising.
|
||||||
|
8. **Prometheus**: Primary source. Fall back to nvidia-smi/rocm-smi direct probes if exporter down.
|
||||||
|
|
||||||
|
## Lessons Learned (2026-07-12)
|
||||||
|
|
||||||
|
### L1: API Key Standardization Is Critical
|
||||||
|
- All GPU llama-servers MUST use the same api-key as the LiteLLM config.
|
||||||
|
- RTX 5070 had `--api-key sk-loc...5678` while LiteLLM sent `not-needed`.
|
||||||
|
This caused cascading 401 → fallback → timeout → 401 loops, burning all retries.
|
||||||
|
- **Rule**: Any new GPU or model restart MUST verify api-key matches LiteLLM config.
|
||||||
|
|
||||||
|
### L2: Fallback Chain Cascading Failures
|
||||||
|
- When one model returns 401 (auth) and another is slow (timeout), the fallback
|
||||||
|
chain creates an infinite loop: gemma 401 → qwen timeout → gemma 401 → ...
|
||||||
|
- **Rule**: If a model returns 401 (auth error), do NOT fall back to it again.
|
||||||
|
Mark it as permanently failed for this request.
|
||||||
|
|
||||||
|
### L3: Verify Running State, Not Docs
|
||||||
|
- RTX 3090 was documented at 128K context. Actually running at 256K.
|
||||||
|
- Parallel count wrong (docs said 2, actual is 1 on RTX 3090).
|
||||||
|
- **Rule**: Before making decisions, check `/proc/PID/cmdline` on GPU hosts.
|
||||||
|
|
||||||
|
### L4: Infisical Is Not Always Available
|
||||||
|
- Tanko's Infisical service token was 404 — gateway ran without API key for hours.
|
||||||
|
- **Rule**: Always keep a local `.env` fallback for `LITELLM_API_KEY`.
|
||||||
|
- Contract hermes-config-template Rule 3 updated.
|
||||||
|
|
||||||
|
### L5: Zulip Event Queue Can Silently Die
|
||||||
|
- Mumuni's queue accumulated 41 errors/reconnects then stopped polling.
|
||||||
|
Gateway was running but ignoring all messages.
|
||||||
|
- **Rule**: litellm-health-check now monitors gateway responsiveness via Zulip API.
|
||||||
@@ -22,30 +22,38 @@ done
|
|||||||
|
|
||||||
## Agent Map
|
## Agent Map
|
||||||
|
|
||||||
| Agent | CT | Node | IP | LiteLLM Key | LiteLLM Alias | Platform |
|
| Agent | CT | Node | IP | LiteLLM Alias | Key Source | Platform |
|
||||||
|-------|-----|------|-----|-------------|---------------|----------|
|
|-------|-----|------|-----|---------------|------------|----------|
|
||||||
| Tanko | 112 | amdpve | .122 | `sk-CggiHWlamQyShxWC3Hx6uw` | `tanko` | Hermes |
|
| Tanko | 112 | amdpve | .122 | `tanko` | Infisical vault | Hermes |
|
||||||
| Mumuni | 114 | minipve | .123 | `sk-VrqCNlwUgzoNGOpikJ7nwQ` | `mumuni` | Hermes |
|
| Mumuni | 114 | minipve | .123 | `mumuni` | Infisical vault | Hermes |
|
||||||
| Tdunna | 111 | amdpve | srv1079750 | `sk-Qvzi4uYQBhlSK_XstEhcyQ` | `tdunna` | **pi** |
|
| Koby | 111 | amdpve | srv1079750 | `koby` | Infisical vault | **Hermes** |
|
||||||
| Baggy | 113 | amdpve | ? | `sk-krnw_zGBwvvL5b7l2t-s-A` | `baggy` | Hermes |
|
| Koonimo | 113 | amdpve | ? | `koonimo` | Infisical vault | Hermes |
|
||||||
|
| Shumba | — | 192.168.68.119 | N/A | N/A (DeepSeek) | Hermes |
|
||||||
|
|
||||||
|
> **Note**: CT hostnames (tdunna→CT111, baggy→CT113) differ from agent identities (koby, koonimo).
|
||||||
|
|
||||||
Access: `pct-run <CT_ID> <command>` — no IPs needed. GPU hosts (.8, .110, .15) use SSH.
|
Access: `pct-run <CT_ID> <command>` — no IPs needed. GPU hosts (.8, .110, .15) use SSH.
|
||||||
|
Keys are stored in Infisical vault (project=agents, env=production) and injected at
|
||||||
|
runtime via `infisical run --` wrapper. Plaintext keys removed from this baseline.
|
||||||
|
|
||||||
## Key Architecture
|
## Key Architecture
|
||||||
|
|
||||||
```
|
```
|
||||||
|
Infisical vault → infisical run -- hermes gateway → LITELLM_API_KEY (runtime)
|
||||||
|
↓
|
||||||
Agent (systemd) → LITELLM_API_KEY → LiteLLM (:116/v1) → Router (:9000) → GPU (llama-server)
|
Agent (systemd) → LITELLM_API_KEY → LiteLLM (:116/v1) → Router (:9000) → GPU (llama-server)
|
||||||
└── Key DB (Postgres)
|
└── Key DB (Postgres)
|
||||||
```
|
```
|
||||||
|
|
||||||
- **Master key**: `sk-litellm-7f96080dd99b15c36bd4b333b58a6796` — ADMIN ONLY, never in agent configs
|
- **Master key**: stored in Infisical vault (project=infrastructure, secret=LITELLM_MASTER_KEY) — ADMIN ONLY
|
||||||
- **Agent keys**: Each agent has a dedicated key in LiteLLM's database with alias matching the agent name
|
- **Agent keys**: Each agent has a dedicated key in LiteLLM's database with alias matching the agent name
|
||||||
- **Key source**: `/etc/environment` → `LITELLM_API_KEY=sk-...` (systemd service sources this)
|
- **Key injection**: `infisical run --project=agents --env=production -- hermes gateway run` injects `LITELLM_API_KEY` at runtime
|
||||||
- **Override**: `/home/jerome/.config/systemd/user/hermes-gateway.service.d/env.conf` (if present, must match)
|
- **Key source**: Infisical vault → runtime env var. /etc/environment is CLEAN (stripped, tagged `# [INFISICAL]`)
|
||||||
|
- **Legacy override** (pre-migration): `/home/jerome/.config/systemd/user/hermes-gateway.service.d/env.conf` — should be REMOVED
|
||||||
|
|
||||||
## Config Pattern — Mandatory Fields
|
## Config Pattern — Mandatory Fields
|
||||||
|
|
||||||
### For Hermes Agents (Tanko, Mumuni, Baggy)
|
### For Hermes Agents (Tanko, Mumuni, Koonimo)
|
||||||
|
|
||||||
Every agent's `/root/.hermes/config.yaml` (or `/home/jerome/.hermes/config.yaml`) MUST have:
|
Every agent's `/root/.hermes/config.yaml` (or `/home/jerome/.hermes/config.yaml`) MUST have:
|
||||||
|
|
||||||
@@ -76,7 +84,7 @@ auxiliary:
|
|||||||
model: gemma-4-12b # or syslog-auto
|
model: gemma-4-12b # or syslog-auto
|
||||||
base_url: http://192.168.68.116/v1
|
base_url: http://192.168.68.116/v1
|
||||||
api_key_env: LITELLM_API_KEY
|
api_key_env: LITELLM_API_KEY
|
||||||
api_key: <ACTUAL_KEY_FROM_/etc/environment> # ← MANDATORY workaround
|
api_key: <value from: infisical secrets get LITELLM_API_KEY --project=agents --env=production> # ← MANDATORY workaround
|
||||||
timeout: 60
|
timeout: 60
|
||||||
download_timeout: 30
|
download_timeout: 30
|
||||||
```
|
```
|
||||||
@@ -91,7 +99,7 @@ auxiliary:
|
|||||||
model: syslog-auto # or gemma-4-12b
|
model: syslog-auto # or gemma-4-12b
|
||||||
base_url: http://192.168.68.116/v1
|
base_url: http://192.168.68.116/v1
|
||||||
api_key_env: LITELLM_API_KEY
|
api_key_env: LITELLM_API_KEY
|
||||||
api_key: <ACTUAL_KEY_FROM_/etc/environment> # ← MANDATORY workaround
|
api_key: <value from: infisical secrets get LITELLM_API_KEY --project=agents --env=production> # ← MANDATORY workaround
|
||||||
timeout: 120
|
timeout: 120
|
||||||
```
|
```
|
||||||
|
|
||||||
@@ -110,8 +118,7 @@ LiteLLM/harness will fail with:
|
|||||||
401: LiteLLM Virtual Key expected. Received=no-k****ired, expected to start with 'sk-'
|
401: LiteLLM Virtual Key expected. Received=no-k****ired, expected to start with 'sk-'
|
||||||
```
|
```
|
||||||
|
|
||||||
**Workaround**: Set `api_key` directly (copy the value from `/etc/environment`) alongside
|
**Workaround**: Set `api_key` directly (copy the value from Infisical vault: `infisical secrets get LITELLM_API_KEY --project=agents --env=production`) alongside `api_key_env` in every auxiliary task config that uses the harness provider.
|
||||||
`api_key_env` in every auxiliary task config that uses the harness provider.
|
|
||||||
|
|
||||||
**Permanent fix**: Patch `_resolve_task_provider_model()` to resolve `api_key_env` when
|
**Permanent fix**: Patch `_resolve_task_provider_model()` to resolve `api_key_env` when
|
||||||
`api_key` is empty:
|
`api_key` is empty:
|
||||||
@@ -129,7 +136,11 @@ if not cfg_api_key:
|
|||||||
```bash
|
```bash
|
||||||
for ct in 112 114 111 113; do
|
for ct in 112 114 111 113; do
|
||||||
echo "=== CT $ct ==="
|
echo "=== CT $ct ==="
|
||||||
pct-run $ct grep LITELLM_API_KEY /etc/environment
|
# Verify /etc/environment is CLEAN (no LITELLM_API_KEY)
|
||||||
|
pct-run $ct "grep -c LITELLM_API_KEY /etc/environment 2>/dev/null || echo '0 (clean)'"
|
||||||
|
# Verify gateway uses infisical run wrapper
|
||||||
|
pct-run $ct "ps aux | grep 'infisical run' | grep -v grep"
|
||||||
|
# Check for hardcoded harness keys
|
||||||
pct-run $ct grep "api_key: sk-" /root/.hermes/config.yaml | grep -v api_key_env
|
pct-run $ct grep "api_key: sk-" /root/.hermes/config.yaml | grep -v api_key_env
|
||||||
echo ""
|
echo ""
|
||||||
done
|
done
|
||||||
@@ -137,15 +148,17 @@ done
|
|||||||
|
|
||||||
### Master Key Leak Check
|
### Master Key Leak Check
|
||||||
```bash
|
```bash
|
||||||
# On every agent:
|
# On every agent — must return empty:
|
||||||
pct-run <CT> grep -rl "sk-litellm-7f96080dd" /root/ /etc/ 2>/dev/null
|
pct-run <CT> grep -rl "sk-litellm" /root/ /etc/ 2>/dev/null
|
||||||
# Must return empty
|
# Vault is the only place the master key should exist
|
||||||
```
|
```
|
||||||
|
|
||||||
### Verify Key Works
|
### Verify Key Works
|
||||||
```bash
|
```bash
|
||||||
|
# Retrieve key from vault and test:
|
||||||
|
KEY=$(infisical secrets get LITELLM_API_KEY --project=agents --env=production --plain)
|
||||||
curl -s http://192.168.68.116:80/v1/models \
|
curl -s http://192.168.68.116:80/v1/models \
|
||||||
-H "Authorization: Bearer <AGENT_KEY>" | grep syslog-auto
|
-H "Authorization: Bearer $KEY" | grep syslog-auto
|
||||||
# Must return model list
|
# Must return model list
|
||||||
```
|
```
|
||||||
|
|
||||||
@@ -156,19 +169,27 @@ pct-run <CT> grep -A8 "vision:" /root/.hermes/config.yaml | grep api_key
|
|||||||
# Must show both api_key: sk-... and api_key_env: LITELLM_API_KEY
|
# Must show both api_key: sk-... and api_key_env: LITELLM_API_KEY
|
||||||
```
|
```
|
||||||
|
|
||||||
### For pi Agents (Tdunna)
|
### For Koby (CT 111 / tdunna)
|
||||||
|
|
||||||
Tdunna (CT111) runs pi 0.80.3 via PM2 with the Zulip extension (router-worker architecture).
|
Koby runs Hermes on CT 111 (tdunna). Config files at `/root/.hermes/config.yaml`.
|
||||||
|
Same Hermes pattern as Tanko/Mumuni/Koonimo — see config sections above.
|
||||||
|
|
||||||
|
**LiteLLM key**: alias `koby` in LiteLLM DB, injected via `infisical run --` wrapper.
|
||||||
|
|
||||||
|
### For pi Agents (Abiba)
|
||||||
|
|
||||||
|
Abiba (CT100) runs pi via PM2 with the Zulip extension.
|
||||||
Config files: `~/.pi/agent/models.json`, `~/.pi/agent/settings.json`.
|
Config files: `~/.pi/agent/models.json`, `~/.pi/agent/settings.json`.
|
||||||
|
|
||||||
**models.json** — Must only list models authorized for the agent's LiteLLM key:
|
**models.json** — Must only list models authorized for the agent's LiteLLM key.
|
||||||
|
Key is injected via `infisical run --` wrapper at PM2 startup:
|
||||||
```json
|
```json
|
||||||
{
|
{
|
||||||
"providers": {
|
"providers": {
|
||||||
"syslog-harness": {
|
"syslog-harness": {
|
||||||
"baseUrl": "http://192.168.68.116/v1",
|
"baseUrl": "http://192.168.68.116/v1",
|
||||||
"api": "openai-completions",
|
"api": "openai-completions",
|
||||||
"apiKey": "sk-...",
|
"apiKey": "${LITELLM_API_KEY}",
|
||||||
"models": [
|
"models": [
|
||||||
{ "id": "syslog-auto" },
|
{ "id": "syslog-auto" },
|
||||||
{ "id": "ornith-1.0-35b" },
|
{ "id": "ornith-1.0-35b" },
|
||||||
@@ -257,6 +278,6 @@ If they differ → ghost detected → kill ghost → start fresh.
|
|||||||
|
|
||||||
| Date | Change |
|
| Date | Change |
|
||||||
|------|--------|
|
|------|--------|
|
||||||
| 2026-07-08 | Tdunna: fixed model mismatch (qwen3.6-35B-A3B→syslog-auto), added pi-specific config section. Key updated to sk-Qvzi4uYQBhlSK_XstEhcyQ. Added Failure Mode #11 to zulip-adapter-lessons. |
|
| 2026-07-08 | Koby: fixed model mismatch (qwen3.6-35B-A3B→syslog-auto), added config section. Key rotated and stored in vault. Added Failure Mode #11 to zulip-adapter-lessons. |
|
||||||
| 2026-07-06 | Port conflict detection added to all 3 GPU wrappers. Consolidated health check script deployed. Zulip streaming edit_message enabled for Tanko/Mumuni. |
|
| 2026-07-06 | Port conflict detection added to all 3 GPU wrappers. Consolidated health check script deployed. Zulip streaming edit_message enabled for Tanko/Mumuni. |
|
||||||
| 2026-07-05 | Baseline created. All 4 agents audited, master key removed, api_key workaround applied |
|
| 2026-07-05 | Baseline created. All 4 agents audited, master key removed, api_key workaround applied |
|
||||||
|
|||||||
@@ -5,35 +5,40 @@ description: >
|
|||||||
Standard Hermes configuration template for Syslog Solution LLC agents.
|
Standard Hermes configuration template for Syslog Solution LLC agents.
|
||||||
Enforces shared infrastructure setup (Firecrawl, SearXNG, local models,
|
Enforces shared infrastructure setup (Firecrawl, SearXNG, local models,
|
||||||
RA-H OS MCP) while keeping agent-specific API keys and model choices.
|
RA-H OS MCP) while keeping agent-specific API keys and model choices.
|
||||||
Updated 2026-07-08: GPU context reduced to 128K on NVIDIA (.8, .110),
|
UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo (ornith-1.0-35b).
|
||||||
parallel 2 on all GPUs, LiteLLM timeouts tuned, context_length guidance added.
|
RTX 3090 context verified at 256K (was incorrectly documented as 128K).
|
||||||
|
RTX 5070 stays at 131K for vision/web. Infisical .env fallback required (Rule 3).
|
||||||
|
Parallel counts corrected (RTX 3090=1, RTX 5070=2, Strix=2).
|
||||||
---
|
---
|
||||||
|
|
||||||
## Maintains
|
## Maintains
|
||||||
|
|
||||||
- template_version: "2.1.0"
|
- template_version: "2.1.0"
|
||||||
- last_applied: timestamp
|
- last_applied: timestamp
|
||||||
- agents_configured: ["tanko", "mumuni", "abiba", "tdunna", "baggy", "kagenz0"]
|
- agents_configured: ["tanko", "mumuni", "abiba", "koby", "koonimo", "kagenz0"]
|
||||||
- agent_keys: map (see Agent Keys section)
|
- agent_keys: map (see Agent Keys section)
|
||||||
- infra_endpoints_verified: array
|
- infra_endpoints_verified: array
|
||||||
|
|
||||||
## Agent Keys (LiteLLM — Current 2026-07-04)
|
## Agent Keys (LiteLLM — Current 2026-07-11)
|
||||||
|
|
||||||
Each agent has a unique LiteLLM API key (virtual key) generated against the LiteLLM
|
Each agent has a unique LiteLLM API key (virtual key) generated against the LiteLLM
|
||||||
PostgreSQL DB via `POST /key/generate` on CT 116. Keys are stored in the DB, not in
|
PostgreSQL DB via `POST /key/generate` on CT 116. Keys are stored in the DB.
|
||||||
config files. The env var `LITELLM_API_KEY` is set in `/etc/environment` on each agent
|
The env var `LITELLM_API_KEY` is injected at runtime via `infisical run --` wrapper
|
||||||
host AND in `~/.hermes/.env` for gateway env propagation.
|
(project=agents, env=production). /etc/environment and ~/.hermes/.env are NO LONGER
|
||||||
|
used for agent keys — stripped and tagged `# [INFISICAL]` post-migration.
|
||||||
Sub-agent profiles inherit auth from the main config — no separate keys needed.
|
Sub-agent profiles inherit auth from the main config — no separate keys needed.
|
||||||
|
|
||||||
| Agent | Key Alias | Host | SSH | Sub-Agents |
|
| Agent | Key Alias | Host | SSH | Sub-Agents |
|
||||||
|-------|-----------|------|-----|-----------|
|
|-------|-----------|------|-----|-----------|
|
||||||
| Tanko | `tanko-*` | 192.168.68.122 | jerome@.122 | — |
|
| Tanko | `tanko-*` | 192.168.68.122 | jerome@.122 | — |
|
||||||
| Mumuni | `mumuni-jul2026` | 192.168.68.123 | root@.123 | 6 profiles ✱ |
|
| Mumuni | `mumuni-jul2026` | 192.168.68.123 | root@.123 | 6 profiles ✱ |
|
||||||
| Abiba | `abiba-*` | 192.168.68.24 | local | — |
|
| Abiba | `abiba-pi` | 192.168.68.24 | local | — |
|
||||||
| Tdunna | `tdunna-*` | ? | Zulip | — |
|
| Koby | `koby` | CT 111 (tdunna) | Zulip | — |
|
||||||
| Baggy | `baggy-*` | ? | Zulip | — |
|
| Koonimo | `koonimo` | CT 113 (baggy) | Zulip | — |
|
||||||
| Kagenz0 | `kagenz0-*` | ? | Zulip | — |
|
| Kagenz0 | `kagenz0-*` | ? | Zulip | — |
|
||||||
|
|
||||||
|
> CT hostnames (tdunna, baggy) differ from agent identities (koby, koonimo).
|
||||||
|
|
||||||
✱ Mumuni sub-agents: syslog-code, syslog-devops, syslog-email, syslog-research,
|
✱ Mumuni sub-agents: syslog-code, syslog-devops, syslog-email, syslog-research,
|
||||||
syslog-review, syslog-writer — all at `/root/.hermes/profiles/<name>/config.yaml`
|
syslog-review, syslog-writer — all at `/root/.hermes/profiles/<name>/config.yaml`
|
||||||
|
|
||||||
@@ -51,11 +56,12 @@ Sub-agent profiles inherit auth from the main config — no separate keys needed
|
|||||||
## API Key Rules
|
## API Key Rules
|
||||||
|
|
||||||
- `api_key_env: LITELLM_API_KEY` — Use env var for main model auth (preferred)
|
- `api_key_env: LITELLM_API_KEY` — Use env var for main model auth (preferred)
|
||||||
|
Key is injected at runtime via `infisical run --` wrapper — never in /etc/environment
|
||||||
- `api_key: ''` — Sub-agents leave empty to inherit from main config's custom_provider
|
- `api_key: ''` — Sub-agents leave empty to inherit from main config's custom_provider
|
||||||
- `api_key: sk-...` — Hardcoded key only as fallback when env var not possible
|
- `api_key: sk-...` — Hardcoded key only as fallback when env var not possible
|
||||||
- Set `LITELLM_API_KEY` in `/etc/environment` on each host
|
- Store `LITELLM_API_KEY` in Infisical vault (project=agents, env=production)
|
||||||
- Sub-agents NEVER get their own key — they share the host agent's key
|
- Sub-agents NEVER get their own key — they share the host agent's key
|
||||||
- Restart Hermes after updating `/etc/environment`
|
- Restart Hermes gateway after updating vault secret (key auto-injected via wrapper)
|
||||||
|
|
||||||
### Sub-Agent Profiles (Mumuni pattern)
|
### Sub-Agent Profiles (Mumuni pattern)
|
||||||
|
|
||||||
@@ -79,7 +85,7 @@ Sub-agent profile rules:
|
|||||||
5. **Auxiliary tasks** (vision, compression, etc.) also leave `api_key` empty
|
5. **Auxiliary tasks** (vision, compression, etc.) also leave `api_key` empty
|
||||||
6. **Never hardcode a key** in sub-agent profiles
|
6. **Never hardcode a key** in sub-agent profiles
|
||||||
|
|
||||||
This ensures all 6 sub-agents use the same LiteLLM key set in `/etc/environment`.
|
This ensures all 6 sub-agents use the same LiteLLM key injected via `infisical run --` wrapper.
|
||||||
When the key is rotated, only the env var needs updating — all 7 configs (main + 6 subs)
|
When the key is rotated, only the env var needs updating — all 7 configs (main + 6 subs)
|
||||||
work immediately after restart.
|
work immediately after restart.
|
||||||
|
|
||||||
@@ -91,10 +97,10 @@ model:
|
|||||||
default: <agent_model> # e.g., ornith-1.0-35b, qwen3.6-27B-code
|
default: <agent_model> # e.g., ornith-1.0-35b, qwen3.6-27B-code
|
||||||
provider: harness
|
provider: harness
|
||||||
base_url: http://192.168.68.116/v1
|
base_url: http://192.168.68.116/v1
|
||||||
api_key_env: LITELLM_API_KEY # Set in /etc/environment AND ~/.hermes/.env
|
api_key_env: LITELLM_API_KEY # Injected via infisical run -- wrapper
|
||||||
max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation
|
max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation
|
||||||
context_length: 262144 # For syslog-auto (ornith route supports 256K).
|
context_length: 262144 # For syslog-auto (all GPUs support 256K).
|
||||||
# Set 131072 if using qwen3.6-27B-code or gemma-4-12b directly.
|
# Set 131072 if using gemma-4-12b directly (12GB VRAM constraint).
|
||||||
|
|
||||||
fallback_providers:
|
fallback_providers:
|
||||||
provider: deepseek
|
provider: deepseek
|
||||||
@@ -123,10 +129,10 @@ mcp_servers:
|
|||||||
# ─── Compression ───
|
# ─── Compression ───
|
||||||
compression:
|
compression:
|
||||||
enabled: true
|
enabled: true
|
||||||
model: gemma-4-12b # ⚠️ Must match auxiliary.compression.model
|
model: ornith-1.0-35b # ⚠️ Must match auxiliary.compression.model
|
||||||
provider: harness
|
provider: harness
|
||||||
max_context_window: 262144 # For syslog-auto (ornith supports 256K).
|
max_context_window: 262144 # For Strix Halo compression (256K ctx).
|
||||||
# Set 131072 if using qwen or gemma directly.
|
# Set 131072 if using gemma directly.
|
||||||
threshold: 0.65 # Fires at ~170K for 262K window, ~85K for 128K
|
threshold: 0.65 # Fires at ~170K for 262K window, ~85K for 128K
|
||||||
target_ratio: 0.30
|
target_ratio: 0.30
|
||||||
protect_last_n: 40
|
protect_last_n: 40
|
||||||
@@ -158,7 +164,7 @@ auxiliary:
|
|||||||
timeout: 30
|
timeout: 30
|
||||||
compression:
|
compression:
|
||||||
provider: harness
|
provider: harness
|
||||||
model: gemma-4-12b
|
model: ornith-1.0-35b
|
||||||
base_url: http://192.168.68.116/v1
|
base_url: http://192.168.68.116/v1
|
||||||
api_key_env: LITELLM_API_KEY
|
api_key_env: LITELLM_API_KEY
|
||||||
timeout: 60
|
timeout: 60
|
||||||
@@ -176,7 +182,7 @@ custom_providers:
|
|||||||
|
|
||||||
When LiteLLM keys are regenerated (e.g., after infrastructure changes):
|
When LiteLLM keys are regenerated (e.g., after infrastructure changes):
|
||||||
|
|
||||||
1. **If SSH available**: `ssh <host> "sudo sed -i 's/LITELLM_API_KEY=.*/LITELLM_API_KEY=sk-<NEW>/' /etc/environment"`
|
1. **If SSH available**: Update Infisical vault: `infisical secrets set LITELLM_API_KEY=sk-<NEW> --project=agents --env=production`, then `ssh <host> "systemctl restart hermes-gateway"`
|
||||||
2. **If SSH unavailable**: Send Zulip DM via abiba-bot with update command
|
2. **If SSH unavailable**: Send Zulip DM via abiba-bot with update command
|
||||||
3. **After update**: Restart Hermes on the agent host
|
3. **After update**: Restart Hermes on the agent host
|
||||||
4. **Verify**: `curl -H "Authorization: Bearer sk-<KEY>" http://192.168.68.116/v1/models`
|
4. **Verify**: `curl -H "Authorization: Bearer sk-<KEY>" http://192.168.68.116/v1/models`
|
||||||
@@ -198,7 +204,14 @@ The following MUST be identical across ALL profiles:
|
|||||||
### Rule 3: API Keys via Environment
|
### Rule 3: API Keys via Environment
|
||||||
- Prefer `api_key_env: LITELLM_API_KEY` over hardcoded keys
|
- Prefer `api_key_env: LITELLM_API_KEY` over hardcoded keys
|
||||||
- Hardcoded keys in config.yaml become stale after key rotation
|
- Hardcoded keys in config.yaml become stale after key rotation
|
||||||
- `/etc/environment` persists across config updates
|
- Infisical vault secrets persist across config updates / reinstalls
|
||||||
|
- **NEW (July 2026): Always keep a local `.env` fallback.** Infisical service tokens
|
||||||
|
can expire/404 (tanko incident: token not found, gateway ran without key for hours).
|
||||||
|
The `.env` file should have the key uncommented as a fallback:
|
||||||
|
```
|
||||||
|
LITELLM_API_KEY=sk-...
|
||||||
|
# [INFISICAL] Also sourced from vault.sysloggh.net
|
||||||
|
```
|
||||||
- Restart Hermes after env var updates
|
- Restart Hermes after env var updates
|
||||||
|
|
||||||
### Rule 4: Sub-Agent Profiles Inherit Auth
|
### Rule 4: Sub-Agent Profiles Inherit Auth
|
||||||
@@ -208,7 +221,7 @@ The following MUST be identical across ALL profiles:
|
|||||||
- Auxiliary tasks: `api_key: ''`, `provider: harness`
|
- Auxiliary tasks: `api_key: ''`, `provider: harness`
|
||||||
- Never hardcode a key in sub-agent profiles
|
- Never hardcode a key in sub-agent profiles
|
||||||
- When main config uses `api_key_env`, sub-agents automatically use it
|
- When main config uses `api_key_env`, sub-agents automatically use it
|
||||||
- This means key rotation only touches ONE file (`/etc/environment`)
|
- This means key rotation only touches ONE vault secret (`LITELLM_API_KEY`)
|
||||||
|
|
||||||
### Rule 5: Main Config Base URL
|
### Rule 5: Main Config Base URL
|
||||||
|- Use direct IP: `http://192.168.68.116/v1`
|
|- Use direct IP: `http://192.168.68.116/v1`
|
||||||
@@ -223,23 +236,42 @@ The following MUST be identical across ALL profiles:
|
|||||||
- Apply to BOTH main config AND all sub-agent profiles
|
- Apply to BOTH main config AND all sub-agent profiles
|
||||||
- For agents needing longer outputs: raise to 8192, but never omit
|
- For agents needing longer outputs: raise to 8192, but never omit
|
||||||
|
|
||||||
### Rule 7: Auxiliary Model Consistency
|
### Rule 7: Auxiliary Model Consistency (UPDATED July 2026)
|
||||||
- All auxiliary services (vision, web_extract, compression) MUST use the same model:
|
- Vision and web_extract use `gemma-4-12b` (RTX 5070 — 12GB, vision-optimized)
|
||||||
- `model: gemma-4-12b`
|
- Compression uses `ornith-1.0-35b` (Strix Halo — 64GB, 256K ctx, compression-optimized)
|
||||||
|
- All auxiliary services MUST use identical routing:
|
||||||
- `base_url: http://192.168.68.116/v1`
|
- `base_url: http://192.168.68.116/v1`
|
||||||
- `api_key_env: LITELLM_API_KEY`
|
- `api_key_env: LITELLM_API_KEY`
|
||||||
- **Do NOT use `syslog-auto`** for auxiliary tasks — it routes to the primary 35B reasoning GPU
|
- **Do NOT use `syslog-auto`** for auxiliary tasks — it routes unpredictably
|
||||||
- gemma-4-12b is a lightweight 12B model on the RTX 5070, keeping the Strix Halo free for reasoning
|
- **Compression moved to Strix Halo (July 2026)**: The ornith-1.0-35b model on Strix Halo
|
||||||
- The `compression:` block's `model` MUST match `auxiliary: compression: model` — they are two different configs for the same service
|
(64GB UMA, 256K context, 72.4 tok/s) is the designated compression GPU. This frees the
|
||||||
|
RTX 5070 for vision and web search, and the RTX 3090 for heavy reasoning.
|
||||||
|
- The `compression:` block's `model` MUST match `auxiliary: compression: model`
|
||||||
|
- The `compression: max_context_window: 262144` MUST match Strix Halo's actual capacity
|
||||||
|
|
||||||
### Rule 8: Compression Threshold for 256K Models
|
### Rule 8: GPU Workload Distribution (July 2026)
|
||||||
|
- **RTX 3090 (24GB, 256K ctx, qwen3.6-27B-code)**: Heavy reasoning, code gen, long conversations
|
||||||
|
- **RTX 5070 (12GB, 131K ctx, gemma-4-12b)**: Vision, web search, quick tasks, web_extract
|
||||||
|
- **Strix Halo (64GB, 256K ctx, ornith-1.0-35b)**: Context compression, summarization, long docs
|
||||||
|
- Agent profiles MUST route auxiliary tasks to the correct GPU:
|
||||||
|
- `auxiliary.vision.model: gemma-4-12b` (RTX 5070)
|
||||||
|
- `auxiliary.web_extract.model: gemma-4-12b` (RTX 5070)
|
||||||
|
- `auxiliary.compression.model: ornith-1.0-35b` (Strix Halo)
|
||||||
|
- Default model (`model.default`) and custom_provider remain `syslog-auto` for auto-routing
|
||||||
- For 262K context window: `threshold: 0.65` (fires at ~170K tokens)
|
- For 262K context window: `threshold: 0.65` (fires at ~170K tokens)
|
||||||
- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss
|
- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss
|
||||||
- Do NOT use `threshold: 0.80` — this delays until 209K, risking the gateway hygiene layer
|
- Do NOT use `threshold: 0.80` — this delays until 209K, risking the gateway hygiene layer
|
||||||
- `max_context_window: 262144` MUST match the model's actual capacity
|
- `max_context_window: 262144` MUST match the model's actual capacity
|
||||||
- See `devops-hermes-compression` skill for full reference
|
- See `devops-hermes-compression` skill for full reference
|
||||||
|
|
||||||
### Rule 9: Default Model Must Be `syslog-auto` (All Agents)
|
### Rule 9: Compression Threshold for 256K Models
|
||||||
|
- For 262K context window: `threshold: 0.65` (fires at ~170K tokens)
|
||||||
|
- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss
|
||||||
|
- Do NOT use `threshold: 0.80` — this delays until 209K, risking the gateway hygiene layer
|
||||||
|
- `max_context_window: 262144` MUST match the model's actual capacity (Strix Halo = 256K)
|
||||||
|
- See `devops-hermes-compression` skill for full reference
|
||||||
|
|
||||||
|
### Rule 10: Default Model Must Be `syslog-auto` (All Agents)
|
||||||
- **Hermes agents**: `model.default: syslog-auto`, `custom_providers[0].model: syslog-auto`
|
- **Hermes agents**: `model.default: syslog-auto`, `custom_providers[0].model: syslog-auto`
|
||||||
- **pi agents**: `defaultModel: syslog-auto` in `settings.json`, first model in `models.json`
|
- **pi agents**: `defaultModel: syslog-auto` in `settings.json`, first model in `models.json`
|
||||||
- `syslog-auto` is the LiteLLM routing model — it load-balances between ornith-1.0-35b
|
- `syslog-auto` is the LiteLLM routing model — it load-balances between ornith-1.0-35b
|
||||||
@@ -250,7 +282,7 @@ The following MUST be identical across ALL profiles:
|
|||||||
- **Exception**: Sub-agent profiles (Mumuni's 6 profiles) may specify explicit models
|
- **Exception**: Sub-agent profiles (Mumuni's 6 profiles) may specify explicit models
|
||||||
for specialized tasks, but MUST validate those models exist in the key's authorized list
|
for specialized tasks, but MUST validate those models exist in the key's authorized list
|
||||||
|
|
||||||
### Rule 10: Validate Model IDs Before Deployment (pi Agents)
|
### Rule 11: Validate Model IDs Before Deployment (pi Agents)
|
||||||
- After configuring a pi agent's `models.json`, verify every model ID:
|
- After configuring a pi agent's `models.json`, verify every model ID:
|
||||||
```bash
|
```bash
|
||||||
curl -s http://192.168.68.116:4000/v1/models \
|
curl -s http://192.168.68.116:4000/v1/models \
|
||||||
|
|||||||
+119
-48
@@ -5,8 +5,10 @@ version: 1.0.0
|
|||||||
description: >
|
description: >
|
||||||
Enforces standardized API key configuration across all Hermes agents. Harness/LiteLLM
|
Enforces standardized API key configuration across all Hermes agents. Harness/LiteLLM
|
||||||
providers MUST use api_key_env indirection. External providers (DeepSeek, OpenAI,
|
providers MUST use api_key_env indirection. External providers (DeepSeek, OpenAI,
|
||||||
Anthropic) may use hardcoded keys. Single source of truth: /etc/environment on each
|
Anthropic) may use hardcoded keys. Single source of truth: Infisical vault
|
||||||
agent host. Designed to make key rotation a one-step operation.
|
(project=agents, env=production) — injected at runtime via `infisical run --` wrapper.
|
||||||
|
/etc/environment is DEPRECATED for agent keys post-migration. Designed to make key
|
||||||
|
rotation a one-step vault operation.
|
||||||
author: Abiba (pi agent)
|
author: Abiba (pi agent)
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -14,7 +16,7 @@ author: Abiba (pi agent)
|
|||||||
|
|
||||||
## Rule (One Sentence)
|
## Rule (One Sentence)
|
||||||
|
|
||||||
**Any `api_key` pointing to a Syslog-hosted LiteLLM/harness provider MUST be replaced with `api_key_env: LITELLM_API_KEY` — hardcoded harness keys are forbidden.**
|
**All harness/litellm providers MUST use `api_key_env: LITELLM_API_KEY` with authenticated path `http://192.168.68.116/litellm/v1/responses` — hardcoded keys AND unauthenticated `/v1` direct access are both forbidden.**
|
||||||
|
|
||||||
## Scope
|
## Scope
|
||||||
|
|
||||||
@@ -26,6 +28,39 @@ Applies to all Hermes agent configs across all hosts. Covers these config sectio
|
|||||||
- `compression.api_key` (when `provider` is `harness` or contains `litellm`)
|
- `compression.api_key` (when `provider` is `harness` or contains `litellm`)
|
||||||
- `fallback_providers[].api_key` (when provider is harness)
|
- `fallback_providers[].api_key` (when provider is harness)
|
||||||
|
|
||||||
|
## Architecture (2026-07-10)
|
||||||
|
|
||||||
|
Syslog is migrating away from **unauthenticated direct access** to the shared inference harness.
|
||||||
|
|
||||||
|
| Path | Auth | Status |
|
||||||
|
|------|------|--------|
|
||||||
|
| `http://192.168.68.116/v1` | None (direct) | ❌ **DEPRECATED** — being phased out |
|
||||||
|
| `http://192.168.68.116/litellm/v1/responses` | Bearer `sk-*` key | ✅ **CURRENT** — authenticated LiteLLM proxy |
|
||||||
|
|
||||||
|
All harness/litellm providers MUST use the authenticated `/litellm/v1/responses` path.
|
||||||
|
Any `base_url` pointing to bare `/v1` on 192.168.68.116 is a **migration violation**.
|
||||||
|
|
||||||
|
### 🔥 CRITICAL: Double-Path Bug (2026-07-10)
|
||||||
|
|
||||||
|
When `api_mode: responses` is set, Hermes **appends `/v1/responses`** to `base_url`.
|
||||||
|
If `base_url` already includes `/litellm/v1/responses`, the result is:
|
||||||
|
|
||||||
|
```
|
||||||
|
http://192.168.68.116/litellm/v1/responses/v1/responses → 404
|
||||||
|
```
|
||||||
|
|
||||||
|
**The `base_url` must end at `/v1` — never include `/responses`:**
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
# ✅ CORRECT — Hermes appends /v1/responses for api_mode: responses
|
||||||
|
base_url: http://192.168.68.116/litellm/v1
|
||||||
|
|
||||||
|
# ❌ WRONG — produces double path
|
||||||
|
base_url: http://192.168.68.116/litellm/v1/responses
|
||||||
|
```
|
||||||
|
|
||||||
|
This applies to ALL sections using the harness provider: `custom_providers`, `delegation`, `auxiliary.*`.
|
||||||
|
|
||||||
## Exemptions
|
## Exemptions
|
||||||
|
|
||||||
External providers are **explicitly exempt** and may use hardcoded keys:
|
External providers are **explicitly exempt** and may use hardcoded keys:
|
||||||
@@ -38,35 +73,42 @@ External providers are **explicitly exempt** and may use hardcoded keys:
|
|||||||
## Standard Pattern
|
## Standard Pattern
|
||||||
|
|
||||||
```yaml
|
```yaml
|
||||||
# ✅ CORRECT — all harness/litellm providers
|
# ✅ CORRECT — all harness/litellm providers (authenticated path, NO /responses suffix)
|
||||||
model:
|
model:
|
||||||
provider: harness # or custom:litellm.sysloggh.net
|
provider: harness
|
||||||
base_url: http://192.168.68.116/v1
|
base_url: http://192.168.68.116/litellm/v1 # ← Hermes appends /v1/responses
|
||||||
api_key_env: LITELLM_API_KEY # ← indirection
|
api_key_env: LITELLM_API_KEY
|
||||||
|
|
||||||
custom_providers:
|
custom_providers:
|
||||||
- name: harness
|
- name: harness
|
||||||
base_url: http://192.168.68.116/v1
|
api_mode: responses
|
||||||
api_key_env: LITELLM_API_KEY # ← indirection
|
base_url: http://192.168.68.116/litellm/v1 # ← NO /responses suffix!
|
||||||
|
api_key_env: LITELLM_API_KEY
|
||||||
|
|
||||||
auxiliary:
|
auxiliary:
|
||||||
compression:
|
compression:
|
||||||
provider: harness
|
provider: harness
|
||||||
api_key_env: LITELLM_API_KEY # ← indirection
|
base_url: http://192.168.68.116/litellm/v1 # ← NO /responses suffix!
|
||||||
|
api_key_env: LITELLM_API_KEY
|
||||||
|
|
||||||
# ✅ ALSO CORRECT — external providers
|
# ✅ ALSO CORRECT — external providers
|
||||||
fallback_providers:
|
fallback_providers:
|
||||||
- provider: deepseek
|
- provider: deepseek
|
||||||
base_url: https://api.deepseek.com
|
base_url: https://api.deepseek.com
|
||||||
api_key: sk-b7d9... # ← hardcoded OK (external)
|
api_key: sk-b7d9... # ← hardcoded OK (external)
|
||||||
api_key_env: DEEPSEEK_API_KEY # ← also OK if set in /etc/environment
|
api_key_env: DEEPSEEK_API_KEY # ← also OK if set in environment (vault or /etc/environment)
|
||||||
```
|
```
|
||||||
|
|
||||||
```yaml
|
```yaml
|
||||||
# ❌ FORBIDDEN — hardcoded harness/litellm key
|
# ❌ FORBIDDEN — hardcoded key (top) OR unauthenticated path (bottom)
|
||||||
model:
|
model:
|
||||||
provider: harness
|
provider: harness
|
||||||
api_key: sk-Flc62smlegyMEaSo1ka8JA # ← RULE VIOLATION
|
api_key: sk-Flc62smlegyMEaSo1ka8JA # ← RULE VIOLATION: hardcoded key
|
||||||
|
|
||||||
|
model:
|
||||||
|
provider: harness
|
||||||
|
base_url: http://192.168.68.116/v1 # ← RULE VIOLATION: unauthenticated path
|
||||||
|
api_key_env: LITELLM_API_KEY
|
||||||
```
|
```
|
||||||
|
|
||||||
## Detection Query
|
## Detection Query
|
||||||
@@ -75,13 +117,18 @@ Run on any Hermes host to detect violations:
|
|||||||
|
|
||||||
```bash
|
```bash
|
||||||
# 1. Check config.yaml for hardcoded harness keys
|
# 1. Check config.yaml for hardcoded harness keys
|
||||||
grep -rn 'api_key: sk-' /home/jerome/.hermes/ \
|
grep -rn 'api_key: sk-' /root/.hermes/ \
|
||||||
--include='config.yaml' \
|
--include='config.yaml' \
|
||||||
| grep -v 'deepseek\|openai\|anthropic\|DEEPSEEK'
|
| grep -v 'deepseek\|openai\|anthropic\|DEEPSEEK'
|
||||||
|
|
||||||
|
# 1b. Check for double-path bug: base_url ending with /responses
|
||||||
|
# (Hermes appends /v1/responses when api_mode=responses, so base_url must end at /v1)
|
||||||
|
grep -rn 'litellm/v1/responses' /root/.hermes/config.yaml
|
||||||
|
# ANY output here = WRONG. Must be 'litellm/v1' without /responses suffix.
|
||||||
|
|
||||||
# 2. Check systemd drop-ins for master key leaks (2026-07-05: Tanko had this)
|
# 2. Check systemd drop-ins for master key leaks (2026-07-05: Tanko had this)
|
||||||
grep -rn 'LITELLM_API_KEY' /home/jerome/.config/systemd/user/ 2>/dev/null
|
grep -rn 'LITELLM_API_KEY' /root/.config/systemd/user/ 2>/dev/null
|
||||||
grep -rn 'LITELLM_API_KEY=sk-litellm-7f96080d' /home/jerome/.config/systemd/ 2>/dev/null
|
grep -rn 'LITELLM_API_KEY=sk-litellm-7f96080d' /root/.config/systemd/ 2>/dev/null
|
||||||
|
|
||||||
# 3. Verify running process env matches dedicated key
|
# 3. Verify running process env matches dedicated key
|
||||||
cat /proc/$(cat /home/jerome/.hermes/gateway.pid | python3 -c "import sys,json; print(json.load(sys.stdin)['pid'])")/environ \
|
cat /proc/$(cat /home/jerome/.hermes/gateway.pid | python3 -c "import sys,json; print(json.load(sys.stdin)['pid'])")/environ \
|
||||||
@@ -92,26 +139,32 @@ If any output from step 2 — **critical violation** (master key leaked). Fix im
|
|||||||
|
|
||||||
## Rotation Procedure
|
## Rotation Procedure
|
||||||
|
|
||||||
With this standard enforced, key rotation is one step:
|
With this standard enforced, key rotation is one vault update:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
# 1. Generate new key in LiteLLM
|
# 1. Generate new key in LiteLLM: POST /key/generate with agent alias
|
||||||
# 2. Update /etc/environment on agent host
|
# 2. Update Infisical vault secret
|
||||||
ssh root@<host> "sed -i 's/LITELLM_API_KEY=.*/LITELLM_API_KEY=sk-NEW_KEY/' /etc/environment"
|
infisical secrets set LITELLM_API_KEY=sk-NEW_KEY \
|
||||||
# 3. Restart agent gateway
|
--project=agents --env=production
|
||||||
ssh root@<host> "pkill -f 'hermes_cli.main gateway run'; sleep 2; nohup ... &"
|
# 3. Restart agent gateway (key auto-injected via infisical run -- wrapper)
|
||||||
|
ssh root@<host> "systemctl restart hermes-gateway"
|
||||||
# 4. Verify
|
# 4. Verify
|
||||||
curl -s -H "Authorization: Bearer sk-NEW_KEY" http://192.168.68.116/v1/models
|
curl -s -H "Authorization: Bearer sk-NEW_KEY" http://192.168.68.116/litellm/v1/models
|
||||||
```
|
```
|
||||||
|
|
||||||
**Done.** No config file changes needed. The agent picks up the new key on restart.
|
**Done.** No config file changes needed. No /etc/environment edits needed.
|
||||||
|
The agent picks up the new key via `infisical run --` at gateway startup.
|
||||||
|
|
||||||
|
> **Post-migration note**: /etc/environment is NO LONGER the key source.
|
||||||
|
> Strip all `LITELLM_API_KEY` lines from /etc/environment (comment out with `# [INFISICAL]`)
|
||||||
|
> and let the `infisical run --` wrapper inject the key at runtime.
|
||||||
|
|
||||||
## Key Longevity Policy (2026-07-04)
|
## Key Longevity Policy (2026-07-04)
|
||||||
|
|
||||||
**Keys are permanent and use bare agent name aliases.**
|
**Keys are permanent and use bare agent name aliases.**
|
||||||
|
|
||||||
- **Duration**: `null` — keys never expire. This is enforced by `default_key_generate_params` in `litellm_config.yaml`.
|
- **Duration**: `null` — keys never expire. This is enforced by `default_key_generate_params` in `litellm_config.yaml`.
|
||||||
- **Alias convention**: bare agent name only (e.g., `tanko`, `mumuni`, `tdunna`, `baggy`). No dates, no versions. The alias IS the identity.
|
- **Alias convention**: bare agent name only (e.g., `tanko`, `mumuni`, `koby`, `koonimo`). No dates, no versions. The alias IS the identity.
|
||||||
- **Rotation triggers**: compromise, personnel departure, or quarterly security hygiene. NOT calendar-driven.
|
- **Rotation triggers**: compromise, personnel departure, or quarterly security hygiene. NOT calendar-driven.
|
||||||
- **Max budget**: $100 per key (config default).
|
- **Max budget**: $100 per key (config default).
|
||||||
|
|
||||||
@@ -128,36 +181,54 @@ litellm_settings:
|
|||||||
|
|
||||||
## Verified Agents (2026-07-05 update)
|
## Verified Agents (2026-07-05 update)
|
||||||
|
|
||||||
| Agent | CT | IP | LiteLLM Alias | Key | Status | Systemd Source | Last Verified |
|
| Agent | CT | IP | LiteLLM Alias | Key Source | Status | Gateway Wrapper | Last Verified |
|
||||||
|-------|-----|-----|---------------|-----|--------|----------------|---------------|
|
|-------|-----|-----|---------------|------------|--------|-----------------|---------------|
|
||||||
| Tanko | 112 | .122 | `tanko` | `sk-CggiHWlamQyShxWC3Hx6uw` | ✅ Fixed | User drop-in `env.conf` | 20:17 UTC Jul 5 |
|
| Tanko | 112 | .122 | `tanko` | Infisical vault | ✅ Fixed | `infisical run` | 20:17 UTC Jul 5 |
|
||||||
| Mumuni | 114 | .123 | `mumuni` | `sk-XY2aUfvy2BIs6kp1ZPh6VA` | ⚠️ Unverified | `/etc/environment` | 23:00 EDT Jul 4 |
|
| Mumuni | 114 | .123 | `mumuni` | Infisical vault | ✅ Fixed | `infisical run` | 01:46 EDT Jul 10 |
|
||||||
| Tdunna | 111 | ? | `tdunna` | `sk-6sbCNjz2T6lTVDBdlNHXsA` | ✅ Fixed | `/etc/environment` + drop-in | 23:30 UTC Jul 5 |
|
| Koby | 111 | ? | `koby` | Infisical vault | ✅ Fixed | `infisical run` | 23:30 UTC Jul 5 |
|
||||||
| Baggy | 113 | ? | `baggy` | `sk-krnw_zGBwvvL5b7l2t-s-A` | ✅ Fixed | `/etc/environment` | 23:30 UTC Jul 5 |
|
| Koonimo | 113 | ? | `koonimo` | Infisical vault | ✅ Fixed | `infisical run` (migrated 2026-07-11) | 2026-07-11 |
|
||||||
| Abiba | 100 | .65 | — | — | ✅ N/A (pi native) | — | 19:44 UTC Jul 5 |
|
| Abiba | 100 | .65 | `abiba-pi` | Infisical vault | ✅ N/A (pi native) | — | 19:44 UTC Jul 5 |
|
||||||
| Kagenz0 | 105 | ? | — | — | ❌ DOWN | — | 19:14 EDT Jul 4 |
|
| Kagenz0 | 105 | ? | — | — | ❌ DOWN | — | 19:14 EDT Jul 4 |
|
||||||
|
|
||||||
### Systemd Service Pattern (2026-07-04 fix)
|
> **Note**: CT hostnames (tdunna, baggy) differ from agent identities (koby, koonimo).
|
||||||
|
> LiteLLM key aliases use agent identity, not CT hostname.
|
||||||
|
|
||||||
All Hermes agents use systemd to manage their gateway. Two issues were fixed:
|
### Migration Status: Authenticated Path
|
||||||
|
|
||||||
1. **Drop-in override** — `/etc/systemd/system/hermes-gateway.service.d/litellm-key.conf` (or user equivalent) had hardcoded `LITELLM_API_KEY` that bypassed `/etc/environment`.
|
| Agent | `/litellm/v1/responses` | Deprecated `/v1` | Status |
|
||||||
2. **Missing EnvironmentFile** — Services did not source `/etc/environment`.
|
|-------|--------------------------|--------------------|--------|
|
||||||
|
| Mumuni | ✅ 5 sections | 0 | ✅ Authenticated |
|
||||||
|
| Tanko | ⚠️ No SSH access | — | Needs check |
|
||||||
|
| Koby | ⚠️ No route to host | — | Needs check |
|
||||||
|
| Koonimo | ⚠️ Connection timed out | — | Needs check |
|
||||||
|
|
||||||
**Correct pattern:**
|
### Systemd Service Pattern (2026-07-11 — vault migration)
|
||||||
|
|
||||||
|
All Hermes agents use systemd to manage their gateway. The gateway service is wrapped
|
||||||
|
with `infisical run --` to inject secrets at runtime.
|
||||||
|
|
||||||
|
**Correct pattern (post-migration):**
|
||||||
```ini
|
```ini
|
||||||
# In service file:
|
# Service file wraps gateway with Infisical:
|
||||||
EnvironmentFile=/etc/environment
|
[Service]
|
||||||
|
ExecStart=/usr/bin/infisical run --project=agents --env=production -- \
|
||||||
|
/usr/bin/hermes gateway run
|
||||||
|
|
||||||
# Drop-in only for overrides, NOT primary key storage.
|
# /etc/environment is CLEAN — no LITELLM_API_KEY present
|
||||||
# If a drop-in exists, it must match /etc/environment.
|
# (strip it and tag with # [INFISICAL] if present)
|
||||||
```
|
```
|
||||||
|
|
||||||
**Rotation procedure** (one step with this standard):
|
**Legacy pattern (deprecated — pre-migration only):**
|
||||||
|
```ini
|
||||||
|
# DO NOT USE post-migration:
|
||||||
|
EnvironmentFile=/etc/environment
|
||||||
|
# This pattern was replaced by infisical run -- wrapper
|
||||||
|
```
|
||||||
|
|
||||||
|
**Rotation procedure** (one vault operation with this standard):
|
||||||
1. Generate new key in LiteLLM: `curl /key/generate` with agent alias
|
1. Generate new key in LiteLLM: `curl /key/generate` with agent alias
|
||||||
2. Update `/etc/environment`: `sed -i 's/LITELLM_API_KEY=.*/LITELLM_API_KEY=sk-NEW/' /etc/environment`
|
2. Update Infisical vault: `infisical secrets set LITELLM_API_KEY=sk-NEW --project=agents --env=production`
|
||||||
3. Update drop-in (if exists): same sed on `litellm-key.conf`
|
3. Restart: `systemctl restart hermes-gateway` (key auto-injected via wrapper)
|
||||||
4. Restart: `systemctl [--user] restart hermes-gateway`
|
|
||||||
|
|
||||||
## Violation Response
|
## Violation Response
|
||||||
|
|
||||||
@@ -167,7 +238,7 @@ EnvironmentFile=/etc/environment
|
|||||||
4. **Restart** — gateway must restart to pick up env var
|
4. **Restart** — gateway must restart to pick up env var
|
||||||
5. **Confirm** — test key against LiteLLM: `curl -H "Authorization: Bearer $KEY" .../v1/models` → 200
|
5. **Confirm** — test key against LiteLLM: `curl -H "Authorization: Bearer $KEY" .../v1/models` → 200
|
||||||
6. **Update** — bump the verified table above
|
6. **Update** — bump the verified table above
|
||||||
7. **Use safe-mutate** — if the fix requires changing `/etc/environment` or restarting the gateway on a remote host, use `safe-mutate` to verify current state before mutating.
|
7. **Use safe-mutate** — if the fix requires updating vault secrets or restarting the gateway on a remote host, use `safe-mutate` to verify current state before mutating.
|
||||||
|
|
||||||
## Related Contracts
|
## Related Contracts
|
||||||
|
|
||||||
@@ -206,13 +277,13 @@ task config:
|
|||||||
```yaml
|
```yaml
|
||||||
auxiliary:
|
auxiliary:
|
||||||
vision:
|
vision:
|
||||||
api_key: sk-CggiHWlamQyShxWC3Hx6uw # ← workaround
|
api_key: sk-<agent-key-from-vault> # ← workaround (get via: infisical secrets get LITELLM_API_KEY --project=agents --env=production --plain)
|
||||||
api_key_env: LITELLM_API_KEY
|
api_key_env: LITELLM_API_KEY
|
||||||
base_url: http://192.168.68.116/v1
|
base_url: http://192.168.68.116/v1
|
||||||
model: gemma-4-12b
|
model: gemma-4-12b
|
||||||
provider: harness
|
provider: harness
|
||||||
compression:
|
compression:
|
||||||
api_key: sk-CggiHWlamQyShxWC3Hx6uw # ← workaround
|
api_key: sk-<agent-key-from-vault> # ← workaround (same as above)
|
||||||
api_key_env: LITELLM_API_KEY
|
api_key_env: LITELLM_API_KEY
|
||||||
base_url: http://192.168.68.116/v1
|
base_url: http://192.168.68.116/v1
|
||||||
model: gemma-4-12b
|
model: gemma-4-12b
|
||||||
|
|||||||
@@ -28,7 +28,7 @@ connectivity recovery including end-to-end DM validation.
|
|||||||
|
|
||||||
| Param | Type | Required | Default | Description |
|
| Param | Type | Required | Default | Description |
|
||||||
|-------|------|----------|---------|-------------|
|
|-------|------|----------|---------|-------------|
|
||||||
| `target` | string | yes | — | Agent name: `mumuni`, `tanko`, or `koby` |
|
| `target` | string | yes | — | Agent name: `mumuni`, `tanko`, `koby`, or `shumba` |
|
||||||
| `branch` | string | no | `master` | Git branch to pull (overridable for pinning) |
|
| `branch` | string | no | `master` | Git branch to pull (overridable for pinning) |
|
||||||
|
|
||||||
## Maintains
|
## Maintains
|
||||||
@@ -58,6 +58,7 @@ connectivity recovery including end-to-end DM validation.
|
|||||||
| Mumuni | CT114 | — | 192.168.68.123 | /root/.hermes | root |
|
| Mumuni | CT114 | — | 192.168.68.123 | /root/.hermes | root |
|
||||||
| Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome |
|
| Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome |
|
||||||
| Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root |
|
| Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root |
|
||||||
|
| Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky |
|
||||||
|
|
||||||
| Field | Value | Trust |
|
| Field | Value | Trust |
|
||||||
|-------|-------|-------|
|
|-------|-------|-------|
|
||||||
|
|||||||
@@ -3,7 +3,7 @@ kind: function
|
|||||||
name: hermes-zulip-restore
|
name: hermes-zulip-restore
|
||||||
description: >
|
description: >
|
||||||
Restores Zulip connectivity for any Hermes agent (Mumuni CT114, Tanko CT112,
|
Restores Zulip connectivity for any Hermes agent (Mumuni CT114, Tanko CT112,
|
||||||
Koby CT111). Deploys the zulip-platform adapter to the correct bundled plugin
|
Koby CT111, Shumba on Lucky's mini PC). Deploys the zulip-platform adapter to the correct bundled plugin
|
||||||
path, verifies env credentials, restarts the gateway, and confirms Zulip
|
path, verifies env credentials, restarts the gateway, and confirms Zulip
|
||||||
connects. Run this whenever a Hermes agent stops responding on Zulip or after
|
connects. Run this whenever a Hermes agent stops responding on Zulip or after
|
||||||
a fresh agent deployment.
|
a fresh agent deployment.
|
||||||
@@ -23,7 +23,7 @@ gateway restart, and connection validation.
|
|||||||
|
|
||||||
| Param | Type | Required | Default | Description |
|
| Param | Type | Required | Default | Description |
|
||||||
|-------|------|----------|---------|-------------|
|
|-------|------|----------|---------|-------------|
|
||||||
| `target` | string | yes | — | Agent name: `mumuni`, `tanko`, or `koby` |
|
| `target` | string | yes | — | Agent name: `mumuni`, `tanko`, `koby`, or `shumba` |
|
||||||
|
|
||||||
## Maintains
|
## Maintains
|
||||||
|
|
||||||
@@ -54,6 +54,7 @@ gateway restart, and connection validation.
|
|||||||
| Mumuni | CT114 | — | 192.168.68.123 | /root/.hermes | root |
|
| Mumuni | CT114 | — | 192.168.68.123 | /root/.hermes | root |
|
||||||
| Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome |
|
| Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome |
|
||||||
| Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root |
|
| Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root |
|
||||||
|
| Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky |
|
||||||
|
|
||||||
| Field | Value | Trust |
|
| Field | Value | Trust |
|
||||||
|-------|-------|-------|
|
|-------|-------|-------|
|
||||||
|
|||||||
@@ -62,19 +62,19 @@ description: >
|
|||||||
|
|
||||||
| Resource | Auth Method | Credential Source | Status |
|
| Resource | Auth Method | Credential Source | Status |
|
||||||
|----------|------------|-------------------|--------|
|
|----------|------------|-------------------|--------|
|
||||||
| Proxmox Cluster | PVE API Token | `monitoring@pve!mumuni=...` | ✅ |
|
| Proxmox Cluster | PVE API Token | Infisical vault (`PROXMOX_API_TOKEN`) | ✅ |
|
||||||
| Proxmox Root | Password via API ticket | `root@pam:kakashi19` | ✅ |
|
| Proxmox Root | Password via API ticket | Infisical vault (`PROXMOX_ROOT_PASSWORD`) | ✅ |
|
||||||
| docker-vm (.7) | SSH root | SSH key | ✅ |
|
| docker-vm (.7) | SSH root | SSH key | ✅ |
|
||||||
| CT 116 (syslog-api) | SSH root | SSH key | ✅ |
|
| CT 116 (syslog-api) | SSH root | SSH key | ✅ |
|
||||||
| Tanko CT (.122) | SSH jerome | id_ed25519 | ✅ |
|
| Tanko CT (.122) | SSH jerome | id_ed25519 | ✅ |
|
||||||
| Mumuni CT (.123) | SSH root | id_ed25519 | ✅ |
|
| Mumuni CT (.123) | SSH root | id_ed25519 | ✅ |
|
||||||
| Baggy CT (113) | SSH jerome | ❌ no key access |
|
| Baggy CT (113) | SSH jerome | ❌ no key access |
|
||||||
| Netbird (.17) | SSH root | SSH key | ✅ |
|
| Netbird (.17) | SSH root | SSH key | ✅ |
|
||||||
| Gitea | API token | abiba-bot token | ✅ |
|
| Gitea | API token | Infisical vault (`GITEA_BOT_TOKEN`) | ✅ |
|
||||||
| Zulip | Bot API key | abiba-bot@chat.sysloggh.net | ✅ |
|
| Zulip | Bot API key | Infisical vault (`ZULIP_BOT_KEY`) | ✅ |
|
||||||
| RA-H OS | MCP bridge | port 3100 | ✅ |
|
| RA-H OS | MCP bridge | port 3100 | ✅ |
|
||||||
| LiteLLM Admin | master key | `sk-litellm-7f96...` | ✅ VERIFY-BEFORE-USE |
|
| LiteLLM Admin | master key | Infisical vault (`LITELLM_MASTER_KEY`) | ✅ VERIFY-BEFORE-USE |
|
||||||
| Grafana | admin password | `syslog-grafana-2026` | ✅ VERIFY-BEFORE-USE |
|
| Grafana | admin password | Infisical vault (`GRAFANA_ADMIN_PASSWORD`) | ✅ VERIFY-BEFORE-USE |
|
||||||
|
|
||||||
> **VERIFY-BEFORE-USE**: Credentials, IPs, ports, and hostnames in this
|
> **VERIFY-BEFORE-USE**: Credentials, IPs, ports, and hostnames in this
|
||||||
> contract are live-state fields. Test them against the live system before
|
> contract are live-state fields. Test them against the live system before
|
||||||
@@ -550,13 +550,13 @@ curl -s http://192.168.68.116/health/unified | jq .status
|
|||||||
curl -s http://192.168.68.24:9100/gpu-data | jq .summary
|
curl -s http://192.168.68.24:9100/gpu-data | jq .summary
|
||||||
|
|
||||||
# Grafana status
|
# Grafana status
|
||||||
curl -s http://admin:syslog-grafana-2026@192.168.68.116:3001/api/health
|
curl -s http://admin:$(infisical secrets get GRAFANA_ADMIN_PASSWORD --project=infrastructure --env=production --plain)@192.168.68.116:3001/api/health
|
||||||
|
|
||||||
# Prometheus targets
|
# Prometheus targets
|
||||||
curl -s http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets[] | {job: .labels.job, health: .health}'
|
curl -s http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets[] | {job: .labels.job, health: .health}'
|
||||||
|
|
||||||
# LiteLLM key check
|
# LiteLLM key check
|
||||||
curl -s -H "Authorization: Bearer sk-litellm-7f96080dd99b15c36bd4b333b58a6796" \
|
curl -s -H "Authorization: Bearer $(infisical secrets get LITELLM_MASTER_KEY --project=infrastructure --env=production --plain)" \
|
||||||
http://192.168.68.116/litellm/key/list | jq '.keys[] | {alias: .key_alias, models: .models}'
|
http://192.168.68.116/litellm/key/list | jq '.keys[] | {alias: .key_alias, models: .models}'
|
||||||
|
|
||||||
# Storage check
|
# Storage check
|
||||||
@@ -586,7 +586,7 @@ ssh root@192.168.68.110 "systemctl restart llama-server"
|
|||||||
| 108 | media | storepve | — | Media | ❌ |
|
| 108 | media | storepve | — | Media | ❌ |
|
||||||
| 109 | docker-vm | storepve | .7 | Docker host | ❌ |
|
| 109 | docker-vm | storepve | .7 | Docker host | ❌ |
|
||||||
| 110 | gitea | minipve | — | Git | ❌ |
|
| 110 | gitea | minipve | — | Git | ❌ |
|
||||||
| 111 | tdunna | amdpve | — | ? | ❌ |
|
| 111 | tdunna | amdpve | .129 | Hermes agent | ✅ |
|
||||||
| 112 | tanko | amdpve | .122 | Hermes agent | ✅ |
|
| 112 | tanko | amdpve | .122 | Hermes agent | ✅ |
|
||||||
| 113 | baggy | amdpve | ? | Hermes agent | ✅ |
|
| 113 | baggy | amdpve | ? | Hermes agent | ✅ |
|
||||||
| 114 | mumuni | minipve | .123 | Hermes agent | ✅ |
|
| 114 | mumuni | minipve | .123 | Hermes agent | ✅ |
|
||||||
|
|||||||
@@ -11,7 +11,7 @@ triggers:
|
|||||||
- on "infra update" command
|
- on "infra update" command
|
||||||
- weekly (Sunday 03:00 EDT) via cron
|
- weekly (Sunday 03:00 EDT) via cron
|
||||||
- on security advisory relay from Mumuni
|
- on security advisory relay from Mumuni
|
||||||
version: 1.1.0
|
version: 1.2.0
|
||||||
---
|
---
|
||||||
|
|
||||||
## Maintains
|
## Maintains
|
||||||
@@ -27,11 +27,12 @@ Before ANY update wave:
|
|||||||
2. ✅ All critical VMs/CTs running (VM 109 docker-vm, CT 116 syslog-api, CT 117 zulip, CT 106 ra-h-os)
|
2. ✅ All critical VMs/CTs running (VM 109 docker-vm, CT 116 syslog-api, CT 117 zulip, CT 106 ra-h-os)
|
||||||
3. ✅ GPU bare-metal hosts reachable: .8 (RTX 3090), .110 (RTX 5070), .15 (Strix Halo)
|
3. ✅ GPU bare-metal hosts reachable: .8 (RTX 3090), .110 (RTX 5070), .15 (Strix Halo)
|
||||||
4. ✅ Docker healthy on VM 109 (.7), CT 116 (.116)
|
4. ✅ Docker healthy on VM 109 (.7), CT 116 (.116)
|
||||||
5. ✅ LiteLLM health check passing
|
5. ✅ LiteLLM health check passing (port 4000, /mcp-rest/tools/list with master key)
|
||||||
6. ✅ Zulip server reachable
|
6. ✅ LiteLLM MCP gateway serving RA-H OS tools (90 tools)
|
||||||
6. ✅ GPU fleet healthy (all 3 GPUs: RTX 3090, RTX 5070, RX 7600)
|
7. ✅ Zulip server reachable
|
||||||
7. ✅ Disk >20% free on all nodes
|
8. ✅ GPU fleet healthy (all 3 GPUs: RTX 3090, RTX 5070, RX 7600)
|
||||||
8. 📋 Snapshot critical configs (LiteLLM, nginx, docker-compose files)
|
9. ✅ Disk >20% free on all nodes
|
||||||
|
10. 📋 Snapshot critical configs (LiteLLM, nginx, docker-compose files)
|
||||||
|
|
||||||
## Wave 1: Storage & Infra Nodes (lowest impact)
|
## Wave 1: Storage & Infra Nodes (lowest impact)
|
||||||
|
|
||||||
@@ -65,6 +66,7 @@ Before ANY update wave:
|
|||||||
**Verify after Wave 2:**
|
**Verify after Wave 2:**
|
||||||
- All VMs/CTs running: check via Proxmox API
|
- All VMs/CTs running: check via Proxmox API
|
||||||
- LiteLLM healthy: `curl localhost:4000/health/liveliness` (via CT 116)
|
- LiteLLM healthy: `curl localhost:4000/health/liveliness` (via CT 116)
|
||||||
|
- LiteLLM MCP tools: `curl localhost:4000/mcp-rest/tools/list -H "Authorization: Bearer $MASTER_KEY"` → 90 tools
|
||||||
- GPU servers responding: check :8080 on VM 101, VM 103; check ornith via router (http://192.168.68.116/health/unified — .15:8080 is firewalled to .116 only)
|
- GPU servers responding: check :8080 on VM 101, VM 103; check ornith via router (http://192.168.68.116/health/unified — .15:8080 is firewalled to .116 only)
|
||||||
- Zulip agents connected: check Mumuni/Tanko gateway state
|
- Zulip agents connected: check Mumuni/Tanko gateway state
|
||||||
- Abiba PM2 processes online: `pm2 status`
|
- Abiba PM2 processes online: `pm2 status`
|
||||||
@@ -83,6 +85,7 @@ Before ANY update wave:
|
|||||||
**Verify after Wave 3:**
|
**Verify after Wave 3:**
|
||||||
- All containers healthy: `docker ps` on each host
|
- All containers healthy: `docker ps` on each host
|
||||||
- End-to-end inference test: `curl localhost:4000/v1/chat/completions` (via CT 116) with syslog-auto
|
- End-to-end inference test: `curl localhost:4000/v1/chat/completions` (via CT 116) with syslog-auto
|
||||||
|
- MCP integration test: `curl localhost:4000/mcp-rest/tools/list -H "Authorization: Bearer $MASTER_KEY"` → 90 tools (23 RA-H OS + 67 GitHub)
|
||||||
- Zulip test: send test message to #agent-hub
|
- Zulip test: send test message to #agent-hub
|
||||||
- Dashboard loading: `curl localhost:3001/` (via CT 116)
|
- Dashboard loading: `curl localhost:3001/` (via CT 116)
|
||||||
- Firecrawl test: `curl :3002/`
|
- Firecrawl test: `curl :3002/`
|
||||||
@@ -112,8 +115,8 @@ If ANY verification fails:
|
|||||||
|
|
||||||
Before Wave 1, snapshot these files:
|
Before Wave 1, snapshot these files:
|
||||||
```
|
```
|
||||||
/opt/inference-harness/docker-compose.yml (CT 116 .116)
|
/opt/inference-harness/docker-compose.yml (CT 116 .116) ⚡ contains MCP_SERVER env vars
|
||||||
/opt/inference-harness/litellm_config.yaml (CT 116 .116)
|
/opt/inference-harness/litellm_config.yaml (CT 116 .116) ⚡ contains mcp_servers.ra_h_os
|
||||||
/opt/monitoring/prometheus.yml (CT 116 .116)
|
/opt/monitoring/prometheus.yml (CT 116 .116)
|
||||||
/etc/nginx/nginx.conf (harness-nginx on CT 116)
|
/etc/nginx/nginx.conf (harness-nginx on CT 116)
|
||||||
/opt/search-stack/firecrawl-source/docker-compose.yaml (VM 109 .7)
|
/opt/search-stack/firecrawl-source/docker-compose.yaml (VM 109 .7)
|
||||||
@@ -123,9 +126,51 @@ Before Wave 1, snapshot these files:
|
|||||||
/root/.pi/agent/extensions/config.yaml (CT 100 .24)
|
/root/.pi/agent/extensions/config.yaml (CT 100 .24)
|
||||||
/etc/systemd/system/ornith-server.service (amdpve .15)
|
/etc/systemd/system/ornith-server.service (amdpve .15)
|
||||||
/etc/systemd/system/llama-server.service (VM 101 .8, VM 103 .110)
|
/etc/systemd/system/llama-server.service (VM 101 .8, VM 103 .110)
|
||||||
|
# Hermes agent configs (key enforcement — 2026-07-10)
|
||||||
|
/root/.hermes/config.yaml (Mumuni CT 114, Tanko CT 112, etc.)
|
||||||
|
/root/.config/systemd/user/hermes-gateway.service (Mumuni CT 114 — EnvironmentFile fixed)
|
||||||
|
/etc/environment (Mumuni CT 114 — LITELLM_API_KEY)
|
||||||
```
|
```
|
||||||
|
|
||||||
Run: `mkdir -p /tmp/infra-update-backup-$(date +%Y%m%d) && rsync -av ...`
|
## MCP Gateway (2026-07-10)
|
||||||
|
|
||||||
|
LiteLLM CT 116 now serves as an authenticated MCP gateway for RA-H OS tools.
|
||||||
|
|
||||||
|
### Configuration
|
||||||
|
|
||||||
|
**litellm_config.yaml** (`/opt/inference-harness/litellm_config.yaml`):
|
||||||
|
```yaml
|
||||||
|
mcp_servers:
|
||||||
|
ra_h_os:
|
||||||
|
url: "http://192.168.68.65:3100/mcp"
|
||||||
|
transport: "http"
|
||||||
|
auth_type: "none"
|
||||||
|
```
|
||||||
|
|
||||||
|
**docker-compose.yml** env vars:
|
||||||
|
```yaml
|
||||||
|
- MCP_SERVER_RAHOS_URL=http://192.168.68.65:3100/mcp
|
||||||
|
- MCP_SERVER_RAHOS_TRANSPORT=http
|
||||||
|
```
|
||||||
|
|
||||||
|
### Access
|
||||||
|
|
||||||
|
| Key | MCP Access |
|
||||||
|
|-----|-----------|
|
||||||
|
| Master key | ✅ Full — 90 tools (vault-injected) |
|
||||||
|
| Agent keys (mumuni, tanko, etc.) | ❌ Per-key grants not supported in v1.90.0-rc.1 |
|
||||||
|
|
||||||
|
### Known Limitations
|
||||||
|
- Per-key MCP server grants not functional — only master key has access
|
||||||
|
- Responses API (`/v1/responses`) with MCP tools broken on llama.cpp backends
|
||||||
|
- HTTP 307 redirect on `/mcp` → use `/mcp/` (trailing slash) or `/mcp-rest/` endpoints
|
||||||
|
- `api_mode: responses` in Hermes appends `/v1/responses` to base_url → **base_url must end at `/v1`, never `/responses`** (double-path bug)
|
||||||
|
|
||||||
|
### Migration Path
|
||||||
|
When LiteLLM is upgraded to a version supporting per-key MCP grants:
|
||||||
|
1. Grant agent keys `mcp_servers: ["ra_h_os"]`
|
||||||
|
2. Update Hermes `mcp_servers.ra-h-os.url` from `http://192.168.68.65:3100/mcp` → `http://192.168.68.116:4000/mcp/`
|
||||||
|
3. Add `headers: {x-litellm-api-key: "Bearer $LITELLM_API_KEY"}` to MCP config
|
||||||
|
|
||||||
## Security-Specific Updates
|
## Security-Specific Updates
|
||||||
|
|
||||||
@@ -144,6 +189,7 @@ Run: `mkdir -p /tmp/infra-update-backup-$(date +%Y%m%d) && rsync -av ...`
|
|||||||
- [ ] LiteLLM inference passing (syslog-auto test)
|
- [ ] LiteLLM inference passing (syslog-auto test)
|
||||||
- [ ] Zulip server + all 3 agents connected
|
- [ ] Zulip server + all 3 agents connected
|
||||||
- [ ] GPU fleet at full capacity (3/3)
|
- [ ] GPU fleet at full capacity (3/3)
|
||||||
|
- [ ] LiteLLM MCP gateway healthy (90 tools via master key)
|
||||||
- [ ] Zero security CVEs remaining
|
- [ ] Zero security CVEs remaining
|
||||||
- [ ] <10 min total downtime per service
|
- [ ] <10 min total downtime per service
|
||||||
|
|
||||||
|
|||||||
@@ -9,9 +9,14 @@ description: >
|
|||||||
not calendar-driven — rotate only on compromise, personnel change, or
|
not calendar-driven — rotate only on compromise, personnel change, or
|
||||||
periodic security hygiene (quarterly/annually).
|
periodic security hygiene (quarterly/annually).
|
||||||
|
|
||||||
|
UPDATED 2026-07-12: Keys are stored in Infisical vault (project=agents, env=production)
|
||||||
|
BUT each agent host MUST keep a local .env fallback. Infisical service tokens can
|
||||||
|
expire/404. The .env fallback prevents agents from running without keys.
|
||||||
|
Tanko incident: token 404 → gateway had no LITELLM_API_KEY for hours.
|
||||||
|
|
||||||
Current key inventory and agent list: see gpu-fleet.prose.md § Agent Keys.
|
Current key inventory and agent list: see gpu-fleet.prose.md § Agent Keys.
|
||||||
Source of truth for LiteLLM config: /opt/inference-harness/litellm_config.yaml
|
Source of truth for LiteLLM config: /opt/inference-harness/litellm_config.yaml
|
||||||
on CT 116. Last verified: 2026-07-09.
|
on CT 116. Last verified: 2026-07-12.
|
||||||
---
|
---
|
||||||
|
|
||||||
## Parameters
|
## Parameters
|
||||||
@@ -19,7 +24,10 @@ description: >
|
|||||||
- agent_name: string — The agent to manage keys for (e.g., "tanko", "mumuni")
|
- agent_name: string — The agent to manage keys for (e.g., "tanko", "mumuni")
|
||||||
- action: "create" | "rotate" | "verify" | "list" — What to do (default: "create")
|
- action: "create" | "rotate" | "verify" | "list" — What to do (default: "create")
|
||||||
- litellm_host: string — LiteLLM admin endpoint (default: "192.168.68.116:4000")
|
- litellm_host: string — LiteLLM admin endpoint (default: "192.168.68.116:4000")
|
||||||
- master_key: string — LiteLLM master key (default from environment)
|
- master_key: string — LiteLLM master key (default from Infisical vault: project=infrastructure, env=production, secret=LITELLM_MASTER_KEY)
|
||||||
|
- vault_url: string — Infisical vault URL (default: "https://vault.sysloggh.net")
|
||||||
|
- vault_project: string — Infisical project slug (default: "infrastructure")
|
||||||
|
- vault_env: string — Infisical environment (default: "production")
|
||||||
- agent_host: string — Agent's IP for SSH (default: resolved from infra)
|
- agent_host: string — Agent's IP for SSH (default: resolved from infra)
|
||||||
- agent_user: string — SSH user (default: "jerome")
|
- agent_user: string — SSH user (default: "jerome")
|
||||||
|
|
||||||
@@ -30,12 +38,13 @@ description: >
|
|||||||
- key_prefix: string — First 10 chars of the new key (for identification)
|
- key_prefix: string — First 10 chars of the new key (for identification)
|
||||||
- previous_key_alias: string | null — Previous key alias if rotating
|
- previous_key_alias: string | null — Previous key alias if rotating
|
||||||
- litellm_response: object — Raw response from LiteLLM /key/generate
|
- litellm_response: object — Raw response from LiteLLM /key/generate
|
||||||
- agent_config_updated: boolean — Whether /etc/environment was updated
|
- vault_updated: boolean — Whether Infisical vault secret was updated
|
||||||
|
- agent_config_updated: boolean — Legacy: whether /etc/environment was updated (deprecated, always false post-migration)
|
||||||
- verification: { status: string, detail: string } — Final health check
|
- verification: { status: string, detail: string } — Final health check
|
||||||
|
|
||||||
## Execution
|
## Execution
|
||||||
|
|
||||||
1. **Authenticate** — Verify master_key works against LiteLLM /key/list
|
1. **Authenticate** — Retrieve master key from Infisical vault via `infisical export --project=<vault_project> --env=<vault_env>`, verify against LiteLLM /key/list
|
||||||
2. **Check existing keys** — List all keys, find any with agent_name alias
|
2. **Check existing keys** — List all keys, find any with agent_name alias
|
||||||
3. **If action == "list"**: Return all keys with their aliases and spend
|
3. **If action == "list"**: Return all keys with their aliases and spend
|
||||||
4. **If action == "create"**:
|
4. **If action == "create"**:
|
||||||
@@ -47,11 +56,14 @@ description: >
|
|||||||
- Return the new key
|
- Return the new key
|
||||||
5. **If action == "rotate"**:
|
5. **If action == "rotate"**:
|
||||||
- Generate new key with same alias (LiteLLM replaces the old key)
|
- Generate new key with same alias (LiteLLM replaces the old key)
|
||||||
- SSH to agent_host, update /etc/environment LITELLM_API_KEY
|
- Update secret in Infisical vault: `infisical secrets set LITELLM_API_KEY=<new_key> --project=<vault_project> --env=<vault_env>`
|
||||||
- Restart agent gateway (hermes gateway restart for Hermes agents)
|
- Restart agent gateway (Hermes: `systemctl restart hermes-gateway`; pi: restart PM2 process)
|
||||||
|
The gateway automatically picks up the new key via `infisical run --` wrapper
|
||||||
- Verify: curl test against /v1/models with new key
|
- Verify: curl test against /v1/models with new key
|
||||||
- Rotation policy: on-demand only (compromise, departure, quarterly hygiene)
|
- Rotation policy: on-demand only (compromise, departure, quarterly hygiene)
|
||||||
|
- Note: /etc/environment is NO LONGER used for LiteLLM keys. Agents inject keys at runtime via vault wrapper.
|
||||||
6. **If action == "verify"**:
|
6. **If action == "verify"**:
|
||||||
- SSH to agent, read /etc/environment
|
- Retrieve key from Infisical vault: `infisical secrets get LITELLM_API_KEY --project=<vault_project> --env=<vault_env>`
|
||||||
- Test the key against LiteLLM /v1/models
|
- Test the key against LiteLLM /v1/models
|
||||||
- Confirm key alias matches agent_name in LiteLLM key list
|
- Confirm key alias matches agent_name in LiteLLM key list
|
||||||
|
- Verify agent gateway uses vault wrapper: `cat /proc/<pid>/cmdline` shows `infisical run`
|
||||||
|
|||||||
@@ -1,18 +1,20 @@
|
|||||||
---
|
---
|
||||||
kind: responsibility
|
kind: responsibility
|
||||||
name: litellm-self-heal
|
name: litellm-self-heal
|
||||||
status: manual-only
|
status: deployed
|
||||||
note: >
|
note: >
|
||||||
Auto-remediation code was removed from the pi Zulip extension (retired 2026-07-04).
|
DEPLOYED 2026-07-12 on CT 116 cron: 0 */6 * * *
|
||||||
This contract is now manual-only — triggers require explicit user request.
|
Auto-remediation code was removed from the pi Zulip extension (retired 2026-07-04),
|
||||||
Consider reimplementing as a standalone cron job or prose contract.
|
now reimplemented as `litellm-health-check.sh` on CT 116.
|
||||||
|
Reports to /var/log/litellm/health-*.json and RA-H OS knowledge graph.
|
||||||
|
GPU monitoring integrated from gpu-monitor on .24:9100.
|
||||||
|
|
||||||
Consolidated from litellm-health + litellm-self-heal on 2026-07-09 to eliminate
|
Consolidated from litellm-health + litellm-self-heal on 2026-07-09 to eliminate
|
||||||
duplication of architecture diagrams, GPU topology, timeout tables, and container
|
duplication of architecture diagrams, GPU topology, timeout tables, and container
|
||||||
lists. Health check is now § Health Check within this contract.
|
lists. Health check is now § Health Check within this contract.
|
||||||
|
|
||||||
Source of truth for GPU topology and keys: gpu-fleet.prose.md
|
Source of truth for GPU topology and keys: gpu-fleet.prose.md
|
||||||
Last verified: 2026-07-09
|
Last verified: 2026-07-12
|
||||||
description: >
|
description: >
|
||||||
LiteLLM inference stack health monitoring + self-healing. Verifies the full
|
LiteLLM inference stack health monitoring + self-healing. Verifies the full
|
||||||
nginx → LiteLLM → GPU chain, 8 containers on CT 116, 3 GPU hosts, model
|
nginx → LiteLLM → GPU chain, 8 containers on CT 116, 3 GPU hosts, model
|
||||||
@@ -54,8 +56,8 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
|||||||
|
|
||||||
| Host | IP | Hardware | Models Served | Engine | Context | Parallel |
|
| Host | IP | Hardware | Models Served | Engine | Context | Parallel |
|
||||||
|------|-----|----------|---------------|--------|---------|----------|
|
|------|-----|----------|---------------|--------|---------|----------|
|
||||||
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd | 128K | 2 |
|
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd | **256K** | 1 |
|
||||||
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd | 128K | 2 |
|
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd | 131K | 2 |
|
||||||
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | ornith-1.0-35b | llama-server systemd (Vulkan) | 256K | 2 |
|
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | ornith-1.0-35b | llama-server systemd (Vulkan) | 256K | 2 |
|
||||||
|
|
||||||
## Model Fallback Chains (LiteLLM)
|
## Model Fallback Chains (LiteLLM)
|
||||||
|
|||||||
@@ -1,7 +1,7 @@
|
|||||||
---
|
---
|
||||||
name: memory-audit-maintenance
|
name: memory-audit-maintenance
|
||||||
kind: responsibility
|
kind: responsibility
|
||||||
description: Shared memory audit and maintenance contract for all Hermes agents (Mumuni, Tanko, Tdunna, Baggy). Each agent runs it against its own isolated memory files — no cross-agent access, no shared state. Detects staleness, enforces writer registry, and rotates canary tokens.
|
description: Shared memory audit and maintenance contract for all Hermes agents (Mumuni, Tanko, Koby, Koonimo). Each agent runs it against its own isolated memory files — no cross-agent access, no shared state. Detects staleness, enforces writer registry, and rotates canary tokens.
|
||||||
id: 067NC4KG01RG50R40M30E20918
|
id: 067NC4KG01RG50R40M30E20918
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -15,13 +15,13 @@ Autonomously audit and reorganize an agent's native memory (MEMORY.md, USER.md,
|
|||||||
|
|
||||||
### Scope
|
### Scope
|
||||||
|
|
||||||
This contract is the **Hermes Agent standard** for memory maintenance. It is shared across all Hermes agents (Mumuni, Tanko, Tdunna, Baggy). Each agent runs it against its own memory files only no cross-agent access, no shared state, no shared ledger, no shared canary. The contract is the standard; each agent enforces it independently with fully isolated data.
|
This contract is the **Hermes Agent standard** for memory maintenance. It is shared across all Hermes agents (Mumuni, Tanko, Koby, Koonimo). Each agent runs it against its own memory files only no cross-agent access, no shared state, no shared ledger, no shared canary. The contract is the standard; each agent enforces it independently with fully isolated data.
|
||||||
|
|
||||||
**Agent Roster:**
|
**Agent Roster:**
|
||||||
- Mumuni
|
- Mumuni
|
||||||
- Tanko
|
- Tanko
|
||||||
- Tdunna
|
- Koby (CT 111 / tdunna)
|
||||||
- Baggy
|
- Koonimo (CT 113 / baggy)
|
||||||
|
|
||||||
**Isolation Principle:** Each agent has its own:
|
**Isolation Principle:** Each agent has its own:
|
||||||
- `MEMORY.md` and `USER.md`
|
- `MEMORY.md` and `USER.md`
|
||||||
|
|||||||
@@ -1,290 +0,0 @@
|
|||||||
---
|
|
||||||
kind: pattern
|
|
||||||
name: mumuni-delegation
|
|
||||||
description: >
|
|
||||||
Mumuni-specific operating doctrine for task decomposition, worker
|
|
||||||
delegation, verification, and delivery. Defines when to delegate, which
|
|
||||||
worker to use for what, how to handle failures, and the kanban board
|
|
||||||
protocol. Enforces context-window discipline and separation of concerns.
|
|
||||||
Runs on Mumuni (CT 118, storepve, .6) via Hermes agent.
|
|
||||||
version: 1.0.0
|
|
||||||
---
|
|
||||||
|
|
||||||
## Maintains
|
|
||||||
|
|
||||||
- Worker roster: 6 profiles (`syslog-code`, `syslog-devops`, `syslog-email`,
|
|
||||||
`syslog-research`, `syslog-review`, `syslog-writer`)
|
|
||||||
- Kanban board state at `~/.hermes/kanban/kanban.json`
|
|
||||||
- Context window budget: ~65K tokens per request (131K total, 60% threshold)
|
|
||||||
|
|
||||||
## Topology
|
|
||||||
|
|
||||||
**Cluster:** 5 Proxmox nodes (ocupve, acerpve, minipve, amdpve, storepve)
|
|
||||||
**Manager:** Mumuni (CT 118, storepve, .6) via Hermes agent
|
|
||||||
**Workers:** 6 profiles, all running on the same agent — no separate hosts needed
|
|
||||||
|
|
||||||
This contract is infrastructure-agnostic in terms of which nodes are used.
|
|
||||||
Workers execute tasks on whatever infrastructure they're given — SSH to .6,
|
|
||||||
.pm, .9, .12, or .15 depending on the task. The contract defines the
|
|
||||||
**who** and **when** — not the **where**.
|
|
||||||
|
|
||||||
## Why This Matters
|
|
||||||
|
|
||||||
Without enforced delegation, the manager consumes the full iteration budget
|
|
||||||
(60 calls) on single-turn tasks — SSH to 5 nodes, check each VM, read logs —
|
|
||||||
leaving no capacity for actual coordination. The result: context overflow
|
|
||||||
(59K tokens in system prompt), iteration exhaustion, and degraded response
|
|
||||||
quality. This contract exists because I blew through my budget checking
|
|
||||||
Proxmox node status instead of delegating to `syslog-devops`.
|
|
||||||
|
|
||||||
## Context Window Discipline
|
|
||||||
|
|
||||||
**The system prompt is ~6.5K tokens (stable: ~4.5K tool schemas + ~2K other guidance).**
|
|
||||||
**Volatile (MEMORY.md + USER.md): ~300 tokens.**
|
|
||||||
**Total base: ~6,800 tokens per request.**
|
|
||||||
|
|
||||||
The remaining budget is the conversation. Every tool call result adds to it.
|
|
||||||
If a single call returns >10K tokens (e.g., `grep` on a large file, SSH output
|
|
||||||
from multiple nodes), the context fills fast. That's why we delegate: workers
|
|
||||||
process in isolation and return compact results.
|
|
||||||
|
|
||||||
## Trigger Conditions
|
|
||||||
|
|
||||||
Delegation is **mandatory** when any of these apply:
|
|
||||||
|
|
||||||
| Condition | Threshold | Example |
|
|
||||||
|-----------|-----------|---------|
|
|
||||||
| Multiple tool calls needed | 2+ calls with intermediate logic | Read file → analyze → write report |
|
|
||||||
| Large data retrieval | Output >5K tokens | `grep -r "pattern" /path` on large dirs |
|
|
||||||
| Cross-domain work | Spans 2+ worker specialties | Infra check + email filter |
|
|
||||||
| Infrastructure changes | Any mutating operation | `qm set`, `systemctl restart`, `git push` |
|
|
||||||
| Research/analysis | Needs browser or deep reading | Web research, code review, data analysis |
|
|
||||||
| Code builds or changes | Writing or modifying code | Scripts, configs, patches |
|
|
||||||
| Sequential dependencies | Worker B needs Worker A's output | Code → Review → Deliver |
|
|
||||||
|
|
||||||
**Single tool calls stay at manager level.** Quick `grep`, `ls`, `cat`,
|
|
||||||
`curl`, `hermes tools list` — these are decision-making tools. The manager
|
|
||||||
reads them directly.
|
|
||||||
|
|
||||||
## Data Source Integrity (CRITICAL)
|
|
||||||
|
|
||||||
**Workers MUST use the data provided in their task context. They MUST NOT
|
|
||||||
fetch their own data from external sources unless explicitly told to.**
|
|
||||||
|
|
||||||
When a task says "Read file X and format it", the worker reads file X. It does
|
|
||||||
not query a separate API, run its own diagnostics, or pull data from a different
|
|
||||||
system. This is the #1 source of cross-worker inconsistency: one worker gathers
|
|
||||||
SSH data, another queries the Proxmox API, and the report merges two incompatible
|
|
||||||
datasets.
|
|
||||||
|
|
||||||
**Rule:** If a worker needs additional data beyond what's in its task description,
|
|
||||||
it asks the manager (via relay) — it doesn't go find it on its own.
|
|
||||||
|
|
||||||
**This is a hard rule, not a recommendation.** Violating it produces the exact
|
|
||||||
type of discrepancy the kanban pipeline exists to prevent: a review worker finds
|
|
||||||
"5 nodes present" in the raw data but "5/5 online" in the report — even though
|
|
||||||
one of those nodes was unreachable. The report lied because it used data the
|
|
||||||
raw data never provided.
|
|
||||||
|
|
||||||
## Worker Selection Matrix
|
|
||||||
|
|
||||||
| Worker | Model | Toolsets | Role | Use When |
|
|
||||||
|--------|-------|----------|------|----------|
|
|
||||||
| `syslog-code` | qwen3.6-27B-code | terminal, file, web, memory, skills | Code patches, automation, scripts | Writing/modifying code, creating scripts, debugging, reading/writing files |
|
|
||||||
| `syslog-devops` | qwen3.6-27B-code | terminal, file, web, memory, skills | Infrastructure, DB, bridge, Proxmox | Server ops, SSH, Docker, Proxmox, DB queries, hardware checks |
|
|
||||||
| `syslog-email` | ornith-1.0-35b | terminal, file, web, memory, skills | Email automation, mail operations | Sending/receiving email, inbox management, SMTP operations |
|
|
||||||
| `syslog-research` | ornith-1.0-35b | terminal, file, web, memory, skills, **browser** | Analysis, classification, data processing | Web research, browser tasks, data analysis, classification, reading docs |
|
|
||||||
| `syslog-review` | ornith-1.0-35b | terminal, file, web, memory, skills | Verification, QA, audit validation | **ALWAYS** verify worker output before delivery — especially for infra changes, code builds, and research findings |
|
|
||||||
| `syslog-writer` | ornith-1.0-35b | terminal, file, web, memory, skills | Docs, content, branding, reports | Writing docs, reports, proposals, content, markdown formatting |
|
|
||||||
|
|
||||||
### Selection Rules
|
|
||||||
|
|
||||||
1. **Match specialty first.** A code task → `syslog-code`. An infra task →
|
|
||||||
`syslog-devops`. Don't put a `syslog-email` worker on a code review.
|
|
||||||
2. **Research tasks with browser needs → `syslog-research`.** Other workers
|
|
||||||
don't have the browser toolset.
|
|
||||||
3. **Verification → `syslog-review`.** Never deliver raw worker output.
|
|
||||||
4. **Documentation/content → `syslog-writer`.** Let them own the prose.
|
|
||||||
5. **If unsure, delegate to `syslog-research`** — it has the broadest toolset
|
|
||||||
(includes browser) and high reasoning effort.
|
|
||||||
|
|
||||||
## Delegation Protocol
|
|
||||||
|
|
||||||
### Step 1: Decompose
|
|
||||||
|
|
||||||
Break the task into lanes. Each lane does ONE thing. Workers are independent —
|
|
||||||
no lane depends on another's output mid-flight. If lanes depend on each other,
|
|
||||||
dispatch sequentially.
|
|
||||||
|
|
||||||
### Step 2: Dispatch
|
|
||||||
|
|
||||||
Fire workers via `delegate_task`:
|
|
||||||
|
|
||||||
**Critical: Pass the data, not just the goal.** When dispatching a worker that
|
|
||||||
processes output from another worker, include the file path AND explicit
|
|
||||||
instructions to use ONLY that source. Example:
|
|
||||||
|
|
||||||
```
|
|
||||||
delegate_task(
|
|
||||||
goal="Format the cluster check into a clean report",
|
|
||||||
context="Source data is at /tmp/proxmox-check-raw.md. Format ONLY the data
|
|
||||||
in that file. Do NOT query the Proxmox API or any other data source. Use the
|
|
||||||
file as your sole source of truth."
|
|
||||||
)
|
|
||||||
```
|
|
||||||
|
|
||||||
**Parallel (independent lanes):**
|
|
||||||
```
|
|
||||||
delegate_task(
|
|
||||||
tasks=[
|
|
||||||
{"goal": "Check all 5 Proxmox nodes for VM status", "context": "SSH to each node via 192.168.68.x, run 'qm list'"},
|
|
||||||
{"goal": "Check Docker container health on .7/.116/.17", "context": "SSH to each host, check container status"},
|
|
||||||
]
|
|
||||||
)
|
|
||||||
```
|
|
||||||
|
|
||||||
**Sequential (dependent lanes):**
|
|
||||||
Dispatch lane 1 → wait for result → dispatch lane 2.
|
|
||||||
|
|
||||||
### Step 3: Verify
|
|
||||||
|
|
||||||
**MANDATORY for:**
|
|
||||||
- Infrastructure changes (any `qm`, `pct`, `systemctl`, `git push`)
|
|
||||||
- Code builds and modifications
|
|
||||||
- Research findings (web data, external sources)
|
|
||||||
- Any output that will reach the user
|
|
||||||
|
|
||||||
**Fire `syslog-review` to verify:**
|
|
||||||
```
|
|
||||||
delegate_task(
|
|
||||||
goal="Review the output of the devops worker. Verify the node status
|
|
||||||
report is accurate, check for inconsistencies, confirm all nodes were
|
|
||||||
reachable.",
|
|
||||||
context="Worker was syslog-devops. Output is at /tmp/node-report.md.
|
|
||||||
Verify against live system."
|
|
||||||
)
|
|
||||||
```
|
|
||||||
|
|
||||||
**If verification fails:**
|
|
||||||
1. Send work back to original worker with review feedback
|
|
||||||
2. Re-verify
|
|
||||||
3. Max 2 re-verify cycles before escalating to Kwame
|
|
||||||
|
|
||||||
### Step 4: Deliver
|
|
||||||
|
|
||||||
Only verified results reach Kwame. Format per channel:
|
|
||||||
- Telegram: Use `telegram-formatting` skill
|
|
||||||
- Zulip: Use Zulip Markdown (CommonMark)
|
|
||||||
- Email: Use `syslog-email` skill
|
|
||||||
|
|
||||||
## Kanban Board Protocol
|
|
||||||
|
|
||||||
**File:** `~/.hermes/kanban/kanban.json`
|
|
||||||
|
|
||||||
```json
|
|
||||||
{
|
|
||||||
"task_id": "unique-id",
|
|
||||||
"title": "Task description",
|
|
||||||
"created": "2026-07-09T01:00:00",
|
|
||||||
"status": "backlog|in_progress|review|done",
|
|
||||||
"lanes": [
|
|
||||||
{
|
|
||||||
"lane_id": "devops-check",
|
|
||||||
"worker": "syslog-devops",
|
|
||||||
"goal": "Check all 5 Proxmox nodes",
|
|
||||||
"status": "dispatched|completed|failed",
|
|
||||||
"output_file": "/tmp/node-report.md"
|
|
||||||
}
|
|
||||||
]
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
**Update the board on every state change.**
|
|
||||||
|
|
||||||
## Failure Handling
|
|
||||||
|
|
||||||
### Worker Timeouts
|
|
||||||
|
|
||||||
- Child timeout: **900 seconds** (15 minutes)
|
|
||||||
- Worker model `syslog-auto` is slow — it can hit the timeout limit with
|
|
||||||
22+ API calls
|
|
||||||
- **If a worker times out:** Re-dispatch with a narrower scope. Break the
|
|
||||||
task into smaller pieces that fit in the timeout window.
|
|
||||||
- **Avoid delegating sequential SSH hops** — each SSH connection adds latency
|
|
||||||
that compounds quickly. Prefer API-based or local approaches when possible.
|
|
||||||
|
|
||||||
### Worker Selection Failures
|
|
||||||
|
|
||||||
- `syslog-devops` is best for infrastructure tasks (SSH, Proxmox, Docker)
|
|
||||||
- `syslog-code` is best for code-level work (reading files, writing scripts)
|
|
||||||
- `syslog-research` has the browser toolset — use for web research
|
|
||||||
- `syslog-review` is the QA gate — always fire before delivery
|
|
||||||
- **Never fire more than 3 parallel workers** (max_concurrent_children: 3)
|
|
||||||
- **Never nest delegation** (max_spawn_depth: 1)
|
|
||||||
|
|
||||||
### Context Overflow
|
|
||||||
|
|
||||||
- If a task requires >10K tokens of output, delegate the processing
|
|
||||||
- Workers return compact summaries, not raw data dumps
|
|
||||||
- Pass file paths and concrete goals — never dump raw data into context
|
|
||||||
|
|
||||||
## Anti-patterns
|
|
||||||
|
|
||||||
- ❌ Reading large files into your own context before deciding → delegate the read
|
|
||||||
- ❌ Carrying SSH/grep/output results in your context → delegate the analysis
|
|
||||||
- ❌ Doing work yourself and then "pretending" to delegate → the user can tell
|
|
||||||
- ❌ Skipping verification → raw worker output never reaches the user
|
|
||||||
- ❌ Delegating single tool calls → keep quick reads/writes at manager level
|
|
||||||
- ❌ Firing more than 3 workers in parallel → hard limit
|
|
||||||
- ❌ **Workers fetching their own data sources** → a writer worker that queries the Proxmox API when told to "format the raw file" is fabricating data. Use the input given, not external sources
|
|
||||||
|
|
||||||
## Emergency Exception
|
|
||||||
|
|
||||||
**In an emergency (server down, service must be restored immediately):**
|
|
||||||
- Delegate the diagnosis (find the problem)
|
|
||||||
- Execute the fix yourself (minimize handoff latency)
|
|
||||||
- Verify the fix after delivery
|
|
||||||
- Log the exception in the kanban board
|
|
||||||
|
|
||||||
The emergency exception exists because the user needs the service back NOW,
|
|
||||||
not after three worker round-trips. But it's an exception — not the rule.
|
|
||||||
|
|
||||||
## What This Contract Doesn't Cover
|
|
||||||
|
|
||||||
1. **Worker profile configuration** — covered by `hermes-config-template.prose.md`
|
|
||||||
2. **SSH key management** — covered by existing SSH/Proxmox contracts
|
|
||||||
3. **Git workflow** — covered by `AGENTS.md` in the prose-contracts repo
|
|
||||||
4. **Cron job management** — covered by individual cron contracts
|
|
||||||
5. **Infra verification** — covered by `verify-before-mutate` protocol
|
|
||||||
|
|
||||||
## Verification
|
|
||||||
|
|
||||||
Run `scripts/worker-audit.py` to verify all 6 profiles are aligned:
|
|
||||||
|
|
||||||
```bash
|
|
||||||
python3 /root/.hermes/skills/kanban-orchestrator/scripts/worker-audit.py
|
|
||||||
```
|
|
||||||
|
|
||||||
## References
|
|
||||||
|
|
||||||
- `kanban-orchestrator` skill: The operational playbook (detailed execution steps)
|
|
||||||
- `worker-profile-audit.md` (skill reference): Worker configuration audit notes
|
|
||||||
- `delegation-timeout-patterns.md` (skill reference): Timeout handling patterns
|
|
||||||
- `verify-before-mutate` protocol: Infrastructure change verification
|
|
||||||
- `hermes-config-template.prose.md`: Worker profile configuration
|
|
||||||
|
|
||||||
## Success Criteria
|
|
||||||
|
|
||||||
This contract succeeds when:
|
|
||||||
|
|
||||||
1. **No context overflow** — single-turn tasks don't exhaust the iteration budget
|
|
||||||
2. **Workers do the work** — manager coordinates, doesn't execute
|
|
||||||
3. **Verification before delivery** — all output passes through `syslog-review`
|
|
||||||
4. **Kanban board is current** — every task has a lane, every lane has a status
|
|
||||||
5. **User gets verified results** — raw worker output never reaches Kwame
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
**Last updated:** 2026-07-09
|
|
||||||
**Author:** Mumuni (with Kwame's input on triggers and exception criteria)
|
|
||||||
**Status:** Draft — awaiting PR review and merge to prose-contracts main
|
|
||||||
@@ -41,7 +41,7 @@ agent: abiba
|
|||||||
|
|
||||||
- User: `monitoring@pve` (cluster-replicated)
|
- User: `monitoring@pve` (cluster-replicated)
|
||||||
- Role: `PVEAuditor` on `/` (read-only, whole cluster)
|
- Role: `PVEAuditor` on `/` (read-only, whole cluster)
|
||||||
- Token: `monitoring@pve!prometheus` = `2c74ceb6-f905-444a-94f9-1c4f7889b68c`
|
- Token: `monitoring@pve!prometheus` — stored in Infisical vault (`PROXMOX_MONITOR_TOKEN`)
|
||||||
- `verify_ssl: false` (proxmoxer uses `verify_ssl`, NOT `verify_tls`)
|
- `verify_ssl: false` (proxmoxer uses `verify_ssl`, NOT `verify_tls`)
|
||||||
|
|
||||||
## Grafana Dashboards (file-provisioned, folder "Syslog Fleet")
|
## Grafana Dashboards (file-provisioned, folder "Syslog Fleet")
|
||||||
@@ -62,7 +62,7 @@ agent: abiba
|
|||||||
|
|
||||||
- **URL**: `http://192.168.68.116:3001/` (LAN, direct — Grafana bound to `0.0.0.0:3001`)
|
- **URL**: `http://192.168.68.116:3001/` (LAN, direct — Grafana bound to `0.0.0.0:3001`)
|
||||||
- **Dashboards**: `http://192.168.68.116:3001/d/gpu-fleet`, `.../d/proxmox-cluster`, `.../d/proxmox-node`, `.../d/docker-containers`
|
- **Dashboards**: `http://192.168.68.116:3001/d/gpu-fleet`, `.../d/proxmox-cluster`, `.../d/proxmox-node`, `.../d/docker-containers`
|
||||||
- **Credentials**: admin / syslog-grafana-2026
|
- **Credentials**: admin / password stored in Infisical vault (`GRAFANA_ADMIN_PASSWORD`)
|
||||||
- Grafana is NOT behind nginx — access port 3001 directly. The `harness-nginx` `/grafana/` sub-path route was tried and reverted (broke the existing `:3001` URL and gpu-fleet path). Do not re-add `GF_SERVER_SERVE_FROM_SUB_PATH` or an nginx `/grafana/` route.
|
- Grafana is NOT behind nginx — access port 3001 directly. The `harness-nginx` `/grafana/` sub-path route was tried and reverted (broke the existing `:3001` URL and gpu-fleet path). Do not re-add `GF_SERVER_SERVE_FROM_SUB_PATH` or an nginx `/grafana/` route.
|
||||||
- grafana compose port mapping: `"3001:3000"` (0.0.0.0, not 127.0.0.1)
|
- grafana compose port mapping: `"3001:3000"` (0.0.0.0, not 127.0.0.1)
|
||||||
|
|
||||||
|
|||||||
@@ -1,210 +0,0 @@
|
|||||||
---
|
|
||||||
kind: pattern
|
|
||||||
name: ra-h-os-custodianship-contract
|
|
||||||
description: >
|
|
||||||
Operational standards for maintaining RA-H OS as a high-functioning shared
|
|
||||||
memory system across all agents. Defines rules for embedding consistency,
|
|
||||||
namespace discipline, staleness management, orphan prevention, agent
|
|
||||||
custodianship, and recovery protocols. Prevents knowledge graph degradation
|
|
||||||
and ensures reliable semantic search capabilities.
|
|
||||||
---
|
|
||||||
|
|
||||||
# RA-H OS Custodianship Contract — Shared Memory Protocol
|
|
||||||
|
|
||||||
## Purpose
|
|
||||||
This contract establishes the operational standards for maintaining RA-H OS as a high-functioning shared memory system across all agents. It defines the rules, responsibilities, and recovery protocols that prevent knowledge graph degradation and ensure consistent, reliable semantic search capabilities.
|
|
||||||
|
|
||||||
## Core Principles
|
|
||||||
|
|
||||||
### 1. Embedding Consistency
|
|
||||||
**Rule:** All nodes MUST use the RA-H OS built-in embedding pipeline via `http://192.168.68.65:8080/v1/embeddings` for consistent vector representation.
|
|
||||||
|
|
||||||
**Requirements:**
|
|
||||||
- The embedding service at `192.168.68.65:8080` is the single source of truth for vector generation
|
|
||||||
- No external embedding APIs (OpenAI, Cohere, etc.) may be used for RA-H OS nodes
|
|
||||||
- All nodes must have `embedding_status: "chunked"` and `chunk_status: "chunked"` upon creation
|
|
||||||
- If embedding fails, the node must be immediately flagged as `state: "unsearchable"` and logged
|
|
||||||
|
|
||||||
### 2. Namespace Discipline
|
|
||||||
**Rule:** All nodes MUST be categorized into one of three namespaces with strict metadata requirements.
|
|
||||||
|
|
||||||
**Namespace Structure:**
|
|
||||||
- `agent-private`: Isolated working notes for individual agents (Mumuni, Tanko, Okyeame, etc.)
|
|
||||||
- `shared`: Policy files, registry nodes, collective knowledge
|
|
||||||
- `syslogsolution`: Business nodes, client data, operational context
|
|
||||||
|
|
||||||
**Metadata Requirements:**
|
|
||||||
Every node MUST have these metadata fields:
|
|
||||||
```json
|
|
||||||
{
|
|
||||||
"agent_id": "<agent_name>",
|
|
||||||
"namespace": "<namespace>",
|
|
||||||
"state": "<state>",
|
|
||||||
"type": "<type>",
|
|
||||||
"owner": "<owner>",
|
|
||||||
"tenant": "<tenant>",
|
|
||||||
"visibility": "<visibility>",
|
|
||||||
"source": "<source>",
|
|
||||||
"captured_by": "<captured_by>"
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### 3. Staleness Management
|
|
||||||
**Rule:** Nodes are categorized by type and have specific staleness thresholds.
|
|
||||||
|
|
||||||
**Staleness Thresholds:**
|
|
||||||
- `policy` type: 30 days (e.g., Shared Memory Policy, Agent Registry)
|
|
||||||
- `registry` type: 30 days (e.g., Node State Management, Health Dashboard)
|
|
||||||
- `template` type: 30 days (e.g., Agent SOUL.md)
|
|
||||||
- `note` type: 60 days (e.g., business context, project notes)
|
|
||||||
- `information` type: 90 days (e.g., reference documentation)
|
|
||||||
- `idea` type: 30 days (e.g., brainstorming, experimental notes)
|
|
||||||
|
|
||||||
**Transition Protocol:**
|
|
||||||
- `active` → `stale`: After days since `updated_at` exceeds threshold
|
|
||||||
- `stale` → `review_pending`: Daily cron check flags for human review
|
|
||||||
- `review_pending` → `active`: Human reviews and updates content
|
|
||||||
- `review_pending` → `archived`: After 15 days in review_pending state (automatic purge)
|
|
||||||
|
|
||||||
### 4. Orphan Prevention
|
|
||||||
**Rule:** No node should remain orphaned (0 edges) for more than 7 days.
|
|
||||||
|
|
||||||
**Orphan Detection:**
|
|
||||||
- Daily cron job identifies nodes with 0 incoming/outgoing edges
|
|
||||||
- Orphans are flagged with `state: "review_pending"` and `metadata.orphan_detected: true`
|
|
||||||
- 7-day grace period: Orphans must be either:
|
|
||||||
- Connected to relevant nodes via edges
|
|
||||||
- Merged into existing parent nodes
|
|
||||||
- Archived after 15 days in review_pending state
|
|
||||||
|
|
||||||
### 5. Agent Custodianship
|
|
||||||
**Rule:** Each agent is responsible for maintaining their own created nodes.
|
|
||||||
|
|
||||||
**Responsibilities:**
|
|
||||||
- **Mumuni**: Primary custodian for `syslogsolution` namespace, health dashboard, and all business nodes
|
|
||||||
- **Tanko**: Custodian for `agent-private` namespace, personal notes, and fitness/creative work
|
|
||||||
- **Okyeame**: Custodian for `shared` namespace, policy files, and collective knowledge nodes
|
|
||||||
- **All Agents**: Must verify embedding status before reporting node creation as complete
|
|
||||||
|
|
||||||
### 6. Embedding Failure Recovery
|
|
||||||
**Rule:** When embedding pipeline is unavailable, nodes are blocked from creation or immediately flagged.
|
|
||||||
|
|
||||||
**Recovery Protocol:**
|
|
||||||
1. **Detection**: Daily health check monitors `http://192.168.68.65:8080/health` endpoint
|
|
||||||
2. **Alert**: If embedding service is down, alert is sent to all agents via relay
|
|
||||||
3. **Mitigation**:
|
|
||||||
- New node creation is suspended
|
|
||||||
- Existing nodes with `chunk_status: "not_chunked"` are flagged as `unsearchable`
|
|
||||||
4. **Recovery**: When service returns:
|
|
||||||
- All `not_chunked` nodes are queued for re-embedding
|
|
||||||
- Re-embedding status logged in `embedding_retry_log` table
|
|
||||||
- 5-minute retry window with exponential backoff (1s, 2s, 4s, 8s, 16s)
|
|
||||||
|
|
||||||
### 7. Chunking Verification
|
|
||||||
**Rule:** All node modifications must verify chunking status within 5 minutes.
|
|
||||||
|
|
||||||
**Verification Protocol:**
|
|
||||||
- After `updateNode` or `createNode`, query `chunks` table for `node_id`
|
|
||||||
- If `chunk_status != "chunked"` after 5 minutes, log failure and flag node
|
|
||||||
- Automated cleanup script runs every 6 hours to retry failed chunks
|
|
||||||
- Maximum 3 retry attempts per node before escalation to human
|
|
||||||
|
|
||||||
### 8. Graph Health Monitoring
|
|
||||||
**Rule:** Health dashboard (Node #74) must reflect real-time graph status.
|
|
||||||
|
|
||||||
**Dashboard Requirements:**
|
|
||||||
- Total nodes, edges, orphans, stale nodes, chunk errors
|
|
||||||
- Embedding service health status
|
|
||||||
- Last successful embedding attempt
|
|
||||||
- Alert thresholds:
|
|
||||||
- Orphan rate > 40%: Warning
|
|
||||||
- Stale node rate > 50%: Critical
|
|
||||||
- Chunk error rate > 30%: Critical
|
|
||||||
- Embedding service down: Critical (immediate alert)
|
|
||||||
|
|
||||||
## Operational Procedures
|
|
||||||
|
|
||||||
### Daily Maintenance (Automated)
|
|
||||||
1. **Staleness Check**: Identify nodes past their threshold
|
|
||||||
2. **Orphan Detection**: Flag nodes with 0 edges
|
|
||||||
3. **Chunk Verification**: Retry failed chunks, log errors
|
|
||||||
4. **Health Update**: Refresh Node #74 dashboard
|
|
||||||
|
|
||||||
### Weekly Maintenance (Human Review)
|
|
||||||
1. **Review Pending Nodes**: Examine flagged nodes
|
|
||||||
2. **Edge Optimization**: Connect related orphaned nodes
|
|
||||||
3. **Archive Cleanup**: Remove nodes past 15-day review period
|
|
||||||
4. **Embedding Audit**: Verify all nodes have valid embeddings
|
|
||||||
|
|
||||||
### Emergency Recovery
|
|
||||||
1. **Embedding Service Down**:
|
|
||||||
- Suspend node creation
|
|
||||||
- Alert all agents
|
|
||||||
- Investigate root cause (check `192.168.68.65:8080`)
|
|
||||||
2. **Database Corruption**:
|
|
||||||
- Restore from latest PBS backup
|
|
||||||
- Verify chunk integrity
|
|
||||||
- Re-run embedding for affected nodes
|
|
||||||
3. **Mass Orphan Creation**:
|
|
||||||
- Identify source agent/namespace
|
|
||||||
- Review recent changes
|
|
||||||
- Reconnect or archive affected nodes
|
|
||||||
|
|
||||||
## Compliance & Enforcement
|
|
||||||
|
|
||||||
### Audit Schedule
|
|
||||||
- **Daily**: Automated health checks
|
|
||||||
- **Weekly**: Human review of review_pending nodes
|
|
||||||
- **Monthly**: Comprehensive graph audit (full schema validation)
|
|
||||||
- **Quarterly**: Contract review and threshold adjustment
|
|
||||||
|
|
||||||
### Violation Consequences
|
|
||||||
- **First**: Alert to responsible agent
|
|
||||||
- **Second**: Node placed in `review_pending` state
|
|
||||||
- **Third**: Agent access suspended until remediation
|
|
||||||
- **Fourth**: Escalation to human (Kwame)
|
|
||||||
|
|
||||||
### Metrics for Success
|
|
||||||
- Orphan rate: < 10%
|
|
||||||
- Stale node rate: < 20%
|
|
||||||
- Chunk error rate: < 5%
|
|
||||||
- Embedding service uptime: > 99.9%
|
|
||||||
- Node creation verification: 100%
|
|
||||||
|
|
||||||
## Implementation Notes
|
|
||||||
|
|
||||||
### Database Schema Requirements
|
|
||||||
- `nodes` table: Standard fields plus `embedding_status`, `chunk_status`, `last_embedding_attempt`
|
|
||||||
- `chunks` table: Must track `chunk_status` and `embedding_status`
|
|
||||||
- `embedding_retry_log` table: Track retry attempts for failed chunks
|
|
||||||
- `node_state_history` table: Log state transitions for audit trail
|
|
||||||
|
|
||||||
### Cron Job Requirements
|
|
||||||
- `ra-h-health-check`: Daily (4h interval)
|
|
||||||
- `ra-h-staleness-detection`: Daily
|
|
||||||
- `ra-h-orphan-detection`: Daily
|
|
||||||
- `ra-h-chunk-retry`: Every 6 hours
|
|
||||||
- `ra-h-review-cleanup`: Weekly (15-day archive)
|
|
||||||
|
|
||||||
### Agent Onboarding
|
|
||||||
All new agents must:
|
|
||||||
1. Read this contract
|
|
||||||
2. Understand namespace responsibilities
|
|
||||||
3. Know how to verify embedding status
|
|
||||||
4. Know the alert escalation path
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Approval
|
|
||||||
This contract is effective immediately upon approval by the primary custodian (Mumuni) and system owner (Kwame).
|
|
||||||
|
|
||||||
**Effective Date:** July 10, 2026
|
|
||||||
**Review Date:** October 10, 2026
|
|
||||||
**Primary Custodian:** Mumuni 🦅
|
|
||||||
**System Owner:** Jerome Tabiri
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
*Contract Version: 1.0*
|
|
||||||
*Last Updated: July 10, 2026*
|
|
||||||
*Storage: `/root/.hermes/skills/ra-h-os-custodianship-contract.prose.md`*
|
|
||||||
@@ -19,13 +19,53 @@ from datetime import datetime
|
|||||||
|
|
||||||
LITELLM = "http://192.168.68.116:80"
|
LITELLM = "http://192.168.68.116:80"
|
||||||
|
|
||||||
|
def _get_agent_key(agent_name):
|
||||||
|
"""Retrieve agent key from Infisical vault."""
|
||||||
|
try:
|
||||||
|
result = subprocess.run(
|
||||||
|
["infisical", "secrets", "get", "LITELLM_API_KEY",
|
||||||
|
"--project=agents", "--env=production", "--plain"],
|
||||||
|
capture_output=True, text=True, timeout=10
|
||||||
|
)
|
||||||
|
if result.returncode == 0:
|
||||||
|
return result.stdout.strip()
|
||||||
|
except Exception:
|
||||||
|
pass
|
||||||
|
|
||||||
|
# Fallback: try exporting all secrets
|
||||||
|
try:
|
||||||
|
result = subprocess.run(
|
||||||
|
["infisical", "export", "--project=agents", "--env=production",
|
||||||
|
"--format=dotenv"],
|
||||||
|
capture_output=True, text=True, timeout=10
|
||||||
|
)
|
||||||
|
if result.returncode == 0:
|
||||||
|
for line in result.stdout.splitlines():
|
||||||
|
if line.startswith(f"LITELLM_API_KEY_{agent_name.upper()}") or \
|
||||||
|
(line.startswith("LITELLM_API_KEY=") and agent_name == os.uname().nodename):
|
||||||
|
return line.split("=", 1)[1].strip().strip('"').strip("'")
|
||||||
|
except Exception:
|
||||||
|
pass
|
||||||
|
|
||||||
|
return None
|
||||||
|
|
||||||
|
# Agent keys are pulled from Infisical vault at runtime.
|
||||||
|
# The 'key' field is populated dynamically below.
|
||||||
AGENTS = {
|
AGENTS = {
|
||||||
"tanko": {"ct": 112, "host": "192.168.68.122", "key": "sk-CggiHWlamQyShxWC3Hx6uw", "user": "jerome"},
|
"tanko": {"ct": 112, "host": "192.168.68.122", "user": "jerome"},
|
||||||
"mumuni": {"ct": 114, "host": "192.168.68.123", "key": "sk-VrqCNlwUgzoNGOpikJ7nwQ", "user": "root"},
|
"mumuni": {"ct": 114, "host": "192.168.68.123", "user": "root"},
|
||||||
"tdunna": {"ct": 111, "host": None, "key": "sk-6sbCNjz2T6lTVDBdlNHXsA", "user": None},
|
"koby": {"ct": 111, "host": None, "user": None},
|
||||||
"baggy": {"ct": 113, "host": None, "key": "sk-krnw_zGBwvvL5b7l2t-s-A", "user": None},
|
"koonimo": {"ct": 113, "host": None, "user": None},
|
||||||
}
|
}
|
||||||
|
|
||||||
|
# Inject keys from vault
|
||||||
|
for agent_name in AGENTS:
|
||||||
|
key = _get_agent_key(agent_name)
|
||||||
|
if key:
|
||||||
|
AGENTS[agent_name]["key"] = key
|
||||||
|
else:
|
||||||
|
AGENTS[agent_name]["key"] = None
|
||||||
|
|
||||||
GPU_HOSTS = {
|
GPU_HOSTS = {
|
||||||
"gpu-rtx3090 (.8)": {"host": "192.168.68.8", "port": 8080, "service": "llama-server"},
|
"gpu-rtx3090 (.8)": {"host": "192.168.68.8", "port": 8080, "service": "llama-server"},
|
||||||
"gpu-rtx5070 (.110)": {"host": "192.168.68.110", "port": 8080, "service": "llama-server"},
|
"gpu-rtx5070 (.110)": {"host": "192.168.68.110", "port": 8080, "service": "llama-server"},
|
||||||
|
|||||||
Reference in New Issue
Block a user