Audit results (all patterns checked across .md, .prose.md, .sh, .py, .js, .ts, .json, .yaml, .yml, .env): - sk-or-v1 (OpenRouter): 0 occurrences - sk- prefix (20+ chars): 0 occurrences - sk_live: 0 occurrences - Bearer <key>: 0 occurrences - api_key: <value>: 0 occurrences - PASSWORD=: 0 occurrences - TOKEN=: 0 occurrences - SECRET=: 0 occurrences Files changed: - agent-zero-fix-summary.md (removed 2 OpenRouter keys) - agent-zero-openrouter-key.prose.md (removed 1 OpenRouter key) - hermes-key-enforcement.prose.md (removed 1 LiteLLM key, 1 external key) - litellm-api-keys.prose.md (removed 1 LiteLLM key) - litellm-self-heal.prose.md (removed 1 stale key reference) - scripts/agent-health-check.py (INFISICAL_TOKEN now required) - scripts/daily-infra-report.py (EMAIL_PASSWORD now required) - zulip-health.prose.md (TOKEN references annotated)
304 lines
13 KiB
Markdown
304 lines
13 KiB
Markdown
---
|
|
report_only_agents:
|
|
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
|
|
kind: responsibility
|
|
name: litellm-self-heal
|
|
status: deployed
|
|
note: >
|
|
DEPLOYED 2026-07-12 on CT 116 cron: 0 */6 * * *
|
|
Auto-remediation code was removed from the pi Zulip extension (retired 2026-07-04),
|
|
now reimplemented as `litellm-health-check.sh` on CT 116.
|
|
Script: `/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116 (cron `0 */6 * * *`).
|
|
Reports to /var/log/litellm/health-*.json and Gitea (SyslogSolution/health-logs).
|
|
GPU monitoring integrated from gpu-monitor on .24:9100.
|
|
|
|
Health probes are owned by litellm-health.prose.md (dispatched as
|
|
`run contract: litellm-health`). This contract owns remediation only — it does not
|
|
re-specify the probes.
|
|
|
|
Source of truth for GPU topology and keys: gpu-fleet.prose.md
|
|
Last verified: 2026-07-12
|
|
description: >
|
|
LiteLLM inference stack remediation. Applies remediation rules for failures detected
|
|
by litellm-health.prose.md (nginx → LiteLLM → GPU chain, CT 116 containers, GPU hosts,
|
|
model inference, and agent keys).
|
|
Reports every action via Zulip DM and Gitea (SyslogSolution/health-logs).
|
|
---
|
|
---
|
|
|
|
# LiteLLM Operations — Self-Heal (Remediation)
|
|
|
|
## Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU)
|
|
|
|
```
|
|
Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
|
│
|
|
Key validation
|
|
Fallback chains
|
|
Budget tracking
|
|
│
|
|
Prometheus ← metrics
|
|
│
|
|
Grafana :3001
|
|
|
|
harness-router :9000 — DECOMMISSIONED 2026-09-11 (container, image
|
|
and config removed). nginx routes /v1 → LiteLLM directly.
|
|
Router slot booking + circuit breakers replaced by
|
|
LiteLLM native fallbacks + timeouts.
|
|
```
|
|
|
|
## Parameters
|
|
|
|
- public_url: string — Public LiteLLM URL (default: "https://litellm.sysloggh.net")
|
|
- backend_host: string — Internal CT host (default: "192.168.68.116")
|
|
- auth_host: string — Authentik OIDC server (default: "192.168.68.11")
|
|
- gpu_hosts: array — GPU inference hosts (default: ["192.168.68.8", "192.168.68.110", "192.168.68.15"])
|
|
- gpu_dashboard_url: string — Fleet dashboard (default: "http://192.168.68.24:9100")
|
|
- grafana_url: string — Grafana dashboards (default: "http://192.168.68.116:3001")
|
|
|
|
## GPU Fleet Topology
|
|
|
|
| Host | IP | Hardware | Role |
|
|
|------|-----|----------|------|
|
|
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | heavy reasoning (`gpu-dense`) |
|
|
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | vision / web extract / light tasks (`gpu-vision`) |
|
|
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | compression (`strix-moe`) |
|
|
|
|
> Verified on the ground 2026-07-16. The legacy name `ornith-1.0-35b` does NOT exist in LiteLLM and must not be referenced.
|
|
|
|
## LiteLLM Model Surface
|
|
|
|
Single source of truth for models, aliases, rpm caps, weights and fallback chains:
|
|
CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those values in
|
|
contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-light`,
|
|
`crew-auto`).
|
|
|
|
### Alias Surface
|
|
|
|
| Alias | Serves | Where | Kind |
|
|
|-------|--------|-------|------|
|
|
| `gpu-dense` | heavy reasoning | RTX 3090 (192.168.68.8) | direct alias |
|
|
| `gpu-vision` | vision / web extract / light tasks | RTX 5070 (192.168.68.110) | direct alias AND `syslog-auto` pool member |
|
|
| `strix-moe` | compression (MoE) | Strix Halo (192.168.68.15) | direct alias |
|
|
| `syslog-auto` | balanced default | weighted pool across the three GPU hosts | pool router |
|
|
|
|
### Context Cap Split (2026-08-20; crew cap RETIRED)
|
|
|
|
- **Abiba (firstmate)**: 128K uncapped — unlimited context for primary workloads
|
|
- **Hermes agents** (mumuni, tanko, koby, koonimo): 128K uncapped
|
|
- **Crewmates** (ops, tune, verify, auth-keys, build): the 64K cap was retired together with the `crew-auto` alias; NO context cap is currently in force.
|
|
|
|
> The 64K crew cap was retired with `crew-auto` (2026-09-12). No limit is currently in force; reinstating one would need per-key model limits as a separate, deliberately-scoped change.
|
|
|
|
- Key scoping: `/v1/models` is key-scoped, so the set a caller sees must be read with a named key rather than assumed — a monitor key, an agent key, and the master key can each return a different set. Agents should use the stable aliases (`strix-moe`, `gpu-vision`, `gpu-dense`) rather than raw model names, so model swaps don't break them.
|
|
- **NEVER use `litellm_proxy_master_key` (the `sk-litellm-...` master key) for inference.** It is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`). All inference — agent traffic, health-check model tests, monitor scripts — uses agent-specific keys. The health-check script's model tests use a dedicated `monitor` agent key stored at `/etc/litellm-monitor.env` on CT 116 (root-only, `chmod 600`); `/key/list` is the only call that legitimately uses `$MASTER_KEY`.
|
|
|
|
> Re-scope note (2026-09-12): the fallback-chain and per-model timeout tables were removed
|
|
> by the single-source-of-truth re-scope; read those values from the CT 116 config named
|
|
> above rather than from this contract.
|
|
|
|
## Containers on CT 116
|
|
|
|
| Container | Image | Port | Health Check |
|
|
|-----------|-------|------|-------------|
|
|
| harness-litellm | ghcr.io/docker.litellm.ai/berriai/litellm:1.99.1 | :4000→:4000 | /health/liveliness |
|
|
| harness-nginx | nginx:alpine | :80 | HTTP 200 on /health |
|
|
| harness-postgres | postgres:16-alpine | :5432 | pg_isready |
|
|
| harness-redis | redis:7-alpine | :6379 | PING |
|
|
| harness-dashboard | inference-harness-dashboard | :3000 | /health |
|
|
| harness-grafana | grafana/grafana | :3000→:3001 | /api/health |
|
|
| harness-prometheus | prom/prometheus | :9090 | /-/healthy |
|
|
| harness-docker-stats | python:3.12-alpine | — | container stats exporter |
|
|
| harness-pve-exporter | prompve/prometheus-pve-exporter | — | Proxmox metrics → Prometheus |
|
|
|
|
## Script Operations (synced 2026-07-16)
|
|
|
|
- **Health-check script** (`/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116): `gpu-fleet` check fails only on **critical** alerts (warnings are informational). Tests `strix-moe` (not `ornith-1.0-35b`).
|
|
- **GPU monitor** (`/root/scripts/gpu-monitor-server.py` on pi .24): runs as **systemd unit `gpu-monitor.service`** (was bare `&` process). `gpu_count` includes Strix Halo (was 2, now 3). VRAM alert thresholds: warning 93%, critical 97% (raised from 90/95 — 128K context steady-state is ~70% on RTX 3090, not a fault).
|
|
- **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): v4 (2026-09-10) — vault-backed agents (tanko/koby/koonimo) read their **agent-specific** `{NAME}_LITELLM_API_KEY` from Infisical vault (not the shared master key); abiba (pi agent) reads `LITELLM_API_KEY` from its local `/root/.pi/agent/env.sh` (#735 — moved out of shared `/root/.bashrc`), not from the vault. Abiba is pi-only since the harness purge, so its Hermes config/wrapper/gateway legs are skipped rather than reported as faults; koby is **report-only** (captain's 2026-08-17 ruling) — its findings go to the `--json` `report_only` array and are never counted as fleet failures or repaired, and its CT 111 liveness is probed on storepve (.6). Covers: LiteLLM keys, GPU ports, agent gateways, CT liveness (pct status on PVE nodes), config.yaml YAML integrity, wrapper/CLI integrity, vault secret non-emptiness checks. Every run/report carries the absolute execution path (`script=` + `cwd=`). The current fleet roster is owned by the script changelog (`scripts/agent-health-check.py`); mumuni is no longer probed from this host. Legacy `tdunna`/`baggy` replaced with canonical agent hostnames.
|
|
- **Stale keys cleaned**: `daily-infra-report.py` SYNTHETIC_API_KEY was stale (`«vault: agents/production LITELLM_API_KEY»` → 401); now reads `LITELLM_MASTER_KEY` from env. Deprecated scripts (`router-original.py`, `router-phase0-backup.py`, `apply-fixes.py`) still reference `sk-syslog-local-master-key` but do not actively poll LiteLLM.
|
|
|
|
## Maintains
|
|
|
|
- litellm-admin-ui: { status: "healthy", last_check: timestamp }
|
|
- litellm-api-docs: { status: "healthy", last_check: timestamp }
|
|
- litellm-containers: { status: "healthy", last_check: timestamp }
|
|
- litellm-oidc: { status: "healthy", last_check: timestamp }
|
|
- litellm-gpu-fleet: { status: "healthy", models: int, alerts: array }
|
|
- litellm-agent-keys: { count: int, valid: int }
|
|
|
|
## Continuity
|
|
|
|
- Self-driven: check every 300 seconds
|
|
- Also wakes on user request
|
|
- On failure: re-check after 30s, escalate after 3 consecutive failures
|
|
|
|
---
|
|
---
|
|
|
|
## Health Check
|
|
|
|
Health probes are owned by `litellm-health.prose.md` (dispatched as `run contract: litellm-health`). This contract owns remediation only — it does not re-specify the probes.
|
|
|
|
---
|
|
---
|
|
|
|
## Remediation Rules
|
|
|
|
### Rule 1: Container Not Healthy
|
|
Detect → `docker compose up -d <container>` → verify → log
|
|
Escalate after 3 failures
|
|
|
|
### Rule 2: Admin UI Non-200
|
|
Detect → restart harness-litellm → verify → log
|
|
Escalate after 2 failures
|
|
|
|
### Rule 3: Auth Unreachable
|
|
Detect → immediate escalate (external dependency)
|
|
|
|
### Rule 4: Docs 404
|
|
Detect → fix DOCS_URL env var → verify → log
|
|
|
|
### Rule 5: GPU Unreachable
|
|
Detect → SSH to GPU host fails or model not responding
|
|
Fix → restart llama-server on host via SSH
|
|
Escalate → immediately (downtime affects inference)
|
|
|
|
### Rule 6: Model Not Responding
|
|
Detect → test prompt via LiteLLM returns non-200
|
|
Fix → restart llama-server on GPU host → verify
|
|
Escalate → after 2 failed restarts
|
|
|
|
### Rule 7: Router Returns 503 (All GPUs Saturated) — DEPRECATED
|
|
Router is no longer in the request path (2026-07-08). LiteLLM proxies
|
|
directly to GPU. This rule is retained for reference but is inactive.
|
|
If 503 errors occur, check LiteLLM timeouts and GPU health directly.
|
|
|
|
### Rule 8: Agent Keys Invalid (401)
|
|
Detect → agents report auth errors
|
|
Root cause → keys not in LiteLLM DB
|
|
Fix → generate keys in LiteLLM via /key/generate → update /etc/environment on agent hosts
|
|
Escalate → if SSH access unavailable, send Zulip DM
|
|
|
|
### Rule 9: Stale Active Counter in Redis — DEPRECATED
|
|
Router no longer in path so router active-slot counters are unused. Rule retained
|
|
for reference but inactive. `harness-redis` now serves only LiteLLM cache and
|
|
rate-limit state; check the container if cache errors appear.
|
|
|
|
---
|
|
---
|
|
|
|
## Reporting
|
|
|
|
Every remediation cycle produces a structured report:
|
|
|
|
### 1. Gitea Log Entry
|
|
Pushed to `SyslogSolution/health-logs/litellm/{run_id}.json` — versioned, searchable, not in graph.
|
|
|
|
### 2. Zulip DM to Owner
|
|
- `issues_fixed > 0` — "🛠 LiteLLM Self-Heal — Fix Applied"
|
|
- `issues_escalated > 0` — "⚠ LiteLLM Self-Heal — Needs Your Attention"
|
|
- Every 10th clean cycle — "✅ All Clear (10 checks passed)"
|
|
|
|
### 3. Daily Digest (end of day)
|
|
Summary of last 24 hours: total cycles, issues found/fixed/escalated,
|
|
top actions, uptime.
|
|
|
|
### 4. Relay Message (if cross-agent)
|
|
If a fix requires another agent (e.g., Authentik restart), relay sent
|
|
to responsible agent with full context.
|
|
|
|
---
|
|
---
|
|
|
|
## Execution
|
|
|
|
```prose
|
|
-- Phase 1: Health Check
|
|
let health = call health-check
|
|
public_url: public_url
|
|
backend_host: backend_host
|
|
gpu_dashboard_url: gpu_dashboard_url
|
|
grafana_url: grafana_url
|
|
|
|
-- Phase 2: Apply remediation for each failure
|
|
let actions = []
|
|
for check in health.failed:
|
|
let fix = apply-remediation-rule
|
|
rule: lookup-rule(check.name)
|
|
target: check.target
|
|
push actions fix
|
|
|
|
-- Phase 3: Generate report
|
|
call report-generator
|
|
health: health
|
|
actions: actions
|
|
|
|
-- Phase 4: Log to Gitea (not knowledge graph — hard rule)
|
|
call gitea-logger
|
|
run_id: run_id
|
|
repo: SyslogSolution/health-logs
|
|
path: litellm/{run_id}.json
|
|
health: health
|
|
actions: actions
|
|
|
|
-- Phase 5: Notify if anything changed
|
|
if actions.length > 0:
|
|
call zulip-notifier
|
|
actions: actions
|
|
health: health
|
|
|
|
-- Wait 300s and repeat
|
|
```
|
|
|
|
## Audit Trail Format
|
|
|
|
```json
|
|
{
|
|
"run_id": "self-heal-20260709-001",
|
|
"timestamp": "2026-07-09T14:00:00Z",
|
|
"duration_ms": 1234,
|
|
"checks_passed": 7,
|
|
"checks_failed": 0,
|
|
"issues_found": 0,
|
|
"issues_fixed": 0,
|
|
"issues_escalated": 0,
|
|
"actions": []
|
|
}
|
|
```
|
|
|
|
With failures:
|
|
```json
|
|
{
|
|
"run_id": "self-heal-20260709-002",
|
|
"issues_found": 1,
|
|
"issues_fixed": 1,
|
|
"actions": [
|
|
{
|
|
"rule": "Container Not Healthy",
|
|
"target": "harness-nginx",
|
|
"action": "restarted container",
|
|
"result": "healthy",
|
|
"duration_ms": 5000
|
|
}
|
|
]
|
|
}
|
|
```
|
|
|
|
## gitea-logger Implementation
|
|
|
|
When this step executes, write the report JSON to a temp file and push to Gitea:
|
|
|
|
```bash
|
|
REPO="https://abiba-bot:${GITEA_PAT}@git.sysloggh.net/SyslogSolution/health-logs"
|
|
DIR="litellm"
|
|
FILE="${run_id}.json"
|
|
echo "${report_json}" > /tmp/${FILE}
|
|
(cd /tmp && git clone --depth 1 "${REPO}" &&
|
|
cp ${FILE} health-logs/${DIR}/${FILE} &&
|
|
cd health-logs && git add ${DIR}/${FILE} &&
|
|
git commit -m "litellm-health: ${run_id}" && git push)
|
|
rm -rf /tmp/health-logs /tmp/${FILE}
|
|
```
|
|
|