Merge pull request 'fix: redirect health check logs from knowledge graph to Gitea (hard rule)' (#36) from fix/health-logs-to-gitea into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s

This commit was merged in pull request #36.
This commit is contained in:
2026-07-28 21:21:46 +00:00
3 changed files with 11 additions and 9 deletions
+2 -2
View File
@@ -251,8 +251,8 @@ call update-gpu-health
## Reporting ## Reporting
### 1. Knowledge Graph ### 1. Gitea Log (not knowledge graph — hard rule)
Every action logged as `[GPU-SELF-HEAL] <run_id>` node with full audit trail. Pushed to `SyslogSolution/health-logs/gpu/{run_id}.json` — versioned, searchable, not in graph.
### 2. Zulip Alerts (#agent-hub → alerts-gpu) ### 2. Zulip Alerts (#agent-hub → alerts-gpu)
- `issues_fixed > 0` → "🛠 GPU Self-Heal — <gpu> <issue> resolved" - `issues_fixed > 0` → "🛠 GPU Self-Heal — <gpu> <issue> resolved"
+8 -6
View File
@@ -7,7 +7,7 @@ note: >
Auto-remediation code was removed from the pi Zulip extension (retired 2026-07-04), Auto-remediation code was removed from the pi Zulip extension (retired 2026-07-04),
now reimplemented as `litellm-health-check.sh` on CT 116. now reimplemented as `litellm-health-check.sh` on CT 116.
Script: `/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116 (cron `0 */6 * * *`). Script: `/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116 (cron `0 */6 * * *`).
Reports to /var/log/litellm/health-*.json and RA-H OS knowledge graph. Reports to /var/log/litellm/health-*.json and Gitea (SyslogSolution/health-logs).
GPU monitoring integrated from gpu-monitor on .24:9100. GPU monitoring integrated from gpu-monitor on .24:9100.
Consolidated from litellm-health + litellm-self-heal on 2026-07-09 to eliminate Consolidated from litellm-health + litellm-self-heal on 2026-07-09 to eliminate
@@ -20,7 +20,7 @@ description: >
LiteLLM inference stack health monitoring + self-healing. Verifies the full LiteLLM inference stack health monitoring + self-healing. Verifies the full
nginx → LiteLLM → GPU chain, 8 containers on CT 116, 3 GPU hosts, model nginx → LiteLLM → GPU chain, 8 containers on CT 116, 3 GPU hosts, model
inference, and agent keys. Applies remediation rules for common failures. inference, and agent keys. Applies remediation rules for common failures.
Reports every action via Zulip DM and RA-H OS knowledge graph. Reports every action via Zulip DM and Gitea (SyslogSolution/health-logs).
--- ---
# LiteLLM Operations — Health Check + Self-Heal # LiteLLM Operations — Health Check + Self-Heal
@@ -211,8 +211,8 @@ for reference but inactive. If Redis issues occur, check harness-redis container
Every remediation cycle produces a structured report: Every remediation cycle produces a structured report:
### 1. RA-H OS Knowledge Graph Node ### 1. Gitea Log Entry
Created as `[LEARN] litellm-self-heal: <run_id>` with full JSON report. Pushed to `SyslogSolution/health-logs/litellm/{run_id}.json` — versioned, searchable, not in graph.
### 2. Zulip DM to Owner ### 2. Zulip DM to Owner
- `issues_fixed > 0` — "🛠 LiteLLM Self-Heal — Fix Applied" - `issues_fixed > 0` — "🛠 LiteLLM Self-Heal — Fix Applied"
@@ -252,9 +252,11 @@ call report-generator
health: health health: health
actions: actions actions: actions
-- Phase 4: Log to knowledge graph -- Phase 4: Log to Gitea (not knowledge graph — hard rule)
call kg-logger call gitea-logger
run_id: run_id run_id: run_id
repo: SyslogSolution/health-logs
path: litellm/{run_id}.json
health: health health: health
actions: actions actions: actions
+1 -1
View File
@@ -52,7 +52,7 @@ description: >
- If status is "online" → pass, log restarts count - If status is "online" → pass, log restarts count
- If status is "stopped" or "errored" → **DO NOT RESTART** — alert owner immediately - If status is "stopped" or "errored" → **DO NOT RESTART** — alert owner immediately
- If restarts > 5 in last hour → alert owner with full diagnostics - If restarts > 5 in last hour → alert owner with full diagnostics
4. **Log results**Create `[LEARN]` node in knowledge graph for any actions taken 4. **Log results**Append to `SyslogSolution/health-logs/pm2/{timestamp}.md` in Gitea (not knowledge graph — hard rule)
5. **Alert** — Send Zulip DM to owner if escalation needed (do NOT run pm2 commands during alerting) 5. **Alert** — Send Zulip DM to owner if escalation needed (do NOT run pm2 commands during alerting)
6. **Wait 5 min** → repeat from step 1 6. **Wait 5 min** → repeat from step 1