Improve memory audit contract: autonomous execution, incremental patches, pre-flight checks, stale state thresholds, dry-run mode, success metrics
This commit is contained in:
@@ -0,0 +1,97 @@
|
|||||||
|
# Prose Contracts — Syslog Solution LLC
|
||||||
|
|
||||||
|
Operating contracts for Hermes agents. Each `.prose.md` file defines a **responsibility** (recurring duty) or **template** (scaffold) that agents can execute via OpenProse or follow manually.
|
||||||
|
|
||||||
|
## How Agents Run These Contracts
|
||||||
|
|
||||||
|
### Option A: Via OpenProse CLI (preferred)
|
||||||
|
|
||||||
|
```bash
|
||||||
|
prose run <contract-name> [param1=value1 param2=value2 ...]
|
||||||
|
```
|
||||||
|
|
||||||
|
**Examples:**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Run a memory audit with default parameters
|
||||||
|
prose run memory-audit-maintenance
|
||||||
|
|
||||||
|
# Run a memory audit with custom threshold
|
||||||
|
prose run memory-audit-maintenance memory_threshold=90 verify_configs=true
|
||||||
|
|
||||||
|
# Configure a new agent with the template
|
||||||
|
prose run hermes-config-template agent_name=syslog-devops default_model=claude-sonnet-4
|
||||||
|
|
||||||
|
# Configure an agent with a different auxiliary model
|
||||||
|
prose run hermes-config-template agent_name=syslog-code default_model=qwen3.6-27B-code auxiliary_model=gemma-4-12b
|
||||||
|
```
|
||||||
|
|
||||||
|
### Option B: Manual Execution
|
||||||
|
|
||||||
|
If `prose` CLI isn't available, any agent can:
|
||||||
|
1. Fetch the raw contract: `curl -sL https://git.sysloggh.net/SyslogSolution/prose-contracts/raw/branch/master/<filename>`
|
||||||
|
2. Read the **Execution** section
|
||||||
|
3. Follow the steps
|
||||||
|
|
||||||
|
**Example for fetching:**
|
||||||
|
```bash
|
||||||
|
curl -sL https://git.sysloggh.net/SyslogSolution/prose-contracts/raw/branch/master/memory-audit-maintenance.prose.md
|
||||||
|
```
|
||||||
|
|
||||||
|
## Available Contracts
|
||||||
|
|
||||||
|
### Responsibilities (Recurring Duties)
|
||||||
|
|
||||||
|
| Contract | What It Does | When To Run |
|
||||||
|
|---|---|---|
|
||||||
|
| `memory-audit-maintenance` | Audits & reorganizes an agent's native memory (MEMORY.md, USER.md). Categorizes entries, moves rules to skills, verifies configs, frees up char budget. | Memory >85% usage or user request: "memory audit" |
|
||||||
|
| `zulip-health` | Checks Zulip connectivity, message flow, and bot responsiveness. | On Zulip issues or periodic health check |
|
||||||
|
| `litellm-health` | Verifies LiteLLM proxy connectivity, model availability, and key rotation. | On inference failures or periodic check |
|
||||||
|
| `litellm-self-heal` | Attempts to fix common LiteLLM issues (key rotation, restart, config reload). | When litellm-health reports failures |
|
||||||
|
| `zulip-mention-reliability` | Diagnoses and fixes @mention detection issues in Zulip. | On missed @mention reports |
|
||||||
|
| `pm2-self-heal` | Restarts crashed PM2 processes and verifies recovery. | On PM2 process failure |
|
||||||
|
|
||||||
|
### Templates (Scaffolds)
|
||||||
|
|
||||||
|
| Contract | What It Does | Parameters |
|
||||||
|
|---|---|---|
|
||||||
|
| `hermes-config-template` | Generates a standard Hermes config for any agent. Enforces shared infra (Firecrawl, SearXNG, RA-H OS MCP, LiteLLM) while keeping model choices flexible. | `agent_name` (required), `default_model`, `default_provider`, `fallback_model`, `auxiliary_model` |
|
||||||
|
|
||||||
|
### Scripts
|
||||||
|
|
||||||
|
| File | What It Does |
|
||||||
|
|---|---|
|
||||||
|
| `pm2-self-heal.sh` | Shell script to restart crashed PM2 processes (companion to the prose contract) |
|
||||||
|
|
||||||
|
## Contract Structure
|
||||||
|
|
||||||
|
Every `.prose.md` file has the same structure:
|
||||||
|
|
||||||
|
```
|
||||||
|
---
|
||||||
|
kind: <responsibility | template>
|
||||||
|
name: <unique-name>
|
||||||
|
description: ...
|
||||||
|
---
|
||||||
|
|
||||||
|
## Maintains — What state the contract tracks
|
||||||
|
## Parameters — Tunable inputs (with defaults)
|
||||||
|
## Continuity — When to run (triggers & cadence)
|
||||||
|
## Success Criteria — How to know it worked
|
||||||
|
## Remembers — What to learn between runs
|
||||||
|
## <Type> Rules — Inviolable rules for execution
|
||||||
|
## Execution — Step-by-step instructions
|
||||||
|
## Example Output — What a good run looks like
|
||||||
|
```
|
||||||
|
|
||||||
|
## Adding New Contracts
|
||||||
|
|
||||||
|
1. Create a `.prose.md` file following the structure above
|
||||||
|
2. Push to this repo
|
||||||
|
3. The contract is immediately available for any agent to `prose run`
|
||||||
|
|
||||||
|
## Repository
|
||||||
|
|
||||||
|
- **URL:** https://git.sysloggh.net/SyslogSolution/prose-contracts
|
||||||
|
- **Branch:** master
|
||||||
|
- **Auth:** mumuni-bot (write access via PAT)
|
||||||
@@ -0,0 +1,118 @@
|
|||||||
|
---
|
||||||
|
kind: responsibility
|
||||||
|
name: build-zulip-plugin
|
||||||
|
description: >
|
||||||
|
Generates and iteratively improves a Hermes Zulip platform plugin.
|
||||||
|
Each run can produce a new version of the plugin, learning from
|
||||||
|
previous runs and incrementally improving the adapter code.
|
||||||
|
The contract itself can be updated to capture new patterns.
|
||||||
|
---
|
||||||
|
|
||||||
|
## Maintains
|
||||||
|
|
||||||
|
- plugin_version: string — Current version of the generated plugin
|
||||||
|
- generation_count: number — How many generations have been produced
|
||||||
|
- last_generation: timestamp — When the last generation was created
|
||||||
|
- known_issues: array — Issues discovered in the current version
|
||||||
|
|
||||||
|
## Parameters
|
||||||
|
|
||||||
|
- zulip_site: string — Zulip server URL (default: "https://chat.sysloggh.net")
|
||||||
|
- bot_email_prefix: string — Prefix for bot emails (default: "abiba-bot")
|
||||||
|
- output_dir: string — Where to write the plugin (default: "plugins/platforms/zulip")
|
||||||
|
- generation_goal: string — What to improve this generation (default: "fix bugs and improve reliability")
|
||||||
|
|
||||||
|
## Continuity
|
||||||
|
|
||||||
|
The improvement cycle is **event-driven**, not just cron-based:
|
||||||
|
|
||||||
|
- **On check failure**: If any Success Criteria fails, wake immediately to fix
|
||||||
|
- **On user request**: `prose run build-zulip-plugin` — explicit upgrade
|
||||||
|
- **On new pattern**: If another plugin or agent reports a better pattern, wake to assimilate
|
||||||
|
- **Retrospective**: Every 24 hours if no other trigger fired — scan for subtle degradation
|
||||||
|
- **On first deploy**: Always run Generation 1 when no plugin exists yet
|
||||||
|
|
||||||
|
## Success Criteria (What Makes a Plugin "Good")
|
||||||
|
|
||||||
|
The plugin is considered GOOD when ALL of these pass:
|
||||||
|
|
||||||
|
### Connectivity
|
||||||
|
- plugin.connected == true — Establishes and maintains Zulip event queue
|
||||||
|
- plugin.reconnect_on_failure == true — Reconnects after queue expiry
|
||||||
|
- plugin.recovers_from_network_loss == true — Handles temporary network drops
|
||||||
|
|
||||||
|
### Message Flow
|
||||||
|
- dm_response_time_ms < 10000 — DMs get a response within 10 seconds
|
||||||
|
- stream_mention_detection == true — @mentions in streams are detected and routed
|
||||||
|
- all_bots_detection == true — @all-bots mentions are detected
|
||||||
|
- echo_loop_prevention == true — Never replies to its own messages
|
||||||
|
|
||||||
|
### Code Quality
|
||||||
|
- syntax_check == "pass" — All Python files are valid
|
||||||
|
- has_register_function == true — Exposes `register(ctx)` entry point
|
||||||
|
- plugin_yaml_valid == true — plugin.yaml parses correctly
|
||||||
|
- follows_base_adapter_pattern == true — Extends BasePlatformAdapter
|
||||||
|
|
||||||
|
### Reliability
|
||||||
|
- graceful_disconnect == true — shutdown doesn't crash
|
||||||
|
- handles_malformed_messages == true — Bad JSON doesn't kill the poll loop
|
||||||
|
- typing_indicators_work == true — Sends typing notifications
|
||||||
|
|
||||||
|
## Remembers
|
||||||
|
|
||||||
|
Each generation stores:
|
||||||
|
- What worked well (success patterns)
|
||||||
|
- What broke (failure modes)
|
||||||
|
- What the user complained about (pain points)
|
||||||
|
- Which Success Criteria passed and failed
|
||||||
|
|
||||||
|
## Generation Rules
|
||||||
|
|
||||||
|
### Rule 1: First Generation
|
||||||
|
- Generate complete plugin.yaml, adapter.py, __init__.py
|
||||||
|
- Follow Hermes BasePlatformAdapter pattern
|
||||||
|
- Include: connect, disconnect, send, edit_message, send_typing
|
||||||
|
- Include: DM-first routing, @mention detection, @all-bots support
|
||||||
|
|
||||||
|
### Rule 2: Subsequent Generations
|
||||||
|
- Read the previous generation's output and known issues
|
||||||
|
- Apply improvements based on generation_goal
|
||||||
|
- Fix any issues from the previous generation
|
||||||
|
- Add new capabilities if needed
|
||||||
|
|
||||||
|
### Rule 3: Testing
|
||||||
|
- After generating, run syntax check on the Python code
|
||||||
|
- Verify plugin.yaml is valid YAML
|
||||||
|
- Log the generation result to knowledge graph
|
||||||
|
|
||||||
|
### Rule 4: Self-Improvement
|
||||||
|
- After N generations, review the contract itself
|
||||||
|
- If new patterns emerged, update the Generation Rules
|
||||||
|
- If old rules are no longer relevant, remove them
|
||||||
|
|
||||||
|
## Execution
|
||||||
|
|
||||||
|
1. **Read state** — Check previous generations from knowledge graph
|
||||||
|
2. **Generate or improve** — Based on generation count:
|
||||||
|
- Generation 1: Scaffold full plugin from scratch
|
||||||
|
- Generation N: Read previous code, apply improvements
|
||||||
|
3. **Validate** — Syntax check, structure check, dependency check
|
||||||
|
4. **Log** — Save generation result to knowledge graph as [LEARN]
|
||||||
|
5. **Report** — Output what was generated, what changed, what needs review
|
||||||
|
|
||||||
|
## Example Output
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"generation": 1,
|
||||||
|
"version": "1.0.0",
|
||||||
|
"files_written": [
|
||||||
|
"plugins/platforms/zulip/plugin.yaml",
|
||||||
|
"plugins/platforms/zulip/__init__.py",
|
||||||
|
"plugins/platforms/zulip/adapter.py"
|
||||||
|
],
|
||||||
|
"syntax_check": "pass",
|
||||||
|
"known_issues": [],
|
||||||
|
"next_generation_goal": "add streaming responses via placeholder->edit pattern"
|
||||||
|
}
|
||||||
|
```
|
||||||
@@ -0,0 +1,15 @@
|
|||||||
|
---
|
||||||
|
kind: function
|
||||||
|
name: hello-world
|
||||||
|
description: A simple hello world contract to test OpenProse on pi
|
||||||
|
---
|
||||||
|
|
||||||
|
## Parameters
|
||||||
|
- name: string — The name to greet (default: "World")
|
||||||
|
|
||||||
|
## Returns
|
||||||
|
- greeting: string — The generated greeting message
|
||||||
|
|
||||||
|
## Ensures
|
||||||
|
- The greeting includes the provided name
|
||||||
|
- The greeting is friendly and warm
|
||||||
@@ -0,0 +1,307 @@
|
|||||||
|
---
|
||||||
|
kind: template
|
||||||
|
name: hermes-config-template
|
||||||
|
description: >
|
||||||
|
Standard Hermes configuration template for Syslog Solution LLC agents.
|
||||||
|
Enforces shared infrastructure setup (Firecrawl, SearXNG, local models,
|
||||||
|
RA-H OS MCP) while keeping model/provider choices flexible per agent role.
|
||||||
|
Every new agent profile should start from this template.
|
||||||
|
---
|
||||||
|
|
||||||
|
## Maintains
|
||||||
|
|
||||||
|
- template_version: string — Current template version (e.g., "1.0.0")
|
||||||
|
- last_applied: timestamp — When any agent was last configured from this template
|
||||||
|
- agents_configured: array — Which agents/profiles were created from this template
|
||||||
|
- infra_endpoints_verified: array — Firecrawl, SearXNG, RA-H OS, LiteLLM endpoints last confirmed working
|
||||||
|
- known_deviations: array — Profiles that diverge from the template (and why)
|
||||||
|
|
||||||
|
## Parameters
|
||||||
|
|
||||||
|
- agent_name: string — Name of the agent/profile (e.g., "syslog-code", "syslog-devops")
|
||||||
|
- default_model: string — Agent's primary model (e.g., "qwen3.6-27B-code", "syslog-auto")
|
||||||
|
- default_provider: string — Inference provider for the primary model (default: "harness")
|
||||||
|
- fallback_model: string — Fallback model when primary is down (default: "gemma-4-12b")
|
||||||
|
- fallback_provider: string — Fallback provider (default: "harness")
|
||||||
|
- auxiliary_model: string — Model for vision/compression/extraction tasks (default: "gemma-4-12b")
|
||||||
|
- enable_firecrawl_searxng: boolean — Whether to configure web search/extract (default: true)
|
||||||
|
- enable_ra_h_os_mcp: boolean — Whether to wire up RA-H OS MCP bridge (default: true)
|
||||||
|
- enable_compression: boolean — Whether to enable context compression (default: true)
|
||||||
|
- custom_provider_name: string — Custom provider name (default: "harness")
|
||||||
|
- custom_provider_model: string — Custom provider model (default: same as default_model)
|
||||||
|
- custom_provider_url: string — LiteLLM base URL (default: "http://litellm.sysloggh.net/litellm/v1")
|
||||||
|
- extra_auxiliary: array — Additional auxiliary tasks to configure (vision, compression, etc.)
|
||||||
|
- output_dir: string — Where to write the config (default: "~/.hermes/profiles/<agent_name>/config.yaml")
|
||||||
|
|
||||||
|
## Quickstart — How to Run This Contract
|
||||||
|
|
||||||
|
### For the Agent Running This Contract
|
||||||
|
|
||||||
|
Read this section first. It tells you exactly what to do.
|
||||||
|
|
||||||
|
**When to run:**
|
||||||
|
- Setting up a **new agent profile** for the first time
|
||||||
|
- An **existing profile** is missing web search, firecrawl, or MCP configs
|
||||||
|
- User says: `"apply the config template"` or `"fix my infra config"`
|
||||||
|
- User says: `"configure <agent_name> with <model>"`
|
||||||
|
|
||||||
|
**How to run (via OpenProse CLI):**
|
||||||
|
```bash
|
||||||
|
# Minimal — only agent name required
|
||||||
|
prose run hermes-config-template agent_name=syslog-code
|
||||||
|
|
||||||
|
# Full control
|
||||||
|
prose run hermes-config-template \
|
||||||
|
agent_name=syslog-devops \
|
||||||
|
default_model=claude-sonnet-4 \
|
||||||
|
auxiliary_model=gemma-4-12b \
|
||||||
|
fallback_model=deepseek-chat \
|
||||||
|
fallback_provider=deepseek
|
||||||
|
```
|
||||||
|
|
||||||
|
### How to Follow the Template (Manual)
|
||||||
|
|
||||||
|
If `prose` CLI is not available:
|
||||||
|
|
||||||
|
1. **Fetch this contract:** `curl -sL https://git.sysloggh.net/SyslogSolution/prose-contracts/raw/branch/master/hermes-config-template.prose.md`
|
||||||
|
2. **Open the target config.yaml** — the profile you're updating
|
||||||
|
3. **Compare** each of the 5 required sections below against what exists in the profile
|
||||||
|
4. **For each section:**
|
||||||
|
- If it says **SHARED INFRA — DO NOT CHANGE**: Copy the value exactly as shown
|
||||||
|
- If it says **USER CHOOSES**: Use the agent's assigned model/provider
|
||||||
|
- If missing entirely: Add the full section
|
||||||
|
5. **Verify** every endpoint responds (see Verification section at the end)
|
||||||
|
6. **Restart** the gateway
|
||||||
|
|
||||||
|
### Decision Tree
|
||||||
|
|
||||||
|
```
|
||||||
|
What kind of profile is this?
|
||||||
|
├── Main profile (~/.hermes/config.yaml)
|
||||||
|
│ └── Replace web.*, mcp_servers.*, compression.*, auxiliary.*, custom_providers.*
|
||||||
|
│ └── model.default + fallback: use current values (USER CHOOSES)
|
||||||
|
│
|
||||||
|
├── Worker profile (~/.hermes/profiles/<name>/config.yaml)
|
||||||
|
│ ├── Does it have web.search_backend?
|
||||||
|
│ │ ├── NO → Add the full web.* section (copy exactly)
|
||||||
|
│ │ └── YES → Verify it points to the right IP
|
||||||
|
│ ├── Does it have mcp_servers.ra-h-os?
|
||||||
|
│ │ ├── NO → Add mcp_servers.* section (copy exactly)
|
||||||
|
│ │ └── YES → Verify URL
|
||||||
|
│ └── model.default: USE the worker's assigned model (don't copy main)
|
||||||
|
│
|
||||||
|
└── Non-Syslog profile (personal/lifestyle agent)
|
||||||
|
└── Skip — use default Hermes config or Tanko's config instead
|
||||||
|
```
|
||||||
|
|
||||||
|
## Infrastructure Stack
|
||||||
|
|
||||||
|
This template provisions the following shared infrastructure. Every agent should point to the same endpoints unless explicitly overridden.
|
||||||
|
|
||||||
|
| Component | Endpoint | Purpose |
|
||||||
|
|---|---|---|
|
||||||
|
| Firecrawl | `http://192.168.68.7:3002/` | Web content extraction |
|
||||||
|
| SearXNG | `http://storepve:8888` | Privacy-respecting web search |
|
||||||
|
| LiteLLM | `http://litellm.sysloggh.net/litellm/v1` | Unified model gateway |
|
||||||
|
| RA-H OS MCP | `http://192.168.68.65:3100/mcp` | Knowledge graph bridge |
|
||||||
|
| Context7 MCP | `http://localhost:8079/mcp` | Documentation queries |
|
||||||
|
|
||||||
|
**API Key Rules:**
|
||||||
|
- `api_key: sk-_SW...B8nw` — Hardcoded LiteLLM master key for custom_providers and auxiliary tasks
|
||||||
|
- `api_key_env: DEEPSEEK_API_KEY` — Env-var based key for fallback provider
|
||||||
|
- Worker profiles should inherit proxy config via `harness` custom_provider, NOT hardcode LiteLLM keys
|
||||||
|
|
||||||
|
**For workers that need API keys overridden:** Set `api_key: ''` in the model section (inherits from custom_providers), or use `api_key_env` for env-based auth.
|
||||||
|
|
||||||
|
## Template Structure
|
||||||
|
|
||||||
|
### Required Sections (every agent MUST have)
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
# ─── Model Selection (USER CHOOSES) ───
|
||||||
|
model:
|
||||||
|
# !! CHANGE THIS for your agent !!
|
||||||
|
default: <agent_default_model>
|
||||||
|
provider: <agent_default_provider>
|
||||||
|
# LiteLLM base URL (shared infra)
|
||||||
|
base_url: http://litellm.sysloggh.net/litellm/v1
|
||||||
|
|
||||||
|
fallback_providers:
|
||||||
|
# !! OPTIONAL: overridable fallback !!
|
||||||
|
provider: <fallback_provider>
|
||||||
|
model: <fallback_model>
|
||||||
|
|
||||||
|
# ─── Web Stack (SHARED INFRA — DO NOT CHANGE) ───
|
||||||
|
web:
|
||||||
|
backend: firecrawl
|
||||||
|
search_backend: searxng
|
||||||
|
extract_backend: firecrawl
|
||||||
|
firecrawl:
|
||||||
|
base_url: http://192.168.68.7:3002/
|
||||||
|
|
||||||
|
# ─── MCP Servers (SHARED INFRA — DO NOT CHANGE) ───
|
||||||
|
mcp_servers:
|
||||||
|
context7:
|
||||||
|
connect_timeout: 60
|
||||||
|
timeout: 300
|
||||||
|
url: http://localhost:8079/mcp
|
||||||
|
ra-h-os:
|
||||||
|
url: http://192.168.68.65:3100/mcp
|
||||||
|
timeout: 120
|
||||||
|
connect_timeout: 60
|
||||||
|
|
||||||
|
# ─── Compression (SHARED — can override model) ───
|
||||||
|
compression:
|
||||||
|
enabled: true
|
||||||
|
provider: harness
|
||||||
|
model: <auxiliary_model> # default: gemma-4-12b
|
||||||
|
threshold: 0.5
|
||||||
|
target_ratio: 0.25
|
||||||
|
protect_last_n: 30
|
||||||
|
hygiene_hard_message_limit: 400
|
||||||
|
|
||||||
|
# ─── Auxiliary Tasks (SHARED INFRA + MODEL) ───
|
||||||
|
auxiliary:
|
||||||
|
vision:
|
||||||
|
provider: harness
|
||||||
|
model: <auxiliary_model>
|
||||||
|
web_extract:
|
||||||
|
provider: harness
|
||||||
|
model: <auxiliary_model>
|
||||||
|
session_search:
|
||||||
|
provider: harness
|
||||||
|
model: <auxiliary_model>
|
||||||
|
max_concurrency: 3
|
||||||
|
|
||||||
|
# ─── Custom Provider (SHARED INFRA — LiteLLM) ───
|
||||||
|
custom_providers:
|
||||||
|
- name: <custom_provider_name>
|
||||||
|
model: <custom_provider_model>
|
||||||
|
base_url: http://litellm.sysloggh.net/litellm/v1
|
||||||
|
api_key: <liteLLM_master_key>
|
||||||
|
api_mode: chat_completions
|
||||||
|
```
|
||||||
|
|
||||||
|
### Optional Sections (agent-specific)
|
||||||
|
|
||||||
|
- **Platform configs** (Telegram, Discord, WhatsApp bot tokens)
|
||||||
|
- **Display/skin** options (compact mode, personality)
|
||||||
|
- **Plugin lists** (zulip-platform, etc.)
|
||||||
|
- **Terminal settings** (container image, timeouts)
|
||||||
|
- **Security/approvals** (command allowlists)
|
||||||
|
- **Kanban/worker** settings (if this profile is a Kanban worker)
|
||||||
|
- **Skills auto_load** list
|
||||||
|
- **Timezone** override
|
||||||
|
|
||||||
|
## Continuity
|
||||||
|
|
||||||
|
Configuration drift detection is **event-driven**:
|
||||||
|
|
||||||
|
- **On agent onboarding**: Always scaffold from this template
|
||||||
|
- **On infra change**: If Firecrawl/SearXNG/LiteLLM endpoints change, update template + all profiles
|
||||||
|
- **On new model**: When a new model is deployed on Litellm, any agent can change their `default` independently
|
||||||
|
- **Periodic**: Every 2 weeks — check if any profile's infra section has drifted from the template
|
||||||
|
- **On connection failure**: If web search/extract fails, check if that agent has missing `web.*` config
|
||||||
|
|
||||||
|
## Success Criteria
|
||||||
|
|
||||||
|
The configuration is CONSIDERED GOOD when ALL of these pass:
|
||||||
|
|
||||||
|
### Infra Connectivity
|
||||||
|
- `web.backend == "firecrawl"` — Firecrawl backend configured
|
||||||
|
- `web.search_backend == "searxng"` — SearXNG search configured
|
||||||
|
- `web.extract_backend == "firecrawl"` — Firecrawl extraction configured
|
||||||
|
- `web.firecrawl.base_url` points to `192.168.68.7:3002`
|
||||||
|
- `mcp_servers.ra-h-os.url` points to `192.168.68.65:3100/mcp`
|
||||||
|
- All endpoints respond to `curl` health checks
|
||||||
|
|
||||||
|
### Model Flexibility
|
||||||
|
- `model.default` is settable per agent (not hardcoded)
|
||||||
|
- `model.provider` is settable per agent
|
||||||
|
- `custom_providers[0].model` matches the agent's workload (code vs general)
|
||||||
|
|
||||||
|
### Completeness
|
||||||
|
- `auxiliary.vision` configured (even if provider=auto)
|
||||||
|
- `compression.enabled` set (true for most agents)
|
||||||
|
- Fallback provider configured (deepseek or another harness model)
|
||||||
|
- `_config_version` correctly set (currently `30`)
|
||||||
|
|
||||||
|
### No Drift
|
||||||
|
- No duplicate infra endpoints (one canonical Firecrawl URL)
|
||||||
|
- No stale endpoints (old IPs, dead services)
|
||||||
|
- All profiles consistent on shared infra
|
||||||
|
|
||||||
|
## Configuration Rules
|
||||||
|
|
||||||
|
### Rule 1: Shared Infra Is Locked
|
||||||
|
The following MUST be identical across ALL profiles:
|
||||||
|
- `web.backend`, `web.search_backend`, `web.extract_backend`
|
||||||
|
- `web.firecrawl.base_url`
|
||||||
|
- `mcp_servers.ra-h-os.url`
|
||||||
|
- `custom_providers[0].base_url`
|
||||||
|
|
||||||
|
### Rule 2: Model Choice Is Free
|
||||||
|
The following are OWNED by each agent and can differ:
|
||||||
|
- `model.default` — the primary model
|
||||||
|
- `model.provider` — where it runs
|
||||||
|
- `fallback_providers.model` — backup model
|
||||||
|
- `custom_providers[0].model` — what this agent uses LiteLLM for
|
||||||
|
|
||||||
|
### Rule 3: Workers Inherit, Not Duplicate
|
||||||
|
Worker profiles (syslog-code, syslog-devops, etc.) generally inherit from the main config. Only override what's DIFFERENT:
|
||||||
|
- Different default model → override `model.default`
|
||||||
|
- Different auxiliary model → override auxiliary sections
|
||||||
|
- No web search needed → still keep the config (tools won't enable without it)
|
||||||
|
|
||||||
|
### Rule 4: API Key Placement
|
||||||
|
- **Custom provider keys** (`custom_providers[0].api_key`): Must be hardcoded (the LiteLLM master key or a virtual key)
|
||||||
|
- **Model section keys** (`model.api_key`): Prefer empty. Workers inherit from custom_providers.
|
||||||
|
- **Environment keys** (`api_key_env`): For fallback providers that need separate auth (DeepSeek, OpenRouter)
|
||||||
|
- **Hardcoded keys in profile**: If a worker's `model.api_key` is set, it overrides custom_providers. Only use when the worker needs a DIFFERENT provider than Litellm.
|
||||||
|
|
||||||
|
### Rule 5: Verify After Every Config Change
|
||||||
|
After changing a profile's config:
|
||||||
|
1. `curl` the model endpoint to confirm auth works
|
||||||
|
2. `curl` the web stack endpoints (Firecrawl, SearXNG)
|
||||||
|
3. Check that `hermes gateway --restart` reloads cleanly
|
||||||
|
|
||||||
|
## Execution
|
||||||
|
|
||||||
|
1. **Check current config** — Read the target agent's config.yaml
|
||||||
|
2. **Compare against template** — Identify missing or divergent sections
|
||||||
|
3. **Apply shared infra** — Lock the web/MCP/compression sections to template values
|
||||||
|
4. **Apply model choice** — Set default/fallback/auxiliary models per agent's workload
|
||||||
|
5. **Preserve agent-specific** — Keep platform configs, plugins, terminal settings, skills
|
||||||
|
6. **Remove stale** — Drop any old infra endpoints that don't match the template
|
||||||
|
7. **Verify** — curl all shared endpoints, test the model
|
||||||
|
8. **Report** — What was changed, what was preserved, what's still custom
|
||||||
|
|
||||||
|
## Example Output
|
||||||
|
|
||||||
|
```
|
||||||
|
## Config Template Applied: syslog-code ✅
|
||||||
|
|
||||||
|
### Shared Infrastructure (Locked)
|
||||||
|
| Section | Before | After |
|
||||||
|
|---|---|---|
|
||||||
|
| web.backend | firecrawl | firecrawl ✅ |
|
||||||
|
| web.search_backend | NOT SET | searxng ✅ |
|
||||||
|
| web.firecrawl.base_url | NOT SET | http://192.168.68.7:3002/ ✅ |
|
||||||
|
| mcp_servers.ra-h-os | present | present ✅ |
|
||||||
|
|
||||||
|
### Agent Model Choices (Preserved)
|
||||||
|
| Setting | Value |
|
||||||
|
|---|---|
|
||||||
|
| model.default | qwen3.6-27B-code (code agent) |
|
||||||
|
| fallback | gemma-4-12b (harness) |
|
||||||
|
| auxiliary | gemma-4-12b (harness) |
|
||||||
|
| compression | gemma-4-12b (harness) |
|
||||||
|
|
||||||
|
### Endpoint Verification
|
||||||
|
| Endpoint | Status |
|
||||||
|
|---|---|
|
||||||
|
| Firecrawl :3002 | ✅ 200 OK |
|
||||||
|
| SearXNG :8888 | ✅ 200 OK |
|
||||||
|
| LiteLLM | ✅ 200 OK |
|
||||||
|
| RA-H OS MCP | ✅ Connected |
|
||||||
|
```
|
||||||
@@ -0,0 +1,333 @@
|
|||||||
|
---
|
||||||
|
kind: pattern
|
||||||
|
name: infrastructure-control
|
||||||
|
description: >
|
||||||
|
Full infrastructure monitoring and control pattern covering the
|
||||||
|
5-node Proxmox cluster, 3 Docker ecosystems (22 containers),
|
||||||
|
NFS storage, and network services. Defines monitors, remediations,
|
||||||
|
and the access matrix for all environments.
|
||||||
|
---
|
||||||
|
|
||||||
|
# Infrastructure Control Pattern
|
||||||
|
|
||||||
|
## Topology
|
||||||
|
|
||||||
|
```
|
||||||
|
┌─────────────────────────┐
|
||||||
|
│ OpenProse Contract │
|
||||||
|
│ (declares what's true) │
|
||||||
|
└──────────┬──────────────┘
|
||||||
|
│
|
||||||
|
┌────────────────────┼────────────────────┐
|
||||||
|
▼ ▼ ▼
|
||||||
|
┌─────────────┐ ┌──────────────┐ ┌──────────────┐
|
||||||
|
│ Abiba │ │ Tanko │ │ Mumuni │
|
||||||
|
│ (pi) │ │ (Hermes) │ │ (Hermes) │
|
||||||
|
│ CT 100 │ │ CT 122 │ │ CT 114 │
|
||||||
|
└──────┬──────┘ └──────┬───────┘ └──────┬───────┘
|
||||||
|
│ │ │
|
||||||
|
└──────────────────┼────────────────────┘
|
||||||
|
▼
|
||||||
|
┌──────────────────────────────────────┐
|
||||||
|
│ Proxmox Cluster API │
|
||||||
|
│ minipve.sysloggh.net:443 │
|
||||||
|
│ (monitoring@pve!mumuni token) │
|
||||||
|
└────┬──────┬──────┬──────┬──────┬─────┘
|
||||||
|
│ │ │ │ │
|
||||||
|
┌────┘ ┌────┘ ┌────┘ ┌────┘ ┌────┘
|
||||||
|
▼ ▼ ▼ ▼ ▼
|
||||||
|
minipve amdpve storepve acerpve ocupve
|
||||||
|
(.12) (.15) (.6) (.9) (.5)
|
||||||
|
|
||||||
|
▼
|
||||||
|
┌─────────────────────────────────────────────┐
|
||||||
|
│ Docker Ecosystems │
|
||||||
|
│ ┌────────────┬────────────┬──────────────┐ │
|
||||||
|
│ │ docker-vm │ CT 116 │ Netbird │ │
|
||||||
|
│ │ (.7) │ (.116) │ (72.61.0.17) │ │
|
||||||
|
│ │ 11 ctrs │ 6 ctrs │ 5 ctrs │ │
|
||||||
|
│ └────────────┴────────────┴──────────────┘ │
|
||||||
|
└─────────────────────────────────────────────┘
|
||||||
|
```
|
||||||
|
|
||||||
|
## Section 1: Access Matrix
|
||||||
|
|
||||||
|
### API Tokens & Credentials
|
||||||
|
|
||||||
|
| Resource | Auth Method | Credential Source | Status |
|
||||||
|
|----------|------------|-------------------|--------|
|
||||||
|
| Proxmox Cluster | PVE API Token | `monitoring@pve!mumuni=...` | ✅ |
|
||||||
|
| Proxmox Root | Password via API ticket | `root@pam:kakashi19` | ✅ |
|
||||||
|
| docker-vm (.7) | SSH root | SSH key | ✅ |
|
||||||
|
| CT 116 (syslog-api) | SSH root | SSH key | ✅ |
|
||||||
|
| Tanko CT (.122) | SSH jerome | SSH key | ✅ |
|
||||||
|
| Netbird (.17) | SSH root | SSH key | ✅ |
|
||||||
|
| Gitea | API token | abiba-bot token | ✅ |
|
||||||
|
| Zulip | Bot API key | per-bot tokens | ✅ |
|
||||||
|
| RA-H OS | MCP bridge | port 3100 | ✅ |
|
||||||
|
|
||||||
|
### Reachability Matrix
|
||||||
|
|
||||||
|
| From / To | PVE API | docker-vm (.7) | CT 116 | Tanko (.122) | Netbird (.17) |
|
||||||
|
|-----------|---------|----------------|--------|-------------|---------------|
|
||||||
|
| **Abiba** (CT 100) | ✅ :443 | ✅ SSH | ✅ SSH | ✅ SSH | ❌ no Netbird |
|
||||||
|
| **Tanko** (CT 122) | ❌ | ❌ | ❌ | ✅ | ❌ |
|
||||||
|
| **docker-vm** (.7) | ❌ | ✅ | ❌ | ❌ | ❌ |
|
||||||
|
|
||||||
|
**Conclusion:** Only Abiba has cross-infrastructure access. All monitoring contracts run from Abiba.
|
||||||
|
|
||||||
|
## Section 2: Proxmox Cluster — Monitoring
|
||||||
|
|
||||||
|
### Nodes (5)
|
||||||
|
|
||||||
|
| Node | IP | CPU | RAM | VMs/CTs | Role |
|
||||||
|
|------|----|-----|-----|---------|------|
|
||||||
|
| minipve | .12 | 16C | 30GB | authentik, gitea, mumuni, syslog-api, jitsi | Auth, git, messaging |
|
||||||
|
| amdpve | .15 | 32C | 62GB | abiba, kagentz, tanko, tdunna, baggy, scottdenya | Agents, compute |
|
||||||
|
| storepve | .6 | 28C | 31GB | docker-vm, ra-h-os, PBS, media, zulip | Docker, storage, chat |
|
||||||
|
| acerpve | .9 | 28C | 31GB | llm-gpu, adguard | GPU VMs |
|
||||||
|
| ocupve | .5 | 12C | 14GB | ocu-llm | GPU VMs |
|
||||||
|
|
||||||
|
### Checks (every 5 min)
|
||||||
|
|
||||||
|
```
|
||||||
|
## Maintains
|
||||||
|
|
||||||
|
- cluster-status: { online_nodes: int, offline_nodes: int, timestamp }
|
||||||
|
- node-cpu-usage: { node: pct, warnings: [] }
|
||||||
|
- node-memory-usage: { node: free_pct, warnings: [] }
|
||||||
|
- node-uptime: { node: seconds, just_rebooted: bool }
|
||||||
|
|
||||||
|
## Checks
|
||||||
|
|
||||||
|
- For each node in [minipve, amdpve, storepve, acerpve, ocupve]:
|
||||||
|
- GET /api2/json/nodes/{node}/status → check status == "online"
|
||||||
|
- GET /api2/json/nodes/{node}/status → cpu < 0.80
|
||||||
|
- GET /api2/json/nodes/{node}/status → free_mem > 10%
|
||||||
|
- GET /api2/json/nodes/{node}/status → uptime > 300 (warn if just rebooted)
|
||||||
|
|
||||||
|
- For each CT in inventory (19 total):
|
||||||
|
- GET /api2/json/cluster/resources → filter by type=lxc
|
||||||
|
- Warn if status != "running"
|
||||||
|
|
||||||
|
- Storage pools:
|
||||||
|
- GET /api2/json/nodes/storepve/storage/mediastore → used < 80%
|
||||||
|
- GET /api2/json/nodes/storepve/storage/pbs-backup → last backup < 48h
|
||||||
|
|
||||||
|
## Remediations
|
||||||
|
|
||||||
|
- Node offline → alert via Zulip DM, escalate after 3x
|
||||||
|
- CT stopped → `pct start <vmid>` via PVE API, verify after 30s
|
||||||
|
- VM stopped → `qm start <vmid>` via PVE API, verify after 60s
|
||||||
|
- Storage > 80% → warn via Zulip DM
|
||||||
|
- Storage > 95% → crit via Zulip + relay to maintainer
|
||||||
|
```
|
||||||
|
|
||||||
|
## Section 3: Docker Ecosystems — Monitoring
|
||||||
|
|
||||||
|
### Ecosystem A: docker-vm (192.168.68.7)
|
||||||
|
|
||||||
|
11 containers across 4 compose stacks:
|
||||||
|
|
||||||
|
| Stack | Path | Containers |
|
||||||
|
|-------|------|-----------|
|
||||||
|
| **Firecrawl** | `/opt/search-stack/firecrawl-source/` | api, rabbitmq, postgres, playwright, redis |
|
||||||
|
| **SearXNG** | `/opt/search-stack/searxng/` | searxng, valkey |
|
||||||
|
| **Home stack** | `/opt/home_stack/` | jdownloader, bentopdf, pulse |
|
||||||
|
| **Audiobookshelf** | `/opt/audiobookshelf/` | audiobookshelf |
|
||||||
|
|
||||||
|
### Ecosystem B: CT 116 syslog-api (192.168.68.116)
|
||||||
|
|
||||||
|
| Container | Image | Port |
|
||||||
|
|-----------|-------|------|
|
||||||
|
| harness-litellm | berriai/litellm:1.90.0-rc.1 | :4001 |
|
||||||
|
| harness-nginx | nginx:alpine | :80 |
|
||||||
|
| harness-router | inference-harness-router | :9000 |
|
||||||
|
| harness-postgres | postgres:16-alpine | :5432 |
|
||||||
|
| harness-redis | redis:7-alpine | :6379 |
|
||||||
|
| harness-dashboard | inference-harness-dashboard | :3000 |
|
||||||
|
|
||||||
|
### Ecosystem C: Netbird (72.61.0.17)
|
||||||
|
|
||||||
|
| Container | Image | Role |
|
||||||
|
|-----------|-------|------|
|
||||||
|
| netbird-server | netbird | VPN controller |
|
||||||
|
| netbird-dashboard | netbird UI | Web management |
|
||||||
|
| netbird-proxy | nginx | TLS termination |
|
||||||
|
| netbird-crowdsec | crowdsecurity/crowdsec:v1.7.7 | WAF |
|
||||||
|
| netbird-traefik | traefik | Reverse proxy |
|
||||||
|
|
||||||
|
### Checks (every 60s)
|
||||||
|
|
||||||
|
```
|
||||||
|
## Maintains
|
||||||
|
|
||||||
|
- docker-health: { ecosystem: string, healthy: int, unhealthy: int, total: int }
|
||||||
|
- container-status: { name: string, status: string, restarts: int, image: string }
|
||||||
|
|
||||||
|
## Checks
|
||||||
|
|
||||||
|
For each Docker host:
|
||||||
|
- docker ps --format "{{.Names}}" → every required container is running
|
||||||
|
- docker inspect <name> --format "{{.State.Health.Status}}" → "healthy"
|
||||||
|
- docker inspect <name> --format "{{.RestartCount}}" → < 3/hour
|
||||||
|
- df -h / | awk '{print $5}' → usage < 80%
|
||||||
|
|
||||||
|
For docker-vm specifically:
|
||||||
|
- mountpoint -q /media/storage → NFS mounted
|
||||||
|
- mountpoint -q /media/mediastore → NFS mounted
|
||||||
|
- docker compose ls → expected stacks present
|
||||||
|
|
||||||
|
## Remediations
|
||||||
|
|
||||||
|
- Container unhealthy → `docker compose up -d <name>`, wait 10s, re-check
|
||||||
|
- Container missing → `docker compose up -d` in its stack directory
|
||||||
|
- Docker daemon down → `systemctl restart docker`
|
||||||
|
- NFS mount lost → `mount -a`, if fails → alert (storepve issue)
|
||||||
|
- Disk > 80% → alert via Zulip DM
|
||||||
|
- Disk > 95% → crit + relay to maintainer
|
||||||
|
```
|
||||||
|
|
||||||
|
## Section 4: Storage — Monitoring
|
||||||
|
|
||||||
|
### NFS Mounts (docker-vm → storepve)
|
||||||
|
|
||||||
|
| Mount | Export | Capacity | Used | Alert |
|
||||||
|
|-------|--------|----------|------|-------|
|
||||||
|
| /media/storage | storepve:/media/storage | 3.6TB | 264GB (8%) | None |
|
||||||
|
| /media/mediastore | storepve:/media/mediastore | 7.3TB | 5.2TB (75%) | ⚠️ warn at 80% |
|
||||||
|
|
||||||
|
### Proxmox Storage Backends
|
||||||
|
|
||||||
|
| Storage | Type | Content | Active |
|
||||||
|
|---------|------|---------|--------|
|
||||||
|
| local | dir | ISO, vztmpl, backup | ✅ |
|
||||||
|
| local-lvm | lvmthin | images, rootdir | ✅ |
|
||||||
|
| zfs-vm | nfs | rootdir, images | ✅ |
|
||||||
|
| mediastore | dir | images, backup, rootdir | ✅ |
|
||||||
|
| zfs-iso | nfs | vztmpl, iso | ✅ |
|
||||||
|
| pbs-backup | pbs | backup | ✅ |
|
||||||
|
|
||||||
|
### Checks (every 10 min)
|
||||||
|
|
||||||
|
```
|
||||||
|
## Maintains
|
||||||
|
|
||||||
|
- storage-usage: { mount: string, used_pct: float, growth_rate: float }
|
||||||
|
- backup-status: { last_success: timestamp, age_hours: int }
|
||||||
|
- snapshot-status: { pool: string, newest_age_hours: int }
|
||||||
|
|
||||||
|
## Specific Alerts
|
||||||
|
|
||||||
|
- mediastore growth rate > 10GB/day → warn (check what's writing)
|
||||||
|
- mediastore > 90% → crit (only 730GB remaining)
|
||||||
|
- No PBS backup in 48h → fail
|
||||||
|
```
|
||||||
|
|
||||||
|
## Section 5: Network Services — Monitoring
|
||||||
|
|
||||||
|
| Service | Domain | Status |
|
||||||
|
|---------|--------|--------|
|
||||||
|
| Proxmox API | minipve.sysloggh.net:443 | ✅ |
|
||||||
|
| Authentik | auth.sysloggh.net:443 | ✅ |
|
||||||
|
| Gitea | git.sysloggh.net:443 | ✅ |
|
||||||
|
| Zulip | chat.sysloggh.net:443 | ✅ |
|
||||||
|
| LiteLLM | litellm.sysloggh.net:443 | ✅ |
|
||||||
|
| Pulse | pulse.sysloggh.net:443 | ✅ |
|
||||||
|
| SearXNG | searxng.sysloggh.net:8888 | ✅ |
|
||||||
|
| Firecrawl | firecrawl.sysloggh.net:3002 | ✅ |
|
||||||
|
|
||||||
|
### Checks (every 2 min)
|
||||||
|
|
||||||
|
```
|
||||||
|
## Maintains
|
||||||
|
|
||||||
|
- dns-resolution: { services: [{name, resolves}] }
|
||||||
|
- ssl-expiry: { services: [{name, days_remaining}] }
|
||||||
|
- endpoint-reachability: { services: [{name, http_code}] }
|
||||||
|
```
|
||||||
|
|
||||||
|
## Section 6: Alert Routing & Escalation
|
||||||
|
|
||||||
|
### Severity Levels
|
||||||
|
|
||||||
|
| Severity | Channel | Format | Rate Limit |
|
||||||
|
|----------|---------|--------|-----------|
|
||||||
|
| **warn** | Zulip DM to owner | "⚠️ {check}: {detail}" | 1x/check/hour |
|
||||||
|
| **crit** | Zulip DM + #agent-hub | "🚨 {check}: {detail}" | Immediate |
|
||||||
|
| **escalated** | Zulip DM + relay to maintainer | "🔥 {check}: {detail}" | Immediate |
|
||||||
|
|
||||||
|
### Escalation Chain
|
||||||
|
|
||||||
|
```
|
||||||
|
Container unhealthy (3 consecutive failures)
|
||||||
|
→ Alert owner
|
||||||
|
→ 15 min no response → Relay to maintainer
|
||||||
|
|
||||||
|
Node offline
|
||||||
|
→ Alert owner
|
||||||
|
→ 5 min → Attempt restart via PVE API
|
||||||
|
→ 2 failures → Escalate
|
||||||
|
|
||||||
|
Storage > 95%
|
||||||
|
→ Alert owner + maintainer relay
|
||||||
|
→ Immediate action required
|
||||||
|
|
||||||
|
SSL cert < 7 days
|
||||||
|
→ Alert owner daily
|
||||||
|
→ 3 days out → Escalate
|
||||||
|
|
||||||
|
ZFS pool degraded
|
||||||
|
→ Cannot auto-fix → Escalate immediately
|
||||||
|
```
|
||||||
|
|
||||||
|
## Appendix A: Quick Health Commands
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Full cluster status
|
||||||
|
PVE="https://minipve.sysloggh.net"
|
||||||
|
AUTH="Authorization: PVEAPIToken=monitoring@pve!mumuni=eafd56c5-93d4-4d40-a41d-e688be0987f3"
|
||||||
|
curl -sfk "$PVE/api2/json/cluster/resources" -H "$AUTH"
|
||||||
|
|
||||||
|
# Docker health from Abiba
|
||||||
|
ssh root@192.168.68.7 "docker ps --format '{{.Names}} {{.Status}}'"
|
||||||
|
ssh root@192.168.68.116 "docker ps --format '{{.Names}} {{.Status}}'"
|
||||||
|
|
||||||
|
# Storage check
|
||||||
|
ssh root@192.168.68.7 "df -h /media/storage /media/mediastore"
|
||||||
|
```
|
||||||
|
|
||||||
|
## Appendix B: CT Inventory
|
||||||
|
|
||||||
|
| CT | Name | Node | IP | Role | Agent |
|
||||||
|
|----|------|------|----|------|-------|
|
||||||
|
| 100 | abiba | amdpve | .24 | Pi agent (this host) | ✅ pi |
|
||||||
|
| 101 | llm-gpu | acerpve | — | GPU VM | ❌ |
|
||||||
|
| 102 | adguard | acerpve | — | DNS | ❌ |
|
||||||
|
| 103 | ocu-llm | ocupve | — | GPU VLM | ❌ |
|
||||||
|
| 104 | authentik | minipve | .11 | OIDC | ❌ |
|
||||||
|
| 105 | kagentz | amdpve | — | Agent Zero | ✅ |
|
||||||
|
| 106 | ra-h-os | storepve | .65 | KG bridge | ✅ MCP |
|
||||||
|
| 107 | pbs | storepve | — | Backups | ❌ |
|
||||||
|
| 108 | media | storepve | — | Media | ❌ |
|
||||||
|
| 109 | docker-vm | storepve | .7 | Docker host | ❌ |
|
||||||
|
| 110 | gitea | minipve | — | Git | ❌ |
|
||||||
|
| 111 | tdunna | amdpve | — | ? | ❌ |
|
||||||
|
| 112 | tanko | amdpve | .122 | Hermes agent | ✅ |
|
||||||
|
| 113 | baggy | amdpve | — | ? | ❌ |
|
||||||
|
| 114 | mumuni | minipve | — | Hermes agent | ✅ |
|
||||||
|
| 115 | scottdenya | amdpve | — | ? | ❌ |
|
||||||
|
| 116 | syslog-api | minipve | .116 | LiteLLM stack | ❌ |
|
||||||
|
| 117 | zulip | storepve | — | Chat | ❌ |
|
||||||
|
| 118 | jitsi | minipve | — | Video | ❌ |
|
||||||
|
|
||||||
|
## Appendix C: Docker Compose Files Location
|
||||||
|
|
||||||
|
| Host | Stack | Compose File |
|
||||||
|
|------|-------|-------------|
|
||||||
|
| docker-vm (.7) | Firecrawl | `/opt/search-stack/firecrawl-source/docker-compose.yaml` |
|
||||||
|
| docker-vm (.7) | SearXNG | `/opt/search-stack/searxng/docker-compose.yml` |
|
||||||
|
| docker-vm (.7) | Home stack | `/opt/home_stack/docker-compose.yml` |
|
||||||
|
| docker-vm (.7) | Audiobookshelf | `/opt/audiobookshelf/docker-compose.yml` |
|
||||||
|
| CT 116 (.116) | LiteLLM | `/root/docker-compose-litellm.yml` |
|
||||||
|
| Netbird (.17) | Netbird | Docker run (not compose) |
|
||||||
@@ -0,0 +1,56 @@
|
|||||||
|
---
|
||||||
|
kind: function
|
||||||
|
name: check-litellm-health
|
||||||
|
description: >
|
||||||
|
Verifies LiteLLM deployment is healthy by checking admin UI, API docs,
|
||||||
|
OIDC auth endpoint, container status, and aggregate health on the
|
||||||
|
backend host. Designed as a reusable contract for any Syslog agent.
|
||||||
|
---
|
||||||
|
|
||||||
|
## Parameters
|
||||||
|
|
||||||
|
- public_url: string — The public LiteLLM URL (default: "https://litellm.sysloggh.net")
|
||||||
|
- backend_host: string — Internal CT host to check containers (default: "192.168.68.116")
|
||||||
|
- auth_host: string — Authentik server for OIDC (default: "192.168.68.11")
|
||||||
|
|
||||||
|
## Returns
|
||||||
|
|
||||||
|
- overall_status: "healthy" | "degraded" | "down"
|
||||||
|
- checks: array of { name: string, status: string, detail: string }
|
||||||
|
- timestamp: string — ISO timestamp of the check run
|
||||||
|
- duration_ms: number — How long the check took
|
||||||
|
|
||||||
|
## Requires
|
||||||
|
|
||||||
|
- SSH key access to backend_host for container checks
|
||||||
|
- Network access to public_url and auth_host
|
||||||
|
- curl and openssl available on the execution host
|
||||||
|
|
||||||
|
## Ensures
|
||||||
|
|
||||||
|
- Each check returns a clear pass/fail status with detail message
|
||||||
|
- If any endpoint returns non-200, overall_status is "degraded"
|
||||||
|
- If backend host unreachable or >2 containers down, overall_status is "down"
|
||||||
|
- If /health/unified reports any non-healthy components, status reflects it
|
||||||
|
- Checks cover at minimum: admin UI, API docs, OIDC, containers, unified health, nginx proxy
|
||||||
|
|
||||||
|
## Execution
|
||||||
|
|
||||||
|
1. **Read parameters** — Use provided values or defaults
|
||||||
|
2. **Check public endpoints**:
|
||||||
|
- GET {{public_url}}/ui/ → expect 200 ("LiteLLM Dashboard")
|
||||||
|
- GET {{public_url}}/docs → expect 200 ("LiteLLM API - Swagger UI")
|
||||||
|
- GET {{public_url}}/openapi.json → expect 200 (valid JSON, 497 paths)
|
||||||
|
- GET {{public_url}}/redoc → expect 200 ("LiteLLM API - ReDoc")
|
||||||
|
3. **Check nginx-proxied endpoints** (internal only — Netbird routes to LiteLLM directly):
|
||||||
|
- GET http://{{backend_host}}/litellm/ui/ → expect 200
|
||||||
|
- GET http://{{backend_host}}/litellm/docs → expect 200
|
||||||
|
4. **Check aggregate health endpoint**:
|
||||||
|
- GET {{backend_host}}:9000/health/unified → expect 200, check all components
|
||||||
|
5. **Check backend container health**:
|
||||||
|
- SSH to {{backend_host}} → `docker ps` → verify all 6 containers are healthy
|
||||||
|
- Containers: harness-litellm, harness-nginx, harness-router, harness-postgres, harness-redis, harness-dashboard
|
||||||
|
6. **Check OIDC auth endpoint**:
|
||||||
|
- Verify auth.sysloggh.net resolves to {{auth_host}}
|
||||||
|
- GET https://auth.sysloggh.net/ → expect login page
|
||||||
|
7. **Compile and report** — Determine overall_status from individual check results
|
||||||
@@ -0,0 +1,126 @@
|
|||||||
|
---
|
||||||
|
kind: responsibility
|
||||||
|
name: litellm-self-heal
|
||||||
|
description: >
|
||||||
|
Standing responsibility that monitors LiteLLM health and proactively
|
||||||
|
fixes common issues. Reports every action via Zulip DM and RA-H OS
|
||||||
|
knowledge graph for full audit trail.
|
||||||
|
---
|
||||||
|
|
||||||
|
## Maintains
|
||||||
|
|
||||||
|
- litellm-admin-ui: { status: "healthy", last_check: timestamp }
|
||||||
|
- litellm-api-docs: { status: "healthy", last_check: timestamp }
|
||||||
|
- litellm-containers: { status: "healthy", last_check: timestamp }
|
||||||
|
- litellm-oidc: { status: "healthy", last_check: timestamp }
|
||||||
|
|
||||||
|
## Requires
|
||||||
|
|
||||||
|
- litellm-health:function
|
||||||
|
|
||||||
|
## Continuity
|
||||||
|
|
||||||
|
- Self-driven: check every 300 seconds
|
||||||
|
- Also wakes on user request
|
||||||
|
- On failure: re-check after 30s, escalate after 3 consecutive failures
|
||||||
|
|
||||||
|
## Reporting
|
||||||
|
|
||||||
|
Every remediation cycle produces a structured report in three channels:
|
||||||
|
|
||||||
|
### 1. RA-H OS Knowledge Graph Node
|
||||||
|
Created as `[LEARN] litellm-self-heal: <run_id>` with:
|
||||||
|
- metadata: { type: "remediation", status: "fixed" | "escalated" | "healthy" }
|
||||||
|
- source: Full JSON report of the cycle
|
||||||
|
- description: Summary of what was found, fixed, and escalated
|
||||||
|
|
||||||
|
### 2. Zulip DM to Owner
|
||||||
|
Sent immediately for:
|
||||||
|
- `issues_fixed > 0` — "🛠 LiteLLM Self-Heal — Fix Applied"
|
||||||
|
- `issues_escalated > 0` — "⚠ LiteLLM Self-Heal — Needs Your Attention"
|
||||||
|
- Every 10th clean cycle — "✅ All Clear (10 checks passed)"
|
||||||
|
|
||||||
|
### 3. Daily Digest (end of day)
|
||||||
|
A summary of the last 24 hours sent as a single Zulip DM:
|
||||||
|
```
|
||||||
|
📋 LiteLLM Self-Heal — Daily Digest (2026-06-26)
|
||||||
|
|
||||||
|
Total cycles: 288 (every 5 min)
|
||||||
|
Issues found: 3
|
||||||
|
├─ Fixed automatically: 3 (nginx restart x2, docs fix x1)
|
||||||
|
└─ Escalated: 0
|
||||||
|
|
||||||
|
Top actions:
|
||||||
|
• harness-nginx restarted — 2026-06-26 02:15
|
||||||
|
• DOCS_URL corrected — 2026-06-26 07:30
|
||||||
|
• harness-nginx restarted — 2026-06-26 14:46
|
||||||
|
|
||||||
|
Uptime: 23h 47m — All healthy now ✅
|
||||||
|
```
|
||||||
|
|
||||||
|
### 4. Weekly Digest (every Monday)
|
||||||
|
Same format as daily but covering 7 days. Sent to both Zulip DM and
|
||||||
|
logged as a `[REPORT]` knowledge graph node for long-term trending.
|
||||||
|
|
||||||
|
### 5. Relay Message (if cross-agent)
|
||||||
|
If a fix required another agent's help (e.g., Authentik restart), a relay
|
||||||
|
message is sent to the responsible agent with full context.
|
||||||
|
|
||||||
|
## Remediation Rules
|
||||||
|
|
||||||
|
### Rule 1: Container Not Healthy
|
||||||
|
Detect → `docker compose up -d <container>` → verify → log
|
||||||
|
Escalate after 3 failures
|
||||||
|
|
||||||
|
### Rule 2: Admin UI Non-200
|
||||||
|
Detect → restart harness-litellm → verify → log
|
||||||
|
Escalate after 2 failures
|
||||||
|
|
||||||
|
### Rule 3: Auth Unreachable
|
||||||
|
Detect → immediate escalate (external dependency)
|
||||||
|
|
||||||
|
### Rule 4: Docs 404
|
||||||
|
Detect → fix DOCS_URL env var → verify → log
|
||||||
|
|
||||||
|
## Execution
|
||||||
|
|
||||||
|
1. Run health check
|
||||||
|
2. Apply remediation for each failure
|
||||||
|
3. **Generate report** with all actions taken
|
||||||
|
4. **Log to knowledge graph** — create `[LEARN]` node
|
||||||
|
5. **Notify user** via Zulip DM if anything changed
|
||||||
|
6. Wait 300s and repeat
|
||||||
|
|
||||||
|
## Audit Trail Format
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"run_id": "self-heal-20260626-001",
|
||||||
|
"timestamp": "2026-06-26T14:00:00Z",
|
||||||
|
"duration_ms": 1234,
|
||||||
|
"checks_passed": 7,
|
||||||
|
"checks_failed": 0,
|
||||||
|
"issues_found": 0,
|
||||||
|
"issues_fixed": 0,
|
||||||
|
"issues_escalated": 0,
|
||||||
|
"actions": []
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
With failures:
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"run_id": "self-heal-20260626-002",
|
||||||
|
"issues_found": 1,
|
||||||
|
"issues_fixed": 1,
|
||||||
|
"actions": [
|
||||||
|
{
|
||||||
|
"rule": "Container Not Healthy",
|
||||||
|
"target": "harness-nginx",
|
||||||
|
"action": "restarted container",
|
||||||
|
"result": "healthy",
|
||||||
|
"duration_ms": 5000
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
```
|
||||||
@@ -0,0 +1,78 @@
|
|||||||
|
---
|
||||||
|
kind: responsibility
|
||||||
|
name: pm2-self-heal
|
||||||
|
description: >
|
||||||
|
Monitors critical PM2 processes (abiba-zulip, abiba-telegram) and
|
||||||
|
auto-restarts any that are stopped or errored. Logs every action to
|
||||||
|
the knowledge graph and alerts the owner via Zulip DM on failures.
|
||||||
|
CRITICAL: Never restart abiba-zulip — it runs this contract.
|
||||||
|
---
|
||||||
|
|
||||||
|
## Maintains
|
||||||
|
|
||||||
|
- abiba-zulip: { status: "online", uptime: string, restarts: number }
|
||||||
|
- abiba-telegram: { status: "online", uptime: string, restarts: number }
|
||||||
|
- last_check: timestamp
|
||||||
|
|
||||||
|
## Continuity
|
||||||
|
|
||||||
|
- Self-driven: check every 300 seconds (5 min)
|
||||||
|
- Also wakes on user request (`prose run pm2-self-heal`)
|
||||||
|
|
||||||
|
## Remediation Rules
|
||||||
|
|
||||||
|
### Rule 1: Process Stopped or Errored
|
||||||
|
- **Detect**: `pm2 status` shows "stopped" or "errored" for a named process
|
||||||
|
- **Fix**: `pm2 restart <name>`
|
||||||
|
- **Verify**: Re-check status after 5 seconds
|
||||||
|
- **Escalate**: If still failed after 2 retries, send Zulip DM to owner
|
||||||
|
|
||||||
|
### Rule 2: Process Restarting Too Often
|
||||||
|
- **Detect**: `pm2 status` shows restarts > 5 in the last hour
|
||||||
|
- **Fix**: `pm2 delete <name> && pm2 start <ecosystem> --only <name>`
|
||||||
|
- **Escalate**: Always — alert owner with restart count
|
||||||
|
|
||||||
|
## Execution
|
||||||
|
|
||||||
|
1. **Check PM2 status** — Run `pm2 status --no-color` and parse the table (5th data column = PID, 8th = restarts, 9th = status)
|
||||||
|
2. **Check abiba-telegram**:
|
||||||
|
- If status is "online" → pass
|
||||||
|
- If status is "stopped" or "errored" → apply Rule 1
|
||||||
|
- If restarts > 5 → alert owner
|
||||||
|
3. **Check abiba-zulip** (self-process, read-only):
|
||||||
|
- If status is "online" → pass, log restarts count
|
||||||
|
- If status is "stopped" or "errored" → **DO NOT RESTART** — alert owner immediately
|
||||||
|
- If restarts > 5 in last hour → alert owner with full diagnostics
|
||||||
|
4. **Log results** — Create `[LEARN]` node in knowledge graph for any actions taken
|
||||||
|
5. **Alert** — Send Zulip DM to owner if escalation needed (do NOT run pm2 commands during alerting)
|
||||||
|
6. **Wait 5 min** → repeat from step 1
|
||||||
|
|
||||||
|
## Example Output (when healthy)
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"run_id": "pm2-heal-001",
|
||||||
|
"checked_at": "2026-06-26T15:00:00Z",
|
||||||
|
"processes": {
|
||||||
|
"abiba-zulip": { "status": "online", "uptime": "2h", "restarts": 0 },
|
||||||
|
"abiba-telegram": { "status": "online", "uptime": "48h", "restarts": 0 }
|
||||||
|
},
|
||||||
|
"actions_taken": [],
|
||||||
|
"overall": "healthy"
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
## Example Output (when fixed)
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"processes": {
|
||||||
|
"abiba-telegram": { "status": "errored → restarted → online" },
|
||||||
|
"abiba-zulip": { "status": "online", "uptime": "2h", "restarts": 0 }
|
||||||
|
},
|
||||||
|
"actions_taken": [
|
||||||
|
{ "target": "abiba-telegram", "action": "pm2 restart", "result": "online" }
|
||||||
|
],
|
||||||
|
"overall": "fixed"
|
||||||
|
}
|
||||||
|
```
|
||||||
Executable
+64
@@ -0,0 +1,64 @@
|
|||||||
|
#!/bin/bash
|
||||||
|
# pm2-self-heal — hourly PM2 process check
|
||||||
|
# Part of the pm2-self-heal prose contract
|
||||||
|
# Only sends Zulip DM when action is taken or escalation needed
|
||||||
|
# Field positions (awk -F'│'): $7=pid $8=uptime $9=restarts $10=status
|
||||||
|
|
||||||
|
ZULIP_BOT="abiba-bot@chat.sysloggh.net"
|
||||||
|
ZULIP_KEY="cKTDMZAPW08dk3zl05sStzO7HRztzyn8"
|
||||||
|
ZULIP_SITE="https://chat.sysloggh.net"
|
||||||
|
OWNER_ID=9
|
||||||
|
LOG="/tmp/pm2-self-heal.log"
|
||||||
|
|
||||||
|
send_alert() {
|
||||||
|
local subject="$1"
|
||||||
|
local body="$2"
|
||||||
|
curl -s -X POST "$ZULIP_SITE/api/v1/messages" \
|
||||||
|
-u "$ZULIP_BOT:$ZULIP_KEY" \
|
||||||
|
-d "type=private" \
|
||||||
|
-d "to=[$OWNER_ID]" \
|
||||||
|
-d "content=$subject\n$body" > /dev/null 2>&1
|
||||||
|
echo "[$(date '+%Y-%m-%d %H:%M:%S')] Alert sent" >> "$LOG"
|
||||||
|
}
|
||||||
|
|
||||||
|
# Parse PM2 status (--no-color to avoid ANSI escape codes in awk columns)
|
||||||
|
STATUS=$(pm2 status --no-color 2>/dev/null)
|
||||||
|
|
||||||
|
ALERTS=""
|
||||||
|
|
||||||
|
# Check abiba-telegram (safe to auto-restart)
|
||||||
|
TEL_LINE=$(echo "$STATUS" | grep "abiba-telegram")
|
||||||
|
TEL_STATUS=$(echo "$TEL_LINE" | awk -F'│' '{print $10}' | xargs)
|
||||||
|
TEL_RESTARTS=$(echo "$TEL_LINE" | awk -F'│' '{print $9}' | xargs)
|
||||||
|
|
||||||
|
if [ "$TEL_STATUS" != "online" ]; then
|
||||||
|
pm2 restart abiba-telegram > /dev/null 2>&1
|
||||||
|
sleep 3
|
||||||
|
TEL_LINE2=$(pm2 status --no-color 2>/dev/null | grep "abiba-telegram")
|
||||||
|
TEL_STATUS2=$(echo "$TEL_LINE2" | awk -F'│' '{print $10}' | xargs)
|
||||||
|
if [ "$TEL_STATUS2" = "online" ]; then
|
||||||
|
ALERTS="${ALERTS}⚠️ abiba-telegram was **$TEL_STATUS** → restarted to online\n"
|
||||||
|
else
|
||||||
|
ALERTS="${ALERTS}🚨 abiba-telegram **failed restart** (was $TEL_STATUS, still $TEL_STATUS2)\n"
|
||||||
|
fi
|
||||||
|
fi
|
||||||
|
|
||||||
|
# Check abiba-zulip (read-only — never restart)
|
||||||
|
ZUL_LINE=$(echo "$STATUS" | grep "abiba-zulip")
|
||||||
|
ZUL_STATUS=$(echo "$ZUL_LINE" | awk -F'│' '{print $10}' | xargs)
|
||||||
|
ZUL_RESTARTS=$(echo "$ZUL_LINE" | awk -F'│' '{print $9}' | xargs)
|
||||||
|
|
||||||
|
if [ "$ZUL_STATUS" != "online" ]; then
|
||||||
|
ALERTS="${ALERTS}🚨 **abiba-zulip is $ZUL_STATUS** — needs investigation!\n"
|
||||||
|
ALERTS="${ALERTS} NOT auto-restarting (runs this contract)\n"
|
||||||
|
elif [ "$ZUL_RESTARTS" -gt 5 ] 2>/dev/null; then
|
||||||
|
ALERTS="${ALERTS}⚠️ abiba-zulip has **$ZUL_RESTARTS restarts** — may need attention\n"
|
||||||
|
fi
|
||||||
|
|
||||||
|
# Log check
|
||||||
|
echo "[$(date '+%Y-%m-%d %H:%M:%S')] tel=$TEL_STATUS zul=$ZUL_STATUS alerts=${ALERTS:+yes}" >> "$LOG"
|
||||||
|
|
||||||
|
# Only DM when there's something to report
|
||||||
|
if [ -n "$ALERTS" ]; then
|
||||||
|
send_alert "**PM2 Self-Heal — $(date '+%H:%M UTC')**" "$ALERTS"
|
||||||
|
fi
|
||||||
Executable
+708
@@ -0,0 +1,708 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""
|
||||||
|
/root/scripts/daily-infra-report.py — v2.0.0
|
||||||
|
Daily infrastructure dashboard emailed to jerome@sysloggh.com
|
||||||
|
Runs at 6:00 AM daily via cron.
|
||||||
|
Integrated with prose contracts: infrastructure-control, litellm-health, zulip-health.
|
||||||
|
|
||||||
|
Usage:
|
||||||
|
python3 daily-infra-report.py # Generate and email report
|
||||||
|
python3 daily-infra-report.py --test-email # Send a test email for review
|
||||||
|
python3 daily-infra-report.py --json # Output raw JSON to stdout
|
||||||
|
"""
|
||||||
|
|
||||||
|
import smtplib, json, subprocess, os, sys, datetime, re
|
||||||
|
from email.mime.text import MIMEText
|
||||||
|
from email.mime.multipart import MIMEMultipart
|
||||||
|
|
||||||
|
PVE = "https://minipve.sysloggh.net"
|
||||||
|
AUTH = "Authorization: PVEAPIToken=monitoring@pve!mumuni=eafd56c5-93d4-4d40-a41d-e688be0987f3"
|
||||||
|
|
||||||
|
# ── Shared credentials —─
|
||||||
|
|
||||||
|
ZULIP_SITE = "https://chat.sysloggh.net"
|
||||||
|
ZULIP_EMAIL = "abiba-bot@chat.sysloggh.net"
|
||||||
|
ZULIP_KEY = "cKTDMZAPW08dk3zl05sStzO7HRztzyn8"
|
||||||
|
ZULIP_AUTH = f"{ZULIP_EMAIL}:{ZULIP_KEY}"
|
||||||
|
|
||||||
|
LITELLM_PUBLIC = "https://litellm.sysloggh.net"
|
||||||
|
LITELLM_BACKEND = "192.168.68.116"
|
||||||
|
AUTH_HOST = "192.168.68.11"
|
||||||
|
|
||||||
|
SYNTHETIC_API_KEY = "sk-U_ydi3B-wfGU-_xESkoU1Q"
|
||||||
|
|
||||||
|
NOW = datetime.datetime.now()
|
||||||
|
DATE_STR = NOW.strftime("%Y-%m-%d")
|
||||||
|
TIME_STR = NOW.strftime("%Y-%m-%d %H:%M UTC")
|
||||||
|
|
||||||
|
# ── Helpers ──
|
||||||
|
|
||||||
|
def pve_get(path):
|
||||||
|
cmd = f'curl -sfk --connect-timeout 10 "{PVE}{path}" -H "{AUTH}"'
|
||||||
|
try:
|
||||||
|
return json.loads(subprocess.check_output(cmd, shell=True))["data"]
|
||||||
|
except: return []
|
||||||
|
|
||||||
|
def ssh(host, cmd):
|
||||||
|
try:
|
||||||
|
return subprocess.check_output(
|
||||||
|
f'ssh -o StrictHostKeyChecking=no -o ConnectTimeout=10 root@{host} "{cmd}"',
|
||||||
|
shell=True, stderr=subprocess.DEVNULL).decode().strip()
|
||||||
|
except: return ""
|
||||||
|
|
||||||
|
def ssh_jerome(host, cmd):
|
||||||
|
try:
|
||||||
|
return subprocess.check_output(
|
||||||
|
f'ssh -o StrictHostKeyChecking=no -o ConnectTimeout=10 jerome@{host} "{cmd}"',
|
||||||
|
shell=True, stderr=subprocess.DEVNULL).decode().strip()
|
||||||
|
except: return ""
|
||||||
|
|
||||||
|
def http_get(url, auth=None, timeout=10):
|
||||||
|
try:
|
||||||
|
cmd = f'curl -sfk --connect-timeout {timeout} -o /dev/null -w "%{{http_code}}" "{url}"'
|
||||||
|
if auth:
|
||||||
|
cmd = cmd.replace('"', '\\"')
|
||||||
|
cmd = f'curl -sfk --connect-timeout {timeout} -u "{auth}" -o /dev/null -w "%{{http_code}}" "{url}"'
|
||||||
|
r = subprocess.run(cmd, shell=True, capture_output=True, text=True, timeout=timeout+2)
|
||||||
|
return r.stdout.strip() or "000"
|
||||||
|
except:
|
||||||
|
return "timeout"
|
||||||
|
|
||||||
|
def http_get_body(url, auth=None, timeout=10):
|
||||||
|
try:
|
||||||
|
cmd = f'curl -sfk --connect-timeout {timeout} "{url}"'
|
||||||
|
if auth:
|
||||||
|
cmd = f'curl -sfk --connect-timeout {timeout} -u "{auth}" "{url}"'
|
||||||
|
r = subprocess.run(cmd, shell=True, capture_output=True, text=True, timeout=timeout+2)
|
||||||
|
if r.returncode == 0:
|
||||||
|
return r.stdout.strip()
|
||||||
|
return ""
|
||||||
|
except:
|
||||||
|
return ""
|
||||||
|
|
||||||
|
def count_in_log(filepath, pattern, minutes=15):
|
||||||
|
"""Count occurrences of a pattern in a log file within the last N minutes."""
|
||||||
|
try:
|
||||||
|
since = (NOW - datetime.timedelta(minutes=minutes)).strftime("%Y-%m-%d %H:%M")
|
||||||
|
cmd = f"""grep -a "{pattern}" {filepath} 2>/dev/null | awk '$0 >= "{since}"' | wc -l"""
|
||||||
|
count = subprocess.check_output(cmd, shell=True).decode().strip()
|
||||||
|
return int(count)
|
||||||
|
except:
|
||||||
|
return 0
|
||||||
|
|
||||||
|
# ── Data Collection ──
|
||||||
|
|
||||||
|
def collect():
|
||||||
|
report = {}
|
||||||
|
|
||||||
|
# ── Proxmox Nodes ──
|
||||||
|
nodes = pve_get("/api2/json/nodes")
|
||||||
|
report["nodes"] = {n["node"]: {
|
||||||
|
"cpu_pct": round(n.get('cpu',0)*100, 1),
|
||||||
|
"ram": f"{n.get('mem',0)//1024//1024}/{n.get('maxmem',0)//1024//1024}MB",
|
||||||
|
"ram_pct": round(n.get('mem',0)/n.get('maxmem',1)*100, 0),
|
||||||
|
"disk": f"{n.get('disk',0)//1024//1024//1024}/{n.get('maxdisk',0)//1024//1024//1024}GB",
|
||||||
|
"disk_pct": round(n.get('disk',0)/n.get('maxdisk',1)*100, 0),
|
||||||
|
"uptime_h": n.get('uptime',0)//3600,
|
||||||
|
"status": n["status"]
|
||||||
|
} for n in nodes}
|
||||||
|
report["node_count"] = len(nodes)
|
||||||
|
report["nodes_online"] = sum(1 for n in nodes if n["status"] == "online")
|
||||||
|
|
||||||
|
# ── VMs/CTs ──
|
||||||
|
resources = pve_get("/api2/json/cluster/resources")
|
||||||
|
vms = [r for r in resources if r.get("type") in ("qemu","lxc")]
|
||||||
|
report["total_vms"] = len(vms)
|
||||||
|
report["running_vms"] = sum(1 for v in vms if v.get("status") == "running")
|
||||||
|
stopped = [v for v in vms if v.get("status") != "running"]
|
||||||
|
report["stopped_vms"] = [f"{v['type']} {v['vmid']} {v.get('name','?')}" for v in stopped]
|
||||||
|
|
||||||
|
report["vms_by_node"] = {}
|
||||||
|
for n in ["amdpve", "minipve", "storepve", "acerpve", "ocupve"]:
|
||||||
|
node_vms = [v for v in vms if v.get("node") == n]
|
||||||
|
node_vms.sort(key=lambda x: int(x.get("vmid",0)))
|
||||||
|
report["vms_by_node"][n] = [{
|
||||||
|
"type": v.get("type"), "vmid": v.get("vmid"), "name": v.get("name","?"),
|
||||||
|
"status": v.get("status"),
|
||||||
|
"ram": f"{v.get('mem',0)//1024//1024}/{v.get('maxmem',0)//1024//1024}MB",
|
||||||
|
"disk": f"{v.get('disk',0)//1024//1024//1024}/{v.get('maxdisk',0)//1024//1024//1024}GB",
|
||||||
|
"cpu": v.get("cpus", v.get("maxcpu","?")),
|
||||||
|
} for v in node_vms]
|
||||||
|
|
||||||
|
# ── Storage ──
|
||||||
|
storages = pve_get("/api2/json/nodes/storepve/storage")
|
||||||
|
report["storage"] = []
|
||||||
|
for s in storages:
|
||||||
|
total = s.get("total",0) or 1
|
||||||
|
used = s.get("used",0)
|
||||||
|
pct = used/total*100
|
||||||
|
report["storage"].append({
|
||||||
|
"name": s.get("storage","?"), "used": f"{used//1024//1024//1024}GB",
|
||||||
|
"total": f"{total//1024//1024//1024}GB", "pct": round(pct, 0), "type": s.get("type","")
|
||||||
|
})
|
||||||
|
|
||||||
|
# ── Docker: docker-vm (.7) ──
|
||||||
|
docker_raw = ssh("192.168.68.7", "docker ps --format '{{.Names}}|{{.Status}}|{{.Image}}'")
|
||||||
|
containers = []
|
||||||
|
for line in docker_raw.strip().split("\n"):
|
||||||
|
if "|" in line:
|
||||||
|
parts = line.split("|", 2)
|
||||||
|
containers.append({"name": parts[0], "status": parts[1], "image": parts[2] if len(parts) > 2 else ""})
|
||||||
|
report["docker_vm"] = {"total": len(containers), "running": sum(1 for c in containers if "Up" in c["status"]),
|
||||||
|
"unhealthy": [c for c in containers if "unhealthy" in c["status"]], "containers": containers}
|
||||||
|
report["docker_vm"]["reclaimable"] = ssh("192.168.68.7", "docker system df 2>/dev/null | tail -1 | awk '{print $5}'")
|
||||||
|
report["docker_vm"]["disk_used"] = ssh("192.168.68.7", "df -h / | tail -1 | awk '{print $5}'")
|
||||||
|
|
||||||
|
# ── Docker: CT 116 (syslog-api — LiteLLM stack) ──
|
||||||
|
docker2_raw = ssh(LITELLM_BACKEND, "docker ps --format '{{.Names}}|{{.Status}}|{{.Image}}'")
|
||||||
|
containers2 = []
|
||||||
|
for line in docker2_raw.strip().split("\n"):
|
||||||
|
if "|" in line:
|
||||||
|
parts = line.split("|", 2)
|
||||||
|
containers2.append({"name": parts[0], "status": parts[1], "image": parts[2] if len(parts) > 2 else ""})
|
||||||
|
report["docker_syslog"] = {"total": len(containers2), "running": sum(1 for c in containers2 if "Up" in c["status"]),
|
||||||
|
"containers": containers2}
|
||||||
|
|
||||||
|
# ── Docker: Netbird (72.61.0.17) ──
|
||||||
|
docker3_raw = ssh("72.61.0.17", "docker ps --format '{{.Names}}|{{.Status}}|{{.Image}}'")
|
||||||
|
containers3 = []
|
||||||
|
for line in docker3_raw.strip().split("\n"):
|
||||||
|
if "|" in line:
|
||||||
|
parts = line.split("|", 2)
|
||||||
|
containers3.append({"name": parts[0], "status": parts[1], "image": parts[2] if len(parts) > 2 else ""})
|
||||||
|
report["docker_netbird"] = {"total": len(containers3), "running": sum(1 for c in containers3 if "Up" in c["status"]),
|
||||||
|
"containers": containers3}
|
||||||
|
|
||||||
|
# ── Public Endpoints ──
|
||||||
|
endpoints = [
|
||||||
|
("LiteLLM Admin UI", f"{LITELLM_PUBLIC}/ui/"),
|
||||||
|
("LiteLLM Swagger", f"{LITELLM_PUBLIC}/docs"),
|
||||||
|
("LiteLLM ReDoc", f"{LITELLM_PUBLIC}/redoc"),
|
||||||
|
("LiteLLM OpenAPI", f"{LITELLM_PUBLIC}/openapi.json"),
|
||||||
|
("Gitea", "https://git.sysloggh.net"),
|
||||||
|
("Authentik", "https://auth.sysloggh.net"),
|
||||||
|
("Zulip", "https://chat.sysloggh.net"),
|
||||||
|
("Pulse", "https://pulse.sysloggh.net"),
|
||||||
|
("Proxmox", "https://minipve.sysloggh.net"),
|
||||||
|
("SearXNG", "http://192.168.68.7:8888"),
|
||||||
|
("Firecrawl", "http://192.168.68.7:3002/health"),
|
||||||
|
]
|
||||||
|
report["endpoints"] = []
|
||||||
|
for name, url in endpoints:
|
||||||
|
code = http_get(url)
|
||||||
|
report["endpoints"].append({"name": name, "code": code})
|
||||||
|
|
||||||
|
# ── LiteLLM Specific Checks (from litellm-health prose contract) ──
|
||||||
|
report["litellm"] = {"checks": []}
|
||||||
|
|
||||||
|
# Check 1: LiteLLM aggregate health endpoint (router binds to 127.0.0.1, check via SSH)
|
||||||
|
health_unified = ssh(LITELLM_BACKEND, "curl -sf http://127.0.0.1:9000/health/unified -o /dev/null -w '%{http_code}' 2>/dev/null")
|
||||||
|
report["litellm"]["health_unified"] = health_unified or "000"
|
||||||
|
report["litellm"]["checks"].append({"name": "unified-health", "status": "pass" if health_unified == "200" else "fail", "code": health_unified or "000"})
|
||||||
|
|
||||||
|
# Check 2: Nginx-proxied internal endpoints
|
||||||
|
for path, name in [("/litellm/ui/", "nginx-ui"), ("/litellm/docs", "nginx-docs")]:
|
||||||
|
code = http_get(f"http://{LITELLM_BACKEND}{path}")
|
||||||
|
report["litellm"]["checks"].append({"name": name, "status": "pass" if code == "200" else "fail", "code": code})
|
||||||
|
|
||||||
|
# Check 3: Docker container health for LiteLLM stack
|
||||||
|
expected_containers = ["harness-litellm", "harness-nginx", "harness-router",
|
||||||
|
"harness-postgres", "harness-redis", "harness-dashboard"]
|
||||||
|
actual_names = [c["name"] for c in containers2]
|
||||||
|
report["litellm"]["expected_containers"] = expected_containers
|
||||||
|
report["litellm"]["missing_containers"] = [e for e in expected_containers if e not in actual_names]
|
||||||
|
report["litellm"]["checks"].append({
|
||||||
|
"name": "container-health",
|
||||||
|
"status": "pass" if not report["litellm"]["missing_containers"] else "fail",
|
||||||
|
"missing": report["litellm"]["missing_containers"],
|
||||||
|
"healthy": sum(1 for c in containers2 if "healthy" in c.get("status","")),
|
||||||
|
"total": len(expected_containers)
|
||||||
|
})
|
||||||
|
|
||||||
|
# Check 4: OIDC Auth endpoint
|
||||||
|
auth_code = http_get(f"https://auth.sysloggh.net")
|
||||||
|
report["litellm"]["auth_status"] = auth_code
|
||||||
|
report["litellm"]["checks"].append({"name": "oidc-auth", "status": "pass" if auth_code in ("200","302") else "fail", "code": auth_code})
|
||||||
|
|
||||||
|
# Check 5: Synthetic API call through LiteLLM
|
||||||
|
api_check = http_get(f"{LITELLM_PUBLIC}/v1/models", auth=SYNTHETIC_API_KEY)
|
||||||
|
report["litellm"]["api_models"] = api_check
|
||||||
|
report["litellm"]["checks"].append({"name": "api-endpoint", "status": "pass" if api_check == "200" else "fail", "code": api_check})
|
||||||
|
|
||||||
|
# ── NFS Mounts ──
|
||||||
|
nfs = ssh("192.168.68.7", "df -h /media/storage /media/mediastore 2>/dev/null | tail -n +2")
|
||||||
|
report["nfs"] = []
|
||||||
|
for line in nfs.strip().split("\n"):
|
||||||
|
parts = line.split()
|
||||||
|
if len(parts) >= 6:
|
||||||
|
report["nfs"].append({"mount": parts[5], "size": parts[1], "used": parts[2], "pct": parts[4], "avail": parts[3]})
|
||||||
|
|
||||||
|
# ── Zulip Extension Health (from zulip-health prose contract) ──
|
||||||
|
report["zulip_ext"] = {}
|
||||||
|
|
||||||
|
# Phase 1: Health endpoint
|
||||||
|
health_body = http_get_body("http://localhost:9200/health")
|
||||||
|
zulip_health = {}
|
||||||
|
try:
|
||||||
|
zulip_health = json.loads(health_body) if health_body else {}
|
||||||
|
except:
|
||||||
|
zulip_health = {}
|
||||||
|
report["zulip_ext"]["connected"] = zulip_health.get("connected", False)
|
||||||
|
report["zulip_ext"]["queue_id"] = zulip_health.get("queue_id")
|
||||||
|
report["zulip_ext"]["last_error"] = zulip_health.get("last_error")
|
||||||
|
report["zulip_ext"]["messages_processed"] = zulip_health.get("messages_processed", 0)
|
||||||
|
report["zulip_ext"]["retry_count"] = zulip_health.get("retry_count", 0)
|
||||||
|
|
||||||
|
# Phase 2: PM2 process check
|
||||||
|
pm2_raw = subprocess.check_output(
|
||||||
|
"pm2 show abiba-zulip --no-color 2>/dev/null | grep -E 'status|restarts|uptime'",
|
||||||
|
shell=True).decode().strip()
|
||||||
|
pm2 = {}
|
||||||
|
for line in pm2_raw.split("\n"):
|
||||||
|
if "│" in line:
|
||||||
|
parts = line.split("│")
|
||||||
|
if len(parts) >= 3:
|
||||||
|
key = parts[1].strip()
|
||||||
|
val = parts[2].strip()
|
||||||
|
pm2[key] = val
|
||||||
|
report["zulip_ext"]["pm2"] = pm2
|
||||||
|
report["zulip_ext"]["pm2_healthy"] = pm2.get("status") == "online"
|
||||||
|
|
||||||
|
# Phase 3: Echo loop detection
|
||||||
|
skipped_count = count_in_log("/root/.pm2/logs/abiba-zulip-out.log", "Skipped.*bot msgs", 15)
|
||||||
|
report["zulip_ext"]["bot_skipped_15min"] = skipped_count
|
||||||
|
|
||||||
|
# Phase 4: Response delivery stats
|
||||||
|
finalized = count_in_log("/root/.pm2/logs/abiba-zulip-out.log", "Finalized", 60)
|
||||||
|
failed = count_in_log("/root/.pm2/logs/abiba-zulip-out.log", "Failed to finalize", 60)
|
||||||
|
report["zulip_ext"]["finalized_1h"] = finalized
|
||||||
|
report["zulip_ext"]["failed_finalize_1h"] = failed
|
||||||
|
report["zulip_ext"]["finalize_fail_pct"] = round(failed / (finalized + failed) * 100, 0) if (finalized + failed) > 0 else 0
|
||||||
|
|
||||||
|
# Phase 5: Zulip server liveness
|
||||||
|
zulip_server_code = http_get(f"{ZULIP_SITE}/api/v1/server_settings", auth=ZULIP_AUTH)
|
||||||
|
report["zulip_ext"]["server_status"] = zulip_server_code
|
||||||
|
|
||||||
|
# ── Agent Status ──
|
||||||
|
report["agents"] = {}
|
||||||
|
|
||||||
|
# Abiba (pi)
|
||||||
|
report["agents"]["abiba"] = {
|
||||||
|
"platform": "pi", "ct": 100, "ip": "192.168.68.24",
|
||||||
|
"zulip_connected": zulip_health.get("connected", False),
|
||||||
|
"zulip_processed": zulip_health.get("messages_processed", 0),
|
||||||
|
"pm2_status": pm2.get("status", "unknown"),
|
||||||
|
"pm2_restarts": pm2.get("restarts", "?"),
|
||||||
|
"pm2_uptime": pm2.get("uptime", "?"),
|
||||||
|
}
|
||||||
|
|
||||||
|
# Tanko (CT 122)
|
||||||
|
tanko_state = ssh_jerome("192.168.68.122", "cat ~/.hermes/gateway_state.json 2>/dev/null")
|
||||||
|
tanko_data = {}
|
||||||
|
try:
|
||||||
|
tanko_data = json.loads(tanko_state) if tanko_state else {}
|
||||||
|
except:
|
||||||
|
tanko_data = {}
|
||||||
|
platforms = tanko_data.get("platforms", {})
|
||||||
|
report["agents"]["tanko"] = {
|
||||||
|
"platform": "hermes", "ct": 112, "ip": "192.168.68.122",
|
||||||
|
"gateway_state": tanko_data.get("gateway_state", "unknown"),
|
||||||
|
"zulip_state": platforms.get("zulip", {}).get("state", "unknown"),
|
||||||
|
"telegram_state": platforms.get("telegram", {}).get("state", "unknown"),
|
||||||
|
"gateway_pid": tanko_data.get("pid"),
|
||||||
|
"updated_at": tanko_data.get("updated_at"),
|
||||||
|
}
|
||||||
|
|
||||||
|
# Mumuni (CT 114, IP 192.168.68.123)
|
||||||
|
mumuni_state = ssh("192.168.68.123", "cat ~/.hermes/gateway_state.json 2>/dev/null")
|
||||||
|
mumuni_data = {}
|
||||||
|
try:
|
||||||
|
mumuni_data = json.loads(mumuni_state) if mumuni_state else {}
|
||||||
|
except:
|
||||||
|
mumuni_data = {}
|
||||||
|
mumuni_platforms = mumuni_data.get("platforms", {})
|
||||||
|
report["agents"]["mumuni"] = {
|
||||||
|
"platform": "hermes", "ct": 114, "ip": "192.168.68.123",
|
||||||
|
"gateway_state": mumuni_data.get("gateway_state", "unknown"),
|
||||||
|
"telegram_state": mumuni_platforms.get("telegram", {}).get("state", "unknown"),
|
||||||
|
"zulip_state": mumuni_platforms.get("zulip", {}).get("state", "not_installed"),
|
||||||
|
"email_state": mumuni_platforms.get("email", {}).get("state", "unknown"),
|
||||||
|
"hermes_version": "",
|
||||||
|
}
|
||||||
|
# Get Hermes version
|
||||||
|
ver = ssh("192.168.68.123", "hermes --version 2>/dev/null | head -1")
|
||||||
|
if ver:
|
||||||
|
report["agents"]["mumuni"]["hermes_version"] = ver.split("·")[0].replace("Hermes Agent ","").strip()
|
||||||
|
|
||||||
|
return report
|
||||||
|
|
||||||
|
|
||||||
|
# ── HTML Dashboard ──
|
||||||
|
|
||||||
|
def build_html(r):
|
||||||
|
issues = []
|
||||||
|
|
||||||
|
# Proxmox issues
|
||||||
|
if r["nodes_online"] < r["node_count"]:
|
||||||
|
issues.append(f"🔴 {r['node_count'] - r['nodes_online']} Proxmox node(s) offline")
|
||||||
|
if r["stopped_vms"]:
|
||||||
|
issues.append(f"🔴 {len(r['stopped_vms'])} VM(s)/CT(s) stopped — {', '.join(r['stopped_vms'][:3])}")
|
||||||
|
|
||||||
|
# Endpoint issues
|
||||||
|
for ep in r["endpoints"]:
|
||||||
|
if ep["code"] in ("000", "timeout"):
|
||||||
|
issues.append(f"🔴 {ep['name']} — unreachable")
|
||||||
|
elif ep["code"] >= "500":
|
||||||
|
issues.append(f"🟡 {ep['name']} — HTTP {ep['code']}")
|
||||||
|
|
||||||
|
# Docker issues
|
||||||
|
for c in r["docker_vm"].get("unhealthy", []):
|
||||||
|
issues.append(f"🟡 docker-vm: {c['name']} — unhealthy")
|
||||||
|
|
||||||
|
# Small cluster mode warning
|
||||||
|
if r["docker_vm"].get("total", 0) == 0:
|
||||||
|
issues.append("🟡 docker-vm: no containers reported — host may be down")
|
||||||
|
|
||||||
|
# LiteLLM issues
|
||||||
|
for check in r.get("litellm", {}).get("checks", []):
|
||||||
|
if check["status"] == "fail":
|
||||||
|
issues.append(f"🔴 LiteLLM: {check['name']} — {'missing: ' + ', '.join(check.get('missing',[])) if check.get('missing') else 'HTTP ' + check.get('code','?')}")
|
||||||
|
|
||||||
|
# Zulip extension issues
|
||||||
|
z = r.get("zulip_ext", {})
|
||||||
|
if not z.get("connected"):
|
||||||
|
issues.append("🔴 Zulip extension: disconnected")
|
||||||
|
if z.get("last_error"):
|
||||||
|
issues.append(f"🟡 Zulip extension: {z['last_error'][:80]}")
|
||||||
|
if z.get("retry_count", 0) >= 3:
|
||||||
|
issues.append(f"🟡 Zulip extension: {z['retry_count']} retries")
|
||||||
|
if z.get("finalize_fail_pct", 0) > 50:
|
||||||
|
issues.append(f"🔴 Zulip extension: {z['finalize_fail_pct']:.0f}% edit failures")
|
||||||
|
if z.get("bot_skipped_15min", 0) > 100:
|
||||||
|
issues.append(f"🟡 Zulip echo loop: {z['bot_skipped_15min']} bot msgs skipped in 15min")
|
||||||
|
if z.get("server_status") and z["server_status"] not in ("200", "302"):
|
||||||
|
issues.append(f"🔴 Zulip server: HTTP {z['server_status']}")
|
||||||
|
|
||||||
|
# Agent issues
|
||||||
|
def agent_gateway(agent):
|
||||||
|
return agent.get("gateway_state") or agent.get("pm2_status") or "unknown"
|
||||||
|
for name, agent in r.get("agents", {}).items():
|
||||||
|
gs = agent_gateway(agent)
|
||||||
|
if gs == "unknown":
|
||||||
|
issues.append(f"🟡 {name}: gateway state unknown (unreachable)")
|
||||||
|
elif gs != "running" and gs != "online":
|
||||||
|
issues.append(f"🔴 {name}: gateway {gs}")
|
||||||
|
|
||||||
|
status = "🟢 All Healthy" if not issues else f"{'🔴' if sum(1 for i in issues if '🔴' in i) > 0 else '🟡'} {len(issues)} Issue(s)"
|
||||||
|
|
||||||
|
# ── Build HTML ──
|
||||||
|
html = f"""<!DOCTYPE html>
|
||||||
|
<html><head><meta charset="utf-8"><style>
|
||||||
|
body {{ font-family: -apple-system, BlinkMacSystemFont, sans-serif; background: #0d1117; color: #c9d1d9; padding: 20px; }}
|
||||||
|
h1 {{ color: #58a6ff; }}
|
||||||
|
h2 {{ color: #8b949e; border-bottom: 1px solid #30363d; padding-bottom: 8px; margin-top: 24px; }}
|
||||||
|
h3 {{ color: #c9d1d9; margin: 12px 0 6px 0; }}
|
||||||
|
.card {{ background: #161b22; border: 1px solid #30363d; border-radius: 8px; padding: 16px; margin: 12px 0; }}
|
||||||
|
.grid {{ display: flex; gap: 12px; flex-wrap: wrap; }}
|
||||||
|
.stat {{ flex: 1; min-width: 140px; background: #21262d; border-radius: 6px; padding: 12px; text-align: center; }}
|
||||||
|
.stat .num {{ font-size: 28px; font-weight: bold; color: #58a6ff; }}
|
||||||
|
.stat .label {{ font-size: 12px; color: #8b949e; }}
|
||||||
|
table {{ width: 100%; border-collapse: collapse; }}
|
||||||
|
td, th {{ padding: 6px 12px; border-bottom: 1px solid #21262d; text-align: left; font-size: 13px; }}
|
||||||
|
th {{ color: #8b949e; font-weight: normal; }}
|
||||||
|
.green {{ color: #3fb950; }}
|
||||||
|
.red {{ color: #f85149; }}
|
||||||
|
.yellow {{ color: #d29922; }}
|
||||||
|
.alert {{ border-radius: 8px; padding: 12px; margin: 12px 0; }}
|
||||||
|
.good {{ background: #1f3d1f; border: 1px solid #3fb950; }}
|
||||||
|
.bad {{ background: #3d1f1f; border: 1px solid #f85149; }}
|
||||||
|
.warn {{ background: #3d3d1f; border: 1px solid #d29922; }}
|
||||||
|
</style></head><body>
|
||||||
|
<h1>🏗️ Syslog Infrastructure Report</h1>
|
||||||
|
<p style="color:#8b949e">{TIME_STR}</p>
|
||||||
|
|
||||||
|
<div class="alert {'good' if not issues else 'bad' if any('🔴' in i for i in issues) else 'warn'}">
|
||||||
|
<p style="margin:0;font-size:16px"><b>{status}</b></p>
|
||||||
|
<p style="margin:4px 0 0 0;font-size:13px">
|
||||||
|
{r['node_count']} PVE nodes · {r['total_vms']} VMs/CTs · {r['running_vms']} running ·
|
||||||
|
{r['docker_vm']['total'] + r['docker_syslog']['total'] + r['docker_netbird']['total']} containers ·
|
||||||
|
{len(r['endpoints'])} endpoints · {len(r.get('agents',{}))} agents
|
||||||
|
</p>
|
||||||
|
</div>
|
||||||
|
"""
|
||||||
|
|
||||||
|
# Issues list
|
||||||
|
if issues:
|
||||||
|
html += '<div class="card"><h2>🚨 Issues ({})</h2>'.format(len(issues))
|
||||||
|
for i in issues:
|
||||||
|
html += f'<p style="margin:4px 0;font-size:13px">{i}</p>'
|
||||||
|
html += '</div>'
|
||||||
|
|
||||||
|
# ── Quick Stats ──
|
||||||
|
html += '<div class="card"><h2>📊 Quick Stats</h2><div class="grid">'
|
||||||
|
stats = [
|
||||||
|
("PVE Nodes", f"{r['nodes_online']}/{r['node_count']}", "green" if r['nodes_online'] == r['node_count'] else "red"),
|
||||||
|
("VMs/CTs", f"{r['running_vms']}/{r['total_vms']}", "green" if r['running_vms'] == r['total_vms'] else "red"),
|
||||||
|
("Containers", f"{r['docker_vm']['running']}/{r['docker_vm']['total']}", "green" if r['docker_vm']['running'] == r['docker_vm']['total'] else "yellow"),
|
||||||
|
("LiteLLM Ctrs", f"{r['docker_syslog']['running']}/{r['docker_syslog']['total']}", "green" if r['docker_syslog']['running'] == r['docker_syslog']['total'] else "red"),
|
||||||
|
("Netbird Ctrs", f"{r['docker_netbird']['running']}/{r['docker_netbird']['total']}", "green" if r['docker_netbird']['running'] == r['docker_netbird']['total'] else "red"),
|
||||||
|
("Endpoints", f"{sum(1 for e in r['endpoints'] if e['code'] in ('200','302','401'))}/{len(r['endpoints'])}", "green",),
|
||||||
|
("Zulip Ext", "✅ Connected" if r.get('zulip_ext',{}).get('connected') else "❌ Disconnected", "green" if r.get('zulip_ext',{}).get('connected') else "red"),
|
||||||
|
]
|
||||||
|
for label, val, color in stats:
|
||||||
|
html += f'<div class="stat"><div class="num {color}">{val}</div><div class="label">{label}</div></div>'
|
||||||
|
|
||||||
|
# VLM/GPU capacity
|
||||||
|
gpu_nodes = [v for vms in r["vms_by_node"].values() for v in vms if "llm" in v["name"].lower() or "gpu" in v["name"].lower() or "ocu" in v["name"].lower()]
|
||||||
|
gpu_running = sum(1 for g in gpu_nodes if g["status"] == "running")
|
||||||
|
html += f'<div class="stat"><div class="num {"green" if gpu_running == len(gpu_nodes) else "yellow"}">{gpu_running}/{len(gpu_nodes)}</div><div class="label">GPU VMs</div></div>'
|
||||||
|
html += '</div></div>'
|
||||||
|
|
||||||
|
# ── LiteLLM Section ──
|
||||||
|
html += '<div class="card"><h2>⚡ LiteLLM Inference Stack</h2>'
|
||||||
|
|
||||||
|
# Endpoint status table
|
||||||
|
html += '<table><tr><th>Check</th><th>Status</th></tr>'
|
||||||
|
for check in r.get("litellm", {}).get("checks", []):
|
||||||
|
color = "green" if check["status"] == "pass" else "red"
|
||||||
|
detail = check.get("code", "")
|
||||||
|
if check.get("missing"):
|
||||||
|
detail = f"missing: {', '.join(check['missing'])}"
|
||||||
|
elif check.get("healthy") is not None:
|
||||||
|
detail = f"{check['healthy']}/{check['total']} healthy"
|
||||||
|
html += f'<tr><td>{check["name"]}</td><td class="{color}">{detail}</td></tr>'
|
||||||
|
html += '</table>'
|
||||||
|
|
||||||
|
# LiteLLM container detail
|
||||||
|
html += '<h3>Containers (CT 116)</h3><table><tr><th>Container</th><th>Status</th><th>Image</th></tr>'
|
||||||
|
for c in r["docker_syslog"]["containers"]:
|
||||||
|
color = "green" if "Up" in c["status"] and "healthy" in c["status"] else ("yellow" if "Up" in c["status"] else "red")
|
||||||
|
html += f'<tr><td>{c["name"]}</td><td class="{color}">{c["status"]}</td><td style="font-size:11px;color:#8b949e">{c["image"]}</td></tr>'
|
||||||
|
html += '</table>'
|
||||||
|
html += '</div>'
|
||||||
|
|
||||||
|
# ── Proxmox Nodes ──
|
||||||
|
html += '<div class="card"><h2>📦 Proxmox Cluster</h2><div class="grid">'
|
||||||
|
for name, nd in sorted(r["nodes"].items()):
|
||||||
|
cpu_color = "green" if nd["cpu_pct"] < 50 else ("yellow" if nd["cpu_pct"] < 80 else "red")
|
||||||
|
ram_color = "green" if nd["ram_pct"] < 70 else ("yellow" if nd["ram_pct"] < 85 else "red")
|
||||||
|
disk_color = "green" if nd["disk_pct"] < 70 else ("yellow" if nd["disk_pct"] < 85 else "red")
|
||||||
|
html += f'<div class="stat"><div class="num {cpu_color}">{nd["cpu_pct"]}%</div><div class="label"><b>{name}</b><br>CPU: {nd["cpu_pct"]}%<br>RAM: <span class="{ram_color}">{nd["ram_pct"]:.0f}%</span><br>DISK: <span class="{disk_color}">{nd["disk_pct"]:.0f}%</span><br>⬆ {nd["uptime_h"]}h</div></div>'
|
||||||
|
html += '</div>'
|
||||||
|
if r["stopped_vms"]:
|
||||||
|
html += f'<p class="red" style="font-size:13px">Stopped: {", ".join(r["stopped_vms"])}</p>'
|
||||||
|
html += '</div>'
|
||||||
|
|
||||||
|
# ── VM/CT Lists per Node ──
|
||||||
|
for node_name in ["amdpve", "minipve", "storepve", "acerpve", "ocupve"]:
|
||||||
|
node_vms = r["vms_by_node"].get(node_name, [])
|
||||||
|
if not node_vms: continue
|
||||||
|
stopped_count = sum(1 for v in node_vms if v["status"] != "running")
|
||||||
|
badge = f'<span class="green">{len(node_vms)-stopped_count} running</span>'
|
||||||
|
if stopped_count:
|
||||||
|
badge += f' <span class="red">{stopped_count} stopped</span>'
|
||||||
|
html += f'<div class="card"><h3>🔹 {node_name}</h3><p style="font-size:12px;color:#8b949e">{badge}</p>'
|
||||||
|
html += '<table><tr><th>ID</th><th>Type</th><th>Name</th><th>Status</th><th>CPU</th><th>RAM</th><th>Disk</th></tr>'
|
||||||
|
for v in node_vms:
|
||||||
|
color = "green" if v["status"] == "running" else "red"
|
||||||
|
html += f'<tr><td>{v["vmid"]}</td><td>{v["type"]}</td><td>{v["name"]}</td><td class="{color}">{v["status"]}</td><td>{v["cpu"]}</td><td>{v["ram"]}</td><td>{v["disk"]}</td></tr>'
|
||||||
|
html += '</table></div>'
|
||||||
|
|
||||||
|
# ── Storage ──
|
||||||
|
html += '<div class="card"><h2>💾 Storage</h2><table><tr><th>Pool</th><th>Used</th><th>Total</th><th>%</th><th>Type</th></tr>'
|
||||||
|
for s in sorted(r["storage"], key=lambda x: -x["pct"]):
|
||||||
|
color = "green" if s["pct"] < 70 else ("yellow" if s["pct"] < 85 else "red")
|
||||||
|
html += f'<tr><td>{s["name"]}</td><td>{s["used"]}</td><td>{s["total"]}</td><td class="{color}">{s["pct"]:.0f}%</td><td>{s["type"]}</td></tr>'
|
||||||
|
html += '</table></div>'
|
||||||
|
|
||||||
|
# ── Docker Ecosystems ──
|
||||||
|
html += '<div class="card"><h2>🐳 Docker Ecosystems</h2>'
|
||||||
|
for host, key, label in [
|
||||||
|
("docker-vm (.7) — storepve", "docker_vm", f"Disk: {r['docker_vm'].get('disk_used','?')}"),
|
||||||
|
("syslog-api (.116) — minipve (LiteLLM)", "docker_syslog", ""),
|
||||||
|
("Netbird (72.61.0.17)", "docker_netbird", ""),
|
||||||
|
]:
|
||||||
|
d = r[key]
|
||||||
|
html += f'<h3>{host} — {d["running"]}/{d["total"]} running {label}</h3>'
|
||||||
|
if d.get("reclaimable"):
|
||||||
|
html += f'<p style="font-size:12px;color:#8b949e">Reclaimable Docker space: {d["reclaimable"]}</p>'
|
||||||
|
html += '<table><tr><th>Container</th><th>Status</th></tr>'
|
||||||
|
for c in sorted(d["containers"], key=lambda x: x["name"]):
|
||||||
|
color = "green" if "Up" in c["status"] and "healthy" in c["status"] else ("yellow" if "Up" in c["status"] else "red")
|
||||||
|
html += f'<tr><td>{c["name"]}</td><td class="{color}">{c["status"][:50]}</td></tr>'
|
||||||
|
html += '</table>'
|
||||||
|
html += '</div>'
|
||||||
|
|
||||||
|
# ── Agent Status ──
|
||||||
|
html += '<div class="card"><h2>🤖 Agent Status</h2><table><tr><th>Agent</th><th>Platform</th><th>CT</th><th>Gateway</th><th>Zulip</th><th>Processed</th></tr>'
|
||||||
|
for name, agent in sorted(r.get("agents", {}).items()):
|
||||||
|
if name == "abiba":
|
||||||
|
zulip_state = "✅" if agent.get("zulip_connected") else "❌"
|
||||||
|
gateway = f"pm2:{agent.get('pm2_status','?')} (↺{agent.get('pm2_restarts','?')})"
|
||||||
|
processed = f"{agent.get('zulip_processed', 0)} msgs"
|
||||||
|
elif name == "tanko":
|
||||||
|
zulip_state = "✅" if agent.get("zulip_state") == "connected" else ("❌" if agent.get("zulip_state") == "disconnected" else "⬜")
|
||||||
|
gateway = agent.get("gateway_state", "?")
|
||||||
|
processed = agent.get("updated_at", "")[:10]
|
||||||
|
elif name == "mumuni":
|
||||||
|
zulip_state = "⬜" if agent.get("zulip_state") == "not_installed" else ("✅" if agent.get("zulip_state") == "connected" else "⬜")
|
||||||
|
gateway = agent.get("gateway_state", "?")
|
||||||
|
tg = "✅" if agent.get("telegram_state") == "connected" else "❌"
|
||||||
|
ver = agent.get("hermes_version", "")
|
||||||
|
processed = f"TG:{tg} v{ver}"
|
||||||
|
else:
|
||||||
|
zulip_state = "⬜"
|
||||||
|
gateway = agent.get("gateway_state", "?")
|
||||||
|
processed = ""
|
||||||
|
html += f'<tr><td><b>{name}</b></td><td>{agent["platform"]}</td><td>{agent["ct"]}</td><td>{gateway}</td><td>{zulip_state}</td><td style="font-size:12px">{processed}</td></tr>'
|
||||||
|
html += '</table></div>'
|
||||||
|
|
||||||
|
# ── Zulip Extension Health ──
|
||||||
|
html += '<div class="card"><h2>💬 Zulip Extension (Abiba)</h2><table><tr><th>Check</th><th>Status</th></tr>'
|
||||||
|
z = r.get("zulip_ext", {})
|
||||||
|
|
||||||
|
checks = [
|
||||||
|
("Connection", "✅ Connected" if z.get("connected") else "❌ Disconnected"),
|
||||||
|
("Queue", f"✅ {z.get('queue_id','none')[:12]}..." if z.get("queue_id") else "❌ No queue"),
|
||||||
|
("Messages", f"{z.get('messages_processed',0)} processed"),
|
||||||
|
("PM2 Status", f"{z.get('pm2',{}).get('status','?')} (↺{z.get('pm2',{}).get('restarts','?')} restarts, up {z.get('pm2',{}).get('uptime','?')})"),
|
||||||
|
("Bot Echo Loop", f"{z.get('bot_skipped_15min',0)} skipped in 15min" if z.get('bot_skipped_15min',0) > 0 else "✅ No echo loop"),
|
||||||
|
("Edit Success", f"{z.get('finalized_1h',0)} ok / {z.get('failed_finalize_1h',0)} fail" if z.get('failed_finalize_1h',0) > 0 else f"✅ {z.get('finalized_1h',0)} ok"),
|
||||||
|
("Zulip Server", f"✅ HTTP {z.get('server_status','?')}" if z.get('server_status') in ('200','302') else f"❌ HTTP {z.get('server_status','?')}"),
|
||||||
|
]
|
||||||
|
for name, val in checks:
|
||||||
|
color = "green" if "✅" in val or "ok" in val.lower() else ("yellow" if "🟡" in val or "skip" in val.lower() else "red")
|
||||||
|
html += f'<tr><td>{name}</td><td class="{color}">{val}</td></tr>'
|
||||||
|
html += '</table></div>'
|
||||||
|
|
||||||
|
# ── Network Endpoints ──
|
||||||
|
html += '<div class="card"><h2>🌐 Network Endpoints</h2><table><tr><th>Service</th><th>Status</th></tr>'
|
||||||
|
for ep in r["endpoints"]:
|
||||||
|
color = "green" if ep["code"] in ("200","302","401") else ("yellow" if ep["code"] >= "400" else "red")
|
||||||
|
html += f'<tr><td>{ep["name"]}</td><td class="{color}">HTTP {ep["code"]}</td></tr>'
|
||||||
|
html += '</table></div>'
|
||||||
|
|
||||||
|
# ── NFS ──
|
||||||
|
html += '<div class="card"><h2>🗄️ NFS Mounts</h2><table><tr><th>Mount</th><th>Size</th><th>Used</th><th>Avail</th><th>%</th></tr>'
|
||||||
|
for m in r["nfs"]:
|
||||||
|
pct_str = m.get("pct","0%")
|
||||||
|
pct_num = int(pct_str.replace("%","")) if pct_str.replace("%","").strip().isdigit() else 0
|
||||||
|
color = "green" if pct_num < 80 else ("yellow" if pct_num < 90 else "red")
|
||||||
|
html += f'<tr><td>{m["mount"]}</td><td>{m["size"]}</td><td>{m["used"]}</td><td>{m.get("avail","?")}</td><td class="{color}">{pct_str}</td></tr>'
|
||||||
|
html += '</table></div>'
|
||||||
|
|
||||||
|
# ── Suggestions ──
|
||||||
|
html += '<div class="card"><h2>💡 Suggestions</h2><ul style="font-size:13px;line-height:1.6">'
|
||||||
|
|
||||||
|
added = False
|
||||||
|
|
||||||
|
# Storage warnings
|
||||||
|
for s in r["storage"]:
|
||||||
|
if s["name"] == "mediastore" and s["pct"] > 70:
|
||||||
|
html += f'<li>📀 mediastore at {s["pct"]:.0f}% ({s["used"]}/{s["total"]}) — plan expansion or cleanup</li>'
|
||||||
|
added = True
|
||||||
|
|
||||||
|
# Stopped VMs
|
||||||
|
if r["stopped_vms"]:
|
||||||
|
html += f'<li>🛑 Stopped VMs/CTs need attention: {", ".join(r["stopped_vms"][:5])}</li>'
|
||||||
|
added = True
|
||||||
|
|
||||||
|
# Endpoint issues
|
||||||
|
for ep in r["endpoints"]:
|
||||||
|
if ep["code"] >= "500":
|
||||||
|
html += f'<li>🌐 {ep["name"]} returning HTTP {ep["code"]}</li>'
|
||||||
|
added = True
|
||||||
|
|
||||||
|
# Zulip extension
|
||||||
|
if not z.get("connected"):
|
||||||
|
html += f'<li>💬 Zulip extension disconnected — run: pm2 restart abiba-zulip</li>'
|
||||||
|
added = True
|
||||||
|
if z.get("finalize_fail_pct", 0) > 50:
|
||||||
|
html += f'<li>💬 Zulip edit failures at {z["finalize_fail_pct"]:.0f}% — check editMessage API</li>'
|
||||||
|
added = True
|
||||||
|
|
||||||
|
# LiteLLM
|
||||||
|
for check in r.get("litellm", {}).get("checks", []):
|
||||||
|
if check["status"] == "fail" and check.get("missing"):
|
||||||
|
html += f'<li>⚡ LiteLLM missing containers: {", ".join(check["missing"])}</li>'
|
||||||
|
added = True
|
||||||
|
|
||||||
|
# Backup compliance
|
||||||
|
html += '<li>🗄️ Verify PBS backups completed within last 48h via Proxmox Backup Server (CT 107)</li>'
|
||||||
|
|
||||||
|
# Heartbeat suggestion
|
||||||
|
html += '<li>💓 All agents now have Zulip heartbeat logging — silence detection active</li>'
|
||||||
|
|
||||||
|
if not added:
|
||||||
|
html += '<li>✅ Nothing flagged — infrastructure is healthy</li>'
|
||||||
|
|
||||||
|
html += '</ul></div>'
|
||||||
|
html += '<p style="color:#484f58;font-size:11px;text-align:center;margin-top:20px">Syslog Infrastructure Monitor v2.0 · Generated ' + TIME_STR + '</p>'
|
||||||
|
html += '</body></html>'
|
||||||
|
return html
|
||||||
|
|
||||||
|
|
||||||
|
# ── Send Email ──
|
||||||
|
|
||||||
|
def send_email(html_content, subject_prefix=""):
|
||||||
|
FROM = "abiba@sysloggh.com"
|
||||||
|
TO = "jerome@sysloggh.com"
|
||||||
|
SUBJECT = f"{subject_prefix}{'🏗️ Infrastructure Report — ' + DATE_STR}"
|
||||||
|
|
||||||
|
msg = MIMEMultipart("alternative")
|
||||||
|
msg["From"] = FROM
|
||||||
|
msg["To"] = TO
|
||||||
|
msg["Subject"] = SUBJECT
|
||||||
|
msg.attach(MIMEText("Infrastructure report in HTML format — enable images to view.", "plain"))
|
||||||
|
msg.attach(MIMEText(html_content, "html"))
|
||||||
|
|
||||||
|
try:
|
||||||
|
EMAIL_PASSWORD = "rgbuomwcydxwbszd"
|
||||||
|
GMAIL_EMAIL = "jtabiri@gmail.com"
|
||||||
|
|
||||||
|
server = smtplib.SMTP("smtp.gmail.com", 587)
|
||||||
|
server.starttls()
|
||||||
|
server.login(GMAIL_EMAIL, EMAIL_PASSWORD)
|
||||||
|
server.sendmail(FROM, [TO], msg.as_string())
|
||||||
|
server.quit()
|
||||||
|
return True, "✅ Email sent to jerome@sysloggh.com"
|
||||||
|
except Exception as e:
|
||||||
|
return False, f"❌ Email failed: {e}"
|
||||||
|
|
||||||
|
|
||||||
|
# ── Main ──
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
is_test = "--test-email" in sys.argv
|
||||||
|
|
||||||
|
print(f"{'🧪 TEST MODE' if is_test else '📊'} Collecting infrastructure data...")
|
||||||
|
report = collect()
|
||||||
|
|
||||||
|
if "--json" in sys.argv:
|
||||||
|
print(json.dumps(report, indent=2, default=str))
|
||||||
|
sys.exit(0)
|
||||||
|
|
||||||
|
print(" Building dashboard...")
|
||||||
|
html = build_html(report)
|
||||||
|
|
||||||
|
if is_test:
|
||||||
|
prefix = "🧪 TEST — "
|
||||||
|
print(" Sending test email...")
|
||||||
|
else:
|
||||||
|
prefix = ""
|
||||||
|
print(" Sending email...")
|
||||||
|
|
||||||
|
ok, msg = send_email(html, subject_prefix=prefix)
|
||||||
|
print(f" {msg}")
|
||||||
|
|
||||||
|
# Show summary
|
||||||
|
issues = sum(1 for i in ["red"] if report.get("zulip_ext", {}).get("connected") == False)
|
||||||
|
print(f"\n📋 Summary:")
|
||||||
|
print(f" Proxmox: {report['nodes_online']}/{report['node_count']} nodes online")
|
||||||
|
print(f" VMs/CTs: {report['running_vms']}/{report['total_vms']} running")
|
||||||
|
print(f" Zulip Ext: {'✅' if report.get('zulip_ext',{}).get('connected') else '❌'}")
|
||||||
|
print(f" LiteLLM: {sum(1 for c in report.get('litellm',{}).get('checks',[]) if c['status']=='pass')}/{len(report.get('litellm',{}).get('checks',[]))} checks pass")
|
||||||
|
agent_parts = []
|
||||||
|
for k,v in report.get('agents',{}).items():
|
||||||
|
agent_parts.append(f"{k}:{v.get('gateway_state',v.get('pm2_status','?'))}")
|
||||||
|
print(f" Agents: {', '.join(agent_parts)}")
|
||||||
Executable
+156
@@ -0,0 +1,156 @@
|
|||||||
|
#!/bin/bash
|
||||||
|
# /root/scripts/zulip-monitor.sh — Zulip Mesh Health Monitor
|
||||||
|
# Implements zulip-health.prose.md v2
|
||||||
|
# Runs every 15 min via cron. Alerts via Telegram.
|
||||||
|
set -euo pipefail
|
||||||
|
|
||||||
|
ZULIP_SITE="https://chat.sysloggh.net"
|
||||||
|
ZULIP_EMAIL="abiba-bot@chat.sysloggh.net"
|
||||||
|
ZULIP_KEY="cKTDMZAPW08dk3zl05sStzO7HRztzyn8"
|
||||||
|
OWNER_ZULIP_ID="9"
|
||||||
|
|
||||||
|
# Email config
|
||||||
|
GMAIL_USER="jtabiri@gmail.com"
|
||||||
|
GMAIL_PASS="rgbuomwcydxwbszd"
|
||||||
|
EMAIL_TO="jerome@sysloggh.com"
|
||||||
|
|
||||||
|
notify() {
|
||||||
|
local severity="$1" msg="$2"
|
||||||
|
echo "[$severity] $msg"
|
||||||
|
|
||||||
|
# Zulip DM to owner
|
||||||
|
local content="${severity} Zulip Monitor: ${msg}"
|
||||||
|
local form="type=private&to=%5B${OWNER_ZULIP_ID}%5D&content=$(python3 -c "import urllib.parse; print(urllib.parse.quote('''${content}'''))")"
|
||||||
|
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
|
||||||
|
-u "${ZULIP_EMAIL}:${ZULIP_KEY}" \
|
||||||
|
-d "${form}" > /dev/null 2>&1 || true
|
||||||
|
|
||||||
|
# Email alert
|
||||||
|
local subject="${severity} Zulip Monitor Alert"
|
||||||
|
python3 -c "
|
||||||
|
import smtplib
|
||||||
|
from email.mime.text import MIMEText
|
||||||
|
m = MIMEText('''${msg}''')
|
||||||
|
m['From'] = 'abiba@sysloggh.com'
|
||||||
|
m['To'] = '${EMAIL_TO}'
|
||||||
|
m['Subject'] = '${subject}'
|
||||||
|
s = smtplib.SMTP('smtp.gmail.com', 587)
|
||||||
|
s.starttls()
|
||||||
|
s.login('${GMAIL_USER}', '${GMAIL_PASS}')
|
||||||
|
s.sendmail('abiba@sysloggh.com', ['${EMAIL_TO}'], m.as_string())
|
||||||
|
s.quit()
|
||||||
|
" 2>/dev/null || true
|
||||||
|
}
|
||||||
|
|
||||||
|
TIMESTAMP=$(date -u '+%Y-%m-%d %H:%M UTC')
|
||||||
|
ISSUES=0
|
||||||
|
LOG="/root/zulip-health-monitor.log"
|
||||||
|
|
||||||
|
echo "=== Zulip Health Check — $TIMESTAMP ===" >> "$LOG"
|
||||||
|
|
||||||
|
# ── Global: Zulip Server ──
|
||||||
|
SERVER_CODE=$(curl -s -o /dev/null -w "%{http_code}" --connect-timeout 10 \
|
||||||
|
https://chat.sysloggh.net/api/v1/server_settings \
|
||||||
|
-u 'abiba-bot@chat.sysloggh.net:cKTDMZAPW08dk3zl05sStzO7HRztzyn8' 2>/dev/null || echo "000")
|
||||||
|
if [ "$SERVER_CODE" != "200" ]; then
|
||||||
|
notify "🔴" "Zulip server returned HTTP $SERVER_CODE"
|
||||||
|
ISSUES=$((ISSUES + 1))
|
||||||
|
else
|
||||||
|
echo " Server: ✅ HTTP 200" >> "$LOG"
|
||||||
|
fi
|
||||||
|
|
||||||
|
# ── Platform A: pi (Abiba) ──
|
||||||
|
PI_HEALTH=$(curl -sf --connect-timeout 5 http://localhost:9200/health 2>/dev/null || echo "{}")
|
||||||
|
PI_CONNECTED=$(echo "$PI_HEALTH" | python3 -c "import sys,json; d=json.load(sys.stdin); print(d.get('connected',False))" 2>/dev/null)
|
||||||
|
PI_ERROR=$(echo "$PI_HEALTH" | python3 -c "import sys,json; d=json.load(sys.stdin); print(d.get('last_error') or '')" 2>/dev/null)
|
||||||
|
PI_RETRIES=$(echo "$PI_HEALTH" | python3 -c "import sys,json; d=json.load(sys.stdin); print(d.get('retry_count',0))" 2>/dev/null)
|
||||||
|
|
||||||
|
if [ "$PI_CONNECTED" != "True" ]; then
|
||||||
|
notify "🔴" "Abiba pi extension DISCONNECTED — restarting"
|
||||||
|
pm2 restart abiba-zulip 2>/dev/null || true
|
||||||
|
ISSUES=$((ISSUES + 1))
|
||||||
|
echo " Abiba: ❌ Disconnected — restarted" >> "$LOG"
|
||||||
|
elif [ -n "$PI_ERROR" ]; then
|
||||||
|
notify "🟡" "Abiba pi extension error: ${PI_ERROR:0:100}"
|
||||||
|
echo " Abiba: 🟡 Error: ${PI_ERROR:0:100}" >> "$LOG"
|
||||||
|
elif [ "$PI_RETRIES" -ge 3 ]; then
|
||||||
|
notify "🟡" "Abiba pi extension: $PI_RETRIES retries — restarting"
|
||||||
|
pm2 restart abiba-zulip 2>/dev/null || true
|
||||||
|
echo " Abiba: 🟡 $PI_RETRIES retries — restarted" >> "$LOG"
|
||||||
|
else
|
||||||
|
echo " Abiba: ✅ Connected (processed=$(echo "$PI_HEALTH" | python3 -c "import sys,json; d=json.load(sys.stdin); print(d.get('messages_processed',0))" 2>/dev/null))" >> "$LOG"
|
||||||
|
fi
|
||||||
|
|
||||||
|
# ── Platform B: Hermes (Tanko) ──
|
||||||
|
TANKO_STATE=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 jerome@192.168.68.122 \
|
||||||
|
"cat ~/.hermes/gateway_state.json 2>/dev/null" 2>/dev/null || echo "{}")
|
||||||
|
TANKO_ZULIP=$(echo "$TANKO_STATE" | python3 -c "
|
||||||
|
import sys,json
|
||||||
|
d=json.load(sys.stdin)
|
||||||
|
p=d.get('platforms',{}).get('zulip',{})
|
||||||
|
print(p.get('state','unknown'))
|
||||||
|
" 2>/dev/null)
|
||||||
|
|
||||||
|
if [ "$TANKO_ZULIP" != "connected" ]; then
|
||||||
|
notify "🔴" "Tanko (Hermes) Zulip state: $TANKO_ZULIP — needs restart"
|
||||||
|
ISSUES=$((ISSUES + 1))
|
||||||
|
echo " Tanko: ❌ state=$TANKO_ZULIP" >> "$LOG"
|
||||||
|
else
|
||||||
|
echo " Tanko: ✅ Zulip connected" >> "$LOG"
|
||||||
|
fi
|
||||||
|
|
||||||
|
# ── Platform B: Hermes (Mumuni) ──
|
||||||
|
MUMUNI_STATE=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.123 \
|
||||||
|
"cat ~/.hermes/gateway_state.json 2>/dev/null" 2>/dev/null || echo "{}")
|
||||||
|
MUMUNI_ZULIP=$(echo "$MUMUNI_STATE" | python3 -c "
|
||||||
|
import sys,json
|
||||||
|
d=json.load(sys.stdin)
|
||||||
|
p=d.get('platforms',{}).get('zulip',{})
|
||||||
|
print(p.get('state','unknown'))
|
||||||
|
" 2>/dev/null)
|
||||||
|
|
||||||
|
if [ "$MUMUNI_ZULIP" != "connected" ]; then
|
||||||
|
notify "🔴" "Mumuni (Hermes) Zulip state: $MUMUNI_ZULIP"
|
||||||
|
ISSUES=$((ISSUES + 1))
|
||||||
|
echo " Mumuni: ❌ state=$MUMUNI_ZULIP" >> "$LOG"
|
||||||
|
else
|
||||||
|
echo " Mumuni: ✅ Zulip connected" >> "$LOG"
|
||||||
|
fi
|
||||||
|
|
||||||
|
# ── Platform C: Agent Zero (kagentz) ──
|
||||||
|
AZ_A2A=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.14 \
|
||||||
|
"docker exec agent-zero curl -s --connect-timeout 5 http://127.0.0.1:8001/.well-known/agent.json 2>/dev/null" 2>/dev/null || echo "")
|
||||||
|
AZ_ALIVE=$(echo "$AZ_A2A" | python3 -c "import sys,json; d=json.load(sys.stdin); print(d.get('name',''))" 2>/dev/null)
|
||||||
|
|
||||||
|
if [ "$AZ_ALIVE" != "kagentz" ]; then
|
||||||
|
notify "🔴" "kagentz A2A server DOWN — restarting"
|
||||||
|
ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.14 \
|
||||||
|
"docker exec agent-zero bash -c 'pkill -9 -f a2a_agent; sleep 1; cd /a0 && /opt/venv-a0/bin/python3 -u /a0/usr/a2a_agent.py > /tmp/a2a.log 2>&1 &'" 2>/dev/null || true
|
||||||
|
ISSUES=$((ISSUES + 1))
|
||||||
|
echo " kagentz: ❌ A2A down — restarted" >> "$LOG"
|
||||||
|
else
|
||||||
|
echo " kagentz: ✅ A2A alive" >> "$LOG"
|
||||||
|
|
||||||
|
# Check adapter process
|
||||||
|
AZ_ADAPTER=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.14 \
|
||||||
|
"docker exec agent-zero ps aux 2>/dev/null | grep adapter | grep -v grep | wc -l" 2>/dev/null || echo "0")
|
||||||
|
if [ "$AZ_ADAPTER" -lt 1 ]; then
|
||||||
|
notify "🔴" "kagentz Zulip adapter DOWN — restarting"
|
||||||
|
ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.14 \
|
||||||
|
"docker exec agent-zero bash -c 'cd /a0/usr/kagentz-zulip && ZULIP_SITE=https://chat.sysloggh.net ZULIP_EMAIL=kagentz-bot@chat.sysloggh.net ZULIP_API_KEY=E9q9PXJTxftPYBkb5pBDWupDO7KK21ty ZULIP_AGENT_NAME=kagentz A2A_URL=http://localhost:8001/a2a A2A_TOKEN=8zNgdOEXzYxjQvTl /opt/venv-a0/bin/python3 -u adapter.py > /tmp/zulip-adapter.log 2>&1 &'" 2>/dev/null || true
|
||||||
|
ISSUES=$((ISSUES + 1))
|
||||||
|
echo " kagentz: ❌ Adapter down — restarted" >> "$LOG"
|
||||||
|
else
|
||||||
|
echo " kagentz: ✅ Adapter running" >> "$LOG"
|
||||||
|
fi
|
||||||
|
fi
|
||||||
|
|
||||||
|
# ── Summary ──
|
||||||
|
if [ "$ISSUES" -eq 0 ]; then
|
||||||
|
echo " Result: ✅ All healthy" >> "$LOG"
|
||||||
|
else
|
||||||
|
echo " Result: 🔴 $ISSUES issue(s) found" >> "$LOG"
|
||||||
|
notify "🔴" "$ISSUES issue(s) found — check /root/zulip-health-monitor.log"
|
||||||
|
fi
|
||||||
|
|
||||||
|
echo "" >> "$LOG"
|
||||||
@@ -0,0 +1,235 @@
|
|||||||
|
---
|
||||||
|
kind: contract
|
||||||
|
name: zulip-health
|
||||||
|
title: Zulip Mesh Health Monitor — Multi-Platform
|
||||||
|
version: 2.0.0
|
||||||
|
agent: abiba
|
||||||
|
triggers:
|
||||||
|
- on startup
|
||||||
|
- every 15 minutes while running
|
||||||
|
- on zulip-status command
|
||||||
|
---
|
||||||
|
|
||||||
|
# Zulip Mesh Health Monitor — Multi-Platform
|
||||||
|
|
||||||
|
Monitors ALL Zulip-connected agents across three platforms.
|
||||||
|
Runs every 15 minutes in the background. Also triggers on session start.
|
||||||
|
|
||||||
|
## Platform Overview
|
||||||
|
|
||||||
|
| Platform | Agents | Bot | Adapter | Health Check |
|
||||||
|
|----------|--------|-----|---------|-------------|
|
||||||
|
| **pi** | Abiba (CT 100) | abiba-bot | `/root/.pi/agent/extensions/zulip/index.js` | `:9200/health` |
|
||||||
|
| **Hermes** | Tanko (CT 122), Mumuni (CT 123) | tanko-bot, mumuni-bot | `~/.hermes/plugins/platforms/zulip/` | `gateway_state.json` |
|
||||||
|
| **Agent Zero** | kagentz (CT 105) | kagentz-bot | Docker container, `/a0/usr/kagentz-zulip/` | A2A endpoint `:8001` |
|
||||||
|
|
||||||
|
## Phase 1: Zulip Server Liveness (All Platforms)
|
||||||
|
|
||||||
|
```bash
|
||||||
|
curl -s -o /dev/null -w "%{http_code}" https://chat.sysloggh.net/api/v1/server_settings \
|
||||||
|
-u 'abiba-bot@chat.sysloggh.net:$ZULIP_API_KEY'
|
||||||
|
```
|
||||||
|
|
||||||
|
Expected: `200`. Anything else → Server issue, alert maintainer.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Platform A: pi (Abiba — CT 100)
|
||||||
|
|
||||||
|
### A1: Health Endpoint
|
||||||
|
|
||||||
|
```js
|
||||||
|
const health = await fetch("http://localhost:9200/health").then(r => r.json());
|
||||||
|
```
|
||||||
|
|
||||||
|
| Field | Healthy | Critical |
|
||||||
|
|-------|---------|----------|
|
||||||
|
| `connected` | `true` | `false` |
|
||||||
|
| `last_error` | `null` | non-null string |
|
||||||
|
| `retry_count` | 0-2 | 3+ |
|
||||||
|
| `queue_id` | non-null string | `null` |
|
||||||
|
|
||||||
|
### A2: PM2 Process
|
||||||
|
|
||||||
|
```bash
|
||||||
|
pm2 show abiba-zulip --no-color 2>/dev/null
|
||||||
|
```
|
||||||
|
|
||||||
|
Check: `status=online`, `restarts < 10/h`, `uptime > 60s`
|
||||||
|
|
||||||
|
### A3: Echo Loop Detection
|
||||||
|
|
||||||
|
```bash
|
||||||
|
grep -a "Skipped.*bot msgs" /root/.pm2/logs/abiba-zulip-out.log | tail -5
|
||||||
|
```
|
||||||
|
|
||||||
|
> 100 skipped in 15min → 🟢 Info only (echo loop prevention working)
|
||||||
|
|
||||||
|
### A4: Response Delivery
|
||||||
|
|
||||||
|
```bash
|
||||||
|
grep -a "Finalized\|Failed to finalize" /root/.pm2/logs/abiba-zulip-out.log | tail -20
|
||||||
|
```
|
||||||
|
|
||||||
|
> 50% fail rate → 🔴 Critical — check editMessage API
|
||||||
|
|
||||||
|
### Actions
|
||||||
|
|
||||||
|
| Condition | Action |
|
||||||
|
|-----------|--------|
|
||||||
|
| `connected: false` | `pm2 restart abiba-zulip` |
|
||||||
|
| `retry_count >= 3` | `pm2 restart abiba-zulip` |
|
||||||
|
| `last_error` set | Log and monitor |
|
||||||
|
| Crash loop >10/h | Alert user |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Platform B: Hermes (Tanko — CT 122, Mumuni — CT 123)
|
||||||
|
|
||||||
|
### B1: Gateway State
|
||||||
|
|
||||||
|
```bash
|
||||||
|
ssh root@192.168.68.122 "cat ~/.hermes/gateway_state.json"
|
||||||
|
ssh root@192.168.68.123 "cat ~/.hermes/gateway_state.json"
|
||||||
|
```
|
||||||
|
|
||||||
|
Check `platforms.zulip.state`:
|
||||||
|
|
||||||
|
| Value | Meaning | Action |
|
||||||
|
|-------|---------|--------|
|
||||||
|
| `"connected"` | ✅ Healthy | None |
|
||||||
|
| `"disconnected"` | ❌ Disconnected | Check `last_error` |
|
||||||
|
| `"error"` | ❌ Error | Check `error_message`, restart gateway |
|
||||||
|
| missing | ❌ Not installed | Run deploy scripts |
|
||||||
|
|
||||||
|
### B2: Agent Process
|
||||||
|
|
||||||
|
```bash
|
||||||
|
ssh root@192.168.68.122 "ps aux | grep 'gateway run' | grep -v grep"
|
||||||
|
```
|
||||||
|
|
||||||
|
Gateway PID should exist and uptime > 60s.
|
||||||
|
|
||||||
|
### B3: Heartbeat Verification
|
||||||
|
|
||||||
|
```bash
|
||||||
|
ssh root@192.168.68.122 "grep Heartbeat ~/.hermes/logs/agent.log | tail -3"
|
||||||
|
```
|
||||||
|
|
||||||
|
Expected: recent heartbeat (within 5 min) showing `polls=N` incrementing.
|
||||||
|
If silence > 300s → 🟡 Warning (queue may be stuck).
|
||||||
|
If silence > 600s → 🔴 Critical (queue expired, adapter needs restart).
|
||||||
|
|
||||||
|
### B4: Response Delivery
|
||||||
|
|
||||||
|
```bash
|
||||||
|
ssh root@192.168.68.122 "grep -E 'Finalized|Failed to finalize|Replied to' ~/.hermes/logs/agent.log | tail -10"
|
||||||
|
```
|
||||||
|
|
||||||
|
> 50% fail rate → 🔴 Critical
|
||||||
|
|
||||||
|
### Actions
|
||||||
|
|
||||||
|
| Condition | Action |
|
||||||
|
|-----------|--------|
|
||||||
|
| `zulip.state != "connected"` | `ssh root@<CT> "pkill -f 'gateway run'; sleep 2; hermes gateway restart"` |
|
||||||
|
| No heartbeat in 10min | `ssh root@<CT> "pkill -f 'gateway run'; sleep 2; hermes gateway restart"` |
|
||||||
|
| `Failed to finalize` > 50% | Check PATCH API, Zulip server |
|
||||||
|
| Response empty/short | Check A2A endpoint / LiteLLM model |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Platform C: Agent Zero (kagentz — CT 105)
|
||||||
|
|
||||||
|
### C1: A2A Server Health
|
||||||
|
|
||||||
|
```bash
|
||||||
|
ssh root@192.168.68.14 "docker exec agent-zero curl -s --connect-timeout 5 http://127.0.0.1:8001/.well-known/agent.json"
|
||||||
|
```
|
||||||
|
|
||||||
|
Expected: `{"name":"kagentz",...}` — JSON response with agent identity.
|
||||||
|
Connection refused → A2A server is down.
|
||||||
|
|
||||||
|
### C2: Adapter Process
|
||||||
|
|
||||||
|
```bash
|
||||||
|
ssh root@192.168.68.14 "docker exec agent-zero ps aux | grep adapter | grep -v grep"
|
||||||
|
```
|
||||||
|
|
||||||
|
Adapter should be running. If missing, restart.
|
||||||
|
|
||||||
|
### C3: Heartbeat & Queue
|
||||||
|
|
||||||
|
```bash
|
||||||
|
ssh root@192.168.68.14 "docker exec agent-zero grep Heartbeat /tmp/zulip-adapter.log | tail -3"
|
||||||
|
```
|
||||||
|
|
||||||
|
Check:
|
||||||
|
- `processed=N` incrementing when DMs arrive
|
||||||
|
- `silence < 600s` (queue expiry timeout)
|
||||||
|
- `reconnects` — should be 0 under normal operation
|
||||||
|
|
||||||
|
### C4: A2A Response Verification
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Test A2A sends a task
|
||||||
|
ssh root@192.168.68.14 "docker exec agent-zero curl -s -X POST http://127.0.0.1:8001/a2a \
|
||||||
|
-H 'Content-Type: application/json' \
|
||||||
|
-d '{\"jsonrpc\":\"2.0\",\"method\":\"tasks/send\",\"params\":{\"message\":{\"role\":\"user\",\"parts\":[{\"text\":\"ping\"}]}},\"id\":1}'"
|
||||||
|
|
||||||
|
# Expected: task ID returned with "working" status
|
||||||
|
# Then poll for completion with tasks/get
|
||||||
|
```
|
||||||
|
|
||||||
|
### Actions
|
||||||
|
|
||||||
|
| Condition | Action |
|
||||||
|
|-----------|--------|
|
||||||
|
| A2A `.well-known/agent.json` fails | `docker exec agent-zero bash -c "pkill -9 -f a2a_agent; cd /a0 && /opt/venv-a0/bin/python3 -u /a0/usr/a2a_agent.py > /tmp/a2a.log 2>&1 &"` |
|
||||||
|
| Adapter process missing | `docker exec agent-zero bash -c "cd /a0/usr/kagentz-zulip && ZULIP_SITE=... ZULIP_EMAIL=... ZULIP_API_KEY=... A2A_URL=http://localhost:8001/a2a /opt/venv-a0/bin/python3 -u adapter.py > /tmp/zulip-adapter.log 2>&1 &"` |
|
||||||
|
| Silence > 600s (queue expiry) | Restart adapter (fix deployed: auto-reconnect on BAD_EVENT_QUEUE_ID) |
|
||||||
|
| LiteLLM 401 | Check API key in a2a_agent.py `LITELLM_KEY` |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Global Checks
|
||||||
|
|
||||||
|
### Zulip Server
|
||||||
|
|
||||||
|
```bash
|
||||||
|
curl -s https://chat.sysloggh.net/api/v1/server_settings \
|
||||||
|
-u 'abiba-bot@chat.sysloggh.net:$ZULIP_API_KEY' -o /dev/null -w "%{http_code}"
|
||||||
|
```
|
||||||
|
|
||||||
|
Expected: `200`
|
||||||
|
|
||||||
|
### Cross-Agent Echo Loop Detection
|
||||||
|
|
||||||
|
Check each agent's log for excessive bot-to-bot chatter:
|
||||||
|
- Abiba: `Skipped.*bot msgs` count
|
||||||
|
- Tanko/Mumuni: Check for repeated DM exchanges between bots
|
||||||
|
- kagentz: Check adapter log for bot DMs being processed
|
||||||
|
|
||||||
|
If any bot is processing >50 bot-originated messages in 15min → 🟡 Warning.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Consolidated Action Matrix
|
||||||
|
|
||||||
|
| Condition | Severity | Action |
|
||||||
|
|-----------|----------|--------|
|
||||||
|
| Zulip server not 200 | 🔴 Critical | Alert maintainer |
|
||||||
|
| All agents silent | 🔴 Critical | Zulip server likely down |
|
||||||
|
| pi: connected false | 🔴 Critical | `pm2 restart abiba-zulip` |
|
||||||
|
| pi: edit fail >50% | 🔴 Critical | Check editMessage API |
|
||||||
|
| Hermes: zulip disconnected | 🔴 Critical | Restart gateway |
|
||||||
|
| Hermes: no heartbeat 10min | 🔴 Critical | Restart gateway |
|
||||||
|
| Az: A2A server down | 🔴 Critical | Restart inside container |
|
||||||
|
| Az: adapter down | 🔴 Critical | Restart adapter |
|
||||||
|
| Az: silence >600s | 🟡 Warning | Queue expired (auto-recover) |
|
||||||
|
| Echo loop detected | 🟢 Info | Auto-mitigated (bot filtering) |
|
||||||
|
|
||||||
|
## Logging
|
||||||
|
|
||||||
|
All diagnostics logged to `/root/zulip-health-monitor.log` with timestamps.
|
||||||
|
Critical alerts sent as relay messages to user.
|
||||||
@@ -0,0 +1,72 @@
|
|||||||
|
---
|
||||||
|
kind: responsibility
|
||||||
|
name: zulip-mention-reliability
|
||||||
|
description: >
|
||||||
|
Monitors whether Abiba Bot is correctly detecting and responding to
|
||||||
|
@mentions in Zulip stream topics. Tests periodically, diagnoses
|
||||||
|
failures, and attempts fixes. Same self-improving pattern as
|
||||||
|
litellm-self-heal.
|
||||||
|
---
|
||||||
|
|
||||||
|
## Maintains
|
||||||
|
|
||||||
|
- mention_response_rate: number — % of @mentions that get a response (target: >95%)
|
||||||
|
- last_test_result: "pass" | "fail" | "degraded"
|
||||||
|
- known_failure_modes: array — History of what went wrong and how it was fixed
|
||||||
|
- queue_health: "healthy" | "expired" | "reconnecting"
|
||||||
|
|
||||||
|
## Success Criteria
|
||||||
|
|
||||||
|
- test_mention_gets_response == true — Send @Abiba Bot test, confirm response within 30s
|
||||||
|
- stream_subscription_active == true — Bot is subscribed to #agent-hub and #general chat
|
||||||
|
- queue_active == true — PM2's event queue is registered and polling
|
||||||
|
- no_echo_loop == true — Bot doesn't respond to its own messages
|
||||||
|
- dm_always_works == true — DMs still respond even if streams are broken
|
||||||
|
|
||||||
|
## Known Failure Modes (from History)
|
||||||
|
|
||||||
|
| Failure | Root Cause | Fix |
|
||||||
|
|---|---|---|
|
||||||
|
| No stream events | Bot had 0 stream subscriptions | Subscribe via API |
|
||||||
|
| No @mention detection | botUserId was null (register endpoint doesn't return it) | Fetch from /api/v1/users/me |
|
||||||
|
| No stream events | zulip-js library sends multipart/form-data instead of form-urlencoded | Use URLSearchParams directly |
|
||||||
|
| Double responses | Both TUI and PM2 loaded the extension | ZULIP_EXTENSION_ACTIVE guard |
|
||||||
|
|
||||||
|
## Remediation Rules
|
||||||
|
|
||||||
|
### Rule 1: Queue Expired
|
||||||
|
- Detect: Poll returns BAD_EVENT_QUEUE_ID
|
||||||
|
- Fix: `pm2 restart abiba-zulip`
|
||||||
|
- Verify: Check logs for "Connected, queue:"
|
||||||
|
|
||||||
|
### Rule 2: Subscription Missing
|
||||||
|
- Detect: `GET /api/v1/users/me/subscriptions` returns 0 for #agent-hub
|
||||||
|
- Fix: Subscribe via `POST /api/v1/users/me/subscriptions`
|
||||||
|
- Verify: Re-check subscriptions
|
||||||
|
|
||||||
|
### Rule 3: Extension Crashed
|
||||||
|
- Detect: PM2 status shows "errored" for abiba-zulip
|
||||||
|
- Fix: `pm2 restart abiba-zulip`
|
||||||
|
- Escalate: If restarts >5 in last hour, alert user
|
||||||
|
|
||||||
|
### Rule 4: No Response to Test Mention
|
||||||
|
- Detect: Send test @mention, wait 30s, check if response appeared
|
||||||
|
- Fix: Run diagnostics (check queue, check code, check logs)
|
||||||
|
- Log: Record failure mode as new entry in Known Failure Modes
|
||||||
|
|
||||||
|
## Continuity
|
||||||
|
|
||||||
|
- **On user report**: User says "you didn't respond" → run diagnostic
|
||||||
|
- **On test failure**: Test mention fails → run remediation
|
||||||
|
- **Retrospective**: Every 6 hours — send a test @mention to self-verify
|
||||||
|
- **On PM2 restart**: Always run a self-test after restart
|
||||||
|
|
||||||
|
## Execution
|
||||||
|
|
||||||
|
1. **Self-test**: Send `@**Abiba Bot** _selftest_` in #agent-hub, wait for response
|
||||||
|
2. **If no response in 30s**: Run diagnostics
|
||||||
|
3. **Diagnose**: Check queue, subscriptions, PM2 status, logs
|
||||||
|
4. **Fix**: Apply matching remediation rule
|
||||||
|
5. **Retest**: Try again
|
||||||
|
6. **Log**: Record success or new failure mode
|
||||||
|
7. **Report**: DM user only if escalation needed
|
||||||
Reference in New Issue
Block a user