Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
9e87927444 | ||
|
|
b0683e9566 | ||
|
|
dd11c8f14f | ||
|
|
0732eed329 | ||
|
|
59ed7cdbf7 | ||
|
|
5d70bbf25b | ||
|
|
ccc916d1ec | ||
|
|
48fb263d4b | ||
|
|
64790ebd19 | ||
|
|
6fa0a255df |
@@ -38,6 +38,7 @@ index:
|
||||
by_category:
|
||||
compliance:
|
||||
- hermes-key-enforcement
|
||||
- litellm-api-keys
|
||||
- hermes-config-template
|
||||
- hermes-agent-baseline
|
||||
monitoring:
|
||||
@@ -101,6 +102,7 @@ index:
|
||||
proxmox:
|
||||
- proxmox-monitor
|
||||
litellm:
|
||||
- litellm-api-keys
|
||||
- litellm-health
|
||||
- litellm-self-heal
|
||||
memory:
|
||||
@@ -1869,6 +1871,58 @@ contracts:
|
||||
drift_alerts: []
|
||||
# Koby Report-Only Registry (2026-08-17 — Captain)
|
||||
# ⛔ KOBY IS NEVER REPAIRED — detect + report, never fix on .129
|
||||
- name: litellm-api-keys
|
||||
file: litellm-api-keys.prose.md
|
||||
kind: function
|
||||
category: compliance
|
||||
sensitivity: critical
|
||||
status: active
|
||||
owner: abiba
|
||||
version: 1.1.0
|
||||
trigger:
|
||||
type: on_demand
|
||||
cadence: null
|
||||
description: "Manual invocation when creating/rotating/verifying agent LiteLLM keys"
|
||||
cron_job_id: null
|
||||
execution:
|
||||
agent: abiba
|
||||
timeout: 120
|
||||
requires: []
|
||||
protocol:
|
||||
- Load contract from prose-contracts/main
|
||||
- Retrieve master key from Infisical (project=infrastructure env=production)
|
||||
- Read live key-scoped model roster from CT 116 /v1/models
|
||||
- Create/rotate/verify the requested agent key with an EXPLICIT models list
|
||||
- 'Never create a key with an empty models list or all-proxy-models (Cloud leak)'
|
||||
verification:
|
||||
postconditions:
|
||||
- check: standard agent key is local-only
|
||||
verify: 'curl -s -H "Authorization: Bearer <KEY>" http://192.168.68.116/litellm/v1/models | jq -r ''.data[].id'' | grep -c /'
|
||||
expect: 0 cloud models
|
||||
- check: key exists with correct alias
|
||||
verify: 'curl -s -H "Authorization: Bearer <MASTER>" http://192.168.68.116/litellm/v1/key/info?key_alias=<AGENT>'
|
||||
expect: 200 with matching alias
|
||||
artifact: key creation/rotation report
|
||||
receipt:
|
||||
format: json
|
||||
storage: ~/.hermes/runs/litellm-api-keys/
|
||||
graph_node: true
|
||||
escalation:
|
||||
info:
|
||||
action: log_to_receipt
|
||||
notify: []
|
||||
warning:
|
||||
action: relay_alert
|
||||
notify:
|
||||
- abiba
|
||||
- mumuni
|
||||
critical:
|
||||
action: relay_alert
|
||||
notify:
|
||||
- abiba
|
||||
- mumuni
|
||||
- ops
|
||||
|
||||
koby_report_only: true
|
||||
koby_host: "CT 111 (tdunna)"
|
||||
koby_ip: ".129"
|
||||
|
||||
@@ -18,6 +18,25 @@ author: Abiba (pi agent)
|
||||
|
||||
**All harness/litellm providers MUST use `api_key_env: LITELLM_API_KEY` with canonical internal path `http://192.168.68.116/litellm/v1` (Hermes appends `/v1/responses`) or public path `https://litellm.sysloggh.net/v1` — hardcoded keys AND direct `:4000` access are both forbidden. Internal `/v1` still works but is non-canonical (WARN, not FAIL).**
|
||||
|
||||
**Cloud provider models (OpenRouter, DeepSeek, Google AI Studio, QwenCloud PAYG/Plan, Tencent TokenHub PAYG/Plan) added to CT 116 on 2026-09-20 are reachable ONLY through the designated cloud-enabled key. Standard agent keys remain LOCAL-ONLY (`strix-moe`, `gpu-dense`, `gpu-vision`, `syslog-auto`) and MUST NOT be granted cloud models unless explicitly approved by the captain.**
|
||||
|
||||
## Model Access Tiers (2026-09-20)
|
||||
|
||||
CT 116 hosts two **access tiers** of model. Tier membership is enforced per virtual key via that key's `models` allowlist.
|
||||
|
||||
| Tier | Models | Who gets it |
|
||||
|------|--------|-------------|
|
||||
| **Local** | `strix-moe`, `gpu-dense`, `gpu-vision`, `syslog-auto` | All standard agent keys (tanko, mumuni, koby, koonimo, abiba-pi) |
|
||||
| **Cloud** | 47 provider models: `openrouter/*`, `deepseek/*`, `google/*`, `qwen-payg/*`, `qwen-plan/*`, `tencent-payg/*`, `tencent-plan/*` | **Only** the designated cloud-enabled key (captain decision) |
|
||||
|
||||
**Rules:**
|
||||
1. A key with an empty `models` list (`{}`) or `all-proxy-models` is UNSCOPED — it silently gains ALL models, including cloud. Never create or leave an agent key in this state.
|
||||
2. Agent keys MUST carry an **explicit local-only** `models` list.
|
||||
3. Granting a cloud model to an agent key requires explicit captain approval and a recorded reason.
|
||||
4. The master key always bypasses scoping — it is admin-only, never for inference.
|
||||
|
||||
See `litellm-api-keys` § Cloud Provider Consolidation for the key-creation procedure.
|
||||
|
||||
## Scope
|
||||
|
||||
Applies to all Hermes agent configs across all hosts. Covers these config sections:
|
||||
|
||||
@@ -71,6 +71,10 @@ description: >
|
||||
- Set models: read the live key-scoped set rather than hardcoding one — `/v1/models` is key-scoped,
|
||||
and the authoritative registry is CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not add
|
||||
retired names (`gemma-4-12b`, `gpu-light`, `crew-auto` — all retired 2026-09-12).
|
||||
- **Standard agent keys are LOCAL-ONLY**: `["strix-moe", "gpu-dense", "gpu-vision", "syslog-auto"]`.
|
||||
Cloud models are granted ONLY to the designated cloud-enabled key — see § Cloud Provider
|
||||
Consolidation. NEVER create a key with an empty `models` list (`{}`) or `all-proxy-models`.
|
||||
In LiteLLM Community both silently grant access to EVERY model, including cloud.
|
||||
- Note: `ornith-1.0-35b` is NOT a valid LiteLLM model name (use `strix-moe`, the stable alias). qwen3.6-35B-A3B removed from fleet (was never deployed).
|
||||
- Return the new key
|
||||
5. **If action == "rotate"**:
|
||||
@@ -87,6 +91,66 @@ description: >
|
||||
- Confirm key alias matches agent_name in LiteLLM key list
|
||||
- Verify agent gateway uses vault wrapper: `cat /proc/<pid>/cmdline` shows `infisical run`
|
||||
|
||||
## Cloud Provider Consolidation (2026-09-20)
|
||||
|
||||
CT 116 LiteLLM (Community v1.99.1) fronts **7 upstream providers** in addition to the local
|
||||
GPU models. Added 2026-09-20 — 47 cloud deployments, 51 unique model names total.
|
||||
|
||||
### Provider map (per-account namespacing)
|
||||
|
||||
Two accounts on the same vendor get **distinct prefixes** so billing, rate limits, and keys
|
||||
stay separate:
|
||||
|
||||
| Prefix | Upstream | Auth | Vault secret |
|
||||
|--------|----------|------|--------------|
|
||||
| `openrouter/` | OpenRouter | API key | `OPENROUTER_API_KEY` |
|
||||
| `deepseek/` | DeepSeek direct | API key | `DEEPSEEK_API_KEY` |
|
||||
| `google/` | Google AI Studio (Gemini) | API key | `GEMINI_API_KEY` |
|
||||
| `qwen-payg/` | QwenCloud / DashScope (pay-as-you-go) | API key | `DASHSCOPE_PAYG_KEY` |
|
||||
| `qwen-plan/` | QwenCloud / DashScope (token plan) | API key | `DASHSCOPE_PLAN_KEY` |
|
||||
| `tencent-payg/` | Tencent TokenHub (PAYG) | Bearer token | `TENCENT_PAYG_KEY` |
|
||||
| `tencent-plan/` | Tencent TokenHub (Plan) | Bearer token | `TENCENT_PLAN_KEY` |
|
||||
|
||||
All cloud `api_key` fields use `os.environ/<NAME>` — the 7 secrets live in Infisical
|
||||
(project=`infrastructure`, env=`production`, folder=`root`) and are injected into the
|
||||
`harness-litellm` container at start. **No literal cloud keys in `litellm_config.yaml`.**
|
||||
|
||||
### Access tiers (MUST be enforced per key)
|
||||
|
||||
| Tier | Model names | Granted to |
|
||||
|------|-------------|------------|
|
||||
| **Local** | `strix-moe`, `gpu-dense`, `gpu-vision`, `syslog-auto` | every standard agent key |
|
||||
| **Cloud** | the 47 provider models (`<prefix>/<model>`) | **only** the designated cloud-enabled key |
|
||||
|
||||
> ⚠️ **Community-edition caveat:** LiteLLM Community does not restrict wildcard access groups
|
||||
> the way Enterprise does. Access is decided by each key's explicit `models` list. A key with
|
||||
> `models = {}` or `models = ["all-proxy-models"]` sees **all** models — a silent cloud leak.
|
||||
> Every key MUST carry an explicit list. The master key always bypasses scoping (admin-only).
|
||||
|
||||
### Creating the cloud-enabled key
|
||||
|
||||
```bash
|
||||
# ALWAYS read the live roster first (key-scoped):
|
||||
curl -s -H "Authorization: Bearer <AGENT_KEY>" http://192.168.68.116/litellm/v1/models \
|
||||
| jq -r '.data[].id'
|
||||
|
||||
# Then generate a key with an EXPLICIT model list (never empty, never a wildcard).
|
||||
# For the cloud-enabled key, list local + cloud. For a standard agent, local only.
|
||||
```
|
||||
|
||||
Verify after any key change: a standard agent key must return **4 models**, and must NOT return
|
||||
any `<prefix>/` cloud model.
|
||||
|
||||
### Adding a new cloud provider
|
||||
|
||||
1. Add the upstream key to Infisical `infrastructure/production/root`.
|
||||
2. Add the deployment(s) to `/opt/inference-harness/litellm_config.yaml` with an `os.environ/` ref
|
||||
and a namespaced `model_name` (`<provider>-<account>/<model>` when a vendor has >1 account).
|
||||
3. Restart the `harness-litellm` container.
|
||||
4. Grant the model to the cloud-enabled key ONLY (explicit list) — never to agent keys without
|
||||
captain approval.
|
||||
5. Update this table and the access-tier section.
|
||||
|
||||
## Production Vault Access Process (canonical, 2026-07-17)
|
||||
|
||||
The non-fail approach to agentic vault access. Deployed on all 4 Hermes agents
|
||||
|
||||
+20
-10
@@ -2,10 +2,10 @@
|
||||
kind: responsibility
|
||||
name: pm2-self-heal
|
||||
description: >
|
||||
PM2 process health check for abiba-telegram, abiba-zulip, gitea-runner, and zulip-watchdog.
|
||||
gpu-monitor is systemd-managed (gpu-monitor.service), NOT PM2.
|
||||
gpu-watchdog is decommissioned and folded into gpu-monitor.service.
|
||||
gitea-runner is KEPT. abiba-zulip is KEPT (online for days).
|
||||
Monitors critical PM2 processes (abiba-telegram, abiba-zulip, gitea-runner, zulip-watchdog)
|
||||
and auto-restarts any that are stopped or errored. Logs every action to
|
||||
Gitea health-logs and alerts the owner via Telegram (primary) or Zulip DM (secondary).
|
||||
Abiba-zulip is the live Zulip bridge and may be restarted; alert owner on failure.
|
||||
---
|
||||
|
||||
## Maintains
|
||||
@@ -16,6 +16,8 @@ description: >
|
||||
- zulip-watchdog: { status: "online", uptime: string, restarts: number }
|
||||
- last_check: timestamp
|
||||
|
||||
> **Status (2026-08-03):** `abiba-zulip` fully restored — live Zulip bridge, heartbeating, monitored. `gpu-monitor` runs via systemd only (gpu-monitor.service); NOT PM2-tracked. `gpu-watchdog` retired from PM2 (folded into gpu-monitor.service). `zulip-watchdog` remains live and PM2-managed.
|
||||
|
||||
|
||||
## Continuity
|
||||
|
||||
@@ -40,7 +42,7 @@ description: >
|
||||
crash-loop that never toggles status to "stopped"/"errored" (e.g. the 10k-restarts
|
||||
spoton incident). Alerts include the restart count.
|
||||
- **AS-BUILT (2026-09-15)**: spoton-service was deleted with its app; the live PM2 set is
|
||||
four processes (abiba-telegram, abiba-zulip, gitea-runner, zulip-watchdog). The spoton
|
||||
four processes (abiba-telegram, abiba-zulip, gitea-runner, gpu-monitor). The spoton
|
||||
reference above is historical context for the crash-loop guard, not a live process.
|
||||
- **Escalate**: Only when restarts > 30 — alerts to Zulip DM
|
||||
- **Historical fix**: Previous cycles were caused by abiba-zulip extension's
|
||||
@@ -51,17 +53,25 @@ description: >
|
||||
## Execution
|
||||
|
||||
1. **Check PM2 status** — Run `pm2 status --no-color` and parse the table (5th data column = PID, 8th = restarts, 9th = status)
|
||||
2. **Check abiba-telegram**:
|
||||
2. **Check abiba-telegram** (safe to auto-restart):
|
||||
- If status is "online" → pass
|
||||
- If status is "stopped" or "errored" → apply Rule 1
|
||||
- If restarts > 5 → alert owner
|
||||
- If restarts > 1000 → apply Rule 2 (crash-loop guard)
|
||||
3. **Check abiba-zulip** (live Zulip bridge, heartbeating):
|
||||
- If status is "online" → pass, log restarts count
|
||||
- If status is "stopped" or "errored" → restart (`pm2 restart abiba-zulip` — fully restored)
|
||||
- If restarts > 5 in last hour → alert owner with full diagnostics
|
||||
4. **Log results** — Append to `SyslogSolution/health-logs/pm2/{timestamp}.md` in Gitea (not knowledge graph — hard rule)
|
||||
5. **Alert** — Send Zulip DM to owner if escalation needed (do NOT run pm2 commands during alerting)
|
||||
6. **Wait 5 min** → repeat from step 1
|
||||
4. **Check gitea-runner**:
|
||||
- If status is "online" → pass
|
||||
- If status is "stopped" or "errored" → apply Rule 1
|
||||
- If restarts > 5 → alert owner
|
||||
5. **Check zulip-watchdog**:
|
||||
- If status is "online" → pass
|
||||
- If status is "stopped" or "errored" → apply Rule 1
|
||||
- If restarts > 5 → alert owner
|
||||
6. **Log results** — Append to `SyslogSolution/health-logs/pm2/{timestamp}.md` in Gitea (not knowledge graph — hard rule)
|
||||
7. **Alert** — Send Telegram (primary) or Zulip DM (secondary) to owner if escalation needed (do NOT run pm2 commands during alerting)
|
||||
8. **Wait 5 min** → repeat from step 1
|
||||
|
||||
## Example Output (when healthy)
|
||||
|
||||
|
||||
@@ -22,10 +22,12 @@ AUTH = "Authorization: PVEAPIToken=«vault: infrastructure/production PVE_API_TO
|
||||
|
||||
ZULIP_SITE = "https://chat.sysloggh.net"
|
||||
ZULIP_EMAIL = "abiba-bot@chat.sysloggh.net"
|
||||
ZULIP_API_KEY = os.environ.get("ZULIP_API_KEY", "")
|
||||
if not ZULIP_API_KEY:
|
||||
raise SystemExit("ZULIP_API_KEY not set — refusing to run with no credential")
|
||||
ZULIP_AUTH = f"{ZULIP_EMAIL}:{ZULIP_API_KEY}"
|
||||
# Note: /api/v1/server_settings is a PUBLIC endpoint (verified HTTP 200 with or without credential).
|
||||
# No Zulip API key is required for this call. If a future leg genuinely needs abiba-bot's key,
|
||||
# it must prove it with a 200 from /api/v1/users/me as abiba-bot and label itself degraded when it cannot.
|
||||
# Never fall back to the vault's shared ZULIP_API_KEY.
|
||||
ZULIP_AUTH = None
|
||||
DEGRADED_LEGS = []
|
||||
|
||||
LITELLM_PUBLIC = "https://litellm.sysloggh.net"
|
||||
LITELLM_BACKEND = "192.168.68.116"
|
||||
@@ -158,7 +160,7 @@ def collect():
|
||||
# ── Storage ──
|
||||
storages = pve_get("/api2/json/nodes/storepve/storage")
|
||||
report["storage"] = []
|
||||
for s in storages:
|
||||
for s in (storages or []):
|
||||
total = s.get("total",0) or 1
|
||||
used = s.get("used",0)
|
||||
pct = used/total*100
|
||||
@@ -673,8 +675,9 @@ def send_email(html_content, subject_prefix=""):
|
||||
try:
|
||||
EMAIL_PASSWORD = os.environ.get("EMAIL_PASSWORD") or os.environ.get("SMTP_PASSWORD") or os.environ.get("MAIL_PASSWORD")
|
||||
if not EMAIL_PASSWORD:
|
||||
print("EMAIL_PASSWORD not set — refusing to send email", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
print(" ⚠️ Degraded leg: credential-missing: EMAIL_PASSWORD (or SMTP_PASSWORD/MAIL_PASSWORD)", file=sys.stderr)
|
||||
DEGRADED_LEGS.append("credential-missing: EMAIL_PASSWORD")
|
||||
return True, "✅ Email leg degraded (no credential) — report still produced"
|
||||
GMAIL_EMAIL = "jtabiri@gmail.com"
|
||||
|
||||
server = smtplib.SMTP("smtp.gmail.com", 587)
|
||||
@@ -713,6 +716,17 @@ if __name__ == "__main__":
|
||||
print(f" {msg}")
|
||||
|
||||
# Show summary
|
||||
if DEGRADED_LEGS:
|
||||
print(f"\n⚠️ Degraded legs ({len(DEGRADED_LEGS)}):")
|
||||
for leg in DEGRADED_LEGS:
|
||||
print(f" - {leg}")
|
||||
else:
|
||||
print("\n✅ All legs fully credentialed")
|
||||
|
||||
# A failed send must exit non-zero; a degraded leg (no credential) must stay exit 0
|
||||
if not ok:
|
||||
sys.exit(1)
|
||||
|
||||
issues = sum(1 for i in ["red"] if report.get("zulip_ext", {}).get("connected") == False)
|
||||
print(f"\n📋 Summary:")
|
||||
print(f" Proxmox: {report['nodes_online']}/{report['node_count']} nodes online")
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
#!/bin/bash
|
||||
# pm2-self-heal — hourly PM2 process check
|
||||
# Part of the pm2-self-heal prose contract
|
||||
# Alerts via Telegram (abiba-zulip decommissioned 2026-07-04)
|
||||
# Alerts via Telegram (primary) and Zulip DM (secondary, abiba-zulip restored 2026-08-03)
|
||||
# Field positions (awk -F'│'): $7=pid $8=uptime $9=restarts $10=status
|
||||
|
||||
TELEGRAM_BOT_TOKEN="$(grep TELEGRAM_BOT_TOKEN /root/.pi/agent/extensions/telegram/.env 2>/dev/null | cut -d= -f2 || echo '')"
|
||||
@@ -46,9 +46,75 @@ if [ "$TEL_STATUS" != "online" ] || [ "$TEL_RESTARTS" -gt 1000 ]; then
|
||||
fi
|
||||
fi
|
||||
|
||||
# Check abiba-zulip (live Zulip bridge, heartbeating)
|
||||
ZULIP_LINE=$(echo "$STATUS" | grep "abiba-zulip")
|
||||
ZULIP_STATUS=$(echo "$ZULIP_LINE" | awk -F'│' '{print $10}' | xargs)
|
||||
ZULIP_RESTARTS=$(echo "$ZULIP_LINE" | awk -F'│' '{print $9}' | xargs)
|
||||
|
||||
if [ "$ZULIP_STATUS" != "online" ]; then
|
||||
pm2 restart abiba-zulip > /dev/null 2>&1
|
||||
sleep 3
|
||||
ZULIP_LINE2=$(pm2 status --no-color 2>/dev/null | grep "abiba-zulip")
|
||||
ZULIP_STATUS2=$(echo "$ZULIP_LINE2" | awk -F'│' '{print $10}' | xargs)
|
||||
if [ "$ZULIP_STATUS2" = "online" ]; then
|
||||
msg="⚠️ abiba-zulip was **$ZULIP_STATUS** → restarted to online"
|
||||
ALERTS="${ALERTS}${msg}\n"
|
||||
else
|
||||
msg="🚨 abiba-zulip **failed restart** (was $ZULIP_STATUS, still $ZULIP_STATUS2)"
|
||||
ALERTS="${ALERTS}${msg}\n"
|
||||
fi
|
||||
elif [ "$ZULIP_RESTARTS" -gt 5 ]; then
|
||||
msg="⚠️ abiba-zulip has **$ZULIP_RESTARTS** restarts (high count)"
|
||||
ALERTS="${ALERTS}${msg}\n"
|
||||
fi
|
||||
|
||||
# Check gitea-runner
|
||||
GITEA_LINE=$(echo "$STATUS" | grep "gitea-runner")
|
||||
GITEA_STATUS=$(echo "$GITEA_LINE" | awk -F'│' '{print $10}' | xargs)
|
||||
GITEA_RESTARTS=$(echo "$GITEA_LINE" | awk -F'│' '{print $9}' | xargs)
|
||||
|
||||
if [ "$GITEA_STATUS" != "online" ]; then
|
||||
pm2 restart gitea-runner > /dev/null 2>&1
|
||||
sleep 3
|
||||
GITEA_LINE2=$(pm2 status --no-color 2>/dev/null | grep "gitea-runner")
|
||||
GITEA_STATUS2=$(echo "$GITEA_LINE2" | awk -F'│' '{print $10}' | xargs)
|
||||
if [ "$GITEA_STATUS2" = "online" ]; then
|
||||
msg="⚠️ gitea-runner was **$GITEA_STATUS** → restarted to online"
|
||||
ALERTS="${ALERTS}${msg}\n"
|
||||
else
|
||||
msg="🚨 gitea-runner **failed restart** (was $GITEA_STATUS, still $GITEA_STATUS2)"
|
||||
ALERTS="${ALERTS}${msg}\n"
|
||||
fi
|
||||
elif [ "$GITEA_RESTARTS" -gt 5 ]; then
|
||||
msg="⚠️ gitea-runner has **$GITEA_RESTARTS** restarts (high count)"
|
||||
ALERTS="${ALERTS}${msg}\n"
|
||||
fi
|
||||
|
||||
# Check zulip-watchdog
|
||||
WATCHDOG_LINE=$(echo "$STATUS" | grep "zulip-watchdog")
|
||||
WATCHDOG_STATUS=$(echo "$WATCHDOG_LINE" | awk -F'│' '{print $10}' | xargs)
|
||||
WATCHDOG_RESTARTS=$(echo "$WATCHDOG_LINE" | awk -F'│' '{print $9}' | xargs)
|
||||
|
||||
if [ "$WATCHDOG_STATUS" != "online" ]; then
|
||||
pm2 restart zulip-watchdog > /dev/null 2>&1
|
||||
sleep 3
|
||||
WATCHDOG_LINE2=$(pm2 status --no-color 2>/dev/null | grep "zulip-watchdog")
|
||||
WATCHDOG_STATUS2=$(echo "$WATCHDOG_LINE2" | awk -F'│' '{print $10}' | xargs)
|
||||
if [ "$WATCHDOG_STATUS2" = "online" ]; then
|
||||
msg="⚠️ zulip-watchdog was **$WATCHDOG_STATUS** → restarted to online"
|
||||
ALERTS="${ALERTS}${msg}\n"
|
||||
else
|
||||
msg="🚨 zulip-watchdog **failed restart** (was $WATCHDOG_STATUS, still $WATCHDOG_STATUS2)"
|
||||
ALERTS="${ALERTS}${msg}\n"
|
||||
fi
|
||||
elif [ "$WATCHDOG_RESTARTS" -gt 5 ]; then
|
||||
msg="⚠️ zulip-watchdog has **$WATCHDOG_RESTARTS** restarts (high count)"
|
||||
ALERTS="${ALERTS}${msg}\n"
|
||||
fi
|
||||
|
||||
# Log check
|
||||
{
|
||||
echo "[$(date '+%Y-%m-%d %H:%M:%S')] tel=$TEL_STATUS alerts=${ALERTS:+yes}"
|
||||
echo "[$(date '+%Y-%m-%d %H:%M:%S')] tel=$TEL_STATUS zulip=$ZULIP_STATUS gitea=$GITEA_STATUS watchdog=$WATCHDOG_STATUS alerts=${ALERTS:+yes}"
|
||||
[ -n "$ALERTS" ] && echo "$ALERTS"
|
||||
} >> "$LOG"
|
||||
|
||||
|
||||
Reference in New Issue
Block a user