fix(agent-health-check): correct unit names, abiba key path, pid crash, uptime math (script-only) #65

Merged
abiba-bot merged 4 commits from fm/agent-health-probe-repoint-20260907 into master 2026-09-09 01:39:30 +00:00
2 changed files with 2 additions and 3 deletions
Showing only changes of commit 2e0b737f2d - Show all commits
+1 -2
View File
@@ -220,8 +220,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
| `/root/scripts/gpu_benchmark.py` | pi (.24) | GPU inference benchmark module (tok/s tracking) |
| `/root/scripts/gpu-saturation-watchdog.py` | pi (.24) | Auto-restart stuck llama-server |
| `/root/dashboard/gpu-fleet.html` | pi (.24) | Live HTML dashboard |
| `/etc/systemd/system/llama-chat-api.service` | .8 | RTX 3090 GPU daemon — active unit since 2026-09-06; the old `llama-server.service` file on .8 is stale/inactive |
| `/etc/systemd/system/llama-server.service` | .110 | llama-server daemon (RTX 5070) |
| `/etc/systemd/system/llama-server.service` | .8, .110 | llama-server daemons (Nvidia GPUs) |
| `/etc/systemd/system/strix-server.service` | .15 (amdpve) | llama-server daemon (Vulkan, Strix Halo) running unsloth/Qwen3.6-35B-A3B-MTP-GGUF. Note: `llama-server.service` and `llama-server@.service` are **masked** on .15 to prevent port 8080 collisions. |
## Prometheus & Grafana
+1 -1
View File
@@ -593,7 +593,7 @@ curl -s -X POST http://192.168.68.116:9000/admin/roster/reload \
-H "Authorization: Bearer sk-admin-ee09fffd04978b61a1569ac670c68814"
# Restart stuck GPU (saturation watchdog alternative)
ssh root@192.168.68.8 "systemctl restart llama-chat-api.service" # .8 active unit since 2026-09-06; llama-server.service file on .8 is stale/inactive
ssh root@192.168.68.8 "systemctl restart llama-server"
ssh root@192.168.68.110 "systemctl restart llama-server"
```