fix(agent-health-check): correct unit names, abiba key path, pid crash, uptime math (script-only) #65
+1
-2
@@ -220,8 +220,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
|
||||
| `/root/scripts/gpu_benchmark.py` | pi (.24) | GPU inference benchmark module (tok/s tracking) |
|
||||
| `/root/scripts/gpu-saturation-watchdog.py` | pi (.24) | Auto-restart stuck llama-server |
|
||||
| `/root/dashboard/gpu-fleet.html` | pi (.24) | Live HTML dashboard |
|
||||
| `/etc/systemd/system/llama-chat-api.service` | .8 | RTX 3090 GPU daemon — active unit since 2026-09-06; the old `llama-server.service` file on .8 is stale/inactive |
|
||||
| `/etc/systemd/system/llama-server.service` | .110 | llama-server daemon (RTX 5070) |
|
||||
| `/etc/systemd/system/llama-server.service` | .8, .110 | llama-server daemons (Nvidia GPUs) |
|
||||
| `/etc/systemd/system/strix-server.service` | .15 (amdpve) | llama-server daemon (Vulkan, Strix Halo) running unsloth/Qwen3.6-35B-A3B-MTP-GGUF. Note: `llama-server.service` and `llama-server@.service` are **masked** on .15 to prevent port 8080 collisions. |
|
||||
|
||||
## Prometheus & Grafana
|
||||
|
||||
@@ -593,7 +593,7 @@ curl -s -X POST http://192.168.68.116:9000/admin/roster/reload \
|
||||
-H "Authorization: Bearer sk-admin-ee09fffd04978b61a1569ac670c68814"
|
||||
|
||||
# Restart stuck GPU (saturation watchdog alternative)
|
||||
ssh root@192.168.68.8 "systemctl restart llama-chat-api.service" # .8 active unit since 2026-09-06; llama-server.service file on .8 is stale/inactive
|
||||
ssh root@192.168.68.8 "systemctl restart llama-server"
|
||||
ssh root@192.168.68.110 "systemctl restart llama-server"
|
||||
```
|
||||
|
||||
|
||||
Reference in New Issue
Block a user