revert: drop co-gated prose edits (gpu-fleet, infrastructure-control) — script-only scope per captain decision
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
The no-mistakes document step had auto-updated the .8 systemd unit name (llama-server -> llama-chat-api.service) in gpu-fleet.prose.md and infrastructure-control.prose.md. Those contracts are co-gated with another agent and cannot be amended inside a script-scope task (firstmate decision 2026-09-08, key prose-doc-step-scope); the unit-name doc sync is separate follow-up work. This restores both files to their origin/master content. The litellm-self-heal.prose.md agent-health-check v3/env.sh description update (same document commit) is intentional and kept.
This commit is contained in:
+1
-2
@@ -220,8 +220,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
|
||||
| `/root/scripts/gpu_benchmark.py` | pi (.24) | GPU inference benchmark module (tok/s tracking) |
|
||||
| `/root/scripts/gpu-saturation-watchdog.py` | pi (.24) | Auto-restart stuck llama-server |
|
||||
| `/root/dashboard/gpu-fleet.html` | pi (.24) | Live HTML dashboard |
|
||||
| `/etc/systemd/system/llama-chat-api.service` | .8 | RTX 3090 GPU daemon — active unit since 2026-09-06; the old `llama-server.service` file on .8 is stale/inactive |
|
||||
| `/etc/systemd/system/llama-server.service` | .110 | llama-server daemon (RTX 5070) |
|
||||
| `/etc/systemd/system/llama-server.service` | .8, .110 | llama-server daemons (Nvidia GPUs) |
|
||||
| `/etc/systemd/system/strix-server.service` | .15 (amdpve) | llama-server daemon (Vulkan, Strix Halo) running unsloth/Qwen3.6-35B-A3B-MTP-GGUF. Note: `llama-server.service` and `llama-server@.service` are **masked** on .15 to prevent port 8080 collisions. |
|
||||
|
||||
## Prometheus & Grafana
|
||||
|
||||
@@ -593,7 +593,7 @@ curl -s -X POST http://192.168.68.116:9000/admin/roster/reload \
|
||||
-H "Authorization: Bearer sk-admin-ee09fffd04978b61a1569ac670c68814"
|
||||
|
||||
# Restart stuck GPU (saturation watchdog alternative)
|
||||
ssh root@192.168.68.8 "systemctl restart llama-chat-api.service" # .8 active unit since 2026-09-06; llama-server.service file on .8 is stale/inactive
|
||||
ssh root@192.168.68.8 "systemctl restart llama-server"
|
||||
ssh root@192.168.68.110 "systemctl restart llama-server"
|
||||
```
|
||||
|
||||
|
||||
Reference in New Issue
Block a user