fix(backups): document the thin-pool headroom and staging-directory preconditions, with the incidents that justify them #97

Merged
abiba-bot merged 3 commits from fix/backup-safety-preconditions-20260915 into master 2026-09-15 12:19:25 +00:00
Owner

Turns two incidents into a precondition the next operator or agent actually runs before a backup can take a host down.

Why this exists

  • acerpve, VM 101, twice on 2026-09-13. A snapshot-mode vzdump filled the LVM thin pool: dmsetup status pve-data-tpool showed thin-pool Error, then Fail; the host root remounted emergency_ro; ordinary commands returned I/O errors; LVM tools returned nothing; VM 101 (the RTX 3090 host) was unreachable while the host still answered ping and ssh. A second run in --mode stop produced the same failure. Both times a reboot rolled the failed transaction back and the pool returned rw (~30% data, ~1.2% metadata). Pool free space was never the obvious explanation (~816G with ~572G free), so metadata/snapshot pressure remains the leading hypothesis and is UNPROVEN - which is why the control is a preflight rather than a claimed root cause.
  • amdpve, 2026-09-14. A custom vzdump tmpdir created 0700 broke a whole night of container backups (fstat ... failed - EACCES) because the archive step runs in an unprivileged user namespace and could not traverse a root-owned directory. Fixed with chmod 1777 and proved with a real backup.

Changes (infrastructure-control.prose.md, +62)

  • PREFLIGHT preconditions before ANY snapshot-mode backup on a thin-pool host: read pool headroom and stop if data_percent >= 90% or metadata_percent >= 70%; verify the pool is not in Error/Fail state. States plainly that a full or errored thin pool fails every volume on the VG at once, including the host root.
  • Staging directory rule: any custom vzdump tmpdir MUST be mode 1777 like /var/tmp, with the exact EACCES symptom named so it is recognised in one line.
  • GPU-host coverage fact: VM 101 (llm-gpu) and VM 103 (ocu-llm) are in NO scheduled job, so their only coverage is one-off runs - deliberately so for VM 101 until the pool is understood.
  • The --output-format json rule when starting a task from a truncating shell (a | head -2 pipe killed a backup task with "broken pipe").

Verification performed by firstmate

Two commands in the first revision were WRONG and were caught by running them on the affected host before this PR:

  • dmsetup status ... | awk '{print $4, $6}' labelled $4 as a data percentage. Real output: ... thin-pool 99 26676/2183168 4005020/13372736 - rw ... - $4 is the transaction id, $5/$6 are used/total pairs. Following it would have recorded 99 as the data figure and blocked backups on a healthy pool.
  • lvs ... pve-data-tpool returns Volume group "pve-data-tpool" not found; the LVM object is pve/data. The corrected primary command was run on acerpve and prints data 29.95 1.22 <816.21g.
    Both are now correct in the branch, and the dmsetup field meaning is documented rather than implied.

Reviewers: run the two commands on a thin-pool host (acerpve) and confirm they produce the numbers the thresholds refer to; confirm the thresholds and the error-state check are unambiguous hard stops; and confirm no claim in the section states the thin-pool cause as proven.

Turns two incidents into a precondition the next operator or agent actually runs before a backup can take a host down. ## Why this exists - **acerpve, VM 101, twice on 2026-09-13.** A snapshot-mode vzdump filled the LVM thin pool: `dmsetup status pve-data-tpool` showed `thin-pool Error`, then `Fail`; the host root remounted `emergency_ro`; ordinary commands returned I/O errors; LVM tools returned nothing; VM 101 (the RTX 3090 host) was unreachable while the host still answered ping and ssh. A second run in `--mode stop` produced the same failure. Both times a reboot rolled the failed transaction back and the pool returned rw (~30% data, ~1.2% metadata). Pool free space was never the obvious explanation (~816G with ~572G free), so **metadata/snapshot pressure remains the leading hypothesis and is UNPROVEN** - which is why the control is a preflight rather than a claimed root cause. - **amdpve, 2026-09-14.** A custom vzdump `tmpdir` created 0700 broke a whole night of container backups (`fstat ... failed - EACCES`) because the archive step runs in an unprivileged user namespace and could not traverse a root-owned directory. Fixed with `chmod 1777` and proved with a real backup. ## Changes (`infrastructure-control.prose.md`, +62) - **PREFLIGHT preconditions** before ANY snapshot-mode backup on a thin-pool host: read pool headroom and stop if `data_percent >= 90%` or `metadata_percent >= 70%`; verify the pool is not in `Error`/`Fail` state. States plainly that a full or errored thin pool fails every volume on the VG at once, including the host root. - **Staging directory rule**: any custom vzdump `tmpdir` MUST be mode 1777 like `/var/tmp`, with the exact `EACCES` symptom named so it is recognised in one line. - **GPU-host coverage fact**: VM 101 (llm-gpu) and VM 103 (ocu-llm) are in NO scheduled job, so their only coverage is one-off runs - deliberately so for VM 101 until the pool is understood. - The `--output-format json` rule when starting a task from a truncating shell (a `| head -2` pipe killed a backup task with "broken pipe"). ## Verification performed by firstmate Two commands in the first revision were WRONG and were caught by running them on the affected host before this PR: - `dmsetup status ... | awk '{print $4, $6}'` labelled `$4` as a data percentage. Real output: `... thin-pool 99 26676/2183168 4005020/13372736 - rw ...` - `$4` is the **transaction id**, `$5`/`$6` are used/total pairs. Following it would have recorded 99 as the data figure and blocked backups on a healthy pool. - `lvs ... pve-data-tpool` returns `Volume group "pve-data-tpool" not found`; the LVM object is `pve/data`. The corrected primary command was run on acerpve and prints `data 29.95 1.22 <816.21g`. Both are now correct in the branch, and the dmsetup field meaning is documented rather than implied. Reviewers: run the two commands on a thin-pool host (acerpve) and confirm they produce the numbers the thresholds refer to; confirm the thresholds and the error-state check are unambiguous hard stops; and confirm no claim in the section states the thin-pool cause as proven.
abiba-bot added 2 commits 2026-09-15 11:45:39 +00:00
Add documented preflight checks for VM/CT backups on LVM thin-pool hosts:
- dmsetup status pve-data-tpool + lvs to verify data_percent < 90% and metadata_percent < 70%
- Exit 1 if pool shows Error/Fail state (takes down entire VG including host root)
- tmpdir must be mode 1777 (world-traversable) for vzdump archive step
- --output-format json for tasks started from truncating shells
- Document two incidents: acerpve thin-pool VM 101 (twice on 2026-09-13) and amdpve 0700 tmpdir (2026-09-14)
- Note metadata/snapshot-pressure hypothesis is UNPROVEN; preflight is the control
- Document GPU-host fact: VM 101 (llm-gpu) and VM 103 (ocu-llm) have no scheduled backup
fix: correct backup preflight commands - lvs pve/data and dmsetup field documentation
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
69940bc9eb
Fix two errors in the PREFLIGHT section (measured on acerpve 2026-09-15):
1. lvs -o ... pve/data (not pve-data-tpool) - this is the PRIMARY check that yields percentages directly
   - Quote the acerpve example: data 29.95% 1.22% <816.21g
2. dmsetup status pve-data-tpool - document fields correctly:
   -  = transaction ID (99), NOT data_percent
   -  = metadata used/total blocks
   -  = data used/total sectors
   - Show how to derive percentages if needed
3. Keep the error-state check (grep -q 'Error|Fail') - this is how the incident presented

Everything else stays: 1777 tmpdir requirement with EACCES symptom, incidents as rationale,
GPU-host fact, --output-format json rule, honest note that metadata/snapshot pressure is unproven.
abiba-bot added 1 commit 2026-09-15 12:10:07 +00:00
fix: add hostname resolution warning and fix dmsetup field documentation
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
b4b5321011
Add warning that bare PVE hostnames (acerpve, amdpve, etc.) resolve to VPS
via *.dns.sysloggh.net wildcard, not to actual nodes. List IP addresses:
- acerpve 192.168.68.9
- amdpve 192.168.68.15
- storepve 192.168.68.6
- minipve 192.168.68.12
- ocupve 192.168.68.5

Update acerpve example in backup preflight to include address (192.168.68.9).

Fix dmsetup comment to show full field order:
=start =length =thin-pool =transaction-id
=metadata_used/metadata_total =data_used/data_total
remaining fields are flags

Make it clear lvs command is the primary source for percentages, dmsetup is only for error-state check.

Incidents now include addresses: acerpve (192.168.68.9) and amdpve (192.168.68.15).
abiba-bot merged commit 40cf057370 into master 2026-09-15 12:19:25 +00:00
Sign in to join this conversation.
No Reviewers
No labels
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SyslogSolution/prose-contracts#97