Turns two incidents into a precondition the next operator or agent actually runs before a backup can take a host down.
Why this exists
acerpve, VM 101, twice on 2026-09-13. A snapshot-mode vzdump filled the LVM thin pool: dmsetup status pve-data-tpool showed thin-pool Error, then Fail; the host root remounted emergency_ro; ordinary commands returned I/O errors; LVM tools returned nothing; VM 101 (the RTX 3090 host) was unreachable while the host still answered ping and ssh. A second run in --mode stop produced the same failure. Both times a reboot rolled the failed transaction back and the pool returned rw (~30% data, ~1.2% metadata). Pool free space was never the obvious explanation (~816G with ~572G free), so metadata/snapshot pressure remains the leading hypothesis and is UNPROVEN - which is why the control is a preflight rather than a claimed root cause.
amdpve, 2026-09-14. A custom vzdump tmpdir created 0700 broke a whole night of container backups (fstat ... failed - EACCES) because the archive step runs in an unprivileged user namespace and could not traverse a root-owned directory. Fixed with chmod 1777 and proved with a real backup.
Changes (infrastructure-control.prose.md, +62)
PREFLIGHT preconditions before ANY snapshot-mode backup on a thin-pool host: read pool headroom and stop if data_percent >= 90% or metadata_percent >= 70%; verify the pool is not in Error/Fail state. States plainly that a full or errored thin pool fails every volume on the VG at once, including the host root.
Staging directory rule: any custom vzdump tmpdir MUST be mode 1777 like /var/tmp, with the exact EACCES symptom named so it is recognised in one line.
GPU-host coverage fact: VM 101 (llm-gpu) and VM 103 (ocu-llm) are in NO scheduled job, so their only coverage is one-off runs - deliberately so for VM 101 until the pool is understood.
The --output-format json rule when starting a task from a truncating shell (a | head -2 pipe killed a backup task with "broken pipe").
Verification performed by firstmate
Two commands in the first revision were WRONG and were caught by running them on the affected host before this PR:
dmsetup status ... | awk '{print $4, $6}' labelled $4 as a data percentage. Real output: ... thin-pool 99 26676/2183168 4005020/13372736 - rw ... - $4 is the transaction id, $5/$6 are used/total pairs. Following it would have recorded 99 as the data figure and blocked backups on a healthy pool.
lvs ... pve-data-tpool returns Volume group "pve-data-tpool" not found; the LVM object is pve/data. The corrected primary command was run on acerpve and prints data 29.95 1.22 <816.21g.
Both are now correct in the branch, and the dmsetup field meaning is documented rather than implied.
Reviewers: run the two commands on a thin-pool host (acerpve) and confirm they produce the numbers the thresholds refer to; confirm the thresholds and the error-state check are unambiguous hard stops; and confirm no claim in the section states the thin-pool cause as proven.
Turns two incidents into a precondition the next operator or agent actually runs before a backup can take a host down.
## Why this exists
- **acerpve, VM 101, twice on 2026-09-13.** A snapshot-mode vzdump filled the LVM thin pool: `dmsetup status pve-data-tpool` showed `thin-pool Error`, then `Fail`; the host root remounted `emergency_ro`; ordinary commands returned I/O errors; LVM tools returned nothing; VM 101 (the RTX 3090 host) was unreachable while the host still answered ping and ssh. A second run in `--mode stop` produced the same failure. Both times a reboot rolled the failed transaction back and the pool returned rw (~30% data, ~1.2% metadata). Pool free space was never the obvious explanation (~816G with ~572G free), so **metadata/snapshot pressure remains the leading hypothesis and is UNPROVEN** - which is why the control is a preflight rather than a claimed root cause.
- **amdpve, 2026-09-14.** A custom vzdump `tmpdir` created 0700 broke a whole night of container backups (`fstat ... failed - EACCES`) because the archive step runs in an unprivileged user namespace and could not traverse a root-owned directory. Fixed with `chmod 1777` and proved with a real backup.
## Changes (`infrastructure-control.prose.md`, +62)
- **PREFLIGHT preconditions** before ANY snapshot-mode backup on a thin-pool host: read pool headroom and stop if `data_percent >= 90%` or `metadata_percent >= 70%`; verify the pool is not in `Error`/`Fail` state. States plainly that a full or errored thin pool fails every volume on the VG at once, including the host root.
- **Staging directory rule**: any custom vzdump `tmpdir` MUST be mode 1777 like `/var/tmp`, with the exact `EACCES` symptom named so it is recognised in one line.
- **GPU-host coverage fact**: VM 101 (llm-gpu) and VM 103 (ocu-llm) are in NO scheduled job, so their only coverage is one-off runs - deliberately so for VM 101 until the pool is understood.
- The `--output-format json` rule when starting a task from a truncating shell (a `| head -2` pipe killed a backup task with "broken pipe").
## Verification performed by firstmate
Two commands in the first revision were WRONG and were caught by running them on the affected host before this PR:
- `dmsetup status ... | awk '{print $4, $6}'` labelled `$4` as a data percentage. Real output: `... thin-pool 99 26676/2183168 4005020/13372736 - rw ...` - `$4` is the **transaction id**, `$5`/`$6` are used/total pairs. Following it would have recorded 99 as the data figure and blocked backups on a healthy pool.
- `lvs ... pve-data-tpool` returns `Volume group "pve-data-tpool" not found`; the LVM object is `pve/data`. The corrected primary command was run on acerpve and prints `data 29.95 1.22 <816.21g`.
Both are now correct in the branch, and the dmsetup field meaning is documented rather than implied.
Reviewers: run the two commands on a thin-pool host (acerpve) and confirm they produce the numbers the thresholds refer to; confirm the thresholds and the error-state check are unambiguous hard stops; and confirm no claim in the section states the thin-pool cause as proven.
Add documented preflight checks for VM/CT backups on LVM thin-pool hosts:
- dmsetup status pve-data-tpool + lvs to verify data_percent < 90% and metadata_percent < 70%
- Exit 1 if pool shows Error/Fail state (takes down entire VG including host root)
- tmpdir must be mode 1777 (world-traversable) for vzdump archive step
- --output-format json for tasks started from truncating shells
- Document two incidents: acerpve thin-pool VM 101 (twice on 2026-09-13) and amdpve 0700 tmpdir (2026-09-14)
- Note metadata/snapshot-pressure hypothesis is UNPROVEN; preflight is the control
- Document GPU-host fact: VM 101 (llm-gpu) and VM 103 (ocu-llm) have no scheduled backup
Fix two errors in the PREFLIGHT section (measured on acerpve 2026-09-15):
1. lvs -o ... pve/data (not pve-data-tpool) - this is the PRIMARY check that yields percentages directly
- Quote the acerpve example: data 29.95% 1.22% <816.21g
2. dmsetup status pve-data-tpool - document fields correctly:
- = transaction ID (99), NOT data_percent
- = metadata used/total blocks
- = data used/total sectors
- Show how to derive percentages if needed
3. Keep the error-state check (grep -q 'Error|Fail') - this is how the incident presented
Everything else stays: 1777 tmpdir requirement with EACCES symptom, incidents as rationale,
GPU-host fact, --output-format json rule, honest note that metadata/snapshot pressure is unproven.
Add warning that bare PVE hostnames (acerpve, amdpve, etc.) resolve to VPS
via *.dns.sysloggh.net wildcard, not to actual nodes. List IP addresses:
- acerpve 192.168.68.9
- amdpve 192.168.68.15
- storepve 192.168.68.6
- minipve 192.168.68.12
- ocupve 192.168.68.5
Update acerpve example in backup preflight to include address (192.168.68.9).
Fix dmsetup comment to show full field order:
=start =length =thin-pool =transaction-id
=metadata_used/metadata_total =data_used/data_total
remaining fields are flags
Make it clear lvs command is the primary source for percentages, dmsetup is only for error-state check.
Incidents now include addresses: acerpve (192.168.68.9) and amdpve (192.168.68.15).
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Turns two incidents into a precondition the next operator or agent actually runs before a backup can take a host down.
Why this exists
dmsetup status pve-data-tpoolshowedthin-pool Error, thenFail; the host root remountedemergency_ro; ordinary commands returned I/O errors; LVM tools returned nothing; VM 101 (the RTX 3090 host) was unreachable while the host still answered ping and ssh. A second run in--mode stopproduced the same failure. Both times a reboot rolled the failed transaction back and the pool returned rw (~30% data, ~1.2% metadata). Pool free space was never the obvious explanation (~816G with ~572G free), so metadata/snapshot pressure remains the leading hypothesis and is UNPROVEN - which is why the control is a preflight rather than a claimed root cause.tmpdircreated 0700 broke a whole night of container backups (fstat ... failed - EACCES) because the archive step runs in an unprivileged user namespace and could not traverse a root-owned directory. Fixed withchmod 1777and proved with a real backup.Changes (
infrastructure-control.prose.md, +62)data_percent >= 90%ormetadata_percent >= 70%; verify the pool is not inError/Failstate. States plainly that a full or errored thin pool fails every volume on the VG at once, including the host root.tmpdirMUST be mode 1777 like/var/tmp, with the exactEACCESsymptom named so it is recognised in one line.--output-format jsonrule when starting a task from a truncating shell (a| head -2pipe killed a backup task with "broken pipe").Verification performed by firstmate
Two commands in the first revision were WRONG and were caught by running them on the affected host before this PR:
dmsetup status ... | awk '{print $4, $6}'labelled$4as a data percentage. Real output:... thin-pool 99 26676/2183168 4005020/13372736 - rw ...-$4is the transaction id,$5/$6are used/total pairs. Following it would have recorded 99 as the data figure and blocked backups on a healthy pool.lvs ... pve-data-tpoolreturnsVolume group "pve-data-tpool" not found; the LVM object ispve/data. The corrected primary command was run on acerpve and printsdata 29.95 1.22 <816.21g.Both are now correct in the branch, and the dmsetup field meaning is documented rather than implied.
Reviewers: run the two commands on a thin-pool host (acerpve) and confirm they produce the numbers the thresholds refer to; confirm the thresholds and the error-state check are unambiguous hard stops; and confirm no claim in the section states the thin-pool cause as proven.