fix: add backup safety preconditions - thin-pool headroom and tmpdir 1777
Add documented preflight checks for VM/CT backups on LVM thin-pool hosts: - dmsetup status pve-data-tpool + lvs to verify data_percent < 90% and metadata_percent < 70% - Exit 1 if pool shows Error/Fail state (takes down entire VG including host root) - tmpdir must be mode 1777 (world-traversable) for vzdump archive step - --output-format json for tasks started from truncating shells - Document two incidents: acerpve thin-pool VM 101 (twice on 2026-09-13) and amdpve 0700 tmpdir (2026-09-14) - Note metadata/snapshot-pressure hypothesis is UNPROVEN; preflight is the control - Document GPU-host fact: VM 101 (llm-gpu) and VM 103 (ocu-llm) have no scheduled backup
This commit is contained in:
@@ -364,6 +364,68 @@ For docker-vm specifically:
|
||||
- No PBS backup in 48h → fail
|
||||
```
|
||||
|
||||
### Backup Safety Preconditions (2026-09-15)
|
||||
|
||||
#### Background & Rationale
|
||||
|
||||
Two incidents from 2026-09-13/14 demonstrate that backup operations can catastrophically fail when storage conditions are not verified first:
|
||||
|
||||
1. **acerpve thin-pool VM 101** (acerpve, 2026-09-13): A snapshot-mode vzdump of VM 101 on acerpve filled the LVM thin pool. `dmsetup status pve-data-tpool` showed `thin-pool Error` (then `Fail`), the host root remounted `emergency_ro`, ordinary commands failed with I/O errors, LVM tools returned nothing and VM 101 (the RTX 3090 host) went unreachable while the host still answered ping and ssh. It happened TWICE in one day with different modes: snapshot at 13:39Z and a `--mode stop` cold run at 18:52Z. Both times a reboot rolled the failed transaction back and the pool returned rw (~30% data, ~1.2% metadata). Pool capacity was NOT the obvious explanation - ~816G with ~572G free - which is why the metadata/snapshot-pressure hypothesis stands unproven. A full or errored thin pool fails EVERY volume on the VG at once, including the host root.
|
||||
|
||||
2. **amdpve 0700 tmpdir** (amdpve, 2026-09-14): A custom vzdump `tmpdir` created with mode 0700 broke a whole night of container backups: `fstat "<dir>/vzdumptmp<n>_<ct>//." failed - EACCES`, because the archive step runs through an unprivileged user namespace and could not traverse a root-owned 0700 directory. Fixed with `chmod 1777` (match /var/tmp) and proved with a real backup.
|
||||
|
||||
3. **acerpve GPU-host fact**: VM 101 (llm-gpu) and VM 103 (ocu-llm) are in NO scheduled job, so their only coverage is one-off runs - and for VM 101 that is deliberate until the thin-pool is understood.
|
||||
|
||||
#### PREFLIGHT Preconditions (Before ANY snapshot-mode backup on thin-pool hosts)
|
||||
|
||||
Before starting ANY snapshot-mode vzdump on a host whose storage is an LVM thin pool, the following checks MUST pass:
|
||||
|
||||
```bash
|
||||
# Check 1: Pool headroom
|
||||
dmsetup status pve-data-tpool | awk '{print $4, $6}' # data_percent metadata_percent
|
||||
lvs -o lv_name,data_percent,metadata_percent,lv_size pve-data-tpool
|
||||
|
||||
# Required thresholds (documented minimum):
|
||||
# data_percent < 90% (80% recommended for safety margin)
|
||||
# metadata_percent < 70% (metadata fills faster than data)
|
||||
|
||||
# Check 2: Verify pool is not in error state
|
||||
dmsetup status pve-data-tpool | grep -q "Error\|Fail" && exit 1
|
||||
```
|
||||
|
||||
**Minimum thresholds**: If either `data_percent >= 90%` or `metadata_percent >= 70%`, the backup MUST NOT start. State explicitly that these are hard stops, not warnings.
|
||||
|
||||
**Why this is a precondition**: A full or errored thin pool fails EVERY volume on the VG at once, including the host root. This is not a soft failure - it takes down the entire Proxmox host.
|
||||
|
||||
#### Staging Directory Requirement (2026-09-14 incident)
|
||||
|
||||
Any custom vzdump `tmpdir` MUST be world-traversable and writable exactly like `/var/tmp` (mode 1777). The archive step of vzdump runs in an unprivileged user namespace and cannot traverse a root-owned 0700 directory.
|
||||
|
||||
**Symptom to recognize**: `fstat "<dir>/vzdumptmp<n>_<ct>//." failed - EACCES` on every container in the backup run.
|
||||
|
||||
**Fix**: `chmod 1777 <custom-tmpdir>` before starting vzdump.
|
||||
|
||||
#### Task Start Rule for Truncating Shells
|
||||
|
||||
When starting a backup task from a shell that may truncate output (e.g., pipes, `head`), always use:
|
||||
|
||||
```bash
|
||||
pvesh create /storage/backup --output-format json -- ... | head -2
|
||||
# ❌ Can kill the backup task ("broken pipe" status)
|
||||
```
|
||||
|
||||
Instead, capture JSON output without piping to truncating commands:
|
||||
|
||||
```bash
|
||||
# Use --output-format json and capture to variable
|
||||
result=$(pvesh create /storage/backup --output-format json -- ...)
|
||||
# Then parse result if needed
|
||||
```
|
||||
|
||||
#### GPU Host Backup Status (acerpve VM 101)
|
||||
|
||||
VM 101 (llm-gpu) and VM 103 (ocu-llm) have NO scheduled backup job. Coverage is manual one-off runs only. This is intentional for VM 101 until the thin-pool failure mechanism is understood and documented.
|
||||
|
||||
## Section 5: Network Services — Monitoring
|
||||
|
||||
### 5.1 Service Inventory
|
||||
|
||||
Reference in New Issue
Block a user