fix(backups): document the thin-pool headroom and staging-directory preconditions, with the incidents that justify them #97
@@ -370,20 +370,23 @@ For docker-vm specifically:
|
|||||||
|
|
||||||
Two incidents from 2026-09-13/14 demonstrate that backup operations can catastrophically fail when storage conditions are not verified first:
|
Two incidents from 2026-09-13/14 demonstrate that backup operations can catastrophically fail when storage conditions are not verified first:
|
||||||
|
|
||||||
1. **acerpve thin-pool VM 101** (acerpve, 2026-09-13): A snapshot-mode vzdump of VM 101 on acerpve filled the LVM thin pool. `dmsetup status pve-data-tpool` showed `thin-pool Error` (then `Fail`), the host root remounted `emergency_ro`, ordinary commands failed with I/O errors, LVM tools returned nothing and VM 101 (the RTX 3090 host) went unreachable while the host still answered ping and ssh. It happened TWICE in one day with different modes: snapshot at 13:39Z and a `--mode stop` cold run at 18:52Z. Both times a reboot rolled the failed transaction back and the pool returned rw (~30% data, ~1.2% metadata). Pool capacity was NOT the obvious explanation - ~816G with ~572G free - which is why the metadata/snapshot-pressure hypothesis stands unproven. A full or errored thin pool fails EVERY volume on the VG at once, including the host root.
|
1. **acerpve thin-pool VM 101** (acerpve, 192.168.68.9, 2026-09-13): A snapshot-mode vzdump of VM 101 on acerpve filled the LVM thin pool. `dmsetup status pve-data-tpool` showed `thin-pool Error` (then `Fail`), the host root remounted `emergency_ro`, ordinary commands failed with I/O errors, LVM tools returned nothing and VM 101 (the RTX 3090 host) went unreachable while the host still answered ping and ssh. It happened TWICE in one day with different modes: snapshot at 13:39Z and a `--mode stop` cold run at 18:52Z. Both times a reboot rolled the failed transaction back and the pool returned rw (~30% data, ~1.2% metadata). Pool capacity was NOT the obvious explanation - ~816G with ~572G free - which is why the metadata/snapshot-pressure hypothesis stands unproven. A full or errored thin pool fails EVERY volume on the VG at once, including the host root.
|
||||||
|
|
||||||
2. **amdpve 0700 tmpdir** (amdpve, 2026-09-14): A custom vzdump `tmpdir` created with mode 0700 broke a whole night of container backups: `fstat "<dir>/vzdumptmp<n>_<ct>//." failed - EACCES`, because the archive step runs through an unprivileged user namespace and could not traverse a root-owned 0700 directory. Fixed with `chmod 1777` (match /var/tmp) and proved with a real backup.
|
2. **amdpve 0700 tmpdir** (amdpve, 192.168.68.15, 2026-09-14): A custom vzdump `tmpdir` created with mode 0700 broke a whole night of container backups: `fstat "<dir>/vzdumptmp<n>_<ct>//." failed - EACCES`, because the archive step runs through an unprivileged user namespace and could not traverse a root-owned 0700 directory. Fixed with `chmod 1777` (match /var/tmp) and proved with a real backup.
|
||||||
|
|
||||||
3. **acerpve GPU-host fact**: VM 101 (llm-gpu) and VM 103 (ocu-llm) are in NO scheduled job, so their only coverage is one-off runs - and for VM 101 that is deliberate until the thin-pool is understood.
|
3. **acerpve GPU-host fact**: VM 101 (llm-gpu) and VM 103 (ocu-llm) are in NO scheduled job, so their only coverage is one-off runs - and for VM 101 that is deliberate until the thin-pool is understood.
|
||||||
|
|
||||||
|
> ⚠️ **Hostname Resolution Warning (2026-09-15)**: The PVE node hostnames (acerpve, amdpve, minipve, storepve, ocupve) all resolve to the VPS (72.61.0.17, the Netbird VPS at srv1079750.hstgr.cloud) via the wildcard `*.dns.sysloggh.net` record, NOT to the actual nodes. So `ssh acerpve` lands on the VPS. **Nodes must be addressed by IP**: acerpve 192.168.68.9, amdpve 192.168.68.15, storepve 192.168.68.6, minipve 192.168.68.12, ocupve 192.168.68.5. Guest CTs are reached through their node (`pct exec`). Guest hostnames that resolve on the LAN (e.g. kagentz = 192.168.68.14) are fine. (The DNS address records are a separate decision — row: dag-daemon-node-hostnames-resolve-to-the-vps-20260915.)
|
||||||
|
|
||||||
#### PREFLIGHT Preconditions (Before ANY snapshot-mode backup on thin-pool hosts)
|
#### PREFLIGHT Preconditions (Before ANY snapshot-mode backup on thin-pool hosts)
|
||||||
|
|
||||||
Before starting ANY snapshot-mode vzdump on a host whose storage is an LVM thin pool, the following checks MUST pass:
|
Before starting ANY snapshot-mode vzdump on a host whose storage is an LVM thin pool, the following checks MUST pass:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
# Check 1: Pool headroom (PRIMARY - yields percentages directly)
|
# Check 1: Pool headroom (PRIMARY - yields percentages directly)
|
||||||
|
# Run on the NODE (e.g. ssh root@192.168.68.9 for acerpve) — NOT by bare hostname, see warning above
|
||||||
lvs -o lv_name,data_percent,metadata_percent,lv_size pve/data
|
lvs -o lv_name,data_percent,metadata_percent,lv_size pve/data
|
||||||
# Example output (acerpve):
|
# Example output (acerpve, 192.168.68.9):
|
||||||
# LV Data% Meta% LSize
|
# LV Data% Meta% LSize
|
||||||
# data 29.95 1.22 <816.21g
|
# data 29.95 1.22 <816.21g
|
||||||
# Required thresholds (documented minimum):
|
# Required thresholds (documented minimum):
|
||||||
@@ -391,12 +394,13 @@ lvs -o lv_name,data_percent,metadata_percent,lv_size pve/data
|
|||||||
# metadata_percent < 70% (metadata fills faster than data)
|
# metadata_percent < 70% (metadata fills faster than data)
|
||||||
|
|
||||||
# Check 2: Verify pool is not in error state (dmsetup shows the raw DM device)
|
# Check 2: Verify pool is not in error state (dmsetup shows the raw DM device)
|
||||||
|
# Run on the NODE (e.g. ssh root@192.168.68.9 for acerpve) — NOT by bare hostname
|
||||||
dmsetup status pve-data-tpool | grep -q "Error\|Fail" && exit 1
|
dmsetup status pve-data-tpool | grep -q "Error\|Fail" && exit 1
|
||||||
# dmsetup output fields: <transaction-id> <metadata_used>/<metadata_total> <data_used>/<data_total>
|
# dmsetup status pve-data-tpool field order (verified on 192.168.68.9):
|
||||||
# $4 is transaction ID (99), NOT a percentage
|
# $1=start $2=length $3="thin-pool" $4=transaction-id
|
||||||
# Example: 0 99 26676/2183168 4005020/13372736
|
# $5=metadata_used/metadata_total (blocks) $6=data_used/data_total (sectors)
|
||||||
# Use $5 and $6 to calculate percentages if needed:
|
# remaining fields are flags ("-", "rw", "discard_passdown", "queue_if_no_space", ...)
|
||||||
# data_percent = $6 / ($6 split by /) [second number in pair]
|
# This is only used for the ERROR-STATE check; use the lvs command above for percentages.
|
||||||
# metadata_percent = $5 / ($5 split by /) [second number in pair]
|
# metadata_percent = $5 / ($5 split by /) [second number in pair]
|
||||||
```
|
```
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user