From 65eaffe1c6683b5f672adb073bff5d9cb1a0aed3 Mon Sep 17 00:00:00 2001 From: root Date: Tue, 15 Sep 2026 11:37:41 +0000 Subject: [PATCH 1/3] fix: add backup safety preconditions - thin-pool headroom and tmpdir 1777 Add documented preflight checks for VM/CT backups on LVM thin-pool hosts: - dmsetup status pve-data-tpool + lvs to verify data_percent < 90% and metadata_percent < 70% - Exit 1 if pool shows Error/Fail state (takes down entire VG including host root) - tmpdir must be mode 1777 (world-traversable) for vzdump archive step - --output-format json for tasks started from truncating shells - Document two incidents: acerpve thin-pool VM 101 (twice on 2026-09-13) and amdpve 0700 tmpdir (2026-09-14) - Note metadata/snapshot-pressure hypothesis is UNPROVEN; preflight is the control - Document GPU-host fact: VM 101 (llm-gpu) and VM 103 (ocu-llm) have no scheduled backup --- infrastructure-control.prose.md | 62 +++++++++++++++++++++++++++++++++ 1 file changed, 62 insertions(+) diff --git a/infrastructure-control.prose.md b/infrastructure-control.prose.md index 9734a1c..26c9a0f 100644 --- a/infrastructure-control.prose.md +++ b/infrastructure-control.prose.md @@ -364,6 +364,68 @@ For docker-vm specifically: - No PBS backup in 48h → fail ``` +### Backup Safety Preconditions (2026-09-15) + +#### Background & Rationale + +Two incidents from 2026-09-13/14 demonstrate that backup operations can catastrophically fail when storage conditions are not verified first: + +1. **acerpve thin-pool VM 101** (acerpve, 2026-09-13): A snapshot-mode vzdump of VM 101 on acerpve filled the LVM thin pool. `dmsetup status pve-data-tpool` showed `thin-pool Error` (then `Fail`), the host root remounted `emergency_ro`, ordinary commands failed with I/O errors, LVM tools returned nothing and VM 101 (the RTX 3090 host) went unreachable while the host still answered ping and ssh. It happened TWICE in one day with different modes: snapshot at 13:39Z and a `--mode stop` cold run at 18:52Z. Both times a reboot rolled the failed transaction back and the pool returned rw (~30% data, ~1.2% metadata). Pool capacity was NOT the obvious explanation - ~816G with ~572G free - which is why the metadata/snapshot-pressure hypothesis stands unproven. A full or errored thin pool fails EVERY volume on the VG at once, including the host root. + +2. **amdpve 0700 tmpdir** (amdpve, 2026-09-14): A custom vzdump `tmpdir` created with mode 0700 broke a whole night of container backups: `fstat "/vzdumptmp_//." failed - EACCES`, because the archive step runs through an unprivileged user namespace and could not traverse a root-owned 0700 directory. Fixed with `chmod 1777` (match /var/tmp) and proved with a real backup. + +3. **acerpve GPU-host fact**: VM 101 (llm-gpu) and VM 103 (ocu-llm) are in NO scheduled job, so their only coverage is one-off runs - and for VM 101 that is deliberate until the thin-pool is understood. + +#### PREFLIGHT Preconditions (Before ANY snapshot-mode backup on thin-pool hosts) + +Before starting ANY snapshot-mode vzdump on a host whose storage is an LVM thin pool, the following checks MUST pass: + +```bash +# Check 1: Pool headroom +dmsetup status pve-data-tpool | awk '{print $4, $6}' # data_percent metadata_percent +lvs -o lv_name,data_percent,metadata_percent,lv_size pve-data-tpool + +# Required thresholds (documented minimum): +# data_percent < 90% (80% recommended for safety margin) +# metadata_percent < 70% (metadata fills faster than data) + +# Check 2: Verify pool is not in error state +dmsetup status pve-data-tpool | grep -q "Error\|Fail" && exit 1 +``` + +**Minimum thresholds**: If either `data_percent >= 90%` or `metadata_percent >= 70%`, the backup MUST NOT start. State explicitly that these are hard stops, not warnings. + +**Why this is a precondition**: A full or errored thin pool fails EVERY volume on the VG at once, including the host root. This is not a soft failure - it takes down the entire Proxmox host. + +#### Staging Directory Requirement (2026-09-14 incident) + +Any custom vzdump `tmpdir` MUST be world-traversable and writable exactly like `/var/tmp` (mode 1777). The archive step of vzdump runs in an unprivileged user namespace and cannot traverse a root-owned 0700 directory. + +**Symptom to recognize**: `fstat "/vzdumptmp_//." failed - EACCES` on every container in the backup run. + +**Fix**: `chmod 1777 ` before starting vzdump. + +#### Task Start Rule for Truncating Shells + +When starting a backup task from a shell that may truncate output (e.g., pipes, `head`), always use: + +```bash +pvesh create /storage/backup --output-format json -- ... | head -2 +# ❌ Can kill the backup task ("broken pipe" status) +``` + +Instead, capture JSON output without piping to truncating commands: + +```bash +# Use --output-format json and capture to variable +result=$(pvesh create /storage/backup --output-format json -- ...) +# Then parse result if needed +``` + +#### GPU Host Backup Status (acerpve VM 101) + +VM 101 (llm-gpu) and VM 103 (ocu-llm) have NO scheduled backup job. Coverage is manual one-off runs only. This is intentional for VM 101 until the thin-pool failure mechanism is understood and documented. + ## Section 5: Network Services — Monitoring ### 5.1 Service Inventory -- 2.54.0 From 69940bc9ebefe8c1eb53699f9af022f209327561 Mon Sep 17 00:00:00 2001 From: root Date: Tue, 15 Sep 2026 11:42:10 +0000 Subject: [PATCH 2/3] fix: correct backup preflight commands - lvs pve/data and dmsetup field documentation Fix two errors in the PREFLIGHT section (measured on acerpve 2026-09-15): 1. lvs -o ... pve/data (not pve-data-tpool) - this is the PRIMARY check that yields percentages directly - Quote the acerpve example: data 29.95% 1.22% <816.21g 2. dmsetup status pve-data-tpool - document fields correctly: - = transaction ID (99), NOT data_percent - = metadata used/total blocks - = data used/total sectors - Show how to derive percentages if needed 3. Keep the error-state check (grep -q 'Error|Fail') - this is how the incident presented Everything else stays: 1777 tmpdir requirement with EACCES symptom, incidents as rationale, GPU-host fact, --output-format json rule, honest note that metadata/snapshot pressure is unproven. --- infrastructure-control.prose.md | 17 ++++++++++++----- 1 file changed, 12 insertions(+), 5 deletions(-) diff --git a/infrastructure-control.prose.md b/infrastructure-control.prose.md index 26c9a0f..852fc86 100644 --- a/infrastructure-control.prose.md +++ b/infrastructure-control.prose.md @@ -381,16 +381,23 @@ Two incidents from 2026-09-13/14 demonstrate that backup operations can catastro Before starting ANY snapshot-mode vzdump on a host whose storage is an LVM thin pool, the following checks MUST pass: ```bash -# Check 1: Pool headroom -dmsetup status pve-data-tpool | awk '{print $4, $6}' # data_percent metadata_percent -lvs -o lv_name,data_percent,metadata_percent,lv_size pve-data-tpool - +# Check 1: Pool headroom (PRIMARY - yields percentages directly) +lvs -o lv_name,data_percent,metadata_percent,lv_size pve/data +# Example output (acerpve): +# LV Data% Meta% LSize +# data 29.95 1.22 <816.21g # Required thresholds (documented minimum): # data_percent < 90% (80% recommended for safety margin) # metadata_percent < 70% (metadata fills faster than data) -# Check 2: Verify pool is not in error state +# Check 2: Verify pool is not in error state (dmsetup shows the raw DM device) dmsetup status pve-data-tpool | grep -q "Error\|Fail" && exit 1 +# dmsetup output fields: / / +# $4 is transaction ID (99), NOT a percentage +# Example: 0 99 26676/2183168 4005020/13372736 +# Use $5 and $6 to calculate percentages if needed: +# data_percent = $6 / ($6 split by /) [second number in pair] +# metadata_percent = $5 / ($5 split by /) [second number in pair] ``` **Minimum thresholds**: If either `data_percent >= 90%` or `metadata_percent >= 70%`, the backup MUST NOT start. State explicitly that these are hard stops, not warnings. -- 2.54.0 From b4b5321011be304449015bb54b7cb6ab975c6969 Mon Sep 17 00:00:00 2001 From: root Date: Tue, 15 Sep 2026 12:09:59 +0000 Subject: [PATCH 3/3] fix: add hostname resolution warning and fix dmsetup field documentation Add warning that bare PVE hostnames (acerpve, amdpve, etc.) resolve to VPS via *.dns.sysloggh.net wildcard, not to actual nodes. List IP addresses: - acerpve 192.168.68.9 - amdpve 192.168.68.15 - storepve 192.168.68.6 - minipve 192.168.68.12 - ocupve 192.168.68.5 Update acerpve example in backup preflight to include address (192.168.68.9). Fix dmsetup comment to show full field order: =start =length =thin-pool =transaction-id =metadata_used/metadata_total =data_used/data_total remaining fields are flags Make it clear lvs command is the primary source for percentages, dmsetup is only for error-state check. Incidents now include addresses: acerpve (192.168.68.9) and amdpve (192.168.68.15). --- infrastructure-control.prose.md | 20 ++++++++++++-------- 1 file changed, 12 insertions(+), 8 deletions(-) diff --git a/infrastructure-control.prose.md b/infrastructure-control.prose.md index 852fc86..87f64b7 100644 --- a/infrastructure-control.prose.md +++ b/infrastructure-control.prose.md @@ -370,20 +370,23 @@ For docker-vm specifically: Two incidents from 2026-09-13/14 demonstrate that backup operations can catastrophically fail when storage conditions are not verified first: -1. **acerpve thin-pool VM 101** (acerpve, 2026-09-13): A snapshot-mode vzdump of VM 101 on acerpve filled the LVM thin pool. `dmsetup status pve-data-tpool` showed `thin-pool Error` (then `Fail`), the host root remounted `emergency_ro`, ordinary commands failed with I/O errors, LVM tools returned nothing and VM 101 (the RTX 3090 host) went unreachable while the host still answered ping and ssh. It happened TWICE in one day with different modes: snapshot at 13:39Z and a `--mode stop` cold run at 18:52Z. Both times a reboot rolled the failed transaction back and the pool returned rw (~30% data, ~1.2% metadata). Pool capacity was NOT the obvious explanation - ~816G with ~572G free - which is why the metadata/snapshot-pressure hypothesis stands unproven. A full or errored thin pool fails EVERY volume on the VG at once, including the host root. +1. **acerpve thin-pool VM 101** (acerpve, 192.168.68.9, 2026-09-13): A snapshot-mode vzdump of VM 101 on acerpve filled the LVM thin pool. `dmsetup status pve-data-tpool` showed `thin-pool Error` (then `Fail`), the host root remounted `emergency_ro`, ordinary commands failed with I/O errors, LVM tools returned nothing and VM 101 (the RTX 3090 host) went unreachable while the host still answered ping and ssh. It happened TWICE in one day with different modes: snapshot at 13:39Z and a `--mode stop` cold run at 18:52Z. Both times a reboot rolled the failed transaction back and the pool returned rw (~30% data, ~1.2% metadata). Pool capacity was NOT the obvious explanation - ~816G with ~572G free - which is why the metadata/snapshot-pressure hypothesis stands unproven. A full or errored thin pool fails EVERY volume on the VG at once, including the host root. -2. **amdpve 0700 tmpdir** (amdpve, 2026-09-14): A custom vzdump `tmpdir` created with mode 0700 broke a whole night of container backups: `fstat "/vzdumptmp_//." failed - EACCES`, because the archive step runs through an unprivileged user namespace and could not traverse a root-owned 0700 directory. Fixed with `chmod 1777` (match /var/tmp) and proved with a real backup. +2. **amdpve 0700 tmpdir** (amdpve, 192.168.68.15, 2026-09-14): A custom vzdump `tmpdir` created with mode 0700 broke a whole night of container backups: `fstat "/vzdumptmp_//." failed - EACCES`, because the archive step runs through an unprivileged user namespace and could not traverse a root-owned 0700 directory. Fixed with `chmod 1777` (match /var/tmp) and proved with a real backup. 3. **acerpve GPU-host fact**: VM 101 (llm-gpu) and VM 103 (ocu-llm) are in NO scheduled job, so their only coverage is one-off runs - and for VM 101 that is deliberate until the thin-pool is understood. +> ⚠️ **Hostname Resolution Warning (2026-09-15)**: The PVE node hostnames (acerpve, amdpve, minipve, storepve, ocupve) all resolve to the VPS (72.61.0.17, the Netbird VPS at srv1079750.hstgr.cloud) via the wildcard `*.dns.sysloggh.net` record, NOT to the actual nodes. So `ssh acerpve` lands on the VPS. **Nodes must be addressed by IP**: acerpve 192.168.68.9, amdpve 192.168.68.15, storepve 192.168.68.6, minipve 192.168.68.12, ocupve 192.168.68.5. Guest CTs are reached through their node (`pct exec`). Guest hostnames that resolve on the LAN (e.g. kagentz = 192.168.68.14) are fine. (The DNS address records are a separate decision — row: dag-daemon-node-hostnames-resolve-to-the-vps-20260915.) + #### PREFLIGHT Preconditions (Before ANY snapshot-mode backup on thin-pool hosts) Before starting ANY snapshot-mode vzdump on a host whose storage is an LVM thin pool, the following checks MUST pass: ```bash # Check 1: Pool headroom (PRIMARY - yields percentages directly) +# Run on the NODE (e.g. ssh root@192.168.68.9 for acerpve) — NOT by bare hostname, see warning above lvs -o lv_name,data_percent,metadata_percent,lv_size pve/data -# Example output (acerpve): +# Example output (acerpve, 192.168.68.9): # LV Data% Meta% LSize # data 29.95 1.22 <816.21g # Required thresholds (documented minimum): @@ -391,12 +394,13 @@ lvs -o lv_name,data_percent,metadata_percent,lv_size pve/data # metadata_percent < 70% (metadata fills faster than data) # Check 2: Verify pool is not in error state (dmsetup shows the raw DM device) +# Run on the NODE (e.g. ssh root@192.168.68.9 for acerpve) — NOT by bare hostname dmsetup status pve-data-tpool | grep -q "Error\|Fail" && exit 1 -# dmsetup output fields: / / -# $4 is transaction ID (99), NOT a percentage -# Example: 0 99 26676/2183168 4005020/13372736 -# Use $5 and $6 to calculate percentages if needed: -# data_percent = $6 / ($6 split by /) [second number in pair] +# dmsetup status pve-data-tpool field order (verified on 192.168.68.9): +# $1=start $2=length $3="thin-pool" $4=transaction-id +# $5=metadata_used/metadata_total (blocks) $6=data_used/data_total (sectors) +# remaining fields are flags ("-", "rw", "discard_passdown", "queue_if_no_space", ...) +# This is only used for the ERROR-STATE check; use the lvs command above for percentages. # metadata_percent = $5 / ($5 split by /) [second number in pair] ``` -- 2.54.0