Commit Graph

199 Commits

Author SHA1 Message Date
MacRimi
9c90722ef9 update 1.2.2.3 beta 2026-07-06 00:59:50 +02:00
MacRimi
7156af1965 update 1.2.2.3 beta 2026-07-06 00:39:43 +02:00
MacRimi
8df8a77bf1 update 1.2.2.3 beta 2026-07-06 00:30:47 +02:00
MacRimi
487ab04a14 update 1.2.2.3 beta 2026-07-05 23:58:58 +02:00
MacRimi
e35ef38fea update 1.2.2.3 beta 2026-07-05 23:22:39 +02:00
MacRimi
5b22039600 Update 1.2.2.3 beta 2026-07-05 17:23:00 +02:00
MacRimi
8bcfcd6059 update 1.2.2.3 beta 2026-07-05 16:58:27 +02:00
MacRimi
7fc7125c71 create 1.2.2.3 beta 2026-07-05 16:50:06 +02:00
MacRimi
29bca610a0 update 1.2.2.2 beta 2026-07-05 09:30:13 +02:00
MacRimi
0d6c7290e5 Update 1.2.2.2 beta 2026-07-04 22:03:45 +02:00
MacRimi
f768d9eff8 Update 1.2.2.2 beta 2026-07-04 21:50:29 +02:00
MacRimi
35cb10ff44 update 1.2.2.2 beta 2026-07-03 18:59:45 +02:00
MacRimi
c455d66b91 update 1.2.2.2 beta 2026-07-02 23:52:04 +02:00
MacRimi
827c88d24f update 1.2.2.2 beta 2026-07-02 20:57:11 +02:00
github-actions[bot]
872f79a9ea Update AppImage beta build (2026-07-02 18:18:20) 2026-07-02 18:18:20 +00:00
github-actions[bot]
efa84b0fa1 Update AppImage beta build (2026-07-02 16:24:06) 2026-07-02 16:24:06 +00:00
github-actions[bot]
f0e79e93b6 Update AppImage beta build (2026-07-01 18:57:10) 2026-07-01 18:57:10 +00:00
github-actions[bot]
8235549e4f Update AppImage beta build (2026-07-01 18:19:00) 2026-07-01 18:19:00 +00:00
github-actions[bot]
d522b9a337 Update AppImage beta build (2026-06-30 16:05:14) 2026-06-30 16:05:14 +00:00
MacRimi
b2753be204 update 1.2.2.2 beta 2026-06-28 15:17:30 +02:00
MacRimi
51b9285980 update 1.2.2.2 beta 2026-06-28 13:06:40 +02:00
MacRimi
ecfdcf1bac update 1.2.2.2 beta 2026-06-26 11:11:39 +02:00
MacRimi
63b9d69f3f update 1.2.2.2 beta 2026-06-26 10:25:06 +02:00
MacRimi
a56afecccf update 1.2.2.2 beta 2026-06-25 17:29:55 +02:00
MacRimi
484f0ce897 update 1.2.2.2 beta 2026-06-25 17:07:58 +02:00
MacRimi
cdb4522e5f update 1.2.2.2 beta 2026-06-25 00:26:03 +02:00
MacRimi
4f2494e135 Update 1.2.2.2 beta 2026-06-25 00:00:12 +02:00
MacRimi
202068124b update 1.2.2.2 beta 2026-06-24 23:40:22 +02:00
MacRimi
87a29f324b Update 1.2.2.2 beta 2026-06-24 21:58:21 +02:00
MacRimi
61b9fd12bb update 1.2.2.2 beta 2026-06-24 18:23:16 +02:00
MacRimi
cd2a075fab update 1.2.2.2 beta 2026-06-24 16:02:35 +02:00
MacRimi
93553574b3 Update 1.2.2.2 beta 2026-06-23 11:19:04 +02:00
MacRimi
6ab9d4ca27 Update 1.2.2.2 beta 2026-06-22 18:52:26 +02:00
MacRimi
68c8c03642 update 1.2.2.2 pre-beta 2026-06-22 11:38:06 +02:00
MacRimi
c4cab77319 update 1.2.2.2 beta 2026-06-22 01:16:49 +02:00
MacRimi
9d099ba358 Update 1.2.2.2 beta 2026-06-22 00:55:34 +02:00
MacRimi
e5669dd982 update 1.2.2.2 beta 2026-06-22 00:26:31 +02:00
MacRimi
09a67bf4e6 update 1.2.2.2 beta 2026-06-21 22:41:54 +02:00
MacRimi
ed924b67fe update 1.2.2.2 beta 2026-06-11 23:08:56 +02:00
MacRimi
61ff665cec update beta 1.2.2.2 2026-06-09 00:13:24 +02:00
MacRimi
6844406cf7 Update 1.2.2.1 2026-06-07 11:31:50 +02:00
MacRimi
61ff98e830 Update beta 1.2.2.1 2026-06-06 18:30:11 +02:00
MacRimi
66419777d8 Update beta 1.2.2.1 2026-06-06 11:37:54 +02:00
MacRimi
d401e5f7de Add new beta 1.2.2.1 2026-06-05 19:45:46 +02:00
github-actions[bot]
3103b87249 Update AppImage release build (2026-06-02 16:52:21) 2026-06-02 16:52:21 +00:00
MacRimi
5a116e77b9 Discord channel: split oversized digests across embeds (#220)
A mass-backup webhook that exceeded ~2 KB used to be silently
truncated by `desc = message[:MAX_EMBED_DESC]` with MAX_EMBED_DESC
set to 2048 — half of Discord's real description limit and far
below what a multi-VM backup digest produces. The trailing jobs
just vanished from the channel.

Bring the channel up to Discord's actual webhook contract:

* description limit raised to the real 4096-char cap
* if the body still doesn't fit, split it on line boundaries into
  one embed per chunk so every backup entry is preserved
* keep title + fields on the first embed only; attach the footer
  and timestamp to the last embed so the rendered card has the
  normal head/tail framing even when split across many embeds
* enforce Discord's 6000-char-per-embed cap (title + description +
  every field name+value) — only kicks in when many large fields
  combine with a chunk already near the description ceiling
* batch up to 10 embeds per webhook POST (Discord's per-message
  limit) and POST additional messages sequentially with a 0.4 s
  gap so a >10-embed digest doesn't trip the 5/2 s webhook rate
  limit

Verified with synthetic mass-backup payloads:
* 14 KB / 200 jobs → 4 embeds, 1 POST
* 60 KB / 60 lines → 15 embeds, 2 POSTs (10 + 5)

New AppImage SHA-256:
  16ad59ea63a64e5be460cd73f87315e8b39b756bf1c61f3cb2019e9fa3e76361

Closes #220.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-02 17:30:59 +02:00
MacRimi
17cae5d3a4 Refresh AppImage binary + sha256 after NVMe-obs-count fix
New build picks up the get_disks_observation_counts NVMe-rename fix.

SHA-256:
  3b44eb1172b4b1b7e6a36d1c9f1cd5a237ec04d52543bb791358525b0653a402

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-01 23:52:39 +02:00
MacRimi
4cd1cb4e39 Refresh AppImage binary + sha256 for v1.2.2
The tracked binary still pointed at the build made before the
last two fixes landed (resolution_reason persistence in
health_persistence and disk-temp breakdown alignment in
storage-overview). Re-build the AppImage so the GitHub-published
binary matches what is actually running on the deploy targets.

New SHA-256:
  d043e2f27f21315931ab53d87f02390b1a66b0c1730e8b7699aafb565809efbb

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-01 23:21:24 +02:00
MacRimi
3c5beb0286 Persist resolution_reason on resolve_error so the audit log is useful
The UPDATE in `_resolve_error_impl` only touched `resolved_at` — the
`reason` argument every caller passes was silently dropped, and the
`resolution_reason` / `resolution_type` columns stayed NULL for every
auto-resolved error. The columns were added back in a previous sprint
for exactly this audit-log purpose, but the writer was never updated
to populate them.

Fix the SQL to write `resolution_reason = ?` and tag
`resolution_type = COALESCE(existing, 'auto')` so admin-cleared
errors (whose type is set elsewhere) keep their value while the
default auto path correctly labels itself.

Verified end-to-end on the lab host: re-injected the `disk_nvme2n1`
warning, waited one scan cycle, the row now reads
`resolution_type='auto'` and
`resolution_reason='Transient I/O cleared, SMART now reports healthy'`
— previously these columns stayed NULL even though the resolve_error
call passed a descriptive reason.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-01 23:02:52 +02:00
MacRimi
9677c5cb19 Health Monitor: reconcile stale disk warnings across reboots
When a host gets transient I/O events on a disk while smartctl is
momentarily unavailable (the canonical case: late in a noisy
shutdown), the disk-scan code records a `disk_<name>` WARNING tagged
"SMART: unavailable" exactly once and trusts the next scan to clear
it. That trust is misplaced: the clear path only fires when the
device shows up in the current dmesg window with zero events. After
a reboot, dmesg is empty for that device — so the device never gets
iterated, resolve_error is never called, and the dashboard stays
orange for a disk whose SMART now reports PASSED.

Caught on a lab host where `disk_nvme2n1` had been stuck as WARNING
for hours after a reboot. SMART was 100% healthy at the moment of
inspection (Critical Warning 0x00, 0 media errors, 100% spare). The
error's first_seen and last_seen were identical and pre-dated the
current boot, confirming a one-shot record that nothing had cleared.

Fix: add a `_reconcile_stale_disk_warnings()` pass at the top of
`_check_disks_optimized()`. For every active `disk_*` error
(skipping `disk_fs_*`, which is already reconciled separately):

  - device gone from /dev/   → resolve "Device no longer present"
  - device present + SMART PASSED → resolve "Transient I/O cleared,
    SMART now reports healthy"
  - device present + SMART UNKNOWN/FAILED → leave active so the
    main loop can re-classify on the next dmesg window

Acknowledged errors are left alone so the user's explicit dismiss
intent isn't overridden.

Verified end-to-end: re-injected the original `disk_nvme2n1`
warning into the persistence DB on the lab host, waited one scan
cycle, error was resolved automatically with `resolved_at` set and
`resolution_reason = 'Transient I/O cleared, SMART now reports
healthy'`.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-01 22:54:14 +02:00