mirror of
https://github.com/MacRimi/ProxMenux.git
synced 2026-10-08 22:46:41 +00:00
feat(oci): recover applications after a Proxmox reinstall, cluster records and AMD GPU profiles
The installation record of an OCI application travels with its container: a copy inside the container and another in /etc/pve, written together with the one kept on the host. A container restored on a newly installed Proxmox, restored with another ID or moved to another node of a cluster is recognised and registered again, with its private network, hookscript, Rclone mount, host firewall rule and NVIDIA runtime. The Monitor offers the same recovery from the Updates tab. AMD GPUs are offered by generation. A GPU the ROCm image supports takes the profile as it is; one of a supported family (Radeon 680M, 780M) is an experimental option that asks for confirmation and is never proposed; an older one is not offered. The GPU is checked with a real inference before the installation accepts it. Recreate changes what runs recognition in an installed Immich, between the CPU and a GPU of the host. Updates: - A failed update that is restored and checked removes its temporary container and the disks of the failed attempt. - Every container volume is part of the backups, so Jellyfin, Plex and Hugo update with their default installation. - An image published with a Docker-format manifest is recognised by its layers and build time and updates. - The Proxmox notes of a multi-container application link to its LAN address.
This commit is contained in:
@@ -68,6 +68,36 @@ def create_parent(path, root):
|
||||
os.chown(path, owner.st_uid, owner.st_gid)
|
||||
|
||||
|
||||
def rebuild(root, vmid, same_gpu=False):
|
||||
"""Point a stopped container at the NVIDIA driver of this host. The caller
|
||||
holds the registry and the container is not started: it is the step a
|
||||
restore on a host with another driver needs before the first start.
|
||||
Returns whether anything had to change."""
|
||||
record = instances.read(root, vmid)
|
||||
config = instances.command('pct', 'config', str(vmid))
|
||||
if not instances.same_config_except_notes(record, config):
|
||||
raise ValueError(translate('The container identity or configuration changed'))
|
||||
plan = nv.refresh_plan(config, record['observed']['gpu_devices'][nv.KEY], same_gpu=same_gpu)
|
||||
if not plan['changed']:
|
||||
return False
|
||||
if instances.command('pct', 'status', str(vmid)).strip() != b'status: stopped':
|
||||
raise ValueError(translate('Stop the container before the NVIDIA refresh'))
|
||||
instances.command('pct', 'mount', str(vmid))
|
||||
try:
|
||||
candidate = prepare(Path(f'/var/lib/lxc/{vmid}/rootfs'), plan)
|
||||
finally:
|
||||
instances.command('pct', 'unmount', str(vmid))
|
||||
Path(f'/etc/pve/lxc/{vmid}.conf').write_bytes(candidate)
|
||||
updated = copy.deepcopy(record)
|
||||
updated['observed'] = instances.observe(vmid, record['installation_id'],
|
||||
record['observed']['archive_path'], record['observed']['resolved_registry_digest'],
|
||||
record['observed']['image'])
|
||||
nv.check_mounts(updated['observed']['config'].encode(), plan['inventory'])
|
||||
nv.check_devices(updated['observed']['config'].encode(), plan['inventory'])
|
||||
instances.write(instances.location(root, vmid), updated)
|
||||
return True
|
||||
|
||||
|
||||
def refresh(root, vmid, apply=False):
|
||||
with instances.locked(root):
|
||||
record = instances.read(root, vmid)
|
||||
|
||||
Reference in New Issue
Block a user