K3S-Deploy: harden HA bring-up and expand README caveats

Add a post-join control-plane readiness gate so a half-succeeded server
join on a degraded etcd can no longer pass as success. Retry the MetalLB
IPAddressPool/L2Advertisement applies to ride out the validating-webhook
startup race. Narrow the k3sup cleanup glob, quiet the final k3sup ready
tip, and clarify two comments.

README: add networking caveats (VIP/LB outside the DHCP pool, same-subnet
ARP/L2, interface naming), certName key placement, a post-run verification
section, machine-count/admin-box notes, and fix the RAW_BASE description.
This commit is contained in:
cyberops7
2026-07-15 22:26:45 -06:00
parent 49313239f4
commit c436d29182
2 changed files with 89 additions and 13 deletions
+51 -7
View File
@@ -38,26 +38,71 @@ duplicate LoadBalancer controllers.
- An **SSH key pair** you can use to reach the nodes (the script distributes
the public key and never clobbers your `~/.ssh/config`).
- At least **3 control-plane nodes** for real HA (etcd needs a quorum).
- Run the script from an **Ubuntu/Debian admin box** (it installs `k3sup`
and `kubectl` locally if missing).
- **5 nodes by default** (3 servers + 2 agents). The minimum for HA is 3
servers (etcd needs a quorum); workers are optional. Adjust the node list in
the config block if you have fewer machines.
- Run the script from a **separate Ubuntu/Debian admin machine** (your
laptop/workstation) that can SSH to every node — not on one of the nodes
themselves. It installs `k3sup` and `kubectl` locally if missing.
## Important caveats
- **The VIP and the MetalLB range must be outside your DHCP pool**, and must
not overlap each other or any node IP. If your router hands out `vip` or an
address in `lbrange` via DHCP, you'll get intermittent, hard-to-debug
failures — reserve them on your router/DHCP server first.
- **Same subnet / L2 only.** kube-vip (ARP mode) and MetalLB (L2) both
advertise via ARP, so the VIP and LoadBalancer IPs must be on the **same
subnet/VLAN** as the nodes and the clients reaching them — they are not
routed across subnets.
- **`interface` must match your nodes' real NIC.** It's `eth0` on many
systems but often `ens18`/`enp0s3` on Proxmox and cloud images. Check with
`ip -o -4 route show default`.
- **Host keys are trusted on first use** (`accept-new`). If you re-image a
node, clear its stale entry first: `ssh-keygen -R <node-ip>`.
- This is a **homelab tutorial, not production-hardened** as-is (aggressive
leader election, no etcd backups, single L2 domain).
## Usage
1. Snapshot your VMs.
2. Copy your SSH key into your home directory (or into `~/.ssh`).
1. Snapshot your VMs (so you can roll back — the script modifies every node).
2. Place your SSH **private** key at `~/<certName>` or `~/.ssh/<certName>`,
with the matching `.pub` beside it. `certName` is the key's *filename only*
— no path, no `.pub` suffix. Modern OpenSSH defaults to `id_ed25519`, so set
`certName` to match the key you actually have.
3. Edit the **"YOU SHOULD ONLY NEED TO EDIT THIS SECTION"** block in
`k3s.sh` — node IPs, `user`, `interface`, `vip`, `lbrange`, `certName`.
The example IPs (`192.168.3.x`, VIP `.50`, range `.60-.80`) and
`interface=eth0` are placeholders — change them for your network.
4. `chmod +x k3s.sh && ./k3s.sh`
5. Review the pre-flight summary and confirm. Grab a coffee.
### After it finishes
The kubeconfig is **merged** into `~/.kube/config` as context `k3s-ha` (set by
`context` in the config block), pointed at the VIP. Select it and check the
cluster:
```bash
kubectl config use-context k3s-ha
kubectl get nodes -o wide
curl http://<load-balancer-ip> # the script prints the assigned IP
```
Worker nodes are labeled `worker=true` and `longhorn=true` (the latter for the
companion Longhorn storage tutorial — harmless if you don't use it).
If a node fails to join, reset just that node and re-run the script (it's
idempotent): `k3s-uninstall.sh` on a server, `k3s-agent-uninstall.sh` on a
worker.
### Options (environment variables)
| Variable | Purpose |
| ---------------- | ----------------------------------------------------- |
| `ASSUME_YES=1` | Skip the pre-flight `[y/N]` prompt (unattended runs). |
| `NO_COLOR=1` | Disable colored output. |
| `RAW_BASE=<url>` | Base URL for the sibling manifests (default: `main`). |
| `RAW_BASE=<url>` | Where to fetch the sibling manifests (kube-vip, ipAddressPool, l2Advertisement); default: upstream `main`. Override for a fork/branch/local copy. |
### Tracking the latest k3s instead of a pinned version
@@ -75,7 +120,6 @@ docker run --rm ghcr.io/kube-vip/kube-vip:<version> manifest daemonset \
--controlplane --arp --leaderElection --taint --inCluster > kube-vip
sed -i 's/value: eth0/value: REPLACE_INTERFACE/' kube-vip
sed -i 's/value: 10.0.0.254/value: REPLACE_VIP/' kube-vip
```
`--inCluster` is required on k3s: it makes kube-vip use the `kube-vip`