Merge pull request #170 from JamesTurland/k3s-deploy-overhaul

Overhaul K3S-Deploy (k3sup + kube-vip + MetalLB)
This commit is contained in:
DefNotJeffrey
2026-08-22 15:16:59 +02:00
committed by GitHub
6 changed files with 508 additions and 194 deletions
+151
View File
@@ -0,0 +1,151 @@
# K3S HA Deploy (k3sup + kube-vip + MetalLB)
Deploys a highly-available k3s cluster over SSH with `k3sup`, a floating
control-plane VIP via **kube-vip**, and `type: LoadBalancer` support via
**MetalLB**.
> **Following the YouTube tutorial?** This script has been modernized since
> the video — see [What changed](#what-changed-from-the-video) below.
## What it builds
- A 3-server (control-plane) + 2-agent (worker) k3s cluster with embedded etcd.
- **kube-vip** advertises a single virtual IP for the Kubernetes API across
all control-plane nodes (survives a node failure).
- **MetalLB** hands out real IPs to `type: LoadBalancer` services.
### Why both kube-vip and MetalLB?
They solve different problems. **kube-vip** provides the *control-plane* VIP
(one address for the API server, HA across masters). **MetalLB** provides
*service* load balancing (`type: LoadBalancer` for your apps). The older
kube-vip cloud-provider that overlapped MetalLB has been removed to avoid
duplicate LoadBalancer controllers.
## Prerequisites
- **Ubuntu/Debian** nodes (the script installs prerequisites with `apt`; on a
non-apt distro it exits with a clear message).
- **Passwordless sudo** for the SSH user on every node. Example cloud-init:
```yaml
#cloud-config
users:
- name: <your-user>
sudo: ["ALL=(ALL) NOPASSWD:ALL"]
groups: [sudo]
```
- An **SSH key pair** you can use to reach the nodes (the script distributes
the public key and never clobbers your `~/.ssh/config`).
- **5 nodes by default** (3 servers + 2 agents). The minimum for HA is 3
servers (etcd needs a quorum); workers are optional. Adjust the node list in
the config block if you have fewer machines.
- Run the script from a **separate Ubuntu/Debian admin machine** (your
laptop/workstation) that can SSH to every node — not on one of the nodes
themselves. It installs `k3sup` and `kubectl` locally if missing.
## Important caveats
- **The VIP and the MetalLB range must be outside your DHCP pool**, and must
not overlap each other or any node IP. If your router hands out `vip` or an
address in `lbrange` via DHCP, you'll get intermittent, hard-to-debug
failures — reserve them on your router/DHCP server first.
- **Same subnet / L2 only.** kube-vip (ARP mode) and MetalLB (L2) both
advertise via ARP, so the VIP and LoadBalancer IPs must be on the **same
subnet/VLAN** as the nodes and the clients reaching them — they are not
routed across subnets.
- **`interface` must match your nodes' real NIC.** It's `eth0` on many
systems but often `ens18`/`enp0s3` on Proxmox and cloud images. Check with
`ip -o -4 route show default`.
- **Host keys are trusted on first use** (`accept-new`). If you re-image a
node, clear its stale entry first: `ssh-keygen -R <node-ip>`.
- This is a **homelab tutorial, not production-hardened** as-is (aggressive
leader election, no etcd backups, single L2 domain).
## Usage
1. Snapshot your VMs (so you can roll back — the script modifies every node).
2. Place your SSH **private** key at `~/<certName>` or `~/.ssh/<certName>`,
with the matching `.pub` beside it. `certName` is the key's *filename only*
— no path, no `.pub` suffix. Modern OpenSSH defaults to `id_ed25519`, so set
`certName` to match the key you actually have.
3. Edit the **"YOU SHOULD ONLY NEED TO EDIT THIS SECTION"** block in
`k3s.sh` — node IPs, `user`, `interface`, `vip`, `lbrange`, `certName`.
The example IPs (`192.168.3.x`, VIP `.50`, range `.60-.80`) and
`interface=eth0` are placeholders — change them for your network.
4. `chmod +x k3s.sh && ./k3s.sh`
5. Review the pre-flight summary and confirm. Grab a coffee.
### After it finishes
The kubeconfig is **merged** into `~/.kube/config` as context `k3s-ha` (set by
`context` in the config block), pointed at the VIP. Select it and check the
cluster:
```bash
kubectl config use-context k3s-ha
kubectl get nodes -o wide
curl http://<load-balancer-ip> # the script prints the assigned IP
```
Worker nodes are labeled `worker=true` and `longhorn=true` (the latter for the
companion Longhorn storage tutorial — harmless if you don't use it).
If a node fails to join, reset just that node and re-run the script (it's
idempotent): `k3s-uninstall.sh` on a server, `k3s-agent-uninstall.sh` on a
worker.
### Options (environment variables)
| Variable | Purpose |
| ---------------- | ----------------------------------------------------- |
| `ASSUME_YES=1` | Skip the pre-flight `[y/N]` prompt (unattended runs). |
| `NO_COLOR=1` | Disable colored output. |
| `RAW_BASE=<url>` | Where to fetch the sibling manifests (kube-vip, ipAddressPool, l2Advertisement); default: upstream `main`. Override for a fork/branch/local copy. |
### Tracking the latest k3s instead of a pinned version
In the config block set `k3sChannel="stable"` and leave `k3sVersion=""`.
The script then always installs the current stable k3s release.
### Upgrading kube-vip
The `kube-vip` manifest is generated from the pinned image (its env schema
changes between releases). To bump:
```bash
docker run --rm ghcr.io/kube-vip/kube-vip:<version> manifest daemonset \
--interface eth0 --address 10.0.0.254 \
--controlplane --arp --leaderElection --taint --inCluster > kube-vip
sed -i 's/value: eth0/value: REPLACE_INTERFACE/' kube-vip
sed -i 's/value: 10.0.0.254/value: REPLACE_VIP/' kube-vip
```
`--inCluster` is required on k3s: it makes kube-vip use the `kube-vip`
ServiceAccount (created by the RBAC the script applies) instead of the
kubeadm `/etc/kubernetes/admin.conf` kubeconfig, which does not exist on k3s.
Then update `KVVERSION` in `k3s.sh` to match.
## What changed from the video
- **Versions:** k3s `v1.35.6+k3s1`, kube-vip `v1.2.1`, MetalLB `v0.16.0`.
- **kube-vip** now runs on **all** control-plane nodes, and your kubeconfig
points at the **VIP** (not master1) — fixes `localhost:8080`/API errors. A
readiness check confirms the VIP is answering before the script continues.
- Removed the redundant kube-vip **cloud-provider**; MetalLB alone handles
`type: LoadBalancer`.
- **MetalLB** installs from a single native manifest (which creates its own
`metallb-system` namespace) — the old separate namespace apply and its
mismatched MetalLB versions are gone.
- **k3sup:** the join token is fetched once and reused across all nodes, and
`k3sup ready` waits for the cluster instead of a hand-rolled poll loop.
- **Safer SSH:** host keys are added via `ssh-keyscan`; the script no longer
overwrites `~/.ssh/config`.
- **Hardening:** `set -euo pipefail`, per-node time sync, `apt` prerequisite
install with a clear message on unsupported distros, arch-aware `kubectl`,
and it is safe to re-run.
- **Cleaner output:** non-blinking step-by-step logging (only the banner
still blinks — for old times' sake), a pre-flight config summary with a
`[y/N]` confirmation (`ASSUME_YES=1` to skip), and a final summary listing
the API VIP and the assigned LoadBalancer IP.
+1 -1
View File
@@ -5,4 +5,4 @@ metadata:
namespace: metallb-system
spec:
addresses:
- $lbrange
- REPLACE_LBRANGE
Regular → Executable
+334 -162
View File
@@ -1,232 +1,404 @@
#!/bin/bash
#!/usr/bin/env bash
#
# JimsGarage — K3S HA deploy (k3sup + kube-vip + MetalLB)
# Tutorial: https://youtube.com/@jims-garage
# See ./README.md for prerequisites, configuration, and what changed
# from the original video.
set -euo pipefail
# ── Banner (kept from the original tutorial; blink preserved) ──────────
echo -e " \033[33;5m __ _ _ ___ \033[0m"
echo -e " \033[33;5m \ \(_)_ __ ___( )__ / _ \__ _ _ __ __ _ __ _ ___ \033[0m"
echo -e " \033[33;5m \ \ | '_ \` _ \/ __| / /_\/ _\` | '__/ _\` |/ _\` |/ _ \ \033[0m"
echo -e " \033[33;5m /\_/ / | | | | | \__ \ / /_\\ (_| | | | (_| | (_| | __/ \033[0m"
echo -e " \033[33;5m \___/|_|_| |_| |_|___/ \____/\__,_|_| \__,_|\__, |\___| \033[0m"
echo -e " \033[33;5m |___/ \033[0m"
echo -e " \033[36;5m _ _________ ___ _ _ _ \033[0m"
echo -e " \033[36;5m | |/ |__ / __| |_ _|_ _ __| |_ __ _| | | \033[0m"
echo -e " \033[36;5m | ' < |_ \__ \ | || ' \(_-| _/ _\` | | | \033[0m"
echo -e " \033[36;5m |_|\_|___|___/ |___|_||_/__/\__\__,_|_|_| \033[0m"
echo -e " \033[36;5m \033[0m"
echo -e " \033[32;5m https://youtube.com/@jims-garage \033[0m"
echo -e " \033[32;5m \033[0m"
# ── Output helpers (non-blinking; auto-disable when not a TTY) ──────────
if [[ -t 1 && -z "${NO_COLOR:-}" ]]; then
c_reset=$'\033[0m'; c_step=$'\033[1;36m'; c_info=$'\033[36m'
c_ok=$'\033[32m'; c_warn=$'\033[33m'; c_err=$'\033[31m'
else
c_reset=; c_step=; c_info=; c_ok=; c_warn=; c_err=
fi
step() { printf '\n%s━━━ %s ━━━%s\n' "$c_step" "$*" "$c_reset"; }
info() { printf ' %s•%s %s\n' "$c_info" "$c_reset" "$*"; }
ok() { printf ' %s✓%s %s\n' "$c_ok" "$c_reset" "$*"; }
warn() { printf ' %s!%s %s\n' "$c_warn" "$c_reset" "$*" >&2; }
err() { printf ' %s✗%s %s\n' "$c_err" "$c_reset" "$*" >&2; }
die() { err "$*"; exit 1; }
#############################################
# YOU SHOULD ONLY NEED TO EDIT THIS SECTION #
#############################################
# Version of Kube-VIP to deploy
KVVERSION="v0.6.3"
# k3s version to install. To always track the latest stable release
# instead, set k3sChannel="stable" and leave k3sVersion empty (see README).
k3sVersion="v1.35.6+k3s1"
k3sChannel=""
# K3S Version
k3sVersion="v1.26.10+k3s2"
# kube-vip: must match the image tag baked into the ./kube-vip manifest.
# To bump, regenerate ./kube-vip (see README "Upgrading kube-vip").
KVVERSION="v1.2.1"
# Set the IP addresses of the master and work nodes
# MetalLB version (used to build the manifest URL).
METALLB_VERSION="v0.16.0"
# Node IP addresses
master1=192.168.3.21
master2=192.168.3.22
master3=192.168.3.23
worker1=192.168.3.24
worker2=192.168.3.25
# User of remote machines
# SSH user on the remote nodes
user=ubuntu
# Interface used on remotes
# Network interface used on the remote nodes
interface=eth0
# Set the virtual IP address (VIP)
# Virtual IP (VIP) for the HA control plane
vip=192.168.3.50
# Array of master nodes
masters=($master2 $master3)
# Array of worker nodes
workers=($worker1 $worker2)
# Array of all
all=($master1 $master2 $master3 $worker1 $worker2)
# Array of all minus master
allnomaster1=($master2 $master3 $worker1 $worker2)
#Loadbalancer IP range
# MetalLB LoadBalancer address range
lbrange=192.168.3.60-192.168.3.80
#ssh certificate name variable
# SSH private key name (in ~/.ssh) used to reach the nodes
certName=id_rsa
#ssh config file
config_file=~/.ssh/config
# kubeconfig context name to create locally
context=k3s-ha
#############################################
# DO NOT EDIT BELOW #
#############################################
# For testing purposes - in case time is wrong due to VM snapshots
sudo timedatectl set-ntp off
sudo timedatectl set-ntp on
# Move SSH certs to ~/.ssh and change permissions
cp /home/$user/{$certName,$certName.pub} /home/$user/.ssh
chmod 600 /home/$user/.ssh/$certName
chmod 644 /home/$user/.ssh/$certName.pub
# Additional control-plane servers joined after master1
masters=("$master2" "$master3")
# Agent (worker) nodes
workers=("$worker1" "$worker2")
# All control-plane nodes (kube-vip manifest is placed on each)
masters_all=("$master1" "${masters[@]}")
# Every node (for SSH prep loops)
all=("$master1" "${masters[@]}" "${workers[@]}")
# Install k3sup to local machine if not already present
if ! command -v k3sup version &> /dev/null
then
echo -e " \033[31;5mk3sup not found, installing\033[0m"
curl -sLS https://get.k3sup.dev | sh
sudo install k3sup /usr/local/bin/
# Base URL for the sibling manifests (kube-vip, ipAddressPool,
# l2Advertisement). Override to test from a branch/fork/local copy, e.g.
# RAW_BASE="file://$HOME/JimsGarage"
RAW_BASE="${RAW_BASE:-https://raw.githubusercontent.com/JamesTurland/JimsGarage/main}"
manifest_base="$RAW_BASE/Kubernetes/K3S-Deploy"
ssh_key="$HOME/.ssh/$certName"
ssh_opts=(-i "$ssh_key" -o StrictHostKeyChecking=accept-new)
# k3s version selector: prefer an explicit version, else a channel.
if [[ -n "$k3sVersion" ]]; then
k3s_selector=(--k3s-version "$k3sVersion"); k3s_display="$k3sVersion"
else
echo -e " \033[32;5mk3sup already installed\033[0m"
k3s_selector=(--k3s-channel "${k3sChannel:-stable}"); k3s_display="channel:${k3sChannel:-stable}"
fi
# Install Kubectl if not already present
if ! command -v kubectl version &> /dev/null
then
echo -e " \033[31;5mKubectl not found, installing\033[0m"
curl -LO "https://dl.k8s.io/release/$(curl -L -s https://dl.k8s.io/release/stable.txt)/bin/linux/amd64/kubectl"
sudo install -o root -g root -m 0755 kubectl /usr/local/bin/kubectl
# Control-plane k3s args: disable bundled traefik + servicelb (we use
# MetalLB), add the VIP to the API serving cert SAN, pin the flannel interface
# + node IP, taint as control-plane. Applied to EVERY server so master2/3
# declare the VIP SAN too (not just master1).
# Explicit --disable is kept (not k3sup --no-extras) for tutorial clarity.
server_extra_args() { # $1 = node ip
printf '%s' "--disable traefik --disable servicelb --tls-san=$vip --flannel-iface=$interface --node-ip=$1 --node-taint node-role.kubernetes.io/control-plane=true:NoSchedule"
}
# k3sup prints a promo tip ("Create clusters on Mac…/slicervm.com") on every
# install/join. Filter it out. Use sed, not grep -v: sed always exits 0, so
# pipefail still surfaces a real k3sup failure (grep would exit 1 on no match
# and spuriously abort under set -e). Only wraps the noisy display commands —
# never node-token (whose stdout is captured for the token).
k3sup_quiet() { command k3sup "$@" 2>&1 | sed '/Create clusters on Mac/d'; }
# Apply a manifest, retrying on transient failure. MetalLB installs a
# validating webhook served by its controller; there is a brief window where
# the controller Deployment reports Available but the webhook endpoints/TLS
# are not serving yet, so the first apply can fail with "no endpoints
# available for service metallb-webhook-service". Retry to ride that out.
kubectl_apply_retry() { # $1 = manifest path or URL
local i
for i in $(seq 1 6); do
kubectl apply -f "$1" && return 0
warn "kubectl apply $1 failed (attempt $i/6) — retrying in 5s"
sleep 5
done
die "kubectl apply $1 failed after 6 attempts"
}
# ── Pre-flight: show config, confirm before touching any node ──────────
step "Pre-flight · review configuration"
printf ' %-16s %s\n' \
"k3s:" "$k3s_display" \
"kube-vip:" "$KVVERSION" \
"metallb:" "$METALLB_VERSION" \
"control plane:" "${masters_all[*]}" \
"workers:" "${workers[*]}" \
"api VIP:" "$vip (interface $interface)" \
"lb range:" "$lbrange" \
"ssh user/key:" "$user / $ssh_key" \
"kube context:" "$context"
if [[ "${ASSUME_YES:-}" == "1" ]]; then
info "ASSUME_YES=1 — proceeding without prompt"
else
echo -e " \033[32;5mKubectl already installed\033[0m"
read -rp "$(printf '\n Proceed? [y/N] ')" reply || reply=""
[[ "$reply" =~ ^[Yy]$ ]] || die "Aborted."
fi
# Check for SSH config file, create if needed, add/change Strict Host Key Checking (don't use in production!)
if [ ! -f "$config_file" ]; then
# Create the file and add the line
echo "StrictHostKeyChecking no" > "$config_file"
# Set permissions to read and write only for the owner
chmod 600 "$config_file"
echo "File created and line added."
# ── Step 1/9 · Local tools (k3sup, kubectl) ────────────────────────────
step "Step 1/9 · Local tools (k3sup, kubectl)"
if ! command -v k3sup &>/dev/null; then
info "Installing k3sup"
curl -sLS https://get.k3sup.dev | sh
# The installer leaves the binary in the CWD — named "k3sup", or
# "k3sup-<arch>" when run unprivileged (it can't self-install to
# /usr/local/bin). Install whichever it produced.
k3sup_bin=""
for f in k3sup k3sup-*; do
[[ -f "$f" ]] && { k3sup_bin="$f"; break; }
done
[[ -n "$k3sup_bin" ]] || die "k3sup installer produced no binary"
sudo install "$k3sup_bin" /usr/local/bin/k3sup
rm -f "$k3sup_bin"
else
# Check if the line exists
if grep -q "^StrictHostKeyChecking" "$config_file"; then
# Check if the value is not "no"
if ! grep -q "^StrictHostKeyChecking no" "$config_file"; then
# Replace the existing line
sed -i 's/^StrictHostKeyChecking.*/StrictHostKeyChecking no/' "$config_file"
echo "Line updated."
else
echo "Line already set to 'no'."
fi
else
# Add the line to the end of the file
echo "StrictHostKeyChecking no" >> "$config_file"
echo "Line added."
fi
ok "k3sup present"
fi
if ! command -v kubectl &>/dev/null; then
info "Installing kubectl"
arch="$(dpkg --print-architecture 2>/dev/null || uname -m)"
case "$arch" in
amd64|x86_64) arch=amd64 ;;
arm64|aarch64) arch=arm64 ;;
*) die "Unsupported architecture for kubectl: $arch" ;;
esac
kver="$(curl -L -s https://dl.k8s.io/release/stable.txt)"
curl -LO "https://dl.k8s.io/release/${kver}/bin/linux/${arch}/kubectl"
sudo install -o root -g root -m 0755 kubectl /usr/local/bin/kubectl
else
ok "kubectl present"
fi
#add ssh keys for all nodes
# ── Step 2/9 · SSH keys and known_hosts ────────────────────────────────
step "Step 2/9 · SSH keys and known_hosts"
mkdir -p "$HOME/.ssh"; chmod 700 "$HOME/.ssh"
# If the key was dropped in $HOME (per the README), move it into ~/.ssh.
if [[ -f "$HOME/$certName" && ! -f "$ssh_key" ]]; then
info "Moving $certName into ~/.ssh"
cp "$HOME/$certName" "$ssh_key"
[[ -f "$HOME/$certName.pub" ]] && cp "$HOME/$certName.pub" "$ssh_key.pub"
fi
[[ -f "$ssh_key" ]] || die "SSH key not found: $ssh_key (see README prerequisites)"
chmod 600 "$ssh_key"; [[ -f "$ssh_key.pub" ]] && chmod 644 "$ssh_key.pub"
# Trust host keys without clobbering ~/.ssh/config (issue #62).
for node in "${all[@]}"; do
ssh-copy-id $user@$node
ssh-keyscan "$node" 2>/dev/null >> "$HOME/.ssh/known_hosts" || \
warn "ssh-keyscan failed for $node (will accept-new on first connect)"
done
[[ -f "$HOME/.ssh/known_hosts" ]] && sort -u "$HOME/.ssh/known_hosts" -o "$HOME/.ssh/known_hosts"
# Distribute the public key to each node.
for node in "${all[@]}"; do
info "Copying SSH key to $user@$node"
ssh-copy-id -i "$ssh_key.pub" -o StrictHostKeyChecking=accept-new "$user@$node" >/dev/null 2>&1 || \
warn "ssh-copy-id to $node failed (key may already be present)"
done
ok "SSH ready for ${#all[@]} nodes"
# ── Step 3/9 · Prepare nodes (time sync + prerequisites) ───────────────
step "Step 3/9 · Prepare nodes (time sync + prerequisites)"
for node in "${all[@]}"; do
info "Preparing $node"
rc=0
ssh "${ssh_opts[@]}" "$user@$node" 'sudo bash -s' <<'REMOTE' || rc=$?
set -e
# Resync time — VM snapshots drift, which breaks k3s/k3sup installs (issue #68).
timedatectl set-ntp off || true
timedatectl set-ntp on || true
if command -v apt-get >/dev/null 2>&1; then
export DEBIAN_FRONTEND=noninteractive NEEDRESTART_MODE=a
ok=0
for i in $(seq 1 10); do
if apt-get update && apt-get install -y iptables sudo policycoreutils; then ok=1; break; fi
echo "apt busy, retry $i/10..."; sleep 3
done
[ "$ok" = 1 ] || exit 91
else
exit 90
fi
REMOTE
case "$rc" in
0) ok "$node prepared" ;;
90) die "Node $node is not apt-based. This script targets Ubuntu/Debian; install iptables, sudo, and policycoreutils manually (see README)." ;;
91) die "Node $node: apt failed after retries." ;;
*) die "Node $node: preparation failed (exit $rc)." ;;
esac
done
# Install policycoreutils for each node
for newnode in "${all[@]}"; do
ssh $user@$newnode -i ~/.ssh/$certName sudo su <<EOF
NEEDRESTART_MODE=a apt-get install policycoreutils -y
exit
EOF
echo -e " \033[32;5mPolicyCoreUtils installed!\033[0m"
done
# Step 1: Bootstrap First k3s Node
mkdir ~/.kube
k3sup install \
--ip $master1 \
--user $user \
--tls-san $vip \
--cluster \
--k3s-version $k3sVersion \
--k3s-extra-args "--disable traefik --disable servicelb --flannel-iface=$interface --node-ip=$master1 --node-taint node-role.kubernetes.io/master=true:NoSchedule" \
--merge \
# ── Step 4/9 · Bootstrap first control-plane node ──────────────────────
step "Step 4/9 · Bootstrap first control-plane node ($master1)"
mkdir -p "$HOME/.kube"
k3sup_quiet install \
--ip "$master1" \
--user "$user" \
--sudo \
--local-path $HOME/.kube/config \
--ssh-key $HOME/.ssh/$certName \
--context k3s-ha
echo -e " \033[32;5mFirst Node bootstrapped successfully!\033[0m"
--cluster \
"${k3s_selector[@]}" \
--k3s-extra-args "$(server_extra_args "$master1")" \
--merge \
--local-path "$HOME/.kube/config" \
--ssh-key "$ssh_key" \
--context "$context"
ok "First node bootstrapped"
# Step 2: Install Kube-VIP for HA
# ── Step 5/9 · Install kube-vip (control-plane VIP) ────────────────────
step "Step 5/9 · Install kube-vip (control-plane VIP)"
# Ensure kubectl targets the cluster we just created (the merge above may
# have left a different current-context in a pre-existing ~/.kube/config).
kubectl config use-context "$context" >/dev/null
info "Applying kube-vip RBAC"
kubectl apply -f https://kube-vip.io/manifests/rbac.yaml
info "Rendering kube-vip manifest (interface=$interface, vip=$vip)"
curl -sfL "$manifest_base/kube-vip" -o "$HOME/kube-vip.src"
sed "s/REPLACE_INTERFACE/$interface/g; s/REPLACE_VIP/$vip/g" \
"$HOME/kube-vip.src" > "$HOME/kube-vip.yaml"
for node in "${masters_all[@]}"; do
info "Placing kube-vip manifest on $node"
# scp to /tmp (absolute) — the mv below runs as root, where ~ is /root,
# not the ssh user's home where scp would otherwise land the file.
scp "${ssh_opts[@]}" "$HOME/kube-vip.yaml" "$user@$node:/tmp/kube-vip.yaml" >/dev/null
ssh "${ssh_opts[@]}" "$user@$node" 'sudo bash -s' <<'REMOTE'
set -e
mkdir -p /var/lib/rancher/k3s/server/manifests
mv /tmp/kube-vip.yaml /var/lib/rancher/k3s/server/manifests/kube-vip.yaml
chown root:root /var/lib/rancher/k3s/server/manifests/kube-vip.yaml
chmod 600 /var/lib/rancher/k3s/server/manifests/kube-vip.yaml
REMOTE
done
ok "kube-vip deployed to ${#masters_all[@]} control-plane nodes"
info "Pointing local kubeconfig at the VIP ($vip)"
# Rewrite the server URL from master1 to the VIP. The dots in the IP are
# unescaped regex metacharacters, but a false match would need the literal
# "https://<master1-ish>:6443" string elsewhere in the kubeconfig, which does
# not occur (cert data is base64 — no scheme or port), so this is safe here.
sed -i "s#https://$master1:6443#https://$vip:6443#" "$HOME/.kube/config"
# Step 3: Download kube-vip
curl -sO https://raw.githubusercontent.com/JamesTurland/JimsGarage/main/Kubernetes/K3S-Deploy/kube-vip
cat kube-vip | sed 's/$interface/'$interface'/g; s/$vip/'$vip'/g' > $HOME/kube-vip.yaml
# Confirm the VIP is actually answering before any VIP-routed kubectl below.
# This is an early tripwire for a broken kube-vip (wrong interface, RBAC, etc.)
# instead of a confusing failure further down.
info "Waiting for the control-plane VIP ($vip) to answer"
vip_ready=""
for _ in $(seq 1 30); do
if kubectl --request-timeout=5s get --raw='/readyz' &>/dev/null; then
vip_ready=1; break
fi
sleep 2
done
[[ -n "$vip_ready" ]] || die "Control-plane VIP $vip is not answering — check kube-vip on the masters: kubectl -n kube-system logs -l app.kubernetes.io/name=kube-vip-ds"
ok "Control-plane VIP is answering"
# Step 4: Copy kube-vip.yaml to master1
scp -i ~/.ssh/$certName $HOME/kube-vip.yaml $user@$master1:~/kube-vip.yaml
# ── Step 6/9 · Fetch the cluster join token ────────────────────────────
step "Step 6/9 · Fetch the cluster join token"
node_token="$(k3sup node-token --ip "$master1" --user "$user" --ssh-key "$ssh_key")" \
|| die "Failed to fetch node token from $master1"
[[ -n "$node_token" ]] || die "Empty node token from $master1"
ok "Join token fetched"
# Step 5: Connect to Master1 and move kube-vip.yaml
ssh $user@$master1 -i ~/.ssh/$certName <<- EOF
sudo mkdir -p /var/lib/rancher/k3s/server/manifests
sudo mv kube-vip.yaml /var/lib/rancher/k3s/server/manifests/kube-vip.yaml
EOF
# Step 6: Add new master nodes (servers) & workers
for newnode in "${masters[@]}"; do
k3sup join \
--ip $newnode \
--user $user \
# ── Step 7/9 · Join control-plane and worker nodes ─────────────────────
step "Step 7/9 · Join control-plane and worker nodes"
# The prefetched --node-token is what keeps these joins from ever SSHing to
# the server: without it, k3sup would try --server-user (default: root) against
# --server-ip (the VIP, not an SSH target) and fail confusingly. Keep the token.
for node in "${masters[@]}"; do
info "Joining control-plane node $node"
k3sup_quiet join \
--ip "$node" \
--user "$user" \
--sudo \
--k3s-version $k3sVersion \
--server \
--server-ip $master1 \
--ssh-key $HOME/.ssh/$certName \
--k3s-extra-args "--disable traefik --disable servicelb --flannel-iface=$interface --node-ip=$newnode --node-taint node-role.kubernetes.io/master=true:NoSchedule" \
--server-user $user
echo -e " \033[32;5mMaster node joined successfully!\033[0m"
--server-ip "$vip" \
--node-token "$node_token" \
"${k3s_selector[@]}" \
--k3s-extra-args "$(server_extra_args "$node")" \
--ssh-key "$ssh_key"
ok "Control-plane node $node joined"
done
# add workers
for newagent in "${workers[@]}"; do
k3sup join \
--ip $newagent \
--user $user \
for node in "${workers[@]}"; do
info "Joining worker node $node"
k3sup_quiet join \
--ip "$node" \
--user "$user" \
--sudo \
--k3s-version $k3sVersion \
--server-ip $master1 \
--ssh-key $HOME/.ssh/$certName \
--k3s-extra-args "--node-label \"longhorn=true\" --node-label \"worker=true\""
echo -e " \033[32;5mAgent node joined successfully!\033[0m"
--server-ip "$vip" \
--node-token "$node_token" \
"${k3s_selector[@]}" \
--k3s-extra-args "--flannel-iface=$interface --node-ip=$node --node-label longhorn=true --node-label worker=true" \
--ssh-key "$ssh_key"
ok "Worker node $node joined"
done
# Step 7: Install kube-vip as network LoadBalancer - Install the kube-vip Cloud Provider
kubectl apply -f https://raw.githubusercontent.com/kube-vip/kube-vip-cloud-provider/main/manifest/kube-vip-cloud-controller.yaml
# Assert the full HA control plane is up before continuing. The VIP gate in
# Step 5 only proved master1 (a single-member etcd); a control-plane join that
# half-succeeds would otherwise sail through as success on a degraded cluster.
info "Waiting for all ${#masters_all[@]} control-plane nodes to be Ready"
kubectl wait --for=condition=Ready node \
-l node-role.kubernetes.io/control-plane --timeout=180s || \
die "Control-plane nodes did not all become Ready — check etcd/kube-vip on the masters: kubectl get nodes; kubectl -n kube-system logs -l app.kubernetes.io/name=kube-vip-ds"
cp_ready="$(kubectl get nodes -l node-role.kubernetes.io/control-plane --no-headers 2>/dev/null | awk '$2=="Ready"{c++} END{print c+0}')"
[[ "$cp_ready" -eq "${#masters_all[@]}" ]] || \
die "Expected ${#masters_all[@]} control-plane nodes Ready, found $cp_ready — etcd quorum may be degraded"
ok "All ${#masters_all[@]} control-plane nodes Ready"
# Step 8: Install Metallb
kubectl apply -f https://raw.githubusercontent.com/metallb/metallb/v0.12.1/manifests/namespace.yaml
kubectl apply -f https://raw.githubusercontent.com/metallb/metallb/v0.13.12/config/manifests/metallb-native.yaml
# Download ipAddressPool and configure using lbrange above
curl -sO https://raw.githubusercontent.com/JamesTurland/JimsGarage/main/Kubernetes/K3S-Deploy/ipAddressPool
cat ipAddressPool | sed 's/$lbrange/'$lbrange'/g' > $HOME/ipAddressPool.yaml
kubectl apply -f $HOME/ipAddressPool.yaml
# ── Step 8/9 · Install MetalLB (service LoadBalancer) ──────────────────
step "Step 8/9 · Install MetalLB (service LoadBalancer)"
kubectl apply -f "https://raw.githubusercontent.com/metallb/metallb/$METALLB_VERSION/config/manifests/metallb-native.yaml"
info "Waiting for the MetalLB controller"
# rollout status waits on the Deployment (created synchronously by apply) and
# avoids the "no matching resources found" race that `wait --for=condition=ready
# pod` hits when the controller pod does not exist yet.
kubectl -n metallb-system rollout status deploy/controller --timeout=120s
info "Configuring address pool ($lbrange)"
curl -sfL "$manifest_base/ipAddressPool" -o "$HOME/ipAddressPool.src"
sed "s|REPLACE_LBRANGE|$lbrange|g" "$HOME/ipAddressPool.src" > "$HOME/ipAddressPool.yaml"
kubectl_apply_retry "$HOME/ipAddressPool.yaml"
curl -sfL "$manifest_base/l2Advertisement.yaml" -o "$HOME/l2Advertisement.yaml"
kubectl_apply_retry "$HOME/l2Advertisement.yaml"
ok "MetalLB configured"
# Step 9: Test with Nginx
kubectl apply -f https://raw.githubusercontent.com/inlets/inlets-operator/master/contrib/nginx-sample-deployment.yaml -n default
kubectl expose deployment nginx-1 --port=80 --type=LoadBalancer -n default
echo -e " \033[32;5mWaiting for K3S to sync and LoadBalancer to come online\033[0m"
while [[ $(kubectl get pods -l app=nginx -o 'jsonpath={..status.conditions[?(@.type=="Ready")].status}') != "True" ]]; do
sleep 1
# ── Step 9/9 · Verify cluster and LoadBalancer ─────────────────────────
step "Step 9/9 · Verify cluster and LoadBalancer"
info "Waiting for the cluster to be ready"
k3sup_quiet ready --context "$context" --kubeconfig "$HOME/.kube/config"
info "Deploying nginx sample"
kubectl apply -n default -f https://raw.githubusercontent.com/inlets/inlets-operator/master/contrib/nginx-sample-deployment.yaml
# `|| true` — expose is imperative and errors with AlreadyExists on a re-run,
# which would abort the script under set -e.
kubectl expose deployment nginx-1 --port=80 --type=LoadBalancer -n default 2>/dev/null || true
info "Waiting for the nginx deployment"
kubectl -n default rollout status deploy/nginx-1 --timeout=120s
info "Waiting for the LoadBalancer IP"
lb_ip=""
for _ in $(seq 1 30); do
lb_ip="$(kubectl get svc nginx-1 -n default -o jsonpath='{.status.loadBalancer.ingress[0].ip}' 2>/dev/null || true)"
[[ -n "$lb_ip" ]] && break
sleep 2
done
# Step 10: Deploy IP Pools and l2Advertisement
kubectl wait --namespace metallb-system \
--for=condition=ready pod \
--selector=component=controller \
--timeout=120s
kubectl apply -f ipAddressPool.yaml
kubectl apply -f https://raw.githubusercontent.com/JamesTurland/JimsGarage/main/Kubernetes/K3S-Deploy/l2Advertisement.yaml
kubectl get nodes
kubectl get svc
kubectl get pods --all-namespaces -o wide
echo -e " \033[32;5mHappy Kubing! Access Nginx at EXTERNAL-IP above\033[0m"
kubectl get nodes -o wide || true
step "Done · cluster ready"
ready_nodes="$(kubectl get nodes --no-headers 2>/dev/null | awk '$2=="Ready"{c++} END{print c+0}')" || ready_nodes=0
ok "$ready_nodes nodes Ready · context '$context'"
ok "API server: https://$vip:6443"
if [[ -n "$lb_ip" ]]; then
ok "nginx LoadBalancer IP: $lb_ip (curl http://$lb_ip)"
else
warn "nginx LoadBalancer IP not assigned yet — check: kubectl get svc -A"
fi
info "Next: kubectl get pods -A"
+20 -24
View File
@@ -1,10 +1,9 @@
apiVersion: apps/v1
kind: DaemonSet
metadata:
creationTimestamp: null
labels:
app.kubernetes.io/name: kube-vip-ds
app.kubernetes.io/version: v0.8.2
app.kubernetes.io/version: v1.2.1
name: kube-vip-ds
namespace: kube-system
spec:
@@ -13,10 +12,9 @@ spec:
app.kubernetes.io/name: kube-vip-ds
template:
metadata:
creationTimestamp: null
labels:
app.kubernetes.io/name: kube-vip-ds
app.kubernetes.io/version: v0.8.2
app.kubernetes.io/version: v1.2.1
spec:
affinity:
nodeAffinity:
@@ -36,36 +34,38 @@ spec:
value: "true"
- name: port
value: "6443"
- name: vip_nodename
valueFrom:
fieldRef:
fieldPath: spec.nodeName
- name: vip_interface
value: $interface
- name: vip_cidr
value: REPLACE_INTERFACE
- name: vip_subnet
value: "32"
- name: dns_mode
value: first
- name: dhcp_mode
value: ipv4
- name: cp_enable
value: "true"
- name: cp_namespace
value: kube-system
- name: vip_ddns
value: "false"
- name: svc_enable
value: "false"
- name: svc_leasename
value: plndr-svcs-lock
- name: vip_leaderelection
value: "true"
- name: vip_leasename
value: plndr-cp-lock
- name: vip_leaseduration
value: "5"
value: "15"
- name: vip_renewdeadline
value: "3"
value: "10"
- name: vip_retryperiod
value: "1"
value: "2"
- name: address
value: $vip
value: REPLACE_VIP
- name: prometheus_server
value: :2112
image: ghcr.io/kube-vip/kube-vip:v0.8.2
imagePullPolicy: Always
image: ghcr.io/kube-vip/kube-vip:v1.2.1
imagePullPolicy: IfNotPresent
name: kube-vip
resources: {}
securityContext:
@@ -73,6 +73,8 @@ spec:
add:
- NET_ADMIN
- NET_RAW
drop:
- ALL
hostNetwork: true
serviceAccountName: kube-vip
tolerations:
@@ -81,9 +83,3 @@ spec:
- effect: NoExecute
operator: Exists
updateStrategy: {}
status:
currentNumberScheduled: 0
desiredNumberScheduled: 0
numberMisscheduled: 0
numberReady: 0
+2 -2
View File
@@ -1,8 +1,8 @@
apiVersion: metallb.io/v1beta1
kind: L2Advertisement
metadata:
name: example
name: l2-first-pool
namespace: metallb-system
spec:
ipAddressPools:
- first-pool
- first-pool
-5
View File
@@ -1,5 +0,0 @@
1. Snapshot your VMs!
1. Add certificates to your home directory
1. copy script to home directory and make executable
1. Run the script, grab a coffee and enjoy :) (hopefully!)