A constitutional Ansible repository for managing the complete lifecycle of k3s Kubernetes clusters: provisioning, configuration updates, scaling, and upgrades.
✨ Core Capabilities
- Provision highly-available k3s clusters with embedded etcd (3-node control-plane)
- Single-node development clusters for testing
- Control-plane VIP via kube-vip for HA API access
- LoadBalancer service support via kube-vip cloud controller
- Idempotent playbooks for safe re-runs and configuration updates
🎯 Platform Add-ons (Optional, modular deployment)
- cert-manager: Provider-agnostic DNS-01 certificate issuers (Cloudflare, Route53, etc.)
- multus CNI: VLAN-based secondary pod networking
- Rancher: Web-based cluster management and monitoring
- rancher-monitoring: Prometheus + Grafana observability stack
- Traefik: Ingress controller with LoadBalancer integration
- Synology CSI: Persistent storage from Synology NAS (optional)
🔧 Operational Support
- Node scaling: Add/remove control-plane and worker nodes
- Minor/patch upgrades: Rolling k3s version upgrades (major upgrades out-of-scope)
- Host prerequisite validation: Fail-fast checks for OS, CPU, memory, ports, network
- Smoke tests and ansible-lint validation
- Control node: Ansible Core 2.15+ with Python 3.8+
- Target hosts: Debian/Ubuntu Linux (systemd, x86_64/arm64) with SSH access
- Minimum resources:
- Control-plane: 2 CPU cores, 2GB RAM, 20GB disk
- Workers: 1 CPU core, 1GB RAM, 20GB disk
- Network: Required ports open between nodes (see docs/ansible-structure.md)
git clone <repository-url>
cd ansible-k3s-clusterCopy an example inventory and customize for your environment:
cp -r ansible/inventories/examples/ha-cluster ansible/inventories/production
vi ansible/inventories/production/hosts.iniEdit cluster configuration in ansible/group_vars/all.yml:
cluster_name: "my-k3s-cluster"
k3s_version: "v1.28.5+k3s1"
control_plane_vip: "192.168.1.100"
kube_vip_interface: "eth0"
# Enable desired add-ons
cert_manager_enabled: true
rancher_enabled: true
traefik_enabled: true
multus_enabled: trueVersion management policy:
- Treat
ansible/group_vars/all.ymlas the canonical source for managed component versions. - Avoid embedding new hard-coded versions in role defaults or playbook commands.
- Use inventory-specific overrides (for example
ansible/inventories/test-cluster/group_vars/all.yml) when environment-specific version differences are required.
For multus VLAN networking, define secondary networks:
# VLAN networks with DHCP-based IP assignment (default)
multus_vlan_networks:
- name: iot-vlan
# Must match an interface that exists on every target node.
# Using ens18 (same NIC as kube_vip_interface) avoids "Link not found" from macvlan.
interface: ens18
# Leave vlan_id unset unless the host already has a matching VLAN subinterface (e.g. ens18.10).
ipam_type: dhcp # dhcp (default) | host-local | static
- name: iot-vlan-2
interface: eth0
vlan_id: 50
ipam_type: dhcp
- name: storage-vlan
interface: eth0
vlan_id: 100
ipam_type: host-local
cidr: 10.10.100.0/24
gateway: 10.10.100.1See the Multus role documentation for full configuration details.
# Recommended: Unified playbook for install and upgrades (same command)
ansible-playbook -i ansible/inventories/production ansible/playbooks/site.ymlAlternatively, use individual playbooks for granular control:
# Deploy k3s core cluster (control-plane + workers + kube-vip)
ansible-playbook -i ansible/inventories/production ansible/playbooks/cluster-core.yml
# Deploy optional platform add-ons
ansible-playbook -i ansible/inventories/production ansible/playbooks/cluster-addons.yml# SSH to first control-plane node
ssh admin@k3s-server-01
# Check cluster health
export KUBECONFIG=/etc/rancher/k3s/k3s.yaml
kubectl get nodes
kubectl get pods -Aansible/
├── inventories/ # Inventory files
│ ├── examples/ # Example HA and single-node inventories
│ └── production/ # Your production inventory
├── group_vars/ # Cluster configuration
│ ├── all.yml # Cluster-wide settings
│ ├── k3s_servers.yml # Control-plane config
│ └── k3s_agents.yml # Worker config
├── roles/ # Ansible roles
│ ├── k3s-common/ # Prerequisites and validation
│ ├── k3s-server/ # Control-plane installation
│ ├── k3s-agent/ # Worker node installation
│ ├── kube-vip/ # VIP and LoadBalancer
│ ├── cert-manager/ # Certificate management
│ ├── multus/ # Secondary networking
│ ├── rancher/ # Cluster management UI
│ ├── rancher-monitoring/ # Observability
│ ├── traefik/ # Ingress controller
│ └── synology-csi/ # Synology persistent storage
└── playbooks/ # Playbook entrypoints
├── site.yml # Unified install/upgrade orchestrator
├── cluster-core.yml # Provision/update core cluster
├── cluster-addons.yml # Deploy platform add-ons
├── scale-nodes.yml # Add/remove nodes
├── upgrade-k3s.yml # (deprecated) Minor/patch k3s upgrades
└── includes/ # Modular upgrade logic
├── detect-versions.yml
├── compute-plan.yml
├── upgrade-k3s-rolling.yml
├── upgrade-rancher.yml
├── upgrade-kube-vip.yml
└── upgrade-addon.yml
ssh-agent bash
ssh-add <path to your private key here>Example:
Note
The project level .ssh is an ignored folder.
ssh-agent bash
ssh-add ./.ssh/Private\ SSH\ Key\ -\ Neoteric\ -\ wade.opensshansible-playbook -i ansible/inventories/production ansible/playbooks/cluster-core.ymlansible-playbook -i ansible/inventories/test-cluster ansible/playbooks/cluster-core.ymlansible-playbook -i ansible/inventories/production ansible/playbooks/cluster-addons.ymlansible-playbook -i ansible/inventories/test-cluster ansible/playbooks/cluster-addons.ymlModify variables in group_vars/ and re-run playbooks to apply changes:
# Update core cluster settings
ansible-playbook -i ansible/inventories/production ansible/playbooks/cluster-core.yml
# Update add-on configuration
ansible-playbook -i ansible/inventories/production ansible/playbooks/cluster-addons.ymlAdd new hosts to inventory, then:
ansible-playbook -i ansible/inventories/production ansible/playbooks/scale-nodes.ymlUpdate k3s_version in group_vars/all.yml, then:
# Recommended: Unified workflow (handles dependency ordering, cordon/drain, constraint validation)
ansible-playbook -i ansible/inventories/production ansible/playbooks/site.ymlTo upgrade Rancher and k3s together (Rancher is upgraded first automatically):
# Update rancher_version and k3s_version in group_vars/all.yml, then:
ansible-playbook -i ansible/inventories/production ansible/playbooks/site.ymlTo allow a version downgrade:
ansible-playbook -i ansible/inventories/production ansible/playbooks/site.yml \
-e "allow_downgrade=true"Legacy upgrade playbook (deprecated, still functional):
ansible-playbook -i ansible/inventories/production ansible/playbooks/upgrade-k3s.yml- Ansible Structure Guide: Directory layout, supported platforms, host prerequisites
- Quickstart Guide: Step-by-step provisioning and usage examples
- Feature Specification: Complete functional requirements
- Implementation Plan: Technical architecture and decisions
- Multus CNI Role: VLAN networking, DHCP IPAM, and NetworkAttachmentDefinition configuration
- Constitution: Project governance and design principles
ansible-lint ansible/playbooks/cluster-core.yml
ansible-lint ansible/playbooks/cluster-addons.ymlansible-playbook -i ansible/inventories/production ansible/playbooks/cluster-core.yml --checkansible-playbook -i tests/ansible/inventories/local tests/ansible/smoke/smoke.ymlansible-playbook -i tests/ansible/inventories/local tests/ansible/smoke/multus-dhcp-test.yml
ansible-playbook -i tests/ansible/inventories/local tests/ansible/smoke/scale-test.yml
ansible-playbook -i tests/ansible/inventories/local tests/ansible/smoke/upgrade-test.yml- Minimal Core: Separate core k3s provisioning from optional platform add-ons
- Idempotent: Safe to re-run playbooks without side effects
- k3s-Specific: Leverage k3s embedded etcd, no kubeadm assumptions
- Variable-Driven: All configuration via inventory and group_vars, no hardcoded values
- Secure Defaults: No plain-text secrets, Ansible Vault recommended
- Control-plane: 3-node embedded etcd cluster (odd number required)
- VIP Access: kube-vip provides floating IP for API server access
- LoadBalancer: kube-vip cloud controller allocates external IPs for LoadBalancer services
- CNI: Flannel VXLAN (k3s default) + optional multus for secondary networks
- Conditional Deployment: Enable/disable via
*_enabledflags ingroup_vars/all.yml - Helm-Based: Rancher, rancher-monitoring use Helm charts
- kubectl-Based: cert-manager, multus, kube-vip use manifests
- Provider-Agnostic: cert-manager DNS-01 supports multiple providers via credentials
- Control-plane nodes: 1-3 (odd number for HA)
- Worker nodes: Up to ~10 nodes
- Cluster size: Small to medium deployments
- Large-scale clusters (dozens/hundreds of nodes)
- Full disaster recovery (complete etcd loss)
- Major version upgrades (e.g., k3s 1.x → 2.x)
- Air-gapped/offline installations
- OS: Debian 11+, Ubuntu 20.04+
- Architectures: x86_64, arm64
- Init System: systemd
- Access: SSH with sudo privileges
- Check prerequisites:
ansible-playbook -i ansible/inventories/production ansible/playbooks/cluster-core.yml --tags prerequisites - Verify SSH access:
ansible -i ansible/inventories/production all -m ping - Check control-plane VIP:
ping <control_plane_vip>
- Verify kube-vip pods:
kubectl get pods -n kube-system -l app.kubernetes.io/name=kube-vip - Check network interface: Ensure
kube_vip_interfacematches your host's network interface - Verify ARP:
ip addr showon control-plane nodes should show VIP
- Check enablement flags in
group_vars/all.yml - Verify cluster is operational:
kubectl get nodes - Check pod status:
kubectl get pods -A
This project has been configured to use GitHub Spec Kit. The project includes a dev container with Spec Kit installed for this purpose so you can avoid installing any tooling locally on your machine.
This project follows constitutional governance. See .specify/memory/constitution.md for design principles and contribution guidelines.
[Specify your license here]