Post

Building My Two-Site K3s Homelab — GitOps, Tailscale & Monitoring

Building My Two-Site K3s Homelab — GitOps, Tailscale & Monitoring

Building My Two-Site K3s Homelab — GitOps, Tailscale & Monitoring

Over the last few days I’ve been building a small Kubernetes homelab.

The idea was simple: I wanted something small enough that I can understand every part of it, but still realistic enough to learn things I can actually use elsewhere.

I ended up with two separate K3s clusters:

  • one at home on a Radxa X4
  • one on an ARM cloud VM

They are connected with Tailscale, managed through Flux GitOps, and monitored with Zabbix and Grafana.

The lab is still very much a work in progress, but the core setup is finally stable enough that I thought it was worth writing down.


Current Architecture

The setup currently looks roughly like this:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
                        GitHub
                           |
                           | Flux GitOps
                +----------+----------+
                |                     |
                v                     v
             Home site             Cloud site
             K3s cluster           K3s cluster
                |                     |
                +------ Tailscale ----+
                                      |
                        +-------------+-------------+
                        |                           |
                     Zabbix                      Grafana
                        |
                 Cross-site monitoring

                ^
                |
        CachyOS workstation

CachyOS is my main daily-driver machine, not just a management box. I use it for normal desktop work, development, this lab and even gaming.

Windows is still around, but only on a small 256 GB SSD and mostly for the few things I still prefer to do there, mainly Office and Adobe Lightroom.

So naturally, CachyOS also became the place where I manage the lab.

From there I use:

  • kubectl
  • Helm
  • Flux
  • SOPS
  • age
  • Git
  • SSH
  • Tailscale

Home Cluster

The home cluster runs on a Radxa X4.

Current hardware:

  • Intel N100
  • 12 GB RAM
  • 512 GB NVMe SSD
  • Debian 13
  • Wi-Fi only
  • single-node K3s control plane and etcd

A quick check from my workstation:

1
kubectl --context home-cluster get nodes

returns a healthy single-node control plane:

1
2
NAME       STATUS   ROLES                VERSION
home-node  Ready    control-plane,etcd   v1.36.4+k3s1

For now this cluster is mostly used for:

  • Kubernetes learning
  • GitOps testing
  • future home services
  • general lab work

Cloud Management Cluster

The second cluster runs on an ARM cloud VM.

Current setup:

  • ARM / Ampere
  • 4 OCPU
  • 24 GB RAM
  • Ubuntu Server
  • separate persistent data volume
  • K3s
  • Tailscale

This cluster is completely separate from the Radxa cluster.

1
kubectl --context cloud-cluster get nodes

Example:

1
2
NAME        STATUS   ROLES                VERSION
cloud-node  Ready    control-plane,etcd   v1.36.4+k3s1

Why Two Clusters?

At first I did think about joining both machines into one Kubernetes cluster over Tailscale.

It can be done, but in my case it did not really make sense.

The control plane, and especially etcd, does not gain anything from sitting across a WAN link. It just adds latency and gives me one more thing to break.

So I kept the two clusters independent:

1
2
3
4
5
Home cluster
  └── independent etcd

Cloud cluster
  └── independent etcd

If the home internet disappears, the Radxa can still keep doing its thing locally.

If the Radxa dies, the cloud cluster keeps running.

That separation also makes the cloud side useful for monitoring and, later, backups.


GitOps with Flux

Both clusters are managed with Flux.

The repository is split so each cluster only reconciles its own path:

1
2
3
4
5
6
7
8
9
10
11
clusters/
├── home-cluster/
│   ├── flux-system/
│   └── infrastructure/
│
└── cloud-cluster/
    ├── flux-system/
    ├── infrastructure/
    └── monitoring/
        ├── grafana/
        └── zabbix/

The basic flow is:

1
2
3
4
5
6
7
8
9
10
Git change
   |
   v
GitHub
   |
   v
Flux
   |
   v
Kubernetes

So if I change a manifest and push it, Flux picks it up and applies it to the right cluster.

I also added GitHub Actions validation so obvious YAML and manifest problems are caught before they get anywhere near the clusters.


Secrets with SOPS + age

Plaintext secrets in Git were a no-go from the start.

For this I use:

  • SOPS
  • age

Each cluster has its own age key pair.

Encrypted files can live safely in the repository while the private keys stay outside Git.

Examples include:

  • DNS provider API tokens
  • Grafana admin credentials
  • PostgreSQL credentials

Typical encrypted files look like:

1
2
3
dns-api-token.sops.yaml
admin-secret.sops.yaml
postgres-access.sops.yaml

Flux decrypts them inside the cluster during reconciliation.


Tailscale for Private Access

Tailscale is doing a lot of the boring networking work for me.

It connects:

1
2
3
management workstation
home node
cloud node

That gives me private access between all three without exposing SSH directly to the internet.

The same Tailscale network is also used for monitoring traffic between home and cloud.


DNS and TLS

The lab uses a domain managed through Cloudflare DNS.

Kubernetes services sit under a dedicated lab subdomain, with certificates handled automatically by:

  • cert-manager
  • Let’s Encrypt
  • Cloudflare DNS-01

Traefik handles ingress inside Kubernetes.

So I get proper HTTPS endpoints without having to manually deal with certificate renewals.


Monitoring with Zabbix and Grafana

Monitoring runs in the cloud cluster, and that was very deliberate.

If Zabbix and Grafana only lived on the Radxa, they would vanish at exactly the same time as the thing I was trying to monitor.

Instead:

1
2
3
4
5
6
7
8
home node
   |
   | Tailscale
   v
Zabbix
   |
   v
Grafana

Zabbix currently monitors both hosts.

Collected data includes:

  • CPU usage
  • memory
  • filesystem usage
  • network traffic
  • system uptime
  • agent health

Grafana then uses Zabbix as a data source so I can see both machines from one dashboard.


Swapping the Radxa SSD

One of the more interesting side jobs was replacing the Radxa SSD.

Originally I used a 1 TB Crucial NVMe.

That was complete overkill.

The Debian + K3s setup was barely using any of that space, so I decided to move it to a 512 GB NVMe and put the 1 TB drive to better use in my main PC.

Because the destination disk was smaller than the source, I did not want to do a raw block clone and hope for the best.

Instead I recreated the disk layout manually:

1
2
3
4
5
6
EFI
/boot
LVM
├── root
├── k3s
└── backup

Then I copied the filesystems with rsync.

The actual data migration was the easy part.

The fun started on the first boot. 😄

The Radxa dropped straight into:

1
grub>

GRUB itself was there, but the EFI-side GRUB config was still pointing at the old /boot UUID.

Manually pointing GRUB at the correct partition was enough to get Debian booting again:

1
2
3
4
set root=(hd0,gpt2)
set prefix=(hd0,gpt2)/grub
insmod normal
normal

Once the system was up, I corrected the EFI GRUB config so it searched for the new /boot UUID.

After that, normal boot was back.

K3s came up, the node returned to Ready, and all the workloads were still there.

The old 1 TB SSD was then wiped and reformatted as Btrfs for reuse.

That little detour probably deserves its own post at some point because migrating Linux + LVM + K3s to a smaller disk turned out to be more interesting than expected.


Managing It All from CachyOS

This part was less of a “move away from Windows” and more of a “put everything where I actually spend my time.”

CachyOS is my main OS and daily driver. I use it for almost everything, including gaming.

Windows lives on a small 256 GB SSD and is mostly there for Office and Adobe Lightroom.

So once the lab started growing, keeping the management tooling on Windows just felt backwards.

I moved over:

  • kubeconfig
  • SSH keys
  • SOPS age keys
  • Git configuration
  • the homelab repository

Then installed the Linux-native tooling I actually need:

1
2
3
4
5
6
kubectl
helm
flux
sops
age
tailscale

Both clusters are now reachable from the same machine:

1
2
kubectl --context home-cluster get nodes
kubectl --context cloud-cluster get nodes

GitHub access also moved to SSH.

It is a much nicer setup now because the machine I use all day is also the machine I use to manage the lab.


Things That Went Wrong

This lab has already produced a few useful lessons.

Windows line endings

When I copied the Git repository from Windows to CachyOS, Git initially showed thousands of changed lines.

Nothing had actually changed.

It was just CRLF vs LF.

Running:

1
git diff --ignore-space-at-eol --stat

showed there were no real content changes.

SOPS validation

My first GitHub Actions validation also tried to validate encrypted SOPS manifests as normal Kubernetes YAML.

That obviously did not work.

The workflow now skips *.sops.yaml files during kubeconform validation and checks them separately.

GRUB after the SSD migration

The migrated filesystem was fine, but GRUB was still looking for the old /boot UUID.

That was enough to drop the machine to a GRUB prompt even though all the data was sitting there perfectly intact.

A good reminder that “the copy finished successfully” and “the machine will boot” are two very different things. 😄


What’s Next?

The lab is working, but it is nowhere near finished.

Next on the list:

  1. Zabbix alerts and notifications
  2. Kubernetes-level monitoring
  3. automated home → cloud backups
  4. restore testing
  5. cloud object-storage backup tier
  6. more real workloads

Possible future services:

  • Uptime Kuma
  • Vaultwarden
  • AdGuard Home
  • Loki
  • ntfy
  • changedetection.io
  • Home Assistant
  • Paperless-ngx
  • Mealie

I’m trying not to install everything just because I can.

The point is to add things when they either teach me something useful or solve an actual problem.


Final Thoughts

The most useful part of this project so far has not really been getting Kubernetes running.

That part is the easy bit.

The interesting part is everything around it:

  • managing two independent environments
  • keeping secrets out of Git
  • monitoring one site from another
  • recovering when a bootloader points at the wrong disk
  • keeping the whole thing reproducible

The lab is still growing, so I’ll probably split some of the more detailed pieces into their own posts later.

For now, the foundation is there.

And more importantly:

1
kubectl get nodes

is green. 😄

This post is licensed under CC BY 4.0 by the author.