Skip to content

Systems

My homelab doesn’t need Kubernetes

Kubernetes solves real problems. They just aren't the ones I have, and my lab is better off boring.

There is a moment in a lot of homelabs when Kubernetes starts to look like the obvious next level. The services multiply, the tooling is everywhere, and running a cluster feels like what a serious setup does. I tried it, and I didn’t stay. That isn’t because I think Kubernetes is bad. Next to the alternatives I already had it felt complex, and I couldn’t name the problem it was solving for me.

Kubernetes would work

I want to get this out of the way. Kubernetes works, and I don’t doubt it would run what I run. It solves real problems: scheduling workloads across interchangeable machines, rolling updates, self-healing, a declarative description of a whole platform. People who need those things need them badly. The argument here isn’t “it’s overkill, therefore ignore it.” It’s narrower: those are real problems, and they are not currently mine.

But what problem would it solve for me?

So I tried asking the question properly: what do I need this infrastructure to do?

The mini PC is the centre of the lab: Docker Compose stacks, the Agent, storage, monitoring and local models. The external server hosts the web panel, and ESP32 devices publish over MQTT. Each has a specific job. The map below is the short version; the Afterlab project page has the service-by-service detail.

Afterlab responsibilities: a mini PC hosts Docker Compose services, the Agent, storage, monitoring and local models; ESP32 devices communicate over MQTT; a separate server hosts the web panel with Cloudflare and Access.
The machines have different responsibilities. This map explains the current infrastructure choice; it does not describe a pool of interchangeable compute nodes.

Put next to each other, what a cluster offers and what I have don’t line up much:

What Kubernetes is good atMy situation
Scheduling workloads across interchangeable nodesOne mini PC at the centre and one external server, each with its own job. GPU work depends on which machine owns the GPU.
Surviving the loss of a nodeIf something breaks I fix it myself; there is no second node waiting to take over.
A declarative description of a whole platformA handful of Compose stacks I can read in one sitting.
Many services and many usersOne person, a few services and a handful of devices.

Containers already give me most of the isolation and packaging I need. The parts of Kubernetes I would benefit from most—automatic placement across a pool and failover to another node—depend on spare, interchangeable capacity. That is not how this lab is organised.

What I’m optimising for instead

If I write down what I want from the lab, the list looks different from a production platform’s: understandability, recoverability, low operational overhead, room to experiment, visibility, and how easily I can change my mind. Those are properties of a personal lab. In some ways they are the opposite of what makes a platform impressive.

Recoverability is the one I care about most, and for me it doesn’t mean the system heals itself. It means I can reason about how to get things back from a short list of components I understand, without first needing a control plane to be healthy. Experimentation is the other one. I add devices, models and services constantly, so every change should be cheap to try and cheap to undo. A setup that makes me write a manifest before I can test an idea works against the whole reason the lab exists.

My bottleneck isn’t orchestration

When I look at what actually slows me down, orchestration isn’t on the list. It’s hardware, and specifically which GPU is in which machine. The training run in AI.001 happened on a laptop. Local speech recognition ran on a different PC. Those are not interchangeable nodes. A scheduler can’t move a workload onto a GPU that isn’t there, and one of those machines is a laptop.

The rest of the list is debugging, device integration, and understanding what state the whole thing is in. That last one is why I built a small Mission Control. A cluster would give me a more elaborate way of describing state, and it wouldn’t help me notice that a sensor had gone quiet. Kubernetes would not magically fix any of this.

Complexity has a maintenance cost

Every abstraction is another thing that can break, and another thing I need to understand when it does. In a company, someone owns that. In my lab it’s me, usually while I’m trying to debug something else.

There is real value in being able to follow the whole path from the host to the container to the service with one hand on the keyboard. When something misbehaves, I want the list of suspects to be short. Adding a control plane, an ingress layer, cluster networking and a storage abstraction lengthens that list before I’ve done anything useful with it.

I apply the same rule inside Afterlab. The Agent that sits between the infrastructure and everything else owns the writes, and it doesn’t run arbitrary shell commands or change Docker. It’s deliberately a small surface. Choosing a platform that adds a very large surface would pull in the opposite direction.

Boring infrastructure is underrated

Infrastructure doesn’t need to be intellectually impressive. It needs to make experiments possible. Docker Compose stacks, Portainer and Coolify are more than enough for what I run, and they have a property I value a lot: I rarely have to think about them.

The interesting problems in this lab are elsewhere: how Pixel decides which sources to read, how a device recovers when an update fails, what a model does when you give it only 4 GB. I would rather spend my limited attention there. Boring infrastructure is what lets the interesting parts be the only parts that break.

When I would reconsider

I don’t want this to turn into dogma, so here is what would change my mind.

  • Several interchangeable compute nodes, where I actually want workloads placed automatically.
  • Availability requirements stronger than “I’ll fix it tonight”.
  • A much larger number of services or users than one person and a handful of devices.
  • Wanting Kubernetes itself as the object of study. That one is legitimate, and I’d treat it as a project of its own: an experiment with a hypothesis and an end date, not a platform I quietly depend on.

None of those is true at the moment. If one becomes true, the decision changes, and I’d be glad to have a reason.

Choosing infrastructure by problems, not prestige

It’s easy to treat infrastructure as a ladder, with Compose stacks at the bottom and a cluster at the top. I’d rather treat it as a fit. The best stack is the one that disappears enough for me to work on the thing I actually care about. For now that’s a few Docker hosts, some boring tools, and a clear picture of what is running where.

The rest of the system is described on the Afterlab project page.

Keep exploring

The questions behind these notes

Each note starts from something I tested. The experiments hold the method, the evidence and what did not work.