← blog · July 21, 2026

Kubernetes from scratch on Talos Linux: operating a cluster without a shell

Talos Linux is an immutable, API managed Kubernetes OS with no shell and no SSH on the node. How machine config works, the bootstrap and upgrade flow, debugging without a shell, how the operational load compares to kubeadm, and when not to pick it.

The operating system underneath a Kubernetes cluster is the layer most teams talk about least and spend the most effort on. Installing a general purpose distribution and bringing up a cluster with kubeadm is easy on day one. By year two you have fifteen nodes that have quietly drifted apart, a package list nobody can account for, and a seven step upgrade runbook. Talos Linux proposes closing that layer entirely: no shell on the node, no SSH, no package manager, no files to edit by hand. One interface, and it is an API.

What an immutable, API managed OS actually means

The Talos root filesystem is read only and there is no general purpose userland on the node. The system carries only what is needed to be a Kubernetes node: kubelet, containerd, etcd on control plane nodes, and Talos's own supervisor that manages all of it. No systemd, no bash, no apt or dnf.

The day to day consequence is that "log in and have a look" stops being an available action. A configuration change is a document you send to an API, and debugging is a stream you read back from that API. The Talos API listens on port 50000 by default and requires mutual TLS; without a client certificate you cannot talk to the node at all.

What you give up is real: tcpdump on a node, strace on a stuck process, editing a file and restarting a service in thirty seconds. What you get is equally real: configuration drift disappears as a category of problem, because the mechanism that produces it is gone. A node is either in the state the document describes or it is broken. There is no "someone fixed that by hand last year" in between, because fixing by hand is not possible.

Machine config: one document, two sections

A node's entire personality lives in a single YAML document with two top level sections. machine holds everything node specific: networking, install disk, kubelet settings, certificates, kernel modules to load. cluster holds what belongs to the cluster: API server settings, etcd, networking, shared secrets.

talosctl gen config produces these. It writes three files: a control plane config, a worker config, and talosconfig, which carries your client identity. Losing the third one means losing access to the cluster, so it belongs in a secrets store from the first day, not in the repository.

version: v1alpha1
machine:
  type: controlplane
  install:
    disk: /dev/sda
    image: ghcr.io/siderolabs/installer:v1.9.0
  network:
    hostname: cp-1
    interfaces:
      - interface: eth0
        addresses:
          - 198.51.100.11/24
        routes:
          - network: 0.0.0.0/0
            gateway: 198.51.100.1
cluster:
  network:
    cni:
      name: none

Keep per node differences as patches rather than as copies of the whole document. You can apply a patch offline and inspect the result before anything touches a node:

talosctl machineconfig patch controlplane.yaml \
  --patch @patches/cp-2.yaml -o cp-2.yaml

Or do it in one step:

talosctl apply-config --nodes 198.51.100.12 \
  --file controlplane.yaml --config-patch @patches/cp-2.yaml

That distinction creates a useful habit: a change is no longer "I edited a file on the server," it is a patch that lives in version control and gets reviewed.

Bringing up a cluster from nothing

The flow is short. Boot the nodes from a Talos image; an unconfigured node comes up in maintenance mode and serves its API without requiring a certificate, which is why the first apply uses --insecure.

talosctl gen config demo https://198.51.100.11:6443 \
  --install-disk /dev/sda --output-dir ./out

talosctl apply-config --insecure --nodes 198.51.100.11 \
  --file ./out/controlplane.yaml
talosctl apply-config --insecure --nodes 198.51.100.21 \
  --file ./out/worker.yaml

export TALOSCONFIG=./out/talosconfig
talosctl config endpoint 198.51.100.11
talosctl config node 198.51.100.11

talosctl bootstrap
talosctl kubeconfig .
talosctl health

bootstrap runs exactly once, on exactly one control plane node. It starts etcd there, and the remaining control plane nodes join on their own. Running it on a second node produces two separate etcd clusters, and untangling that takes longer than rebuilding.

Upgrades

An OS upgrade is an API call. You hand the node the address of a new installer image; it pulls that image, writes it to the alternate partition and reboots:

talosctl upgrade --nodes 198.51.100.11 \
  --image ghcr.io/siderolabs/installer:v1.9.0

Because the partition scheme is A/B, the previous kernel and image stay in place. If the new version fails to boot, the node falls back on its own. talosctl rollback exists for doing it deliberately, but it only goes one step back: if you skipped two versions, getting to the older one means naming it explicitly in an upgrade.

The Kubernetes version is upgraded separately:

talosctl --nodes 198.51.100.11 upgrade-k8s --to 1.31.4 --dry-run

This walks the control plane components and the kubelets in the right order, waiting for health at each step, and --dry-run shows the plan first. The operational value of the split: an OS security patch and a Kubernetes minor bump no longer share a maintenance window. They become independent decisions.

Debugging without a shell

This is the part that takes longest to internalize. Your reflexes get replaced:

talosctl -n 198.51.100.11 logs -f kubelet
talosctl -n 198.51.100.11 dmesg
talosctl -n 198.51.100.11 get members
talosctl -n 198.51.100.11 get staticpods
talosctl -n 198.51.100.11 dashboard
talosctl -n 198.51.100.11 support -O support.zip

get queries Talos's internal resource model: addresses, links, service states, rendered static pod definitions. This is usually where you find out why a node is in the state it is in, because those resources are derived from configuration and you can read the derivation backwards. When you genuinely need packet or kernel level visibility, the answer is a privileged debug pod on the host network. It works, but it is slower than SSH, so keep that manifest ready before you need it.

How the operational load compares to kubeadm

  • OS patching: a classic setup coordinates security updates, reboots and drains separately. Talos gives you one API call, an atomic image, and automatic fallback.
  • Drift: instead of a configuration management tool attempting convergence, a model where divergence cannot be produced.
  • Access control: no SSH key distribution, sudo rules or session auditing; one client certificate and a role.
  • Attack surface: an attacker escaping a container finds no shell and no download tool to work with.
  • The price: adding a driver or an unusual agent means entering an image build flow. You do not install packages on Talos, you add system extensions and generate an image from a schematic that lists them. Straightforward once established, unfamiliar the first time.

Habits and the learning curve

For a team that already knows Kubernetes, the real curve is not Talos's concepts, it is the missing escape hatch. Instead of connecting to a node to "have a look," you look for the hypothesis in the configuration first. Pleasant side effect: problems become documentation rather than personal knowledge, because nobody carries "that node has a special tweak" around in their head.

The second shift is that authority splits in two. What you do with kubectl and what you do with talosctl run on different certificates and different roles. Granting a developer cluster access no longer implies node access, which is good, but your on call procedures have to account for both.

Traps

  • Picking the wrong install disk. On machines with several disks, /dev/sda is not stable; matching by size or serial is safer.
  • Forgetting extensions. System extensions are part of the image. If you upgrade to a new version without generating its image from the same schematic, your drivers vanish with the upgrade.
  • Committing talosconfig and the generated secrets. Those files are authoritative over the whole cluster.
  • Exposing the Talos API port to the public internet. Mutual TLS helps, but that port belongs on a management network.
  • CNI assumptions. The default config ships its own network plugin; if you plan to run a different one, disable it from the start rather than swapping later.
  • Backup scope. Your application backup tool captures Kubernetes objects, but what restores a control plane is an etcd snapshot, and that has to be taken separately and regularly.

When not to choose Talos

Do not choose it if anything other than Kubernetes has to run on the machine. If you need a database, a backup agent or a legacy service alongside, Talos is designed to make that hard.

Do not choose it if a compliance requirement forces an agent onto the node that cannot run as a container. Packaging it as an extension may be possible, but vendor support probably will not cover it.

Do not choose it if the team is still learning Kubernetes. Two unknowns at once makes it hard to tell which layer broke. Running a cluster on a familiar distribution first and migrating later is cheaper.

For a single node development environment, lighter options produce less friction. And if a managed Kubernetes service covers your needs, never owning the OS layer is still the least work; Talos pays off precisely where that layer is yours.

The cheapest way to decide: build a cluster on three virtual machines, upgrade one node and roll it back deliberately, then diagnose a real problem without a shell. Half a day tells you how the flow feels and how comfortable your team is with the model.