← blog · August 8, 2026

Reproducible service deployment with NixOS

What declarative configuration, atomic switching and flake pinning actually buy you, where NixOS beats Docker and conventional configuration management, which traps show up in the first month, and when it is the wrong choice.

Same config, three machines, three different results

Installing one service on three servers and getting identical behaviour on all three is harder than most teams expect. The provisioning script is the same and the package names are the same, but one machine has a stale repository index, someone hand-edited a config file on another six months ago, and dependency resolution pulled a different minor version on the third. The real problem is not that any single step is wrong; it is that the steps describe a history rather than a result. A conventional configuration management task says "install this package, add this line" and knows nothing about what the machine looked like before that line existed. NixOS starts from the other end: you write what the system will be, not what should be done to it.

Declarative is not the same as idempotent

Most configuration management tools are convergent. They read the current state of a machine and fill in whatever is missing relative to the target. That only works in one direction: deleting a task from the playbook does not remove the package it installed months ago. Over time you end up with a server full of software that has no counterpart in any source file, and nobody dares remove any of it because nobody knows what depends on what.

In NixOS the configuration is not a script but an expression. Evaluating it produces a complete description of the system: which packages at which versions, which systemd units, which users, which kernel parameters. Delete a line and everything that line produced is absent from the next generation.

The store is what makes this possible. Every package lives at a path derived from a hash of its full dependency graph, something like /nix/store/<hash>-nginx-1.28.0. Two versions of the same library can sit side by side, so the question "will upgrading this break that service" stops being interesting. There is no global /usr/lib, therefore no global conflict.

Reading configuration.nix

On a single-machine install everything starts in one file:

{ config, pkgs, ... }:
{
  imports = [ ./hardware-configuration.nix ];

  networking.hostName = "app01";
  time.timeZone = "UTC";

  services.openssh = {
    enable = true;
    settings.PasswordAuthentication = false;
  };

  services.postgresql = {
    enable = true;
    package = pkgs.postgresql_16;
    ensureDatabases = [ "appdb" ];
  };

  networking.firewall.allowedTCPPorts = [ 80 443 ];

  system.stateVersion = "26.05";
}

What you are doing here is assigning options to a module system. services.postgresql.enable = true does not trigger an install command; it tells the PostgreSQL module in nixpkgs to generate its systemd unit, service user, data directory and initialisation script. You can read the module and see exactly which option produces what.

A service and the proxy in front of it

You do not need an upstream module to run your own application. Describe the unit directly:

{ config, pkgs, ... }:
{
  systemd.services.myapp = {
    description = "Application server";
    wantedBy = [ "multi-user.target" ];
    after = [ "network.target" "postgresql.service" ];
    serviceConfig = {
      ExecStart = "${pkgs.myapp}/bin/myapp --listen 127.0.0.1:8080";
      DynamicUser = true;
      Restart = "on-failure";
      ProtectSystem = "strict";
      PrivateTmp = true;
    };
  };

  services.nginx = {
    enable = true;
    recommendedProxySettings = true;
    recommendedTlsSettings = true;
    virtualHosts."app.example.com" = {
      enableACME = true;
      forceSSL = true;
      locations."/" = {
        proxyPass = "http://127.0.0.1:8080";
        proxyWebsockets = true;
      };
    };
  };

  security.acme = {
    acceptTerms = true;
    defaults.email = "[email protected]";
  };
}

enableACME wires up certificate issuance, the renewal timer and the directory permissions shared with nginx. Domain validation needs ports 80 and 443 reachable, so if you forget the firewall line the certificate simply never arrives, and the error will not say so in plain language. One of the most common first-week failures.

Atomic switching and rollback

nixos-rebuild switch builds the entire new system, then flips a symlink and restarts whichever units changed. There is no half-applied state: you are either on the old generation or the new one. Every generation stays in the bootloader menu as its own entry, so even a kernel upgrade is reversible.

Three commands earn their keep. nixos-rebuild test applies the configuration to the running system without touching the boot default, so a reboot returns you to the previous state; use it for anything that could cut your own network path. nixos-rebuild boot does the opposite and takes effect on next boot. nixos-rebuild switch --rollback returns to the previous generation.

To see what an upgrade will actually change before you commit to it:

nixos-rebuild build --flake .#app01
nix store diff-closures /run/current-system ./result

The second command lists packages added, removed and version-bumped between the two systems. Reviewing a Nix diff in a pull request usually tells you very little; this output tells you everything.

One honest caveat: rollback restores the system, not your data. If a schema migration ran, or the application started writing a new on-disk format, going back a generation will not save you. Generations are not backups.

Pinning versions with flakes

With channels, nixos-rebuild switch --upgrade pulls whatever nixpkgs snapshot is current on that machine at that moment, which is the most common way teams break the promise that identical configuration yields identical results. Flakes close the hole by recording the exact commit hash of every input in a flake.lock file that lives in the repository.

{
  description = "Server configurations";

  inputs.nixpkgs.url = "github:NixOS/nixpkgs/nixos-26.05";

  outputs = { self, nixpkgs }: {
    nixosConfigurations.app01 = nixpkgs.lib.nixosSystem {
      system = "x86_64-linux";
      modules = [ ./hosts/app01/configuration.nix ];
    };
  };
}

You deploy with nixos-rebuild switch --flake .#app01. Upgrading becomes a deliberate act: nix flake update nixpkgs refreshes the lock, the change appears as a reviewable commit, and only then does it reach a machine. Build a new server six months later from the same lock and it gets the same package versions as today.

Flakes are still formally an experimental feature, which in daily practice means adding nix.settings.experimental-features = [ "nix-command" "flakes" ]; to your configuration. They have been in production use for years, but that flag is why some commands still ask for it explicitly.

Compared with Docker and with configuration management

Container images are less reproducible than they look. Build a Dockerfile containing FROM debian:12 and apt-get install twice, three months apart, and you get two different images. That rarely hurts, because the tagged image is itself a pinned artefact, but what you pinned is the output, not the recipe. The moment you have to rebuild, for a security patch say, you no longer know what you are getting. Nix pins the recipe. Containers still win on distribution and on having an interface every engineer already knows, and the two are not exclusive either, since Nix can build container images that come out reproducible without depending on layer ordering.

Configuration management tools solve a different problem: running a fleet you did not create and cannot fully replace, often spanning several operating systems. On machines you build yourself they are the source of drift rather than the cure. The rule is simple. If you provision the machines and can change all of them, NixOS wins. If you inherited a heterogeneous estate, it does not.

Traps

Secrets do not belong in the store. The Nix store is world-readable, so a plaintext password written into your configuration ends up at a path any user on the machine can read, regardless of file permissions, and it is committed in version control besides. Use encrypted files decrypted at activation time, with agenix or sops-nix, and set it up on day one; retrofitting means rotating everything.

system.stateVersion is not a version number. It records the NixOS release you first installed on that machine and decides which defaults are used by software that cannot migrate its own data. Bumping it during an upgrade updates nothing, it only changes those data-format assumptions. Leave it alone.

Disks fill up quietly. Every generation keeps its whole closure alive, and after a few unattended months the store reaches tens of gigabytes. nix.gc.automatic with an option like --delete-older-than 30d handles it, but you have to enable it explicitly.

Software outside nixpkgs is expensive. Language-level package managers such as pip, npm and cargo expect to download their own binaries, and because NixOS does not provide a standard filesystem hierarchy those binaries cannot find the dynamic linker. Workarounds exist, and each one is extra work.

Error messages are rough. A typo can produce a stack trace pointing anywhere but your file, capped with something like "infinite recursion encountered". This is the steepest part of the learning curve. The language itself is small and takes a few days; understanding how the module system evaluates takes considerably longer.

What happens once there is a team

For one person, NixOS is a pleasure. On a team, a pattern emerges: one engineer learns it and everyone else asks that engineer. Changes queue up behind their review, and while they are on holiday nobody touches infrastructure. Three things actually help: keep every host configuration in one repository with shared modules, attach diff-closures output to code reviews, and have each new joiner add a real service end to end during their first week. Documentation does not substitute for any of the three.

When not to choose it

There is no payoff on short-lived or one-off machines; if a VM gets built once and thrown away, the learning cost never comes back. The gain is also limited for teams already living on a container orchestrator and barely touching the node OS, because the variability is not in the operating system anymore. If a vendor support matrix mandates a specific distribution, the discussion ends before it starts. And most importantly, do not do it if nobody on the team wants to. A server that is half NixOS and half hand-edited files is worse than a plain distribution: you get the drift, and you also get tooling that lies to you about it.

The decision usually comes down to one thing. If your machine count is growing and rebuilding any one of them turns into an archaeology session, the price is a few weeks of learning. You pay that once. Drift you pay for every month.