← blog · September 29, 2026

DNS Resolution in Kubernetes: ndots, Search Domains, CoreDNS Caching and the Five-Second Stalls

Why a pod turns one external lookup into eight DNS queries, what lowering ndots breaks, when autopath and NodeLocal DNSCache are worth it, where the exact five-second stalls come from, and which fix belongs at the application, pod or CoreDNS level.

Nobody thinks about DNS when an application inside Kubernetes calls an external API or another service, until tail latency jumps by an unexplained 5 seconds or the CoreDNS pods buckle under load as the cluster grows. The root cause is rarely the application. It is the resolver configuration written into the pod. This post explains how a pod resolves a name, why the defaults stop scaling, and which fix belongs where.

How a pod resolves a name

With the default dnsPolicy: ClusterFirst, the kubelet writes a resolver file into every pod containing the cluster DNS service address, three search domains (<namespace>.svc.cluster.local, svc.cluster.local, cluster.local) and options ndots:5. The search domains are what let you reach a service in the same namespace as just database and one in another namespace as database.other-ns.

The ndots:5 rule says: if the queried name has fewer than five dots, treat it as relative, try it with each search domain in order, and only then ask for the absolute name. api.example.com has two dots, so the resolver first asks for api.example.com.<namespace>.svc.cluster.local, then with svc.cluster.local, then with cluster.local, and only after three NXDOMAIN answers does the real name go out. With A and AAAA requested for every attempt, a single call to an external host becomes eight DNS queries. An HTTP request that never touches the cluster starts with four round trips to CoreDNS.

Verify it inside your own pod:

kubectl exec -it <pod> -- cat /etc/resolv.conf
kubectl exec -it <pod> -- dig +search api.example.com

dig does not apply search domains by default; the +search flag shows the behaviour your application actually sees. The query count and timing in that output are the concrete justification for everything that follows.

Three fixes, three different places

In the application: write the name as absolute. A trailing dot (api.example.com.) tells the resolver "this name is complete, do not append search domains" and removes every extra query. If you own the code or the configuration, this is the cheapest fix. The trap: some HTTP clients and TLS libraries carry the trailing dot into the Host header or SNI, and the server fails to match a virtual host. Test it with your library rather than assuming.

In the pod: lower ndots. If you cannot touch the source, change resolver options from the pod spec:

spec:
  dnsPolicy: ClusterFirst
  dnsConfig:
    options:
      - name: ndots
        value: "2"

With ndots:2, api.example.com is queried as absolute immediately, while database and database.other-ns still go through the search list. The only thing it breaks is the three-part shorthand: a library that uses service.ns.svc now has two dots, is queried as absolute without search, and fails to resolve. Grep for that pattern across your workloads before rolling this out.

In the server: let CoreDNS walk the search path. The CoreDNS autopath plugin performs the search walk on the server instead of the client: it answers the first query with a CNAME pointing at the correct result and the client never makes the remaining attempts. The price is that the kubernetes plugin must run in pods verified mode, meaning CoreDNS watches every pod to know which IP belongs to which namespace. In a large cluster that watch noticeably raises memory use and API server load. It makes sense for third-party workloads where you can change neither the source nor the pod, not as a general default.

Cache and timeouts

The CoreDNS cache plugin stores both positive and negative answers; the NXDOMAINs produced by the ndots walk come back from the negative cache, which sharply reduces load for workloads that keep asking for the same external name. The default cache 30 caps positive answers at a 30-second TTL. A short TTL reflects in-cluster service changes quickly but raises query volume; keeping a longer cache for external names in a separate server block is a reasonable balance.

example.com:53 {
    forward . /etc/resolv.conf
    cache 300
}
.:53 {
    errors
    health { lameduck 5s }
    ready
    kubernetes cluster.local in-addr.arpa ip6.arpa {
        pods insecure
        fallthrough in-addr.arpa ip6.arpa
        ttl 30
    }
    prometheus :9153
    forward . /etc/resolv.conf
    cache 30
    loop
    reload
    loadbalance
}

The first block of this Corefile answers only names under example.com with a long cache; the second is the default that kubeadm installs. The reload plugin picks up ConfigMap changes without a restart. The loop plugin detects CoreDNS forwarding to itself and stops the pod, so when you see CrashLoopBackOff, look at the node's own resolver file first.

Do not tune without measuring. coredns_dns_requests_total, coredns_dns_request_duration_seconds and coredns_cache_hits_total are labelled by query type and zone; the effect of an ndots change shows up in those three within minutes.

The five-second stalls

The classic symptom of DNS latency is a broken tail, not a broken average: most requests take milliseconds, some take exactly 5 seconds. That number is the default timeout of the glibc resolver and almost always points at a lost UDP packet. The best-known cause is the A and AAAA queries leaving the same socket at the same moment and racing to insert two conntrack entries during the DNAT to the cluster DNS service; the losing packet is silently dropped and the client waits out the timeout before retrying.

There are three remedies. On glibc-based images the single-request-reopen option sends the two queries from separate sockets and removes the race; add it via dnsConfig.options. musl-based images such as Alpine do not recognise the option and silently ignore it, so the problem persists on Alpine and many teams simply accept it as "DNS is sometimes slow". The durable fix is NodeLocal DNSCache: a DaemonSet on every node serves pod queries from a local cache, talks to the upstream over TCP, and never touches conntrack. Because the local hit rate is high, load on CoreDNS drops as well. The cost is one more component and a change to the kubelet's cluster DNS address; forgetting that DaemonSet during upgrades takes down DNS for the whole cluster.

Common traps

A pod with hostNetwork: true and a ClusterFirst policy inherits the node's resolver and cannot resolve cluster services; the correct value is ClusterFirstWithHostNet. dnsPolicy: None only works when dnsConfig supplies a complete nameservers list, otherwise the pod never starts.

If the cluster domain is anything other than cluster.local, every component with cluster.local hard-coded in source or Helm templates produces unresolvable names, and the symptom is usually connection refused rather than a DNS error. Before installing, check that the template reads the cluster domain from a value and that the value is not empty; an empty value produces a half name ending in svc. and warns nowhere.

Pushing the TTL towards zero for frequently queried workloads sends every request to CoreDNS and creates exactly the load you were trying to avoid. Instead of lowering the TTL, reduce the source of churn, for example headless service pods that are recreated far too often.

When to leave it alone

A small cluster, a single namespace and few external dependencies are fine on the defaults; ndots:5 is a genuine convenience and two CoreDNS replicas handle everything. The trouble starts as the cluster grows, external API calls multiply and the latency budget tightens. At that point the order is always the same: first see what is happening with dig +search, then fix it with the least intervention at the application or pod level, and only if the numbers are still bad move on to CoreDNS and the node-level cache.