← blog · October 3, 2026

Zero-Downtime Secret Rotation: Dual Keys, Alternating Users and the Consumers Left Behind

Rotation outages are caused by consumers, not by the value. Overlap windows, dual keys and alternating users, measuring the cutover, and the long-running processes that keep the old credential in memory.

Changing a password, an API key or a signing key looks like a one-line job: generate a new value, put it where the old one was. In practice it is one of the maintenance tasks most likely to cause an outage. The value is never the hard part. The consumers are. Who reads this secret, how long do they keep it in memory, and when do they notice it changed? Rotate without answering those questions and you either take a service down or, worse, leave one consumer failing quietly for weeks.

This post covers the approaches that make rotation zero-downtime and verifiable, the concrete steps, and the traps people keep falling into.

The real problem: you cannot add and revoke in the same instant

Rotation has two separate steps: making the new value valid and making the old value invalid. Outages happen when both are done at once. The second the old password is revoked, every consumer that has not picked up the new value fails authentication.

The correct model is an overlap window in which both values are valid. Add the new value, move the consumers over, measure that they moved, and only then revoke the old one. There are three common ways to get there.

Three approaches

1. Dual keys (two valid values for the same identity). If the provider lets an identity hold more than one key at a time, this is the cleanest path. Cloud access keys, HMAC webhook signing secrets and JWT signing keys all fall here. AWS IAM, for instance, allows at most two access keys per user, and that limit exists precisely so you can rotate.

For JWTs the model is built on the key ID (kid). Publish the new public key in the JWKS first, wait until verifiers have refreshed their caches, then switch signing to the new key. Remove the old key only after the longest-lived token it signed has expired. Do it the other way round, signing with the new key first, and every verifier holding a stale JWKS cache starts rejecting perfectly valid tokens.

2. Alternating users. Systems like databases give a user exactly one password. Instead of changing it in place, keep two users and switch to the other one on every rotation. This is the strategy AWS Secrets Manager's rotation templates call "alternating users".

In PostgreSQL, grant privileges to a shared group role and make both login users members of it:

CREATE ROLE app_owner NOLOGIN;
CREATE ROLE app_a LOGIN PASSWORD 'first-value' IN ROLE app_owner;
CREATE ROLE app_b LOGIN PASSWORD 'second-value' IN ROLE app_owner;
GRANT USAGE ON SCHEMA public TO app_owner;
GRANT SELECT, INSERT, UPDATE, DELETE ON ALL TABLES IN SCHEMA public TO app_owner;

The rotation order: set a new password on the idle user, point the application at it, confirm the old user's connections have drained, then disable the old user.

SELECT usename, count(*) FROM pg_stat_activity GROUP BY usename;
ALTER ROLE app_a NOLOGIN;

3. Change in place and restart. When the system supports a single value and a second user is not an option, you change the value and restart every consumer right away. A short outage is unavoidable here, so it needs a maintenance window.

If I have to pick: dual keys whenever the provider allows it, alternating users when it does not. Change-in-place is acceptable only when there are few consumers, one team can restart all of them, and a few seconds of errors genuinely do not matter. Teams that make it the default turn rotation into something everyone dreads, and eventually stop doing it.

Steps that work

Build the consumer inventory before you rotate. Who reads this secret is the hardest question in rotation. Reconstructing the answer by scanning file systems on the day is slow and incomplete. Keep the list of consumers next to the secret's record and update it on every rotation, so the next rotation starts from the previous one's notes.

Distribute from a single source. If the same secret sits in plain text in five config files, one of them will be forgotten. Keeping it in a vault and delivering it to consumers from there reduces the number of places to update to one.

Measure the cutover, do not assume it. After rolling out the new value and before revoking the old one, check whether the old value is still in use. Session lists in the database, the last-used timestamp that cloud providers record per key, and per-key request counts at the API gateway all give you that answer. If the old key was used in the last hour, your inventory is missing a consumer.

Measure after revoking, too. Once the old value is revoked, send a request with it and confirm it is actually rejected. A successful revoke call does not mean the value stopped working; some systems apply revocation only after a cache expires.

Traps

Environment variables are read when the process starts. A process that receives the secret through an environment variable will not see the change until it restarts. Containers make this sneakier: restarting the container is often not enough, because environment variables are fixed when the container is created. With Docker Compose, editing .env and running docker restart brings the container back with the old value; it has to be recreated with docker compose up -d.

Kubernetes Secret updates do not reach everything. Secrets mounted as volumes are refreshed by the kubelet after a delay, but those exposed as environment variables stay stale until the pod is recreated, and files mounted with subPath are never updated at all. Even when the file does change, the application has to reread it; for an app that reads its config once at startup, a changed file means nothing. That is why a kubectl rollout restart deployment/<name> after rotation is usually required.

Long-running processes carry the old credential in memory. A process that mounts object storage as a file system, a tunnel client, or a connection pool that never closes reads its credentials once at startup. Update its config file and it keeps using the old value, and the failure surfaces only when the old value is revoked. Worse, these processes usually sit underneath a nightly backup or sync job, so the breakage is noticed days later rather than on the first run. Put the long-running processes that read a config file into the inventory, not just the file itself.

Existing sessions survive revocation. In PostgreSQL, changing a user's password or setting NOLOGIN does not close connections already opened by that user. A connection pool can keep old sessions alive for hours, and you will believe the rotation "worked". The errors start the next day, when the pool tries to open a fresh connection. To make the cutover definitive, terminate the old user's sessions deliberately and watch the application reconnect.

Alternating users split object ownership. If migrations run as app_a and create a table, app_a owns it. At the next rotation, when you switch to app_b, you have lost the right to ALTER that table. Start migration scripts with SET ROLE app_owner; so objects belong to the shared role.

The error message does not mention rotation. A consumer stuck on the old key reports "403 Forbidden", "authentication failed" or just "could not connect". Those messages can be chased for days without anyone linking them to the rotation. Record the rotation date and read authentication errors against it for the following days; the root cause then takes minutes to find.

Generate the new value where nobody sees it. Printing a fresh secret to a terminal and copying it from there leaves it in shell history, screen recordings or a chat window. The value should travel straight from where it is generated into the vault and to the consumers. Ideally even the person doing the rotation never sees it.

When not to rotate more often

Frequent rotation does not improve security if the process is manual; it increases outage risk. A monthly manual rotation means a forgotten consumer every month. Automate the process and build verification into it first, then raise the frequency.

A better goal is to have fewer long-lived secrets to rotate at all. Workload identity, short-lived certificates, and tokens issued by an identity provider for minutes at a time remove the rotation problem at the root. Not holding a long-lived static key is usually less work and less risk than rotating one every week.

Do not ask how often you rotate. Ask how fast you could rotate after a leak. If you cannot rotate a leaked key without downtime within an hour, the problem is not your schedule, it is your process.