← blog · September 9, 2026

How to Design a Client That Keeps Working Offline

Keeping an app from crashing when the network drops is one problem; making sure the user's changes survive that drop is a different one. We compare caching, optimistic UI and local-first architectures, and walk through queue design and conflict resolution.

Offline is two different problems

"Make it work offline" sounds like a single requirement, but it splits into two. The first is reading: a user who has already seen some content should still be able to see it without a connection. The second is writing: a user should be able to fill in a form, update a record, or add an item while offline, and that change should not disappear. The solutions to these two problems are unrelated, and most "offline support" efforts fail because they solve the first and quietly skip the second, leaving the interface saying "saved" while nothing was actually saved anywhere.

Here are three approaches, when each one is the right call, and the concrete failure modes of each.

Approach 1: read-only caching

This is the cheapest and simplest option. Through a service worker you set up a stale-while-revalidate strategy: on a request, you return the cached response immediately, then fetch fresh data in the background and update the cache for next time. When the connection drops, the user still sees the last page they loaded.

The limit is obvious: this only covers reads. It does nothing for a form submission or a button click, because there is no "response" to cache, only a "request" that still has to reach the server. For content sites, documentation, catalogs, and anything that is mostly read rather than written, this is enough on its own and needs no further complexity.

The most common mistake here is not versioning the cache. If the service worker's cache name stays fixed (cache-v1) across deployments, users can keep seeing months-old JavaScript and stale API responses long after you shipped a fix; a bug report comes in that you cannot reproduce, because your own cache is already current. Changing the cache name on every deploy and clearing old caches in the activate event is the standard fix for this class of bug.

Approach 2: optimistic UI with a local queue

When users need to perform an action while offline (create a task, add a comment, change a setting), you move to the second approach. The interface marks the action as successful right away (an optimistic update), while the actual request is written to a durable queue in the browser, in IndexedDB rather than localStorage, since the latter is synchronous and blocks the main thread. When connectivity returns, the queue is drained against the server in order.

Three things have to be right for this to hold up.

First, every queued item needs a client-generated idempotency key, a UUID created before the request is sent. If a request times out on the network but actually reached the server, the client may wrongly assume it failed and resend it. Unless the server recognizes the same key as "already processed" and discards the duplicate, an order gets created twice or a balance gets debited twice.

Second, you need a reliable trigger to drain the queue. The online event and the Background Sync API exist for this, but Background Sync support is inconsistent across browsers; in practice you end up combining the online event, a check on visibilitychange, and a simple periodic poll, because relying on a single trigger leaves the queue stuck indefinitely in some browser.

Third, session expiry. If a user stays offline long enough, their auth token may expire before the queue gets a chance to drain. If you send the queued requests with the stale token instead of refreshing it first, every item comes back with a 401, and because the interface already told the user "sent," they will never notice anything went wrong.

Approach 3: local-first architecture

For collaborative tools, shared documents, whiteboards, notebooks edited by more than one person at once, the second approach breaks down, because there is no single "request" but a continuous stream of small changes, and multiple users can edit the same data at the same time. This is where local-first architecture comes in: data always lives in local storage first, the server is a sync point rather than a central authority, and merging is automated with CRDTs (conflict-free replicated data types) or operational transformation.

CRDTs give a mathematical guarantee, for specific data structures like counters, sets, ordered lists, and text, that both sides converge to the same result no matter what order the changes are applied in. That lets you write without asking the server first and merge without conflict later. The cost is complexity: your data model has to fit the shapes CRDTs support, deletions are usually represented as tombstone records that accumulate over time, and cleaning those up is its own engineering problem that needs a garbage collection pass.

What a queued item needs to carry

Every mutation you push into the queue should carry at least an idempotency key, the operation type, the target resource, the payload, and a retry count. A simplified example:

{
  "id": "6f2b6e2a-6c31-4e7a-9e2e-2a6f2b6e2a6c",
  "type": "update-task",
  "resource_id": "task-482",
  "payload": { "title": "Send report", "done": true },
  "created_at": "2026-09-09T08:14:00Z",
  "attempts": 0
}

Tracking the retry count matters: instead of retrying a request that the server keeps rejecting with a 400 forever, you should drop it from the queue after three or four failed attempts and tell the user plainly that it could not be sent. Otherwise the queue grows silently in the background, the user notices nothing, and once storage quota fills up the browser starts evicting the oldest entries on its own, at which point nobody can tell which operations were actually lost.

Don't wait for a real outage to test this. The offline toggle in browser devtools, along with artificial latency and packet-loss settings, is the fastest way to see whether the queue actually fills up, drains in the right order once connectivity returns, and whether the idempotency key genuinely prevents duplicate submissions. Beyond exercising these scenarios manually once, the queue logic deserves plain unit tests that don't depend on a real server; a test that fakes the network layer is always faster and more repeatable than the real thing.

When not to reach for this at all

For anything that needs strong consistency, balances, payments, inventory decrements, the right call is usually to not allow offline writes at all, and instead show a clear message that the action requires a connection. The risk here isn't technical: if a user places an offline order from two devices, an optimistic interface will tell both devices "order received," the server may only be able to honor one, and explaining that after the fact becomes a support problem, not an engineering one.

If your application is mostly used in an office, on a reliably connected network, an admin panel or an internal tool, the queueing, conflict resolution, and testing surface that offline support adds will cost more than it returns. In that case the better investment is simply giving a clear error on network failure without losing the user's unsaved input; that alone is a more trustworthy experience than a half-built claim of offline support.

The right question is never "should this work offline," but "for this specific piece of data, how much consistency am I willing to trade away." As the answer changes across your application, you will likely end up combining all three layers: caching for reads, a queue for non-critical writes, and CRDTs for anything genuinely co-edited.