← blog · August 25, 2026

Designing a multi-tenant SaaS: isolation strategies and what they really cost

Database per tenant, schema per tenant and shared tables with a tenant column carry very different migration, backup, noisy neighbour and leak costs. Which one to start with, and how to push the tenant boundary out of application code and into the database.

Where the tenant boundary gets drawn

One question decides most of the architecture of a multi-tenant application: at which layer does one customer's data become invisible to another customer's query? A system whose answer is "a predicate in the SQL" and a system whose answer is "a separate database" may sell the identical product, but they do not share an operational cost, a migration procedure, or a blast radius. You make this call in the first few months and you live with it for years, because switching isolation models afterwards is a data migration project, not a refactor.

Three models, three different bills

Database per tenant gives you the strongest separation. One tenant's rows are technically invisible to another tenant's query, and no single badly written statement can turn into a leak. The cost lands in operations: every database has its own connection pool, its own maintenance window, its own statistics. PostgreSQL ships with a max_connections default of 100, and that limit is server-wide. An application that opens a pool per tenant can hit it with a hundred tenants before a single user sends a request. Past a few hundred tenants, connection management becomes your actual job.

Schema per tenant opens a namespace per tenant inside one database. The connection is shared and the separation is logical. It looks attractive because it feels isolated while staying on one server. Its hidden cost is in the catalog: every schema means new rows per table and per index in the system catalogs. A hundred-table application at a thousand tenants pushes the catalog into hundreds of thousands of entries, which shows up in dump times and in autovacuum scheduling.

The second cost is the friction with connection pooling. If you select the schema with search_path and you have a PgBouncer running in transaction mode in front of you, a session-scoped SET is not safe: PgBouncer recycles the backend connection exactly at the transaction boundary, so the next transaction can open against a different tenant's schema. Either use SET LOCAL, or pin it at the role level, or have PgBouncer track the parameter with track_extra_parameters. Teams that do not know this see the hardest class of production bug there is: the wrong tenant's data, intermittently.

Shared tables with row-level separation put a tenant_id column on every table and filter every query by it. This is the cheapest, the most flexible and the most dangerous. Cheap because there is one schema, one migration, one pool. Dangerous because isolation depends entirely on the discipline of application code, and discipline does not scale.

Migrations break differently in each model

In shared tables a migration is a single operation with a single risk: locking a large table. A thousand tenants live in the same rows, so your ALTER TABLE touches all of them at once. The upside is that rollback is also a single operation.

In schema-per-tenant and database-per-tenant a migration becomes a loop, and loops fail partially. The migration passes on 380 of 500 schemas and hits a unique constraint violation on the 381st. Now you have production running two schema versions simultaneously, and your application code has to work against both. The only sane way to manage that is to write migrations backwards-compatible from the start: add the column nullable, dual-write for a while, backfill, move the code, drop the old column last. In a single-tenant application that discipline is a luxury. In schema-per-tenant it is the entry fee. You also need per-tenant migration state stored somewhere, or you will not know where the loop stopped.

There is one more thing nobody budgets for on day one: tenant provisioning time. Creating a hundred-table schema from scratch takes seconds, and seconds do not fit inside a signup form. Push provisioning onto a queue and show the user that the account is being prepared.

The real backup question is how one tenant comes back

Everyone takes backups. The question that matters is different: when a customer says they deleted four thousand records by accident yesterday at noon, can you restore only them without touching the other nine hundred and ninety nine tenants?

With a database per tenant the answer is easy. With a schema per tenant it stays reasonable: pg_dump -n pulls a single schema, you restore it under a temporary name, compare and move the rows back. With shared tables the answer hurts. You restore the full backup onto a separate server, filter that tenant's rows out of it, and resolve foreign key ordering and sequence values by hand. That is measured in hours, and a real incident is a bad time to try it for the first time.

This is why most teams on the shared-table model stop deleting for real. Soft deletes, an event log or an audit table move the majority of restore requests out of the backup path and into the application. Put it in the first release; adding it later does not recover the history you already lost.

Noisy neighbours

On shared infrastructure a single tenant can slow down everyone. The classic shape is the big customer's report query: a ten million row scan sweeps the shared buffer cache and everybody's latency goes up. The nastier shape is the queue. One tenant's import job produces fifty thousand tasks, the workers chew on it for hours, and another tenant's password reset email waits in line.

In order of effect: split queues by job type and expected duration rather than by tenant, then add a per-tenant concurrency cap. Route heavy reports to a read replica. Set a statement timeout in the database, because one query allowed to run forever is everyone's problem regardless of your isolation model. Moving a large tenant to its own database is also a tool, but you can only reach for it if the migration path already exists.

Forgetting the predicate is a structural defect, not a personal mistake

In row-level separation the rule is as simple as everyone says: put the tenant_id predicate on every query. The rule does not hold. Expecting it to hold across a hundred files, three years and five developers is not realistic. A report query, a batch job, an admin screen, a one-off data fix script; one of them will miss it.

The fix is not to remind people. It is to take the predicate out of the code and push it into the database. Row-level security in PostgreSQL exists for exactly this:

ALTER TABLE invoices ENABLE ROW LEVEL SECURITY;
ALTER TABLE invoices FORCE ROW LEVEL SECURITY;

CREATE POLICY tenant_isolation ON invoices
  USING (tenant_id = current_setting('app.tenant_id', true)::uuid)
  WITH CHECK (tenant_id = current_setting('app.tenant_id', true)::uuid);

Two details here save you. The first is FORCE: a table owner normally bypasses policies, so if your application connects as the role that owns the tables, the policy does nothing at all while you believe you are protected. The second is WITH CHECK: without it reads are filtered but writing a row under another tenant's identity stays legal.

Set the variable at the start of the transaction and bind it to that transaction:

BEGIN;
SELECT set_config('app.tenant_id', '3f2b6c1e-...', true);
SELECT id, total FROM invoices;
COMMIT;

The third argument of set_config scopes the setting to the transaction. On pooled connections that is the only correct usage, because a value written to the session leaks into whichever request picks up that connection after you.

Give it a single owner on the application side too: the layer that hands out connections sets the variable, and no code path gets a raw connection. Superusers and roles carrying BYPASSRLS skip policies unconditionally, so pin down with an actual test that your application role has neither. And watch that test go red once: if you have never removed the policy and measured the test failing, what you own is a feeling, not a guarantee.

Identity belongs to the user and tenant pair, not the user

The most common authorization mistake is attaching the role to the user. The same email address can hold different permissions at two customers, and eventually it will. Put the active tenant inside the session token, and after validating the token check separately that this user actually has a membership in that tenant. Any endpoint that reads a tenant identifier from the request body or the URL and trusts it is an open door to switching tenants from the browser.

Support will eventually ask for an impersonation feature. When it arrives, make it arrive through a separate path: a distinct role, an expiring session, an audit record for every use. An admin flag bolted onto the normal login flow puts your most dangerous capability on your least guarded route.

Per-tenant customization rots the architecture

The first large customer shows up and wants their own approval flow, their own invoice layout, their own field names. Serving that with tenant-specific branches in code is easy and it quietly splits the product in two. Six months later three tenants have three flows, no migration works across all of them, and the tests only cover the default path.

The healthy boundary is this: flexibility in the data model is allowed, branching in the code path is not. Per-tenant configuration, per-tenant field definitions, per-tenant templates are all data, and they all run through one code path. If a request cannot be expressed as configuration, it either becomes a product feature available to everyone or the answer is no. There is no durable third option.

Where to start

For most teams the answer is unambiguous: start with shared tables and row-level separation, and push the isolation down into the database with row-level security on day one. Those two have to arrive together, because row-level separation without row-level security is a data leak spread out over time.

Choose schema-per-tenant only if your contract says the data will be kept separate and your tenant count will not reach three digits. Choose database-per-tenant when customers are few, revenue per customer is high, or you carry data residency obligations. These are not the next stage of the same journey; they are the answer to a different business.

The preparation that actually pays off is narrow: make tenant_id part of the primary key on every table, or at minimum the leading column of every composite index, and write a path that exports one tenant's data and restores it into another database. With that path in hand, moving a large tenant to a dedicated database is a weekend. Without it, the same work turns into a quarter.