Multi-tenant SaaS: the isolation boundaries nobody tests until a customer finds them
Every multi-tenancy article is an argument about the database schema. Six of the eight places tenant isolation actually breaks are nowhere near it.
Every article about multi-tenant SaaS is an argument about the database. Shared tables with a tenant_id column, a schema per tenant, or a database per tenant. Pick a side, defend it at length.
That argument is worth having exactly once, and it takes about an hour. It is also the least likely place your isolation will actually fail.
The cross-tenant leaks I have been called in to look at were never schema problems. The tenant_id column was present and correct in every table. One was a queued job that ran against whichever tenant the worker happened to have loaded. One was a cache key. One was a signed URL from 2023 that still worked. One was an error tracker helpfully attaching the request payload, which contained somebody else's invoice.
None of those are database design. They are the same mistake in different clothing: tenancy was treated as a property of the request, when it is a property of every piece of work the request causes.
The schema question, settled in one table
Decide this, then stop thinking about it, because it is the one boundary you will enforce by default.
| Model | Isolation strength | What it costs you |
|---|---|---|
Shared tables, tenant_id column | Weakest. One missing predicate is a breach. | Cheapest to run and migrate. Only safe if enforcement lives in the ORM or the database, never in application code. |
| Shared database, schema per tenant | Middling. A connection pointed at the wrong schema is a breach. | Migrations multiply. Cross-tenant reporting gets awkward. Connection pooling needs real thought. |
| Database per tenant | Strongest. A breach needs two sets of credentials, not one bad query. | Operationally heaviest. Migrating 2,000 databases is a project, not a deploy. Per-tenant restore arrives free, which enterprise buyers ask about. |
There is no correct answer, only a correct answer for your risk profile. Regulated data and a handful of large enterprise tenants points at database per tenant. Thousands of self-serve accounts on a free tier points at shared tables. Most products I see should take shared tables with a tenant_id and spend everything they saved on the eight boundaries below, because that is where they will actually get hurt.
Isolation is eight boundaries, not one
Tenancy has to hold everywhere data comes to rest or moves. Here is the full surface, and what each failure looks like from the outside.
| Boundary | How it leaks | What the customer sees |
|---|---|---|
| Request context | Tenant resolved from a subdomain, header or path, then never re-checked against the authenticated user. | A valid session reads another account by editing a subdomain. |
| Data access | A raw query, an export or a reporting view that bypasses the ORM scope. | A CSV download containing other people's rows. |
| Background jobs | Tenant read from the worker's ambient state instead of the job payload. | Tenant A's monthly report, emailed with Tenant B's figures. |
| Cache | A key like dashboard:metrics with no tenant in it. | Branding, feature flags or counts belonging to whoever warmed the cache first. |
| Files and object storage | A flat bucket with guessable paths, or signed URLs that never expire. | A document reachable by anyone holding a link, permanently. |
| Logs and telemetry | Request bodies and stack traces captured with payloads attached. | Nothing. This one leaks to your own staff and your vendors, which is why it survives for years. |
| Outbound integrations | One shared credential or webhook endpoint serving every tenant. | Tenant A's data written into Tenant B's Slack, CRM or accounting system. |
| Search and vector indexes | One shared index, filtered after retrieval rather than during it. | An AI assistant citing a competitor's document. The newest of these, and the worst. |
Six of those eight live outside your database. No choice of schema defends any of them.
The handoff where tenant context dies
An HTTP request knows who it is for. A queue worker does not. It is a long-running process that handled Tenant A thirty milliseconds ago and will handle Tenant C next, and anything it reads from global state is whatever the previous job left lying there.
Request arrives on acme.yourapp.com
Middleware resolves Tenant A and sets it as current.
Does this user belong to Tenant A?
โณ no? return 404, not 403 โ a 403 confirms the account exists
Dispatch GenerateMonthlyReport
The job is serialised here. Anything not in the payload is gone.
The job waits four minutes in the queue
A worker picks it up
That process last ran a job for Tenant B. Current tenant is still Tenant B.
Report builds against whichever tenant is current
โณ and it succeeds, quietly, producing an entirely plausible PDF
Nothing here throws. The report generates, the email sends, and the bug surfaces when a customer recognises a number that is not theirs.
The fix is a rule rather than a library: a job carries its tenant in its payload and establishes it on entry. Never from ambient state.
// The leak. Tenancy::currentId() returns whatever the worker last set.
class GenerateMonthlyReport implements ShouldQueue
{
public function __construct(private Carbon $month) {}
public function handle(): void
{
$invoices = Invoice::whereMonth('issued_at', $this->month)->get();
// ... builds a report for the wrong tenant, successfully
}
}
// The fix. Tenancy is a constructor argument, established before any query runs.
class GenerateMonthlyReport implements ShouldQueue
{
public function __construct(
public readonly string $tenantId,
public readonly Carbon $month,
) {}
public function handle(): void
{
Tenancy::for($this->tenantId, function () {
$invoices = Invoice::whereMonth('issued_at', $this->month)->get();
// ...
});
}
}Enforce at the data layer, because people forget
A developer writing a controller should not have to remember to filter by tenant. If isolation depends on remembering, it holds for roughly four months, and then somebody ships a report that joins two tables and does not. This is the same instinct as keeping controllers boring: put the rule somewhere it cannot be skipped, rather than somewhere it has to be recalled.
// A trait every tenant-owned model uses.
trait BelongsToTenant
{
protected static function bootBelongsToTenant(): void
{
static::addGlobalScope('tenant', function (Builder $query) {
if ($tenantId = Tenancy::currentId()) {
$table = $query->getModel()->getTable();
$query->where($table.'.tenant_id', $tenantId);
}
});
static::creating(function (Model $model) {
$model->tenant_id ??= Tenancy::currentId();
});
}
}Note the creating hook. Half the tenant-scoping bugs I find are not reads at all. They are writes that land with a null tenant_id and become invisible to everyone, including the customer who created them.
A global scope has one hard limit: it only covers queries that go through the ORM. Raw SQL, a reporting view, a bulk insert-select, a DB::table() call in a console command, and anything your BI tool runs, all walk straight past it. Where the data genuinely must not mix, put the boundary in the database itself:
ALTER TABLE invoices ENABLE ROW LEVEL SECURITY;
CREATE POLICY tenant_isolation ON invoices
USING (tenant_id = current_setting('app.tenant_id')::uuid);
-- Set per transaction, not per session.
BEGIN;
SET LOCAL app.tenant_id = '8f14e45f-ea8d-4f29-9c1f-6f4a1b2c3d4e';
SELECT sum(total) FROM invoices; -- policy applies, scope or no scope
COMMIT;The four boundaries nobody draws in the diagram
Cache keys
Every cache key needs the tenant in it, including the keys you did not write: rate limiters, session stores, computed permission sets, settings objects, feature flags.
// Leaks. The first tenant to hit this endpoint warms it for everyone.
Cache::remember('dashboard:metrics', 300, fn () => $this->metrics());
// Scoped. Trivial to do now, unpleasant to retrofit at 2am.
Cache::tags(['tenant:'.Tenancy::currentId()])
->remember('dashboard:metrics', 300, fn () => $this->metrics());Tagging also buys you something you will want later: flushing one tenant's cache without touching anybody else's.
Files and signed URLs
Prefix every object with the tenant id, so a path traversal or an off-by-one in a filename is contained rather than catastrophic. Give signed URLs an expiry measured in minutes, and re-authorise on download instead of trusting a URL issued once. A permanent signed URL is a password you have emailed to someone and cannot rotate.
Logs and telemetry
This is the boundary that leaks for years unnoticed, because it leaks inwards. Tag every log line with the tenant id so you can investigate one customer's problem without reading another's data, and strip request bodies before they reach your error tracker. That tracker is a third-party processor holding your customers' data, and your data processing agreement almost certainly promises you are not sending it payloads.
Search and vector indexes
The newest boundary, and the one I now check first on any AI feature. A retrieval pipeline over tenant documents has the same boundary as a database table and far weaker defaults. It is not on the usual list of what breaks in a production RAG system, but on a multi-tenant product it is the first thing I look at.
The dangerous pattern is retrieve-then-filter: fetch the top twenty chunks by similarity from one shared index, then discard the ones belonging to other tenants. It looks correct and it is wrong twice over. The top twenty may be entirely other tenants' documents, so the tenant with the best-matching content silently gets a thin result set or none at all. And the filter is now application code sitting in the hot path of an LLM call, which is exactly where somebody later adds a caching layer and quietly drops it.
Filter inside the index, not after it: a namespace per tenant where your vector database supports one, or a pre-filter on a tenant field that the engine applies before scoring rather than after. Not every engine does both well, which is worth weighing while choosing a vector database rather than after the migration. Then test it the way you would test a table. Ask Tenant A's assistant a question only Tenant B's documents can answer, and assert that it says it does not know.
Prove it with tests that try to break in
Every multi-tenancy article tells you to write tests. Almost none say what the tests should assert, so here is the shape. The useful ones are negative: set up two tenants, then assert that the thing you fear does not work.
it('refuses to expose an invoice belonging to another tenant', function () {
[$acme, $globex] = Tenant::factory()->count(2)->create();
$invoice = Invoice::factory()->for($globex)->create();
$user = User::factory()->for($acme)->create();
actingAs($user)
->get("/invoices/{$invoice->id}")
->assertNotFound(); // 404, never 403
});Then the edge cases, which is where the real bugs live:
- Invitations and role changes. A user removed from a tenant loses access on the next request, not at the next login.
- Exports and reports. The paths most likely to use raw SQL, and the ones that hand data to a customer in a file you cannot recall.
- Retries and scheduled jobs. Run the job twice, from a worker that handled a different tenant in between.
- Soft deletes and restores. A restored record returns to its own tenant, and a deleted tenant's rows stop appearing in global counts.
- Support impersonation. Assert that it is logged, and assert that it expires.
Run these in CI from the first week. A negative test that has never failed is still doing work, because the day somebody adds a DB::table() shortcut to a controller it fails in the pull request instead of in a support ticket.
Where to start if this is already live
Do not start with a migration. Start by finding out whether you have a problem, in this order:
- 1Grep for query builders that bypass the ORM:
DB::table,DB::select, raw joins, anything in an export or reporting path. Each one is a place isolation is currently unenforced. - 2Grep for cache keys with no tenant in them, including inside packages you did not write.
- 3Check your error tracker for captured request bodies. If they are there, that is a live leak into a third party, and the cheapest one to stop today.
- 4List every queueable job and scheduled command, and mark the ones that do not take a tenant id. That list is your backlog, in priority order.
- 5Write the two-tenant negative test for your most sensitive read path. One test, today. It either passes, which is reassuring, or fails, which is more valuable.
That sequence takes an afternoon and tells you whether you are looking at a tidy-up or a rebuild. The gap between finding it now and finding it when an enterprise buyer's security questionnaire asks how you enforce tenant isolation is the difference between a week of work and a lost deal. Of everything in what a SaaS MVP actually costs to build, this is the line item people cut first and regret most.
Multi-tenancy is not a feature you add. It is a property that every line of the system either preserves or breaks, and the schema is only the first of eight places it can break.
Pick the simplest schema your risk profile allows. Then spend what you saved on the queue, the cache, the files, the logs, the integrations and the index, because that is where the leak will be.
Building something like this?
I design and ship these systems for clients: retrieval over private data, agents that complete real tasks, and the Laravel platforms underneath them.