Multi-tenancy · Case Study · 7 min read
One query that forgot the tenant
The platform had tenant scoping in its data layer, applied consistently, on every query but one. That one was in a service written by a different team for a feature nobody considered sensitive.
The client was a multi-tenant SaaS platform serving mid-market customers, preparing for a security review from a prospect substantially larger than their existing base. They had done the right things. Tenant scoping was enforced in the data access layer rather than in the controllers, there were tests for it, and the tests passed.
We asked for two separate tenant organisations with two accounts in each, which is the minimum that makes cross-tenant testing possible, and they provisioned them the same day.
Where we looked
Not at the main application. If tenant scoping is in the data layer and the tests cover it, the main application is usually the least interesting place to spend a week.
We went looking for the code that does not go through the data layer. Every platform has some: a reporting job, a webhook dispatcher, a search indexer, an export worker, an admin tool built for support staff. These are written by whoever needed them, often quickly, and they talk to the database directly because that was simpler.
We found the notifications service by reading the JavaScript bundle. It exposed an endpoint the interface used to fetch the notification detail behind a bell icon, taking a notification identifier as a parameter.
The finding
The identifier was sequential. We requested one below our own and received a notification belonging to another organisation entirely, with the record it referenced embedded in the payload.
Notifications on this platform carried context. A notification about an approval included the approval, which included the document, which included the customer record it concerned. We were able to walk the identifier space and read across every tenant on the platform from a standard user account in a trial organisation.
We stopped after confirming it worked across three organisations we could identify as distinct, recorded the evidence, and called the client the same afternoon rather than waiting for the report.
Why the tests passed
The tests covered the data access layer, and the data access layer was correct. The notifications service had been written eighteen months earlier by a team that has since moved on, it used its own database client, and it predated the scoping work by several months.
Nobody had done anything careless. The control was built, it was built well, and one component was simply outside it. That is the ordinary shape of a multi-tenancy failure: not an absent control, but a component the control never reached.
What we recommended
Two things, in order. First, scope the notifications service query immediately, which they deployed within four hours of our call.
Second, and more usefully, find every other component that talks to the database without going through the data access layer, and either route it through or give it the same scoping. They found four. Two were fine, one had the same class of issue on a narrower surface, and one was a scheduled job that had been writing to the wrong tenant for a period they were then able to bound and notify on.
We also suggested moving from sequential identifiers to unguessable ones. They asked whether that would have prevented the issue. It would not have. Unguessable is not unauthorised, and identifiers leak through exports, notifications and referral links regardless. It raises the cost of finding the bug without removing it, which is worth doing and is not a fix.
The question to ask your own platform
Not whether tenant scoping exists. Ask which components reach your data without going through it. The list is usually longer than the team expects, and it is where we start on every multi-tenant engagement.
Then ask whether your last test was given two tenants. Without a genuine second organisation, cross-tenant isolation can only be asserted, and it is the single most important control on a SaaS platform.
In short
- Point 1
- Tenant scoping in the data layer is correct and insufficient. Find what bypasses it.
- Point 2
- Reporting jobs, webhook dispatchers, search indexers and support tools are where the exception lives.
- Point 3
- Two tenants are the minimum for a meaningful test. One tenant proves nothing about isolation.
- Point 4
- Unguessable identifiers raise the cost of finding the flaw. They do not remove it.