Skip to content

The root-tenant model: how every request resolves to a tenant in a 42-module SaaS

6/7/2026Backend Development•Next Boilerplate•8 min read

How modules/tenant/ and modules/db/ implement per-tenant DataSource isolation behind a single UUID constant.

Multi-tenancy is one of those decisions that's easy to defer until it's impossible to retrofit. The architecture described here grew out of a concrete problem: 42 independent business modules — auth, billing, webhooks, SCIM, e-signatures, notifications — all needing to coexist in a single codebase while keeping tenant data physically or logically separate. The solution is built on two files, one UUID constant, and a deliberate choice not to invent a separate "system" scope.

The root tenant is a real tenant row

The first thing that surprises engineers looking at this codebase is that there is no "system" scope. No isSystem boolean, no systemDbConnection, no special branch in the service layer for platform-level operations. Instead, there is a root tenant:

// modules/tenant/tenant.constants.ts
export const ROOT_TENANT_ID = '00000000-0000-4000-8000-000000000000';

This UUID is a real row in the tenants table with name = "Platform". Super-admins are TenantMember rows pointing to that tenant with memberRole = 'ADMIN'. Platform-level resources — global subscription plans, system-wide coupon templates, audit logs for root operations — are owned by this tenant, not by a parallel data model.

The implication is that every service method that accepts tenantId works uniformly whether it's called for a regular customer tenant or for the root platform. There are no if-statements guarding against tenantId === ROOT_TENANT_ID scattered through the business logic. The root tenant is just another tenant, which is what makes the architecture composable across 42 modules.

// modules/tenant/tenant.service.ts (excerpt)
static async getById(tenantId: string): Promise<SafeTenant | null> {
  const base = await getDataSource();
  const repo = base.getRepository(Tenant);
  const cached = await redis.get(`tenant:${tenantId}`);
  if (cached) return JSON.parse(cached);
  const row = await repo.findOne({ where: { tenantId } });
  if (!row) return null;
  await redis.set(`tenant:${tenantId}`, JSON.stringify(row), 'EX', 300);
  return row;
}

The same getById is called for a customer lookup and for a root-tenant config resolution. No special case. This is the practical payoff of the root-tenant model: it keeps the module layer free of platform-awareness.

getDataSource() and tenantDataSourceFor(): the two-tier connection model

Every database operation in the app goes through one of two exported functions from modules/db/db.ts. The default one:

// modules/db/db.ts
const defaultDataSource = new DataSource({
  type: 'postgres',
  url: DEFAULT_DB_URL,
  schema: DEFAULT_SCHEMA,
  synchronize: env.NODE_ENV === 'development',
  logging: typeormLogging(env.NODE_ENV),
  entities: ENTITIES,
  migrations: [],
});

let defaultInitialized = false;

export async function getDataSource(): Promise<DataSource> {
  if (!defaultInitialized) {
    await defaultDataSource.initialize();
    defaultInitialized = true;
  }
  return defaultDataSource;
}

getDataSource() returns a lazily initialized singleton pointing at DATABASE_URL. It carries the full ENTITIES array — every entity from every module, 80+ types registered once. This is the connection used for shared-database tenants, which is the majority case.

The second function is what makes per-tenant physical isolation possible without changing any service code:

// modules/db/db.ts (continued)
const MAX_CACHED = 100;
const tenantCache = new Map<string, DataSource>();

function evictOldest(): void {
  const [key, ds] = tenantCache.entries().next().value!;
  tenantCache.delete(key);
  ds.destroy().catch(() => {});
}

export async function tenantDataSourceFor(tenantId: string): Promise<DataSource> {
  if (tenantCache.has(tenantId)) return tenantCache.get(tenantId)!;

  const base = await getDataSource();
  const row = await base.getRepository(TenantDatabase).findOne({ where: { tenantId } });
  if (!row) return base;          // ← no row = shared DB, falls back

  const { url, schema } = parseDbUrl(row.databaseUrl);
  if (tenantCache.size >= MAX_CACHED) evictOldest();

  const ds = new DataSource({
    type: 'postgres',
    url,
    schema,
    synchronize: env.NODE_ENV === 'development',
    logging: typeormLogging(env.NODE_ENV),
    entities: ENTITIES,
    migrations: [],
  });
  await ds.initialize();
  tenantCache.set(tenantId, ds);
  return ds;
}

The lookup path: check the in-process LRU cache, then check the system tenant_databases table for a tenantId → databaseUrl row, then fall back to the shared connection. The cache is capped at 100 entries; the oldest is evicted and its connection destroyed when the cap is hit. Most calls return in microseconds after the first warm-up.

Services that need per-tenant isolation call tenantDataSourceFor(tenantId) instead of getDataSource(). SCIM provisioning, tenant-scoped webhooks, subscription state resolution — they all go through this path. Services that operate across tenants (plan catalog, root-tenant audit, cron purge) use getDataSource() directly.

The tenant_databases table: opt-in isolation

Physical database isolation is not the default. It's a deliberate operator decision, encoded in a single table:

// modules/db/entities/tenant_database.entity.ts
@Entity('tenant_databases')
export class TenantDatabase {
  @PrimaryColumn('uuid')
  tenantId!: string;

  @Column('text')
  databaseUrl!: string;
}

To move a tenant to its own database, an operator inserts one row:

INSERT INTO tenant_databases (tenant_id, database_url)
VALUES (
  'f47ac10b-58cc-4372-a567-0e02b2c3d479',
  'postgresql://tenant_db_user:secret@dedicated-host:5432/tenant_db?schema=public'
);

From that point forward, every tenantDataSourceFor('f47ac10b-...') call returns a connection to the dedicated database instead of the shared one. No deployment, no code change, no service restart — the cache simply warms up on the next request. clearTenantDsCache(tenantId) is available for cases where the mapping row is updated and the running instance needs to reconnect immediately.

This is a meaningful operational capability. You can start a product on a shared database, move a high-volume or compliance-constrained tenant to dedicated hardware when it makes sense, and do it live. The business modules never know which path they took.

Row-Level Security as defense-in-depth

The two-tier connection model is the first line of isolation. The second is a Postgres RLS policy applied at the database level, defined in modules/db/migrations/001_tenant_rls.sql:

-- Every connection sets this once per request:
-- SET LOCAL app.current_tenant = '<tenant-uuid>';

CREATE OR REPLACE FUNCTION app_current_tenant() RETURNS UUID AS $$
DECLARE v TEXT;
BEGIN
  v := current_setting('app.current_tenant', true);
  IF v IS NULL OR v = '' THEN RETURN NULL; END IF;
  RETURN v::UUID;
END;
$$ LANGUAGE plpgsql STABLE;

-- Policy on tenant-scoped tables:
CREATE POLICY tenant_isolation ON tenant_members
  USING (
    current_setting('app.bypass_rls', true) = 'on'
    OR tenant_id = app_current_tenant()
  );

Every service-layer where: { tenantId } clause is already doing the right thing, but a forgotten clause now returns an empty result set rather than a cross-tenant data leak. The RLS policy is not the primary enforcement mechanism — it's a backstop. Cron jobs and platform-admin tooling that legitimately need cross-tenant access connect with a role that has BYPASSRLS and set app.bypass_rls = 'on'.

The combination of three layers — service-layer where clauses, the tenantDataSourceFor routing, and Postgres RLS — is what makes this architecture defensible at enterprise scale. Any single layer failing doesn't create a cross-tenant breach; the other two still hold.

The lifecycle: creation to deletion

Tenant lifecycle runs through four status values — ACTIVE, SUSPENDED, PENDING_DELETION, DELETED — with an explicit 30-day soft-delete grace period:

// modules/tenant/tenant.deletion.service.ts (condensed)
static async requestDeletion(tenantId: string): Promise<void> {
  const ds = await getDataSource();
  await ds.getRepository(Tenant).update(tenantId, {
    tenantStatus: 'PENDING_DELETION',
    deletionRequestedAt: new Date(),
    deleteAfter: addDays(new Date(), 30),
  });
}

static async purgeExpiredTenants(): Promise<number> {
  const ds = await getDataSource();
  const expired = await ds.getRepository(Tenant).find({
    where: { tenantStatus: 'PENDING_DELETION', deleteAfter: LessThan(new Date()) },
  });
  for (const t of expired) {
    await ds.getRepository(Tenant).softRemove(t);
  }
  return expired.length;
}

The purge is driven by a BullMQ cron job (tenant-purge, default 0 4 * * *) defined in modules/tenant/tenant.deletion.job.ts. Serverless deployments can skip the queue and hit a protected cron endpoint instead — the service method is the same either way.

Trade-off

The root-tenant model simplifies the service layer significantly but requires discipline at every call site: services that should only run in root-tenant context need explicit checks, and services that fan out across all tenants need to handle the root tenant deliberately (usually exempting it from billing gates and subscription limits). The shared-database default also means the ENTITIES array is registered on every DataSource — per-tenant databases carry all 80+ entity definitions even if they only use a subset. That's an acceptable cost for code uniformity, but it's a cost worth knowing about before the tenant count reaches the thousands.

Business impact

Physical database isolation, configurable at the row level with no deployment, means that a standard SaaS contract can become a dedicated-infrastructure contract for a single enterprise customer without forking the codebase. The root-tenant model means platform features — plan management, system audit, super-admin tooling — are built once and reused, not rebuilt in a separate "admin" application. For a solo operator running 42 modules, that's the difference between a manageable surface area and an unmanageable one.

What to do next

If you're designing multi-tenancy from scratch, the first question to answer is whether your compliance requirements dictate physical isolation at launch or whether you can start shared and migrate. The architecture here supports both, but starting shared and migrating later requires that every service call already goes through tenantDataSourceFor — retrofitting that after the fact is painful. Make the routing decision a structural one from day one, even if the fallback path always returns the shared connection.

Related Articles

Same Category

Comments (0)

Newsletter

Stay updated! Get all the latest and greatest posts delivered straight to your inbox