Building Multi-Tenant SaaS on AWS: 12 Lessons We Learned the Hard Way
Shipping a multi-tenant SaaS on AWS sounds like a solved problem until you start drawing the boxes. Here are twelve patterns — and four anti-patterns — we keep coming back to when designing isolation, billing, and noisy-neighbour defences for B2B platforms in production.
TL;DR Pick a tenancy model before you write a single line of code, push the tenant identifier into your auth tokens (not your URL paths), use a single DynamoDB table partitioned by tenant, measure cost per tenant from day one, and never let a free-tier customer share a Lambda concurrency pool with an enterprise one. The rest of this post is the long version.
OrbitNexa has built a dozen multi-tenant B2B SaaS platforms on AWS over the last four years — invoice automation, expense management, an HRMS, and a few we can't name publicly. Across all of them, the same handful of decisions kept showing up at architecture reviews. Get them wrong early and you pay for two years. Get them right and almost everything else becomes a refactor instead of a rewrite.
This is the cheat sheet we hand to our consulting engagements on day one.
What "multi-tenant" actually means
Before we get into patterns, it's worth being explicit. A SaaS platform is multi-tenant when a single deployment serves multiple customer organisations whose data and configuration must stay isolated from each other. That definition splits cleanly into three sub-questions:
- How much infrastructure do tenants share? (Database, compute, network)
- How much identity do tenants share? (Cognito pool, user directory)
- How is the boundary enforced? (Application code, IAM, network)
Most "multi-tenant" debates are really about question 1. The other two get answered as side effects and then bite you eighteen months later.
1. Choose your tenancy model on a napkin, in pencil
There are exactly three real choices, and you have to commit to one per data layer:
| Model | Database | Compute | Cost | Blast radius | Migrate later? |
|---|---|---|---|---|---|
| Silo | Per-tenant tables / accounts | Per-tenant Lambdas | High | Tiny | Hard |
| Pool | Single shared table, tenant key | Shared Lambdas | Low | Wide | Easy → Bridge |
| Bridge | Shared table + tenant overrides | Shared Lambdas, isolated by quota | Medium | Medium | Easy ↔ either |
The pool model is right for almost every B2B SaaS until you have a customer who legally cannot share infrastructure with anyone else. At that point you graduate that *one tenant* to a silo and leave the rest in the pool.
Don't try to support all three from day one. The abstraction layer needed to switch transparently is *itself* a six-month project.
2. The tenant ID belongs in the JWT, not the URL
The single most common anti-pattern we see: GET /api/tenants/acme/invoices or GET /api/v1/orgs/acme/.... It looks RESTful and it is *catastrophically* exploitable.
Here is the rule: the tenant identifier is an authorisation claim, not a request parameter.
Put it in the Cognito ID token, validate it at the API Gateway authoriser, and never read it from the request path or body for anything that affects data access.
// backend/functions/shared/auth.tsimport { CognitoJwtVerifier } from "aws-jwt-verify";const verifier = CognitoJwtVerifier.create({userPoolId: process.env.USER_POOL_ID!,tokenUse: "id",clientId: process.env.CLIENT_ID!,});export interface AuthContext {sub: string;email: string;tenantId: string;groups: string[];}export async function authenticate(token: string): Promise<AuthContext> {const claims = await verifier.verify(token);// `custom:tenantId` is set by the pre-token-generation Lambda on// first sign-in. After that it's immutable from the client's// perspective — admins can rewrite it but the user can't.const tenantId = claims["custom:tenantId"] as string | undefined;if (!tenantId) {throw new Error("Token has no tenant claim — refuse the request.");}return {sub: claims.sub,email: claims.email as string,tenantId,groups: (claims["cognito:groups"] as string[] | undefined) ?? [],};}
In the handler, you build the DDB query from auth.tenantId *only*. The client cannot influence which tenant's data they read, ever, even if they spoof every other field in the request.
// backend/functions/invoices-api/index.tsimport { QueryCommand } from "@aws-sdk/lib-dynamodb";export async function listInvoices(auth: AuthContext, status?: string) {return docClient.send(new QueryCommand({TableName: process.env.INVOICES_TABLE,KeyConditionExpression: "PK = :pk",FilterExpression: status ? "#s = :status" : undefined,ExpressionAttributeNames: status ? { "#s": "status" } : undefined,ExpressionAttributeValues: {":pk": `TENANT#${auth.tenantId}`, // ← derived from JWT, not request...(status ? { ":status": status } : {}),},}));}
The same rule applies to clientId, organizationId, workspaceId — whatever you call your top-level grouping. It belongs in the token.
3. Single-table DynamoDB, partitioned by tenant
For the pool and bridge models, the right DynamoDB layout is a single table where the partition key encodes the tenant. We use:
PK = TENANT#<tenantId>
SK = <entityType>#<entityId>So an invoice for ACME looks like:
PK: "TENANT#acme-corp"
SK: "INVOICE#inv_01HXG3K9..."This gives you:
- Locality: every tenant's data lives on the same partition. Range scans for one tenant's invoices are cheap.
- Isolation by routing: a bug that builds a query without the tenant prefix simply returns no results instead of silently leaking another tenant's data.
- Throttling per tenant: hot-partition warnings light up *which tenant* is hot, not "the invoices table".
Add a few GSIs for the access patterns that cross tenants (admin views, billing reconciliation) and resist the urge to add per-entity GSIs you'll never use.
Things that **don't** work as PK
customerId(changes with re-orgs, M&A — you can't change a PK)email(users change email; you don't want to rewrite history)uuid(fine as an ID, useless as a partition strategy)
A short, slug-style tenant ID that you generate at signup time and never expose to the user works best.
4. One Cognito pool, custom claims on every token
Resist the temptation to create one user pool per tenant. It looks like better isolation; it gives you a permissions nightmare.
Run a single Cognito user pool, and use the *pre-token-generation* trigger to stamp tenant + role claims onto every issued token. That trigger is the one place in your system that decides "who is this user, and which tenant do they belong to?".
// backend/functions/cognito-pre-token/index.tsimport type { PreTokenGenerationTriggerHandler } from "aws-lambda";import { GetCommand } from "@aws-sdk/lib-dynamodb";export const handler: PreTokenGenerationTriggerHandler = async (event) => {const { userName } = event;const profile = await docClient.send(new GetCommand({TableName: process.env.USERS_TABLE,Key: { PK: `USER#${userName}` },}));if (!profile.Item) {// First sign-in — defer to your invite/signup flow. Returning the// event unchanged means the user gets a token with no tenant// claim, which your API layer should reject as unauthorised.return event;}event.response.claimsOverrideDetails = {claimsToAddOrOverride: {"custom:tenantId": String(profile.Item.tenantId ?? ""),"custom:role": String(profile.Item.role ?? "member"),},};return event;};
Now every downstream service — API Gateway authoriser, AppSync resolver, Lambda — sees the same tenant + role claims and doesn't need a second database round-trip to figure out who's calling.
5. Per-tenant Lambda concurrency, the cheap way
This one bites later. By default, all your Lambdas share a regional concurrency pool of 1,000. One tenant's bug can starve every other tenant.
You don't need separate Lambdas per tenant. You need *reserved* concurrency on the Lambdas that handle hot paths, sized for your worst tenant:
// backend/lib/stacks/api-stack.tsconst invoicesApiFunction = new NodejsFunction(this, "InvoicesAPI", {entry: path.join(__dirname, "../../functions/invoices-api/index.ts"),runtime: lambda.Runtime.NODEJS_20_X,timeout: cdk.Duration.seconds(30),// Cap the function at 200 concurrent executions. Any single tenant// can burst to ~150 without affecting the other endpoints. The// remaining 800 from the regional pool stays available for every// other Lambda in the account.reservedConcurrentExecutions: 200,});
For the *truly* noisy-neighbour problem (a tenant generating 10× the API traffic of all others combined), look at Lambda's Provisioned Concurrency for that tenant's reserved set, or split the offending tenant onto a silo'd Lambda alias. The pool model gracefully degrades to a hybrid when you need it to.
6. Cost per tenant, measured from day one
Pick *any* AWS billing post-mortem from your favourite SaaS company and search the page for "we couldn't tell which customer was costing us money". It is always there.
Two things make this tractable:
- **Tag every resource with
tenant=*where you can.** Lambdas, S3 prefixes, DDB items via client request tags, CloudWatch log streams. - Sample structured logs. Every API call writes a single JSON line with
tenantId,endpoint,durationMs,bytesOut. CloudWatch → Athena → a dashboard that breaks down spend by tenant by day.
Cost per tenant is the single best input to your pricing model. Without it you are guessing. With it, you can do things like:
- Confidently tell a tenant *why* their enterprise tier is priced that way
- Spot the next at-risk tenant before they churn
- Identify free-tier users who are actually costing you money and convert (or off-board) them
7. Stripe metering — bill the thing your customer values, not the thing AWS charges you for
We have made this mistake. Here is the lesson, condensed: decouple your unit of billing from your unit of cost.
If you bill by "number of invoices processed" but your cost driver is "S3 storage for PDF copies", a tenant who imports 10,000 historical PDFs in one batch will cost you $400 of S3 transfer and pay you $0 because they didn't process anything *new*.
The fix is two metering streams:
// backend/functions/invoices-api/handlers/create.tsimport { stripeClient } from "../shared/stripe";export async function createInvoice(auth: AuthContext, invoice: Invoice) {// 1. Save the invoice as usualawait docClient.send(new PutCommand({ /* ... */ }));// 2. Meter the billable event (what the customer pays for)await stripeClient.subscriptionItems.createUsageRecord(auth.stripeSubscriptionItemId,{ quantity: 1, timestamp: Math.floor(Date.now() / 1000) },);// 3. Emit an internal cost signal (what AWS charges YOU for)await emitInternalMetric({tenantId: auth.tenantId,metric: "invoice.created",value: 1,storageBytes: invoice.pdfBytes,computeMs: 240,});}
You can change either dial independently. Pricing experiments don't require code changes — only Stripe configuration. Cost optimisation doesn't change what you charge — only how you allocate spend internally.
8. Idempotency keys on every write that crosses a network boundary
A retry from a flaky mobile network charges the customer twice and pages you at 2am. Avoid it with a single header.
// frontend/src/lib/api/client.tsimport { v4 as uuid } from "uuid";export async function createInvoice(input: CreateInvoiceInput) {// The client generates the idempotency key once per logical// operation. Retries on the same logical operation reuse it.const idempotencyKey = input.idempotencyKey ?? uuid();return apiClient.post("/invoices", input, {headers: { "Idempotency-Key": idempotencyKey },});}
The server stores the key + the response for ~24 hours. A retry with the same key returns the cached response without executing the side effect.
This is the highest-ROI 50 lines of code in any SaaS backend, and yet most teams skip it until the first 2am page.
9. The four anti-patterns
These are the ones we see most often during the architecture audit phase of a new engagement.
❌ Per-tenant subdomain DNS
acme.yourapp.com looks slick. It costs:
- A wildcard ACM cert + Route 53 chaos
- One DNS propagation outage per tenant on first signup
- Eight failed cookie-domain debugging sessions
Use path-based routing (yourapp.com/t/acme) or — better — don't put the tenant in the URL at all (see lesson 2).
❌ "Just use the customer's AWS account"
A few enterprise prospects will ask for it. Don't do it for fewer than five. The cost of running fundamentally different deploy pipelines per tenant compounds — every feature ships at the speed of your slowest tenant's review board.
If you do build it, build it as a *silo* of the same codebase, not a fork. The shared-codebase rule is non-negotiable.
❌ Soft-deleting tenants by flipping a boolean
tenant.active = false is not GDPR compliant, not auditable, and not actually saving you any storage. Either:
- Archive (move to S3 Glacier with a 7-year retention policy + key rotation), or
- Hard-delete with a 30-day grace period and an irreversible terminal state
Pick one. Document it. Tell your customer which it is at signup.
❌ The "we'll add isolation later" promise
Multi-tenancy is a property of your data model. You cannot bolt it on after the fact without rewriting every query in your codebase. The pool→bridge migration we mentioned earlier is *one* table change away if you've used the PK pattern from lesson 3. It is *six months* away if you stored everything by userId and trusted the application layer.
If you're starting today, put tenantId on every entity from the first migration.
10. What a healthy multi-tenant deploy looks like
Here's what a production deploy of an OrbitNexa-built B2B SaaS looks like, redacted to the architecture:
# CloudFormation outputs, abbreviatedSaasCoreTable:PartitionKey: TENANT#<tenantId>SortKey: <entityType>#<entityId>GSI1: BillingByDate # Cross-tenant adminGSI2: SearchByEmail # Per-tenant lookupsBillingMode: PAY_PER_REQUESTPointInTimeRecovery: enabledCognitoUserPool:Triggers:PreSignUp: validate-email-domainPreTokenGeneration: stamp-tenant-and-role-claimsMFA: SOFTWARE_TOKEN_MFAApiGateway:Authorizer: cognito-user-poolEndpoints: 47WAF: rate-limit-per-tenant-per-minuteLambdas:ReservedConcurrency:invoices-api: 200expense-api: 150payroll-api: 50 # Long-running, expensivecontent-api: 100 # Public + adminProvisioned:invoices-api: 10 # Smooth out cold starts on hot pathS3:Buckets:attachments:LifecyclePolicy: standard → IA after 90 days → Glacier after 1yServerSideEncryption: KMS-CMK-per-tenant # ← importantpublic-assets:CloudFront: enabledBilling:Stripe: usage-based + flat tierMetering: internal cost signal + external billable event
Nothing exotic. Nothing custom-built. The discipline is in the *combination*.
11. The migration playbook (when you inherit something worse)
About half of our engagements start with: "we already shipped, and it's a mess". Here is the order of operations that has worked across six different rescues:
- Add tenant tagging to every entity write. Backfill is the next step, but don't add to the mess.
- Backfill in batches. Existing data → derive tenant from owning user → write. Verify a sample.
- Add the tenant claim to JWT tokens. Stop reading tenant from URLs.
- Refactor one endpoint at a time. Start with the hottest read path. Each refactored endpoint validates tenant from JWT + ignores URL tenant.
- Once 100% of endpoints are migrated, delete the URL-tenant fallback. Not before.
- Add cost-per-tenant logging. Use that signal to find your worst tenants.
- Tackle noisy-neighbour after the cost data is in. Otherwise you optimise the wrong thing.
The migration takes 4–8 weeks per platform. It is almost always cheaper than the rewrite the team initially proposes.
12. What we'd do differently
In hindsight, on every project, we'd do these things on day one:
- Stand up the cost-per-tenant dashboard *before* the first customer signs up
- Wire an SLO dashboard with per-tenant error rates and p95 latencies
- Make the tenant-suspend operation a single CLI command, with audit log, on day one
- Document the exact tenancy model decision in the repo's
ARCHITECTURE.md - Reserve the right Cognito custom attributes from the very first deploy — adding them later requires a pool migration
The hardest one is the last. Every Cognito user pool we have ever migrated has burned a week. Future-proof the schema by adding custom:tenantId, custom:role, custom:locale, custom:onboardingStage at pool creation. You will use most of them within a year.
Closing — the reframe
Multi-tenancy is not a feature. It is not a single architectural decision. It is a property your data, your auth, your billing, your monitoring, and your incident response *all need to share*. Get one of those out of step with the others and you create a category of bug that's hard to find and harder to fix.
The good news: the patterns are well-known. The bad news: every B2B SaaS team learns them by stepping on the same six rakes in the same order. This post exists so your team can skip ahead to the seventh.
If you'd like a second pair of eyes on your multi-tenant architecture — whether you're shipping in 90 days or rescuing something that's already in production — reach out to OrbitNexa and we'll walk through it.
About the author
Surya Kenguva is a software engineer at OrbitNexa, where he leads architecture for B2B SaaS engagements and platform engineering. He has spent the last six years building multi-tenant systems on AWS for fintech, healthtech, and HR platforms. You can find him on LinkedIn or the OrbitNexa blog.
Further reading
- AWS SaaS Factory: Multi-Tenant Architectures — the official AWS reference patterns
- DynamoDB Single-Table Design: Alex DeBrie — the canonical resource on the table-design lesson above
- *Designing Data-Intensive Applications* — Martin Kleppmann (chapters 6 and 7 in particular)
Got a multi-tenant story of your own? Reply to this email — a real human reads every response, and the best ones land in our next post.