· NERVICO · cloud-architecture · 11 min read
AWS for SaaS: Implementing the Multi-Tenant Pattern
Technical guide to implementing multi-tenant architectures on AWS: isolation models, data patterns, per-tenant identity management, and real costs of each approach.
Multi-tenancy is what makes a SaaS product profitable. Without it, each customer requires its own infrastructure instance, its own deployment, and its own operational management. With 10 customers, that is inconvenient. With 1,000, it is unsustainable.
But implementing multi-tenancy correctly on AWS is more complex than it appears. The decisions you make in the first weeks (how to isolate data between tenants, how to manage identities, how to distribute load) determine whether your SaaS scales to thousands of customers or becomes an operational nightmare past the hundredth.
This article covers the three main multi-tenancy models on AWS, with their real trade-offs in cost, complexity, and isolation, and the concrete patterns we have seen work in real SaaS products.
What Multi-Tenant Means in Practice
Technical Definition
A multi-tenant system is a software instance that serves multiple customers (tenants) from shared infrastructure. Each tenant has its own data, its own configuration, and in many cases, its own customization. But they share the same code, the same servers, and potentially the same database.
The alternative is single-tenant: a complete instance of the system for each customer. It is simpler to implement but exponentially more expensive to operate.
The Three Decision Axes
Every multi-tenant architecture is defined by three fundamental decisions:
- Compute isolation: Each tenant has its own application instance, or they share the same instance.
- Data isolation: Each tenant has its own database, its own schema within a shared database, or they share tables with a tenant_id column.
- Network isolation: Each tenant has its own VPC, or they share subnets within the same VPC.
The level of isolation you choose directly affects three things: cost per tenant, operational complexity, and security level. There is no universally correct option.
Isolation Models on AWS
Silo Model: Full Isolation per Tenant
Each tenant has its own dedicated infrastructure: its own VPC, its own compute instances, its own database, and its own message queues.
Tenant A: VPC-A -> ECS/EKS-A -> RDS-A -> S3 bucket-A
Tenant B: VPC-B -> ECS/EKS-B -> RDS-B -> S3 bucket-B
Tenant C: VPC-C -> ECS/EKS-C -> RDS-C -> S3 bucket-CAdvantages:
- Maximum isolation: A problem in Tenant A (load spike, data error, security breach) does not affect Tenant B.
- Complete customization: Each tenant can have its own software version, its own infrastructure configuration, even its own AWS region.
- Regulatory compliance: Required when customers demand their data resides on dedicated infrastructure (financial, healthcare, government sectors).
Disadvantages:
- High cost: The minimum cost per tenant includes at least one compute instance and one database instance. On AWS, that is $50-200/month minimum per tenant.
- Operational complexity: Each new tenant requires provisioning complete infrastructure. With 100 tenants, you have 100 RDS instances to maintain, 100 ECS clusters, 100 VPCs.
- Slow deployments: Updating the software requires deploying to N independent environments. Without robust automation, deployments become the bottleneck.
When to use it: Enterprise customers paying more than $5,000/month who demand contractual isolation. Regulated sectors that require demonstrating physical data separation.
Automation with Terraform:
# Tenant module with full isolation
module "tenant" {
source = "./modules/tenant-silo"
for_each = var.tenants
tenant_id = each.key
tenant_name = each.value.name
region = each.value.region
db_instance = each.value.db_instance_class
vpc_cidr = "10.${each.value.network_index}.0.0/16"
tags = {
TenantId = each.key
Environment = "production"
ManagedBy = "terraform"
}
}Pool Model: Fully Shared Infrastructure
All tenants share the same infrastructure. Separation is implemented exclusively at the application level, using a tenant identifier (tenant_id) in every query, every message, and every operation.
All tenants: Shared VPC -> Shared ECS -> Shared RDS -> Shared S3
|
tenant_id in every rowAdvantages:
- Minimal cost per tenant: The marginal cost of adding a new tenant is practically zero. No additional infrastructure to provision.
- Simplified operations: A single cluster, a single database, a single deployment. Updating the software affects all tenants simultaneously.
- Efficient scaling: Infrastructure is sized by total load, not by individual tenant. Resources are shared between tenants naturally.
Disadvantages:
- Noisy neighbor: A tenant generating a load spike affects performance for all others. Without per-tenant throttling, one customer can saturate the database for everyone.
- Data leakage risk: A bug in a query that forgets the WHERE tenant_id = X filter can expose one tenant’s data to another. This is the most serious risk.
- Application complexity: Every layer of the application must be tenant_id-aware. Every query, every cache key, every queue message must include it.
When to use it: Products with many tenants paying small amounts (plans up to $100/month). Early-stage startups that need to minimize infrastructure costs.
Data pattern with DynamoDB:
PK (Partition Key) | SK (Sort Key) | Data
-----------------------|----------------------|------------------
TENANT#acme | USER#u001 | {name, email...}
TENANT#acme | USER#u002 | {name, email...}
TENANT#acme | ORDER#o001 | {total, status...}
TENANT#globex | USER#u001 | {name, email...}
TENANT#globex | ORDER#o001 | {total, status...}With this design, a query with PK = TENANT#acme only returns data for that tenant. Isolation is guaranteed by the primary key design.
Pattern with PostgreSQL (Row-Level Security):
-- Enable RLS on the table
ALTER TABLE orders ENABLE ROW LEVEL SECURITY;
-- Create policy that filters by tenant_id automatically
CREATE POLICY tenant_isolation ON orders
USING (tenant_id = current_setting('app.current_tenant')::uuid);
-- In the application, before each query:
SET app.current_tenant = 'acme-uuid-here';
SELECT * FROM orders; -- Only returns orders from 'acme'Row-Level Security in PostgreSQL is the safest way to implement isolation in a pool model with relational databases. The filter is applied at the database engine level, not at the application level. Even if the application code has a bug, the database will not return data from another tenant.
Bridge Model: Selective Isolation
The bridge model combines elements of silo and pool. The compute layer is shared, but data is isolated (database per tenant or schema per tenant).
All tenants: Shared VPC -> Shared ECS -> RDS-A (tenant A)
-> RDS-B (tenant B)
-> RDS-C (tenant C)Advantages:
- Real data isolation: Each tenant has its own database. No risk of data leakage from a query bug.
- Intermediate cost: The compute layer is shared (reduces cost), but each tenant has its own database (increases isolation).
- Per-tenant backup and restore: You can restore a specific tenant’s database without affecting the others.
Disadvantages:
- Connection management: The application must maintain a connection pool for each tenant database. With 500 tenants, that is 500 connection pools.
- Schema migrations: Every schema update must run on all tenant databases. Without automation, this is an error-prone process.
- Database cost: An RDS db.t4g.micro instance costs $12.41/month. With 100 tenants, that is $1,241/month in databases alone.
When to use it: Products with medium-sized tenants ($100-5,000/month) that require data isolation but do not justify complete dedicated infrastructure.
Multi-Tenant Identity Management
Amazon Cognito as Identity Provider
Cognito enables managing user authentication and authorization within a multi-tenant context. There are two main approaches:
One User Pool per tenant (silo model):
Tenant A -> User Pool A -> App Client A
Tenant B -> User Pool B -> App Client BEach tenant has its own User Pool with its own password policies, its own identity providers (SAML, OIDC), and its own MFA configuration. Maximum isolation, but the default limit is 1,000 User Pools per AWS account.
One shared User Pool with custom attributes (pool model):
Shared User Pool -> custom:tenant_id = "acme"
-> custom:tenant_id = "globex"All users are in the same User Pool. The tenant_id is stored as a custom attribute and included in the JWT token. The application reads the tenant_id from the token and filters data accordingly.
{
"sub": "a1b2c3d4-e5f6-7890-abcd-ef1234567890",
"email": "[email protected]",
"custom:tenant_id": "acme",
"custom:role": "admin",
"iss": "https://cognito-idp.eu-west-1.amazonaws.com/eu-west-1_XXXXX",
"exp": 1694300000
}Recommendation: For most SaaS products, start with a shared User Pool. It is simpler to operate and scales to tens of thousands of tenants. Migrate to one User Pool per tenant only when customers demand federated identity providers (SAML with their own Active Directory).
Tenant-Aware API Gateway
Every request to your API must identify the tenant. The most common options:
- Subdomain:
acme.yourproduct.com,globex.yourproduct.com. The tenant is extracted from the Host header. - Path:
/api/v1/tenants/acme/orders. Explicit but more verbose. - Header:
X-Tenant-ID: acme. Clean but requires the client to send the header. - JWT token: The tenant_id is in the token. The API extracts it automatically. This is the most secure approach because the tenant_id is signed and cannot be tampered with.
Multi-Tenant Data Patterns on AWS
DynamoDB: Isolation by Partition Key
DynamoDB is ideal for multi-tenancy because the partition key defines the access scope. If each tenant has its own partition key prefix, queries are isolated by design.
Advantages for multi-tenancy:
- No fixed cost per tenant (pay per use).
- Scaling is automatic.
- Isolation by partition key is natural.
Limitation: A single DynamoDB partition supports 3,000 RCU and 1,000 WCU. If a tenant concentrates all operations on a single partition key (hot partition), it can affect performance. The solution is to use more granular partition keys.
Aurora: Isolation by Schema
Amazon Aurora allows creating a schema per tenant within the same database instance.
-- Create schema per tenant
CREATE SCHEMA tenant_acme;
CREATE SCHEMA tenant_globex;
-- Create tables in each schema
CREATE TABLE tenant_acme.orders (
id SERIAL PRIMARY KEY,
product_name VARCHAR(255),
total DECIMAL(10,2),
created_at TIMESTAMP DEFAULT NOW()
);
CREATE TABLE tenant_globex.orders (
id SERIAL PRIMARY KEY,
product_name VARCHAR(255),
total DECIMAL(10,2),
created_at TIMESTAMP DEFAULT NOW()
);Advantages: Real isolation at the database level. You can back up an individual schema. Schema migrations are applied per tenant.
Limitations: Aurora PostgreSQL supports thousands of schemas, but each schema consumes metadata. With more than 5,000 schemas, catalog operations slow down. For larger volumes, consider database-per-tenant or pool model with RLS.
S3: Isolation by Prefix and Bucket Policy
For files and objects, S3 allows isolation by prefix or by bucket:
# Isolation by prefix (pool model)
s3://my-saas-data/tenants/acme/uploads/
s3://my-saas-data/tenants/globex/uploads/
# Isolation by bucket (silo model)
s3://my-saas-acme/uploads/
s3://my-saas-globex/uploads/Prefix isolation is more economical (single bucket), but requires carefully configured bucket policies or IAM policies to prevent one tenant from accessing another tenant’s files.
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Action": ["s3:GetObject", "s3:PutObject"],
"Resource": "arn:aws:s3:::my-saas-data/tenants/${aws:PrincipalTag/TenantId}/*"
}
]
}Automated Tenant Onboarding
Each new tenant must have its infrastructure provisioned automatically. A typical flow:
Registration -> Cognito (create user) -> Step Functions -> Provision resources
|
+-- Create schema in Aurora
+-- Create prefix in S3
+-- Create entry in configuration table
+-- Configure DNS (if subdomain)
+-- Send welcome emailStep Functions for orchestration:
{
"StartAt": "CreateDatabaseSchema",
"States": {
"CreateDatabaseSchema": {
"Type": "Task",
"Resource": "arn:aws:lambda:eu-west-1:123456789:function:create-tenant-schema",
"Next": "CreateS3Prefix"
},
"CreateS3Prefix": {
"Type": "Task",
"Resource": "arn:aws:lambda:eu-west-1:123456789:function:create-tenant-s3",
"Next": "ConfigureTenantSettings"
},
"ConfigureTenantSettings": {
"Type": "Task",
"Resource": "arn:aws:lambda:eu-west-1:123456789:function:configure-tenant",
"Next": "SendWelcomeEmail"
},
"SendWelcomeEmail": {
"Type": "Task",
"Resource": "arn:aws:lambda:eu-west-1:123456789:function:send-welcome",
"End": true
}
}
}Onboarding automation is not optional. If provisioning a new tenant requires manual steps, you will not scale beyond dozens of customers without a dedicated operations team.
Real Costs by Model
| Item | Pool (100 tenants) | Bridge (100 tenants) | Silo (100 tenants) |
|---|---|---|---|
| Compute (ECS Fargate) | $200-500/month | $200-500/month | $5,000-15,000/month |
| Database (RDS/Aurora) | $50-200/month | $1,200-5,000/month | $5,000-20,000/month |
| S3 | $10-50/month | $10-50/month | $10-50/month |
| Cognito | $0/month (50K MAUs free) | $0-275/month | $0-275/month |
| Networking | $50-100/month | $50-100/month | $500-2,000/month |
| Total | $310-850/month | $1,460-5,650/month | $10,510-37,325/month |
| Cost per tenant | $3.10-8.50/month | $14.60-56.50/month | $105-373/month |
These numbers explain why most SaaS products start with a pool model and migrate to bridge or silo only for enterprise customers who pay enough to justify it.
Throttling and Fair Usage per Tenant
In a pool model, you need to protect tenants from each other. Without throttling, one tenant can consume all resources and degrade the service for everyone else.
API Gateway usage plans: Configure request limits per API key or per tenant.
Application-level throttling: Implement rate limiting in your application using Redis or DynamoDB as a counter.
import boto3
from datetime import datetime
dynamodb = boto3.resource('dynamodb')
table = dynamodb.Table('tenant-rate-limits')
def check_rate_limit(tenant_id, limit_per_minute=100):
current_minute = datetime.utcnow().strftime('%Y-%m-%dT%H:%M')
key = f"{tenant_id}#{current_minute}"
response = table.update_item(
Key={'pk': key},
UpdateExpression='ADD request_count :inc',
ExpressionAttributeValues={':inc': 1, ':limit': limit_per_minute},
ConditionExpression='attribute_not_exists(request_count) OR request_count < :limit',
ReturnValues='UPDATED_NEW'
)
return True # Within limitCommon Multi-Tenancy Mistakes
Mistake 1: Not Including tenant_id in Every Layer
If the tenant_id is only in the database but not in the cache, one tenant can see another tenant’s cached data. If it is not in message queues, a worker can process messages from one tenant with another tenant’s context.
Rule: The tenant_id must propagate through the entire call chain, including cache keys, queue messages, logs, and metrics.
Mistake 2: Starting with Silo Model Without Need
The silo model is attractive because it is conceptually simple: each tenant has everything separated. But the operational cost grows linearly with the number of tenants. We have seen SaaS products spending 80% of their engineering time managing infrastructure instead of developing product.
Mistake 3: Not Measuring Consumption per Tenant
If you do not measure resource consumption per tenant, you cannot identify noisy neighbors, you cannot bill by usage, and you cannot make informed decisions about when to move a tenant to dedicated infrastructure.
Conclusion
Multi-tenancy is not a binary decision. It is a spectrum between full isolation and fully shared resources. The correct position on that spectrum depends on your business model, the size of your customers, and their security and compliance requirements.
For most SaaS products on AWS, the practical path is to start with the pool model, implement Row-Level Security in the database and per-tenant throttling from the beginning, and offer the silo model as a premium option for enterprise customers.
If you are designing a multi-tenant architecture and need help making the right decisions from the start, our AWS consulting team has implemented these patterns in real SaaS products. You can also request a free audit to evaluate your current architecture.