· NERVICO · cloud-architecture · 12 min read
AWS Migration: A Step-by-Step Plan with Zero Downtime
A detailed plan for migrating your infrastructure to AWS with zero downtime: migration strategies, native tools, real costs, and mistakes to avoid at each phase.
Migrating to AWS is not moving files from one server to another. It is redesigning how your application runs, scales, and recovers from failures. A poorly planned migration results in downtime, data loss, or at best, a cloud bill that triples what you were paying for your previous infrastructure.
According to a Gartner report, 60% of organizations migrating to the cloud experience significant cost overruns in the first 18 months. The primary cause is not technology: it is insufficient planning.
This article details a step-by-step migration plan, with the tools AWS provides, the real costs of each phase, and the decisions that determine whether your migration finishes in weeks or drags on for months.
Why Migrate to AWS
Real Reasons vs Marketing Reasons
AWS offers more than 200 services, global availability across 33 regions, and infrastructure that scales from a prototype to millions of users. But those are sales arguments. The real reasons to migrate are more concrete:
Variable costs vs fixed costs: With on-premise infrastructure, you pay for peak capacity. A server you only need during Black Friday traffic spikes runs all year. On AWS, you pay for what you use. If your traffic is variable, the difference can be 40-70%.
Provisioning speed: Creating an on-premise server takes weeks (ordering, shipping, mounting, configuring). On AWS, an EC2 instance is ready in 90 seconds. For startups that need to iterate fast, that difference is existential.
Managed services: On-premise, your team manages the operating system, security patches, backups, monitoring, and scaling. On AWS, services like RDS, Aurora, or DynamoDB eliminate that burden. Your team focuses on the application, not on keeping the infrastructure running.
Regulatory compliance: AWS holds SOC 2, ISO 27001, HIPAA, GDPR, and PCI DSS certifications, among others. Achieving those certifications for your own datacenter costs hundreds of thousands of euros and months of auditing.
When Migration Does Not Make Sense
Migrating to AWS is not the right answer for everyone:
- Constant, predictable workloads: If your application consumes exactly the same resources 24 hours a day, 365 days a year, a dedicated server may be more economical.
- Extreme latency requirements: High-frequency trading applications or real-time signal processing may need dedicated hardware in specific locations.
- Specialized hardware dependencies: Specific GPUs, FPGAs, or proprietary hardware that AWS does not offer.
- Regulatory restrictions: Some sectors (defense, government) have data localization requirements that limit cloud options.
The 7 Migration Strategies (The 7 Rs)
AWS defines seven migration strategies. The right choice depends on each application, not on the organization as a whole. It is normal to use different strategies for different components of the same system.
Rehost (Lift and Shift)
Move the application as-is to AWS, without code changes. Copy the virtual machine to an EC2 instance, the database to RDS, and the files to S3.
When to use it: When you need to migrate fast, have dozens of servers, and do not have the budget to refactor.
Primary tool: AWS Application Migration Service (MGN). It replicates your servers in real time to AWS. When you are ready, you perform the cutover: traffic redirects to AWS with minimal downtime of minutes.
# Install the replication agent on the source server
wget -O ./aws-replication-installer-init \
https://aws-application-migration-service-us-east-1.s3.amazonaws.com/latest/linux/aws-replication-installer-init
chmod +x aws-replication-installer-init
sudo ./aws-replication-installer-init \
--region us-east-1 \
--aws-access-key-id AKIAIOSFODNN7EXAMPLE \
--aws-secret-access-key wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEYAdvantage: Speed. You can migrate hundreds of servers in weeks.
Disadvantage: You do not leverage AWS-native services. You are paying for EC2 the same amount (or more) than you were paying in your datacenter.
Replatform (Lift, Tinker, and Shift)
Move the application with minor adjustments to take advantage of managed services. For example: the application moves to EC2, but the MySQL database migrates to RDS MySQL.
When to use it: When you want to reduce operational overhead without rewriting code.
Practical example: Your application uses PostgreSQL on a dedicated server. Instead of installing PostgreSQL on EC2, you use RDS PostgreSQL. The application code does not change (same connection string, same driver), but now AWS manages backups, engine updates, and high availability with Multi-AZ.
Refactor (Re-Architect)
Redesign the application to leverage AWS-native services. Convert a monolith to microservices, migrate from a relational database to DynamoDB, or replace a message queue server with SQS.
When to use it: When the application has scalability problems that the current architecture cannot solve, or when the costs of operating the current architecture are unsustainable.
Practical example: A monolithic Java application that processes images. Instead of scaling vertically (more CPU, more RAM), you extract image processing to Lambda functions that scale horizontally automatically. The cost goes from one m5.4xlarge instance ($615/month) to paying only for actual invocations.
Repurchase
Replace the existing application with a SaaS product. Migrate from a self-hosted mail server to Amazon WorkMail or from an on-premise CRM to Salesforce.
Retire
Identify applications that are no longer used and shut them down. In migration audits, it is common to find that 10% to 20% of servers run obsolete applications.
Retain
Keep the application in its current location. Not everything needs to migrate to AWS. Some legacy applications with specific hardware dependencies or short-term decommission plans do not justify the migration cost.
Relocate
Move VMware virtual machines directly to VMware Cloud on AWS. Useful for organizations with large VMware investments that want AWS benefits without changing the virtualization platform.
Step-by-Step Migration Plan
Phase 1: Assessment and Discovery (2-4 Weeks)
Before migrating a single server, you need to understand what you have. In most organizations, the server inventory is incomplete or outdated.
AWS Application Discovery Service scans your network and catalogs servers, applications, dependencies, and communication patterns. Two modes are available:
- Agent-based: Installs an agent on each server. Captures processes, network connections, CPU/memory/disk usage. More detailed data but requires access to each machine.
- Agentless: Scans the network from a virtual appliance. Captures basic information about VMware servers without installing anything. Less detail, faster.
Deliverable from this phase: A dependency map showing which applications communicate with each other, which ports they use, and which databases they query. Without this map, you will migrate an application and discover in production that it depended on a service you left in the datacenter.
# Export discovery data for analysis
aws discovery describe-agents --region us-east-1
# Generate dependency report
aws discovery list-configurations \
--configuration-type SERVER \
--region us-east-1Common mistakes in this phase:
- Ignoring hidden dependencies: An application that calls an internal SOAP service that nobody documented.
- Not measuring bandwidth: If you transfer 50 TB of data to AWS over the internet, it will take weeks. For large volumes, consider AWS Snowball.
- Underestimating licenses: Oracle, SQL Server, or SAP licenses may have different terms in the cloud.
Phase 2: Planning and Design (2-3 Weeks)
With the complete inventory, classify each application by the appropriate migration strategy. A practical approach:
| Criteria | Rehost | Replatform | Refactor |
|---|---|---|---|
| Code changes | None | Minimal | Significant |
| Migration time | Days | Weeks | Months |
| Cost reduction | 10-20% | 20-40% | 40-70% |
| Risk | Low | Medium | High |
| Long-term benefit | Low | Medium | High |
Network design: Define the VPC structure before migrating:
# Terraform: VPC for migration
resource "aws_vpc" "main" {
cidr_block = "10.0.0.0/16"
enable_dns_hostnames = true
enable_dns_support = true
tags = {
Name = "migration-vpc"
Environment = "production"
}
}
# Public subnets for load balancers
resource "aws_subnet" "public" {
count = 3
vpc_id = aws_vpc.main.id
cidr_block = "10.0.${count.index}.0/24"
availability_zone = data.aws_availability_zones.available.names[count.index]
tags = {
Name = "public-${count.index}"
Tier = "public"
}
}
# Private subnets for applications
resource "aws_subnet" "private" {
count = 3
vpc_id = aws_vpc.main.id
cidr_block = "10.0.${count.index + 10}.0/24"
availability_zone = data.aws_availability_zones.available.names[count.index]
tags = {
Name = "private-${count.index}"
Tier = "private"
}
}Hybrid connectivity: During the migration, your infrastructure will be split between on-premise and AWS. You need reliable connectivity:
- AWS Site-to-Site VPN: Encrypted connection over the internet. Up to 1.25 Gbps per VPN tunnel. Cost: $0.05/hour per connection.
- AWS Direct Connect: Dedicated connection between your datacenter and AWS. 1 Gbps or 10 Gbps. Predictable, consistent latency. Cost: from $0.30/hour for 1 Gbps.
For most migrations, VPN is sufficient. Direct Connect is necessary when transferring large data volumes or when you need low, consistent latency.
Phase 3: Pilot Migration (1-2 Weeks)
Do not migrate everything at once. Select a non-critical application as a pilot. Ideally:
- An application with few external dependencies
- One that is not business-critical
- One with an available team to validate the result
- One that is representative of most applications in your portfolio
Pilot process:
- Replication: Configure AWS MGN to replicate the server in real time.
- Testing: Launch a test instance on AWS. Run functional, performance, and security tests.
- Test cutover: Simulate the cutover by redirecting test traffic to the AWS instance.
- Validation: Verify the application works correctly: correct HTTP responses, acceptable response times, operational integrations.
- Rollback: Verify you can return to the original server in under 15 minutes.
Pilot metrics: The pilot is successful if it meets these criteria:
- Cutover time under 30 minutes
- No data loss during replication
- Performance equal to or better than the original
- Rollback verified and documented
Phase 4: Wave Migration (4-12 Weeks)
Group applications into waves of 5-15 applications. Grouping criteria:
- Applications sharing the same database migrate together.
- Applications with mutual dependencies migrate in the same wave.
- Critical applications migrate in later waves, when the team has experience.
Zero-downtime cutover pattern:
1. Continuous replication (MGN keeps the replica synchronized)
2. Change freeze (code freeze)
3. Synchronization verification (replication lag = 0)
4. DNS cutover (low TTL configured 48 hours prior)
5. Functional validation on AWS
6. Intensive monitoring (24-72 hours)
7. Decommission source server (after grace period)DNS and zero downtime: The trick to achieving zero downtime is in the DNS. Before the migration:
- Reduce DNS record TTL to 60 seconds (do this 48 hours before so caches update).
- During the cutover, update the DNS record to point to the new IP on AWS.
- Clients with cached DNS will continue reaching the original server for a maximum of 60 seconds. The original server remains active and redirects traffic or responds normally.
# Route 53: update record with low TTL
aws route53 change-resource-record-sets \
--hosted-zone-id Z1234567890 \
--change-batch '{
"Changes": [{
"Action": "UPSERT",
"ResourceRecordSet": {
"Name": "app.yourdomain.com",
"Type": "A",
"TTL": 60,
"ResourceRecords": [{"Value": "3.120.45.67"}]
}
}]
}'Phase 5: Post-Migration Optimization (Ongoing)
The migration does not end when you shut down the last on-premise server. The first weeks on AWS are the most expensive because instances are over-provisioned (nobody wants the application to fail right after migration).
Right-sizing: Use AWS Compute Optimizer to analyze actual CPU, memory, and network usage of your instances. After 14 days of data, Compute Optimizer recommends the optimal instance type.
Reserved Instances and Savings Plans: Once your consumption is stable (3-6 months after migration), purchase reserved capacity. Savings are 30-60% compared to on-demand pricing.
Managed services: Identify components you can migrate to native services. An NGINX server acting as a reverse proxy can be replaced by ALB. A Redis installation can migrate to ElastiCache. Each managed service you adopt reduces your team’s operational burden.
Real Migration Costs
Direct Costs
| Item | Typical Range |
|---|---|
| AWS MGN (replication) | Free for 90 days per server |
| Replica storage (EBS) | $0.10/GB/month |
| Data transfer (ingress) | Free |
| Test instances during migration | Variable (same cost as production) |
| AWS Direct Connect (if needed) | $0.30/hour + port fees |
Indirect Costs
Indirect costs are usually larger than direct costs:
- Dedicated team: One senior engineer full-time during the migration. In Europe, EUR 6,000-10,000/month.
- Training: The team needs to learn AWS. Courses, certifications, learning time. Budget EUR 2,000-5,000 per person.
- Dual running: During the migration, you pay for on-premise infrastructure and AWS simultaneously. This phase can last 2 to 6 months.
- Licenses: Some software licenses (Oracle, SQL Server) have different costs in the cloud. Verify before migrating.
Total Cost Example
For a startup with 10 servers, 5 TB of data, and a 3-month migration:
| Item | Estimated Cost |
|---|---|
| AWS infrastructure (3 months) | EUR 3,000-6,000 |
| On-premise infrastructure (overlap) | EUR 2,000-4,000 |
| Dedicated team (50% of 1 engineer) | EUR 9,000-15,000 |
| Training | EUR 4,000-10,000 |
| Total | EUR 18,000-35,000 |
Mistakes We Have Seen in Real Migrations
Mistake 1: Migrating Without a Dependency Map
A company migrated its main application to AWS. It worked correctly in tests. In production, the application failed intermittently. The cause: an internal microservice that queried a legacy database still in the datacenter. The latency between AWS and the datacenter (40ms) caused timeouts on queries that previously responded in 2ms.
Mistake 2: Not Adjusting Instance Types
Migrating a server with 32 GB of RAM that uses a maximum of 4 GB. On-premise, the server was already purchased. On AWS, you pay for every GB every hour. Over-provisioning instances is the biggest waste of money on AWS.
Mistake 3: Ignoring Data Transfer Costs
AWS charges for outbound data transfer. If your application serves 10 TB of data per month, you will pay approximately $900 in transfer alone. This cost does not exist on-premise and surprises many teams after the first invoice.
Mistake 4: No Rollback Plan
If the cutover fails, you need to return to the original server in minutes, not hours. Every migration needs a documented and tested rollback plan before the cutover.
AWS Ecosystem Tools for Migration
| Tool | Function | When to Use |
|---|---|---|
| AWS MGN | Server replication and cutover | Rehosting any server |
| AWS DMS | Database migration | Migrating to RDS, Aurora, DynamoDB |
| AWS SCT | Schema conversion | Changing database engines (Oracle to PostgreSQL) |
| AWS DataSync | File transfer | Migrating data to S3, EFS, FSx |
| AWS Snowball | Offline transfer | Volumes greater than 10 TB |
| AWS Transfer Family | Managed SFTP/FTP | Maintaining existing SFTP integrations |
AWS Database Migration Service (DMS)
DMS deserves special mention because database migration is the most delicate part of any migration. DMS continuously replicates data from the source database to the target, including changes that occur during the migration (CDC, Change Data Capture).
# Create DMS replication task
aws dms create-replication-task \
--replication-task-identifier my-migration-task \
--source-endpoint-arn arn:aws:dms:us-east-1:123456789:endpoint:source \
--target-endpoint-arn arn:aws:dms:us-east-1:123456789:endpoint:target \
--migration-type full-load-and-cdc \
--replication-instance-arn arn:aws:dms:us-east-1:123456789:rep:instance \
--table-mappings file://table-mappings.jsonSupported engines: MySQL, PostgreSQL, Oracle, SQL Server, MariaDB, MongoDB, Amazon Aurora. DMS even allows engine changes during migration (for example, from Oracle to PostgreSQL) using AWS Schema Conversion Tool (SCT) to adapt the schema and queries.
Migration Checklist
Before starting each migration wave, verify:
- Updated dependency map for all applications in the wave
- VPC and subnets configured with restrictive security groups
- Connectivity verified between AWS and datacenter (VPN or Direct Connect)
- DNS TTL reduced to 60 seconds (48 hours before cutover)
- Rollback plan documented and tested for each application
- CloudWatch monitoring alerts configured
- On-call team assigned for the first 72 hours post-cutover
- Backups verified before cutover
- User communication about maintenance window (if applicable)
Conclusion
A well-executed AWS migration reduces operational costs, improves scalability, and frees your team from managing physical infrastructure. But the benefits are not automatic. They require planning, disciplined execution, and continuous optimization.
The most frequent mistake is not technical. It is trying to migrate everything at once, without a complete inventory, without rollback tests, and without a team with AWS experience.
If you are planning a migration and need a team that has executed real migrations, with documented errors and lessons learned, contact our AWS consulting team. We also offer a free infrastructure audit to evaluate the current state of your environment and define the optimal migration strategy.