AWS Databases
AWS database services are managed offerings covering relational, key-value, document, graph, time-series, and in-memory workloads. They handle provisioning, patching, backups, and replication so teams can focus on schema design and query patterns rather than infrastructure.
itCloud computing | OpenSkills.info
Course pathWalk it in order
Look it upDip in anytime
Go furtherLeaves this page
Don't Panic
Don't Panic — AWS Databases
AWS Databases is a collection of managed database services, not a very large database wearing several hats. That is good news, because applications do not all ask the same questions. It is also bad news, because the service list is happy to look like a buffet when what you needed was a map.
Start with the workload. A data model is the way records and relationships are represented. Tables and joins point toward relational services such as Amazon RDS or Amazon Aurora. Known lookups by a key point toward Amazon DynamoDB. Relationship traversal points toward Amazon Neptune. Measurements arranged around time point toward Amazon Timestream. The product name comes after that evidence, not before it, which saves an impressive amount of confident wandering.
The surprise is that managed does not mean unattended. AWS takes on infrastructure work, but the application still owns schema choices, queries, identities, network paths, connections, recovery behavior, and the awkward moment when a real client meets a failover. An access pattern, the specific way an application reads or writes data, is therefore more useful than a vague requirement for "scale." It gives the database something concrete to be good at.
Keep three ideas in separate boxes. Availability keeps a service reachable through a defined failure. Scale adds useful capacity. Recovery restores service and data after damage. An RDS Multi-AZ standby supports availability, but it does not serve read traffic. A read replica can add read capacity, but it can lag. Backups support recovery, but only a restore rehearsal tells you whether the application returns in time. These are not interchangeable talismans, however much a diagram may wish otherwise.
More than one database can be the sensible answer. Orders might need relational transactions, while predictable session lookups fit a key-value store and a cache serves hot reusable reads. That arrangement is called polyglot persistence. It also creates another schema, security boundary, monitor, backup path, and failure mode for every new service. The sensible choice is not the most exotic service; it is the one whose model and failure behavior match the evidence.
Read the Intro for the portfolio and the decision questions. Use the Slides when the service relationships need a compact map. Keep the Cheatsheet nearby for the selection matrix, recovery distinctions, and monitoring signals. Then use the Practice Reference and exercise to turn a fictional workload into a decision record before any console clicks acquire the confidence of a plan.
Where this skill leads
Relevant careers
See how this topic contributes to broader role-level skill maps.
Sources
- https://docs.aws.amazon.com/databases-on-aws-how-to-choose/
Supports
- AWS provides relational, key-value, document, in-memory, graph, time-series, wide-column, and vector database options
- Relational services support structured data, joins, and flexible query patterns
- DynamoDB, DocumentDB, Neptune, Timestream, Keyspaces, ElastiCache, and MemoryDB map to distinct workload and data-model needs
- One application can combine best-fit database types
- Amazon Redshift targets analytical data-warehouse workloads rather than the default OLTP path
- Serverless database options scale capacity and use pay-for-use models
- https://aws.amazon.com/products/databases/
Supports
- AWS maintains a portfolio of managed and purpose-built database services
- AWS database services provide service-specific availability, security, scaling, and vector capabilities
- Managed database services reduce common database infrastructure administration
- https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/Welcome.html
Supports
- RDS supports Db2, MariaDB, Microsoft SQL Server, MySQL, Oracle Database, and PostgreSQL
- RDS manages common tasks including installation, patching, backups, failure detection, and recovery
- Customers retain responsibility for database design, query tuning, monitoring, identities, and workload-specific behavior
- RDS supports Multi-AZ deployments, read replicas, VPC placement, CloudWatch monitoring, and several storage and billing models
- https://docs.aws.amazon.com/AmazonRDS/latest/AuroraUserGuide/CHAP_AuroraOverview.html
Supports
- Aurora is a managed relational database compatible with MySQL and PostgreSQL
- Aurora clusters separate shared storage from database compute
- Aurora provides cluster endpoints for connection roles and service-specific scaling and availability features
- https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/Introduction.html
Supports
- DynamoDB is a serverless fully managed NoSQL database
- DynamoDB supports key-value and document data models
- Primary key design controls item identification and data distribution
- DynamoDB provides on-demand and provisioned capacity modes, consistency options, backups, and global tables
- https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/Concepts.MultiAZSingleStandby.html
Supports
- A single-standby RDS Multi-AZ DB instance deployment uses synchronous replication to another Availability Zone
- RDS can fail over to the standby for high availability
- A single standby does not serve read traffic and is not a read-scaling mechanism
- Synchronous replication can increase write and commit latency
- https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/USER_ReadRepl.html
Supports
- RDS DB instance read replicas are read-only copies used to reduce read load
- RDS DB instance read replicas receive primary changes asynchronously and can expose stale data
- Read replicas and Multi-AZ standbys have different purposes and replication behavior
- Read replicas can be promoted and can support cross-Region recovery patterns
- https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/USER_WorkingWithAutomatedBackups.html
Supports
- RDS creates automated backups during the backup window and retains them for the configured period
- RDS supports point-in-time recovery within the backup retention period
- Manual snapshots have a lifecycle independent from automated backups
- Backup deletion and retention behavior depends on database deletion choices
- https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/Overview.Encryption.html
Supports
- RDS encryption at rest covers storage, logs, automated backups, read replicas, and snapshots
- RDS uses AWS KMS keys and supports AWS-managed or customer-managed keys
- Disabling or losing access to a KMS key can make encrypted RDS resources inaccessible
- Key choice and key policy belong in database protection and recovery planning
- https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/MonitoringOverview.html
Supports
- RDS integrates with CloudWatch metrics and alarms
- RDS provides Performance Insights and Enhanced Monitoring
- RDS monitoring includes connections, read and write operations, storage, CPU, memory, network traffic, logs, and events
- Monitoring supports reliability, availability, and performance work but requires workload-specific interpretation
- https://docs.aws.amazon.com/whitepapers/latest/aws-risk-and-compliance/shared-responsibility-model.html
Supports
- AWS and customers share security and compliance responsibilities
- AWS operates the infrastructure from the host operating system and virtualization layer through physical facilities
- Customer responsibilities vary with the selected service and include configured controls, applications, data, and compliance requirements
- https://docs.aws.amazon.com/dms/latest/userguide/Welcome.html
Supports
- AWS DMS discovers sources, converts schemas, and migrates relational, warehouse, NoSQL, and other data stores
- DMS supports one-time migration and ongoing replication
- DMS Fleet Advisor builds an inventory for migration planning
- DMS migrations use source and target endpoints and replication tasks
- https://docs.aws.amazon.com/dms/latest/userguide/schema-conversion.html
Supports
- DMS Schema Conversion assesses heterogeneous migration complexity
- DMS Schema Conversion converts schemas and many database code objects to a target-compatible format
- Conversion reports identify objects that require manual action
- Converted code can be reviewed, saved, edited, or applied to the target
- https://aws.amazon.com/about-aws/whats-new/2009/10/27/introducing-amazon-relational-database-service/
Supports
- Amazon RDS was introduced in October 2009 as a managed relational database service
- https://aws.amazon.com/about-aws/whats-new/2012/01/18/aws-announces-dynamodb/
Supports
- Amazon DynamoDB was announced in January 2012
- https://docs.aws.amazon.com/whitepapers/latest/data-warehousing-on-aws/introducing-amazon-redshift.html
Supports
- Amazon Redshift launched in February 2013 as an AWS data warehouse
- https://aws.amazon.com/blogs/aws/highly-scalable-mysql-compat-rds-db-engine/
Supports
- Amazon Aurora was announced in November 2014 as a MySQL-compatible relational service
- https://aws.amazon.com/blogs/aws/aws-database-migration-service/
Supports
- AWS Database Migration Service was introduced in October 2015
- https://aws.amazon.com/blogs/aws/amazon-neptune-fast-reliable-graph-database-built-for-the-cloud/
Supports
- Amazon Neptune became generally available in May 2018
- https://aws.amazon.com/about-aws/whats-new/2019/01/amazon-documentdb-with-mongodb-compatibility-generally-available/
Supports
- Amazon DocumentDB became generally available in January 2019
- https://aws.amazon.com/about-aws/whats-new/2020/09/amazon-timestream-now-generally-available/
Supports
- Amazon Timestream became generally available in September 2020
- https://aws.amazon.com/about-aws/whats-new/2021/08/amazon-memorydb-redis/
Supports
- Amazon MemoryDB was announced in August 2021 as a durable in-memory database service
- https://docs.aws.amazon.com/dms/latest/userguide/CHAP_BestPractices.html
Supports
- DMS creates target tables and primary keys but not secondary indexes, foreign keys, users, privileges, stored procedures, or triggers
- For a full-load-plus-CDC task: drop secondary indexes, referential integrity constraints, and DML triggers for the full load, add secondary indexes before the CDC phase, and enable triggers right before cutover
- Pausing the task at the end of full load allows building indexes and constraints before CDC resumes
- Turn off backups and Multi-AZ on an RDS target until ready to cut over
- DMS loads eight tables in parallel by default, controlled by the MaxFullLoadSubTasks task setting
- A replication instance failover during full load fails the task, which can be restarted from the point of failure for the tables that did not complete
- Use multiple tasks only when the table sets do not share transactions; transactional consistency is preserved within a task, not across tasks
- AWS recommends a Multi-AZ replication instance when using DMS for ongoing replication
- Limited LOB mode is the default and truncates LOB values above the configured maximum size
- https://docs.aws.amazon.com/dms/latest/userguide/CHAP_Tasks.Creating.html
Supports
- The migration type "Migrate existing data and replicate ongoing changes" is called full-load-and-cdc in the API
- "Migrate existing data" is Full Load and "Replicate data changes only" is the CDC-only option
- A premigration assessment warns of issues such as unsupported data types before the task runs
- Table mapping selection rules choose the schemas and tables in scope
- https://docs.aws.amazon.com/dms/latest/userguide/CHAP_Task.CDC.html
Supports
- DMS CDC reads the source engine's transaction log using engine-native APIs and does not provide real-time replication or a latency SLA
- MySQL sources need row-based binary logging, PostgreSQL needs a logical replication slot with test_decoding, Oracle needs supplemental logging, SQL Server needs MS-CDC or MS-Replication
- For an RDS source, automated backups must be enabled and change logs retained long enough for the full load to complete
- A CDC task can start from a custom start time or a native start point and can be stopped at a source commit time or a replication-instance server time
- Bidirectional replication runs a second CDC task from the new target back to the original source, with LoopbackPreventionSettings.EnableLoopbackPrevention true and BatchApplyEnabled false, giving a post-cutover fallback path
- DMS bidirectional replication has no conflict detection or resolution
- https://docs.aws.amazon.com/dms/latest/userguide/CHAP_Validating.html
Supports
- Setting EnableValidation true in ValidationSettings compares source and target row by row after full load and re-checks rows as CDC changes them
- Data validation requires a primary key or unique index on each table and does not validate views
- Validation mismatches are written to the awsdms_control.awsdms_validation_failures_v1 table on the target with a FAILURE_TYPE of RECORD_DIFF, MISSING_SOURCE, MISSING_TARGET, or TABLE_WARNING
- A per-table ValidationState of Validated means all rows matched; other states include Mismatched records and No primary key
- A full-load validation-only task gives a fast count of mismatched rows just before or during a cutover
- https://docs.aws.amazon.com/dms/latest/userguide/CHAP_Monitoring.html
Supports
- CDCLatencySource is the gap in seconds between the last event captured from the source and the replication instance clock
- CDCLatencyTarget is the gap in seconds between the oldest change waiting to commit on the target and the replication instance clock
- CDCLatencySource resets to zero when there are no new source events, and latency spikes are expected during heavy source activity
- CDCIncomingChanges is the number of change events waiting to be applied to the target
- A high CDCLatencyTarget with a low CDCLatencySource points at missing target indexes or target-side resource or network bottlenecks
- https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/USER_PIT.html
Supports
- RDS point-in-time recovery creates a new DB instance and never modifies the source
- You can restore to any point within the backup retention period, and RDS uploads transaction logs to S3 every five minutes
- LatestRestorableTime from describe-db-instances shows how current a restore can be
- The AWS CLI command is restore-db-instance-to-point-in-time with --source-db-instance-identifier, --target-db-instance-identifier, and either --restore-time or --use-latest-restorable-time
- Restored instances associate with the default DB parameter and option groups unless custom groups are passed at restore time; the console allows changing them after the restore
- You can choose the default VPC security group or apply a custom one
- Restore to the same or similar instance size, and matching Provisioned IOPS, or the restore can fail
- Allocated storage cannot be reduced on a restore, an increase must be at least 10 percent, and allocated storage cannot be increased at all when restoring RDS for SQL Server
- Restored instances can have their parameter group and option group modified after the restore completes
- After the instance is available, RDS keeps loading storage blocks from S3 in the background, shown as StorageOperationStatus Initializing in describe-db-instances
- For SQL Server, databases are restored within one second of each other and cross-database transactions may be inconsistent
- https://docs.aws.amazon.com/AmazonRDS/latest/AuroraUserGuide/aurora-pitr.html
Supports
- Aurora restore-db-cluster-to-point-in-time creates a new DB cluster and does not modify the source
- Aurora uploads log records for a DB cluster to Amazon S3 continuously
- EarliestRestorableTime and LatestRestorableTime from describe-db-clusters bound the restore window
- Restoring a cluster with the AWS CLI or RDS API requires a separate create-db-instance call with --db-cluster-identifier to create the primary writer instance; the console creates it automatically
- Restored clusters associate with the default DB cluster and DB parameter groups unless custom groups are specified during the restore
- The restored cluster keeps the source cluster's backup retention period
