DEA-C01 / Printable Review
Night Before the Exam
A compressed review sheet for AWS Certified Data Engineer - Associate. Use it to refresh decisions, not to learn the exam from scratch.
Domain Weights
| Domain | Weight | Study page |
|---|---|---|
| Data Ingestion and Transformation | 34% | Open Domain 1 |
| Data Store Management | 26% | Open Domain 2 |
| Data Operations and Support | 22% | Open Domain 3 |
| Data Security and Governance | 18% | Open Domain 4 |
Essential Services
Critical Differences
- Kinesis Data Streams vs Data Firehose
- Glue vs EMR
- Athena vs Redshift
- RDS vs Aurora vs DynamoDB
- Step Functions vs SQS/EventBridge Orchestration
Security Concepts
- Roles over long-lived keys where possible.
- Encryption does not replace authorization.
- Resource policies and identity policies can both participate in access decisions.
- CloudTrail answers API activity; Config answers configuration state.
Cost Concepts
- Match capacity model to usage pattern.
- Watch storage access frequency and lifecycle.
- Data transfer and NAT can dominate architecture cost.
- Managed services reduce operations but are not automatically cheapest.
Decision Words
- Least operational overhead: Prefer managed and serverless services when they satisfy the requirement. Exceptions appear when the scenario needs host control, unsupported runtimes, specialized network behavior, or exact migration compatibility.
- Highly available: Identify the failure boundary. One instance is not HA. Multiple instances in one AZ help capacity but not AZ failure. Multi-AZ handles regional AZ faults. Multi-Region handles regional events but adds complexity and cost.
- Durable: Durability is about preserving data. Use replication, versioning, backups, point-in-time recovery, and tested restore plans. A durable backup does not guarantee a low RTO.
- Decouple the application: Use SQS for buffering work, SNS for fanout, EventBridge for event routing, and Step Functions for visible workflow state. Add retries, DLQs, and idempotent consumers.
- Least privilege: Prefer roles and temporary credentials, scope actions/resources/conditions, watch explicit denies, and remember that resource policies may also be required.
- Most cost-effective: Read usage pattern, duration, access frequency, scaling behavior, data transfer, and operations. Cheapest unit price is not always lowest total cost.
- Lowest latency: Move content or compute closer to users, cache aggressively, choose the right database access pattern, and avoid unnecessary cross-Region or NAT paths.
- Private connectivity: Use private subnets, VPC endpoints, PrivateLink, VPN, Direct Connect, Transit Gateway, and tight DNS/routing design instead of public exposure.
- Minimum downtime: Separate deployment downtime, failure recovery, and data restore time. Use blue/green, canary, Multi-AZ, replication, and tested rollback where appropriate.
- Automatic remediation: Pair a reliable signal with EventBridge or CloudWatch, a scoped Systems Manager Automation or Lambda action, and a validation step.
- RPO and RTO: RPO is acceptable data loss. RTO is acceptable recovery time. Backups, replication, failover, and architecture all affect them differently.
- Ordered processing: Use FIFO queues or ordered stream partition keys where ordering really matters. Otherwise preserve throughput and idempotency with standard queues/events.
- Encryption and key management: Encryption protects data confidentiality. KMS key policies and IAM decide who can use keys. Authorization still needs separate design.
- Centralized governance: Use Organizations, SCPs, delegated admin, organization trails, Config aggregators, Security Hub, and account vending for multi-account control.
Common Traps
- Using Redshift for every analytics question even when S3 plus Athena fits better.
- Ignoring partitioning, file format, and small-file cost/performance issues.
- Treating ingestion and transformation as the same design decision.
- Forgetting that data quality and schema drift are operational concerns.
Worth Memorizing
- Data Ingestion and Transformation: 34%
- Data Store Management: 26%
- Data Operations and Support: 22%
- Data Security and Governance: 18%
Understand, Do Not Memorize
- Whether you can select batch or streaming ingestion patterns that match latency, ordering, replay, and scale needs.
- Whether you understand data lakes, warehouses, catalogs, lifecycle, schema evolution, and query engines.
- Whether you can operate data pipelines with monitoring, quality checks, retries, and failure handling.
- Whether you can secure and govern data using IAM, encryption, masking, catalog permissions, and audit logs.
Sources and Review Metadata
This independent training application is not affiliated with or endorsed by Amazon Web Services. AWS, Amazon Web Services, and AWS certification names are trademarks of Amazon.com, Inc. or its affiliates.