DEA-C01 / Printable Review
Night Before the Exam
A compressed review sheet for AWS Certified Data Engineer - Associate. Use it to refresh decisions, not to learn the exam from scratch.
Domain Weights
| Domain | Weight | Study page |
|---|---|---|
| Data Ingestion and Transformation | 34% | Open Domain 1 |
| Data Store Management | 26% | Open Domain 2 |
| Data Operations and Support | 22% | Open Domain 3 |
| Data Security and Governance | 18% | Open Domain 4 |
Essential Services
Critical Differences
- Kinesis Data Streams vs Data Firehose
- Glue vs EMR
- Athena vs Redshift
- RDS vs Aurora vs DynamoDB
- Step Functions vs SQS/EventBridge Orchestration
Security Concepts
- Roles over long-lived keys where possible.
- Encryption does not replace authorization.
- Resource policies and identity policies can both participate in access decisions.
- CloudTrail answers API activity; Config answers configuration state.
Cost Concepts
- Match capacity model to usage pattern.
- Watch storage access frequency and lifecycle.
- Data transfer and NAT can dominate architecture cost.
- Managed services reduce operations but are not automatically cheapest.
Decision Words
- Least operational overhead: Prefer managed and serverless services when they satisfy the requirement. Exceptions appear when the scenario needs host control, unsupported runtimes, specialized network behavior, or exact migration compatibility.
- Highly available: Identify the failure boundary. One instance is not HA. Multiple instances in one AZ help capacity but not AZ failure. Multi-AZ handles regional AZ faults. Multi-Region handles regional events but adds complexity and cost.
- Durable: Durability is about preserving data. Use replication, versioning, backups, point-in-time recovery, and tested restore plans. A durable backup does not guarantee a low RTO.
- Decouple the application: Use SQS for buffering work, SNS for fanout, EventBridge for event routing, and Step Functions for visible workflow state. Add retries, DLQs, and idempotent consumers.
- Least privilege: Prefer roles and temporary credentials, scope actions/resources/conditions, watch explicit denies, and remember that resource policies may also be required.
- Most cost-effective: Read usage pattern, duration, access frequency, scaling behavior, data transfer, and operations. Cheapest unit price is not always lowest total cost.
- Lowest latency: Move content or compute closer to users, cache aggressively, choose the right database access pattern, and avoid unnecessary cross-Region or NAT paths.
- Private connectivity: Use private subnets, VPC endpoints, PrivateLink, VPN, Direct Connect, Transit Gateway, and tight DNS/routing design instead of public exposure.
- Minimum downtime: Separate deployment downtime, failure recovery, and data restore time. Use blue/green, canary, Multi-AZ, replication, and tested rollback where appropriate.
- Automatic remediation: Pair a reliable signal with EventBridge or CloudWatch, a scoped Systems Manager Automation or Lambda action, and a validation step.
- RPO and RTO: RPO is acceptable data loss. RTO is acceptable recovery time. Backups, replication, failover, and architecture all affect them differently.
- Ordered processing: Use FIFO queues or ordered stream partition keys where ordering really matters. Otherwise preserve throughput and idempotency with standard queues/events.
- Encryption and key management: Encryption protects data confidentiality. KMS key policies and IAM decide who can use keys. Authorization still needs separate design.
- Centralized governance: Use Organizations, SCPs, delegated admin, organization trails, Config aggregators, Security Hub, and account vending for multi-account control.
Common Traps
- Using Redshift for every analytics question even when S3 plus Athena fits better.
- Ignoring partitioning, file format, and small-file cost/performance issues.
- Treating ingestion and transformation as the same design decision.
- Forgetting that data quality and schema drift are operational concerns.
Worth Memorizing
- Data Ingestion and Transformation: 34%
- Data Store Management: 26%
- Data Operations and Support: 22%
- Data Security and Governance: 18%
Understand, Do Not Memorize
- Whether you can select batch or streaming ingestion patterns that match latency, ordering, replay, and scale needs.
- Whether you understand data lakes, warehouses, catalogs, lifecycle, schema evolution, and query engines.
- Whether you can operate data pipelines with monitoring, quality checks, retries, and failure handling.
- Whether you can secure and govern data using IAM, encryption, masking, catalog permissions, and audit logs.
DJames617