AWS Service
AWS Glue
Glue is a serverless data integration service for cataloging, ETL, crawlers, and jobs.
What This Service Solves
Glue is a serverless data integration service for cataloging, ETL, crawlers, and jobs.
- Use Glue for serverless ETL, schema discovery, data cataloging, and scheduled data transformations.
When You Should Not Use It
- Avoid Glue when you need full cluster control and a managed Spark/Hadoop cluster such as EMR fits better.
What AWS Manages and What You Manage
AWS manages serverless job infrastructure. You manage jobs, scripts, crawlers, catalog databases, permissions, and data quality.
Security Implications
Use IAM, Lake Formation integration, KMS, network connections, and careful access to source data.
Availability and Scaling
Job retries and bookmarks help, but bad data and schema drift still need handling.
Worker type, worker count, partitioning, and file sizes drive performance.
Cost Behavior
DPU time, crawlers, data quality, and downstream storage/query costs matter.
Common Integrations
- S3
- Glue Data Catalog
- Athena
- Redshift
- Lake Formation
- Step Functions
How AWS Might Present It
Certification-Specific Depth
Data Engineer Associate
Know configuration choices, integrations, failure modes, security, operations, and cost tradeoffs.
Machine Learning Engineer Associate
Know configuration choices, integrations, failure modes, security, operations, and cost tradeoffs.
Related Comparisons
Sources and Review Metadata
This independent training application is not affiliated with or endorsed by Amazon Web Services. AWS, Amazon Web Services, and AWS certification names are trademarks of Amazon.com, Inc. or its affiliates.