Data Engineer
Build data pipelines and stores with reliable ingestion, quality controls, governance and operational visibility.
Who this path is for
Know SQL, data formats and basic data-pipeline concepts.
Study each topic, complete the applied exercise, and keep evidence of what you learned. These are original study milestones, not a claim to reproduce the full exam blueprint. Use the linked AWS exam guide to check all in-scope domains before booking.
Your learning roadmap
01Ingest and transform
3 focus areas · applied exercise
Ingest and transform
3 focus areas · applied exercise- Batch and streaming
- Glue ETL
- Schema evolution
Design ingestion for a daily order export with late-arriving corrections.
Specify the deduplication key and bad-record handling.
02Store and query
3 focus areas · applied exercise
Store and query
3 focus areas · applied exercise- S3 and Parquet
- Glue Catalog
- Athena and warehouse choices
Complete the Athena/Iceberg order analytics lab.
Prove the corrected value and unchanged row count.
03Operate the pipeline
3 focus areas · applied exercise
Operate the pipeline
3 focus areas · applied exercise- Orchestration
- Data quality
- Monitoring
Define completeness and freshness checks for each pipeline stage.
Describe how a failed load is detected and replayed.
04Secure and govern data
3 focus areas · applied exercise
Secure and govern data
3 focus areas · applied exercise- Lake Formation
- KMS
- Data lifecycle
Create a data-access matrix for raw and curated datasets.
Explain retention, ownership and who can publish curated data.
Put the learning into practice
Related labs cover selected skills. Completing them does not establish exam readiness.
