Analytics / Machine Learning
RecognitionAmazon SageMaker AI
Data preparation, lineage, and governed project surfaces used by data engineering workflows.
Key points
- SageMaker AI contributes preparation, lineage, and governed project surfaces to data-engineering workflows.
- Keep the data-engineer scope on datasets, processing artifacts, lineage, and access rather than model training or inference science.
Best-known use cases
- Prepare and transform datasets for machine-learning training.
- Track datasets and processing artifacts across governed ML workflows.
What candidates often confuse it with
- SageMaker AI project and lineage surfaces complement Glue catalog and ETL capabilities rather than replacing them.
Key takeaway
Use SageMaker AI when governed ML data preparation or artifact lineage is part of the pipeline boundary.
Relevant exam tasks
- D2.4 — Task 2.4: Design data models and schema evolution
- 2.4.4 — Establish data lineage by using AWS tools (for example, Amazon SageMaker ML Lineage Tracking and Amazon SageMaker Catalog).
- D3.1 — Task 3.1: Automate data processing by using AWS services
- 3.1.6 — Prepare data for transformation (for example, AWS Glue DataBrew and Amazon SageMaker Unified Studio).
- D3.2 — Task 3.2: Analyze data by using AWS services
- 3.2.2 — Verify and clean data (for example, Lambda, Athena, QuickSight, Jupyter Notebooks, Amazon SageMaker Data Wrangler).
- D4.1 — Task 4.1: Apply authentication mechanisms
- 4.1.7 — Use domain, domain units, and projects for SageMaker Unified Studio.
- D4.5 — Task 4.5: Understand data privacy and governance
- 4.5.6 — Manage data access through Amazon SageMaker Catalog projects.