Analytics
CoreAWS Glue DataBrew
Visual profiling and repeatable preparation of inconsistent or incomplete datasets.
Key points
- DataBrew provides visual profiling and reusable preparation recipes for cleaning data without transformation code.
- Recipes make preparation repeatable, but the resulting data still needs validation and governed publication.
Best-known use cases
- Visually profile and clean datasets without writing transformation code.
- Build reusable preparation recipes for inconsistent or missing values.
What candidates often confuse it with
- DataBrew emphasizes visual preparation; Glue ETL is the broader code-based serverless transformation surface.
Key takeaway
Choose DataBrew for repeatable visual profiling and cleaning of inconsistent datasets.
Relevant exam tasks
- D3.1 — Task 3.1: Automate data processing by using AWS services
- 3.1.6 — Prepare data for transformation (for example, AWS Glue DataBrew and Amazon SageMaker Unified Studio).
- D3.2 — Task 3.2: Analyze data by using AWS services
- 3.2.1 — Visualize data by using AWS services and tools (for example, DataBrew, Amazon QuickSight).
- 3.2.2 — Verify and clean data (for example, Lambda, Athena, QuickSight, Jupyter Notebooks, Amazon SageMaker Data Wrangler).
- D3.4 — Task 3.4: Ensure data quality
- 3.4.2 — Define data quality rules (for example, DataBrew).
- 3.4.3 — Investigate data consistency (for example, DataBrew).