Analytics
CoreAWS Lake Formation
Central governance for cataloged data-lake permissions across integrated analytics engines.
Key points
- Lake Formation governs access to cataloged data-lake resources across integrated analytics engines.
- Its permissions complement rather than replace IAM, catalog visibility, and access to the underlying storage.
Best-known use cases
- Centrally govern access to data-lake tables, columns, and rows.
- Share curated lake data across accounts with consistent permissions.
- Apply fine-grained permissions across Athena, EMR, Glue, and Redshift access.
What candidates often confuse it with
- Lake Formation is a governance layer over cataloged lake data, not a query engine or database.
Key takeaway
Use Lake Formation when table-, column-, or row-level lake permissions must remain consistent across engines or accounts.
Relevant exam tasks
- D2.1 — Task 2.1: Choose a data store
- 2.1.1 — Implement the appropriate storage services for specific cost and performance requirements (for example, Amazon Redshift, Amazon EMR, AWS Lake Formation, Amazon RDS, Amazon DynamoDB, Amazon Kinesis Data Streams, Amazon Managed Streaming for Apache Kafka [Amazon MSK]).
- D2.4 — Task 2.4: Design data models and schema evolution
- 2.4.1 — Design schemas for Amazon Redshift, DynamoDB, and Lake Formation.
- D4.2 — Task 4.2: Apply authorization mechanisms
- 4.2.4 — Manage permissions through AWS Lake Formation (for Amazon Redshift, Amazon EMR, Amazon Athena, and Amazon S3).
- D4.5 — Task 4.5: Understand data privacy and governance
- 4.5.2 — Implement PII identification (for example, Amazon Macie with Lake Formation).