Analytics
CoreAmazon EMR
Customizable distributed processing for batch, transformation, and large-scale log workloads.
Key points
- EMR runs distributed frameworks such as Spark and Hadoop with control over libraries and cluster configuration.
- Its flexibility brings more compute and runtime ownership than serverless ETL services.
Best-known use cases
- Run distributed Spark or Hadoop transformations over large datasets.
- Operate data-processing frameworks that require custom libraries or cluster configuration.
- Process large log archives with parallel compute.
What candidates often confuse it with
- Use EMR when a distributed workload needs framework or cluster control that Glue or Lambda does not provide.
Key takeaway
Choose EMR for large distributed processing that justifies explicit framework and cluster choices.
Relevant exam tasks
- D1.1 — Task 1.1: Perform data ingestion
- 1.1.2 — Read data from batch sources (for example, Amazon S3, AWS Glue, Amazon EMR, AWS DMS, Amazon Redshift, AWS Lambda, Amazon AppFlow).
- D1.2 — Task 1.2: Transform and process data
- 1.2.5 — Implement data transformation services based on requirements (for example, Amazon EMR, AWS Glue, Lambda, Amazon Redshift).
- D2.1 — Task 2.1: Choose a data store
- 2.1.1 — Implement the appropriate storage services for specific cost and performance requirements (for example, Amazon Redshift, Amazon EMR, AWS Lake Formation, Amazon RDS, Amazon DynamoDB, Amazon Kinesis Data Streams, Amazon Managed Streaming for Apache Kafka [Amazon MSK]).
- D3.1 — Task 3.1: Automate data processing by using AWS services
- 3.1.4 — Use the features of AWS services to process data (for example, Amazon EMR, Amazon Redshift, AWS Glue).
- D3.3 — Task 3.3: Maintain and monitor data pipelines
- 3.3.6 — Troubleshoot and maintain pipelines (for example, AWS Glue, Amazon EMR).
- D4.4 — Task 4.4: Prepare logs for audit
- 4.4.5 — Integrate various AWS services to perform logging (for example, Amazon EMR in cases of large volumes of log data).