GregLab | Exam Prep

Storage

Core

Amazon S3

Batch source, lake storage, events, lifecycle, versioning, load/unload, governance, and formats recur.

Key points

  • S3 stores raw, curated, and consumption-zone objects and integrates with ingestion, query, lifecycle, and governance services.
  • Event notifications, versioning, replication, lifecycle actions, formats, and permissions are independent design controls.

Best-known use cases

  • Store raw, curated, and consumption-zone data in a data lake.
  • Trigger processing when new source objects arrive.
  • Manage retention and recovery with lifecycle rules and versioning.
  • Exchange columnar datasets with Athena, Glue, EMR, and Redshift.

What candidates often confuse it with

  • S3 is object storage and a data-lake foundation; EFS provides shared files and Redshift provides warehouse processing.

Key takeaway

Choose S3 for durable object-based lake data whose layout, lifecycle, and access are explicitly governed.

Relevant exam tasks

  • D1.1 — Task 1.1: Perform data ingestion
  • 1.1.2 — Read data from batch sources (for example, Amazon S3, AWS Glue, Amazon EMR, AWS DMS, Amazon Redshift, AWS Lambda, Amazon AppFlow).
  • D2.1 — Task 2.1: Choose a data store
  • 2.1.1 — Implement the appropriate storage services for specific cost and performance requirements (for example, Amazon Redshift, Amazon EMR, AWS Lake Formation, Amazon RDS, Amazon DynamoDB, Amazon Kinesis Data Streams, Amazon Managed Streaming for Apache Kafka [Amazon MSK]).
  • 2.1.5 — Implement data migration or remote access methods (for example, Amazon Redshift federated queries, Amazon Redshift materialized views, Amazon Redshift Spectrum).
  • 2.1.7 — Manage open table formats (for example Apache Iceberg).
  • D2.3 — Task 2.3: Manage the lifecycle of data
  • 2.3.1 — Perform load and unload operations to move data between Amazon S3 and Amazon Redshift.
  • 2.3.2 — Manage S3 Lifecycle policies to change the storage tier of S3 data.
  • 2.3.3 — Expire data when it reaches a specific age by using S3 Lifecycle policies.
  • 2.3.4 — Manage S3 versioning and DynamoDB TTL.
  • 2.3.5 — Delete data to meet business and legal requirements.
  • 2.3.6 — Protect data with appropriate resiliency and availability.
  • D3.1 — Task 3.1: Automate data processing by using AWS services
  • 3.1.7 — Query data (for example, Amazon Athena).
  • D4.1 — Task 4.1: Apply authentication mechanisms
  • 4.1.5 — Apply IAM policies to roles, endpoints, and services (for example, S3 Access Points, AWS PrivateLink).
  • D4.5 — Task 4.5: Understand data privacy and governance
  • 4.5.2 — Implement PII identification (for example, Amazon Macie with Lake Formation).
  • 4.5.3 — Implement data privacy strategies to prevent backups or replications of data to disallowed AWS Regions.

Learn more

Free AWS Certified Data Engineer - Associate prep

Build focused DEA-C01 quizzes from skill areas, topics, and product references.

Practice with exam-style multiple-choice and multiple-response questions, score breakdowns, explanations, and a compact reference for this lane's official exam domains.

Read Topics Build a quiz

Exam Weights

Exam snapshot

DEA-C01 at a glance

Category
Associate
Duration
130 minutes
Questions
65 total; 50 scored and 15 unidentified unscored
Formats
Multiple choice and multiple response
Scoring
100–1,000 scaled score; 720 minimum passing score

Quiz builder

Choose your practice set

Mode

Exam fidelity: AWS documents 65 questions in 130 minutes: 50 scored and 15 unidentified unscored, using multiple-choice and multiple-response formats. This site's practice accuracy and readiness do not reproduce AWS's 100–1,000 scaled scoring or identify unscored items. Difficulty labels describe this site's Associate-level scenario complexity, not an AWS-published question rating.

Reference

DEA-C01 topics and reference map

Study links

DEA-C01 resources