Introduction: The Dual Pillars of Cloud Engineering
Every enterprise digital workload, from microservices and mobile backends to petabyte-scale data analytics pipelines, fundamentally requires two primitives: Compute (processing cycles to execute application logic) and Storage (durable media to persist data). AWS offers an expansive portfolio of specialized compute and storage engines optimized for distinct performance, durability, and cost profiles.
Understanding when to provision virtual machines via Amazon Elastic Compute Cloud (EC2) versus executing serverless functions via AWS Lambda, and how to pair them with Amazon Simple Storage Service (S3) and Amazon Elastic Block Store (EBS), is essential for cloud engineering. This module examines the architectural characteristics of core compute and storage services.
Core Concepts: Compute Modalities (EC2 vs. Lambda)
AWS compute spans from full operating system control to serverless event abstractions:
- Amazon EC2 (Infrastructure as a Service – IaaS): Provides resizable virtual machines in the cloud. You select the instance family (General Purpose, Compute Optimized, Memory Optimized, Storage Optimized), guest OS (Linux, Windows), CPU architecture (x86_64, ARM-based AWS Graviton), and storage attachment. EC2 provides complete operating system-level control, making it ideal for legacy enterprise applications, long-running monolithic workloads, and container orchestrators.
- AWS Lambda (Function as a Service – FaaS): A serverless compute engine that executes code in response to events (e.g., file uploads to S3, HTTP requests via API Gateway, database mutations in DynamoDB). You upload your code; AWS manages all server provisioning, OS patching, auto-scaling, and high availability. You are billed purely for execution time calculated in milliseconds.
Deep Dive: Storage Architectural Paradigms (S3 vs. EBS vs. EFS)
Selecting the incorrect storage engine can degrade application throughput and inflate monthly AWS expenditures tenfold. AWS provides three distinct storage models:
- Amazon S3 (Object Storage): S3 stores data as objects (files and metadata) within flat logical containers called Buckets. Objects are accessible over standard HTTPS APIs from anywhere. S3 provides extreme durability (99.999999999% / 11 9s of durability) by replicating data across multiple Availability Zones. It is the premier choice for media assets, data lakes, backups, and static website hosting.
- Amazon EBS (Block Storage): EBS volumes provide high-performance, raw unformatted block-level storage designed to be attached to a single EC2 instance (analogous to a physical SSD or NVMe hard drive). EBS persists independently from the running state of the EC2 instance, making it suitable for relational databases, file systems, and boot volumes. EBS volumes exist in a specific Availability Zone and replicate within that single AZ.
- Amazon EFS (Elastic File Storage): A fully managed network file system implementing the standard NFSv4 protocol. Unlike EBS, EFS can be mounted concurrently by hundreds of EC2 instances and Lambda functions across multiple Availability Zones.
Deep Dive: Amazon S3 Storage Classes and Lifecycle Rules
Enterprise data lakes accumulate petabytes of historical logs and documents that are queried frequently during the first 30 days but rarely accessed after 90 days. S3 enables automated cost tiering via S3 Lifecycle Policies:
- S3 Standard: High throughput, low latency storage for active, frequently accessed data.
- S3 Standard-Infrequent Access (Standard-IA): Lower storage fees with a per-GB retrieval fee. Designed for data accessed less than once a month (e.g., disaster recovery backups).
- S3 Glacier Flexible & Glacier Deep Archive: Ultra-low-cost archival storage for compliance records where retrieval times ranging from minutes to 12 hours are acceptable. Costs can be as low as $0.00099 per GB per month.
Case Study: Cost-Optimized Serverless Image Processing Pipeline
Consider a digital media enterprise that receives 500,000 high-resolution user image uploads daily. The legacy architecture ran a cluster of 8 Amazon EC2 m5.xlarge instances running 24/7 with attached 500 GB EBS GP3 volumes, costing over $1,800 monthly, with CPU utilization averaging under 12% during off-peak night hours.
The cloud engineering team re-architected the solution into an event-driven serverless pipeline:
- Decoupled Ingestion: User uploads are directed to an Amazon S3 bucket with S3 Transfer Acceleration enabled.
- Event-Driven Processing: The S3 object creation event automatically triggers an AWS Lambda function running Python with the Pillow imaging library, resizing images into thumbnail, mobile, and desktop formats in under 450 milliseconds.
- Lifecycle Archival: Resized images are written to an egress S3 bucket, while original raw images transition automatically to S3 Glacier Deep Archive after 30 days via S3 Lifecycle rules.
Monthly infrastructure costs collapsed from $1,800 to under $95, while processing throughput scaled automatically during peak breaking-news events with zero server provisioning overhead.
Common Mistakes & Practical Pitfalls
- Treating S3 as a Local File System: S3 is an object store, not a POSIX block file system. It does not support byte-level in-place edits (updating a single byte in a 10 GB file requires re-uploading the entire 10 GB object) or atomic file locking. Applications requiring POSIX compliance must use EBS or EFS.
- Unintentional Public S3 Bucket Exposure: Failing to enable S3 Block Public Access has caused high-profile corporate data breaches. S3 buckets are private by default; public access should be locked down at the AWS Account and Organization level.
- Ignoring EC2 Instance Store Volatility: Some EC2 instance types provide temporary physical disk storage called Instance Store. When an instance is stopped or terminated, all data on the Instance Store is permanently lost. Permanent operational data must always reside on EBS, EFS, or S3.
- Serverless Execution Time Limit Violations: AWS Lambda has a strict maximum execution timeout of 15 minutes. Attempting to run long-running batch extraction jobs that take hours on Lambda will result in abrupt function termination. Long-running jobs belong on EC2, AWS Batch, or AWS Fargate.
Exam Connection: Certification Blueprint Alignment
This module aligns directly with competencies evaluated on the AWS Cloud Practitioner Assessment:
- Identifying the primary use cases for EC2, Lambda, S3, EBS, and EFS.
- Selecting optimal S3 Storage Classes based on access frequency and retrieval latency requirements.
- Evaluating EC2 pricing models: On-Demand, Spot Instances (up to 90% discount for fault-tolerant workloads), Reserved Instances, and Savings Plans.
- Understanding EBS snapshot backups stored durably in S3.
Key Takeaways
- Amazon EC2 provides full OS virtual servers; AWS Lambda provides serverless event-driven code execution billed by millisecond.
- Amazon S3 is high-durability (11 9s) object storage; Amazon EBS is single-AZ block storage for EC2 boot drives.
- S3 Lifecycle policies automate transitions between S3 Standard, S3 Standard-IA, and S3 Glacier to optimize storage costs.
- Lambda functions have a hard 15-minute maximum execution timeout.
Knowledge Check
- Which AWS storage service provides an NFS-compatible network file system that can be mounted simultaneously by multiple EC2 instances across different Availability Zones?
Answer: Amazon Elastic File System (Amazon EFS). - What happens to the data on an EBS volume when an EC2 instance is stopped (not terminated)?
Answer: The data on the EBS volume is fully preserved because EBS is durable network-attached block storage that persists independently of the EC2 instance lifecycle. - Which EC2 pricing model offers the highest discount (up to 90%) for workloads that can tolerate unexpected interruptions, such as background video rendering or big data batch processing?
Answer: Amazon EC2 Spot Instances.
Next Step in Curriculum
Advance to the final AWS cloud module: Networking, VPC Architecture, and Security in AWS, or test your readiness in the practice arena.
