# AWS Solutions Architect - Professional

*2024-07-03*

> Notes from studying for the AWS Solutions Architect - Professional certification


After earning the [Advanced Networking Specialty]({{< relref "/posts/2024/aws-advanced-networking-specialty" >}}), the [Solutions Architect – Professional](https://aws.amazon.com/certification/certified-solutions-architect-professional/) exam was the natural next rung on the ladder – and, honestly, one of the more demanding ones. It's less about memorising services and more about knowing *which* service trade-offs to recommend when a scenario hands you a broken architecture and five ways to fix it.

These notes come from working through the [AWS Solutions Architect – Professional path](https://www.pluralsight.com/paths/aws-certified-solutions-architect-associate-saa-c03) on [Pluralsight](https://www.pluralsight.com). They're organised around the exam's major themes: data stores, networking, security, migration, scaling, resilience, operations, and cost. I've kept them flash-card dense on purpose – that's how I learn – but I've flagged the traps and decision points that seem to show up repeatedly.

Whether you're coming from the Associate level or another specialty, I hope these save you some late nights.

> **Before you dive in:** This post reflects the exam as I studied for it in 2024. AWS rotates exam content regularly - always cross-check the official exam guide before scheduling.

## Contents

1. [Data Stores](#data-stores)
2. [Networking](#networking)
3. [Security](#security)
4. [Migration Strategies](#migration-strategies)
5. [Architecting to Scale](#architecting-to-scale)
6. [Business Continuity & Disaster Recovery](#business-continuity-and-disaster-recovery)
7. [Deployment & Operations](#deployment-and-operations)
8. [Cost Management](#cost-management)

## 1. Data Stores{#data-stores}

This domain rewards you for thinking in terms of *data characteristics* – persistence, consistency, access patterns – before picking a service. Almost every scenario question here is really asking "which of these five services fits this workload shape?"

### Core concepts worth having cold

**Persistence model.** Persistent (survives anything – RDS, Glacier), transient (passed along and gone – SQS, SNS), ephemeral (gone on stop - instance store, Memcached). Expect at least one question where the cheapest correct answer hinges on recognising you *don't* need persistence.

**IOPS vs. throughput.** IOPS is how fast you can read/write; throughput is how much at once. This distinction drives EBS volume-type questions and RAID configurations later.

**ACID vs. BASE.** Atomicity, Consistency, Isolation, Durability vs. Basically Available, Soft state, Eventual consistency. This is the classic relational-vs-NoSQL decision axis, and the exam leans on it hard.

### S3

The exam's S3 questions almost always come down to three things: **limits, encryption, and lifecycle behavior**. Knowing the numbers cold saves you real exam time.

**Limits.** Max object size is 5 TB; the largest object in a single `PUT` is 5 GB, so anything bigger needs multipart upload (AWS recommends it from 100 MB).

**Security model.** Resource-based (bucket policies, object ACLs) and user-based (IAM policies), both evaluated together. Optional MFA delete for extra protection on destructive operations.

**Data protection.**
- **Versioning** – new version on every write; enables rollback and undelete. Old versions keep billing until permanently deleted, and versioning is a prerequisite for cross-region replication. Expect a question pairing versioning + lifecycle rules.
- **Cross-region replication** – for compliance, latency, or durability. Remember it requires versioning enabled on both sides.
- **Storage classes** – Standard, Standard-IA, One Zone-IA, Intelligent-Tiering, Glacier, and Glacier Deep Archive. (Reduced Redundancy is effectively deprecated – see the callout below.) Intelligent-Tiering now extends to automatic archival tiers for infrequently and rarely accessed data.
- **Lifecycle management** – transitions between classes and expirations on a schedule; the tool of choice for "optimise storage costs" and "adhere to retention policy" scenario answers.

**Analytics integrations.** S3 as a data lake (Athena, Redshift Spectrum, QuickSight), as a Kinesis Firehose landing zone for IoT streaming data, and as ML/AI storage backing Rekognition, Lex, and MXNet.

**Encryption at rest** – memorise all four variants and, more importantly, *who manages the key*:
- **SSE-S3** - AWS manages the keys, AES-256, simplest option
- **SSE-KMS** - keys generated and managed in KMS, adds an audit trail via CloudTrail
- **SSE-C** - you upload your own AES-256 key with each request
- **Client-side** - you encrypt before upload; AWS never sees plaintext

**Exam tricks.** Transfer Acceleration (CloudFront in reverse for fast uploads), Requester Pays buckets, object tags for cost allocation and security policy, event notifications to SNS/SQS/Lambda, and static website hosting.

> **Deprecated but still quizzed:** Reduced Redundancy Storage is no longer recommended – AWS positions Standard for nearly everything it covered. You may still see it in older practice exams; treat "RRS is the answer" answers as suspect.

### Glacier

Glacier is cheap, cold, and deliberately awkward to retrieve from. The things worth remembering: archives are immutable, max 40 TB, typically zip/tar, stored in **vaults**.

**Vault Lock** is the star feature for compliance scenarios: write a lock policy (e.g., "no deletes ever" or "MFA required to delete") that is *itself immutable once confirmed*. The flow has a deliberate safety catch – initiate the lock, you get a **24-hour window** to either complete or abort it, and if you do nothing within 24 hours the lock simply never takes effect.

### Elastic Block Store (EBS)

Think "virtual hard drive, but with exam-flavoured gotchas." EBS volumes attach only to EC2, and are **locked to a single AZ** – an instance can't mount a volume from another AZ. That constraint is the setup for half the snapshot-related questions.

**Snapshots** are how you escape those constraints: share data sets across users or accounts, migrate a system to a new AZ or region, and convert an unencrypted volume into an encrypted one. **Data Lifecycle Manager** schedules snapshots on your chosen interval with retention rules to prune stale ones.

### Elastic File System (EFS)

NFS-as-a-service. Metadata and data are replicated **across multiple AZs** - the direct contrast with EBS's single-AZ limitation, and the reason "shared file storage across AZs" answers point here. You can mount it from EC2 instances across many AZs, and even from on-premises systems (with caution – that's usually the wrong answer; for on-prem integration, look at **Amazon DataSync** or the Storage Gateway family instead).

### FSx

Distributed, managed file systems that aren't plain NFS – the names basically hand you the answer in the exam. Four flavours:

- **FSx for Windows File Server** – SMB, Active Directory-integrated; the answer whenever you see Windows LoB apps, Windows CMSs, media processing, or Windows analytics workloads
- **FSx for Lustre** - HPC; scales reliably to thousands of concurrent EC2 instances; big data and ML. Watch for it paired with S3-backed data sets
- **FSx for OpenZFS** and **FSx for NetApp ONTAP** - round out the family; ONTAP shows up in multiprotocol enterprise scenarios

### Storage Gateway

A VM (VMware/Hyper-V) or hardware appliance on-premises, backed by S3 and Glacier – the classic first step for hybrid cloud, DR, and migration scenarios. Know the four modes cold, because the exam tests "which gateway type?" constantly:

| Mode                 | Interface | Behavior                                                                      |
|----------------------|-----------|-------------------------------------------------------------------------------|
| File Gateway         | NFS, SMB  | On-prem/EC2 clients store and retrieve **objects in S3** via a file interface |
| Volume – Stored mode | iSCSI     | Primary data stays **on-prem**, asynchronously replicated to S3               |
| Volume – Cached mode | iSCSI     | Primary data lives in **S3**, frequently accessed data cached locally         |
| Tape Gateway         | iSCSI     | Virtual tape library/media changer for existing backup software               |

*Memory hook:* **Stored** = the truth is **stored on-prem**; **Cached** = the truth is in the cloud, cache is local.

### RDS

Managed MySQL, MariaDB, PostgreSQL, MSSQL, Oracle, and Aurora. Drop-in replacement for on-prem relational databases, with automated backups, patching, push-button scaling, replication, and redundancy thrown in.

**When RDS is the *wrong* answer** - the antipatterns are practically a mandatory exam question:
- Lots of BLOBs → **S3**
- Name/value or unpredictable data structures → **DynamoDB**
- Automated scalability beyond what RDS offers → **DynamoDB**
- IBM DB2 or SAP HANA → **EC2** (not RDS-supported)
- Need complete OS-level control → **EC2**

**Replication model.** Multi-AZ gives you a synchronous standby in another AZ; read replicas serve read-heavy workloads (and can span regions). If one AZ fails, the standby is promoted and replicas keep working. If a whole *region* fails, promote a read replica to standalone and reconfigure it as Multi-AZ. Note for MySQL: MyISAM doesn't replicate – you need InnoDB (or XtraDB on Maria).

### Aurora

Fully managed, MySQL- and PostgreSQL-compatible, multi-AZ by design, with storage-level replication. The details the exam loves:

- **Endpoints** – a cluster (writer) endpoint and a reader endpoint, so applications don't track individual instances
- **Global Databases** – one primary region, up to 5 secondary regions, replication at the storage layer with typical lag under a second; secondary regions can be promoted during an outage
- **Serverless** - billed in **ACUs** (Aurora Capacity Units), scaling between **0.5 and 128** ACUs; ideal for variable or infrequent workloads with low operational overhead

### DynamoDB

Multi-AZ NoSQL with optional global tables for cross-region replication, priced on throughput (RCUs/WCUs), with provisioned (auto-scaled between your min/max), and on-demand capacity modes. Defaults to eventual consistency, but you can request strongly consistent reads via an SDK parameter. ACID guarantees arrive via **DynamoDB Transactions**.

The heart of the DynamoDB questions is **indexes** – specifically which kind solves the query you're being asked about:

| Index                      | Keys                                                     | When                                                                                                                                                  |
|----------------------------|----------------------------------------------------------|-------------------------------------------------------------------------------------------------------------------------------------------------------|
| **Global Secondary Index** | Partition key *and* sort key can differ from the table's | Query efficiently on attributes outside the primary key (e.g., look up orders by customer number)                                                     |
| **Local Secondary Index**  | Same partition key, different sort key                   | You already know the partition key and want to filter/refine by another attribute (e.g., all orders *for this customer* with a given material number) |

*Memory hook:* a Local Secondary Index has to stay **local** – it respects the table's partition key. A Global one can go anywhere.

**Projection counts too.** Sparse projections (keys + a few attributes) cost the least and give the lowest latency for narrowly shaped queries; projecting everything doubles storage and write costs but maximises flexibility. And rare queries on frequently written data → keys-only projection to keep writes fast.

**Capacity math (know it, it appears in calculation questions):**
- Partitions by capacity: `(total RCU ÷ 3,000) + (total WCU ÷ 1,000)`
- Partitions by size: `total size ÷ 10 GB`
- Actual partitions = **round up** `MAX(capacity, size)`

Example: 2,000 RCU / 2,000 WCU on a 10 GB table → MAX(0.67 + 2.0, 1.0) = 2.67 → **3 partitions**, each carrying a share of the throughput.

### DocumentDB

MongoDB-compatible NoSQL JSON document store with AWS-managed scaling and write-then-read consistency. Think user profiles, content management, real-time big data.

### Redshift

Petabyte-scale managed data warehouse, PostgreSQL-compatible via JDBC/ODBC, built on columnar storage and massively parallel processing for complex analytical queries. **Redshift Spectrum** queries data directly in S3 without loading it. Data-lake architecture questions – "query raw data with minimal pre-processing," "shorten time from collection to insight" – point here.

### AWS Glue

Serverless ETL. **Crawlers** discover data (S3, Redshift, RDS, databases on EC2), populate the **Data Catalog**, which feeds Athena, EMR, and Redshift, and powers ETL jobs that land data in S3, Redshift, or Lake Formation. Glue Data Quality layers on predefined rules, metrics, and alerts to keep the pipeline trustworthy. Know the crawler → catalog → job flow diagram; it shows up in data-lake questions.

### Neptune

Fully managed **graph database** speaking both Gremlin and SPARQL. Any scenario about relationships - fraud rings, social networks, knowledge graphs - is Neptune.

### ElastiCache

The final answer to "how do I reduce database load / survive instance loss / power a leaderboard?"

| Use case                         | Right choice  | Why                                                           |
|----------------------------------|---------------|---------------------------------------------------------------|
| Web session store                | **Redis**     | Sessions survive the loss of any one load-balanced web server |
| Database caching in front of RDS | **Memcached** | Simple, multicore, caches popular query results               |
| Leaderboards                     | **Redis**     | Sorted sets give live rankings at scale                       |
| Streaming dashboards             | -             | Sensor data landing spot feeding real-time displays           |

**Memcached vs. Redis** is a reliable exam discriminator. **Memcached** is straightforward, scales in and out, uses multiple cores, and is object-cache-shaped. **Redis** brings encryption, HIPAA eligibility, clustering, complex data types, replication/HA, pub/sub, geospatial indexing, and backup/restore. Rule of thumb: if the question implies any *durability, availability, or richness* requirement, it's Redis.

### Athena

Serverless SQL over S3, built on Presto. Convert your data to **Parquet** for a dramatic performance and cost win – that pairing is practically a guaranteed exam answer. Athena also runs Spark code, and federated queries reach 25+ data sources.

**Athena vs. Redshift Spectrum:** choose Athena when the data lives mostly in S3 and stands alone; choose Spectrum when you need to **join S3 data with existing Redshift tables**.

### Quantum Ledger Database (QLDB)

Not quite blockchain – a **centralised, append-only, immutable ledger** with a verifiable hash chain, which centralisation makes fast and scalable. Every record contributes to the integrity of the chain. Use when you need provable transaction history with an audit log - think banking records - *without* multi-party consensus.

QLDB vs. **Amazon Managed Blockchain**: QLDB is centralised and single-owner; Managed Blockchain is a distributed, consensus-based framework (Hyperledger Fabric or Ethereum) involving multiple network members and nodes across AWS accounts. The exam's ledger questions hinge entirely on that distinction.

### Timestream

Purpose-built **time-series** database - sensor networks, industrial machinery, equipment telemetry – with built-in analytics like interpolation and smoothing. The alternative when DynamoDB or Redshift feel heavyweight for timestamped telemetry.

### OpenSearch

Mostly a search engine, though it can store documents (with caution). AWS's managed flavor is the ELK-ish stack: **OpenSearch** for search/storage, **Logstash** (or CloudWatch/Firehose/IoT) for intake, and dashboards via OpenSearch Dashboards (née Kibana). You'll still hear people say "Elasticsearch" – same lineage, renamed.

## 2. Networking{#networking}

If the Pro exam has a "home turf," this is it. Networking questions dominate, and they're rarely "what is X?" – they're "given this connectivity requirement and these constraints, which option?" The decision matrices below (VPN vs. Direct Connect vs. Transit Gateway, placement groups, Route 53 policies) are where I'd spend the most review time.

### Foundations worth reviewing fast

**The OSI model** – you likely know it, but the exam frames load balancer and firewall questions by layer, so keep the mnemonic handy: *Please Do Not Throw Sausage Pizza Away* (Layers 7→1). Remember ALB = Layer 7, NLB = Layer 4, WAF = Layer 7, Network Firewall = Layers 3–4.

**Unicast vs. multicast.** AWS doesn't support multicast at the network level – that's occasionally the entire trick behind a question.

**TCP / UDP / ICMP.** TCP (L4) is connection-based and stateful with acknowledgements; UDP (L4) is connectionless and loss-tolerant (streaming, DNS); ICMP (L3) is the routers' own health check language (ping, traceroute). Know which protocol belongs to which layer - hybrid questions love mixing L3/L4/L7 facts.

**Ephemeral ports.** The suggested range is **49152–65535**. Linux typically uses 32768–61000 and Windows starts around 1025. This matters for security groups and NACLs: if you write rules that forget return traffic on ephemeral ports, that's a classic exam trap.

**Reserved IPs in every subnet.** Five addresses, always: network address (.0), VPC router (.1), Amazon DNS (.2), future use (.3), broadcast (.255). A /28 subnet gives you 16 addresses → only **11 usable**. Capacity-planning questions exploit this constantly.

### Connecting your network to a VPC

This is the highest-yield table in the whole domain. Learn the "when" column - the exam scenario always maps to one of these:

| Option              | What it is                                                                    | When it's the answer                                                                 | Trade-off                                         |
|---------------------|-------------------------------------------------------------------------------|--------------------------------------------------------------------------------------|---------------------------------------------------|
| **AWS Managed VPN** | IPsec VPN over the internet                                                   | Quick setup, or a redundant backup for Direct Connect                                | Depends on internet quality                       |
| **Direct Connect**  | Private dedicated circuit into AWS                                            | Large, steady bandwidth needs; predictable performance; up to 10 Gbps; BGP required  | No built-in redundancy; may need telecom partners |
| **DX + VPN**        | IPsec tunnel over the private circuit                                         | Compliance-driven encryption on top of DX                                            | Added complexity                                  |
| **VPN CloudHub**    | Hub-and-spoke VPN between multiple remote sites via a Virtual Private Gateway | Linking branch offices; can prefer MPLS with VPN as backup                           | Internet-dependent, no inherent redundancy        |
| **Software VPN**    | You run both endpoints                                                        | Compliance scenarios where you must control the whole chain, or unsupported VPN tech | You own all redundancy                            |
| **Transit VPC**     | Legacy pattern for multi-region mesh                                          | Mostly historical - Transit Gateway replaced it                                      | You maintain it                                   |

**Migration tip:** when transitioning from VPN to Direct Connect, run both within the same BGP prefix advertisement – from AWS's side the DX path is always preferred, making the cutover nearly seamless.

### VPC-to-VPC connectivity

Options: VPC peering, Transit Gateway, or the VPN/software approaches above.

**VPC Peering.** Uses AWS-managed networking, non-transitive - every pair must be explicitly connected, which becomes painful at scale ("N×(N−1)/2 peering connections" is a classic exam calculation).

**Transit Gateway.** The modern answer for scale: a regional hub connecting up to **5,000 attachments** across accounts; cross-region requires **peering multiple TGWs**; can associate a DX Gateway. Watch for the CIDR overlap constraint - overlapping CIDRs can't be peered, period.

### Internet access patterns

**Internet Gateway.** Horizontally scaled, redundant by default, no bandwidth constraint. A subnet with a route to an IGW = *public subnet*. It performs NAT for instances with public IPs – but **not** for private-IP-only instances, which is the setup line for NAT devices.

**Egress-only Internet Gateway** – the IPv6 equivalent of NAT. Since IPv6 addresses are globally unique/public by default, this gives outbound-only access while blocking inbound initiation. Stateful. Route `::/0` at it.

**NAT Instance vs. NAT Gateway** – know this table cold; it's near-guaranteed:

|                 | NAT Gateway                                                           | NAT Instance                        |
|-----------------|-----------------------------------------------------------------------|-------------------------------------|
| Availability    | Managed, highly available **within its AZ**                           | On you                              |
| Bandwidth       | Scales up to **45 Gbps**                                              | Bounded by instance type            |
| Maintenance     | AWS-managed                                                           | Yours                               |
| Public IP       | Elastic IP, can't be detached                                         | Elastic IP, detachable              |
| Security groups | Cannot attach                                                         | Can use them (and can be a bastion) |
| Restriction     | Can't be used for peering/VPN/DX traffic - those need explicit routes | Routes are yours                    |

Two placement facts that anchor scenario questions: NAT Gateways live in a **public subnet**, and for multi-AZ resilience you deploy **one per AZ** with private-subnet routes pointing at their local gateway. (Note: single-AZ NAT Gateways remain a notable exam gotcha – an AZ failure takes out any private instances routed through that gateway unless you've architected per-AZ.)

### Routing

Every VPC has an implicit router and a main route table; the **most specific route wins**. That principle resolves half the tricky routing questions – especially when a NAT route (0.0.0.0/0) competes with a more specific peering or DX route.

**BGP.** Dynamic routing; required for Direct Connect, optional for VPN. Needs TCP **port 179** plus ephemeral ports. ASNs identify endpoints; weight is local to the router with **higher weight preferred** for outbound traffic.

### Enhanced Networking & Placement Groups

Enhanced Networking uses **SR-IOV (Single Root I/O Virtualization)** for higher performance – may need a driver on non-Amazon-Linux HVM AMIs. The **Intel 82599 VF** interface delivers 10 Gbps; the **Elastic Network Adapter (ENA)** reaches 25 Gbps.

Placement groups are a favorite exam topic – the three-way contrast:

|          | Cluster                                     | Spread                                                  | Partition                                                                         |
|----------|---------------------------------------------|---------------------------------------------------------|-----------------------------------------------------------------------------------|
| Layout   | Tight low-latency grouping in **one AZ**    | Instances on distinct underlying hardware               | Instances divided into rack-level partitions                                      |
| Use when | Maximise network throughput (HPC)           | Minimise simultaneous hardware failure for small groups | Large distributed systems (HDFS, Cassandra) worried about correlated rack failure |
| Limits   | Finite capacity - launch everything upfront | **Max 7 instances per group per AZ**                    | Not for Dedicated Hosts                                                           |

*Memory hook:* cluster = speed, spread = "don't share hardware," partition = "don't share racks."

### PrivateLink

Private connectivity to services without traversing the internet, and - crucially - **without requiring VPC peering**. Use cases: marketplace SaaS consumption, exposing services to other VPCs with tight control, third-party app access with simplified networking. Contrast with peering whenever a scenario says "only need to reach one service, not the whole VPC."

### Global Accelerator

Users hit AWS edge locations and ride the **AWS private backbone** the rest of the way. You get static anycast IPs, performance gains, and resilience. Use cases: improving global app performance, hybrid networking (cheaper middle ground than DX), origin masking/DDoS defense (entry points protected by Shield), and multi-region failover across ~10 regional endpoints. **Key contrast:** CloudFront caches *content*; Global Accelerator accelerates *TCP/UDP traffic* to your own endpoints.

### Route 53 routing policies

Straightforward table, but the exam disguises the differences in scenario prose:

| Policy            | Thinking                                                                |
|-------------------|-------------------------------------------------------------------------|
| Simple            | "Here's the destination"                                                |
| Failover          | Primary failed its health check → route to the secondary                |
| Geolocation       | *Where is the user?* → route by geography (hard boundaries)             |
| Geoproximity      | *Where is the user relative to my resources?* → route/bias by proximity |
| Latency           | Which endpoint gives the lowest latency → route there                   |
| Multivalue answer | Return several healthy IPs (poor man's load balancer)                   |
| Weighted          | Split traffic by percentage across resources                            |

Geolocation vs. latency is the classic confusion pair: geolocation is *about the user's location* (for compliance/content rules); latency is *about speed*, wherever that lands.

**Cross-account subdomain delegation**: create a hosted zone for the subdomain in the child account, copy its NS records, then add those NS records to the parent account's hosted zone. Straightforward process – worth having walked through once before exam day.

### CloudFront

A CDN distributing static *and* dynamic content via edge caching (dynamic content handled by forwarding cookies/headers to the origin). Supports up to 4K live and on-demand video. Origins can be S3, EC2, ELBs, or any web server - multiple origins with **behaviours** routing by URL path.

- **Invalidation:** either delete at origin and let TTL expire, or submit an invalidation request via console, API, or third-party tools.
- **Geo-restrictictions** by country at the distribution level.
- **Zone apex** records supported (Route 53 alias to CloudFront).
- **SNI** lets many certificates share one IP - older clients that lack SNI force the pricier dedicated-IP option.
- The RTMP distribution type is **deprecated** – modern streaming runs over HTTP(S) web distributions, sometimes with ACM-issued certificates.

### Elastic Load Balancers

|                                | Application Load Balancer       | Network Load Balancer | Classic Load Balancer |
|--------------------------------|---------------------------------|-----------------------|-----------------------|
| OSI Layer                      | Layer 7                         | Layer 4               | Layer 4 or 7          |
| Zonal Failover                 | Yes                             | Yes                   | Yes                   |
| Platform                       | VPC Only                        | VPC Only              | EC2-Classic or VPC    |
| Health Checks                  | Yes                             | Yes                   | Yes                   |
| Cross-Zone Load Balancing      | Yes                             | Yes                   | Yes                   |
| CloudWatch Metrics             | Yes                             | Yes                   | Yes                   |
| SSL Offloading                 | Yes                             | Yes                   | Yes                   |
| Resource-based IAM Permissions | Yes                             | Yes                   | Yes                   |
| Protocols                      | HTTPS. HTTP                     | TCP, UDP, TLS         | TCP, SSL, HTTP, HTTPS |
| Path or Host-based Routing     | Yes                             | No                    | No                    |
| WebSockets                     | Yes                             | Yes                   | No                    |
| Server Name Indication (SNI)   | Yes                             | Yes, as of 9/2019     | No                    |
| Sticky Sessions                | Yes                             | Yes, as of 3/2020     | Yes                   |
| Static IP, Elastic IP          | Only through Global Accelerator | Yes                   | No                    |
| User Authentication            | Yes                             | No                    | No                    |

- **ALB (L7)** - path/host/header/method/query-string/source-CIDR routing, user authentication, WebSockets. The default for HTTP workloads.
- **NLB (L4)** - static/Elastic IPs, extreme performance, TCP/UDP/TLS passthrough (including non-HTTP protocols), TLS target offload, long-lived connections preserved.
- **CLB** - legacy; mainly appears in exam questions as "migrate this away."

Notable facts: sticky sessions are supported by all three (NLB gained them in 3/2020, ALB via duration or application cookie); SNI is supported on ALB and NLB but not CLB; and an ALB can only get a static IP **through Global Accelerator**.

## 3. Security{#security}

Security questions on the Pro exam follow a predictable arc: *who is asking, how do they prove it, what are they allowed to do, and how do I detect when something went wrong*.

### The four questions

**Identity** – who are you? **Authentication** – prove it. **Authorisation** – are you allowed? **Trust** – do entities I already trust vouch for you? Nearly every security scenario on the exam reduces to deciding which of these is missing or misdesigned.

### Multi-account organisations

The "why multiple accounts?" list is worth internalising because scenario questions *presuppose* it: group workloads by business purpose and ownership, apply the least privilege per environment, restrict access to sensitive data, limit the blast radius of a security event, and consolidate billing. Management accounts provision the org; member accounts hold the workloads; organisational units group accounts by application or service.

**AWS Control Tower** automates the best-practice landing zone on top of Organisations: guardrails enforced via SCPs (preventive) and AWS Config rules (detective), plus a dashboard for compliance visibility. Vocabulary to know cold:
- **Landing zone** – the customisable multi-account starting point
- **Guardrail** – a governance rule backed by an SCP or Config rule
- **Baseline** – a bundle of blueprints (CloudFormation) + guardrails
- Special accounts it provisions: **log archive** (aggregates logs) and **audit** (cross-account audit roles)

### Service Control Policies (SCPs)

SCPs use IAM policy syntax but **never grant permissions** – they only filter what's possible. Applied at org, OU, or account level, they inherit downward. The exam tests the interaction:

> **Effective permissions = allowed by IAM ∧ not denied by any SCP.**

Deny-list SCPs explicitly forbid specific actions; allow-list SCPs implicitly deny everything unlisted (with `FullAWSAccess` attached by default so nothing breaks until you replace it). Pair SCPs with **AWS Config** for the detective side - compliant/noncompliant resource tracking with the history of what changed.

**IAM Identity Center** (the successor to AWS SSO) maps users and groups from your identity provider, speaks SAML 2.0 for federation with providers like Azure AD, or acts as a standalone directory. It supports MFA organisation-wide.

**IAM Users vs. Identity Center** – the exam-contrasting table:

|             | IAM Users                   | IAM Identity Center     |
|-------------|-----------------------------|-------------------------|
| Credentials | Static access keys          | Rotating keys via roles |
| Scope       | One user, one account       | One user, many accounts |
| Permissions | One permission set per user | Many assumable roles    |
| Federation  | Not supported               | The design goal         |

Rule of thumb: any scenario mentioning federation, a corporate IdP, or "hundreds of users across dozens of accounts" is Identity Center. Any mention of AD Connector feeding Identity Center means SSO for on-prem employees with their existing credentials (including RADIUS-based MFA).

### Directory services decision table

| Option                              | What it is                                  | Best for                                             |
|-------------------------------------|---------------------------------------------|------------------------------------------------------|
| **AWS Managed Microsoft AD**        | Full managed AD on Windows Server           | Enterprises wanting hosted AD or LDAP for Linux apps |
| **AD Connector**                    | Proxy to your existing on-prem AD           | SSO for on-prem staff; joining EC2 to your domain    |
| **~~Simple AD~~** *(retired)*       | Samba-based standalone directory            | Historical; no MFA, no trust relationships           |
| **~~Cloud Directory~~** *(retired)* | Hierarchical cloud-native directory         | Historical                                           |
| **Amazon Cognito**                  | Consumer sign-up/sign-in, social federation | Customer-facing apps and SaaS                        |

### Roles, trust, and cross-account access

Anything can assume a role: IAM users/groups, AWS services, non-AWS workloads (via **IAM Roles Anywhere**), and federated identities. Cross-account sharing is a three-step dance the exam tests mechanically:

1. **Define permissions** in the target account's role policy
2. **Create the trust policy** naming the trusted principals (`"Action": "sts:AssumeRole"`, optionally constrained by conditions like a source IP)
3. **Grant AssumeRole** permission to the trusted account's users

Canonical use cases: external auditors getting scoped access, Lambda in one account assuming a role in another, and workload permissions inside or outside AWS.

**Token Vending Machine** – the classic pattern for issuing temporary credentials to mobile clients: anonymous TVM grants service access only; identity TVM adds registration/login. Older pattern, but it still surfaces in exam lore.

**AWS Resource Access Manager (RAM)** shares *resources* (not roles) across accounts or OUs – the sharing account retains ownership. Flow: choose the resource, attach a managed policy, define principals, send the invitation. Prime use cases: shared foundational infrastructure (subnets, Transit Gateways), centralised certificate authorities, App Mesh networking, and cross-account Aurora/RDS cluster cloning.

### Secrets

**Secrets Manager** stores passwords, API keys, SSH/PGP keys, and the like, retrievable via API with fine-grained IAM control, and can **automatically rotate RDS credentials** for MySQL, PostgreSQL, and Aurora using Lambda + KMS. Contrast with **Parameter Store** (Systems Manager) – simpler, cheaper, and fine for non-secret config; Secrets Manager is for things that rotate.

### Encryption

**At rest vs. in transit** – the exam loves hiding which one a scenario actually needs.

**KMS** provides key storage, management, and auditing with either AWS-generated or imported keys; access controlled via IAM, usage audited via CloudTrail. It sits behind most AWS-service-native encryption (remember SSE-KMS from the S3 section). Validated against FIPS 140-2 Level 3 and PCI DSS Level 1.

**CloudHSM** is the single-tenant, hardware answer: dedicated FIPS 140-2 Level 3 devices living *inside your VPC*, accessed via VPC peering, clustered for HA. The trade-off: no native AWS service integration – you script applications against it. Classic uses: acting as your own issuing CA, offloading SSL from web servers, TDE for Oracle databases.

|               | KMS                        | CloudHSM                |
|---------------|----------------------------|-------------------------|
| Tenancy       | Multi-tenant AWS service   | Single-tenant hardware  |
| Root of trust | AWS-managed                | **Customer-managed**    |
| Availability  | AWS-managed                | You cluster it          |
| Integration   | Native across AWS services | Custom application work |

The discriminator in exam questions is almost always the **root of trust**: any hint of "we cannot let AWS hold the keys" or strict regulatory key custody → CloudHSM.

**ACM** provisions and deploys SSL/TLS certificates, natively integrated with CloudFront, ELB, and API Gateway. Public certificates are free; renewal is managed; wildcards are supported; third-party certs can be imported (but then *you* renew them – a classic trap). **ACM Private CA** issues private certificates for internal apps and devices.

### Network-level defenses

**Security Groups** are stateful, instance-level firewalls (inbound source / outbound destination rules referencing IPs, subnets, or other SGs). **NACLs** are stateless, subnet-level packet filters – the default NACL allows everything, so NACLs are where you write explicit denies. Together they form layered least privilege; NACLs are also the backup line of defense when someone fat-fingers an SG rule open.

**WAF vs. Network Firewall** – remember by *what it attaches to*:
- **WAF** → ALB, CloudFront, API Gateway, AppSync – Layer 7 rules (SQLi, XSS, rate limiting)
- **Network Firewall** → VPC-level via Transit Gateway, internet gateways, DX/VPN gateways – Layer 3–4 inspection

**Shield.** Standard is free-automatic protection against common L3/L4 DDoS. **Shield Advanced** adds application-layer coverage, protects named resources (ELBs, CloudFront, Route 53, etc.), provides detailed attack forensics, and includes 24×7 access to the SRT (DDoS response team). **Firewall Manager** then scales WAF/Shield/Network Firewall/SG policies across accounts and the org via Organisations.

**GuardDuty** is intelligent threat detection across workloads, S3, accounts, and users – findings flow to EventBridge and into **Security Hub**, which correlates GuardDuty, Inspector, Macie, Firewall Manager, and third-party findings into prioritised recommendations.

**IDS/IPS and SIEM.** IDS watches for suspicious activity; IPS sits inline to prevent it. Neither is an AWS service – you compose them (agents + collection system), typically shipping logs to a SIEM via CloudWatch, S3, or third parties like Splunk/Sumo Logic. Expect "how would you architect detection for this hybrid environment" answers to include that composition.

### CloudWatch vs. CloudTrail

The perennial confusion pair, settled once and for all:

|           | CloudWatch                            | CloudTrail                           |
|-----------|---------------------------------------|--------------------------------------|
| Records   | Metrics, logs, events across services | **API activity** across services     |
| Role      | Higher-level monitoring and alerting  | Granular audit trail of who did what |
| Alarming  | Native                                | Via CloudWatch alarms on trail logs  |
| Retention | Logs indefinitely (configurable)      | Logs to S3/CloudWatch indefinitely   |

And **Security Hub + Firewall Manager + Config + GuardDuty** tie it together into the "managed security posture" answer family – when a scenario asks for centralised, multi-account security visibility, that's the stack.

### Service Catalog

Last piece of identity/governance: admins publish curated products (built on CloudFormation) with launch constraints, so end users self-provision without broad IAM permissions.

| Constraint       | What                                     | Why                                          |
|------------------|------------------------------------------|----------------------------------------------|
| **Launch**       | IAM role the catalog assumes on launch   | Users need no underlying service permissions |
| **Template**     | Rules narrowing allowed parameter values | E.g., only t3.micro in dev                   |
| **Notification** | SNS topic for stack events               | Visibility on launches and failures          |

Products can be versioned and removed without shutting down already-running deployments – worth remembering as the "governed self-service" exam answer.

## 4. Migration Strategies{#migration-strategies}

Migration questions rarely ask "what is DMS?" – they hand you a portfolio of aging applications and ask which strategy applies to each, then which *tool* executes it. So this section runs in that order: strategy vocabulary first, then the toolchain.

### The 6 Rs

Memorise this table until it's reflexive – the exam will describe an application and expect you to name its R in about twenty seconds:

| Strategy        | Description                                                | Example                         | Effort | Optimisation upside |
|-----------------|------------------------------------------------------------|---------------------------------|--------|---------------------|
| **Re-host**     | "Lift and shift" - move as-is                              | On-prem MySQL → EC2             | 2      | 1                   |
| **Re-platform** | "Lift and reshape" - move and swap the platform underneath | On-prem MySQL → RDS MySQL       | 4      | 3                   |
| **Re-purchase** | "Drop and shop" - abandon and buy SaaS                     | Legacy CRM → Salesforce         | 3      | 1                   |
| **Rearchitect** | Cloud-native redesign                                      | Monolith → serverless           | 5      | 5                   |
| **Retire**      | Decommission what nobody uses                              | Kill the label-printing app     | 0      | 0                   |
| **Retain**      | "Do nothing" - revisit later                               | Servers live to see another day | 1      | 0                   |

*Tip:* when a scenario says "minimal change, deadline-driven," that's re-host. When it says "take advantage of cloud elasticity," that's rearchitect. When finance is asking about licence costs on a dying app, check whether the answer is actually *retire* – it's the cheapest R and easy to overlook under exam pressure.

### The Cloud Adoption Framework (CAF)

TOGAF (The Open Group Architecture Framework, developed since 1995) is the de facto enterprise-architecture standard among Fortune 500s, but it's a *framework*, not a cookbook – open to local adaptation and frequently blamed for unreasonable expectations. The CAF, by contrast, is AWS's prescriptive lens for the adoption journey. Classic phases: **Project → Foundation → Migration → Reinvention**.

Then the six CAF perspectives, which map neatly onto "which perspective is this question about?":

- **Business** – the case for adoption; measurable benefits (TCO, ROI)
- **People** - roles, skills, training, incentives
- **Governance** – portfolio management, agile program management, cloud-aligned KPIs
- **Platform** - standardised provisioning, cloud-native architecture patterns
- **Security** – IAM model changes, logging/audit evolution, the shared responsibility shift
- **Operations** - automated monitoring, scalable performance management, cloud-era BC/DR

*Exam hook:* questions quote an organisational symptom (e.g., "no skills plan exists") and expect you to name the perspective – People, in that case.

### Hybrid architectures

Common first step before full migration, and the exam treats it accordingly:

- **Storage Gateway** bridging on-prem and S3 – seamless to users, low-risk, appealing economics; the textbook "first step into cloud"
- **Middleware integration** – loosely coupled, canonical-message-based connections
- **VMware integration** – vCenter plug-in for transparent VM migration; VMware Cloud on AWS deepens it

The unifying principle: loosely coupled integrations where either end can exist without knowledge of the other.

### Discovering and planning

**AWS Migration Hub** is the coordination dashboard: it leverages the **Application Discovery Service** to inventory and group servers, recommends right-sizing, suggests per-application migration strategies, and tracks lift-and-shift versus refactor-first progress. Know that discovery feeds recommendations – that flow is the scenario.

### Migrating applications

- **Containerised workloads** → AWS Batch or Fargate
- **Web applications** → Elastic Beanstalk or Lightsail
- **Message-handling microservices** → event-driven/serverless architectures (decoupled queues)
- **Windows/Linux servers wholesale** → **Application Migration Service (MGN)** - agent-based rehosting from on-prem, other clouds, or within AWS, with test instances and minimal-interruption cutover. Free to use; the agent right-sizes EC2 recommendations.
- **Bringing AWS *to* the data** → **Outposts** (preconfigured hardware running AWS services on-prem) or **Snowball Edge** (compute-capable offline transfer, e.g., EC2/EKS Anywhere)

### Migrating data

Three tools, chosen by *volume and connectivity* – the exam always gives you those two facts:

| Tool                | When                             | Details                                                                                                                                                                      |
|---------------------|----------------------------------|------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| **DMS**             | Database migrations              | Source relational DB → RDS/Aurora/EC2; **Schema Conversion Tool (SCT)** for heterogeneous engines; can also target S3, Snowball, and even MongoDB/DynamoDB for smaller moves |
| **DataSync**        | Online, ongoing file/object sync | SMB, HDFS, NFS → S3/EFS/FSx; automates the transfer securely                                                                                                                 |
| **Transfer Family** | Third-party file exchange        | Managed SFTP/FTPS/FTP/AS2 endpoints, multi-AZ; File Transfer Workflows automate encryption, filtering, tagging, compression                                                  |

And when the data is *too big or the connection too slow* – the **Snow Family**:

| Generation    | What it is                                                   |
|---------------|--------------------------------------------------------------|
| Import/Export | Ship your own hard drive to AWS (historical)                 |
| Snowball      | Ruggedized appliance, up to 80 TB, copied to S3              |
| Snowball Edge | Snowball + onboard compute (Lambda) and clustering           |
| Snowmobile    | A literal shipping container - up to 100 PB - towed by truck |

All encrypted at rest and in transit; transfer speed bounded by your courier, not your WAN. Rule of thumb: **bandwidth math beats postal math only when data fits through the pipe** – the exam's "10 Gbps link vs. 60 PB of video" questions resolve themselves once you compute weeks-versus-days.

### Network migration and cutover

The checklist that keeps a migration from becoming an outage:

1. **No IP overlap** between VPC and on-prem - VPC IPv4 netmasks run /16 to /28 (65,536 → 16 addresses), with 5 reserved IPs per subnet (a /28 nets you 11 usable)
2. **Start with VPN** – most organisations begin here
3. **Layer in DX with the VPN as backup**
4. **Cut over with BGP** - run VPN and DX advertising the same prefixes; AWS prefers the DX path, making the transition near-seamless

## 5. Architecting to Scale{#architecting-to-scale}

Scaling questions are really cost-plus-resilience questions in disguise: the exam describes a spiky workload, and you pick the mechanism that absorbs it without waste. Everything in this section hangs off one decision first – up or out.

### Scaling up vs. scaling out

|            | Scaling up                         | Scaling out                 |
|------------|------------------------------------|-----------------------------|
| What       | More CPU/RAM on existing instances | More instances              |
| Downtime   | Restart required                   | None                        |
| Automation | Manual scripting                   | Native for most AWS compute |
| Ceiling    | Instance size limits               | Effectively unlimited       |

Horizontal wins on exam day unless the scenario explicitly demands a monolithic resource (a legacy vertical app, a licensed single-node database). From here on, we're talking scale-out.

### Auto Scaling

Three layers – don't conflate them: **EC2 Auto Scaling** (ASGs), **Application Auto Scaling** (the API behind DynamoDB/ECS/EMR/etc. table capacity), and **AWS Auto Scaling** (the cross-service planning console).

**ASG scaling options:**

| Option    | What                             | When                  |
|-----------|----------------------------------|-----------------------|
| Maintain  | Keep N instances alive           | "I always need 3"     |
| Manual    | Adjust desired capacity yourself | Rarely-changing needs |
| Scheduled | Scale by clock                   | "Monday-morning rush" |
| Dynamic   | React to live metrics            | CPU crossing 70%      |

**Dynamic scaling policies**:

| Policy          | Behavior                                                 | Vibe                                          |
|-----------------|----------------------------------------------------------|-----------------------------------------------|
| Target tracking | Hold a metric near a target value                        | "Keep CPU at 70%" - the default modern answer |
| Simple          | Wait for health check + cooldown before evaluating again | Slow and steady                               |
| Step scaling    | Escalating adjustments as the alarm breaches further     | "AGG! Add ALL the instances!"                 |

**Cooldowns:** default **300 seconds**, the pause that lets new instances come up and absorb load before further scaling decisions. Applies to dynamic and (optionally) manual scaling – **not** scheduled scaling.

**Predictive scaling** uses ML on your traffic history to scale *ahead* of demand. You can use its forecast to tune your own policies, or opt out of data collection entirely.

### Compute Optimiser

ML-powered rightsizing fed by CloudWatch metrics. Covers EC2 instance types, Auto Scaling groups, EBS volume types/sizes, ECS-on-Fargate task sizing, and Lambda memory/CPU allocation. The "reduce cost without redesign" answer family.

### Event-driven architecture

The serverless backbone – know each player's role on sight:

**Lambda** (serverless compute) · **EventBridge** (event bus and rules) · **Step Functions** (orchestration) · **SQS** (buffer/decouple) · **SNS** (fan-out to subscribers) · **API Gateway** (entry point for external producers) · **DynamoDB** and **S3** (stateful endpoints that also emit events).

- **SNS** – topics (publish channel), subscriptions (endpoint per topic), protocols including HTTPS, email, SMS, SQS, Lambda, mobile push
- **SQS** - KMS-encrypted, transient, **4-day default / 14-day max retention**, optional FIFO, **256 KB** message ceiling (extended-client libraries can push larger payloads via S3)
- **Amazon MQ** - managed Apache ActiveMQ (JMS, NMS, MQTT, WebSocket); the drop-in answer when *migrating an existing broker*; use SQS/SNS for greenfield

**Lambda + SAM + EventBridge.** Lambda supports Node.js, Python, Java, Go, C#; stateless; scales without fundamental limits. **SAM** is CloudFormation's serverless extension - YAML templates, `sam deploy`, local testing via a Docker-based emulator. **EventBridge** connects AWS and third-party SaaS events to rules.

**Step Functions** turns your app into a state machine (Amazon States Language, JSON): tasks, sequences, parallels, branches, timers, with a visual console showing live execution. The workflow-engineer's tool of choice – see the comparison table below.

|                    | When                                             | Example                            |
|--------------------|--------------------------------------------------|------------------------------------|
| **Step Functions** | Out-of-the-box AWS service orchestration         | Order-processing flow              |
| ~~SWF~~ *(legacy)* | Historically: external/manual steps in workflows | Loan application with human review |
| **SQS**            | Store-and-forward queueing                       | Image resize pipeline              |
| **AWS Batch**      | Scheduled or recurring compute jobs, light logic | Rotate firewall logs nightly       |

*(SWF is legacy/closed to new customers – I'd drop it or footnote it; it survives in older question banks.)*

### Streaming with Kinesis

Sharded ingestion: **1,000 records/sec and 1 MB per record** per shard, default **500-shard limit**. Records carry a partition key, sequence number, and data blob. Retention: **24 hours by default, extendable to 365 days**

### DynamoDB scaling (continued from Section 1)

Section 1 covered partition math; here's the elasticity layer. **Auto Scaling** uses target tracking against a utilisation target, driven through Application Auto Scaling, and also applies to GSIs. **On-demand capacity** trades a price premium for instant, provision-nothing capacity – the answer when workloads are unpredictable or you can't tolerate scaling lag. **DAX** is the in-memory cache: microseconds-on reads for read-intensive, repeatedly-hit datasets. It's the *wrong* answer for write-heavy apps or apps that already cache client-side.

### Scaling containers

|                | ECS                                   | EKS                               |
|----------------|---------------------------------------|-----------------------------------|
| Model          | AWS-opinionated                       | Kubernetes-compatible             |
| Integration    | Native with Route 53, ALB, CloudWatch | Needs generalised LB abstractions |
| Learning curve | Simpler                               | Richer, steeper                   |
| Extensibility  | Limited                               | Broad third-party ecosystem       |
| Unit           | Tasks                                 | Pods                              |

Terminology transfer: ECS **tasks** ≈ K8s **pods**; lift-and-shift from an existing K8s estate points at EKS.

**Launch types.** EC2 (you provision, patch, optimise, retain control) vs. **Fargate** (serverless containers – AWS provisions and optimises; you trade granularity for zero-touch). Both are available on ECS and EKS, plus Outposts, Local Zones, and Wavelength.

Around them: **App Runner** (synchronous HTTP apps from source or images, scales to zero - proof-of-concept darling) and **AWS Batch** (plans, schedules, and runs jobs on ECS/EKS/Fargate, with spot-instance cost reduction – create a compute environment, a job queue with priorities, a job definition, then schedule).

### Big data at scale: EMR

Managed Hadoop – also Spark, HBase, Presto, Flink – for log analysis, financial analysis, ETL. An EMR **step** is a programmatic unit of work; a **cluster** is the EC2 fleet assembled to run steps. The Zoo bestiary (Ambari, Sqoop, Flume, ZooKeeper, Oozie, Pig, Hive, Mahout, HBase, MapReduce, HDFS) still appears in distractor answers – recognise the names even if the ecosystem has largely moved to Spark-on-EMR.

### Visualising at scale: QuickSight and OpenSearch

**QuickSight** - serverless, pay-per-user BI, powered by the SPICE in-memory engine; builds hybrid datasets across Athena, Redshift, S3, RDS, Timestream, Snowflake, Salesforce, and dozens more; ML-driven natural-language querying and anomaly detection. The "dashboard for business users, cheaply" answer.

**OpenSearch** - search + real-time analytics, OpenSearch Dashboards (née Kibana); richer querying and lower cost than CloudWatch for log-heavy workloads – the trade being more operational ownership.

## 6. Business Continuity & Disaster Recovery{#business-continuity-and-disaster-recovery}

This is the domain where the exam stops asking "what does this service do" and starts asking "how much does it hurt when things break." Two numbers rule everything here, so we start with them.

### Vocabulary – and why RTO/RPO are the whole game

- **Business continuity (BC)** – minimising disruption to business activity when something unexpected happens
- **Disaster recovery (DR)** – the act of responding to an event that threatens BC
- **High availability (HA)** – designing in redundancy to *reduce the chance* of impacting service levels
- **Fault tolerance (FT)** – designing in the ability to *absorb* failures without impacting service levels
- **SLA** – an agreed target for a service's performance or availability
- **RTO** – time to restore business processes after a disruption. *T is for Time.*
- **RPO** – acceptable amount of data loss, measured in time. *P is for the data that goes "poof."*

Nearly every DR scenario question resolves to one of these two numbers: "we can lose up to 15 minutes of data" → RPO; "we must be serving customers within an hour" → RTO. The architecture you pick is downstream of those two figures – and, always, of the budget.

### What counts as a disaster

| Category              | Example                                              |
|-----------------------|------------------------------------------------------|
| Hardware failure      | Switch power supply dies, LAN goes down              |
| Deployment failure    | A patch breaks a key ERP process                     |
| Load-induced          | DDoS on the website                                  |
| Data-induced          | Ariane 5 explosion, June 4 1996                      |
| Credential expiration | An SSL/TLS certificate quietly lapses                |
| Dependency            | An S3 subsystem failure cascades into other services |
| Infrastructure        | Backhoe through the fiber line                       |
| Identifier exhaustion | "Insufficient capacity in requested AZ"              |

Note how many of these are *self-inflicted* (deployment, credential, identifier) – a hint that DR planning isn't only about acts of God.

### The HA continuum – the four DR patterns

The core of the section. Same workload, four escalating price/risk trade-offs - the exam gives you an RTO/RPO and a budget, and you place the architecture:

| Pattern                        | What's running                            | Failover                             | RTO ballpark  | Cost    |
|--------------------------------|-------------------------------------------|--------------------------------------|---------------|---------|
| **Backup & Restore**           | Nothing - just data in S3/Glacier         | Rebuild everything from scratch      | Hours         | Lowest  |
| **Pilot Light**                | Core data replicated; everything else off | Spin up the rest on demand           | Minutes–hours | Low     |
| **Warm Standby**               | A scaled-down full environment, always on | Scale up, cut over                   | Minutes       | Medium  |
| **Multi-Site (active/active)** | Full production capacity everywhere       | Near-instant, little/no intervention | Seconds       | Highest |

Backup-and-restore is the most common entry point into AWS and needs minimal configuration; pilot-light keeps AMIs synchronised with on-prem counterparts and usually needs manual failover; warm-standby doubles nicely as a testing/staging shadow environment but needs scaling to take production load; multi-site is "effectively a mirrored data center" and can feel wasteful precisely because it's paying for insurance.

### Storage HA, per service

**EBS.** Annualised failure rate under 0.2% (vs. ~4% for commodity HDDs), availability target of 99.999%, replicated automatically **within a single AZ** – which means the AZ itself is the failure domain. Snapshots go to S3 and can be copied across regions; volumes support RAID.

| Config     | Volume size (×each) | IOPS/volume | Total IOPS | Usable space | Throughput     |
|------------|---------------------|-------------|------------|--------------|----------------|
| None       | 1,000 GB            | 4,000       | 4,000      | 1,000 GB     | 500 MB/s       |
| **RAID 0** | 500 GB × 2          | 4,000       | **8,000**  | 1,000 GB     | **1,000 MB/s** |
| **RAID 1** | 500 GB × 2          | 4,000       | 4,000      | 500 GB       | 500 MB/s       |

The one-line takeaway: RAID 0 buys performance (at zero redundancy); RAID 1 buys redundancy (halving usable space). Both occasionally appear as the answer to "I need more IOPS than one volume can give me."

**S3.** Durability is **eleven 9s (99.999999999%)** across storage classes; availability differs: Standard at 99.99% (~52.6 min/year), Standard-IA at 99.9% (~8.8 hrs/year), One Zone-IA at 99.5% (~1.8 days/year) - and One Zone-IA drops the multi-AZ replication, which is exactly why its availability figure slides. Standard and Standard-IA replicate across AZs. Also the backing store for EBS snapshots and half of AWS's internals – remember dependency-failure risk cuts both ways.

**EFS.** True file system semantics (locking, strong consistency), files and metadata spread across multiple AZs, concurrently mountable everywhere – the anti-EBS.

**Storage Gateway / Snowball / Glacier.** Respectively: continuous sync for offsite backup, batch transfers only, and safe long-term archival with rare retrieval.

### Compute HA

Keep AMIs current – a stale AMI is a slow RTO. AMIs copy cross-region for DR staging. Horizontal architectures spread risk across smaller machines; **Reserved Instances get launch priority** in an AZ, and **On-Demand Capacity Reservations guarantee** an instance type in a specific AZ – the "my failover region must actually have capacity" answer. Route 53 health checks provide self-healing DNS redirection.

### Database HA

- **DynamoDB** - data and traffic spread across partitions, synchronously replicated across **three AZs** per region; **Global Tables** add multi-region active-active
- **RDS** - Multi-AZ standby, read replicas (region and cross-region), snapshot recovery
- **Aurora** - multi-AZ copies by design; **Global Databases**: one primary region, up to **5 secondary regions**, storage-level replication with sub-second typical lag, promotable during an outage
- **Redshift** - historically single-node meant restore-from-snapshot on failure and the multi-node cluster was the HA answer; multi-AZ support arrived on **RA3 nodes**

### Networking HA

Subnets across AZs create multi-AZ VPC presence by construction. **Two VPN tunnels** to a Virtual Private Gateway is best practice – AWS gives you both for a reason. **Direct Connect is not HA by itself**: redundancy needs a second DX connection or a VPN fallback (tying back to Section 4's BGP cutover). Route 53 health checks redirect DNS at the layer below your load balancers; Elastic IPs let you swap backing assets without touching DNS. And per Section 2: one **NAT Gateway per AZ**, with private-subnet routes pinned to their local gateway.

### FMEA – scoring your risks

Failure Mode and Effects Analysis is the formal process the exam occasionally asks about: for each failure mode, assess what could go wrong, its impact, its likelihood, and your ability to detect it. The score:

> **Risk Priority Number = Severity × Probability × Detection**

Interesting subtlety worth knowing: a higher *detection* score means *harder to detect* – so a severe, probable failure that hides quietly scores worst. That inversion trips people up.

## 7. Deployment & Operations{#deployment-and-operations}

Two halves: *how you ship changes* (deployment patterns, CI/CD, Elastic Beanstalk, CloudFormation) and *how you run what you shipped* (Systems Manager, Config, enterprise apps). The exam weights both, with a particular fondness for choosing the right deployment strategy for a given downtime tolerance.

### Deployment patterns

Four ways to roll out v2, in ascending order of safety:

- **Rolling** - swap instances batch-by-batch behind your ELB; no downtime, but you're briefly running mixed versions
- **A/B testing** – split traffic by percentage between v1 and v2 to compare behavior
- **Canary** – a *single* v2 instance alongside the fleet; smallest possible blast radius
- **Blue-Green** - stand up the *entire* v2 environment, verify, then flip traffic wholesale

Blue-green implementations the exam expects you to recognise: DNS cutover to a new ELB, swapping the ASG behind the ELB, updating the launch configuration and rolling the ASG, swapping Elastic Beanstalk environment URLs, or cloning an OpsWorks stack (historical flavor).

**Blue-green's contraindication** – the detail that separates the pros: if your data-store schema changes are tightly coupled to code changes, or the upgrade needs special routines run mid-deployment, blue-green breaks down, because both versions must be simultaneously compatible with one database.

### CI/CD on AWS

- **CI** - merge to main frequently
- **CD (delivery)** – automated release process, deploy on demand
- **CD (deployment)** – every change ships to production with no human intervention

The toolchain: **CodeCommit** (managed Git) · **CodePipeline** (orchestration) · **CodeBuild** (compile/test/package) · **CodeDeploy** (deploy to EC2, Beanstalk, ECS, Lambda) · **Cloud9** (cloud IDE - now closed to new customers, but still fair game as a legacy answer) · **CodeGuru** (automated code review) · **CodeStar** (prebuilt CI/CD ecosystems) · **X-Ray** (distributed tracing) · **CodeArtifact** (package management).

### Elastic Beanstalk

Orchestration for scalable web apps – Docker, PHP, Java, Node, and more – with multiple environments (dev/QA/prod) per application. Ease of deployment at the cost of control. Its deployment options table is prime exam material:

| Option                     | What                                                              | Downtime | Rollback                   |
|----------------------------|-------------------------------------------------------------------|----------|----------------------------|
| All-at-once                | Update every existing instance simultaneously                     | **Yes**  | Manual                     |
| Rolling                    | Batches through existing instances                                | No       | Manual                     |
| Rolling + additional batch | Adds new-version instances *before* retiring old ones             | No       | Manual                     |
| Immutable                  | Fresh ASG of new instances; cutover only after health checks pass | No       | Terminate new instances    |
| Traffic splitting          | Route a percentage to new instances for canary testing            | No       | Reroute DNS, terminate new |
| Blue/green                 | Full new environment, swap CNAME                                  | No       | Swap URL back              |

*Pattern to notice:* immutable and traffic-splitting are the safest answers; all-at-once is the trap answer for any scenario that mentions uptime.

### CloudFormation

Infrastructure as code: **templates** (JSON/YAML) describe environments; **stacks** create/update/delete them atomically; **change sets** preview proposed changes before you commit; **StackSets** deploy across multiple accounts and regions. Over 300 resource types, custom resources via SNS or Lambda.

**Stack policies** are the exam favorite: protect specific resources from accidental update/deletion. Applied at creation via console or CLI; adding one to an existing stack is **CLI-only**; once applied it can't be removed – only modified via CLI. Best practices: use Change Sets to spot trouble, make changes through CloudFormation rather than clicking the console, keep templates in version control.

(Don't forget the abstractions layered on top: **CDK**, **SAM** from Section 5, and third-party frameworks like Terraform.)

### Running the fleet: Systems Manager

This is the heart of modern ops on AWS:

| Capability          | What                                                 | Example                                        |
|---------------------|------------------------------------------------------|------------------------------------------------|
| Inventory           | Collect OS/app/instance metadata                     | "Which instances run Apache 2.2.x or earlier?" |
| State Manager       | Desired-state associations                           | Track which instances got patched to stable    |
| Parameter Store     | Shared secure config storage                         | Pull RDS credentials at boot                   |
| Maintenance Windows | Scheduled patch/script windows                       | 00:00–02:00 for Patch Manager                  |
| Automation          | Routine maintenance via automation documents         | Stop dev/QA instances Friday night             |
| Run Command         | Execute commands without SSH/RDP                     | Shell script across 53 instances at once       |
| Patch Manager       | Fleet-wide patching                                  | Keep everyone at the same patch level          |
| Insight Dashboards  | Account-level Config/CloudTrail/Trusted Advisor view | Single compliance viewport                     |
| Resource Groups     | Tag-driven groupings                                 | Dashboard for all Production ERP assets        |

Plus **SSM documents**: command (Run Command/State Manager), policy (enforced state), and automation documents. The SSM agent ships on AWS base AMIs.

**AWS Config** (covered fully in Section 3's security posture) – here it plays configuration management: baselines, drift tracking, compliance rules. One mental slot: *Config evaluates, SSM remediates.*

### API Gateway

Backends via Lambda, AWS-service proxies, or any HTTP endpoint; regional, private, or edge-optimised (CloudFront-backed) deployments. API keys and usage plans give you identification, throttling, and quotas; custom domains with SNI work; and APIs can even be monetised on the AWS Marketplace.

### Enterprise end-user apps

- **WorkSpaces** (managed DaaS) and **AppStream 2.0** (application streaming) – regulated industries, remote/seasonal workers, product demos without installs
- **Connect** - managed cloud contact center: call handling, IVR, chatbots, analytics, CRM integration
- **Chime** - meetings and video conferencing
- ~~**WorkLink**~~ (retired) and ~~**Alexa for Business**~~ (deprecated) – dropped; they occasionally haunt older practice questions only

### Machine learning service tiers

The three-layer taxonomy is the exam-relevant frame – know *which tier* a scenario implies:

1. **AI services** (app developers, no ML background): Comprehend (NLP/sentiment), Lex (chatbots), Polly (text-to-speech), Rekognition (image/video analysis), Translate, Transcribe (speech-to-text), Textract (document text extraction), Personalise (recommendations), Forecast (time-series prediction)
2. **ML services** (data scientists): SageMaker end-to-end, Ground Truth, training/hosting, Marketplace
3. **Frameworks & infrastructure** (researchers): Deep Learning AMIs, Greengrass, Gluon/Keras, MXNet/TensorFlow

Scenario-matching examples worth keeping: sentiment analysis of social posts → Comprehend; upsell recommendations at checkout → Personalise; digitising paper forms → Textract; seasonal demand forecasting → Forecast.

### IoT

The remaining, current core: **IoT Core** (MQTT message broker to/from devices), **IoT Events** (conditional logic on sensor data), **Greengrass** (deploy Lambdas/Docker/ML to edge devices), **Device Management** (registry, groups, OTA firmware), **Device Defender** (ML-based configuration auditing), **IoT Analytics** (aggregation, time-series SQL), **SiteWise** (edge data collection and modeling). Retired: **Things Graph** and **1-Click** - footnote only.

---

## 8. Cost Management{#cost-management}

The shortest domain, but never skipped - FinOps-flavored questions appear reliably, and they're fast points if the vocabulary is cold.

**CapEx vs. OpEx.** Capital expenditure buys long-term assets (buildings, hardware); operational expenditure is ongoing, usually variable spend. Cloud shifts the needle from CapEx to OpEx – often the *business* motivation behind the technical migrations in Sections 4–6.

**TCO and ROI.** Total cost of ownership captures the full cost model of a decision; return on investment is what you get back within a timeframe. If a scenario asks "how do we justify this migration to finance," the answer starts here.

**Cost optimisation levers**, roughly in order of exam frequency:

1. **Right-sizing** - lowest-cost resource meeting specs; CloudWatch utilisation data drives it; loosely coupled architectures smooth demand enough to size against
2. **Purchase options** – Reserved Instances for steady workloads (Standard, Convertible, Scheduled), Spot for interruptible horizontal scale, **EC2 Fleet** for blended On-Demand/RI/Spot
3. **Appropriate provisioning** - don't over-provision; consolidate; watch utilisation
4. **Managed services** – RDS over self-run MySQL; Fargate; the operational savings *are* the saving
5. **Geographic selection** – pricing varies by region; pair with Route 53/CloudFront to counteract latency
6. **Optimised data transfer** - egress and inter-region traffic add up; Direct Connect can win at volume

**RI mechanics** – the details: attributes are instance type, platform, tenancy, and (optionally) AZ; zonal RIs reserve capacity in an AZ, regional RIs apply discounts anywhere in the region (and can be converted zonal→regional). Shared across consolidated billing, sellable on the RI Marketplace.

**Dedicated Instances vs. Dedicated Hosts.** Dedicated Instances run on hardware dedicated to *you* but may share hardware with your other non-dedicated instances – a ~$2/hr/region premium, purchasable as On-Demand, RI, or Spot. Dedicated Hosts are entire physical servers you control placement on - the answer for per-core/per-socket licensing – and each host runs only one instance family/type.

**Spot mechanics:** specify a max price; instances stop/terminate/hibernate when the market exceeds it; requests can be one-time, maintained, or duration-based.

**Tagging** is the cheapest cost-management win and the bridge into governance: name/value pairs on every resource, powering cost allocation, IAM conditions, and automation. Enforce via Config rules or scripts (e.g., untagged EC2 instances stop nightly). **Resource Groups** then turn tags into custom consoles – consolidated views by environment, project, or cost center.

**The reporting toolkit:**
- **Cost and Usage Reports (CUR)** – CSV granularity down to hourly, analysable via Athena/Redshift/QuickSight
- **Budgets** – alerts as you approach limits, based on cost, usage, or RI utilisation/coverage
- **Consolidated Billing** – one payer account, economies of scale across the org
- **Trusted Advisor** – automated checks (core checks free; full set behind Business/Enterprise support) recommending optimisations like RIs and scaling adjustments

---

## Closing: how I studied, and what I'd tell you

Now the part the intro promised. Coming from the Advanced Networking Specialty, the biggest adjustment wasn't depth – it was *breadth*. The Specialty asks you to know a few things deeply; the Pro exam asks you to know everything broadly and choose *correctly under constraints*. Three study habits made the difference:

**Decision tables over definitions.** I ended up reconstructing my notes as comparison tables – you've just read the result. When a practice question described a scenario, I wanted the answer to be a row lookup, not a recall struggle. If you take one thing from this post, rebuild your own tables by hand; the rebuilding *is* the studying.

**Read the constraint sentence twice.** Pro questions are long, and the discriminating fact - "no downtime acceptable," "must encrypt at rest with customer-held keys," "budget is fixed" - is almost always a single clause. I lost more practice questions to skimming than to ignorance.

**Check deprecations before the exam, not after.** I tripped over more retired services in practice exams than in any other category of error. Anything here that's still marked legacy - worth thirty minutes with the AWS "service history" pages the week before your exam.

I used the Pluralsight Solutions Architect – Professional path as the spine, these notes as the skeleton, and practice exams for calibration. If you're coming from an Associate cert rather than another Specialty, budget extra time for the multi-account organisation and DR material – those are the topics with no Associate-level counterpart.

Good luck – and if these notes helped, pay it forward like every blog post that helped me did.
