AWS Solutions Architect - Professional

·45 min read·Graham Mace
AWS Solutions Architect - Professional

After earning the Advanced Networking Specialty, the Solutions Architect – Professional exam was the natural next rung on the ladder – and, honestly, one of the more demanding ones. It’s less about memorising services and more about knowing which service trade-offs to recommend when a scenario hands you a broken architecture and five ways to fix it.

These notes come from working through the AWS Solutions Architect – Professional path on Pluralsight. They’re organised around the exam’s major themes: data stores, networking, security, migration, scaling, resilience, operations, and cost. I’ve kept them flash-card dense on purpose – that’s how I learn – but I’ve flagged the traps and decision points that seem to show up repeatedly.

Whether you’re coming from the Associate level or another specialty, I hope these save you some late nights.

Before you dive in: This post reflects the exam as I studied for it in 2024. AWS rotates exam content regularly - always cross-check the official exam guide before scheduling.

Contents

  1. Data Stores
  2. Networking
  3. Security
  4. Migration Strategies
  5. Architecting to Scale
  6. Business Continuity & Disaster Recovery
  7. Deployment & Operations
  8. Cost Management

1. Data Stores

This domain rewards you for thinking in terms of data characteristics – persistence, consistency, access patterns – before picking a service. Almost every scenario question here is really asking “which of these five services fits this workload shape?”

Core concepts worth having cold

Persistence model. Persistent (survives anything – RDS, Glacier), transient (passed along and gone – SQS, SNS), ephemeral (gone on stop - instance store, Memcached). Expect at least one question where the cheapest correct answer hinges on recognising you don’t need persistence.

IOPS vs. throughput. IOPS is how fast you can read/write; throughput is how much at once. This distinction drives EBS volume-type questions and RAID configurations later.

ACID vs. BASE. Atomicity, Consistency, Isolation, Durability vs. Basically Available, Soft state, Eventual consistency. This is the classic relational-vs-NoSQL decision axis, and the exam leans on it hard.

S3

The exam’s S3 questions almost always come down to three things: limits, encryption, and lifecycle behavior. Knowing the numbers cold saves you real exam time.

Limits. Max object size is 5 TB; the largest object in a single PUT is 5 GB, so anything bigger needs multipart upload (AWS recommends it from 100 MB).

Security model. Resource-based (bucket policies, object ACLs) and user-based (IAM policies), both evaluated together. Optional MFA delete for extra protection on destructive operations.

Data protection.

  • Versioning – new version on every write; enables rollback and undelete. Old versions keep billing until permanently deleted, and versioning is a prerequisite for cross-region replication. Expect a question pairing versioning + lifecycle rules.
  • Cross-region replication – for compliance, latency, or durability. Remember it requires versioning enabled on both sides.
  • Storage classes – Standard, Standard-IA, One Zone-IA, Intelligent-Tiering, Glacier, and Glacier Deep Archive. (Reduced Redundancy is effectively deprecated – see the callout below.) Intelligent-Tiering now extends to automatic archival tiers for infrequently and rarely accessed data.
  • Lifecycle management – transitions between classes and expirations on a schedule; the tool of choice for “optimise storage costs” and “adhere to retention policy” scenario answers.

Analytics integrations. S3 as a data lake (Athena, Redshift Spectrum, QuickSight), as a Kinesis Firehose landing zone for IoT streaming data, and as ML/AI storage backing Rekognition, Lex, and MXNet.

Encryption at rest – memorise all four variants and, more importantly, who manages the key:

  • SSE-S3 - AWS manages the keys, AES-256, simplest option
  • SSE-KMS - keys generated and managed in KMS, adds an audit trail via CloudTrail
  • SSE-C - you upload your own AES-256 key with each request
  • Client-side - you encrypt before upload; AWS never sees plaintext

Exam tricks. Transfer Acceleration (CloudFront in reverse for fast uploads), Requester Pays buckets, object tags for cost allocation and security policy, event notifications to SNS/SQS/Lambda, and static website hosting.

Deprecated but still quizzed: Reduced Redundancy Storage is no longer recommended – AWS positions Standard for nearly everything it covered. You may still see it in older practice exams; treat “RRS is the answer” answers as suspect.

Glacier

Glacier is cheap, cold, and deliberately awkward to retrieve from. The things worth remembering: archives are immutable, max 40 TB, typically zip/tar, stored in vaults.

Vault Lock is the star feature for compliance scenarios: write a lock policy (e.g., “no deletes ever” or “MFA required to delete”) that is itself immutable once confirmed. The flow has a deliberate safety catch – initiate the lock, you get a 24-hour window to either complete or abort it, and if you do nothing within 24 hours the lock simply never takes effect.

Elastic Block Store (EBS)

Think “virtual hard drive, but with exam-flavoured gotchas.” EBS volumes attach only to EC2, and are locked to a single AZ – an instance can’t mount a volume from another AZ. That constraint is the setup for half the snapshot-related questions.

Snapshots are how you escape those constraints: share data sets across users or accounts, migrate a system to a new AZ or region, and convert an unencrypted volume into an encrypted one. Data Lifecycle Manager schedules snapshots on your chosen interval with retention rules to prune stale ones.

Elastic File System (EFS)

NFS-as-a-service. Metadata and data are replicated across multiple AZs - the direct contrast with EBS’s single-AZ limitation, and the reason “shared file storage across AZs” answers point here. You can mount it from EC2 instances across many AZs, and even from on-premises systems (with caution – that’s usually the wrong answer; for on-prem integration, look at Amazon DataSync or the Storage Gateway family instead).

FSx

Distributed, managed file systems that aren’t plain NFS – the names basically hand you the answer in the exam. Four flavours:

  • FSx for Windows File Server – SMB, Active Directory-integrated; the answer whenever you see Windows LoB apps, Windows CMSs, media processing, or Windows analytics workloads
  • FSx for Lustre - HPC; scales reliably to thousands of concurrent EC2 instances; big data and ML. Watch for it paired with S3-backed data sets
  • FSx for OpenZFS and FSx for NetApp ONTAP - round out the family; ONTAP shows up in multiprotocol enterprise scenarios

Storage Gateway

A VM (VMware/Hyper-V) or hardware appliance on-premises, backed by S3 and Glacier – the classic first step for hybrid cloud, DR, and migration scenarios. Know the four modes cold, because the exam tests “which gateway type?” constantly:

ModeInterfaceBehavior
File GatewayNFS, SMBOn-prem/EC2 clients store and retrieve objects in S3 via a file interface
Volume – Stored modeiSCSIPrimary data stays on-prem, asynchronously replicated to S3
Volume – Cached modeiSCSIPrimary data lives in S3, frequently accessed data cached locally
Tape GatewayiSCSIVirtual tape library/media changer for existing backup software

Memory hook: Stored = the truth is stored on-prem; Cached = the truth is in the cloud, cache is local.

RDS

Managed MySQL, MariaDB, PostgreSQL, MSSQL, Oracle, and Aurora. Drop-in replacement for on-prem relational databases, with automated backups, patching, push-button scaling, replication, and redundancy thrown in.

When RDS is the wrong answer - the antipatterns are practically a mandatory exam question:

  • Lots of BLOBs → S3
  • Name/value or unpredictable data structures → DynamoDB
  • Automated scalability beyond what RDS offers → DynamoDB
  • IBM DB2 or SAP HANA → EC2 (not RDS-supported)
  • Need complete OS-level control → EC2

Replication model. Multi-AZ gives you a synchronous standby in another AZ; read replicas serve read-heavy workloads (and can span regions). If one AZ fails, the standby is promoted and replicas keep working. If a whole region fails, promote a read replica to standalone and reconfigure it as Multi-AZ. Note for MySQL: MyISAM doesn’t replicate – you need InnoDB (or XtraDB on Maria).

Aurora

Fully managed, MySQL- and PostgreSQL-compatible, multi-AZ by design, with storage-level replication. The details the exam loves:

  • Endpoints – a cluster (writer) endpoint and a reader endpoint, so applications don’t track individual instances
  • Global Databases – one primary region, up to 5 secondary regions, replication at the storage layer with typical lag under a second; secondary regions can be promoted during an outage
  • Serverless - billed in ACUs (Aurora Capacity Units), scaling between 0.5 and 128 ACUs; ideal for variable or infrequent workloads with low operational overhead

DynamoDB

Multi-AZ NoSQL with optional global tables for cross-region replication, priced on throughput (RCUs/WCUs), with provisioned (auto-scaled between your min/max), and on-demand capacity modes. Defaults to eventual consistency, but you can request strongly consistent reads via an SDK parameter. ACID guarantees arrive via DynamoDB Transactions.

The heart of the DynamoDB questions is indexes – specifically which kind solves the query you’re being asked about:

IndexKeysWhen
Global Secondary IndexPartition key and sort key can differ from the table’sQuery efficiently on attributes outside the primary key (e.g., look up orders by customer number)
Local Secondary IndexSame partition key, different sort keyYou already know the partition key and want to filter/refine by another attribute (e.g., all orders for this customer with a given material number)

Memory hook: a Local Secondary Index has to stay local – it respects the table’s partition key. A Global one can go anywhere.

Projection counts too. Sparse projections (keys + a few attributes) cost the least and give the lowest latency for narrowly shaped queries; projecting everything doubles storage and write costs but maximises flexibility. And rare queries on frequently written data → keys-only projection to keep writes fast.

Capacity math (know it, it appears in calculation questions):

  • Partitions by capacity: (total RCU ÷ 3,000) + (total WCU ÷ 1,000)
  • Partitions by size: total size ÷ 10 GB
  • Actual partitions = round up MAX(capacity, size)

Example: 2,000 RCU / 2,000 WCU on a 10 GB table → MAX(0.67 + 2.0, 1.0) = 2.67 → 3 partitions, each carrying a share of the throughput.

DocumentDB

MongoDB-compatible NoSQL JSON document store with AWS-managed scaling and write-then-read consistency. Think user profiles, content management, real-time big data.

Redshift

Petabyte-scale managed data warehouse, PostgreSQL-compatible via JDBC/ODBC, built on columnar storage and massively parallel processing for complex analytical queries. Redshift Spectrum queries data directly in S3 without loading it. Data-lake architecture questions – “query raw data with minimal pre-processing,” “shorten time from collection to insight” – point here.

AWS Glue

Serverless ETL. Crawlers discover data (S3, Redshift, RDS, databases on EC2), populate the Data Catalog, which feeds Athena, EMR, and Redshift, and powers ETL jobs that land data in S3, Redshift, or Lake Formation. Glue Data Quality layers on predefined rules, metrics, and alerts to keep the pipeline trustworthy. Know the crawler → catalog → job flow diagram; it shows up in data-lake questions.

Neptune

Fully managed graph database speaking both Gremlin and SPARQL. Any scenario about relationships - fraud rings, social networks, knowledge graphs - is Neptune.

ElastiCache

The final answer to “how do I reduce database load / survive instance loss / power a leaderboard?”

Use caseRight choiceWhy
Web session storeRedisSessions survive the loss of any one load-balanced web server
Database caching in front of RDSMemcachedSimple, multicore, caches popular query results
LeaderboardsRedisSorted sets give live rankings at scale
Streaming dashboards-Sensor data landing spot feeding real-time displays

Memcached vs. Redis is a reliable exam discriminator. Memcached is straightforward, scales in and out, uses multiple cores, and is object-cache-shaped. Redis brings encryption, HIPAA eligibility, clustering, complex data types, replication/HA, pub/sub, geospatial indexing, and backup/restore. Rule of thumb: if the question implies any durability, availability, or richness requirement, it’s Redis.

Athena

Serverless SQL over S3, built on Presto. Convert your data to Parquet for a dramatic performance and cost win – that pairing is practically a guaranteed exam answer. Athena also runs Spark code, and federated queries reach 25+ data sources.

Athena vs. Redshift Spectrum: choose Athena when the data lives mostly in S3 and stands alone; choose Spectrum when you need to join S3 data with existing Redshift tables.

Quantum Ledger Database (QLDB)

Not quite blockchain – a centralised, append-only, immutable ledger with a verifiable hash chain, which centralisation makes fast and scalable. Every record contributes to the integrity of the chain. Use when you need provable transaction history with an audit log - think banking records - without multi-party consensus.

QLDB vs. Amazon Managed Blockchain: QLDB is centralised and single-owner; Managed Blockchain is a distributed, consensus-based framework (Hyperledger Fabric or Ethereum) involving multiple network members and nodes across AWS accounts. The exam’s ledger questions hinge entirely on that distinction.

Timestream

Purpose-built time-series database - sensor networks, industrial machinery, equipment telemetry – with built-in analytics like interpolation and smoothing. The alternative when DynamoDB or Redshift feel heavyweight for timestamped telemetry.

OpenSearch

Mostly a search engine, though it can store documents (with caution). AWS’s managed flavor is the ELK-ish stack: OpenSearch for search/storage, Logstash (or CloudWatch/Firehose/IoT) for intake, and dashboards via OpenSearch Dashboards (née Kibana). You’ll still hear people say “Elasticsearch” – same lineage, renamed.

2. Networking

If the Pro exam has a “home turf,” this is it. Networking questions dominate, and they’re rarely “what is X?” – they’re “given this connectivity requirement and these constraints, which option?” The decision matrices below (VPN vs. Direct Connect vs. Transit Gateway, placement groups, Route 53 policies) are where I’d spend the most review time.

Foundations worth reviewing fast

The OSI model – you likely know it, but the exam frames load balancer and firewall questions by layer, so keep the mnemonic handy: Please Do Not Throw Sausage Pizza Away (Layers 7→1). Remember ALB = Layer 7, NLB = Layer 4, WAF = Layer 7, Network Firewall = Layers 3–4.

Unicast vs. multicast. AWS doesn’t support multicast at the network level – that’s occasionally the entire trick behind a question.

TCP / UDP / ICMP. TCP (L4) is connection-based and stateful with acknowledgements; UDP (L4) is connectionless and loss-tolerant (streaming, DNS); ICMP (L3) is the routers’ own health check language (ping, traceroute). Know which protocol belongs to which layer - hybrid questions love mixing L3/L4/L7 facts.

Ephemeral ports. The suggested range is 49152–65535. Linux typically uses 32768–61000 and Windows starts around 1025. This matters for security groups and NACLs: if you write rules that forget return traffic on ephemeral ports, that’s a classic exam trap.

Reserved IPs in every subnet. Five addresses, always: network address (.0), VPC router (.1), Amazon DNS (.2), future use (.3), broadcast (.255). A /28 subnet gives you 16 addresses → only 11 usable. Capacity-planning questions exploit this constantly.

Connecting your network to a VPC

This is the highest-yield table in the whole domain. Learn the “when” column - the exam scenario always maps to one of these:

OptionWhat it isWhen it’s the answerTrade-off
AWS Managed VPNIPsec VPN over the internetQuick setup, or a redundant backup for Direct ConnectDepends on internet quality
Direct ConnectPrivate dedicated circuit into AWSLarge, steady bandwidth needs; predictable performance; up to 10 Gbps; BGP requiredNo built-in redundancy; may need telecom partners
DX + VPNIPsec tunnel over the private circuitCompliance-driven encryption on top of DXAdded complexity
VPN CloudHubHub-and-spoke VPN between multiple remote sites via a Virtual Private GatewayLinking branch offices; can prefer MPLS with VPN as backupInternet-dependent, no inherent redundancy
Software VPNYou run both endpointsCompliance scenarios where you must control the whole chain, or unsupported VPN techYou own all redundancy
Transit VPCLegacy pattern for multi-region meshMostly historical - Transit Gateway replaced itYou maintain it

Migration tip: when transitioning from VPN to Direct Connect, run both within the same BGP prefix advertisement – from AWS’s side the DX path is always preferred, making the cutover nearly seamless.

VPC-to-VPC connectivity

Options: VPC peering, Transit Gateway, or the VPN/software approaches above.

VPC Peering. Uses AWS-managed networking, non-transitive - every pair must be explicitly connected, which becomes painful at scale (“N×(N−1)/2 peering connections” is a classic exam calculation).

Transit Gateway. The modern answer for scale: a regional hub connecting up to 5,000 attachments across accounts; cross-region requires peering multiple TGWs; can associate a DX Gateway. Watch for the CIDR overlap constraint - overlapping CIDRs can’t be peered, period.

Internet access patterns

Internet Gateway. Horizontally scaled, redundant by default, no bandwidth constraint. A subnet with a route to an IGW = public subnet. It performs NAT for instances with public IPs – but not for private-IP-only instances, which is the setup line for NAT devices.

Egress-only Internet Gateway – the IPv6 equivalent of NAT. Since IPv6 addresses are globally unique/public by default, this gives outbound-only access while blocking inbound initiation. Stateful. Route ::/0 at it.

NAT Instance vs. NAT Gateway – know this table cold; it’s near-guaranteed:

NAT GatewayNAT Instance
AvailabilityManaged, highly available within its AZOn you
BandwidthScales up to 45 GbpsBounded by instance type
MaintenanceAWS-managedYours
Public IPElastic IP, can’t be detachedElastic IP, detachable
Security groupsCannot attachCan use them (and can be a bastion)
RestrictionCan’t be used for peering/VPN/DX traffic - those need explicit routesRoutes are yours

Two placement facts that anchor scenario questions: NAT Gateways live in a public subnet, and for multi-AZ resilience you deploy one per AZ with private-subnet routes pointing at their local gateway. (Note: single-AZ NAT Gateways remain a notable exam gotcha – an AZ failure takes out any private instances routed through that gateway unless you’ve architected per-AZ.)

Routing

Every VPC has an implicit router and a main route table; the most specific route wins. That principle resolves half the tricky routing questions – especially when a NAT route (0.0.0.0/0) competes with a more specific peering or DX route.

BGP. Dynamic routing; required for Direct Connect, optional for VPN. Needs TCP port 179 plus ephemeral ports. ASNs identify endpoints; weight is local to the router with higher weight preferred for outbound traffic.

Enhanced Networking & Placement Groups

Enhanced Networking uses SR-IOV (Single Root I/O Virtualization) for higher performance – may need a driver on non-Amazon-Linux HVM AMIs. The Intel 82599 VF interface delivers 10 Gbps; the Elastic Network Adapter (ENA) reaches 25 Gbps.

Placement groups are a favorite exam topic – the three-way contrast:

ClusterSpreadPartition
LayoutTight low-latency grouping in one AZInstances on distinct underlying hardwareInstances divided into rack-level partitions
Use whenMaximise network throughput (HPC)Minimise simultaneous hardware failure for small groupsLarge distributed systems (HDFS, Cassandra) worried about correlated rack failure
LimitsFinite capacity - launch everything upfrontMax 7 instances per group per AZNot for Dedicated Hosts

Memory hook: cluster = speed, spread = “don’t share hardware,” partition = “don’t share racks.”

Private connectivity to services without traversing the internet, and - crucially - without requiring VPC peering. Use cases: marketplace SaaS consumption, exposing services to other VPCs with tight control, third-party app access with simplified networking. Contrast with peering whenever a scenario says “only need to reach one service, not the whole VPC.”

Global Accelerator

Users hit AWS edge locations and ride the AWS private backbone the rest of the way. You get static anycast IPs, performance gains, and resilience. Use cases: improving global app performance, hybrid networking (cheaper middle ground than DX), origin masking/DDoS defense (entry points protected by Shield), and multi-region failover across ~10 regional endpoints. Key contrast: CloudFront caches content; Global Accelerator accelerates TCP/UDP traffic to your own endpoints.

Route 53 routing policies

Straightforward table, but the exam disguises the differences in scenario prose:

PolicyThinking
Simple“Here’s the destination”
FailoverPrimary failed its health check → route to the secondary
GeolocationWhere is the user? → route by geography (hard boundaries)
GeoproximityWhere is the user relative to my resources? → route/bias by proximity
LatencyWhich endpoint gives the lowest latency → route there
Multivalue answerReturn several healthy IPs (poor man’s load balancer)
WeightedSplit traffic by percentage across resources

Geolocation vs. latency is the classic confusion pair: geolocation is about the user’s location (for compliance/content rules); latency is about speed, wherever that lands.

Cross-account subdomain delegation: create a hosted zone for the subdomain in the child account, copy its NS records, then add those NS records to the parent account’s hosted zone. Straightforward process – worth having walked through once before exam day.

CloudFront

A CDN distributing static and dynamic content via edge caching (dynamic content handled by forwarding cookies/headers to the origin). Supports up to 4K live and on-demand video. Origins can be S3, EC2, ELBs, or any web server - multiple origins with behaviours routing by URL path.

  • Invalidation: either delete at origin and let TTL expire, or submit an invalidation request via console, API, or third-party tools.
  • Geo-restrictictions by country at the distribution level.
  • Zone apex records supported (Route 53 alias to CloudFront).
  • SNI lets many certificates share one IP - older clients that lack SNI force the pricier dedicated-IP option.
  • The RTMP distribution type is deprecated – modern streaming runs over HTTP(S) web distributions, sometimes with ACM-issued certificates.

Elastic Load Balancers

Application Load BalancerNetwork Load BalancerClassic Load Balancer
OSI LayerLayer 7Layer 4Layer 4 or 7
Zonal FailoverYesYesYes
PlatformVPC OnlyVPC OnlyEC2-Classic or VPC
Health ChecksYesYesYes
Cross-Zone Load BalancingYesYesYes
CloudWatch MetricsYesYesYes
SSL OffloadingYesYesYes
Resource-based IAM PermissionsYesYesYes
ProtocolsHTTPS. HTTPTCP, UDP, TLSTCP, SSL, HTTP, HTTPS
Path or Host-based RoutingYesNoNo
WebSocketsYesYesNo
Server Name Indication (SNI)YesYes, as of 9/2019No
Sticky SessionsYesYes, as of 3/2020Yes
Static IP, Elastic IPOnly through Global AcceleratorYesNo
User AuthenticationYesNoNo
  • ALB (L7) - path/host/header/method/query-string/source-CIDR routing, user authentication, WebSockets. The default for HTTP workloads.
  • NLB (L4) - static/Elastic IPs, extreme performance, TCP/UDP/TLS passthrough (including non-HTTP protocols), TLS target offload, long-lived connections preserved.
  • CLB - legacy; mainly appears in exam questions as “migrate this away.”

Notable facts: sticky sessions are supported by all three (NLB gained them in 3/2020, ALB via duration or application cookie); SNI is supported on ALB and NLB but not CLB; and an ALB can only get a static IP through Global Accelerator.

3. Security

Security questions on the Pro exam follow a predictable arc: who is asking, how do they prove it, what are they allowed to do, and how do I detect when something went wrong.

The four questions

Identity – who are you? Authentication – prove it. Authorisation – are you allowed? Trust – do entities I already trust vouch for you? Nearly every security scenario on the exam reduces to deciding which of these is missing or misdesigned.

Multi-account organisations

The “why multiple accounts?” list is worth internalising because scenario questions presuppose it: group workloads by business purpose and ownership, apply the least privilege per environment, restrict access to sensitive data, limit the blast radius of a security event, and consolidate billing. Management accounts provision the org; member accounts hold the workloads; organisational units group accounts by application or service.

AWS Control Tower automates the best-practice landing zone on top of Organisations: guardrails enforced via SCPs (preventive) and AWS Config rules (detective), plus a dashboard for compliance visibility. Vocabulary to know cold:

  • Landing zone – the customisable multi-account starting point
  • Guardrail – a governance rule backed by an SCP or Config rule
  • Baseline – a bundle of blueprints (CloudFormation) + guardrails
  • Special accounts it provisions: log archive (aggregates logs) and audit (cross-account audit roles)

Service Control Policies (SCPs)

SCPs use IAM policy syntax but never grant permissions – they only filter what’s possible. Applied at org, OU, or account level, they inherit downward. The exam tests the interaction:

Effective permissions = allowed by IAM ∧ not denied by any SCP.

Deny-list SCPs explicitly forbid specific actions; allow-list SCPs implicitly deny everything unlisted (with FullAWSAccess attached by default so nothing breaks until you replace it). Pair SCPs with AWS Config for the detective side - compliant/noncompliant resource tracking with the history of what changed.

IAM Identity Center (the successor to AWS SSO) maps users and groups from your identity provider, speaks SAML 2.0 for federation with providers like Azure AD, or acts as a standalone directory. It supports MFA organisation-wide.

IAM Users vs. Identity Center – the exam-contrasting table:

IAM UsersIAM Identity Center
CredentialsStatic access keysRotating keys via roles
ScopeOne user, one accountOne user, many accounts
PermissionsOne permission set per userMany assumable roles
FederationNot supportedThe design goal

Rule of thumb: any scenario mentioning federation, a corporate IdP, or “hundreds of users across dozens of accounts” is Identity Center. Any mention of AD Connector feeding Identity Center means SSO for on-prem employees with their existing credentials (including RADIUS-based MFA).

Directory services decision table

OptionWhat it isBest for
AWS Managed Microsoft ADFull managed AD on Windows ServerEnterprises wanting hosted AD or LDAP for Linux apps
AD ConnectorProxy to your existing on-prem ADSSO for on-prem staff; joining EC2 to your domain
Simple AD (retired)Samba-based standalone directoryHistorical; no MFA, no trust relationships
Cloud Directory (retired)Hierarchical cloud-native directoryHistorical
Amazon CognitoConsumer sign-up/sign-in, social federationCustomer-facing apps and SaaS

Roles, trust, and cross-account access

Anything can assume a role: IAM users/groups, AWS services, non-AWS workloads (via IAM Roles Anywhere), and federated identities. Cross-account sharing is a three-step dance the exam tests mechanically:

  1. Define permissions in the target account’s role policy
  2. Create the trust policy naming the trusted principals ("Action": "sts:AssumeRole", optionally constrained by conditions like a source IP)
  3. Grant AssumeRole permission to the trusted account’s users

Canonical use cases: external auditors getting scoped access, Lambda in one account assuming a role in another, and workload permissions inside or outside AWS.

Token Vending Machine – the classic pattern for issuing temporary credentials to mobile clients: anonymous TVM grants service access only; identity TVM adds registration/login. Older pattern, but it still surfaces in exam lore.

AWS Resource Access Manager (RAM) shares resources (not roles) across accounts or OUs – the sharing account retains ownership. Flow: choose the resource, attach a managed policy, define principals, send the invitation. Prime use cases: shared foundational infrastructure (subnets, Transit Gateways), centralised certificate authorities, App Mesh networking, and cross-account Aurora/RDS cluster cloning.

Secrets

Secrets Manager stores passwords, API keys, SSH/PGP keys, and the like, retrievable via API with fine-grained IAM control, and can automatically rotate RDS credentials for MySQL, PostgreSQL, and Aurora using Lambda + KMS. Contrast with Parameter Store (Systems Manager) – simpler, cheaper, and fine for non-secret config; Secrets Manager is for things that rotate.

Encryption

At rest vs. in transit – the exam loves hiding which one a scenario actually needs.

KMS provides key storage, management, and auditing with either AWS-generated or imported keys; access controlled via IAM, usage audited via CloudTrail. It sits behind most AWS-service-native encryption (remember SSE-KMS from the S3 section). Validated against FIPS 140-2 Level 3 and PCI DSS Level 1.

CloudHSM is the single-tenant, hardware answer: dedicated FIPS 140-2 Level 3 devices living inside your VPC, accessed via VPC peering, clustered for HA. The trade-off: no native AWS service integration – you script applications against it. Classic uses: acting as your own issuing CA, offloading SSL from web servers, TDE for Oracle databases.

KMSCloudHSM
TenancyMulti-tenant AWS serviceSingle-tenant hardware
Root of trustAWS-managedCustomer-managed
AvailabilityAWS-managedYou cluster it
IntegrationNative across AWS servicesCustom application work

The discriminator in exam questions is almost always the root of trust: any hint of “we cannot let AWS hold the keys” or strict regulatory key custody → CloudHSM.

ACM provisions and deploys SSL/TLS certificates, natively integrated with CloudFront, ELB, and API Gateway. Public certificates are free; renewal is managed; wildcards are supported; third-party certs can be imported (but then you renew them – a classic trap). ACM Private CA issues private certificates for internal apps and devices.

Network-level defenses

Security Groups are stateful, instance-level firewalls (inbound source / outbound destination rules referencing IPs, subnets, or other SGs). NACLs are stateless, subnet-level packet filters – the default NACL allows everything, so NACLs are where you write explicit denies. Together they form layered least privilege; NACLs are also the backup line of defense when someone fat-fingers an SG rule open.

WAF vs. Network Firewall – remember by what it attaches to:

  • WAF → ALB, CloudFront, API Gateway, AppSync – Layer 7 rules (SQLi, XSS, rate limiting)
  • Network Firewall → VPC-level via Transit Gateway, internet gateways, DX/VPN gateways – Layer 3–4 inspection

Shield. Standard is free-automatic protection against common L3/L4 DDoS. Shield Advanced adds application-layer coverage, protects named resources (ELBs, CloudFront, Route 53, etc.), provides detailed attack forensics, and includes 24×7 access to the SRT (DDoS response team). Firewall Manager then scales WAF/Shield/Network Firewall/SG policies across accounts and the org via Organisations.

GuardDuty is intelligent threat detection across workloads, S3, accounts, and users – findings flow to EventBridge and into Security Hub, which correlates GuardDuty, Inspector, Macie, Firewall Manager, and third-party findings into prioritised recommendations.

IDS/IPS and SIEM. IDS watches for suspicious activity; IPS sits inline to prevent it. Neither is an AWS service – you compose them (agents + collection system), typically shipping logs to a SIEM via CloudWatch, S3, or third parties like Splunk/Sumo Logic. Expect “how would you architect detection for this hybrid environment” answers to include that composition.

CloudWatch vs. CloudTrail

The perennial confusion pair, settled once and for all:

CloudWatchCloudTrail
RecordsMetrics, logs, events across servicesAPI activity across services
RoleHigher-level monitoring and alertingGranular audit trail of who did what
AlarmingNativeVia CloudWatch alarms on trail logs
RetentionLogs indefinitely (configurable)Logs to S3/CloudWatch indefinitely

And Security Hub + Firewall Manager + Config + GuardDuty tie it together into the “managed security posture” answer family – when a scenario asks for centralised, multi-account security visibility, that’s the stack.

Service Catalog

Last piece of identity/governance: admins publish curated products (built on CloudFormation) with launch constraints, so end users self-provision without broad IAM permissions.

ConstraintWhatWhy
LaunchIAM role the catalog assumes on launchUsers need no underlying service permissions
TemplateRules narrowing allowed parameter valuesE.g., only t3.micro in dev
NotificationSNS topic for stack eventsVisibility on launches and failures

Products can be versioned and removed without shutting down already-running deployments – worth remembering as the “governed self-service” exam answer.

4. Migration Strategies

Migration questions rarely ask “what is DMS?” – they hand you a portfolio of aging applications and ask which strategy applies to each, then which tool executes it. So this section runs in that order: strategy vocabulary first, then the toolchain.

The 6 Rs

Memorise this table until it’s reflexive – the exam will describe an application and expect you to name its R in about twenty seconds:

StrategyDescriptionExampleEffortOptimisation upside
Re-host“Lift and shift” - move as-isOn-prem MySQL → EC221
Re-platform“Lift and reshape” - move and swap the platform underneathOn-prem MySQL → RDS MySQL43
Re-purchase“Drop and shop” - abandon and buy SaaSLegacy CRM → Salesforce31
RearchitectCloud-native redesignMonolith → serverless55
RetireDecommission what nobody usesKill the label-printing app00
Retain“Do nothing” - revisit laterServers live to see another day10

Tip: when a scenario says “minimal change, deadline-driven,” that’s re-host. When it says “take advantage of cloud elasticity,” that’s rearchitect. When finance is asking about licence costs on a dying app, check whether the answer is actually retire – it’s the cheapest R and easy to overlook under exam pressure.

The Cloud Adoption Framework (CAF)

TOGAF (The Open Group Architecture Framework, developed since 1995) is the de facto enterprise-architecture standard among Fortune 500s, but it’s a framework, not a cookbook – open to local adaptation and frequently blamed for unreasonable expectations. The CAF, by contrast, is AWS’s prescriptive lens for the adoption journey. Classic phases: Project → Foundation → Migration → Reinvention.

Then the six CAF perspectives, which map neatly onto “which perspective is this question about?”:

  • Business – the case for adoption; measurable benefits (TCO, ROI)
  • People - roles, skills, training, incentives
  • Governance – portfolio management, agile program management, cloud-aligned KPIs
  • Platform - standardised provisioning, cloud-native architecture patterns
  • Security – IAM model changes, logging/audit evolution, the shared responsibility shift
  • Operations - automated monitoring, scalable performance management, cloud-era BC/DR

Exam hook: questions quote an organisational symptom (e.g., “no skills plan exists”) and expect you to name the perspective – People, in that case.

Hybrid architectures

Common first step before full migration, and the exam treats it accordingly:

  • Storage Gateway bridging on-prem and S3 – seamless to users, low-risk, appealing economics; the textbook “first step into cloud”
  • Middleware integration – loosely coupled, canonical-message-based connections
  • VMware integration – vCenter plug-in for transparent VM migration; VMware Cloud on AWS deepens it

The unifying principle: loosely coupled integrations where either end can exist without knowledge of the other.

Discovering and planning

AWS Migration Hub is the coordination dashboard: it leverages the Application Discovery Service to inventory and group servers, recommends right-sizing, suggests per-application migration strategies, and tracks lift-and-shift versus refactor-first progress. Know that discovery feeds recommendations – that flow is the scenario.

Migrating applications

  • Containerised workloads → AWS Batch or Fargate
  • Web applications → Elastic Beanstalk or Lightsail
  • Message-handling microservices → event-driven/serverless architectures (decoupled queues)
  • Windows/Linux servers wholesale → Application Migration Service (MGN) - agent-based rehosting from on-prem, other clouds, or within AWS, with test instances and minimal-interruption cutover. Free to use; the agent right-sizes EC2 recommendations.
  • Bringing AWS to the data → Outposts (preconfigured hardware running AWS services on-prem) or Snowball Edge (compute-capable offline transfer, e.g., EC2/EKS Anywhere)

Migrating data

Three tools, chosen by volume and connectivity – the exam always gives you those two facts:

ToolWhenDetails
DMSDatabase migrationsSource relational DB → RDS/Aurora/EC2; Schema Conversion Tool (SCT) for heterogeneous engines; can also target S3, Snowball, and even MongoDB/DynamoDB for smaller moves
DataSyncOnline, ongoing file/object syncSMB, HDFS, NFS → S3/EFS/FSx; automates the transfer securely
Transfer FamilyThird-party file exchangeManaged SFTP/FTPS/FTP/AS2 endpoints, multi-AZ; File Transfer Workflows automate encryption, filtering, tagging, compression

And when the data is too big or the connection too slow – the Snow Family:

GenerationWhat it is
Import/ExportShip your own hard drive to AWS (historical)
SnowballRuggedized appliance, up to 80 TB, copied to S3
Snowball EdgeSnowball + onboard compute (Lambda) and clustering
SnowmobileA literal shipping container - up to 100 PB - towed by truck

All encrypted at rest and in transit; transfer speed bounded by your courier, not your WAN. Rule of thumb: bandwidth math beats postal math only when data fits through the pipe – the exam’s “10 Gbps link vs. 60 PB of video” questions resolve themselves once you compute weeks-versus-days.

Network migration and cutover

The checklist that keeps a migration from becoming an outage:

  1. No IP overlap between VPC and on-prem - VPC IPv4 netmasks run /16 to /28 (65,536 → 16 addresses), with 5 reserved IPs per subnet (a /28 nets you 11 usable)
  2. Start with VPN – most organisations begin here
  3. Layer in DX with the VPN as backup
  4. Cut over with BGP - run VPN and DX advertising the same prefixes; AWS prefers the DX path, making the transition near-seamless

5. Architecting to Scale

Scaling questions are really cost-plus-resilience questions in disguise: the exam describes a spiky workload, and you pick the mechanism that absorbs it without waste. Everything in this section hangs off one decision first – up or out.

Scaling up vs. scaling out

Scaling upScaling out
WhatMore CPU/RAM on existing instancesMore instances
DowntimeRestart requiredNone
AutomationManual scriptingNative for most AWS compute
CeilingInstance size limitsEffectively unlimited

Horizontal wins on exam day unless the scenario explicitly demands a monolithic resource (a legacy vertical app, a licensed single-node database). From here on, we’re talking scale-out.

Auto Scaling

Three layers – don’t conflate them: EC2 Auto Scaling (ASGs), Application Auto Scaling (the API behind DynamoDB/ECS/EMR/etc. table capacity), and AWS Auto Scaling (the cross-service planning console).

ASG scaling options:

OptionWhatWhen
MaintainKeep N instances alive“I always need 3”
ManualAdjust desired capacity yourselfRarely-changing needs
ScheduledScale by clock“Monday-morning rush”
DynamicReact to live metricsCPU crossing 70%

Dynamic scaling policies:

PolicyBehaviorVibe
Target trackingHold a metric near a target value“Keep CPU at 70%” - the default modern answer
SimpleWait for health check + cooldown before evaluating againSlow and steady
Step scalingEscalating adjustments as the alarm breaches further“AGG! Add ALL the instances!”

Cooldowns: default 300 seconds, the pause that lets new instances come up and absorb load before further scaling decisions. Applies to dynamic and (optionally) manual scaling – not scheduled scaling.

Predictive scaling uses ML on your traffic history to scale ahead of demand. You can use its forecast to tune your own policies, or opt out of data collection entirely.

Compute Optimiser

ML-powered rightsizing fed by CloudWatch metrics. Covers EC2 instance types, Auto Scaling groups, EBS volume types/sizes, ECS-on-Fargate task sizing, and Lambda memory/CPU allocation. The “reduce cost without redesign” answer family.

Event-driven architecture

The serverless backbone – know each player’s role on sight:

Lambda (serverless compute) · EventBridge (event bus and rules) · Step Functions (orchestration) · SQS (buffer/decouple) · SNS (fan-out to subscribers) · API Gateway (entry point for external producers) · DynamoDB and S3 (stateful endpoints that also emit events).

  • SNS – topics (publish channel), subscriptions (endpoint per topic), protocols including HTTPS, email, SMS, SQS, Lambda, mobile push
  • SQS - KMS-encrypted, transient, 4-day default / 14-day max retention, optional FIFO, 256 KB message ceiling (extended-client libraries can push larger payloads via S3)
  • Amazon MQ - managed Apache ActiveMQ (JMS, NMS, MQTT, WebSocket); the drop-in answer when migrating an existing broker; use SQS/SNS for greenfield

Lambda + SAM + EventBridge. Lambda supports Node.js, Python, Java, Go, C#; stateless; scales without fundamental limits. SAM is CloudFormation’s serverless extension - YAML templates, sam deploy, local testing via a Docker-based emulator. EventBridge connects AWS and third-party SaaS events to rules.

Step Functions turns your app into a state machine (Amazon States Language, JSON): tasks, sequences, parallels, branches, timers, with a visual console showing live execution. The workflow-engineer’s tool of choice – see the comparison table below.

WhenExample
Step FunctionsOut-of-the-box AWS service orchestrationOrder-processing flow
SWF (legacy)Historically: external/manual steps in workflowsLoan application with human review
SQSStore-and-forward queueingImage resize pipeline
AWS BatchScheduled or recurring compute jobs, light logicRotate firewall logs nightly

(SWF is legacy/closed to new customers – I’d drop it or footnote it; it survives in older question banks.)

Streaming with Kinesis

Sharded ingestion: 1,000 records/sec and 1 MB per record per shard, default 500-shard limit. Records carry a partition key, sequence number, and data blob. Retention: 24 hours by default, extendable to 365 days

DynamoDB scaling (continued from Section 1)

Section 1 covered partition math; here’s the elasticity layer. Auto Scaling uses target tracking against a utilisation target, driven through Application Auto Scaling, and also applies to GSIs. On-demand capacity trades a price premium for instant, provision-nothing capacity – the answer when workloads are unpredictable or you can’t tolerate scaling lag. DAX is the in-memory cache: microseconds-on reads for read-intensive, repeatedly-hit datasets. It’s the wrong answer for write-heavy apps or apps that already cache client-side.

Scaling containers

ECSEKS
ModelAWS-opinionatedKubernetes-compatible
IntegrationNative with Route 53, ALB, CloudWatchNeeds generalised LB abstractions
Learning curveSimplerRicher, steeper
ExtensibilityLimitedBroad third-party ecosystem
UnitTasksPods

Terminology transfer: ECS tasks ≈ K8s pods; lift-and-shift from an existing K8s estate points at EKS.

Launch types. EC2 (you provision, patch, optimise, retain control) vs. Fargate (serverless containers – AWS provisions and optimises; you trade granularity for zero-touch). Both are available on ECS and EKS, plus Outposts, Local Zones, and Wavelength.

Around them: App Runner (synchronous HTTP apps from source or images, scales to zero - proof-of-concept darling) and AWS Batch (plans, schedules, and runs jobs on ECS/EKS/Fargate, with spot-instance cost reduction – create a compute environment, a job queue with priorities, a job definition, then schedule).

Big data at scale: EMR

Managed Hadoop – also Spark, HBase, Presto, Flink – for log analysis, financial analysis, ETL. An EMR step is a programmatic unit of work; a cluster is the EC2 fleet assembled to run steps. The Zoo bestiary (Ambari, Sqoop, Flume, ZooKeeper, Oozie, Pig, Hive, Mahout, HBase, MapReduce, HDFS) still appears in distractor answers – recognise the names even if the ecosystem has largely moved to Spark-on-EMR.

Visualising at scale: QuickSight and OpenSearch

QuickSight - serverless, pay-per-user BI, powered by the SPICE in-memory engine; builds hybrid datasets across Athena, Redshift, S3, RDS, Timestream, Snowflake, Salesforce, and dozens more; ML-driven natural-language querying and anomaly detection. The “dashboard for business users, cheaply” answer.

OpenSearch - search + real-time analytics, OpenSearch Dashboards (née Kibana); richer querying and lower cost than CloudWatch for log-heavy workloads – the trade being more operational ownership.

6. Business Continuity & Disaster Recovery

This is the domain where the exam stops asking “what does this service do” and starts asking “how much does it hurt when things break.” Two numbers rule everything here, so we start with them.

Vocabulary – and why RTO/RPO are the whole game

  • Business continuity (BC) – minimising disruption to business activity when something unexpected happens
  • Disaster recovery (DR) – the act of responding to an event that threatens BC
  • High availability (HA) – designing in redundancy to reduce the chance of impacting service levels
  • Fault tolerance (FT) – designing in the ability to absorb failures without impacting service levels
  • SLA – an agreed target for a service’s performance or availability
  • RTO – time to restore business processes after a disruption. T is for Time.
  • RPO – acceptable amount of data loss, measured in time. P is for the data that goes “poof.”

Nearly every DR scenario question resolves to one of these two numbers: “we can lose up to 15 minutes of data” → RPO; “we must be serving customers within an hour” → RTO. The architecture you pick is downstream of those two figures – and, always, of the budget.

What counts as a disaster

CategoryExample
Hardware failureSwitch power supply dies, LAN goes down
Deployment failureA patch breaks a key ERP process
Load-inducedDDoS on the website
Data-inducedAriane 5 explosion, June 4 1996
Credential expirationAn SSL/TLS certificate quietly lapses
DependencyAn S3 subsystem failure cascades into other services
InfrastructureBackhoe through the fiber line
Identifier exhaustion“Insufficient capacity in requested AZ”

Note how many of these are self-inflicted (deployment, credential, identifier) – a hint that DR planning isn’t only about acts of God.

The HA continuum – the four DR patterns

The core of the section. Same workload, four escalating price/risk trade-offs - the exam gives you an RTO/RPO and a budget, and you place the architecture:

PatternWhat’s runningFailoverRTO ballparkCost
Backup & RestoreNothing - just data in S3/GlacierRebuild everything from scratchHoursLowest
Pilot LightCore data replicated; everything else offSpin up the rest on demandMinutes–hoursLow
Warm StandbyA scaled-down full environment, always onScale up, cut overMinutesMedium
Multi-Site (active/active)Full production capacity everywhereNear-instant, little/no interventionSecondsHighest

Backup-and-restore is the most common entry point into AWS and needs minimal configuration; pilot-light keeps AMIs synchronised with on-prem counterparts and usually needs manual failover; warm-standby doubles nicely as a testing/staging shadow environment but needs scaling to take production load; multi-site is “effectively a mirrored data center” and can feel wasteful precisely because it’s paying for insurance.

Storage HA, per service

EBS. Annualised failure rate under 0.2% (vs. ~4% for commodity HDDs), availability target of 99.999%, replicated automatically within a single AZ – which means the AZ itself is the failure domain. Snapshots go to S3 and can be copied across regions; volumes support RAID.

ConfigVolume size (×each)IOPS/volumeTotal IOPSUsable spaceThroughput
None1,000 GB4,0004,0001,000 GB500 MB/s
RAID 0500 GB × 24,0008,0001,000 GB1,000 MB/s
RAID 1500 GB × 24,0004,000500 GB500 MB/s

The one-line takeaway: RAID 0 buys performance (at zero redundancy); RAID 1 buys redundancy (halving usable space). Both occasionally appear as the answer to “I need more IOPS than one volume can give me.”

S3. Durability is eleven 9s (99.999999999%) across storage classes; availability differs: Standard at 99.99% (~52.6 min/year), Standard-IA at 99.9% (~8.8 hrs/year), One Zone-IA at 99.5% (~1.8 days/year) - and One Zone-IA drops the multi-AZ replication, which is exactly why its availability figure slides. Standard and Standard-IA replicate across AZs. Also the backing store for EBS snapshots and half of AWS’s internals – remember dependency-failure risk cuts both ways.

EFS. True file system semantics (locking, strong consistency), files and metadata spread across multiple AZs, concurrently mountable everywhere – the anti-EBS.

Storage Gateway / Snowball / Glacier. Respectively: continuous sync for offsite backup, batch transfers only, and safe long-term archival with rare retrieval.

Compute HA

Keep AMIs current – a stale AMI is a slow RTO. AMIs copy cross-region for DR staging. Horizontal architectures spread risk across smaller machines; Reserved Instances get launch priority in an AZ, and On-Demand Capacity Reservations guarantee an instance type in a specific AZ – the “my failover region must actually have capacity” answer. Route 53 health checks provide self-healing DNS redirection.

Database HA

  • DynamoDB - data and traffic spread across partitions, synchronously replicated across three AZs per region; Global Tables add multi-region active-active
  • RDS - Multi-AZ standby, read replicas (region and cross-region), snapshot recovery
  • Aurora - multi-AZ copies by design; Global Databases: one primary region, up to 5 secondary regions, storage-level replication with sub-second typical lag, promotable during an outage
  • Redshift - historically single-node meant restore-from-snapshot on failure and the multi-node cluster was the HA answer; multi-AZ support arrived on RA3 nodes

Networking HA

Subnets across AZs create multi-AZ VPC presence by construction. Two VPN tunnels to a Virtual Private Gateway is best practice – AWS gives you both for a reason. Direct Connect is not HA by itself: redundancy needs a second DX connection or a VPN fallback (tying back to Section 4’s BGP cutover). Route 53 health checks redirect DNS at the layer below your load balancers; Elastic IPs let you swap backing assets without touching DNS. And per Section 2: one NAT Gateway per AZ, with private-subnet routes pinned to their local gateway.

FMEA – scoring your risks

Failure Mode and Effects Analysis is the formal process the exam occasionally asks about: for each failure mode, assess what could go wrong, its impact, its likelihood, and your ability to detect it. The score:

Risk Priority Number = Severity × Probability × Detection

Interesting subtlety worth knowing: a higher detection score means harder to detect – so a severe, probable failure that hides quietly scores worst. That inversion trips people up.

7. Deployment & Operations

Two halves: how you ship changes (deployment patterns, CI/CD, Elastic Beanstalk, CloudFormation) and how you run what you shipped (Systems Manager, Config, enterprise apps). The exam weights both, with a particular fondness for choosing the right deployment strategy for a given downtime tolerance.

Deployment patterns

Four ways to roll out v2, in ascending order of safety:

  • Rolling - swap instances batch-by-batch behind your ELB; no downtime, but you’re briefly running mixed versions
  • A/B testing – split traffic by percentage between v1 and v2 to compare behavior
  • Canary – a single v2 instance alongside the fleet; smallest possible blast radius
  • Blue-Green - stand up the entire v2 environment, verify, then flip traffic wholesale

Blue-green implementations the exam expects you to recognise: DNS cutover to a new ELB, swapping the ASG behind the ELB, updating the launch configuration and rolling the ASG, swapping Elastic Beanstalk environment URLs, or cloning an OpsWorks stack (historical flavor).

Blue-green’s contraindication – the detail that separates the pros: if your data-store schema changes are tightly coupled to code changes, or the upgrade needs special routines run mid-deployment, blue-green breaks down, because both versions must be simultaneously compatible with one database.

CI/CD on AWS

  • CI - merge to main frequently
  • CD (delivery) – automated release process, deploy on demand
  • CD (deployment) – every change ships to production with no human intervention

The toolchain: CodeCommit (managed Git) · CodePipeline (orchestration) · CodeBuild (compile/test/package) · CodeDeploy (deploy to EC2, Beanstalk, ECS, Lambda) · Cloud9 (cloud IDE - now closed to new customers, but still fair game as a legacy answer) · CodeGuru (automated code review) · CodeStar (prebuilt CI/CD ecosystems) · X-Ray (distributed tracing) · CodeArtifact (package management).

Elastic Beanstalk

Orchestration for scalable web apps – Docker, PHP, Java, Node, and more – with multiple environments (dev/QA/prod) per application. Ease of deployment at the cost of control. Its deployment options table is prime exam material:

OptionWhatDowntimeRollback
All-at-onceUpdate every existing instance simultaneouslyYesManual
RollingBatches through existing instancesNoManual
Rolling + additional batchAdds new-version instances before retiring old onesNoManual
ImmutableFresh ASG of new instances; cutover only after health checks passNoTerminate new instances
Traffic splittingRoute a percentage to new instances for canary testingNoReroute DNS, terminate new
Blue/greenFull new environment, swap CNAMENoSwap URL back

Pattern to notice: immutable and traffic-splitting are the safest answers; all-at-once is the trap answer for any scenario that mentions uptime.

CloudFormation

Infrastructure as code: templates (JSON/YAML) describe environments; stacks create/update/delete them atomically; change sets preview proposed changes before you commit; StackSets deploy across multiple accounts and regions. Over 300 resource types, custom resources via SNS or Lambda.

Stack policies are the exam favorite: protect specific resources from accidental update/deletion. Applied at creation via console or CLI; adding one to an existing stack is CLI-only; once applied it can’t be removed – only modified via CLI. Best practices: use Change Sets to spot trouble, make changes through CloudFormation rather than clicking the console, keep templates in version control.

(Don’t forget the abstractions layered on top: CDK, SAM from Section 5, and third-party frameworks like Terraform.)

Running the fleet: Systems Manager

This is the heart of modern ops on AWS:

CapabilityWhatExample
InventoryCollect OS/app/instance metadata“Which instances run Apache 2.2.x or earlier?”
State ManagerDesired-state associationsTrack which instances got patched to stable
Parameter StoreShared secure config storagePull RDS credentials at boot
Maintenance WindowsScheduled patch/script windows00:00–02:00 for Patch Manager
AutomationRoutine maintenance via automation documentsStop dev/QA instances Friday night
Run CommandExecute commands without SSH/RDPShell script across 53 instances at once
Patch ManagerFleet-wide patchingKeep everyone at the same patch level
Insight DashboardsAccount-level Config/CloudTrail/Trusted Advisor viewSingle compliance viewport
Resource GroupsTag-driven groupingsDashboard for all Production ERP assets

Plus SSM documents: command (Run Command/State Manager), policy (enforced state), and automation documents. The SSM agent ships on AWS base AMIs.

AWS Config (covered fully in Section 3’s security posture) – here it plays configuration management: baselines, drift tracking, compliance rules. One mental slot: Config evaluates, SSM remediates.

API Gateway

Backends via Lambda, AWS-service proxies, or any HTTP endpoint; regional, private, or edge-optimised (CloudFront-backed) deployments. API keys and usage plans give you identification, throttling, and quotas; custom domains with SNI work; and APIs can even be monetised on the AWS Marketplace.

Enterprise end-user apps

  • WorkSpaces (managed DaaS) and AppStream 2.0 (application streaming) – regulated industries, remote/seasonal workers, product demos without installs
  • Connect - managed cloud contact center: call handling, IVR, chatbots, analytics, CRM integration
  • Chime - meetings and video conferencing
  • WorkLink (retired) and Alexa for Business (deprecated) – dropped; they occasionally haunt older practice questions only

Machine learning service tiers

The three-layer taxonomy is the exam-relevant frame – know which tier a scenario implies:

  1. AI services (app developers, no ML background): Comprehend (NLP/sentiment), Lex (chatbots), Polly (text-to-speech), Rekognition (image/video analysis), Translate, Transcribe (speech-to-text), Textract (document text extraction), Personalise (recommendations), Forecast (time-series prediction)
  2. ML services (data scientists): SageMaker end-to-end, Ground Truth, training/hosting, Marketplace
  3. Frameworks & infrastructure (researchers): Deep Learning AMIs, Greengrass, Gluon/Keras, MXNet/TensorFlow

Scenario-matching examples worth keeping: sentiment analysis of social posts → Comprehend; upsell recommendations at checkout → Personalise; digitising paper forms → Textract; seasonal demand forecasting → Forecast.

IoT

The remaining, current core: IoT Core (MQTT message broker to/from devices), IoT Events (conditional logic on sensor data), Greengrass (deploy Lambdas/Docker/ML to edge devices), Device Management (registry, groups, OTA firmware), Device Defender (ML-based configuration auditing), IoT Analytics (aggregation, time-series SQL), SiteWise (edge data collection and modeling). Retired: Things Graph and 1-Click - footnote only.


8. Cost Management

The shortest domain, but never skipped - FinOps-flavored questions appear reliably, and they’re fast points if the vocabulary is cold.

CapEx vs. OpEx. Capital expenditure buys long-term assets (buildings, hardware); operational expenditure is ongoing, usually variable spend. Cloud shifts the needle from CapEx to OpEx – often the business motivation behind the technical migrations in Sections 4–6.

TCO and ROI. Total cost of ownership captures the full cost model of a decision; return on investment is what you get back within a timeframe. If a scenario asks “how do we justify this migration to finance,” the answer starts here.

Cost optimisation levers, roughly in order of exam frequency:

  1. Right-sizing - lowest-cost resource meeting specs; CloudWatch utilisation data drives it; loosely coupled architectures smooth demand enough to size against
  2. Purchase options – Reserved Instances for steady workloads (Standard, Convertible, Scheduled), Spot for interruptible horizontal scale, EC2 Fleet for blended On-Demand/RI/Spot
  3. Appropriate provisioning - don’t over-provision; consolidate; watch utilisation
  4. Managed services – RDS over self-run MySQL; Fargate; the operational savings are the saving
  5. Geographic selection – pricing varies by region; pair with Route 53/CloudFront to counteract latency
  6. Optimised data transfer - egress and inter-region traffic add up; Direct Connect can win at volume

RI mechanics – the details: attributes are instance type, platform, tenancy, and (optionally) AZ; zonal RIs reserve capacity in an AZ, regional RIs apply discounts anywhere in the region (and can be converted zonal→regional). Shared across consolidated billing, sellable on the RI Marketplace.

Dedicated Instances vs. Dedicated Hosts. Dedicated Instances run on hardware dedicated to you but may share hardware with your other non-dedicated instances – a ~$2/hr/region premium, purchasable as On-Demand, RI, or Spot. Dedicated Hosts are entire physical servers you control placement on - the answer for per-core/per-socket licensing – and each host runs only one instance family/type.

Spot mechanics: specify a max price; instances stop/terminate/hibernate when the market exceeds it; requests can be one-time, maintained, or duration-based.

Tagging is the cheapest cost-management win and the bridge into governance: name/value pairs on every resource, powering cost allocation, IAM conditions, and automation. Enforce via Config rules or scripts (e.g., untagged EC2 instances stop nightly). Resource Groups then turn tags into custom consoles – consolidated views by environment, project, or cost center.

The reporting toolkit:

  • Cost and Usage Reports (CUR) – CSV granularity down to hourly, analysable via Athena/Redshift/QuickSight
  • Budgets – alerts as you approach limits, based on cost, usage, or RI utilisation/coverage
  • Consolidated Billing – one payer account, economies of scale across the org
  • Trusted Advisor – automated checks (core checks free; full set behind Business/Enterprise support) recommending optimisations like RIs and scaling adjustments

Closing: how I studied, and what I’d tell you

Now the part the intro promised. Coming from the Advanced Networking Specialty, the biggest adjustment wasn’t depth – it was breadth. The Specialty asks you to know a few things deeply; the Pro exam asks you to know everything broadly and choose correctly under constraints. Three study habits made the difference:

Decision tables over definitions. I ended up reconstructing my notes as comparison tables – you’ve just read the result. When a practice question described a scenario, I wanted the answer to be a row lookup, not a recall struggle. If you take one thing from this post, rebuild your own tables by hand; the rebuilding is the studying.

Read the constraint sentence twice. Pro questions are long, and the discriminating fact - “no downtime acceptable,” “must encrypt at rest with customer-held keys,” “budget is fixed” - is almost always a single clause. I lost more practice questions to skimming than to ignorance.

Check deprecations before the exam, not after. I tripped over more retired services in practice exams than in any other category of error. Anything here that’s still marked legacy - worth thirty minutes with the AWS “service history” pages the week before your exam.

I used the Pluralsight Solutions Architect – Professional path as the spine, these notes as the skeleton, and practice exams for calibration. If you’re coming from an Associate cert rather than another Specialty, budget extra time for the multi-account organisation and DR material – those are the topics with no Associate-level counterpart.

Good luck – and if these notes helped, pay it forward like every blog post that helped me did.