# AWS Advanced Networking Specialty

*2024-06-25*

> Notes from studying for the AWS Advanced Networking Specialty certification


After passing the [AWS Certified DevOps Engineer – Professional]({{< relref "/posts/2024/aws-devops-engineer-professional/" >}}), I decided to take [the AWS Certified Solutions Architect – Professional](https://aws.amazon.com/certification/certified-solutions-architect-professional/) route – but with a detour first. Networking underpins almost everything in the SA Pro exam, so I sat the [**AWS Certified Advanced Networking – Specialty**](https://aws.amazon.com/certification/certified-advanced-networking-specialty/) first to shore up the foundations.

These are my study notes from the [AWS Certified Advanced Networking Specialty path](https://www.pluralsight.com/paths/aws-certified-advanced-networking-specialty-ans-c01) on [Pluralsight](https://www.pluralsight.com).

Fair warning: they're the longest set of notes I've ever taken, because the exam domain is genuinely enormous – from BGP path selection to GENEVE encapsulation. I've organised them by the exam's four domains, and I've flagged the things that tripped me up along the way.

---

## Contents

1. [Domain 1: Design and implement AWS networks](#domain1)
2. [Domain 2: Configure network integration with application services](#domain2)
3. [Domain 3: Hybrid networking – VPN, Direct Connect, and transit architectures](#domain3)
4. [Domain 4: Design and implement for security and compliance](#domain4)

## Domain 1: Design and implement AWS networks{#domain1}

### Public vs. private services

- **Public services** have publicly reachable endpoints by default – think S3, IAM, and SQS.
- **Private services** aren't reachable by default: anything you deploy inside a custom VPC.

### Availability vocabulary (know this cold)

The exam loves the distinction between these two terms:

- **Highly available** – the system maintains uptime *almost* always, even if degraded. Designed to remove single points of failure.
- **Fault tolerant** - the system *always* maintains uptime, typically active/active across multiple AZs or regions.

### Subnetting math

IPv4 addresses are calculated from the left in bits. For a `/N` subnet, usable hosts = 2^(32−N) − 5 (the five reserved addresses):

| CIDR | Total | Usable |
|------|-------|--------|
| /24  | 256   | 251    |
| /27  | 32    | 27     |

That "minus 5" matters: network address, VPC router (.1), DNS (.2), future use (.3), and broadcast (.255 of the block).

### IPv4 vs. IPv6 in VPCs

- **IPv4**: required on every VPC resource; you choose the CIDR (between /16 and /28); addresses can be public or private.
- **IPv6**: optional; AWS assigns a fixed-size /56 block per VPC (or you can BYOIP from /48 down to /60); every address is publicly routable – security comes from routing and security groups, not address privacy. IPv6-only subnets are supported.
- Remember the pairing: **you choose IPv4 CIDRs; AWS assigns IPv6** (unless BYOIP).

> **Exam tip:** IPv6 operations in a VPC are bound to the IPv4 structure – the subnet's /64 IPv6 block hangs off the VPC's /56, the same way IPv4 subnets subdivide the VPC's primary CIDR.

### Choosing your CIDR blocks

Subnets must fall between /28 and /16 (65,536 addresses max, 16 minimum). Before picking, consider: overlap with other VPCs you might peer (you **cannot** resize a VPC CIDR later), on-premises networks for hybrid connectivity, and how many subnets you'll need.

For private space, 10.0.0.0/8 gives you 256 × /16s to play with – the roomiest canvas. And when splitting subnets across AZs, divide binary (2, 4, 8…) and size for *more* AZs than you currently use. Example: splitting `172.31.0.0/16` into four /18s covers four AZs with room to grow:

```
172.31.0.0/18
172.31.64.0/18
172.31.128.0/18
172.31.192.0/18
```

# Domain 2: Configure network integration with application services{#domain2}

## DNS fundamentals

DNS translates domain names to IP addresses. The exam expects you to understand hierarchy and delegation:

- **FQDN** (Fully Qualified Domain Name): `subdomain.second-level-domain.top-level-domain.` (note the trailing root dot)
- **Record types you'll see on the exam:**
    - `A` / `AAAA`: host to IPv4 / IPv6
    - `CNAME`: alias for a host (can't be used at zone apex)
    - `NS`: name server delegation
    - `MX`, `TXT`, `PTR`, `SOA`, `CAA`: mail, text, reverse lookup, authority, certificate authority authorisation
- **Name servers:** authoritative ones know the answers; non-authoritative ones point elsewhere or serve cached content.

> **Exam tip:** You cannot put a CNAME record at the zone apex (root domain). This is why Route 53 Alias records exist-they return an IP address to the client, not an FQDN, and can sit at the apex.

## Route 53 hosted zones and routing policies

### CNAME vs. Alias (know the difference cold)

| Aspect            | CNAME | Alias                 |
|-------------------|-------|-----------------------|
| Record type?      | Yes   | No-Route 53 extension |
| Returns to client | FQDN  | IP address            |
| Query charges     | Yes   | No                    |
| Zone apex allowed | No    | Yes                   |
| Name reuse        | No    | Yes                   |

Alias records are specifically for AWS service objects (ELBs, CloudFront distributions, S3 buckets). They're the right choice for production workloads.

### Routing policies (each solves a different problem)

| Policy      | Use case                                                       |
|-------------|----------------------------------------------------------------|
| Simple      | Plain round-robin, single destination                          |
| Weighted    | Percentage-based traffic splitting                             |
| Failover    | Primary/secondary (primary unhealthy → returns primary anyway) |
| Latency     | Fastest connection based on resolver location                  |
| Geolocation | Location-based routing (most specific match wins)              |
| Multivalue  | Up to 8 random healthy hosts (with health checks)              |

> **Exam tip:** With geolocation routing, some IP addresses won't have a location. If you need to lock traffic to a specific region, don't set a default location-the unrouted queries will fall through.

### Health checks

- **Endpoint monitoring:** public IP or domain
- **CloudWatch alarm:** useful for private hosted zones
- **Pricing:** 50 free AWS endpoints, then $0.50/check; non-AWS endpoints are $0.75/check

## Private hosted zones and split-view DNS

- **Split-view/split-horizon DNS**: separate internal and external zones with fully overlapping namespaces.
- Private hosted zones are associated with VPCs; the default VPC router performs DNS resolution.
- Public zones **cannot** delegate subdomains to private zones (and vice versa)-they live in separate DNS silos.

> **Exam tip:** For hybrid networks, Route 53 Resolver endpoints let on-premises systems reach private hosted zones (inbound endpoints) and let AWS resources reach on-prem DNS (outbound endpoints). Each ENI handles up to 10,000 queries/second-far better than a single EC2 DNS resolver's 1,024 qps limit.

## VPC DHCP and Route 53 Resolver

AWS reserves specific addresses within each VPC CIDR:

| Address     | Purpose                        |
|-------------|--------------------------------|
| `x.x.x.0`   | Network address                |
| `x.x.x.1`   | VPC router                     |
| `x.x.x.2`   | DNS server (Route 53 Resolver) |
| `x.x.x.3`   | Future use                     |
| `x.x.x.255` | Broadcast (for /24 blocks)     |

The Route 53 Resolver lives at **+2**, **not** the router (+1). Don't confuse them-they're separate components even though the IPs are adjacent.

DHCP option sets are immutable once created. If you need custom DNS servers, configure them explicitly rather than relying on defaults.

## Elastic Load Balancers (ELBs)

### Types at a glance

| Type | Layer | Protocols      | Key feature                         |
|------|-------|----------------|-------------------------------------|
| ALB  | 7     | HTTP/HTTPS     | Path/header routing, rules engine   |
| NLB  | 4     | TCP/UDP/TLS    | Static IPs, millions of RPS         |
| CLB  | 4 & 7 | HTTP/HTTPS/TCP | Legacy, EC2-Classic only            |
| GWLB | 3 & 4 | GENEVE         | Security appliances, inline traffic |

> **Exam tip:** CLB is legacy and only available for EC2-Classic (accounts created after Dec 2013 don't have it). Don't pick CLB for new designs-ALB or NLB depending on your layer.

### Cross-zone load balancing

| ELB type | Default            | Charges               |
|----------|--------------------|-----------------------|
| ALB      | Enabled            | None                  |
| CLB      | Disabled (CLI/API) | None                  |
| NLB      | Disabled           | Data transfer applies |
| GWLB     | Disabled           | Data transfer applies |

Enabling cross-zone lets ALBs forward to targets in any AZ regardless of where the request came in. For NLBs, it matters more-you'll incur data transfer charges between AZs if enabled.

## Target groups

Target groups define *who* gets the traffic:

- **Target types:** EC2 instance IDs, IP addresses (VPC CIDR, RFC 1918, RFC 6598), or a single Lambda function
- **Health checks:** optional for Lambda, mandatory for others
    - HTTP/HTTPS: path-based, configurable success codes
    - TCP: no path, timeout fixed at 10 seconds
- **Deregistration delay:** puts targets in draining state so in-flight requests complete
- **Session stickiness:** disabled by default, can use ELB-generated cookies (AWSALB/AWSELB) or application-based cookies

> **Exam tip:** You can register the same target in multiple target groups-useful for path-based routing to different backends. But target group attributes (like health check paths) can't be changed after creation.

## Application Load Balancers (ALB)

Key features that show up on the exam:

- **Listeners:** protocol/port combinations; each can have rules with conditions (host header, path, headers, methods, query string, source IP)
- **Rules:** must have a default rule (lowest priority), conditions are ANDed, comparisons within a condition are ORed
- **X-Forwarded headers:** `X-Forwarded-For`, `X-Forwarded-Proto`, `X-Forwarded-Port` (pass client info to targets)
- **Scheme:** internet-facing (public subnets) or internal (private subnets)
- **Dualstack IPv6:** only works with internet-facing scheme; subnets must support IPv6

> **Exam tip:** You cannot change the ALB scheme after creation. Plan internet-facing vs. internal correctly upfront. Also, ALBs require at least two subnets across different AZs for high availability.

## Network Load Balancers (NLB)

- **Layer 4**, routes on protocol/port, supports **static IPs** (Elastic IPs for IPv4, manual assignment for IPv6)
- **Static IPs cannot be changed** after assignment-plan carefully
- **Proxy Protocol v2:** delivers source IP to targets (separate from "preserve client IP" setting)
- **TLS termination:** NLB can terminate TLS, but session stickiness by source IP doesn't work with TLS targets
- **No IPv6** support for GWLB (important distinction!)

> **Exam tip:** NLB preserves the client IP at the packet level. Unlike ALB's X-Forwarded-For headers, NLB targets see the actual source IP unless you disable "preserve client IP."

## Gateway Load Balancers (GWLB)

This one's weird enough to deserve its own section:

- **Not a traditional LB**-it forwards traffic to security/appliance targets, then sends it back to the original destination
- Uses **GENEVE** encapsulation; original IP data stays intact
- **No FQDN**, accessible only as a VPC endpoint service
- Requires route table configuration to send traffic to GWLB endpoints
- **IPv6 not supported**
- Targets must support GENEVE protocol and allow traffic from GWLB

> **Exam tip:** GWLB traffic flows like this: client → GWLB endpoint → target (security appliance) → original destination. The endpoint has no DNS name; you route traffic there via VPC route tables. It's for inline inspection, not serving client requests directly.

## CloudFront

CloudFront is AWS's global CDN. Understanding cache behaviors and security is exam-critical.

### Core concepts

| Term           | Definition                                                   |
|----------------|--------------------------------------------------------------|
| Origin         | Where content is stored (S3, ELB, EC2, Lambda, MediaPackage) |
| Distribution   | Configuration linking origins to edge locations              |
| Edge location  | Geographically dispersed PoPs that cache content             |
| Cache behavior | Path-pattern rules (e.g., `*.png` → S3)                      |

### Origin Access Control (OAC) vs. OAI

|                      | OAI (legacy)   | OAC (successor)                                |
|----------------------|----------------|------------------------------------------------|
| Identity type        | IAM user       | Service principal (`cloudfront.amazonaws.com`) |
| Permissions          | `s3:GetObject` | `s3:GetObject` + `s3:PutObject`                |
| Supports S3 KMS?     | No             | Yes                                            |
| Supports PUT/DELETE? | No             | Yes                                            |

OAC is the modern choice-it grants more permissions and works with SSE-KMS buckets.

### Viewer protocol policies (HTTPS enforcement)

- **HTTP and HTTPS:** default, allows both
- **Redirect HTTP to HTTPS:** returns 302, but pre-HTTP/1.1 clients get 403
- **HTTPS only:** rejects plain HTTP with 403

> **Exam tip:** CloudFront certificates **must** be issued in `us-east-1`, even if your distribution serves users globally. Custom origins (EC2/on-prem) need externally trusted certs; S3 handles this automatically.

### Signed URLs and cookies

- **Signed URLs:** individual files, good for clients that don't support cookies
- **Signed cookies:** multiple files, better UX for video streaming
- Both expire after a configurable duration

### Cache key customisation

The cache key determines whether CloudFront serves a cached object or fetches from origin. Default includes:
- Distribution FQDN
- URL path

Customising it (headers, query strings, cookies) can cause **duplicate caching** or **reduced hit ratios**. Be deliberate-invalidation is expensive (first 1,000 paths/month free, then charged per path).

> **Exam tip:** Use **versioned filenames** (e.g., `style.v3.css`) instead of invalidation for frequent updates. Invalidation takes time to propagate; versioning returns immediately.

## Lambda@Edge vs. CloudFront Functions

Both run code at edge locations, but they're not interchangeable.

| Feature      | Lambda@Edge                       | CloudFront Functions |
|--------------|-----------------------------------|----------------------|
| Language     | Node.js, Python                   | JavaScript **only**  |
| Memory       | 3 GB                              | 2 MB                 |
| Code size    | ~50 MB                            | 10 KB                |
| Startup      | Sub-second                        | **Sub-millisecond**  |
| Requests/sec | 10k per region                    | Millions             |
| Events       | Viewer, Origin (request/response) | Viewer events only   |

> **Exam tip:** Choose CloudFront Functions for lightweight transformations (header manipulation, A/B testing). Use Lambda@Edge when you need more compute power or access to origin response logic.

## Global Accelerator

Unlike CloudFront (HTTP-only), Global Accelerator handles **TCP/UDP** with static Anycast IPs.

- **Two static IPs** (standard) or **four** (dual-stack IPv4 + IPv6)
- Routes traffic based on user location, health, and weights
- **Standard accelerator:** ALBs, NLBs, EC2, EIPs
- **Custom accelerator:** VPC subnets with custom application logic

> **Exam tip:** Global Accelerator is perfect for non-HTTP protocols (game servers, VoIP, real-time apps). The static IPs don't change, which matters for firewalls and allow-lists.

## Lambda in VPCs

By default, Lambda runs in AWS-owned VPC with public internet access. Deploying to your VPC requires:

- **Hyperplane ENIs:** managed resources that provide VPC-to-VPC NAT
- **Security Groups:** must be assigned to Hyperplane ENIs
- **NAT Gateway:** required for internet access (Lambda gets no public IP)

> **Exam tip:** Functions sharing the same subnet and security group can **share Hyperplane ENIs**-this reduces ENI count and improves performance.

## EKS networking

- Clusters are **always** created in a VPC with ≥2 subnets in different AZs
- DNS resolution and hostnames must be enabled on the VPC
- Control plane lives in AWS-managed VPC; worker nodes in your VPC
- Traffic flows between VPCs; ALB for ingress, NLB for LoadBalancer services

> **Exam tip:** EKS admin endpoint access is public by default. Lock it down with VPC endpoint policies if your security posture requires it.

## WorkSpaces and AppStream 2.0

### WorkSpaces (Desktop-as-a-Service)

- Operates from AWS-managed VPC, not your VPC
- Requires **Multi-AZ directory services** (Simple AD, Managed Microsoft AD, or AD Connector)
- Reserves 5 IPs per subnet + 1 for directory service per subnet
- CIDR ranges are region-dependent: `172.31.0.0/16`, `192.168.0.0/16`, `198.19.0.0/16`

### AppStream 2.0 (App streaming)

- Streams desktop apps via HTML5 browser
- Instances spawn fresh for each user
- Internet access depends on concurrent users:
    - **>100 users:** private subnets + NAT gateway
    - **<100 users:** public subnets + IGW acceptable
- Management ENI uses `198.19.0.0/16` range

> **Exam tip:** AppStream port 443 is for HTTPS; port 8433 is for UDP over HTTPS (Windows-native clients). Don't forget port 53 if you're using custom DNS.

# Domain 3: Hybrid networking – VPN, Direct Connect, and transit architectures{#domain3}

This is the heaviest domain. The exam loves BGP path selection, VIF types, and Transit Gateway routing rules-and it loves mixing them up in scenario questions.

## Virtual Private Gateways (VGW)

- One VGW attaches to **one VPC at a time**, but can carry multiple VPN or Direct Connect connections
- **ASNs:** AWS's default is 64512 in all regions (7224 in most regions prior to June 30, 2018)
- **Private ASN ranges** for your side: 64512–65534 (16-bit) or 4200000000+ (32-bit). Public ASNs are controlled by IANA-don't try to use one you don't own.

> **Exam tip:** The single most important VGW limitation: **VGWs are not transitive.** Traffic cannot traverse a VPC to reach another VPC or another VPN/DX connection. This is exactly why Transit Gateway exists.

## Route learning and propagation

| Method        | How it works                 | Best for               |
|---------------|------------------------------|------------------------|
| Static        | Prefixes manually configured | Small, stable networks |
| Dynamic (BGP) | Peers exchange prefixes      | Anything real-world    |

**AWS routing conflict logic** (the exam WILL test this):

1. **Longest prefix match** wins first (VPC route beats overlapping propagated route)
2. VPC **static routes** beat matching **propagated** routes
3. Among propagated routes: **DX → Static VPN → BGP VPN**

So Direct Connect is preferred over VPN when both are propagated. Keep this order memorised-you'll need it for failover scenarios.

## BGP path selection (memorise the sequence)

BGP builds tables of prefixes (routes) shared between configured peers over TCP 179. Nothing is advertised automatically-you control exactly what your peers learn. Path selection order:

1. **Highest Weight** (local to this router only; Cisco default 32,768 for own prefixes, 0 for learned)
2. **Highest Local Preference** (default 100; shared with iBGP peers)
3. **Shortest AS Path**
4. **eBGP over iBGP**
5. **Lowest MED/Metric** (not passed beyond the neighboring AS)

**The three ways you influence path selection on your side:**

- **Local preference** – raise it for prefixes you want preferred on outbound iBGP
- **AS path prepending** – add your ASN multiple times to a path (e.g., `65301 65301 65301 i`) to make it *less* attractive. This is how you signal "backup path" to the outside world.
- **Multi-Exit-Discriminator (MED)** – advertise a lower metric to your peer to make yourself look better (remember: peers don't forward it onward)

> **Exam tip:** AS path prepending makes a route **less** preferred, not more. If a question asks how to prefer a primary DX link over a backup VPN, you prepend on the *backup* path-or use local pref on AWS-managed routes via BGP communities.

**Rules worth knowing:**
- Prefixes learned from an iBGP peer are never re-advertised to other iBGP peers (that's why route reflectors exist on-prem)
- VGWs automatically advertise everything they learn from BGP peers
- You cannot directly configure BGP on AWS's BGP services-you influence them through communities and path attributes

## VPN and IPSec fundamentals

**How a tunnel establishes** (IKE phases matter for troubleshooting questions):

1. **Interesting traffic** triggers tunnel establishment (traffic initiated on-prem → AWS, by default)
2. **IKE Phase 1**: peers authenticate, negotiate key-exchange settings, run Diffie-Hellman, then re-authenticate with the derived key. One bidirectional SA for management traffic.
3. **IKE Phase 2**: negotiates new SAs and keys-two one-way SAs for actual site-to-site traffic
4. Tunnels eventually expire (Phase 2 has shorter lifetimes)

**Policy-based vs. route-based VPNs** (classic exam topic):

|                   | Policy-based                                                        | Route-based                 |
|-------------------|---------------------------------------------------------------------|-----------------------------|
| SAs               | One pair per matched rule set                                       | Single pair for all traffic |
| Traffic selection | Admin-defined rules ("interesting traffic")                         | Destinations in route table |
| Caveat            | Diverse security policies need multiple SA pairs-which wait in line | Simplest option             |

> **Exam tip:** AWS VPN tunnels support only a **single pair** of one-way IPSec SAs. If you configure policy-based VPN with multiple rule sets, subsequent SAs have to wait until existing tunnels expire. The fix: use one catch-all policy, or switch to route-based VPN. This exact scenario shows up constantly.

## Site-to-Site VPN architecture

**Customer gateway (CGW) requirements:**
- Support IKE (v2 supported since Feb 2019) and IPSec
- Public IP or ACM certificate (pre-shared key is default authentication)
- **Dead Peer Detection** (required)
- BGP for dynamic routing

**Ports to allow (both directions):** UDP 500, IP protocol 50 (ESP), and UDP 4500 for NAT traversal.

**Every VPN connection gets:**
- **Two tunnel endpoints in different AZs**, active/passive by default
- BGP peering over a `/30` from `169.254.0.0/16` (IPv4) or `/126` from `fd00::/8` (IPv6)
- Both endpoints must be configured on your CGW

> **Exam tip:** VGW VPN throughput caps at **1.25 Gbps** across all connections, and the VGW uses a single tunnel endpoint for return traffic. Need more throughput or true active/active? That's Transit Gateway with ECMP territory.

**Accelerated Site-to-Site VPN:**
- Routes on-prem traffic into the nearest Global Accelerator edge, then over the AWS backbone
- **Requires Transit Gateway** – not compatible with VGW
- To "convert" an existing VGW VPN, create a new accelerated VPN (there's no in-place upgrade)
- Pricing: connection hours + accelerator hourly + data transfer premium

## Client VPN and VPN CloudHub

### Client VPN (AWS-managed OpenVPN)
- Endpoint attaches to **one VPC**; use DNS to reach others
- Client IP CIDR cannot overlap target networks
- Authentication: Active Directory, mutual (certificate), or SAML 2.0
- **Split tunnel** avoids hairpinning all client traffic through AWS - full tunnel sends internet-bound traffic out the VPC's IGW instead
- Optional: client connect handler (Lambda, runs after auth) and CloudWatch logging

### VPN CloudHub
- Hub-and-spoke VPN between **your own sites** via a VGW – with or without a VPC attachment (detached mode)
- Each CGW needs a **unique BGP ASN**
- No overlapping IP ranges; up to 10 CGWs
- Cheaper than site-to-site meshes; sometimes used as office-to-office backup

> **Exam tip:** The common thread across VPN questions: **overlapping CIDRs are unsupported everywhere** - peering, CloudHub, DX. When a scenario offers a "third-party VPN appliance" answer, the trigger conditions are usually non-IPSec protocols (GRE, DMVPN), overlapping CIDRs, or needing transitivity.

## Direct Connect

### Connection types

|               | Dedicated                          | Hosted                      |
|---------------|------------------------------------|-----------------------------|
| Who orders?   | You, directly                      | A DX Partner via their APIs |
| Speeds        | 1 or 10 Gbps                       | 50 Mbps – 10 Gbps           |
| Cross-connect | LOA-CFA sent to DX location        | Handled by partner          |
| VIFs          | Up to 50 public/private, 1 transit | **One VIF only**            |

Dedicated requests take up to 72 hours. The **LOA-CFA** (valid 90 days) authorises the DX location to physically patch your equipment to AWS hardware.

**Hardware requirements:** single-mode fiber, 1000BASE-LX or 10GBASE-LR transceivers, auto-negotiation disabled, 802.1Q VLANs, BGP with MD5. **BFD** is supported (milliseconds-fast failure detection vs. BGP's default 270-second timeout) but not required.

### Virtual interfaces (VIFs) – THE exam table

DX connections are **Layer 2**; VIFs provide the Layer 3. Each VIF rides its own VLAN.

| VIF type    | Connects to                   | Reachability                                                       |
|-------------|-------------------------------|--------------------------------------------------------------------|
| **Private** | One VGW **or** one DX Gateway | Your VPC(s), private-space                                         |
| **Transit** | One DX Gateway ↔ TGW          | VPCs behind a TGW                                                  |
| **Public**  | (AWS side)                    | All AWS public services in all public regions - no internet needed |

**DX Gateway pairing rule:** a DXGW connects with **either** private VIFs → VGWs (up to 10 VGWs, 30 VIFs) **or** transit VIFs → TGWs (up to 3 TGWs). Never a mix.

**Prefix limits (customer → AWS):** 100 prefixes on private VIFs, 1,000 on public. Exceed them and BGP sessions drop to IDLE.

**AWS → customer advertisements depend on the gateway:**
- VGWs attached to private VIFs advertise **all** known routes
- VGWs behind a **DX Gateway** advertise only your declared allowed prefixes
- Public VIFs: AWS advertises **everything public**, plus non-regional services like CloudFront and Route 53 – filter what you learn with ACLs or BGP communities

### BGP communities (worth memorising verbatim)

**Applied by AWS to prefixes advertised TO your public VIF:**

| Community   | Meaning                             |
|-------------|-------------------------------------|
| `7224:8100` | Same region as the DX connection    |
| `7224:8200` | Same continent as the DX connection |
| *(none)*    | Global services                     |

**Applied by YOU to control where your prefixes propagate:**

| Community   | Scope              |
|-------------|--------------------|
| `7224:9100` | Local AWS region   |
| `7224:9200` | Continent          |
| `7224:9300` | All public regions |

**Applied by you on private/transit VIFs to steer AWS→you return traffic:**

| Community   | Preference |
|-------------|------------|
| `7224:7300` | **High**   |
| `7224:7200` | Medium     |
| `7224:7100` | Low        |

Equal-preference prefixes are load-balanced - this is your primary/backup and load-sharing toolkit across multiple DX connections. All AWS-advertised prefixes also carry the `no_export` community.

### Link Aggregation Groups (LAGs)

- Up to **4 connections** per LAG, all same bandwidth, same DX location
- Max **10 LAGs per region**
- **Minimum links** setting: if active links drop below the threshold, the whole LAG goes down (default 0 = no minimum)
- Adding existing connections interrupts connectivity briefly; you can't remove a connection if it breaches the minimum-links threshold

### MACsec

- Layer 2 encryption on the DX link itself: adds 8-byte header and 16-byte trailer per Ethernet frame
- Configure via a CKN/CAK key pair on both the DX interface/LAG and your router
- Options per connection: `should_encrypt`, `must_encrypt`, or `no_encrypt`
- Doesn't replace IPSec for end-to-end encryption – but covers the "AWS doesn't encrypt DX traffic" gap at Layer 2

### Resiliency tiers and troubleshooting

Well-Architected resiliency ladder: single connection → LAG (one location, multi-link) → multiple connections at one DX location → multiple DX locations → optionally VPN as BGP-preferred failover (create VPN backups with automated routing that flips when DX paths withdraw prefixes).

**Troubleshooting flowchart** (symptom → layer):
- **Can't reach the AWS device from the LOA-CFA?** Layer 1: cables, ports, transceivers, auto-negotiation
- **Connection up, VIFs down?** Layer 2: 802.1Q / VLAN configuration mismatch
- **VIF up, BGP won't establish?** Wrong ASN, wrong peer IP, MD5 key mismatch, prefix limit exceeded, or TCP 179 blocked
- **Everything up, no reachability?** Layer 3+: route tables, NACLs, security groups, application config

**Monitoring note:** DX connections have CloudWatch metrics (`ConnectionState`, `ConnectionBpsEgress/Ingress`, `ConnectionPpsEgress/Ingress`, `ConnectionCRCErrorCount`, `ConnectionLightLevelTx/Rx`) - but **VIFs have no CloudWatch metrics**; check them manually.

### MTU reference (fixed and consolidated)

I had inconsistencies in my original notes – here's the corrected, single table:

| Path                      | Max MTU |
|---------------------------|---------|
| Within VPC (jumbo frames) | 9,001   |
| DX private VIF            | 9,001   |
| DX transit VIF            | 8,500   |
| DX public VIF             | 1,500   |
| VPC peering               | 1,500   |
| Internet                  | 1,500   |

Jumbo frames give you more payload per packet – private VIFs move from 1,500 to 9,001. Caveats: toggling jumbo frames **disrupts all VIFs on the connection for up to 30 seconds**, and it's unsupported via VGW static routes. Set the "Do Not Fragment" flag and undersized paths will drop oversized packets rather than fragment them.

### The break-even math

When does DX beat internet egress pricing? With internet DTO at $0.05/GB, DX DTO at $0.02/GB, and port hours added on top: a 1 Gbps port breaks even around **7,200 TB** of monthly transfer; a 10 Gbps port around **54 PB**. The point of the exercise: DX is rarely justified on price alone until your egress is enormous – the real drivers are consistency, latency, and privacy.

## Hybrid DNS

The core problem: **EC2 instances always use the Route 53 Resolver (VPC+2 address) by default, and on-prem systems can't reach it.**

Resolver search order within a VPC: **private hosted zones → VPC domain name → public DNS.**

The fix ladder, worst to best:

1. **EC2-hosted DNS resolvers** – you run and HA-design them yourself; a single ENI tops out at 1,024 queries/second and most clients won't spread load across servers
2. **Simple AD as a forwarder** – better, but still self-managed
3. **Route 53 Resolver endpoints** – the real answer:
    - **Inbound endpoints:** on-prem → AWS, so on-prem resolvers can query your private hosted zones
    - **Outbound endpoints:** AWS → on-prem, governed by **forwarding rules** matched on FQDN patterns
    - 2–8 ENIs per endpoint, **10,000 qps per ENI**, one AZ-scoped SG set at creation, priced per ENI-hour

Rule conflict resolution: **most specific FQDN wins, and forward rules beat system rules.** System rules exist automatically for private hosted zones, VPC domain names, and publicly reserved domains.

> **Exam tip:** For inbound endpoints to resolve private hosted zones, the zones must be associated with the **same VPC where the endpoint ENIs live**. People lose points assuming the association can be anywhere in the Region.

## Inter-VPC connectivity and the transit decision tree

| Option                          | Transitive?       | Scale limit                                                                              | Character                                |
|---------------------------------|-------------------|------------------------------------------------------------------------------------------|------------------------------------------|
| VPC peering                     | No                | 50 peering connections per VPC                                                           | Non-transitive, mesh gets ugly fast      |
| VPC sharing (RAM)               | N/A (same subnet) | Same AWS Organization                                                                    | Owner manages, others consume            |
| PrivateLink / endpoint services | No                | ~55k concurrent conns per NLB                                                            | One-way consumer→provider, same region   |
| Transit VPC                     | Yes               | Bounded by route table limits (50 static/100 propagated per VPC, 100 dynamic VGW routes) | DIY EC2 VPN hub; you manage availability |
| Transit Gateway                 | Yes               | **5,000 attachments, 255,000 interconnected networks**                                   | The AWS-native answer                    |

Other peering notes: intra-region peering supports SG ID references (cross-region must use CIDR blocks), and request/accept with manual route entries on both sides is mandatory. PrivateLink consumers must be in the **same region** as the provider, and providers approve access per AWS account, IAM user, or role.

## Transit Gateway (TGW)

The centerpiece of Domain 3. A TGW interconnects thousands of VPCs, VPNs, DX gateways, and even peer TGWs across regions/accounts.

### Creation-time settings

- **TGW ASN** (AWS side of BGP): 64512–65534 or 4200000000–4294967294
- **DNS support**, **VPN ECMP**, **multicast** toggles
- **Default route table association/propagation** - leave these on and every attachment can reach every other attachment; the moment you need isolation, turn them off and create per-segment route tables
- Soft limit of 5 TGWs per account; 20 route tables per TGW

### Attachments and routing

Attachments are billed hourly (VPC owner pays for VPC attachments, TGW owner for VPN, DX owner for DX, both sides for peering) plus **per-GB data processed**.

- **Associations:** each attachment associates with exactly **one** TGW route table
- **Propagations:** an attachment can propagate its CIDRs to *multiple* route tables (association not required for propagation)
- **VPC attachments** propagate their local CIDR; VPN/DX attachments propagate BGP-learned prefixes
- **Static routes:** one default route per table, blackhole routes for deliberate drops, and static routes are *required* to reach peered TGWs
- **TGW → VPC:** routes must be manually added to VPC route tables pointing at the TGW

### TGW routing conflict logic (different from VGW!)

Longest prefix wins → **static beats propagated** → among propagated: **VPC > DX > VPN**.

Note the flip compared to VGW conflict logic (where DX beat VPN): at TGW level, the VPC attachment's route is preferred when things collide.

> **Exam tip:** The isolation pattern the exam keeps testing: default settings = full mesh, zero security segmentation. The correct secure design disables default association/propagation and gives each workload segment (prod/dev/shared services) its own TGW route table. Check for Cloud WAN on the exam too – it wraps this same policy model in a declarative document, with segments, segment actions, and core network edges (TGW under the hood per region).

### Limits worth memorising

- 5,000 attachments per TGW, of which max **20 DX Gateway** attachments and **50 peering** attachments
- 10,000 static routes; 20 route tables
- DX Gateway side: 3 TGWs per DXGW, 30 transit VIFs
- VPN fleet per region: 5 VGWs, 10 VPNs per VGW, 50 site-to-site VPNs total

### Monitoring

TGW CloudWatch metrics include `BytesIn/Out`, `PacketsIn/Out`, `PacketDropCountBlackhole`, `BytesDropCountBlackhole`, and `BytesDropCountNoRoute` / `PacketDropCountNoRoute`. Flow logs only work *within* attached VPCs – **TGW itself has no flow logs**, so deeper visibility into transitive traffic means a transit VPC or Network Firewall in the path. And remember: **CloudTrail does not monitor network traffic**, anywhere.

# Domain 4: Design and implement for security and compliance{#domain4}

This domain overlaps heavily with the Security specialty exam, but AN-C01 tests it through a networking lens – perimeter control, traffic inspection, and encryption in flight.

## Traffic control services

### AWS Shield

- **Standard:** free, automatic, protects Route 53, CloudFront, and Global Accelerator at every edge location. Covers roughly 96% of known Layer 3/4 attacks.
- **Advanced:** $3,000/month with a 1-year commitment. Adds protection for EC2, EIPs, and ELBs, volumetric bot attack coverage, DDoS cost-protection (EDoS coverage), 24/7 SRT engagement, WAF charges included, and centralised management via Organisations/Firewall Manager.

> **Exam tip:** Shield is hosted at **CloudFront, Global Accelerator, and Route 53 edge locations only** – an unprotected ALB sitting behind no CDN doesn't benefit until you buy Advanced. Also memorise the Advanced finding metrics: `DDoSDetected`, `DDoSAttackBitsPerSecond`, `DDoSAttackPacketsPerSecond`, `DDoSAttackRequestsPerSecond` (viewed by `AttackVector` dimension; CloudFront/Route 53 metrics land in us-east-1).

### AWS WAF

WAF inspects HTTP/HTTPS requests and allows, denies, or counts them against **ALBs, API Gateway, and CloudFront**. Structure matters:

- **Conditions** contain one or more filters (ORed together): cross-site scripting, country of origin, IP address, request-property size, SQL injection, string/regex matching
- **Rules** contain one or more conditions (ANDed), and can be regular or rate-based
- **ACLs** contain rules processed in listed order, plus a default action
- **Rate-based rules** trigger only after a request threshold, with counters resetting every 5 minutes

Quotas to remember: 10 conditions per rule, 10 rules per ACL, 5 rate-based rules per account/region, and 10 regex conditions (non-raisable).

> **Exam tip:** Once an ACL is associated with an object type, it can only be associated with other objects of that same type – and ALB/API Gateway associations make the ACL **region-restricted**. Web ACLs for CloudFront metrics live in us-east-1.

### Network Firewall

The managed stateful/stateless firewall for VPC perimeters - sits in-line at IGWs, NAT gateways, VPNs, and DX entry points. Uses **Suricata rules** for stateful inspection, supports domain-name filtering (via TLS **SNI**, not decryption), and does smart protocol detection (recognises HTTPS even off port 443).

**Architecture:** traffic from your private subnets gets routed to a **firewall endpoint** (its own dedicated subnets, referenced as a VPC endpoint ID in route tables), passes the rules engine, and permitted traffic continues to its destination.

Components: **firewall** (per VPC) ↔ **firewall policy** (one per firewall, many firewalls per policy) ↔ **rule groups**.

|            | Stateless rules                 | Stateful rules                  |
|------------|---------------------------------|---------------------------------|
| Context    | Examines packet alone (5-tuple) | Examines packet in flow context |
| Processing | Ordered, first match wins       | Pass > Drop > Alert priority    |
| Default    | -                               | Allow all                       |
| Analogy    | NACLs (but smarter)             | Security groups                 |

Stateless actions are **pass, drop, or forward** (forward hands the packet to the stateful engine) – no analog in NACL-land. Stateful rule types: standard rules (translated to Suricata), Suricata rules proper, and domain list rules.

> **Exam tip:** Unlike an EC2-based firewall appliance, Network Firewall endpoints scale and are managed – but traffic mirroring-based IDS and marketplace appliances still have their place. The exam tests knowing which fits "managed perimeter inspection with domain filtering."

### Firewall Manager

Centralised, cross-account deployment and compliance enforcement of firewall policy types: **WAF, Shield Advanced, security groups, Network Firewall, Route 53 Resolver DNS Firewall**, plus the third-party CNF options (Palo Alto Cloud NGFW, FortiGate CNF). Out-of-compliance resources generate **findings sent to Security Hub**.

## Detection services

### GuardDuty
Continuous threat detection using threat intelligence and ML; findings include credential exposure, escalated privileges, malicious IP interactions, and cryptocurrency mining. Centrally managed via Organisations. It consumes VPC Flow Logs, DNS logs, and CloudTrail events as input sources.

### Inspector
Automated vulnerability and **network reachability** scanning. Network assessments (agentless, for EC2 and Lambda) run every 24 hours and report overly permissive security groups/NACLs; host assessments use SSM Agent. Centralise findings with a delegated administrator.

### Traffic Mirroring
Copies ENI traffic out-of-band to a security or monitoring appliance.

- **Source** (the monitored ENI) → **filter** (inbound/outbound rules, numbered, processed low-to-high) → **target** (another ENI, an **NLB**, or a **GWLB**) via **session**
- Targets behind an NLB/GWLB may receive **out-of-order packets**
- Mirror targets must support **VXLAN decapsulation**
- Supported on Nitro-based instances; enhanced networking metrics to watch: `NetworkMirrorIn/Out`, `NetworkSkipMirrorIn/Out`

> **Exam tip:** Traffic Mirroring captures traffic at the ENI level – inside security groups but you control scope via filters. It complements VPC Flow Logs (which only see L3/L4 metadata, and notably do NOT capture Route 53 resolver traffic, DHCP, instance metadata, Windows license activation, time sync, or traffic to the default VPC router).

## Encryption in flight

### Certificate landscape

- **ACM:** free, integrates with ELB/CloudFront/API Gateway/etc., auto-renews, **keys and certificates cannot be exported**, per-region, CloudFront requires **us-east-1**, public certs valid 13 months, domain-validated via DNS or email
- **IAM certificate store:** legacy path - can't create certs, no console management
- **ACM Private CA:** $400/month per CA plus per-certificate charges – for internal PKI
- **CloudHSM:** FIPS 140-2 **Level 3** hardware, single-tenant, ENIs into your subnets – not integrated with AWS services, but it's the answer for private CA root keys, Oracle TDE master keys, and offloading TLS crypto

> **Exam tip:** The recurring trap is cert **location**: CloudFront and edge-optimised API Gateway custom domains need us-east-1; regional ALB/API Gateway certs must be in their own region; custom origins (EC2/on-prem) need publicly trusted certs (ACM or otherwise); **S3 static website endpoints don't support HTTPS at all**.

### Protocol enforcement quick reference

- **CloudFront:** Viewer Protocol Policy = HTTP+HTTPS (default) / Redirect / HTTPS-only (403s otherwise)
- **S3:** `aws:SecureTransport: false` bucket-policy condition to deny non-HTTPS
- **ELBs:** ALB = HTTPS listeners; NLB = TLS listeners; CLB = HTTPS or SSL
- **RDS:** engine-specific enforcement; apply the AWS root cert bundle to clients
- **Intra-instance:** EBS encryption at rest, EFS mount-time TLS, FSx encryption with SMB 3/4.2+

## Governance and compliance

Keep the framing crisp: **governance is control, compliance is proof.**

- **Organisations + SCPs** set the maximum effective permissions across child accounts
- **CloudTrail** audits every API call (multi-account trails into one bucket) but **never network traffic**
- **AWS Config** tracks resource configuration drift and auto-remediates via rules
- **Service Catalog** forces approved CloudFormation products so nobody hand-builds a VPC
- **Security Hub** aggregates findings from GuardDuty, Inspector, Firewall Manager, and third parties
- **EventBridge + Lambda + SNS** close the loop with automated response

> **Exam tip:** When a question mixes "we need to detect" with "we need to prevent," map services accordingly: Config/Inspector/GuardDuty detect, SCPs/WAF/Network Firewall/SCPs prevent, and Security Hub consolidates. Least-privilege IAM plus separate accounts is always at least half the right answer to governance questions.

# Automation appendix (small domain, quick points)

## CloudFormation essentials for networkers

**`Fn::Cidr`** generates subnet CIDRs from a parent block - perfect for consistent subnetting:

```yaml
# !Cidr [ ipBlock, count, cidrBits ]
!Cidr [ '10.0.0.0/16', 6, 8 ]
# → 10.0.0.0/24 ... 10.0.5.0/24
# cidrBits 8 means subnet mask = 32 - 8 = /24
```

**Deployment ordering:** CloudFormation deploys in parallel by default; **`DependsOn`** forces sequencing when one resource must exist first (e.g., route table entries that need a peering connection ID).

**Stack policies:**
- `CreationPolicy` - pause resource completion until signals (typically from cfn-init/ASG instances)
- `DeletionPolicy` - Retain/Snapshot/Delete on stack teardown
- `UpdatePolicy` - controls in-place replacement behaviour during updates

**Cross-account peering:** create a role in the accepting account (`PeerRoleArn`) with `sts:AssumeRole` from the requester, then pass `PeerVpc` + `PeerRoleArn` (+ `PeerRegion` for cross-region) on `AWS::EC2::VPCPeeringConnection`.

**Troubleshooting ladder:** stack Events tab → `Get System Log` for instance boot logs → `/var/log` and `c:\cfn\log` for cfn-init output → CloudWatch log streams.

## Cloud Map

Service discovery for dynamic backends: create a **namespace**, create a **service** per resource type, instances **self-register** via `RegisterInstance`, and consumers resolve via **`DiscoverInstances`**. Deeply integrated with ECS; EKS works via the ExternalDNS connector. Supports DynamoDB, S3, SQS, and more as discoverable services.

## CDK
Infrastructure-as-code in TypeScript, JavaScript, Python, Java, C#, and Go – it synthesises to CloudFormation under the hood.

---

The Advanced Networking Specialty rewards two things above all else: **pattern recognition across hybrid architectures** and **tolerance for very specific numbers**. When I sat down for the exam, the questions weren't really "what is a transit VIF"-they were "given THIS topology with THESE constraints, which design satisfies ALL the requirements?" The answers that feel obviously wrong get eliminated instantly; the pain lives in choosing between the two plausible remainder options.

Good luck – and if these notes saved you some head-scratching, pass them on to the next person staring down the same blueprint.
