# AWS Generative AI Developer - Professional

*2026-07-01*

> Notes from studying for the AWS Generative AI Developer - Professional certification


Almost two years after finishing my [original eight-certification journey]({{< relref "/posts/2024/getting-aws-certified" >}}), I've added the [**AWS Certified Generative AI Developer – Professional**](https://aws.amazon.com/certification/certified-generative-ai-developer-professional/) to the list. With how quickly GenAI has evolved, this felt like the natural next step.

I prepared using Stephane Maarek's [*Ultimate AWS Certified Generative AI Developer Professional*](https://www.udemy.com/course/ultimate-aws-certified-generative-ai-developer-professional/?couponCode=MT260922G1A) course on [Udemy](https://www.udemy.com), switching platforms midway for deeper hands-on labs with emerging services.

These notes focus on the GenAI-specific concepts – the places where traditional AWS knowledge doesn't fully carry you. If you're tackling this exam, I hope they save you some pain.

---

## Contents

1. [Part 1: Foundations – Bedrock and Model Customisation](#part-1-foundations-bedrock-and-model-customisation)
2. [Part 2: RAG Deep Dive](#part-2-rag-deep-dive)
3. [Part 3: Guardrails, Prompts, and Flows](#part-3-guardrails-prompts-and-flows)
4. [Part 4: Data Preparation Services](#part-4-data-preparation-services)
5. [Part 5: OpenSearch, Vector Stores, and Databases](#part-5-opensearch-vector-stores-and-databases)
6. [Part 6: Agentic AI](#part-6-agentic-ai)
7. [Part 7: Performance, Cost and Optimization](#part-7-performance-cost-and-optimization)
8. [Part 8: SageMaker AI](#part-8-sagemaker-ai)
9. [Part 9: Integration and Infrastructure](#part-9-integration-and-infrastructure)
10. [Part 10: Governance and Quality Assurance](#part-10-governance-and-quality-assurance)
11. [Part 11: Architectural Decision Guide](#part-11-architectural-decision-guide)

## Part 1: Foundations – Bedrock and Model Customisation{#part-1-foundations-bedrock-and-model-customisation}

### Foundation models on AWS

- **Jurassic-2 (AI21 Labs)** – multilingual text generation
- **Claude (Anthropic)** – conversation, Q&A, workflow automation
- **Stable Diffusion (Stability AI)** – image, art, logo, and design generation
- **Llama (Meta)** - open-weight LLMs
- **Amazon Titan** – summarisation, generation, Q&A, embeddings, personalisation, search
- **Amazon Nova Pro** – the Nova LLM portfolio
- **Amazon Nova Reels** – video generation

> **Model catalogs churn constantly – verify the current lineup against the exam guide before sitting it.**

### Amazon Bedrock

A serverless API for generative foundation models, integrating with SageMaker Canvas. Its API splits across four endpoints, and knowing *which one does what* is exam gold:

| Endpoint                | Purpose                           | Key operations                                                               |
|-------------------------|-----------------------------------|------------------------------------------------------------------------------|
| `bedrock`               | Manage, deploy, train models      | -                                                                            |
| `bedrock-runtime`       | Inference against models          | `Converse`, `ConverseStream`, `InvokeModel`, `InvokeModelWithResponseStream` |
| `bedrock-agent`         | Manage agents and knowledge bases | -                                                                            |
| `bedrock-agent-runtime` | Inference against agents/KBs      | `InvokeAgent`, `Retrieve`, `RetrieveAndGenerate`                             |

The **Converse API** takes a prompt in the `messages` field plus a `modelId`, optionally with guardrails, inference config (max tokens, temperature), prompt variables, and tools for agentic workflows. Access control is straightforward IAM (`AmazonBedrockFullAccess` / `AmazonBedrockReadOnly`).

### Fine-tuning vs. continued pre-training vs. LoRA

Three customisation techniques, easy to confuse:

- **Fine-tuning** - additional supervised training on *your labeled data* (prompt/completion pairs). Ideal for chatbots with a specific personality, fresher data, proprietary corpora (emails, support transcripts), or narrow tasks like classification. You can fine-tune a fine-tuned model. Use VPC/PrivateLink for sensitive data – and budget carefully, custom models get expensive.
- **Continued pre-training** - same idea but with *unlabeled* data, essentially injecting domain knowledge (business documents, jargon) into the model itself.
- **Low-Rank Adaptation (LoRA)** – injects small trainable "low-rank matrices" into the attention weights and merges them at inference. The base model stays unchanged on disk, making it storage-, training-, and inference-efficient. Note: this differs from bolting an *adapter layer* on top of a frozen model.

Customisation is supported for Titan, Cohere, and Meta models.

## Part 2: RAG Deep Dive{#part-2-rag-deep-dive}

RAG is an "open-book exam" for LLMs: pull relevant context from an external store, weave it into the prompt, and let the model answer from it. Compared to fine-tuning it's faster and cheaper – updating knowledge means updating the database, not retraining – and it enables semantic search without touching model weights. But it's not a silver bullet: RAG is sensitive to prompt templates and retrieval relevance, it's non-deterministic, and it *can still hallucinate*. (One reviewer memorably called it *"the world's most overcomplicated search engine."*)

### Embeddings and vector search

An embedding maps data to a point in a high-dimensional space (hundreds to thousands of dimensions), positioned so similar items sit close together. Retrieval is then:

1. Embed the query,
2. Search the vector store for nearest neighbors,
3. Return the top-N most similar items.

Vector store options: OpenSearch, SQL databases, Neptune, Redis, MongoDB, Cassandra – plus purpose-built engines (Pinecone, Weaviate commercially; Chroma, Marqo, Vespa, Qdrant, LanceDB, Milvus open source).

**Sparse vs. dense embeddings:** sparse vectors are mostly empty (like one-hot encodings) but yield stronger similarity signals; dense vectors pack semantic meaning into fewer dimensions and are far more efficient. Similarity is typically measured with **cosine similarity** – the angle between vectors.

### Chunking strategies (a favorite exam topic)

Chunking splits documents before embedding, keeping within the model's context limit:

- **Fixed-size** - set tokens per chunk with an overlap percentage. Default is 300 tokens, honoring sentence boundaries.
- **Default** – 300 tokens per chunk, sentence-boundary aware.
- **No chunking** – each document is one chunk (pre-process yourself if needed).
- **Hierarchical** - nested parent/child chunks; search hits precise small child chunks, then swaps in the parent for fuller context.
- **Semantic** - the model splits on meaning shifts. Parameters: max tokens, buffer size (surrounding sentences considered; 1 = 3 sentences), and breakpoint percentile threshold (higher = fewer, larger, more distinct chunks).

Worked example – fixed chunking with size 5, 20% overlap:

> "Space, the final frontier. These are the voyages of the Starship Enterprise."
> → *Space, the final frontier. These* / *These are the voyages of* / *of the Starship Enterprise.*

### Retrieval optimization

- **Vector size ≠ chunk size.** Smaller vectors (fewer dimensions) cut cost at some retrieval-quality expense - balance dimensionality against how semantically complex your domain really is. Titan defaults to 1024+ dimensions.
- **Metadata** enriches retrieval: Bedrock Knowledge Bases let you pass a `metadata.json` distinguishing content from metadata, so you don't waste embeddings on things like creation dates but can still filter and rank on them (document IDs, categories, access controls, lineage).
- **Pre-retrieval** work (chunking granularity, data extraction, query rewriting), retrieval, and **post-retrieval** work all matter.
- Keep KBs fresh with a Lambda-triggered embedding regeneration, batched on a schedule rather than real-time.
- **Re-ranker models** (Amazon or Cohere, limited regions) score retrieved chunks against the query and reorder them – invoked via the `rerank` operation or when hitting a KB.

### Bedrock Knowledge Bases

Upload documents (via S3, optionally with a JSON schema), pick an embedding model (Cohere or Titan – you control vector dimensions), and choose a vector store: serverless OpenSearch by default, or MemoryDB, Aurora, MongoDB Atlas, Pinecone, Redis Enterprise Cloud.

### Measuring RAG quality

Correctness, completeness, helpfulness, logical coherence, faithfulness (does the response stick to retrieved text?), citation precision and coverage, harmfulness/stereotyping, and refusal behavior. Evaluation typically uses a JSON prompt dataset with reference responses and contexts, with **LLM-as-judge** scoring defined metrics – remembering that different judge models score differently.

## Part 3: Guardrails, Prompts, and Flows{#part-3-guardrails-prompts-and-flows}

### Bedrock Guardrails

Filters prompts and responses for text foundation models: word/topic filtering, profanities, and PII removal or masking. The **Contextual Grounding Check** scores grounding (response similarity to supplied context) and relevance (response to query) – a key hallucination defense. Guardrails plug into agents and knowledge bases, and you can customise the blocked-response message.

**Automated Reasoning Checks** enforce complex policies (mortgage eligibility, medical protocols) by converting a well-organised policy PDF into structured rules - via the `CreateAutomatedReasoningPolicy` API.

**Token-level redaction** handles cases where guardrails fall short: pre/post-processing Lambdas around the inference endpoint use pattern matching or NER (e.g., Amazon Comprehend) to strip sensitive tokens before the model ever sees them.

### Prompt engineering

Anatomy: instructions, context, input data, output indicator. Best practices: be clear and concise, provide context, specify the response format, place the desired output last, phrase inputs as questions, and provide example responses. Break complex tasks into sub-tasks - and don't be shy about asking the model to "think step by step."

Prompt types:

- **Zero-shot** - no examples; lean on model scale.
- **Few-shot** - provide worked examples.
- **Chain-of-thought** - explicit step-by-step reasoning.

**Prompt injection** ("## ignore the above and..."), roleplay jailbreaks around guardrails, and **prompt leaking** (extracting system prompts or PII) are the main misuse vectors – mitigated with guardrails and hardened system prompts.

### Mitigating bias

- **Disambiguate prompts** – have users specify attributes rather than letting the model default (see TIED/TAB frameworks for text-to-image).
- **Fix the data** - audit and rebalance training data; analyse outputs for imbalances.
- **Counterfactual data augmentation** - detect, segment, augment.

### Prompt management and Flows

Bedrock **Prompt Management** stores reusable, versionable prompts with `{{variable}}` placeholders, variants per model, and optional tool/cache associations. **Bedrock Flows** chain prompts and models via nodes and (possibly conditional) connections – built visually or defined as JSON. **Structured outputs** come via schema-referencing prompts, Tool Use in the Converse API, or response format templates.

## Part 4: Data Preparation Services{#part-4-data-preparation-services}

### Bedrock Data Automation (BDA)

The modern path for multimodal ingestion: documents, images, video, audio → structured output ready for vector stores.

- **Standard output** guesses format from input type (documents → JSON, audio → transcript).
- **Custom output** uses **Blueprints** - basic/table/grouped/custom-typed fields, even creatable from a prompt. Blueprints support classification, extraction, normalisation (keys and values), transformation, and validation. Configurations live in a **Project**, invoked via `InvokeDataAutomationAsync`.
- Documents (PDF, TIFF, JPEG, PNG, DOCX) → JSON ± files (CSV, markdown, HTML), with page/element/word-level granularity and bounding boxes.
- Images → summaries, IAB taxonomy, logos, text, moderation.
- Video (MP4, MOV, AVI, MKV, WEBM) → chapter summaries, transcripts, text, logos, moderation.
- Audio (AMR, FLAC, M4A, MP3, Ogg, WAV) → transcripts, speaker labeling, topics, moderation.

### Other data tools

- **Unstructured text**: preserve structure by converting to HTML (pandoc, Textract, Comprehend), pipeline via Glue. **Divider strings** inserted by a Lambda preprocessor improve chunking.
- **Converse API conversations** need JSON with `role` and `content`.
- **SageMaker Data Wrangler**: visual import/transform (300+ transforms) with a "Quick Model" sanity check. Troubleshooting: IAM permissions (`AmazonSageMakerFullAccess`) and EC2 instance-quota increases.
- **AWS Glue**: Crawlers populate the Data Catalog from S3 (schemas, partitions by `yyyy/mm/dd/device` etc.), enabling Athena/Redshift Spectrum queries. **Glue Studio** builds visual ETL DAGs; **Glue Data Quality** uses DQDL rules that can fail jobs or report to CloudWatch.
- **Transcribe**: ASR with PII redaction and automatic language identification. Boost accuracy with **Custom Vocabularies** (specific words: brands, acronyms) and **Custom Language Models** (domain-trained context) – use both for best results. **Toxicity Detection** reads tone, pitch, and text cues.
- **Comprehend**: NLP for entities, key phrases, sentiment, topics. Custom classification (real-time or async), custom NER for domain entities (policy numbers, escalation phrases). Comprehend + Lambda is a common pre-Bedrock data-quality gate. **Comprehend Medical** is HIPAA-eligible with a separate `DetectPHI` API.
- **CloudWatch metrics** basics (namespaces, dimensions - max 30 per metric, timestamps) and **Metric Streams** (→ Kinesis Firehose → Datadog, Dynatrace, New Relic, Splunk, Sumo Logic) show up on the exam.

## Part 5: OpenSearch, Vector Stores, and Databases{#part-5-opensearch-vector-stores-and-databases}

### OpenSearch essentials

Elasticsearch/Kibana fork: documents live in indices, split into shards (self-contained Lucene indexes) distributed across nodes; writes go primary-then-replica, reads hit primaries or replicas. Managed domains support dedicated master nodes, snapshots to S3, and zone awareness.

**Storage tiers:** hot (instance/EBS), **UltraWarm** (S3-backed, for rarely-written log data, requires dedicated masters), and **cold** (even cheaper, for forensic analysis – requires both dedicated masters *and* UltraWarm, incompatible with T2/T3 data nodes).

**Index State Management** automates deletion, read-only transitions, hot→UltraWarm→cold migration, replica reduction, and snapshots (runs every 30–48 minutes with jitter). Rollups summarise old indices; transforms create alternate analytical views. **Cross-cluster replication** follows leader indices for HA/geo-latency, needing fine-grained access control and node-to-node encryption.

**Sizing rules** (frequently examined):
- Stability: 3 dedicated master nodes to prevent split-brain; minimum storage = `Source Data × (1 + Replicas) × 1.45`
- Shards: `(Source Data + Growth) × (1 + Indexing Overhead) / Desired Shard Size`
- Anti-patterns: OLTP → use RDS/DynamoDB; ad-hoc queries → use Athena.

### OpenSearch as a vector store

- **Serverless collections** (search or time-series type), KMS-encrypted, measured in OCUs (lower bound: 2 indexing + 2 search).
- **Semantic vs. hybrid search**: hybrid adds keyword search on filterable metadata fields – often better relevance.
- ANN engines: FAISS, NMSLib, Lucene. Algorithms - HNSW (fast, high-quality, RAM-hungry) and IVF (scales to huge datasets, trades recall for speed/memory). Tuning knobs: `M` (edges/node - higher = more recall, more memory), `ef_construction` (graph build accuracy vs. speed), `ef_search` (recall vs. query latency).
- **Compression**: binary vectors (32× vs. float32), FP16 scalar quantization.
- **Sharding**: fewer, larger shards (30–50 GB) for semantic search; 10–30 GB for hybrid.
- The **neural plugin** lets OpenSearch call Bedrock directly for embeddings during ingestion/search.
- Watch JVM memory pressure - fewer, better-balanced shards fix it.

### Amazon S3 Vectors

The new cheap-and-simple option: create a vector bucket → index (dimensions + distance metric) → `put_vectors` / `query_vectors`. Roughly 90% cheaper, 100 ms–1 s latency, strongly consistent, max 10,000 indices per bucket (2B vectors per index). Best practices: batch puts (500/call), concurrency (2,500 vectors/sec), retries before 429s, multiple indexes for tenancy. For performance-critical queries, pair with OpenSearch via a tiered strategy – noting the S3 Vectors engine in OpenSearch managed clusters can mean paying for both.

### Relational and NoSQL

- **RDS SQL Server** has vector capabilities; a common pattern couples RDS for structured queries with S3 holding the unstructured blobs.
- **Aurora PostgreSQL + pgvector** - `vector` column type, cosine/L2/inner-product distances, IVF indexing. Wins when you need complex filtering, transactions, and joins alongside similarity search; suits small-to-medium RAG.
- **DynamoDB** – not a vector store, but great for chat history and agent memory. Zero-ETL integration pipelines into OpenSearch let DynamoDB data feed Bedrock KBs.

Capacity math (memorise the formulas):

- **WCU**: one write/sec for items up to 1 KB. E.g., 10 writes/sec of 2 KB items = 10 × 2 = **20 WCUs**.
- **RCU**: one strongly consistent read (or two eventually consistent) per sec for items up to 4 KB. E.g., 16 EC reads/sec of 12 KB items = 8 × 3 = **24 RCUs**.
- **Partitions**: `ceil(max(RCU/3000 + WCU/1000, Size/10GB))`.

**Maintenance**: vectors drift and fragment - schedule EventBridge triggers to AWS Batch jobs that rebuild indexes, validate, and swap atomically. DAX adds caching (5-min TTL, up to 10 nodes, multi-AZ).

## Part 6: Agentic AI{#part-6-agentic-ai}

### Bedrock Agents

An agent = LLM + memory (chat history, external stores) + planning guidance + tools. **Action Groups** define tools with Lambda functions and parameters (name, description, type, required) – the description tells the model *when* to use the tool and what to extract. Knowledge bases attach similarly ("use this for questions about X") – **agentic RAG** treats retrieval as just another tool. A **code interpreter** lets agents write their own code and charts. Deploy via an **alias** (snapshot) using `InvokeAgent`, with on-demand or provisioned throughput.

### Multi-agent patterns

- **Routing** – a router LLM classifies and dispatches to specialist agents.
- **Parallelisation** – independent subtasks at once (multiple guardrails, multiple evaluations) or **voting** (same prompt, several models).
- **Prompt chaining** - discrete sequential steps, with gates for early exit.
- **Evaluator–optimiser** – a generator model plus a critic that iterates until good enough.
- **Orchestrator–synthesiser** - delegating and merging results.

Use multiple agents when tool count overwhelms selection or logic turns into spaghetti conditionals – otherwise, keep it simple with deterministic workflows.

### Memory

Short-term: session context in memory, ElastiCache/MemoryDB, or DynamoDB. Long-term: extracted insights, preferences, session summaries as **memory records** – via AgentCore Memory, Mem0, or Knowledge Bases.

### The modern agent stack

- **Strands Agents** – Python SDK for single- and multi-agent systems with strong AWS integration (Bedrock, Lambda, Step Functions), multimodal and MCP support, and no AWS lock-in. Agent loop: input → tool selection → execution → LLM reasoning → repeat.
- **Agent Squad** - open-source multi-agent orchestration (Python/TypeScript): intent-classification routing, supervisor agents, shared context, Bedrock Flows integration.
- **Bedrock AgentCore** - serverless, framework-agnostic agent deployment: runtime (endpoints from ECR containers, observability via CloudWatch GenAI dashboards), built-in browser/code-interpreter tools, a **Gateway** converting APIs/Lambda/OpenAPI/Smithy into MCP tools with OAuth management, **Identity** (agent-specific identities via Cognito), and **Cedar-policy** enforcement (deny-by-default, contextual validation, enforce or log-only modes). `agentcore import-agent` converts Bedrock agents to Strands code.
- **Model Context Protocol (MCP)** – the "USB-C port for AI applications": JSON-RPC 2.0 over stdio or HTTP streaming. Deploy stateless servers on Lambda, complex ones on ECS/Fargate, front with API Gateway.
- **Humans in the loop**: draft-and-refine patterns, escalation on confidence/risk, and feedback collection (API Gateway → DynamoDB) to measure model/variant preference.

### Amazon Q Business

Managed GenAI assistant over company data with fully managed RAG (S3, SharePoint, Confluence, M365, Salesforce, Slack...), plugins (Jira, ServiceNow, Zendesk), **IAM Identity Center** for document-permission-aware responses, and admin-level guardrail controls (global and topic-level). **Q Apps** builds no-code GenAI apps over that data.

## Part 7: Performance, Cost, and Optimization{#part-7-performance-cost-and-optimization}

### Token efficiency

- `CountTokens` estimates request tokens free-of-charge; CloudWatch tracks `InputTokenCount`/`OutputTokenCount`, TTFT, latency, throttles.
- **Techniques**: context pruning (fewer, metadata-filtered chunks; summarising old history), response limiting (`MaxTokens`, few-shot verbosity examples, JSON forcing), prompt compression (summarise via a small model first).

### Model selection and caching

- Right-size models – put the smarts in RAG and tools, use small models for preprocessing (summarisation, classification, chunking). **Dynamic routing** by query complexity via Bedrock Intelligent Prompt Routing, Lambda, Agent Squad, or Strands.
- Measure price:performance with **Bedrock Evaluations** (human or LLM judges, model comparison) plus token counting.
- **Semantic caching** – cache prompt *embeddings* + responses in a vector store (Valkey/MemoryDB); return cached answers when cosine similarity exceeds a tuned threshold. Tune carefully and mind the overhead.
- **Prompt caching** - Bedrock-native caching of static prefixes (instructions, few-shot examples) behind a checkpoint; cached reads are cheaper per token (writes may cost more).
- **Edge caching** - fingerprint deterministic requests for CloudFront with aggressive TTLs.

### Throughput and resilience

- **Batch inference** (S3 in/out) for embeddings; capacity planning via TPM/RPM quotas; tensor parallelism shards weights across GPUs.
- **Provisioned throughput** (model units/token throughput) is mandatory for customised models – pass the provisioned model ARN.
- **Latency-optimised inference**: `performanceConfig={'latency': 'optimized'}` improves TTFT, OTPS, and end-to-end latency.
- **Cross-region inference profiles** (geography-scoped or global) boost throughput with ~10% savings - but don't mix with provisioned throughput, and check your org SCPs allow the regions.
- Resilience patterns: **exponential backoff** with jitter (100 ms start, factor 2, 3–5 retries), **connection pooling** (10–20 connections, 60–300 s TTL), parallel requests via multi-agent or Step Functions.
- Retrieval tuning: hybrid search, normalised queries, splitting multi-part questions.

### Inference parameters

Temperature (randomness), Top_p (nucleus sampling threshold – specify this *or* temperature), Top_k (candidate pool). A/B test with Bedrock Evaluations or CloudWatch Evidently.

## Part 8: SageMaker AI{#part-8-sagemaker-ai}

Covers the full ML workflow: data prep (RecordIO/Protobuf, Athena/EMR/Redshift ingestion) → processing jobs → training → deployment (persistent endpoints or Batch Transform, with JumpStart's 150+ models, Inference Pipelines, Neo for edge, Elastic Inference, shadow testing).

- **Optimized FM deployment**: single/multi-model and multi-container endpoints, inference components with per-model scaling, **Bedrock Custom Model Import** (train in SageMaker, serve serverlessly in Bedrock), model servers (TorchServe, DJL, Triton), async endpoints with SNS/SQS queues, and compression (quantization, pruning, knowledge distillation). Large models: `ml.p4d.24xlarge` GPUs; small: `ml.c5.9xlarge` CPUs. Models up to 500 GB with adjusted download-timeout quotas.
- **Ground Truth** - managed human labeling (Mechanical Turk, internal, vendor teams); Ground Truth Plus is the turnkey expert-managed version.
- **Model Monitor + Clarify** - CloudWatch alerts for data drift, model-quality drift, bias drift, and feature-attribution drift (NDCG-ranked). Clarify's pre-training bias metrics are worth memorising: Class Imbalance (CI), Difference in Proportions of Labels (DPL), KL/JS divergence, Lp-norm, TVD, Kolmogorov-Smirnov, and Conditional Demographic Disparity.
- **Model Registry** (versions, approval status, CI/CD) and **ML Lineage Tracking** (trials, trial components, experiments, contexts, actions, artifacts, associations – queryable via the `LineageQuery` API, cross-account via `AddAssociation`).
- **SageMaker Neo** - compile once for ARM/Intel/Nvidia edge targets, deploy via IoT Greengrass.
- **Unified Studio** consolidates data, analytics, AI, and ML; **Pipelines** gives visual DAG workflows; **MLflow** (managed) adds experiment tracking and model management.

## Part 9: Integration and Infrastructure{#part-9-integration-and-infrastructure}

- **Lambda** - connects agents to tools (validation, error handling), invokes FMs on demand without capacity provisioning, handles webhooks/API Gateway events, aggregates multi-model output.
- **API Gateway** - fronts models and feedback collection; usage plans with throttling (~10–50 rps) and burst (~2–3×), request validators + JSON schemas for token limits, routing via request transformations.
- **AppConfig** - feature-flag model switching and rollback without code changes; pairs with Bedrock Evaluations/CloudWatch Evidently for A/B tests.
- **Step Functions** - circuit-breaker and fallback-model patterns, ReAct-style structured reasoning, model approval workflows (mind the 256 KB inter-step limit). Bedrock-native (`InvokeModel` with guardrails, `CreateModelCustomizationJob`).
- **Outposts/Wavelength** - GenAI where data can't move (jurisdictional compliance, on-prem privacy) or at the 5G edge for ultra-low latency.
- **Networking**: IGW/NAT distinctions, NACLs (stateless, subnet-level, allow/deny, IP-only) vs. security groups (stateful, ENI-level, allow-only), PrivateLink (NLB + ENI, no peering), ALB/NLB/GLB layers, Global Accelerator vs. CloudFront.
- **Extras**: DataSync (10 Gbps per agent, preserves POSIX/SMB metadata), Transfer Family (FTP/FTPS/SFTP into S3/EFS), Athena SPICE, Kinesis (shards: 1 MB/s in, 2 MB/s out; on-demand vs. provisioned), App Runner, EKS node types (managed/self-managed/Fargate), Lex+Connect, ElastiCache for Valkey vector search (>99% recall, microsecond latency), Neptune Analytics `topKByEmbedding`.

## Part 10: Governance and Quality Assurance{#part-10-governance-and-quality-assurance}

### Agent tracing

Every response includes a trace: PreProcessing, Orchestration, PostProcessing, CustomOrchestration, RoutingClassifier, Failure, and Guardrail types – your primary debugging surface.

### Evaluation methods

- **Human evaluation** – best for UX feel, contextual sensitivity, and creativity; hard to scale given GenAI's non-determinism.
- **Benchmark datasets** - SME-authored prompts scoring accuracy, speed, scalability, and context retrieval.
- **LLM-as-judge** - beware shared blind spots between judge and judged; works well for screening small models.
- **Metrics to know cold**: **ROUGE** (recall-oriented, overlapping n-grams; ROUGE-1/2, ROUGE-L via longest common subsequence for coherence), **BLEU** (precision-oriented n-grams for translation, with brevity penalty), **BERTScore** (embedding-based semantic similarity, robust to synonyms).
- **Bedrock Model Evaluations**: automatic, human (Cognito/Ground Truth/A2I teams), LLM-judge, and RAG-specific jobs (retrieve-only: relevance/coverage; retrieve-and-generate: correctness, completeness, helpfulness).

### Deployment validation

Unit tests don't cut it for non-deterministic systems: simulate synthetic user workflows end-to-end (CloudWatch canaries, Step Functions, EventBridge, Lambda → S3/Athena/QuickSight), validate hallucination rate and faithfulness, and check response consistency against a prompt dataset.

### Responsible AI

The eight dimensions: fairness, explainability, privacy & security, safety, controllability, veracity & robustness, governance, transparency. Tools: Bedrock Evaluations, SageMaker Clarify/Model Monitor/ML Governance, Amazon A2I.

### Monitoring and compliance

Composite CloudWatch alarms (AND/OR conditions to cut noise), prompt-regression testing in logs, KPI dashboards, hallucination rates, anomaly detection (token bursts, drift), Bedrock invocation logs, cost anomaly detection, CloudWatch RUM for mobile apps, and CloudTrail for the full who-called-what compliance trail. X-Ray troubleshooting (IAM roles, daemon on EC2, active tracing on Lambda). Lake Formation for data-lake governance: governed tables with ACID, row/column/cell-level security via data filters, cross-account permissions, blueprint-driven ingestion workflows.

Macie (ML-based PII discovery) rounds out the security toolset.

## Part 11: Architectural Decision Guide{#part-11-architectural-decision-guide}

### Choosing a vector store

| Option                | Best fit                  | Why                                               | Cost             | Ops burden | Key limitations                              |
|-----------------------|---------------------------|---------------------------------------------------|------------------|------------|----------------------------------------------|
| OpenSearch (managed)  | Custom RAG at scale       | Mature, deeply tunable                            | $$–$$$ always-on | Medium     | You pay for provisioned capacity             |
| OpenSearch Serverless | Spiky traffic, low ops    | Managed collections                               | $–$$ usage-based | Low        | Less tuning control, variable latency        |
| Kendra                | Enterprise doc search     | Document-level ACLs (SharePoint, Confluence, Box) | $$$              | Very low   | Less tinkering; pricey at scale              |
| Aurora + pgvector     | RAG + relational in one   | Joins, transactions, filters + vectors            | $$               | Medium     | Your scaling/ANN tuning                      |
| Neptune Analytics     | Graph + vector reasoning  | Fraud rings, lineage, supply chains               | $$–$$$           | Medium     | Analytics engine; vector-loading constraints |
| S3 Vectors            | Massive scale, cost-first | Cheap, simple, S3-native                          | $                | Very low   | Metadata/filtering limits, higher latency    |

**Decision shortcuts:**
- SharePoint/Confluence/ACL requirements → **Kendra**
- Graph relationships → **Neptune Analytics**
- Already on Postgres, need joins/transactions → **pgvector**
- Huge corpus + cost pressure → **S3 Vectors**
- Full search-platform control → **OpenSearch managed**
- Unpredictable traffic, minimise ops → **OpenSearch Serverless**

### Orchestration: when Step Functions?

Choose it when you need auditable state transitions, retry/failure isolation, or explicit (especially human) approval steps.

---

## Final thoughts

Back when I took the Machine Learning Specialty, I had little real data science background – and I figured it was as hard as certifications got. I was wrong.

The GenAI Developer – Professional challenged me just as much, if not more. Not because of the math, but the breadth: Bedrock, half a dozen vector store options, agent frameworks that barely existed last year. It's less one exam than ten small ones stapled together. Ironically, my ML Specialty prep mostly paid off on the SageMaker questions.

My advice: don't underestimate the classic AWS content hiding in there (DynamoDB capacity math, networking, OpenSearch sizing), build comparison tables for everything, and accept that the ecosystem moves faster than you can study – focus on trade-offs, not version numbers.

Worth it? Yes. Good luck!
