# AWS Machine Learning Specialty

*2024-07-11*

> Notes from studying for the AWS Machine Learning Specialty certification


Earlier this year I passed the [AWS Certified Solutions Architect – Professional exam]({{< relref "/posts/2024/aws-solutions-architect-professional/" >}}). Then I did something arguably harder: I tackled the [AWS Certified Machine Learning – Specialty](https://aws.amazon.com/certification/certified-machine-learning-specialty/) with almost no prior machine learning experience.

Of the six certifications I've earned, this was the one I felt least prepared for. Not because the material is impossible – but because unlike networking or compute services, I couldn't lean on day-to-day engineering experience to fill the gaps. Everything here came from studying the official [AWS Certified Machine Learning – Specialty learning path](https://www.pluralsight.com/paths/aws-certified-machine-learning-engineer-associate-mlac01) on [Pluralsight](https://www.pluralsight.com/), and these notes are what I captured along the way.

If you're an infrastructure or development person approaching this exam without a data science background, I hope these notes bridge that gap the way I wish something had for me.

*(Quick disclaimer: exam blueprints evolve – always verify against AWS's official certification page.)*

---

## Contents

1. [Data Collection](#data-collection)
2. [Streaming Data Collection](#streaming-data-collection)
3. [Data Preparation](#data-preparation)
4. [Data Analysis and Evaluation](#data-analysis-and-evaluation)
5. [Modelling](#modelling)
6. [SageMaker Modelling](#sagemaker-modelling)
7. [Algorithms](#algorithms)
8. [Evaluation and Optimisation](#evaluation-and-optimisation)
9. [Implementation and Operations](#implementation-and-operations)
9. [Security](#security)
10. [Monitor and Evaluate](#monitor-and-evaluate)

## Data Collection{#data-collection}

### What makes data "good"?

Before touching AWS services, the exam cares about whether you can *recognise* quality data – because garbage in, garbage out.

| Good data                     | Bad data                    | Why it matters                                   |
|-------------------------------|-----------------------------|--------------------------------------------------|
| Large datasets                | Small datasets (<100 rows)  | More data generally means better model training  |
| Precise, relevant attributes  | Useless attributes          | Models need to train on the features that matter |
| Complete fields               | Missing values / nulls      | Missing data skews results                       |
| Consistent values             | Inconsistent values         | Models depend on clean, consistent inputs        |
| Balanced outcome distribution | Skewed outcome distribution | Models can't learn when one outcome dominates    |
| Fair sampling                 | Biased sampling             | Bias in, bias out                                |

### Core terminology

The exam loves definitions – know these cold:

- **Dataset** – your input data; used for training and testing. Columns are called *attributes* or *features*; rows are *observations*, *samples*, or *data points*.
- **Structured data** - has a defined schema that explains how to interpret it.
- **Unstructured data** – no defined schema (raw text, images, audio).
- **Semi-structured data** – no rigid schema, but some organisational structure (CSV, JSON, XML).

> **Exam tip:** Be able to distinguish a **database** (transactional, strict schema), a **data warehouse** (schema-on-*write*, BI-ready), and a **data lake** (schema-on-*read*, many formats, raw). The write/read distinction is a classic question trap.

### Labelled vs unlabelled data

- **Labelled (supervised)** – you know the target attribute: spam/not-spam, churn labels, housing prices.
- **Unlabelled (unsupervised)** – no target attribute: tweets, logs, news articles. The model must find structure itself.

Two related terms that show up constantly:

- **Ground truth** - observed/measured data trusted as factual. (SageMaker Ground Truth the *service* helps you build labelled datasets through human and automated labelling.)
- **Corpus data** - text datasets used for NLP, speech recognition, and text-to-speech.

## Streaming Data Collection{#streaming-data-collection}

### The Kinesis family – the exam's favourite matching game

Know these four cold; questions almost always present a scenario and ask which one fits.

| Scenario                                                | Service            | Why                                                               |
|---------------------------------------------------------|--------------------|-------------------------------------------------------------------|
| Stream Apache logs from 100 EC2 instances into Redshift | **Firehose**       | Loads to a final destination; S3 first, then copies into Redshift |
| Stream live sporting-event video to customers           | **Video Streams**  | Real-time video/audio processing, feeds other AWS services        |
| Transform streaming data and feed a custom ML app       | **Data Streams**   | Huge ingest volumes, process/transform, feed into apps            |
| Run SQL over real-time data, output metrics to S3       | **Data Analytics** | SQL queries on streams, output to AWS destinations                |

The simplest way I kept them straight: **Data Streams** if *consumers* must act on data in real time (retention matters); **Firehose** if the goal is just getting data *to storage* (retention doesn't matter, processing optional); **Video Streams** for video/audio; **Data Analytics** for SQL.

## Data Preparation{#data-preparation}

Once you've gathered your data, you need to prepare it before feeding it to any ML algorithm. The AWS exam tests whether you understand *how* to transform data for different scenarios – not just what the tools are, but *when* to use them.

### Categorical Encoding

ML algorithms generally expect numerical inputs. When you have categorical values (categories, groups, classes), you need to encode them appropriately.

| Scenario                               | Algorithm                    | Encoding Required? |
|----------------------------------------|------------------------------|--------------------|
| Predicting house price                 | Linear Regression            | Yes                |
| Classifying text as sports/not-sports  | Naïve Bayes                  | No                 |
| Detecting malignancy in medical images | Convolutional Neural Network | Yes                |

#### Two main encoding strategies:

**One-hot encoding** – Creates new binary columns (0 or 1) for each category. Best for nominal features where order doesn't matter.

Example: House type = condo, house, apartment  
→ Type_condo (1/0), Type_house (1/0), Type_apartment (1/0)

**Mapping** – Assigns numbers to categories directly. Only use for ordinal features where order matters (S=5, M=7, L=10).

**Warning:** Don't map categories arbitrarily. If house type = condo(1), house(2), apartment(3), the model may incorrectly infer condo < house < apartment.

> **Exam tip:** When there are *many* categories, consider grouping rare values into "other" or using similarity-based grouping to reduce the number of new columns created.

### Text Feature Engineering

Transforming raw text so ML algorithms can analyse it effectively.

| Technique                          | What it does                                                             | When to use                                                        |
|------------------------------------|--------------------------------------------------------------------------|--------------------------------------------------------------------|
| **Bag-of-Words**                   | Tokenises text by whitespace into individual words                       | Simple frequency analysis                                          |
| **N-Gram**                         | Groups n consecutive words                                               | Matching phrases (e.g., "click here now", "you're a winner")       |
| **Orthogonal Sparse Bigram (OSB)** | Creates word pairs including the first word + delimiters                 | Finding common word combinations across documents                  |
| **TF-IDF**                         | Weights words by importance; downweights common terms like "the" and "a" | Determining document subject matter, filtering noise               |
| **Remove Punctuation**             | Strips punctuation characters                                            | When punctuation isn't semantically important                      |
| **Lowercase transformation**       | Converts all text to lowercase                                           | Standardising text before processing                               |
| **Cartesian Product**              | Combines two categorical/text features into new feature                  | Creating new features from combinations (e.g., textbook + binding) |
| **Date Feature Engineering**       | Extracts attributes from dates (day, month, weekday, holiday flag)       | Time-based analysis, business logic patterns                       |

Example of TF-IDF intuition:
- If "the" appears in both Document 1 and Document 2, it's less important for distinguishing them.
- If "Jedi" only appears in Document 1, it carries more weight for that document.

### Numeric Feature Engineering

Numeric values often need scaling to work well with ML algorithms.

#### Scaling techniques:

**Normalisation** – Rescales values from 0 to 1

Formula: `x = (x - min) / (max - min)`

**Problem:** Outliers can throw off normalisation significantly.

**Standardisation** – Rescales values so mean = 0 (zero-centred)

Formula: `z = (x - mean) / std`

**Advantage:** Less affected by outliers than normalisation.

**Quick example:**

| Price   | Normalised | Standardised |
|---------|------------|--------------|
| 250,555 | 0.578559   | 0.394923     |
| 125,700 | 0.021419   | -1.080890    |
| 120,900 | 0.000000   | -1.137627    |
| 345,000 | 1.000000   | 1.511283     |

#### Binning

Grouping numeric values into ranges reduces minor observation errors and makes patterns clearer.

**Quantile binning** creates bins with equal numbers of observations (not necessarily equal ranges).

Example:
| Age | Bin | Youth | Young Adult | Adult |
|-----|-----|-------|-------------|-------|
| 24 | Youth | 1 | 0 | 0 |
| 45 | Young Adult | 0 | 1 | 0 |
| 18 | Youth | 1 | 0 | 0 |
| 76 | Adult | 0 | 0 | 1 |

> **Exam tip:** There's no single "optimal" number of bins – it depends on the variable characteristics and relationship to the target. Experimentation is key.

### Handling Missing Values

Missing data can appear as null, NaN, NA, None, etc. How you handle it affects model performance.

| Mechanism                               | Description                                                               |
|-----------------------------------------|---------------------------------------------------------------------------|
| **Missing Completely at Random (MCAR)** | The fact that a value is missing has nothing to do with the data itself   |
| **Missing at Random (MAR)**             | Missingness relates to observed data, not the missing data itself         |
| **Missing Not at Random (MNAR)**        | Missingness depends on the hypothetical missing value or another variable |

| Technique               | Pros                                                       | Cons                            |
|-------------------------|------------------------------------------------------------|---------------------------------|
| **Supervised learning** | Best results - predicts missing values from other features | Most difficult to implement     |
| **Mean/Median/Mode**    | Quick and easy                                             | Results can vary significantly  |
| **Dropping rows**       | Easiest                                                    | Can dramatically change dataset |

Replacing missing values is called **imputation**.

> **Exam tip:** For Linear Learner specifically, missing values create sparse datasets. Use Factorisation Machines instead if you have lots of missing data points.

### Feature Selection

Selecting the most relevant features prevents over-complicated analysis and removes irrelevant or duplicate information.

**Principal Component Analysis (PCA)** – Unsupervised algorithm that reduces feature count while retaining as much information as possible.

| Problem           | Solution          | Why                                                      |
|-------------------|-------------------|----------------------------------------------------------|
| Too many features | PCA               | Reduces total feature count while preserving information |
| Useless features  | Feature selection | Removes features that don't help solve the problem       |

### AWS Data Preparation Tools

| Data Source                                | Recommended Tool        | Why                                         |
|--------------------------------------------|-------------------------|---------------------------------------------|
| S3, Redshift, RDS, DynamoDB, on-premise DB | **AWS Glue**            | Python or Scala for flexible transformation |
| S3                                         | **Athena**              | Query data and output results to S3         |
| EMR                                        | **PySpark/Hive in EMR** | Handle petabytes of distributed data        |
| RDS, EMR, DynamoDB, Redshift               | **Data Pipeline**       | Setup EC2 instances for ETL                 |

**AWS Glue vs SageMaker:**
- SageMaker + Jupyter notebooks → ad hoc analysis and experimentation
- AWS Glue jobs → reusable, production ETL pipelines

## Data Analysis and Evaluation{#data-analysis-and-evaluation}

Before training a model, you need to understand your data visually.

### QuickSight

AWS's BI tool for creating visualisations directly from your data in the console.

### Choosing the right visualisation

| Goal              | Visualisation                                                                       |
|-------------------|-------------------------------------------------------------------------------------|
| **Comparisons**   | Bar chart (lookup values), Line chart (change over time)                            |
| **Distributions** | Histogram (single distribution), Box plots / Scatter plots (multiple distributions) |
| **Compositions**  | Pie chart (static), Stacked bar / area chart (changing over time)                   |
| **Relationships** | Scatter plot (2 variables), Bubble chart (3 variables)                              |

| Chart type          | Best for                                                           |
|---------------------|--------------------------------------------------------------------|
| Scatter plot        | Relationships between two values                                   |
| Bubble plot         | Relationships between three values (bubble size = third dimension) |
| Bar chart           | Lookup and compare single variable values                          |
| Line chart          | Variables changing over time                                       |
| Histogram           | Distribution across binned values                                  |
| Box plot            | Lowest/highest values, outliers, quartiles                         |
| Pie chart           | Parts of a whole (static composition)                              |
| Stacked area column | Quantity over shorter time periods                                 |

## Modelling{#modelling}

### Understanding model types

| Learning Type     | Training Input               | Discrete Output               | Continuous Output        |
|-------------------|------------------------------|-------------------------------|--------------------------|
| **Supervised**    | Training + Testing data      | Classification                | Regression               |
| **Unsupervised**  | No training labels           | Clustering                    | Dimensionality reduction |
| **Reinforcement** | Trial and error with rewards | Simulation-based optimisation | Autonomous devices       |

### Choosing the right approach

| Problem                           | Approach                   | Why                                  |
|-----------------------------------|----------------------------|--------------------------------------|
| Detect financial fraud            | Binary Classification      | Only two outcomes: fraud or not      |
| Predict car deceleration rate     | Heuristic (no ML)          | Physics formulas already exist       |
| Find optimal path for lunar rover | Reinforcement Learning     | Trial, error, improvement needed     |
| Identify dog breed in photo       | Multi-class Classification | Many breed categories to choose from |

### The Confusion Matrix

This is genuinely one of the most tested concepts. Memorise this:

|                  | Predicted TRUE                         | Predicted FALSE                       |
|------------------|----------------------------------------|---------------------------------------|
| **Actual TRUE**  | Correct prediction                     | :cross: False Positive (Type I Error) |
| **Actual FALSE** | :cross: False Negative (Type II Error) | Correct prediction                    |

**Fraud detection analogy:**
- **False Positive:** Bank flags legitimate transaction → Angry customer, happy bank (false alarm)
- **False Negative:** Bank misses actual fraud → Happy customer, angry bank (missed loss)

> **Exam tip:** Different thresholds affect precision/recall trade-offs. Spam filters want few false positives (blocking legitimate email), while fraud detection may prioritise catching all fraud even with more false alarms.

### Data Preparation for Training

Remember: We want **generalisation**, not memorisation.

**Split strategy:**
1. Randomise the dataset
2. Split into training and testing sets
3. Train on training set
4. Test on test set

**K-Fold Cross-Validation:**
- Takes random splits multiple times
- Each fold becomes test data once
- Better estimate of model performance than a single split

## SageMaker Modelling{#sagemaker-modelling}

### Creating models

| Option                       | Use case                |
|------------------------------|-------------------------|
| SageMaker Console → Jupyter  | Interactive development |
| SageMaker SDK (Apache Spark) | Production pipelines    |

### Supported data formats

| Content-Type                    | Accept                          |
|---------------------------------|---------------------------------|
| application/x-recordio-protobuf | application/x-recordio-protobuf |
| text/csv                        | application/json                |
| application/jsonlines           | application/jsonlines           |
| application/x-image             | image/*                         |

**Performance note:** Protobuf recordIO format gives optimal performance. Pipe mode allows streaming data from S3 to training instances – uses less EBS space and starts faster.

### Training API basics

CreateTrainingJob API uses SageMaker SDK for Python:

- Specify training algorithm
- Supply algorithm-specific hyperparameters
- Specify input and output configuration

**Hyperparameter vs Parameter:**
- **Hyperparameter:** Set *before* learning (learning rate, number of epochs)
- **Parameter:** Derived *during* learning (weights, biases)

### SageMaker Training

GPU vs CPU performance hierarchy:
> CPU < FPGA < GPU < ASIC (for performance)  
> ASIC < FPGA < GPU < CPU (for flexibility)

**Common training parameters:**
- Channel name – Named input source for algorithm
- Training/inference image registry path - Versioned ECR images (`:1` for stable, `:latest` for bleeding edge)
- Training input mode – File or Pipe (Pipe performs better)
- File format – recordIO protobuf for best performance
- Instance class - GPU/CPU depending on algorithm requirements

### CloudWatch integration

CloudWatch logs:
- Arguments provided
- Errors during training
- Algorithm accuracy statistics
- Timing information

Common errors:
- Extra or invalid hyperparameters
- Incorrect protobuf file format

## Algorithms{#algorithms}

### Supervised vs Unsupervised vs Reinforcement Learning

| Category          | Training Data       | Output                                             |
|-------------------|---------------------|----------------------------------------------------|
| **Supervised**    | Labels required     | Classification (discrete), Regression (continuous) |
| **Unsupervised**  | No labels           | Clustering, Dimensionality reduction               |
| **Reinforcement** | Rewards/punishments | Policy that maximises reward                       |

Algorithms can be:
- SageMaker built-in
- Purchased from AWS Marketplace
- Built custom via Docker

### Regression

**Linear Learner**
- Supervised learning for regression, binary classification, or multi-class classification
- Maps vector x to approximate label y
- Minimises sum of distances from training points to prediction line (using Stochastic Gradient Descent)

**Use cases:**
- Predicting quantitative values (ROI based on marketing spend)
- Binary classification (should I mail this customer?)
- Multi-class classification (email, phone, or mail contact?)

**Factorisation Machines**
- Handles high-dimensional sparse data with missing values ("holes")
- Considers only pair-wise feature relationships
- Works on binary classification or regression only (not multi-class)
- Requires 10,000–10,000,000+ dimensions
- AWS recommends CPUs (better for sparse data)

> **Exam tip:** Factorisation Machines need recordIO-protobuf format – CSV is not supported.

### Clustering

**K-Means** (Unsupervised)
- Groups observations based on attribute similarity (Euclidean distance)
- Members of a group are as similar as possible to each other, as different as possible from other groups

**SageMaker K-Means specifics:**
- Modified web-scale version (more accurate than standard)
- CPU instances recommended (GPU can only use one core anyway)
- Must define number of clusters beforehand
- Training still occurs – model accuracy matters even without labels

> **Exam tip:** K-Means is great for finding hidden patterns in tabular data when you don't have ground truth labels.

### Classification

**K-Nearest Neighbours (KNN)**
- Index-based, non-parametric method
- Finds k closest points to sample and returns most frequent label (classification) or average value (regression)
- Lazy algorithm – doesn't learn during training, just stores data in memory

**Use cases:**
- Credit ratings (group by shared risk attributes)
- Product recommendations (similar items based on likeability)

### Image Analysis

| Algorithm                 | Task                       | Notes                                                            |
|---------------------------|----------------------------|------------------------------------------------------------------|
| **Image Classification**  | Hotdog/not hotdog          | Uses Convolutional Neural Networks (ResNet)                      |
| **Object Detection**      | Clock + lamp in scene      | Assigns classification + confidence                              |
| **Semantic Segmentation** | Pixel-level edge detection | Accepts PNG, requires GPU for training, CPU/GPU OK for inference |

> **Exam tip:** Semantic segmentation is computationally intensive - only GPU instances for training.

### Anomaly Detection

**Random Cut Forest (RCF)**
- Unsupervised anomaly detection
- Gives anomaly score (low = normal, high = outlier)
- Works well with n-dimensional data
- Doesn't benefit from GPU (use regular instances: ml.m4, ml.c5)

**Use cases:**
- Quality control (unusual audio frequencies)
- Fraud detection (unusual amounts, times, locations)

**IP Insights**
- Learns usage patterns for IP addresses paired with entities/user IDs
- Returns anomaly scores for entity/IP combinations
- Uses neural networks for latent vector representations
- GPUs recommended for training (distributed CPUs may be more cost-effective for large datasets)

**Use cases:**
- Tiered authentication (trigger 2FA for anomalous IPs)
- Fraud detection (allow/restrict activities based on login patterns)

### Text Analytics

**Latent Dirichlet Allocation (LDA)** (Unsupervised)
- Discovers topics within document corpora
- Documents are observations, words are features, topics are categories

**Neural Topic Model (NTM)** (Unsupervised)
- Similar to LDA but uses different algorithm → may produce different results

**Sequence-to-Sequence (seq2seq)** (Supervised)
- Input sequence → output sequence (translation, speech-to-text)
- Uses embedding, encoding, decoding layers (RNN/CNN)
- Pre-trained word vectors common (FastText, GloVe)
- GPU-only, single-machine training only

**BlazingText**
- Optimised Word2Vec and text classification
- ~20x faster than FastText
- Modes: Single CPU, Single GPU, Multiple CPU
- Word2Vec = unsupervised, Text Classification = supervised

**Object2Vec** (Supervised)
- Learns embeddings of high-dimensional objects
- Requires pairs of items with relationship labels
- Embeddings usable for downstream tasks

### Reinforcement Learning

"Carrot and stick" approach - positive/negative rewards to optimise agent policy.

**Markov Decision Process (MDP):**
1. Agent
2. Environment
3. Reward
4. State
5. Action
6. Observation
7. Episodes
8. Policy

**Use cases:**
- Autonomous vehicles (stay on road through trial/error)
- Intelligent HVAC control (learn building patterns, optimise energy)

### Forecasting

**DeepAR** (Supervised)
- Scalar time series forecasting using RNN
- Outperforms ARIMA and ETS by training single model over multiple time series
- Supports point forecasts ("X sneakers sold") and probabilistic forecasts ("X-Y with Z% probability")

**Requirements:**
- Minimum 300 observations across all time series
- Hyperparameters: context length, epochs, prediction length, time frequency
- Automatic backtest evaluation after training

**Use cases:**
- New product performance forecasting
- Labor needs prediction for special events

### Ensemble Learning

**XGBoost** (Supervised)
- Gradient boosted trees implementation
- Virtual "Swiss army knife" for regression, classification, ranking
- 2 required + 35 optional hyperparameters

**Technical constraints:**
- Accepts CSV and libsvm formats
- CPU-only training, memory-bound
- Needs lots of memory for entire training data
- Spark integration via SageMaker Spark SDK

**Use cases:**
- E-commerce search ranking (relevance scores)
- Fraud detection (transaction → fraud probability)

## Evaluation and Optimisation{#evaluation-and-optimisation}

### Feedback loop

1. **Define evaluation** – What metrics determine success?
2. **Evaluate** – Review during/after training (manual or automatic)
3. **Tune** – Adjust hyperparameters, data, or algorithm

### Validation types

| Type        | Description           | Example                         |
|-------------|-----------------------|---------------------------------|
| **Offline** | Using test data sets  | Validation sets, K-Fold         |
| **Online**  | Real-world conditions | Canary deployments, A/B testing |

### Monitoring training jobs

**SageMaker → CloudWatch**
- Algorithm metrics: `train:*` and `validation:*`
- Training job logs

### Model accuracy

|                | Underfitting                     | Overfitting                                                                    |
|----------------|----------------------------------|--------------------------------------------------------------------------------|
| **Definition** | Model doesn't reflect data shape | Model memorises training data                                                  |
| **Symptoms**   | Poor on training AND test data   | Great on training, poor on test                                                |
| **Fix**        | More data, train longer          | More data, early stopping, noise, regularisation, ensembles, feature selection |

| Training Error | Testing Error | Diagnosis              |
|----------------|---------------|------------------------|
| Low            | Low           | Ideal!                 |
| Low            | High          | Overfitting            |
| High           | High          | Try different approach |
| High           | Low           | Run away! (data leak)  |

### Accuracy metrics

**Regression:**
- **RMSE** (Root Mean Square Error) – lower is better

**Binary classification:**
- **Recall** = Right guesses / (Right + missed ones)
- **Precision** = Right guesses / (Right + wrong flags)
- **F1 Score** = Harmonic mean of precision and recall

> **Exam tip:** Different thresholds change F1 scores. Spam filters want minimal false positives; fraud detection may tolerate more to catch all fraud.

**Multiclass classification:**
- Per-class accuracy + macro/micro averaging
- F1 scores calculated per class

### Model tuning (hyperparameter tuning)

1. Choose tunable hyperparameter (not all can be auto-tuned)
2. Choose range of values
3. Choose objective metric to optimise

**Optimisation methods:**
- **Random search** – samples randomly across range
- **Bayesian optimisation** – focuses on promising regions based on previous results (avoids costly iterations)

## Implementation and Operations{#implementation-and-operations}

### Deployment types

| Type                  | Time | Risk | Cost                 |
|-----------------------|------|------|----------------------|
| **Big Bang**          | 1    | 3    | function(risk, time) |
| **Phased Rollout**    | 3    | 1    | function(risk, time) |
| **Parallel Adoption** | 4    | 1*   | function(risk, time) |

*Sometimes parallel adoption amplifies risk due to sync issues, concurrency, temp integrations.

**Rolling deployment:** Upgrade resources one by one
**Canary deployment:** Deploy to small traffic portion first, evaluate, then expand
**A/B testing:** Set percentage of traffic to new version, measure outcomes

### CI/CD/CD

**Continuous Integration:** Merge to main branch frequently with automated testing
**Continuous Delivery:** Automated to deploy-ready state
**Continuous Deployment:** Every change passes all stages and releases automatically (no human intervention)

**Typical lifecycle:**
1. Get latest from repo
2. Make changes
3. Unit testing ← automated
4. Commit to repo
5. Integration testing ← automated
6. Acceptance testing ← automated
7. Deploy to production ← human
8. Smoke testing ← automated

### AI Developer Services (managed APIs)

| Service                | What                                       | When                        |
|------------------------|--------------------------------------------|-----------------------------|
| **Amazon Comprehend**  | NLP - finds insights/relationships in text | Sentiment analysis          |
| **Amazon Forecast**    | Time-series forecasting with variables     | Seasonal demand prediction  |
| **Amazon Lex**         | Conversational interfaces                  | Customer service chatbots   |
| **Amazon Personalize** | Recommendation engine                      | Upsell products at checkout |
| **Amazon Polly**       | Text-to-speech                             | Dynamic voice responses     |
| **Amazon Rekognition** | Image/video analysis                       | Facial recognition auth     |
| **Amazon Textract**    | Extract text from scanned docs             | Digitise paper forms        |
| **Amazon Transcribe**  | Speech-to-text                             | Auto-transcribe recordings  |
| **Amazon Translate**   | Multi-language text translation            | Localised web content       |

### SageMaker Deployments

**Ground Truth** – Manage labeling jobs with active learning + human labeling
**Notebook** - Managed Jupyter environment
**Training** – Train and tune models
**Inference** – Package and deploy models at scale

**Two deployment modes:**
| Mode | Usage | Method |
|------|-------|--------|
| **Offline** | Batch/asynchronous | Batch Transform |
| **Online** | Real-time/low-latency | Hosting Services |

**Hosting Services workflow:**
1. Create a Model (inference engine)
2. Create Endpoint Configuration (model, instance type, variants, weights)
3. Create Endpoint (publish via InvokeEndpoint API)

**Production variants:** Control traffic distribution (weight / sum of all weights = % of traffic)

**Inference Pipelines:**
- Sequence of 2–5 containers processing data flow
- Built-in or custom Docker algorithms
- Same EC2 instance for speed
- Works for both real-time and batch

**SageMaker Neo:** Compile ML models for various architectures (ARM, Intel, NVIDIA)

**Elastic Inference:** Attach GPU acceleration to CPU instances for real-time inference (cost-effective alternative to full GPU)

**Auto-scaling:**
- Add/remove instances based on workload
- Define: min/max instance count, target metric, cooldown period
- Cooldown prevents rapid scaling fluctuations

## Security

**Visibility** – VPC Endpoints, NACLs, Security Groups
**Authentication** – IAM (who can enter)
**Access Control** – IAM (what they can access)
**Encryption** - KMS (secret decoder ring)

### VPC Endpoints
- Access AWS services over AWS network (not public internet)
- Interface endpoints: DNS entry
- Gateway endpoints: Route table entry (S3, DynamoDB only)

### Notebook Instances
- Internet-enabled by default
- Can disable (need NAT gateway, routes, security groups)
- Designed for single-user (root access)

### IAM Policies

| Type               | Attached to            | Purpose                             |
|--------------------|------------------------|-------------------------------------|
| **Identity-based** | IAM users/roles        | Allow/deny access to entities       |
| **Resource-based** | Resources (S3 buckets) | Allow/deny access at resource level |

**Critical:** SageMaker actions (CreateModel, CreateTrainingJob) require IAM role via `iam:PassRole`

### Encryption

| What           | Generic Example          | AWS Example                         |
|----------------|--------------------------|-------------------------------------|
| **At rest**    | Encrypt stored files     | S3-KMS encryption on bucket config  |
| **In transit** | TLS for HTTP connections | ACM + CloudFront for custom domains |

## Monitor and Evaluate{#monitor-and-evaluate}

**Amazon CloudWatch**
- Endpoint invocation, instance, training, transform, Ground Truth metrics
- Near real-time (1-minute frequency)
- Metrics retained 15 months
- 2-week limit visible in console (collected every 5 mins for training/prediction jobs)
- stdout/stderr → CloudWatch Logs

**Amazon CloudTrail**
- Logs all API calls for SageMaker
- 90 days visible in Event History
- Store indefinitely on S3 with lifecycle processes
- Query with Athena

## Wrapping Up

These notes represent the core material I worked through to pass the AWS Machine Learning Specialty exam. The breadth is significant – from data fundamentals to model deployment to security – but the patterns repeat: understand when to use each tool, know the differences between similar services (especially the Kinesis family), and focus heavily on the evaluation metrics.

If you're coming from infrastructure or software backgrounds like I was, expect to spend extra time on the math-adjacent concepts (precision/recall, bias-variance trade-offs, embedding theory). But none of it is fundamentally inaccessible with focused study.

My biggest surprise: the exam wasn't about memorising formulas, but understanding trade-offs. "Should I use this algorithm or that one?" "What happens if my data is skewed?" "How do I handle missing values?" Those are the questions that matter.

Onward to applying these skills in production!
