AWS Machine Learning Specialty

·18 min read·Graham Mace
AWS Machine Learning Specialty

Earlier this year I passed the AWS Certified Solutions Architect – Professional exam. Then I did something arguably harder: I tackled the AWS Certified Machine Learning – Specialty with almost no prior machine learning experience.

Of the six certifications I’ve earned, this was the one I felt least prepared for. Not because the material is impossible – but because unlike networking or compute services, I couldn’t lean on day-to-day engineering experience to fill the gaps. Everything here came from studying the official AWS Certified Machine Learning – Specialty learning path on Pluralsight, and these notes are what I captured along the way.

If you’re an infrastructure or development person approaching this exam without a data science background, I hope these notes bridge that gap the way I wish something had for me.

(Quick disclaimer: exam blueprints evolve – always verify against AWS’s official certification page.)


Contents

  1. Data Collection
  2. Streaming Data Collection
  3. Data Preparation
  4. Data Analysis and Evaluation
  5. Modelling
  6. SageMaker Modelling
  7. Algorithms
  8. Evaluation and Optimisation
  9. Implementation and Operations
  10. Security
  11. Monitor and Evaluate

Data Collection

What makes data “good”?

Before touching AWS services, the exam cares about whether you can recognise quality data – because garbage in, garbage out.

Good dataBad dataWhy it matters
Large datasetsSmall datasets (<100 rows)More data generally means better model training
Precise, relevant attributesUseless attributesModels need to train on the features that matter
Complete fieldsMissing values / nullsMissing data skews results
Consistent valuesInconsistent valuesModels depend on clean, consistent inputs
Balanced outcome distributionSkewed outcome distributionModels can’t learn when one outcome dominates
Fair samplingBiased samplingBias in, bias out

Core terminology

The exam loves definitions – know these cold:

  • Dataset – your input data; used for training and testing. Columns are called attributes or features; rows are observations, samples, or data points.
  • Structured data - has a defined schema that explains how to interpret it.
  • Unstructured data – no defined schema (raw text, images, audio).
  • Semi-structured data – no rigid schema, but some organisational structure (CSV, JSON, XML).

Exam tip: Be able to distinguish a database (transactional, strict schema), a data warehouse (schema-on-write, BI-ready), and a data lake (schema-on-read, many formats, raw). The write/read distinction is a classic question trap.

Labelled vs unlabelled data

  • Labelled (supervised) – you know the target attribute: spam/not-spam, churn labels, housing prices.
  • Unlabelled (unsupervised) – no target attribute: tweets, logs, news articles. The model must find structure itself.

Two related terms that show up constantly:

  • Ground truth - observed/measured data trusted as factual. (SageMaker Ground Truth the service helps you build labelled datasets through human and automated labelling.)
  • Corpus data - text datasets used for NLP, speech recognition, and text-to-speech.

Streaming Data Collection

The Kinesis family – the exam’s favourite matching game

Know these four cold; questions almost always present a scenario and ask which one fits.

ScenarioServiceWhy
Stream Apache logs from 100 EC2 instances into RedshiftFirehoseLoads to a final destination; S3 first, then copies into Redshift
Stream live sporting-event video to customersVideo StreamsReal-time video/audio processing, feeds other AWS services
Transform streaming data and feed a custom ML appData StreamsHuge ingest volumes, process/transform, feed into apps
Run SQL over real-time data, output metrics to S3Data AnalyticsSQL queries on streams, output to AWS destinations

The simplest way I kept them straight: Data Streams if consumers must act on data in real time (retention matters); Firehose if the goal is just getting data to storage (retention doesn’t matter, processing optional); Video Streams for video/audio; Data Analytics for SQL.

Data Preparation

Once you’ve gathered your data, you need to prepare it before feeding it to any ML algorithm. The AWS exam tests whether you understand how to transform data for different scenarios – not just what the tools are, but when to use them.

Categorical Encoding

ML algorithms generally expect numerical inputs. When you have categorical values (categories, groups, classes), you need to encode them appropriately.

ScenarioAlgorithmEncoding Required?
Predicting house priceLinear RegressionYes
Classifying text as sports/not-sportsNaïve BayesNo
Detecting malignancy in medical imagesConvolutional Neural NetworkYes

Two main encoding strategies:

One-hot encoding – Creates new binary columns (0 or 1) for each category. Best for nominal features where order doesn’t matter.

Example: House type = condo, house, apartment
→ Type_condo (1/0), Type_house (1/0), Type_apartment (1/0)

Mapping – Assigns numbers to categories directly. Only use for ordinal features where order matters (S=5, M=7, L=10).

Warning: Don’t map categories arbitrarily. If house type = condo(1), house(2), apartment(3), the model may incorrectly infer condo < house < apartment.

Exam tip: When there are many categories, consider grouping rare values into “other” or using similarity-based grouping to reduce the number of new columns created.

Text Feature Engineering

Transforming raw text so ML algorithms can analyse it effectively.

TechniqueWhat it doesWhen to use
Bag-of-WordsTokenises text by whitespace into individual wordsSimple frequency analysis
N-GramGroups n consecutive wordsMatching phrases (e.g., “click here now”, “you’re a winner”)
Orthogonal Sparse Bigram (OSB)Creates word pairs including the first word + delimitersFinding common word combinations across documents
TF-IDFWeights words by importance; downweights common terms like “the” and “a”Determining document subject matter, filtering noise
Remove PunctuationStrips punctuation charactersWhen punctuation isn’t semantically important
Lowercase transformationConverts all text to lowercaseStandardising text before processing
Cartesian ProductCombines two categorical/text features into new featureCreating new features from combinations (e.g., textbook + binding)
Date Feature EngineeringExtracts attributes from dates (day, month, weekday, holiday flag)Time-based analysis, business logic patterns

Example of TF-IDF intuition:

  • If “the” appears in both Document 1 and Document 2, it’s less important for distinguishing them.
  • If “Jedi” only appears in Document 1, it carries more weight for that document.

Numeric Feature Engineering

Numeric values often need scaling to work well with ML algorithms.

Scaling techniques:

Normalisation – Rescales values from 0 to 1

Formula: x = (x - min) / (max - min)

Problem: Outliers can throw off normalisation significantly.

Standardisation – Rescales values so mean = 0 (zero-centred)

Formula: z = (x - mean) / std

Advantage: Less affected by outliers than normalisation.

Quick example:

PriceNormalisedStandardised
250,5550.5785590.394923
125,7000.021419-1.080890
120,9000.000000-1.137627
345,0001.0000001.511283

Binning

Grouping numeric values into ranges reduces minor observation errors and makes patterns clearer.

Quantile binning creates bins with equal numbers of observations (not necessarily equal ranges).

Example:

AgeBinYouthYoung AdultAdult
24Youth100
45Young Adult010
18Youth100
76Adult001

Exam tip: There’s no single “optimal” number of bins – it depends on the variable characteristics and relationship to the target. Experimentation is key.

Handling Missing Values

Missing data can appear as null, NaN, NA, None, etc. How you handle it affects model performance.

MechanismDescription
Missing Completely at Random (MCAR)The fact that a value is missing has nothing to do with the data itself
Missing at Random (MAR)Missingness relates to observed data, not the missing data itself
Missing Not at Random (MNAR)Missingness depends on the hypothetical missing value or another variable
TechniqueProsCons
Supervised learningBest results - predicts missing values from other featuresMost difficult to implement
Mean/Median/ModeQuick and easyResults can vary significantly
Dropping rowsEasiestCan dramatically change dataset

Replacing missing values is called imputation.

Exam tip: For Linear Learner specifically, missing values create sparse datasets. Use Factorisation Machines instead if you have lots of missing data points.

Feature Selection

Selecting the most relevant features prevents over-complicated analysis and removes irrelevant or duplicate information.

Principal Component Analysis (PCA) – Unsupervised algorithm that reduces feature count while retaining as much information as possible.

ProblemSolutionWhy
Too many featuresPCAReduces total feature count while preserving information
Useless featuresFeature selectionRemoves features that don’t help solve the problem

AWS Data Preparation Tools

Data SourceRecommended ToolWhy
S3, Redshift, RDS, DynamoDB, on-premise DBAWS GluePython or Scala for flexible transformation
S3AthenaQuery data and output results to S3
EMRPySpark/Hive in EMRHandle petabytes of distributed data
RDS, EMR, DynamoDB, RedshiftData PipelineSetup EC2 instances for ETL

AWS Glue vs SageMaker:

  • SageMaker + Jupyter notebooks → ad hoc analysis and experimentation
  • AWS Glue jobs → reusable, production ETL pipelines

Data Analysis and Evaluation

Before training a model, you need to understand your data visually.

QuickSight

AWS’s BI tool for creating visualisations directly from your data in the console.

Choosing the right visualisation

GoalVisualisation
ComparisonsBar chart (lookup values), Line chart (change over time)
DistributionsHistogram (single distribution), Box plots / Scatter plots (multiple distributions)
CompositionsPie chart (static), Stacked bar / area chart (changing over time)
RelationshipsScatter plot (2 variables), Bubble chart (3 variables)
Chart typeBest for
Scatter plotRelationships between two values
Bubble plotRelationships between three values (bubble size = third dimension)
Bar chartLookup and compare single variable values
Line chartVariables changing over time
HistogramDistribution across binned values
Box plotLowest/highest values, outliers, quartiles
Pie chartParts of a whole (static composition)
Stacked area columnQuantity over shorter time periods

Modelling

Understanding model types

Learning TypeTraining InputDiscrete OutputContinuous Output
SupervisedTraining + Testing dataClassificationRegression
UnsupervisedNo training labelsClusteringDimensionality reduction
ReinforcementTrial and error with rewardsSimulation-based optimisationAutonomous devices

Choosing the right approach

ProblemApproachWhy
Detect financial fraudBinary ClassificationOnly two outcomes: fraud or not
Predict car deceleration rateHeuristic (no ML)Physics formulas already exist
Find optimal path for lunar roverReinforcement LearningTrial, error, improvement needed
Identify dog breed in photoMulti-class ClassificationMany breed categories to choose from

The Confusion Matrix

This is genuinely one of the most tested concepts. Memorise this:

Predicted TRUEPredicted FALSE
Actual TRUECorrect prediction:cross: False Positive (Type I Error)
Actual FALSE:cross: False Negative (Type II Error)Correct prediction

Fraud detection analogy:

  • False Positive: Bank flags legitimate transaction → Angry customer, happy bank (false alarm)
  • False Negative: Bank misses actual fraud → Happy customer, angry bank (missed loss)

Exam tip: Different thresholds affect precision/recall trade-offs. Spam filters want few false positives (blocking legitimate email), while fraud detection may prioritise catching all fraud even with more false alarms.

Data Preparation for Training

Remember: We want generalisation, not memorisation.

Split strategy:

  1. Randomise the dataset
  2. Split into training and testing sets
  3. Train on training set
  4. Test on test set

K-Fold Cross-Validation:

  • Takes random splits multiple times
  • Each fold becomes test data once
  • Better estimate of model performance than a single split

SageMaker Modelling

Creating models

OptionUse case
SageMaker Console → JupyterInteractive development
SageMaker SDK (Apache Spark)Production pipelines

Supported data formats

Content-TypeAccept
application/x-recordio-protobufapplication/x-recordio-protobuf
text/csvapplication/json
application/jsonlinesapplication/jsonlines
application/x-imageimage/*

Performance note: Protobuf recordIO format gives optimal performance. Pipe mode allows streaming data from S3 to training instances – uses less EBS space and starts faster.

Training API basics

CreateTrainingJob API uses SageMaker SDK for Python:

  • Specify training algorithm
  • Supply algorithm-specific hyperparameters
  • Specify input and output configuration

Hyperparameter vs Parameter:

  • Hyperparameter: Set before learning (learning rate, number of epochs)
  • Parameter: Derived during learning (weights, biases)

SageMaker Training

GPU vs CPU performance hierarchy:

CPU < FPGA < GPU < ASIC (for performance)
ASIC < FPGA < GPU < CPU (for flexibility)

Common training parameters:

  • Channel name – Named input source for algorithm
  • Training/inference image registry path - Versioned ECR images (:1 for stable, :latest for bleeding edge)
  • Training input mode – File or Pipe (Pipe performs better)
  • File format – recordIO protobuf for best performance
  • Instance class - GPU/CPU depending on algorithm requirements

CloudWatch integration

CloudWatch logs:

  • Arguments provided
  • Errors during training
  • Algorithm accuracy statistics
  • Timing information

Common errors:

  • Extra or invalid hyperparameters
  • Incorrect protobuf file format

Algorithms

Supervised vs Unsupervised vs Reinforcement Learning

CategoryTraining DataOutput
SupervisedLabels requiredClassification (discrete), Regression (continuous)
UnsupervisedNo labelsClustering, Dimensionality reduction
ReinforcementRewards/punishmentsPolicy that maximises reward

Algorithms can be:

  • SageMaker built-in
  • Purchased from AWS Marketplace
  • Built custom via Docker

Regression

Linear Learner

  • Supervised learning for regression, binary classification, or multi-class classification
  • Maps vector x to approximate label y
  • Minimises sum of distances from training points to prediction line (using Stochastic Gradient Descent)

Use cases:

  • Predicting quantitative values (ROI based on marketing spend)
  • Binary classification (should I mail this customer?)
  • Multi-class classification (email, phone, or mail contact?)

Factorisation Machines

  • Handles high-dimensional sparse data with missing values (“holes”)
  • Considers only pair-wise feature relationships
  • Works on binary classification or regression only (not multi-class)
  • Requires 10,000–10,000,000+ dimensions
  • AWS recommends CPUs (better for sparse data)

Exam tip: Factorisation Machines need recordIO-protobuf format – CSV is not supported.

Clustering

K-Means (Unsupervised)

  • Groups observations based on attribute similarity (Euclidean distance)
  • Members of a group are as similar as possible to each other, as different as possible from other groups

SageMaker K-Means specifics:

  • Modified web-scale version (more accurate than standard)
  • CPU instances recommended (GPU can only use one core anyway)
  • Must define number of clusters beforehand
  • Training still occurs – model accuracy matters even without labels

Exam tip: K-Means is great for finding hidden patterns in tabular data when you don’t have ground truth labels.

Classification

K-Nearest Neighbours (KNN)

  • Index-based, non-parametric method
  • Finds k closest points to sample and returns most frequent label (classification) or average value (regression)
  • Lazy algorithm – doesn’t learn during training, just stores data in memory

Use cases:

  • Credit ratings (group by shared risk attributes)
  • Product recommendations (similar items based on likeability)

Image Analysis

AlgorithmTaskNotes
Image ClassificationHotdog/not hotdogUses Convolutional Neural Networks (ResNet)
Object DetectionClock + lamp in sceneAssigns classification + confidence
Semantic SegmentationPixel-level edge detectionAccepts PNG, requires GPU for training, CPU/GPU OK for inference

Exam tip: Semantic segmentation is computationally intensive - only GPU instances for training.

Anomaly Detection

Random Cut Forest (RCF)

  • Unsupervised anomaly detection
  • Gives anomaly score (low = normal, high = outlier)
  • Works well with n-dimensional data
  • Doesn’t benefit from GPU (use regular instances: ml.m4, ml.c5)

Use cases:

  • Quality control (unusual audio frequencies)
  • Fraud detection (unusual amounts, times, locations)

IP Insights

  • Learns usage patterns for IP addresses paired with entities/user IDs
  • Returns anomaly scores for entity/IP combinations
  • Uses neural networks for latent vector representations
  • GPUs recommended for training (distributed CPUs may be more cost-effective for large datasets)

Use cases:

  • Tiered authentication (trigger 2FA for anomalous IPs)
  • Fraud detection (allow/restrict activities based on login patterns)

Text Analytics

Latent Dirichlet Allocation (LDA) (Unsupervised)

  • Discovers topics within document corpora
  • Documents are observations, words are features, topics are categories

Neural Topic Model (NTM) (Unsupervised)

  • Similar to LDA but uses different algorithm → may produce different results

Sequence-to-Sequence (seq2seq) (Supervised)

  • Input sequence → output sequence (translation, speech-to-text)
  • Uses embedding, encoding, decoding layers (RNN/CNN)
  • Pre-trained word vectors common (FastText, GloVe)
  • GPU-only, single-machine training only

BlazingText

  • Optimised Word2Vec and text classification
  • ~20x faster than FastText
  • Modes: Single CPU, Single GPU, Multiple CPU
  • Word2Vec = unsupervised, Text Classification = supervised

Object2Vec (Supervised)

  • Learns embeddings of high-dimensional objects
  • Requires pairs of items with relationship labels
  • Embeddings usable for downstream tasks

Reinforcement Learning

“Carrot and stick” approach - positive/negative rewards to optimise agent policy.

Markov Decision Process (MDP):

  1. Agent
  2. Environment
  3. Reward
  4. State
  5. Action
  6. Observation
  7. Episodes
  8. Policy

Use cases:

  • Autonomous vehicles (stay on road through trial/error)
  • Intelligent HVAC control (learn building patterns, optimise energy)

Forecasting

DeepAR (Supervised)

  • Scalar time series forecasting using RNN
  • Outperforms ARIMA and ETS by training single model over multiple time series
  • Supports point forecasts (“X sneakers sold”) and probabilistic forecasts (“X-Y with Z% probability”)

Requirements:

  • Minimum 300 observations across all time series
  • Hyperparameters: context length, epochs, prediction length, time frequency
  • Automatic backtest evaluation after training

Use cases:

  • New product performance forecasting
  • Labor needs prediction for special events

Ensemble Learning

XGBoost (Supervised)

  • Gradient boosted trees implementation
  • Virtual “Swiss army knife” for regression, classification, ranking
  • 2 required + 35 optional hyperparameters

Technical constraints:

  • Accepts CSV and libsvm formats
  • CPU-only training, memory-bound
  • Needs lots of memory for entire training data
  • Spark integration via SageMaker Spark SDK

Use cases:

  • E-commerce search ranking (relevance scores)
  • Fraud detection (transaction → fraud probability)

Evaluation and Optimisation

Feedback loop

  1. Define evaluation – What metrics determine success?
  2. Evaluate – Review during/after training (manual or automatic)
  3. Tune – Adjust hyperparameters, data, or algorithm

Validation types

TypeDescriptionExample
OfflineUsing test data setsValidation sets, K-Fold
OnlineReal-world conditionsCanary deployments, A/B testing

Monitoring training jobs

SageMaker → CloudWatch

  • Algorithm metrics: train:* and validation:*
  • Training job logs

Model accuracy

UnderfittingOverfitting
DefinitionModel doesn’t reflect data shapeModel memorises training data
SymptomsPoor on training AND test dataGreat on training, poor on test
FixMore data, train longerMore data, early stopping, noise, regularisation, ensembles, feature selection
Training ErrorTesting ErrorDiagnosis
LowLowIdeal!
LowHighOverfitting
HighHighTry different approach
HighLowRun away! (data leak)

Accuracy metrics

Regression:

  • RMSE (Root Mean Square Error) – lower is better

Binary classification:

  • Recall = Right guesses / (Right + missed ones)
  • Precision = Right guesses / (Right + wrong flags)
  • F1 Score = Harmonic mean of precision and recall

Exam tip: Different thresholds change F1 scores. Spam filters want minimal false positives; fraud detection may tolerate more to catch all fraud.

Multiclass classification:

  • Per-class accuracy + macro/micro averaging
  • F1 scores calculated per class

Model tuning (hyperparameter tuning)

  1. Choose tunable hyperparameter (not all can be auto-tuned)
  2. Choose range of values
  3. Choose objective metric to optimise

Optimisation methods:

  • Random search – samples randomly across range
  • Bayesian optimisation – focuses on promising regions based on previous results (avoids costly iterations)

Implementation and Operations

Deployment types

TypeTimeRiskCost
Big Bang13function(risk, time)
Phased Rollout31function(risk, time)
Parallel Adoption41*function(risk, time)

*Sometimes parallel adoption amplifies risk due to sync issues, concurrency, temp integrations.

Rolling deployment: Upgrade resources one by one Canary deployment: Deploy to small traffic portion first, evaluate, then expand A/B testing: Set percentage of traffic to new version, measure outcomes

CI/CD/CD

Continuous Integration: Merge to main branch frequently with automated testing Continuous Delivery: Automated to deploy-ready state Continuous Deployment: Every change passes all stages and releases automatically (no human intervention)

Typical lifecycle:

  1. Get latest from repo
  2. Make changes
  3. Unit testing ← automated
  4. Commit to repo
  5. Integration testing ← automated
  6. Acceptance testing ← automated
  7. Deploy to production ← human
  8. Smoke testing ← automated

AI Developer Services (managed APIs)

ServiceWhatWhen
Amazon ComprehendNLP - finds insights/relationships in textSentiment analysis
Amazon ForecastTime-series forecasting with variablesSeasonal demand prediction
Amazon LexConversational interfacesCustomer service chatbots
Amazon PersonalizeRecommendation engineUpsell products at checkout
Amazon PollyText-to-speechDynamic voice responses
Amazon RekognitionImage/video analysisFacial recognition auth
Amazon TextractExtract text from scanned docsDigitise paper forms
Amazon TranscribeSpeech-to-textAuto-transcribe recordings
Amazon TranslateMulti-language text translationLocalised web content

SageMaker Deployments

Ground Truth – Manage labeling jobs with active learning + human labeling Notebook - Managed Jupyter environment Training – Train and tune models Inference – Package and deploy models at scale

Two deployment modes:

ModeUsageMethod
OfflineBatch/asynchronousBatch Transform
OnlineReal-time/low-latencyHosting Services

Hosting Services workflow:

  1. Create a Model (inference engine)
  2. Create Endpoint Configuration (model, instance type, variants, weights)
  3. Create Endpoint (publish via InvokeEndpoint API)

Production variants: Control traffic distribution (weight / sum of all weights = % of traffic)

Inference Pipelines:

  • Sequence of 2–5 containers processing data flow
  • Built-in or custom Docker algorithms
  • Same EC2 instance for speed
  • Works for both real-time and batch

SageMaker Neo: Compile ML models for various architectures (ARM, Intel, NVIDIA)

Elastic Inference: Attach GPU acceleration to CPU instances for real-time inference (cost-effective alternative to full GPU)

Auto-scaling:

  • Add/remove instances based on workload
  • Define: min/max instance count, target metric, cooldown period
  • Cooldown prevents rapid scaling fluctuations

Security

Visibility – VPC Endpoints, NACLs, Security Groups Authentication – IAM (who can enter) Access Control – IAM (what they can access) Encryption - KMS (secret decoder ring)

VPC Endpoints

  • Access AWS services over AWS network (not public internet)
  • Interface endpoints: DNS entry
  • Gateway endpoints: Route table entry (S3, DynamoDB only)

Notebook Instances

  • Internet-enabled by default
  • Can disable (need NAT gateway, routes, security groups)
  • Designed for single-user (root access)

IAM Policies

TypeAttached toPurpose
Identity-basedIAM users/rolesAllow/deny access to entities
Resource-basedResources (S3 buckets)Allow/deny access at resource level

Critical: SageMaker actions (CreateModel, CreateTrainingJob) require IAM role via iam:PassRole

Encryption

WhatGeneric ExampleAWS Example
At restEncrypt stored filesS3-KMS encryption on bucket config
In transitTLS for HTTP connectionsACM + CloudFront for custom domains

Monitor and Evaluate

Amazon CloudWatch

  • Endpoint invocation, instance, training, transform, Ground Truth metrics
  • Near real-time (1-minute frequency)
  • Metrics retained 15 months
  • 2-week limit visible in console (collected every 5 mins for training/prediction jobs)
  • stdout/stderr → CloudWatch Logs

Amazon CloudTrail

  • Logs all API calls for SageMaker
  • 90 days visible in Event History
  • Store indefinitely on S3 with lifecycle processes
  • Query with Athena

Wrapping Up

These notes represent the core material I worked through to pass the AWS Machine Learning Specialty exam. The breadth is significant – from data fundamentals to model deployment to security – but the patterns repeat: understand when to use each tool, know the differences between similar services (especially the Kinesis family), and focus heavily on the evaluation metrics.

If you’re coming from infrastructure or software backgrounds like I was, expect to spend extra time on the math-adjacent concepts (precision/recall, bias-variance trade-offs, embedding theory). But none of it is fundamentally inaccessible with focused study.

My biggest surprise: the exam wasn’t about memorising formulas, but understanding trade-offs. “Should I use this algorithm or that one?” “What happens if my data is skewed?” “How do I handle missing values?” Those are the questions that matter.

Onward to applying these skills in production!