Cloud Design Patterns

Touring the Azure Cloud Design Patterns Library
29 Patterns in 14 Lunchtimes
A few years ago, I ran a Design Patterns Community of Practice at work. Over fourteen sessions, we worked our way through nearly the entire Azure Cloud Design Patterns catalog – one of the best-curated collections of distributed systems wisdom available for free.
This post is the full tour: every pattern we covered, what it’s for, and when to reach for it.
Why bother with a patterns library at all? As we kept reminding ourselves in the CoP: patterns form a common language between developers, they’re proven “in the wild”, and they stop you reinventing a wheel that someone else has already stress-tested.
Don’t reinvent the wheel.
I’ve grouped the patterns roughly the way the Azure catalog does. Caveat up front: none of these are free. Almost every one trades something – consistency, latency, resource efficiency, or simplicity – for something else. The “when not to use” sections in the docs are arguably more valuable than the “when to use” ones.
Data Management
Sharding – Split a data store into horizontal partitions (shards) so you can scale past a single server’s storage, compute, and network limits. The three key-selection strategies each have trade-offs: Range (great for range queries, terrible for hotspots), Hash (even distribution, painful rebalancing), and Lookup (maximum control and easy rebalancing, but extra routing overhead). Avoid volatile or auto-incrementing shard keys.
Index Table – Build secondary indexes over fields you query frequently but that aren’t your primary/shard key. Essential in NoSQL stores that lack native secondary indexes. Decide between fully denormalised, normalised (two lookups), or partially normalised variants based on read/write balance.
Materialized View - Pre-compute query-ready views over data stored in a shape that’s awkward to query. The view is disposable – it can always be rebuilt from source. Pair with Event Sourcing, where it’s often the only way to get data out.
Static Content Hosting – Serve static assets straight from cheap blob storage (optionally via CDN) instead of burning compute cycles. Watch out for deployment/versioning when content and app live in different places.
Caching
Cache-Aside - Load data into the cache on demand when your cache doesn’t do read-through/write-through natively. As Phil Karlton said, cache invalidation is one of the two hard problems; this pattern embraces eventual consistency rather than promising coherence. Pay attention to expiration tuning - too short hammers the data store, too long serves stale data.
Design & Implementation
Anti-Corruption Layer – Put a translating façade between your clean new system and a legacy or external system whose models, protocols, or assumptions you don’t want leaking into your design.
Strangler Fig – Incrementally replace a legacy system by routing requests through a façade that gradually shifts traffic from old to new. Martin Fowler’s framing is sobering: “all we are doing is writing tomorrow’s legacy software today” – so design today’s work to be easily strangled later.
Pipes & Filters – Decompose monolithic processing pipelines into composable single-purpose stages with standardised interfaces. Slow filters can be scaled in parallel; stages run on hardware sized to their needs. Filters should be idempotent and the pipeline should deduplicate messages. Its cousins include Chain of Responsibility, Specification, Rules – and, if you squint, neural networks.
Backends for Frontends – Give each frontend (mobile, web, TV…) its own dedicated backend rather than contorting a general-purpose API to serve them all. Costs you code duplication; buys each interface team autonomy.
Messaging
Publisher-Subscriber - Broadcast events to interested consumers via a broker, decoupling senders from receivers entirely. Use topics or content filtering so subscribers only get what they care about.
Priority Queue - Message ordering does sometimes matter. Route higher-priority messages to be processed first, via prioritised queues or queue-per-priority with separate consumer pools.
Queue-Based Load Levelling – Put a queue between a task and a service to buffer bursts, letting the service process at its own pace. Deploy for average load instead of peak load. Calvin Coolidge: “If you see ten troubles coming down the road, you can be sure that nine will run into the ditch before they reach you.”
Competing Consumers – Multiple consumers pull from one queue to process messages concurrently, giving you throughput, scalability, and load balancing. At-least-once delivery means processing must be idempotent, and you need a story for poison messages.
Resiliency
Retry – Handle transient faults by transparently retrying failed operations: immediately for rare glitches, after a delay (often exponential backoff) for busy services. Don’t retry non-transient failures (bad credentials stay bad), and know when to escalate to-
Circuit Breaker - Stop hammering a failing service. Once failures persist, fail fast for a cooldown period before tentatively letting requests through again.
Compensating Transaction – Undo the work of a failed long-running distributed transaction by applying explicit reverse operations per step, rather than relying on distributed transactions nobody wants. Steps must be idempotent in case the compensation itself fails partway. Avoid needing it in the first place if you can – complexity is the price.
Scheduler Agent Supervisor - Orchestrate a distributed multi-step task: the Scheduler runs the workflow and tracks step state durably, Agents wrap remote calls with retry/timeouts, and the Supervisor periodically detects stuck or failed steps and triggers recovery.
Leader Election – Coordinate peer instances by electing a leader (via distributed mutex, lowest ID, or algorithms like Bully/Ring). Beware: a shared mutex service is a single point of failure, and autoscaling can terminate your leader.
Bulkhead – Partition resources (connection pools, instances, containers) so that one failing service or noisy client can’t exhaust resources needed by everything else. Named for ship compartments. Combine with retry, circuit breaker, and throttling for layered defence.
Health Endpoint Monitoring – Expose an endpoint external monitors can poll to verify the system is truly healthy, not merely up. One “OK” tells you very little; check dependencies too – and protect the endpoint itself.
Security
Gatekeeper - Place a dedicated, low-privilege validation front-end between clients and your trusted hosts. If the gatekeeper is compromised, the attacker gets nothing of value. It’s a firewall for your architecture.
Valet Key – Hand clients time-limited, narrowly-scoped tokens so they can read/write a specific storage resource directly - offloading data transfer from your app. Think AWS S3 pre-signed URLs. Keep the validity window and scope as tight as possible.
Federated Identity – Delegate authentication to an external identity provider to get single sign-on, fewer credential sprawls, and less user admin. Claims-based access control decouples who you are (IdP’s job) from what you can do (your job).
Operational Concerns
Throttling – Cap resource consumption per tenant/user/service so a burst of demand degrades gracefully rather than taking everything down. Return explicit error codes so clients know to back off. Pairs beautifully with autoscaling: throttle now, scale out meanwhile.
External Configuration Store – Centralise configuration outside deployment packages for easier management, sharing across instances, and runtime updates. Manage changes to centrally-stored config with the same rigour as code deployments – a bad flag flip is an outage.
Sidecar - Co-locate cross-cutting concerns (logging, networking, monitoring) in a separate process/container beside your app, independent of language and lifecycle. Familiar territory if you know service meshes.
Ambassador – An out-of-process proxy co-located with the client, handling connectivity concerns (TLS, retry, routing, metering) in a language-agnostic way. Essentially the outbound twin of the Sidecar.
Compute Resource Consolidation – Pack multiple tasks into one compute unit to raise utilisation and cut cost – but group by compatible scalability profiles, lifetimes, and resource usage, and accept the security/fault-tolerance trade-off.
The Gateway Family
Four patterns, one theme: put a gateway in the middle.
Gateway Routing – One public endpoint, Layer 7 routing to many backend services. Clients stay stable while you add, split, or reorganise services behind it.
Gateway Aggregation – Bundle multiple backend calls into a single client request, cutting chattiness – critical on high-latency mobile networks.
Gateway Offloading – Concentrate cross-cutting concerns (SSL termination, auth, logging, throttling) in the gateway so individual services don’t each reimplement them.
Gatekeeper (see Security) – the security-flavoured sibling of the family.
The recurring caveat across all four: your gateway can become a single point of failure, a bottleneck, or an accidental dumping ground for business logic it should never contain.
Closing Thoughts
Twenty-nine patterns later, a few lessons stood out:
- Idempotency is everywhere. Competing Consumers, Retry, Pipes & Filters, Compensating Transactions, Pub/Sub - nearly every messaging pattern leans on it. If you internalise one concept from this post, make it that one.
- Eventual consistency is the cloud-native default. A surprising number of patterns exist specifically because strong consistency across distributed systems is a cost too high to pay.
- The “when not to use” matters more than the “when to use.” Every pattern here adds complexity; the skill is knowing when the problem justifies it.
