03Professional services

Architecture you can rebuild from the repository.

Reviews, migrations, and landing zones on AWS, Azure, and GCP. The deliverable is not a diagram in a slide deck. It is the Terraform that produces the diagram, in your repository, with a plan you can run on the day we leave.

0
Availability zones in every production design
0%
Infrastructure handed over as code
0
SSH paths open to the public internet
0 to 40%
Typical spend recovered on a first cost review
The problem

The cloud bill is usually the second problem.

Most estates we are called into were built quickly, correctly for the time, and then never revisited. One account holds everything. The database is a virtual machine somebody installed PostgreSQL on. There is a security group with 0.0.0.0/0 in it that predates the current team. The whole thing runs in a single availability zone, which is fine right up until it is not.

The spend is what gets noticed, because it arrives monthly with a number on it. But oversized instances are rarely the expensive part. The expensive part is the afternoon nobody can deploy because the one person who knows how the network was wired is on holiday, and the state of the environment exists only in the console.

So the first deliverable is a review that separates the two: what is costing money, and what is costing you the ability to change anything. They usually have the same fix, which is putting the environment into code so that a change is a pull request rather than an act of memory.

Target state

A shape that survives losing a zone.

Drawn at the level a real review happens: where the state lives, where the failure domains are, and which hops are allowed.

Target architectureprivate subnetsUserspublic internetCDN + WAFedge, TLS 1.3Load balancerhealth-checkedApp tier, AZ-aautoscaledApp tier, AZ-bautoscaledPrimary DBencryptedStandbycross-AZ
What the diagram commits to
Failure domain
One AZ can go dark without a page
State
Managed, replicated, restore-tested
Access
No SSH path from the internet
Cost
Right-sized against real utilisation
Delivered as code

The diagram is the readable version. What you actually receive is the Terraform that produces it, in your repository, with a plan you can run yourself.

Scope

What we take on

01

Landing zone and account structure

Separate accounts or subscriptions, with guardrails that apply from the top.

Production, staging, and shared services separated at the account boundary rather than by naming convention, with organisation-level policy, centralised logging, and a billing structure that tells you which team spent what. Retrofitting this later is possible, but it is always more work than doing it first.

AWS OrganizationsAzure landing zoneSCPGCP folders
02

Network and connectivity

Private subnets, controlled egress, and a path in that is not SSH from anywhere.

Application and data tiers in private subnets, egress through a gateway you can log, and administrative access through a session service or a bastion with recorded sessions. Hybrid connectivity, peering, and DNS resolution across accounts included where you have on-premise estate to reach.

VPCTransit GatewayPrivateLinkSSM
03

Migration

A move with a rehearsal, a cutover window, and a rollback that has been tested.

Lift and shift where that is the honest answer, re-platforming where it pays for itself. Databases move to managed services with replication and a rehearsed switch, so the cutover is measured in minutes and the fallback is real rather than theoretical.

DMSreplicationcutover planrollback
04

Reliability and disaster recovery

Multi-zone by default, and a recovery objective you have actually measured.

Autoscaling and health checks that remove a bad instance rather than paging a person, backups replicated across regions, and a restore performed against a stopwatch so your RTO is a number from a test rather than a number from a policy document.

multi-AZautoscalingbackupRTO / RPO
05

Identity, secrets, and audit

Roles instead of long-lived keys, and an audit trail that is written elsewhere.

Workload identity in place of access keys checked into a repository, short-lived credentials for humans through SSO, secrets in a managed store with rotation, and audit logs shipped to an account the workload cannot write to. That last part is what makes the logs worth having after an incident.

IAMOIDCSecrets ManagerVaultCloudTrail
06

Cost engineering

Right-sizing against real utilisation, then commitments once the shape is stable.

Utilisation measured before anything is resized, storage tiered, idle environments scheduled off, and only then commitment discounts applied, because buying a three year reservation for an instance type you are about to stop using is a common and expensive mistake.

right-sizingsavings planstaggingbudgets
How it runs

Review, design, move, hand over.

Weeks 1 to 2

Review

Read-only access, an inventory of what exists, and a written report of the risks, the single points of failure, and the spend.

Weeks 3 to 4

Design

Target architecture agreed with your team, written as a document and then as Terraform modules in your repository.

Project

Migrate

Built in staging, rehearsed, then cut over in an agreed window with a rollback that was tested rather than described.

Handover

Operate or leave

Runbooks, a walkthrough with your engineers, and the choice of holding it yourselves or keeping us on retainer.

Tools we work in
AWSAzureGoogle CloudTerraformOpenTofuCloudFormationKubernetesEKSAKSGKERDSCloud SQLS3CloudFrontRoute 53VaultGitHub ActionsArgo CDDatadogCloudWatch

Get the review before the renewal.

Two weeks of read-only access produces a written picture of the risks and the spend, and it is yours whether or not we do the work.