AWS Landing Zone
A multi-account AWS organization built as code to parity with the Azure and GCP zones: inherited SCP guardrails, an immutable org audit trail, one inspected path to the internet, a central monitoring account, and a private EKS plus Multi-AZ PostgreSQL paved road. Deployed live, verified, and the hourly layers destroyed with the foundation kept.
/01Problem
A landing zone has to settle two things before any workload arrives: what every account inherits, and how anything leaves the network. This one settles them with a persistent multi-account organization whose guardrails are inherited from the OUs, audit and security controls centralized outside the accounts they watch, and an inspected private network that is the only way a workload reaches the internet.
The reference workload on top is private EKS and Multi-AZ PostgreSQL. Its job is to prove the zone can host a real workload tier, not to run an application. Application integration is outside the completion scope, and the write-up keeps that line visible.
/02Five roots, split by how long they live
- accounts is permanent: the organization, OUs, eight member accounts, SCPs, the tag policy, and RAM sharing. Every account carries close_on_deletion false and prevent_destroy, because a closed account sits SUSPENDED for 90 days holding org quota and its email alias, and the next deploy collides with it. Idle accounts cost nothing.
- governance and observability are nearly free and stay up: the audit trail, detective services, identity, budgets, and the monitoring plane.
- network and workload are the hourly layers. They apply and destroy on their own, workload first and network second, so the expensive tiers exist only for a demo.
- Moving the org into its own root was done live with import and removed blocks: 28 resources adopted, zero added, zero destroyed, and both roots re-planned to no changes.
/03Guardrails a workload cannot turn off
- SCPs on the OUs deny root use, leaving the organization, unapproved regions, public or unencrypted S3, and disabling GuardDuty, Config, or CloudTrail. The management account is SCP-exempt by design, so its root hardening is a separate step: a strict password policy and an EventBridge alarm on any root sign-in.
- An organization CloudTrail writes to a KMS-encrypted, GOVERNANCE-mode Object Lock bucket in log-archive, with log-file validation, so the record of what happened is safe from the account that did it.
- GuardDuty, Security Hub with CIS AWS Foundations 1.4.0, and AWS Config are delegated to the security account, so security operations run outside the account that can change the org.
- IAM Identity Center personas (admin, platform engineer, junior engineer, manager, FinOps, security, break-glass) carry permission boundaries and short sessions. There are no IAM users and no standing prod write.
/04One inspected path out, one place to look
The network account shares a Transit Gateway to the org over RAM with explicit attachment acceptance. Separate spoke and inspection route tables force egress and its return through AWS Network Firewall with a default-deny domain allowlist, then NAT, and the allowlist defaults to AWS and Cognito domains. The demo inspection path runs in one availability zone, which is a cost choice and not a production availability design. Admin access is SSM over interface endpoints rather than a bastion.
Shared-services is a CloudWatch OAM monitoring account with prod and network linked into it, so alarms live where no workload team can change them while the data stays in the account that produced it. GuardDuty severity 7 and above and Security Hub HIGH and CRITICAL findings route to one topic, and five cross-account alarms (RDS CPU and storage, EKS failed nodes, firewall drops, failed backups) route to another.
/05The paved road on top
- A private prod VPC with no internet gateway or NAT of its own. Its only way out is the Transit Gateway to the hub firewall.
- EKS with a private API endpoint, KMS-encrypted secrets, IRSA, and two AL2023 managed nodes, plus the account's own interface endpoints for ECR, STS, EKS, Logs, SSM, and Secrets Manager.
- PostgreSQL on RDS, Multi-AZ, encrypted, private, with an RDS-managed secret so no password lands in state. Only the app tier security group reaches 5432, backed by a data-subnet NACL.
- Daily AWS Backup into a Vault Lock vault, with cross-region copy held off until a destination region is approved. An ECR supply chain with immutable tags, scan on push, a pull-through cache, and Inspector, plus an on-demand Image Builder pipeline; the demo nodes ran the standard AL2023 EKS image.
/06What the live apply found
- CreateTrail failed with InsufficientEncryptionPolicyException. kms:DescribeKey sat under a kms:EncryptionContext condition, and DescribeKey carries no context, so it moved into its own statement.
- Root-activity alerts were not arriving. EventBridge and CloudWatch cannot publish to a topic under the AWS-managed aws/sns key, so every alert topic now has its own customer-managed key granting only the service that publishes to it.
- The tag policy was rejected as malformed because rds:db does not support enforcement. CostCenter is enforced on the types that do and still evaluated for compliance on RDS.
- New OAM links return Forbidden for about 3.7 minutes, so the alarms wait on them. On teardown, disabling Inspector timed out after 84 of 85 workload resources were gone, and a reviewed Inspector-only recovery removed the last one.
/07Verified, then torn down
Before teardown, live checks confirmed Network Firewall READY and IN_SYNC, both RAM associations ASSOCIATED, prod seeing the shared Transit Gateway with egress routed through inspection, EKS 1.35 with both nodes ACTIVE, and PostgreSQL 16.14 private, encrypted, and Multi-AZ. The observability suite raised a GuardDuty sample finding and confirmed its publish, saw 15 prod and 9 network metrics from the monitoring account, and forced an alarm to confirm its action fired.
Those checks prove infrastructure state and routing, not application traffic, and the write-up says so: a forced RDS failover and an end-to-end firewall traffic test were not run. Workload and network were then destroyed, and live API checks confirmed zero EKS, RDS, endpoints, NAT gateways, Transit Gateways, and firewalls, with every account still ACTIVE, the trail logging, and the monitoring plane in place.
/08Sandbox-first compute baseline queued
The next phase will plan IMDSv2 and EBS encryption SCPs plus an EC2 declarative policy for the Sandbox OU. The operator applies governance before a reduced workload network or compute root can be deployed. Public repository links use aws-landing-zone; existing state keys retain their internal names.
The proposed Image Builder bake stages the shared Ansible role in S3 for the private build subnet. One golden-AMI management instance will be reached only through SSM. Allowed-image denials and live guest-hardening proof are pending; EKS will retain its managed AL2023 node image.