AWS Provisioning
Optional AWS provisioning for the production infrastructure — creates the EC2 instances (Reverse Proxy, Compute, Storage, and the Backup node), security groups, Elastic IP, and writes provision-output
The bundled AWS provisioning is a separate, optional step that creates the EC2 instances — Reverse Proxy, Compute, Storage, and (with backup_node.enabled: true) the Backup node — and the supporting AWS resources, then writes provision-output.yaml for the orchestrator to consume. Lives at automation/production/aws/.
Use it if you don't already have VMs. If you have your own VMs (other clouds, on-prem, manual EC2), skip this page and go straight to Step 1 of the infrastructure automation.
Prerequisites
AWS CLI
v2 installed on your laptop. aws --version should print aws-cli/2.x.
AWS credentials
Configured via aws configure, environment variables, or an AWS_PROFILE. The script honours AWS_REGION, AWS_PROFILE, and AWS_DEFAULT_REGION.
jq
Not required (we deliberately avoid the dependency).
Permissions
The IAM user/role needs the EC2 permissions listed below.
EIP quota
At least one Elastic IP free in the target region. AWS's default per-region quota is 5 EIPs. The provisioner allocates one EIP for the RP (Wireguard endpoint stability across stop/start — see About the Elastic IP). If you're at quota, free one first (see About the Elastic IP) before running the provisioner.
IAM permissions
The provisioning script needs a moderately broad set of EC2 permissions. The minimal set:
{
"Version": "2012-10-17",
"Statement": [{
"Effect": "Allow",
"Action": [
"ec2:DescribeVpcs",
"ec2:DescribeSubnets",
"ec2:DescribeImages",
"ec2:DescribeInstances",
"ec2:DescribeInstanceStatus",
"ec2:DescribeKeyPairs",
"ec2:DescribeSecurityGroups",
"ec2:DescribeAddresses",
"ec2:DescribeNetworkInterfaces",
"ec2:CreateKeyPair",
"ec2:DeleteKeyPair",
"ec2:CreateSecurityGroup",
"ec2:DeleteSecurityGroup",
"ec2:AuthorizeSecurityGroupIngress",
"ec2:RevokeSecurityGroupIngress",
"ec2:AllocateAddress",
"ec2:ReleaseAddress",
"ec2:AssociateAddress",
"ec2:DisassociateAddress",
"ec2:RunInstances",
"ec2:TerminateInstances",
"ec2:ModifyInstanceAttribute",
"ec2:CreateTags",
"sts:GetCallerIdentity",
"tag:GetResources"
],
"Resource": "*"
}]
}If you have full EC2 admin (AmazonEC2FullAccess managed policy + sts:GetCallerIdentity + tag:GetResources), that works too.
What gets created
All resources are tagged with Project=<project> so the destroy script can find and remove them later.
Key pair
openg2p-prod-key
key_name
Created if missing; .pem saved to aws/keys/ mode 0400
SG: RP
openg2p-prod-reverse-proxy
rp_sg_name
Single SG: admin SSH (admin_cidr, default 0.0.0.0/0), Wireguard UDP, and all intra-VPC. Public 80/443 are not opened here — admin tools stay private; env automation opens them later. Reused if exists; rules added if missing.
SG: Compute
openg2p-prod-k8s-node
compute_sg_name
Same (admin SSH from admin_cidr)
SG: Storage
openg2p-prod-storage
storage_sg_name
Same (admin SSH from admin_cidr)
Elastic IP
tagged Role=reverse-proxy-eip
—
One EIP allocated and associated with the RP. Script hard-fails on AddressLimitExceeded. See below for why.
Instance: RP
openg2p-prod-reverse-proxy
rp_name
t3a.medium, 64 GB gp3, single ENI
Instance: Compute
openg2p-prod-k8s-node-1
compute_name
m5a.4xlarge, 128 GB gp3
Instance: Storage
openg2p-prod-storage
storage_name
t3a.2xlarge, 256 GB gp3
Default sizing
Matches the OpenG2P resource minimums.
Reverse Proxy
t3a.medium
2
4 GB
64 GB
Compute / K8s
m5a.4xlarge
16
64 GB
128 GB
Storage
t3a.2xlarge
8
32 GB
256 GB
All sizes are configurable in aws-config.yaml via *_instance_type, *_disk_gb, *_disk_iops, *_disk_throughput. Larger is fine; smaller may fail the orchestrator's preflight.
About the Elastic IP
Only the reverse-proxy node gets an Elastic IP. Compute and storage use AWS's auto-assigned dynamic public IPs (which is fine — those public IPs are only used for SSH from your laptop).
Why an EIP, not a dynamic public IP? The Wireguard Endpoint line in every peer config is the RP's public IP. An EIP survives instance stop/start; a dynamic IP would change after any stop/start and break every peer config you've already distributed. AWS single-NIC launches do technically support auto-assigned public IPs, but for production stability we always use an EIP.
Behaviour when the EIP quota is exhausted: if your AWS account is at the default 5-EIP per-region limit, AllocateAddress returns AddressLimitExceeded and the provisioner hard-fails at step 2 ("Allocating Elastic IP for RP"). No instances have been launched yet, so there's nothing to clean up — just free an EIP (or request a quota increase) and re-run.
To check your current EIP usage and free unused ones:
During teardown, openg2p-aws-destroy.sh automatically releases every EIP tagged Project=<project> (step 2 of the destroy flow), so the EIP returns to your pool for reuse on the next provision. No manual cleanup needed.
Workflow
admin_cidr — SSH / ping from the admin laptop
Security groups allow inbound SSH (TCP/22) and ICMP from admin_cidr on every node (RP, compute, storage, and the backup node when enabled).
Value in aws-config.yaml
Behaviour
"" (blank)
Defaults to 0.0.0.0/0 — SSH reachable from any public IP. Survives ISP / network changes.
"0.0.0.0/0"
Same as blank, explicit.
Custom CIDR (e.g. "203.0.113.0/24")
Restricts admin SSH/ping to that range (office, VPN egress, etc.).
Blank no longer auto-detects the laptop's current public /32 (that behaviour broke login whenever the operator changed networks). Prefer 0.0.0.0/0 for bring-up and labs; tighten to a known CIDR for stable production once Wireguard / VPN admin access is in place.
Intra-VPC traffic is always allowed separately (VPC CIDR), independent of admin_cidr.
Interactive selection (default)
When vpc_id, subnet_id, or key_mode are blank in aws-config.yaml, the script queries AWS and presents a numbered menu — no need to leave the terminal to look anything up. Example:
Your selection is written back to aws-config.yaml, so the next run is fully non-interactive. The same applies to subnet selection and key-pair selection (existing AWS key pairs are listed alongside a "Create new" option).
For CI / automation, pass --non-interactive and pre-fill all required values. The script will fail loudly (with a list of options) on anything ambiguous.
Reusing existing security groups
If your infra team has already created security groups, point the *_sg_name fields in aws-config.yaml at their names. The script:
Reuses the existing SG (no new SG created).
Verifies the required ingress rules — adds any that are missing, leaves the rest alone.
Never removes rules.
Per-rule status is logged so you can see what was added vs already present:
provision-output.yaml — what the orchestrator consumes
After AWS provisioning succeeds, the orchestrator's prod-config.yaml does not need any IPs or SSH paths. Those live in a sibling file provision-output.yaml:
The orchestrator auto-detects this file next to prod-config.yaml and loads it as an overlay — its keys win on conflict. Re-running AWS provisioning regenerates it cleanly (single .prev archive, no timestamped backup churn). Your hand-edited prod-config.yaml is never touched.
Tearing down
Confirms by asking you to type the project name back, then deletes everything tagged Project=<project>:
1
EC2 instances
terminate-instances + wait for terminated
1
Root EBS volumes
Auto-deleted with the instance (created with DeleteOnTermination: true)
1
Primary ENI
Auto-deleted with the instance
2
Elastic IPs
release-address
3
Security groups
delete-security-group
4
Key pair
Only deleted if WE created it (tagged Project=<project> ManagedBy=openg2p-aws-provision). Pre-existing keys imported by the user are kept. Force-keep with --keep-key.
5
Stray EBS volumes in available / creating / error state
Explicit delete-volume — catches volumes detached from instances or extras attached after provisioning
5
Stray snapshots owned by you, tagged with the project
Explicit delete-snapshot
5
Stray ENIs in available state, tagged
Explicit delete-network-interface
6
../provision-output.yaml
Removed (stale after teardown)
7
Final sweep
Lists anything still tagged Project=<project> so leaks are visible
A clean teardown ends with Nothing left tagged Project=<project>.
The destroy script only touches resources tagged Project=<project>. If you've created VPC peering, NAT gateways, EFS file systems, or anything else that you tagged with the same project, those will appear in the final sweep — review the list before assuming "all clean."
Costs (rough, us-east-1, on-demand)
t3a.medium (RP)
$0.0376
~$27
m5a.4xlarge (Compute)
$0.688
~$502
t3a.2xlarge (Storage)
$0.301
~$220
EIP (attached)
free
$0
EIP (released-but-unattached)
$0.005
~$3.65
EBS gp3 storage (64 + 128 + 256 = 448 GB)
$0.08/GB-month
~$36
Total
~$785/month if running 24/7
Stop instances when not using them to drop EC2 charges to near-zero (you still pay for EBS). The EIP stays attached to the (stopped) RP, so the Wireguard endpoint survives stop/start when present.
Troubleshooting
SSH times out after changing Wi‑Fi / ISP — with the default admin_cidr (0.0.0.0/0) this should not happen. If you locked admin_cidr to a previous /32, either set admin_cidr: "0.0.0.0/0" (or your new network's CIDR) in aws-config.yaml and re-run the provisioner so it adds the missing ingress rule, or update the security group in the AWS console. The provisioner never removes existing rules.
AWS provision: "VPC not found" — some accounts have no default VPC. Either create one (aws ec2 create-default-vpc), set vpc_id and subnet_id explicitly in aws-config.yaml, or run with the default vpc_id: "" and pick interactively.
AWS provision: EIP AddressLimitExceeded — the script hard-fails at step 2 (no instances launched yet). The provisioner allocates one EIP for the RP so the Wireguard endpoint IP survives instance stop/start. Free an unused EIP and re-run, or request a quota increase:
See About the Elastic IP for why we use an EIP.
Multiple environments on the same AWS account — use a different project: value in each aws-config.yaml (e.g., openg2p-prod, openg2p-staging). Resources are isolated by tag; the destroy script only touches the configured project.
Last updated
Was this helpful?