Getting Started¶
Prerequisites¶
| Tool | Version | Install |
|---|---|---|
| AWS CLI | ≥ 2.15 | brew install awscli |
| Terraform | ≥ 1.5 | brew install terraform |
| kubectl | ≥ 1.29 | brew install kubectl |
| Helm | ≥ 3.14 | brew install helm |
| Docker | ≥ 24 | docker.com |
| Node.js | ≥ 22 | brew install node (for Backstage build) |
How long does it take?¶
Local (Kind / Rancher Desktop) — ~15–20 minutes¶
./scripts/bootstrap-local.sh runs end-to-end without AWS credentials.
| Phase | What happens | Time |
|---|---|---|
| Kind cluster creation | kind create cluster, load balancer, kubeconfig |
~2 min |
| nginx ingress controller | Helm install + wait for pods | ~1 min |
| ArgoCD | Helm install + wait for pods + register GitHub credentials | ~3 min |
| Prometheus + Grafana | kube-prometheus-stack Helm install |
~4 min |
| Backstage | Docker Compose build + container start + DB migrations | ~4 min |
| DORA exporter + seed metrics | CronJob apply + one-shot job | ~1 min |
| Total | ~15–20 min |
--skip-obs (skip Prometheus/Grafana) saves ~4 minutes.
Adding the AI/ML stack (./scripts/bootstrap-ai.sh) takes an additional 10–15 minutes:
| Component | Time |
|---|---|
| KAgent controller + idp-assistant agent | ~4 min |
| MLflow tracking server | ~3 min |
| IDP + QA + Contract MCP servers (3×) | ~4 min |
/etc/hosts entries + port-forward |
<1 min |
--skip-mlflow, --skip-mcp, --skip-kagent each save ~3–4 minutes from the AI stack.
AWS (EKS) — ~40–70 minutes¶
./scripts/bootstrap.sh provisions from scratch. Most of the time is AWS control-plane and ALB provisioning, which cannot be parallelised.
| Phase | What happens | Time |
|---|---|---|
| Terraform | VPC, EKS control plane + node groups, RDS, ECR, IAM/OIDC, Crossplane IRSA role | ~20–25 min |
| ArgoCD + app-of-apps | Helm install + GitHub credentials + first sync | ~5 min |
| External Secrets Operator | Helm install + ClusterSecretStore ready | ~3 min |
| Prometheus + Grafana | kube-prometheus-stack + Grafana ALB provisioning |
~5 min |
| OPA/Gatekeeper | CRDs + constraints | ~2 min |
| Crossplane | Core + AWS providers healthy + compositions applied | ~5 min |
| Backstage | ECR image push + K8s deploy + ExternalSecret sync + ALB | ~8 min |
| hello-service + ALB | Helm install + ALB DNS propagation | ~3 min |
| Total | ~40–70 min |
EKS control plane creation (~10 min) and ALB provisioning (~3–5 min per ingress) are the longest waits and are entirely AWS-side — no script change can speed them up.
Adding the AI/ML stack on AWS takes an additional 15–20 minutes:
| Component | Time |
|---|---|
| KAgent controller + idp-assistant + ALB ingresses | ~5 min |
| MLflow (S3 backend) + ALB | ~5 min |
| IDP + QA + Contract MCP servers + ALBs | ~5 min |
| Backstage proxy config patch + restart | ~2 min |
Re-bootstrap (cluster already exists): If you're re-running bootstrap.sh against an existing EKS cluster, Terraform applies only the diff (usually <2 min) and the rest of the phases take ~15–20 minutes total — ALB re-provisioning is skipped if the ingresses already exist.
Quick comparison¶
| Local (full) | Local + AI/ML | AWS (full) | AWS + AI/ML | |
|---|---|---|---|---|
| First run | ~15–20 min | ~25–35 min | ~40–70 min | ~60–80 min |
| Re-bootstrap | ~5–8 min | ~10–15 min | ~15–20 min | ~20–25 min |
| Prerequisites | Docker, Kind | + ANTHROPIC_API_KEY |
AWS account, Terraform | + ANTHROPIC_API_KEY |
| Cost | Free | Free | ~$19/day | ~$25/day |
Local Setup (no AWS needed)¶
See docs/local-setup.md for the full local walkthrough including Backstage, the idp:deploy-local action, and Kind deployment.
Run
setup.sh— not both. On a fresh clone you run./scripts/setup.shand nothing else. When you answer local, it callsbootstrap-local.shfor you and then offers to start Backstage. Runningbootstrap-local.shyourself afterwards just repeats a 15–20 minute install for no benefit.
bootstrap-local.shis what you run standalone later, for day-2 work: recreating the cluster,--destroy,--start-backstage,--print-urls. You do not re-runsetup.shfor those.
Personalisation has to happen before bootstrapping: setup.sh replaces moatazeldebsy and the other placeholders across the repo, and without it the ArgoCD ApplicationSet points at an unresolved placeholder and generates no apps.
If ArgoCD shows no apps: check that
local/argocd/app-of-apps-local.yamlcontains your GitHub org rather thanmoatazeldebsy, then re-runsetup.sh.
AWS Setup¶
🔐 CRITICAL - First: Read docs/PRE_DEPLOYMENT_CHECKLIST.md and verify all API keys are set correctly. Run
./scripts/verify-secrets.shto validate before deployment.⚠️ NEW: Then read docs/DEPLOYMENT_GUIDE.md for a complete step-by-step guide, pre-flight checklist, known issues with solutions, and troubleshooting. Estimated deployment time: 40–70 minutes.
1. Configure AWS¶
1 2 | |
2. Bootstrap the platform¶
1 2 3 4 5 6 | |
Or, if you have already run setup.sh for personalisation and want to re-run the AWS bootstrap directly:
1 2 3 | |
3. Validate deployment¶
After bootstrap.sh completes, run the validation script to verify all components:
1 | |
This runs ~40 automated checks across 10 categories (AWS infrastructure, Kubernetes, Backstage, observability, GitOps, AI/ML, security, networking, storage, and cost). Exit code 0 = success; 1 = failure with debug suggestions.
What gets provisioned:
- EKS cluster (4× t3.medium nodes, 1.32)
- RDS PostgreSQL (for Backstage)
- ECR repository + S3 bucket (for artifacts)
- IAM roles + OIDC (for GitHub Actions and Crossplane)
- All platform components (Prometheus, Grafana, ArgoCD, OPA/Gatekeeper, External Secrets Operator)
- Crossplane with AWS providers (for per-service resources like S3, RDS, DynamoDB)
- hello-service reference deployment
Cost: ~$565/month for the core platform, ~$760/month with the AI/ML layer — measured against a running cluster, not estimated. Full per-component breakdown and the ways to spend less: README → What it costs on AWS.
4. GitHub Actions secrets¶
Add these secrets to any scaffolded service repo to enable AWS CD:
| Secret | Value |
|---|---|
AWS_ROLE_ARN |
cd terraform && terraform output github_actions_role_arn |
AWS_REGION |
us-east-1 |
ECR_REGISTRY |
<account>.dkr.ecr.us-east-1.amazonaws.com |
EKS_CLUSTER |
idp-mvp |
Add these to the platform repo to enable the auto-merge workflow (recommended over a PAT):
| Secret | Value |
|---|---|
APP_ID |
Numeric GitHub App ID (see docs/github-app-setup.md) |
APP_PRIVATE_KEY |
PEM contents of the App's private key |
Add this to the platform repo to enable Datadog deployment markers in build-and-deploy.yml
(optional — the pipeline runs fine without it, just skips the marker step):
| Secret | Value |
|---|---|
DD_API_KEY |
Datadog API key — https://app.datadoghq.eu/organization-settings/api-keys |
5. Team namespace setup¶
After the platform is bootstrapped, onboard teams using the Provision Team Namespace scaffold template (tagged blessed). Each team gets:
- An isolated team-<name> namespace with quota, LimitRange, NetworkPolicy
- A per-team ArgoCD AppProject + ApplicationSet scanning teams/<name>/services/*
- A scoped SecretStore (access only to /<name>/* in Secrets Manager)
- A Grafana folder
For the full walkthrough, see docs/team-management.md.
Path convention: Team service values files go under
teams/<teamName>/services/<serviceName>/, notservices/<teamName>/. Theservices/path is reserved for legacy platform-owned services.
6. Verify¶
1 2 3 4 5 6 7 | |
Visit the ALB hostname:
1 | |
Observability note:
bootstrap.shinstalls the fullkube-prometheus-stack(Prometheus + Grafana + AlertManager + Pushgateway) on AWS at parity with the local Kind setup. Grafana is pre-configured with the CloudWatch datasource using IRSA — no static AWS credentials needed.OPA/Gatekeeper enforces all five golden-path policies (
require-health-probes,require-resource-limits,require-labels,deny-latest-tag,require-cost-tags). The bootstrap waits for CRDs to be established before applying constraints rather than sleeping.
7. Backstage on AWS¶
Nothing to do — bootstrap.sh already did it. Phase 5.6 builds the image, pushes it
to ECR under a content-hash tag, applies aws/backstage/deployment.yaml with that
tag substituted, waits for the ExternalSecret to sync and for the ALB hostname, then
patches the config with the real URLs.
This section used to describe a manual docker build ending in
# Deploy (Kubernetes manifests TBD). The manifests have existed for some time and
the whole flow is automated; following the old steps would have pushed an image that
nothing deployed.
To rebuild and roll out after changing Backstage, re-run the bootstrap — the image fingerprint changes, so it rebuilds and redeploys, and skips everything else that is unchanged:
1 | |
Adding AWS CD to a Scaffolded Service¶
Scaffolded service repos ship with CI only (test job). To add AWS deployment:
- Add the four secrets above to the GitHub repo
- Add a
deployjob to.github/workflows/build-and-deploy.yml:
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 | |
Teardown¶
1 | |
Use cleanup.sh rather than a bare terraform destroy. It deletes orphaned load
balancers first — AWS Load Balancer Controller creates ALBs that Terraform does not
own, and they hold the subnets and security groups destroy is trying to remove, so
it fails partway and leaves the account in a half-torn-down state. The script also
verifies the teardown afterwards.