global-capacity-orchestrator-on-aws

mcp
Security Audit
Fail
Health Pass
  • License — License: MIT-0
  • Description — Repository has a description
  • Active repo — Last push 0 days ago
  • Community trust — 49 GitHub stars
Code Fail
  • rm -rf — Recursive force deletion command in .github/actions/apt-install-with-retry/action.yml
  • rm -rf — Recursive force deletion command in .github/actions/free-disk-space/action.yml
Permissions Pass
  • Permissions — No dangerous permissions requested

No AI report is available for this listing yet.

SUMMARY

AI/ML and HPC at planetary scale. One API. Every Accelerator. Any Region. MIT-0 licensed

README.md

Global Capacity Orchestrator (GCO) runs accelerated workloads — LLM training and inference, batch ML, HPC — on EKS Auto Mode clusters in as many AWS Regions as you configure, behind one IAM-authenticated API, CLI and MCP server. It finds where NVIDIA GPU, Trainium, Inferentia and CPU capacity actually is, places jobs there, and serves inference endpoints with automatic cross-Region failover.

A recorded gco session: fleet status, capacity discovery, four schedulers, shared storage, a vector search and live LLM inference

A real gco session, reproducible from demo/live_demo.sh: fleet status, capacity discovery, four schedulers plus KEDA, shared storage, a globally replicated vector store and live LLM inference. More recordings are under See it running.

Why GCO?

Running accelerated workloads across Regions means finding capacity, standing up clusters, wiring authentication, handling failover, and keeping job outputs after the pods exit. GCO packages all of it as one platform you deploy with a single command:

  • Capacity-aware placement. gco capacity reads spot placement scores, spot price history, On-Demand Capacity Reservations and Capacity Blocks for ML, and auto-Region submission builds on them.
  • One authenticated front door. A SigV4 API, the gco CLI and an MCP server reach every Region with the AWS credentials you already have, with no kubeconfig distribution.
  • Accelerators on demand. EKS Auto Mode NodePools for NVIDIA GPU (x86 and Arm), EFA, Trainium, Inferentia and CPU launch nodes only when a workload needs them.
  • Inference in every Region. One command deploys vLLM, SGLang or Triton endpoints everywhere, with automatic failover through Global Accelerator in the commercial aws partition.
  • A curated ecosystem. KEDA, Volcano, Kueue, KubeRay and Kubeflow Trainer are on by default and Slurm, YuniKorn, Argo CD and Crossplane are opt-in; EFS ships with every cluster, and FSx for Lustre, Valkey and Aurora pgvector are a toggle away; Prometheus, Grafana, OpenCost and MLflow are on by default for observability and cost tracking.
  • Agent-ready. gco autopilot opens Claude Code, OpenAI Codex or OpenCode on Amazon Bedrock with the GCO MCP server already wired in.
How GCO compares with running the clusters yourself
Challenge Traditional Approach With GCO
GPU availability Manually check each region Capacity tools and auto-region workflows compare configured regions
Node provisioning Pre-provision or wait for scaling EKS Auto Mode provisions on-demand
Multi-region ops Manage clusters separately One platform across unlimited SDK-known Regions in one partition
Authentication Configure per-cluster access IAM-based, uses existing AWS credentials
Job outputs Lost unless persisted EFS/FSx and per-region S3 available to every job that mounts or writes them
Inference serving Deploy and manage per-region Deploy once across selected Regions; global failover in aws
Failover Manual intervention required Automatic via Global Accelerator in aws; explicit regional selection elsewhere

When to use GCO:

  • You need to run GPU workloads (training, inference, batch processing)
  • You want to deploy inference endpoints across multiple regions with a single command
  • You want multi-region redundancy without managing multiple clusters
  • You prefer IAM authentication over kubeconfig management
  • You need job outputs to persist after completion

What it does. Spins up EKS Auto Mode clusters across any number of SDK-known CloudFormation Regions in one AWS partition. In commercial aws, Global Accelerator provides latency-aware anycast routing and automatic failover behind the global workload API; other partitions (aws-cn and aws-us-gov) use IAM-authenticated regional workload APIs while retaining the aggregate global API. Capacity tools and auto-region queue/CLI workflows select a target Region, EKS Auto Mode provisions matching nodes from the built-in system and general-purpose NodePools plus project-managed GPU x86, GPU ARM, inference, EFA, Mooncake EFA, Neuron, and CPU NodePools, and shared storage persists workload outputs. Network routing never substitutes for live GPU-capacity placement.

Why it's different. Capacity-aware placement tools and auto-region workflows, partition-aware authenticated routing, full-stack observability (CloudWatch dashboards, alarms, SNS), and a CDK app validated across the full curated configuration matrix in CI. Read Core Concepts for the ideas behind it and the Learning Path if Kubernetes is new to you.

Get started

Prerequisites

  • An AWS account you can create infrastructure in, with credentials the AWS CLI can use
  • Git and a container runtime: Docker, Finch or Podman (Colima also works)

The dev container brings everything else — Python 3.14, Node.js 24, CDK, kubectl, the AWS CLI, and Docker CLI + Buildx — at pinned versions. A host install is possible but advanced; see More ways to run GCO below.

Deploy

git clone https://github.com/aws-solutions-library-samples/global-capacity-orchestrator-on-aws.git
cd global-capacity-orchestrator-on-aws
./scripts/setup-dev-alias.sh   # builds the dev container and installs the `gco` shell function
source ~/.zshrc                # or ~/.bashrc — the script prints which file it updated
gco stacks deploy-all -y       # stands up every Region in cdk.json; billing starts here

deploy-all bootstraps CDK where needed, deploys the global control plane and one regional stack per Region, and then installs the cluster add-ons in the background.

Run your first workload

gco capacity check --instance-type g4dn.xlarge --region us-east-1   # where is capacity right now?
gco jobs submit-sqs examples/simple-job.yaml --region us-east-1     # submit a job over SQS
gco jobs list --all-regions                                         # watch it run
gco jobs logs hello-gco -n gco-jobs -r us-east-1                    # read its output

Serve a model from every Region with one command:

gco inference deploy my-llm -i vllm/vllm-openai:v0.31.0 --gpu-count 1   # OpenAI-compatible endpoint
gco inference status my-llm                                             # rollout state in every Region
gco inference scale my-llm --replicas 3                                 # scale it out

Let an agent drive

gco autopilot                     # default Claude Code session
gco autopilot --engine codex      # or an OpenAI Codex session
gco autopilot --engine opencode   # or an OpenCode session (Kimi K3)

Every engine runs on Amazon Bedrock with your AWS credentials and the GCO MCP server plus the recommended companion MCP servers wired in, so you can simply ask: "deploy everything", "where is p5 capacity cheapest right now?" The Autopilot Guide covers sessions, opt-in tool groups and dry runs.

Clean up

gco stacks destroy-all -y   # destroys every stack, then sweeps known leftovers

Teardown is not an account-wide emptiness guarantee: retained ECR repositories, resources configured for retention, and unexpected resources can remain. See gco stacks destroy-all for the exact cleanup scope.

More ways to run GCO: the dev container in depth, an interactive shell, or a host install

The dev container. GCO pins exact versions of a lot of Python packages (CDK, AWS SDKs, FastAPI, mypy, Ruff, etc.), and installing them on top of an existing Python environment is the most common source of "it doesn't install" reports. The dev container ships a fully resolved environment so you skip the whole problem, and scripts/setup-dev-alias.sh means you never hand-write a docker run … or live inside an interactive container shell:

./scripts/setup-dev-alias.sh   # builds gco-dev from Dockerfile.dev + installs the `gco` shell function
source ~/.zshrc                # or ~/.bashrc — the script prints which file it updated
gco --help                     # every command now runs inside the container, against your checkout

The script detects your container runtime, builds the gco-dev image from Dockerfile.dev (re-running rebuilds it, so a stale image is refreshed automatically — pass --no-build to skip), wires up the correct socket pass-through, and installs an idempotent gco shell function (not a bare alias, so arguments and pipes forward correctly and a TTY is attached only when one is present) into your profile. Re-run it whenever you switch runtimes, or use --print to preview the function, --runtime <name> to force one, and --rc <path> to target a specific profile. To move an existing deployment to a newer release later, gco upgrade checks out the release tag, rebuilds this image, and cycles the stacks in one pass.

The function shares your host Docker socket with every gco call because gco stacks deploy-all needs it: through your host daemon it builds and bundles the Lambda function assets (as linux/amd64, cross-built via Buildx so this works on Apple Silicon / arm64) and — when the Volcano image mirror is enabled in cdk.json — mirrors third-party images from Docker Hub into your ECR before the Helm install runs. The same socket backs gco images build / push. This is host-socket pass-through, not Docker-in-Docker: anyone with access to the container has root-equivalent access to the host Docker daemon, so keep the container on a trusted host. Finch users: Finch runs in its own VM with no host Docker socket to share, so the function omits the socket mount — everyday commands work as-is, while build-heavy ones like deploy-all run on the host with Finch as the CDK builder.

An interactive container shell. Build the image yourself and drop into it, running gco from inside — handy for ad-hoc tools and exploration. The -v /var/run/docker.sock:/var/run/docker.sock mount gives the container's Docker CLI access to your host daemon for the asset builds and image mirroring described above (Colima users: see the header of Dockerfile.dev for the socket path):

docker build -f Dockerfile.dev -t gco-dev .
docker run -it --rm \
  -v ~/.aws:/root/.aws:ro \
  -v $(pwd):/workspace \
  -v /var/run/docker.sock:/var/run/docker.sock \
  -w /workspace \
  gco-dev

A host install (advanced). Host installs are the advanced, non-recommended path: GCO's exact pins frequently fail to resolve on top of an existing Python environment (ResolutionImpossible). If you still want one, use a clean virtual environment or pipx — the Quick Start has both recipes:

pipx install -e .

A host install additionally needs:

  • Python 3.14+ and Node.js 24 (use .nvmrc)
  • npm 12.2.0 and the repository's locked tooling graph: run npm ci --ignore-scripts --no-audit --no-fund at the repository root; gco prefers its local node_modules/.bin/cdk over a global CLI
  • A clean Python virtual environment or pipx — GCO pins exact versions of many packages, so installing into an existing environment commonly fails with dependency-resolver errors. If you hit ResolutionImpossible, switch to the dev container instead of debugging your local environment.
After you deploy: add-ons converge in the background, and kubectl is optional

Heads up — Helm charts finish installing in the background. When deploy-all reports the cluster CREATE_COMPLETE, the scheduler/operator Helm charts (KEDA, Volcano, KubeRay, cert-manager, Kueue, …) have only been kicked off; they converge asynchronously and can take 10–30+ minutes to all become ready. This is intentional — a slow chart never rolls back the cluster. Track progress with gco stacks addons status -r <region> and re-converge any failures with gco stacks addons install -r <region>. See docs/CUSTOMIZATION.md.

Optional: kubectl access. The EKS API endpoint is PRIVATE by default and most users never need kubectl — submit jobs through SQS or API Gateway instead. When you do, gco cluster tunnel --via-ssm auto reaches the private endpoint from your laptop over SSM and gco stacks access -r <region> grants your IAM principal an EKS access entry; gco cluster doctor checks both. See docs/CUSTOMIZATION.md.

Every way to submit a job: SQS, the global queue, the REST API, or kubectl

Submit a job using whichever path fits your setup — via SQS (recommended), via the global DynamoDB queue, via API Gateway, or directly through kubectl:

gco jobs submit-sqs examples/simple-job.yaml --region us-east-1
gco queue submit examples/simple-job.yaml --region us-east-1
gco jobs submit examples/simple-job.yaml -n gco-jobs
gco jobs submit-direct examples/simple-job.yaml -r us-east-1

The Quick Start Guide walks through every step with the output to expect, and the CLI Reference lists every command.

See it running

Real terminal recordings, not mock-ups. Each one is reproducible from the script linked beneath it; the live demo itself plays at the top of this page.

📦 Deploy recording

GCO Deploy

Fresh gco stacks deploy-all -y --enable fsx_lustre,valkey,aurora_pgvector,vector_store,slurm,yunikorn from a clean account — the six add-ons ship disabled in cdk.json because each bills continuously, so the recording enables them for one run through run-scoped overrides (re-record)

🗑️ Destroy recording

GCO Destroy

Full teardown with gco stacks destroy-all -y --enable fsx_lustre,valkey,aurora_pgvector,vector_store,slurm,yunikorn, repeating the deploy's overrides so it evaluates the same app (re-record)

🤖 Claude Code Autopilot recording — the default engine, ready in one command

GCO Autopilot with Claude Code

A real session: gco autopilot launches Claude Code on Amazon Bedrock (GCO's default Claude Opus profile) with the GCO MCP server + companion MCPs wired in, and it answers from the project's own MCP tools (docs · re-record).

🤖 OpenAI Codex Autopilot recording — the Codex engine, ready in one command

GCO Autopilot with OpenAI Codex

A real Bedrock-backed gco autopilot --engine codex --no-companions session using a least-privilege recording profile: required GCO MCP, only find_docs/read_resource, no built-in shell, and no trust or approval prompts. A normal gco autopilot --engine codex launch includes the recommended companions (docs · re-record with DEMO_ENGINE=codex DEMO_MODE=live).

🤖 OpenCode Autopilot recording — the OpenCode engine on Kimi K3, ready in one command

GCO Autopilot with OpenCode

A real Bedrock-backed gco autopilot --engine opencode --no-companions session (Moonshot AI's Kimi K3) using a least-privilege recording profile: only the GCO MCP find_docs/read_resource tools, every built-in OpenCode tool denied, and no permission prompts. A normal gco autopilot --engine opencode launch includes the recommended companions and OpenCode's ordinary tool set behind an ask-first permission floor (docs · re-record with DEMO_ENGINE=opencode DEMO_MODE=live).

Install the MCP server

GCO ships with the GCO MCP server — an MCP server exposing 142 tools by default (up to 200 with feature flags) that index the whole project: docs, examples, source code, K8s manifests, and scripts. Connect it to an AI-powered IDE with MCP support and explore GCO conversationally — "How does region recommendation work?", "Walk me through the inference deployment flow" — or let it drive real operations. One click adds it (pinned to the latest release, no clone needed) to your client; uv and AWS credentials are the only prerequisites. Other clients and options: setup guide.

Kiro Cursor VS Code
Add to Kiro Add to Cursor Install on VS Code

Architecture Overview

GCO multi-region reference architecture

Figure 1: GCO multi-region reference architecture — global control plane and workload entry

One CDK app deploys a global control plane (the API Gateway entry point, Global Accelerator in aws, shared DynamoDB state, the model bucket and monitoring) and one regional stack per Region: an EKS Auto Mode cluster with GCO's platform services, NodePools, storage and job queue. Architecture Details is the full deep dive, and the AWS Solutions Guidance presents the reference architecture in the AWS Solutions Library.

The multi-Region workflow, step by step

The generated reference architecture shows the commercial aws workload path. Other partitions retain the global aggregate API but route workload control and inference through each Region's IAM-authenticated bridge.

  1. DevOps / Platform engineers own the deployment. They configure the platform through cdk.json and drive everything from the gco CLI.
  2. The AWS CDK app synthesises and deploys the GCO stacks with a single gco stacks deploy-all, provisioning the global control plane and one regional stack per target region.
  3. Users submit jobs and inference requests through the gco CLI, which signs every call with AWS SigV4 credentials.
  4. In commercial aws, Amazon API Gateway is edge-optimized and is the global workload and aggregate entry point. In other partitions it is regional and aggregate-only. Every exposed method enforces IAM (SigV4) authentication before integration.
  5. In aws, route-specific AWS Lambda proxies sign workload requests with a short-lived HMAC envelope derived from a rotating AWS Secrets Manager key; /api/v1/* stays buffered while /inference/* streams. Other partitions omit these global workload proxies and use equivalent VPC proxies behind the regional APIs.
  6. In aws, AWS Global Accelerator routes workload requests over the AWS backbone to a healthy registered Region. Other partitions create no accelerator resources.
  7. A regional internal AWS Application Load Balancer terminates deployment-local private-root TLS from either Global Accelerator (aws) or the regional VPC proxy, then re-encrypts to TLS-only proxy sidecars on the platform API pods. Each sidecar hot-reloads its projected certificate and forwards decrypted traffic only over pod loopback.
  8. Each region runs an Amazon EKS Auto Mode cluster with built-in system and general-purpose NodePools plus project-managed GPU, inference, EFA, Mooncake EFA, Neuron, and CPU NodePools. Platform services include the Cost Monitor, Health Monitor, Manifest Processor, Queue Processor, Inference Monitor, and dedicated Inference Proxy.
The regional architecture: one Region's EKS Auto Mode data plane, step by step

GCO regional EKS reference architecture

Figure 2: GCO regional reference architecture — EKS Auto Mode data plane and regional services

  1. An internal Application Load Balancer created from the shared gco-system/gco-gateway Gateway API resources accepts only HTTPS/443 with a rotating regional ACM leaf, then re-encrypts target traffic to cert-manager-backed HTTPS listeners on the cluster API services. ALB target TLS provides confidentiality; HMAC proves trusted-proxy key possession and request integrity on protected paths, while API Gateway IAM authenticates the original caller.
  2. The Amazon EKS Auto Mode cluster is the heart of the regional stack, hosting platform services and user workloads with a private API endpoint by default.
  3. NodePools provision capacity on demand: built-in system and general-purpose, plus gpu-x86-pool, gpu-arm-pool, gpu-inference-pool, gpu-efa-pool, mooncake-efa-pool, neuron-pool, and cpu-general-pool.
  4. Workloads and platform services run across namespaces: gco-system (Health Monitor, Manifest Processor, Queue Processor, Inference Monitor, Inference Proxy) and gco-jobs / gco-inference (training and batch jobs, inference endpoints, and job DAG pipelines).
  5. Storage and data services back workloads: Amazon EFS, optional FSx for Lustre, optional Valkey, optional Aurora pgvector, and Amazon S3 for KMS-encrypted model weights.
  6. An always-deployed Regional API Gateway bridge gives the aggregator a SigV4-authenticated path to the VPC Lambda and internal ALB. Direct same-account access is optional through regional_api_enabled in aws and enabled automatically as the required workload ingress elsewhere.
  7. Regional AWS services complete the stack: Amazon SQS for job ingestion, DynamoDB-backed state where applicable, and Amazon CloudWatch metrics and logs.
The generated CDK diagram of every deployed resource, and how the diagrams are built

Generated GCO infrastructure architecture

Figure 3: Generated CDK architecture for the global control plane and regional EKS data planes

Regenerate the full architecture and every per-stack view with python diagrams/infra_diagrams/generate.py. The generator synthesizes the current CDK app through cdk-dia so the committed diagrams track the deployed resource graph. See diagrams/infra_diagrams/README.md for per-stack flags (--stack global|api-gateway|regional|regional-api|monitoring|analytics|all).

Flowcharts of Lambda handlers, CLI commands, stack constructors, and MCP control paths live under diagrams/code_diagrams/. Regenerate them through the canonical two-commit workflow, which records an exact source commit without creating a self-referential SHA. Add newly charted functions to diagrams/code_diagrams/_targets.py.

The HTTP API surface has its own catalogue: diagrams/api_specs/ holds one spec sheet per surface — the two AWS API Gateways, the in-cluster Gateway (the internal ALB) and the four FastAPI services — with endpoint tables, parameters, request bodies, responses, component schemas and, for the gateways, each route's Lambda backend and the hops to the Service that answers; all rendered from generated OpenAPI documents (docs/openapi/: the services' own exports, the API Gateway stacks read at CDK synthesis, and the HTTPRoute composed with the service documents). The catalogue index embeds an interaction diagram of how the gateways, the Lambda proxies and the services fit together, drawn from the same documents. The sheets are also the wiki's API reference, and a Swagger UI console for each document is published at /swagger/. Regenerate with python diagrams/generate.py --api-only after refreshing the documents (python scripts/generate_openapi.py, python scripts/generate_api_gateway_openapi.py, python scripts/generate_cluster_gateway_openapi.py); python diagrams/generate.py --check fails when a route, rule or model changed without the sheets.

Security Model

Six complementary controls protect every backend request: IAM authentication at API Gateway, TLS trust separation, a request-bound HMAC, private backend exposure, freshness and integrity validation in the backend middleware, and IRSA / EKS Pod Identity for pod-level AWS access.

The six controls in detail

Six complementary controls protect backend requests:

  1. IAM authentication — API Gateway validates AWS credentials with SigV4.
  2. TLS trust separation — API Gateway uses AWS-managed TLS; proxy-to-ALB traffic uses a deployment-local private root and explicit backend.<project>.gco.internal SNI/hostname verification; the in-cluster hops behind the ALB (cost monitor, OpenCost, model endpoints, metrics scrapes) use a cluster-local cert-manager CA that every client verifies against (In-cluster TLS).
  3. Request-bound HMAC — a trusted Lambda signs the version, timestamp, nonce, method, exact target, and body digest with a rotating key that is never transmitted. HMAC provides integrity/freshness/replay defense, not encryption.
  4. Private backend exposure — regional ALBs are internal and the EKS API endpoint is private by default.
  5. Freshness and integrity validation — backend middleware rejects stale, altered, or replayed envelopes.
  6. IRSA / EKS Pod Identity — pods receive scoped AWS permissions without static workload credentials.
The authenticated request flow, per partition

GCO security controls and request flow

Figure 4: GCO security model — layered controls and the authenticated request flow

Commercial `aws` request flow:
User → API Gateway (AWS TLS + SigV4) → Lambda (HMAC)
  → Global Accelerator (TCP/443 pass-through) → internal ALB (private-root TLS)
  → pod TLS proxy (re-encrypted HTTPS) → AuthenticationMiddleware (pod loopback)

Other partitions:
User → Regional API Gateway (AWS TLS + SigV4) → VPC Lambda (HMAC)
  → internal ALB (private-root TLS)
  → pod TLS proxy (re-encrypted HTTPS) → AuthenticationMiddleware (pod loopback)

Every Region's API bridge is required for aggregator fan-out. Direct regional
access

is optional for same-account callers in aws and enabled automatically as the
supported workload ingress in other partitions.

Key Features

Everything GCO does, grouped by area; expand one for the details and its guide.

Compute and orchestration: NodePools for every accelerator, four submission paths, distributed training, pipelines and Helm-managed schedulers Inference serving: multi-Region vLLM, SGLang and Triton endpoints, canaries, model weights, Spot and autoscaling
  • Multi-region inference: Deploy endpoints across regions with a single command. Each supported framework ships with a ready-to-run example manifest: vLLM (example), SGLang (example), and Triton (example)
  • Canary deployments: A/B test new model versions with weighted traffic routing
  • Model weight management: Central S3 bucket with KMS encryption, automatic sync to each region
  • Spot instance support: Run inference on spot GPUs for significant cost savings
  • Autoscaling: HPA-based scaling with CPU/memory metrics
Networking and security: Global Accelerator, SigV4, cdk-nag policy packs, NetworkPolicies, verified in-cluster HTTPS and EFA
  • Global Accelerator: Single anycast endpoint with automatic failover
  • IAM authentication: SigV4 at the API Gateway — no kubeconfig distribution
  • Infrastructure policy validation: cdk-nag v3 rule packs for AWS Solutions, HIPAA, NIST 800-53, PCI DSS, and Serverless findings (these checks are not certifications)
  • Network policies: Default-deny with explicit allow rules for all service communication
  • Verified in-cluster HTTPS: every hop GCO owns inside a cluster (cost monitor, OpenCost, model endpoints, prefill/decode, Grafana's admin API, metrics scrapes) is TLS from a sidecar with a cert-manager leaf, verified by the client against one cluster-local CA — see In-cluster TLS
  • EFA support: Optional Elastic Fabric Adapter for high-bandwidth distributed training and NIXL-based inference (toggle on/off)
Storage and data: EFS, FSx for Lustre, Valkey, Aurora pgvector and a globally replicated vector store
  • EFS: Shared elastic storage for job outputs that persist after pod termination
  • FSx for Lustre: Optional high-performance parallel file system for ML training (toggle on/off)
  • Valkey cache: Optional serverless key-value cache for prompt caching and session state
  • Aurora pgvector: Optional serverless vector database for RAG, semantic search, and embedding storage
  • Vector store: Optional globally replicated DynamoDB vector index over an S3-ingested document corpus — drop files in, search from every region (gco vector)
Operations: Autopilot, cost visibility and analytics, Spot-aware scheduling, MLflow, monitoring, tracing, GitOps, Crossplane and EKS Capabilities
  • Multi-engine Autopilot: launch Claude Code by default with gco autopilot, OpenAI Codex with gco autopilot --engine codex, or OpenCode with gco autopilot --engine opencode; every engine uses Amazon Bedrock with the GCO MCP server and recommended companion MCPs preconfigured
  • Cost visibility: Track spend by service, region, and workload via Cost Explorer integration
  • Cost monitoring & analytics (on by default): per-cluster OpenCost with a Grafana cost dashboard, scheduled Parquet cost reports to a central S3 bucket, and cross-region Athena analytics via gco costs k8s — see Cost Monitoring Guide
  • Spot price-aware scheduling: central-queue jobs can set a max spot price per instance type and dispatch only when the market clears it
  • MLflow experiment tracking (on by default with observability): an in-cluster MLflow tracking server per region — run artifacts to S3 via a prefix-scoped IAM role, metadata on EBS, reached with gco monitoring open --service mlflow — see MONITORING.md
  • Auto-bootstrap: CDK bootstrap runs automatically for new regions during deploy
  • Multi-region monitoring: the gco-monitoring stack's cross-region CloudWatch dashboards, alarms, and SNS alerts, complemented by per-cluster Prometheus/Grafana cluster observability
  • Distributed tracing (on by default, 5% sampled): the four API services export OpenTelemetry spans straight to AWS X-Ray with no collector, searchable in CloudWatch Transaction Search, and every service log line carries the matching trace id — see Distributed tracing
  • GitOps with Argo CD (off by default): a self-managed Argo CD per regional cluster, installed from the upstream chart in namespaced mode and fenced to the job namespaces by a gco-tenants AppProject and matching RBAC; point it at a repository path per cluster from cdk.json, let the repo server autoscale, and open the UI over the private endpoint with gco gitops open — see GitOps with Argo CD
  • Crossplane (off by default): a self-managed Crossplane v2 and the Crossview dashboard, whose namespaced composite resources compose tenant workloads in the job namespaces only (gco crossplane open) — see Crossplane
  • EKS Capabilities (off by default): attach the AWS-managed ACK and kro capabilities per regional cluster from cdk.json, with IAM roles that carry only the permissions you configure and the tenant RBAC kro composes with; gco stacks capabilities status reports drift — see EKS Capabilities
ML and analytics environment: SageMaker Studio, EMR Serverless and Cognito for notebook analytics
  • ML & Analytics Environment: Optional SageMaker Studio domain + EMR Serverless + Cognito user pool for interactive notebook analytics, with an always-on Cluster_Shared_Bucket that all cluster jobs can read and write. Off by default — enable with gco analytics enable. See Analytics Guide.
Mission: goal-directed iteration loops with deterministic verdicts and budget caps

Goal-directed iteration loop for orchestrated workflows. The operator declares a natural-language directive plus machine-checkable success criteria, a tool allowlist, and a budget; Mission runs five-phase iterations (propose → execute → observe → evaluate → decide) until a verdict is reached. Off by default — enable with GCO_ENABLE_MISSION=true. See Mission Guide.

  • Deterministic verdict cascade with optional advisory LLM sampling (MCP host or Amazon Bedrock). Sampling shapes only the next strategy; it never moves the verdict.
  • Budget caps on iterations and wall clock — the engine terminates cleanly when any cap fires. Cost guardrails live out-of-band via AWS Budgets and Cost Anomaly Detection at the account level.
  • Scripted strategies opt-in: an AST-validated Python sandbox with bounded duration and memory limits.
  • CLI + MCP surface: the gco mission command group (including the chained gco mission run that scaffolds criteria and drives a session to completion in one call, and gco mission memory for the session-memory index) with matching MCP tools, plus three mission://sessions/{id} resource templates.

AWS Services in this Guidance

GCO is published in the AWS Solutions Library as the Guidance for EKS AutoMode Clusters with Global Capacity Orchestrator on AWS. It is built from 31 AWS services.

Every AWS service GCO uses, and what for
AWS Service Usage
Amazon API Gateway IAM-authenticated (SigV4) REST entry point for job submission and inference
Amazon Athena Cross-region cost analytics — a KMS-enforced workgroup queried by gco costs k8s
Amazon Aurora Optional Serverless v2 PostgreSQL with pgvector for RAG and semantic search
Amazon Bedrock Multi-engine Autopilot (gco autopilot for Claude Code, --engine codex for OpenAI Codex, or --engine opencode for OpenCode), the optional AI capacity advisor (gco capacity ai-recommend / predict), and Mission strategy sampling
Amazon CloudWatch Metrics, logs, alarms, dashboards, Container Insights for GPU utilization, and Transaction Search for the API services' trace spans
Amazon Cognito Optional user pool authenticating analytics users to presigned Studio sessions
Amazon DynamoDB Inference endpoint desired-state store, job queue state, and template storage
Amazon EC2 Accelerated instance fleet plus the capacity APIs behind gco capacity — spot placement scores, spot price history, On-Demand Capacity Reservations, and Capacity Blocks for ML
Amazon ECR Container image registry with cross-region replication for platform and user images
Amazon EFS Shared elastic storage for job outputs, model weights, and inter-pod data sharing
Amazon EKS Kubernetes control plane and Auto Mode compute (GPU, Trainium, Inferentia, CPU nodepools); optional EKS Capabilities — the AWS-managed ACK and kro — attached per cluster
Amazon ElastiCache (Valkey) Optional serverless key-value cache for prompt caching and session state
Amazon EMR Serverless Optional Spark application paired with the Studio domain for large-scale notebook analytics
Amazon FSx for Lustre Optional high-performance parallel file system for ML training workloads
Amazon S3 Model weight storage (KMS-encrypted), cluster shared bucket, CDK asset staging
Amazon SageMaker AI Optional Studio domain for interactive notebook analytics (gco analytics enable)
Amazon SNS Alert notifications for drift detection, health issues, and capacity events
Amazon SQS Regional job ingestion queue with dead-letter queue and KEDA-driven scale-to-zero consumer
Amazon VPC Network isolation with public/private subnets, NAT Gateways, and VPC endpoints
AWS CDK Infrastructure as code — synthesizes, validates (cdk-nag), and deploys all stacks
AWS Certificate Manager Stable regional certificate ARNs; rotating deployment-local ALB leaf certificates are reimported into them
AWS Cost Explorer Cost tracking by service, region, and workload via the gco costs commands
AWS Global Accelerator Anycast endpoint with health-based cross-region routing and automatic failover
AWS Glue Data Catalog database and table (partition projection) over the Parquet cost reports — no crawlers or scheduled repair jobs
AWS IAM IRSA roles for pod-level AWS access, service roles, the per-type roles the optional EKS Capabilities assume, and SigV4 authentication
AWS KMS Encryption keys for S3 model buckets, EFS, application secrets, and the backend TLS root secret
AWS Lambda HMAC-signing proxy functions, Global Accelerator registration, manifest application, and Helm chart installation orchestration
AWS Secrets Manager Rotating HMAC signing key plus the KMS-encrypted deployment-local TLS root state
AWS Step Functions Orchestrates Helm chart installs — one state per chart with per-chart retry and backoff
AWS X-Ray OpenTelemetry traces of the four API services (OTLP endpoint, stored through CloudWatch Transaction Search) and active tracing for the Lambda functions
Elastic Load Balancing Internal Application Load Balancers provisioned from the shared Gateway API resources; terminate deployment-local private-root TLS

Sample Cost Table

A single-Region deployment with default settings carries roughly $210 a month of fixed platform cost before any workload runs. GPU instances dominate real spend and scale with the hours they run (one on-demand g5.xlarge around the clock is about $734 a month in us-east-1), and multi-Region deployments scale linearly. gco costs summary tracks what you actually spend.

Itemized monthly estimate, US East (N. Virginia) pricing

The following estimates are for a single-region deployment with default settings. Multi-region deployments scale linearly. Costs vary by region, instance type, and utilization.

Resource Configuration Estimated Monthly Cost (USD)
EKS cluster 1 cluster (Auto Mode) ~$73
NAT Gateways 2 (high availability) ~$65
Application Load Balancer 1 (shared by all services) ~$22
Global Accelerator 1 accelerator + data transfer ~$18 + transfer
Lambda functions 17 functions, minimal invocations < $1 (often $0 within free tier)
Step Functions ~10 state transitions per deploy < $1
DynamoDB On-demand, low throughput ~$5
SQS Standard queue, low message volume < $1
S3 Model storage (varies with model size) ~$2 (10 GB + API requests)
EFS Elastic storage (varies with usage) ~$3 (10 GB stored)
CloudWatch Logs, metrics, Container Insights ~$15
ECR Image storage + replication ~$5
Secrets Manager 2 secrets with managed lifecycle (HMAC key and backend TLS root) < $2
Subtotal (platform, no GPU workloads) ~$210/month
GPU instances (example) 1× g5.xlarge on-demand, 24/7 (us-east-1) ~$734
GPU instances (spot) 1× g5.xlarge spot, 24/7 (us-east-1) ~$250

Notes:

  • Platform costs (~$210/month) are fixed regardless of workload volume.
  • GPU costs dominate and scale with the number of instances and hours run. Use gco costs summary to track actual spend.
  • GPU estimates assume an on-demand g5.xlarge in us-east-1 at ~$1.006/hr (~$734/month over 730 hours); rates vary by region and instance type.
  • Optional services (FSx, Valkey, Aurora, EKS Capabilities — billed per capability-hour) add additional cost depending on configuration.
  • The cost table above uses US East (N. Virginia) pricing as of June 2025.

Supported AWS Regions

GCO deploys to any AWS Region in the aws, aws-cn or GovCloud partitions, as many as you list in cdk.json under deployment_regions.regional. GPU availability varies by Region, so gco capacity recommend-region --gpu is a good first stop.

Adding a Region
// cdk.json
{
  "context": {
    "deployment_regions": {
      "regional": ["us-east-1", "eu-west-1", "ap-northeast-1"]
    }
  }
}

Then redeploy: gco stacks deploy-all -y. CDK bootstrap runs automatically for new regions.

GPU instance availability varies by region. Use gco capacity check -i <instance-type> -r <region> or gco capacity recommend-region --gpu to find regions with available GPU capacity before deploying workloads.

A regional stack can be deployed to any CloudFormation Region known to the installed AWS SDK. Add or remove Regions in deployment_regions.regional; all configured Regions must belong to one AWS partition, and GCO imposes no count limit.

Documentation

Your Goal Read This
Deploy GCO and run a first job Quick Start Guide
Understand what GCO does and the ideas behind it Core Concepts
Learn Kubernetes and GCO along a guided path Learning Path
See how the system fits together Architecture Details
Read the official AWS Solutions Library overview and reference architecture AWS Solutions Guidance
Let Claude Code, OpenAI Codex, or OpenCode drive GCO from your terminal Autopilot Guide
Understand the project's north star and decision priorities Project Tenets
Browse every guide in one place Documentation Index

Prefer a website? The project wiki
is a short orientation site — what GCO is, how it works, what you can run, and
where to go deeper — published from this repository with the live coverage
reports for
Python,
Bash and
Node.js
embedded.

Day-to-day operations
Your Goal Read This
CLI commands and usage CLI Reference
Deploy inference endpoints Inference Guide
Run multi-node distributed training Distributed Training Guide
Use the REST API directly API Reference
Fix issues Troubleshooting
Respond to incidents Operational Runbooks
Track and analyze workload cost Cost Monitoring Guide
Run interactive notebook analytics Analytics Guide
Drive a goal-directed iteration loop Mission Guide
Perform routine maintenance & upgrades Maintenance Guide
Customization and development
Your Goal Read This
Add regions, tune nodepools, enable FSx Customization Guide
Reconcile tenant workloads from Git with Argo CD GitOps with Argo CD
Publish your own platform APIs with Crossplane Crossplane
Attach AWS-managed ACK or kro EKS Capabilities
Choose a scheduler for your workload Schedulers & Orchestrators
Mirror Docker Hub images into ECR (Volcano) Image Mirror
Configure the SQS queue processor Queue Processor Config
Contribute to the project Contributing
Take GCO into your own repository Forking Guide
API client examples (Python, curl, AWS CLI) Client Examples
IAM policy templates IAM Policies
Demo walkthroughs, recordings, and re-record scripts Demo Starter Kit

Project Tenets

GCO is guided by the prioritized project tenets, beginning with workload,
data, and account safety and anchored by the north star One API. Every Accelerator.
Any Region.
The tenets define how the project resolves trade-offs across truthful
state, security, regional behavior, accelerator policy, automation, recovery,
operations, cost, and maintainability. Earlier tenets outrank later ones; durable
exceptions require an Architecture Decision Record.

Project Structure

Repository layout
.
├── app.py                               # CDK app entry point
├── TENETS.md                            # Prioritized project principles and north-star guidance
├── cdk.json                             # CDK configuration (regions, features, thresholds)
├── pyproject.toml                       # Project metadata, dependencies, and CLI installation
│
├── cli/                                 # GCO CLI (jobs, stacks, capacity, inference, costs, DAGs)
├── demo/                                # Recorded CLI demos (GIFs + asciinema sources) with walkthroughs and re-record scripts
├── diagrams/                            # Auto-generated architecture diagrams (infra_diagrams/), code flowcharts (code_diagrams/), API spec sheets (api_specs/)
├── dockerfiles/                         # Distroless container images for the in-cluster GCO services
├── docs/                                # Documentation (architecture, CLI, API, inference, customization, analytics)
├── examples/                            # Example manifests (jobs, inference, Ray, Volcano, Kueue, Slurm, YuniKorn)
├── gco/
│   ├── config/                          # Configuration loader with validation
│   ├── models/                          # Data models for k8s clusters, health monitor, inference monitor and manifest processor
│   ├── services/                        # K8s services (health/inference monitors, inference proxy, manifest/queue processors)
│   └── stacks/                          # CDK stacks (global, API gateway, regional, regional API gateway, monitoring, analytics)
│       └── constants.py                 # Pinned versions: EKS addons, Lambda runtime, Aurora engine
│
├── lambda/                              # Lambda functions
│   ├── analytics-cleanup/               # Custom resource that deletes Studio user profiles + EFS access points on stack destroy
│   ├── analytics-presigned-url/         # Generates presigned SageMaker Studio URLs for Cognito-authenticated users
│   ├── api-gateway-proxy/               # API Gateway → Global Accelerator proxy
│   ├── capacity-poller/                 # Scheduled EC2 capacity snapshots for the historical-capacity surface
│   ├── cross-region-aggregator/         # Cross-region job/health aggregation
│   ├── drift-detection/                 # Scheduled drift checks against deployed CDK stacks
│   ├── ga-registration/                 # Global Accelerator endpoint registration
│   ├── helm-installer/                  # Installs Helm charts (schedulers, cert-manager)
│   │   └── charts.yaml                  # Helm chart configuration (schedulers, cert-manager)
│   ├── helm-orchestrator/               # Custom-resource provider that starts and polls the Helm-install state machine
│   ├── image-lookup/                    # Adopt-or-create custom resource for the project's gco/* ECR repositories
│   ├── inference-streaming-proxy/       # Node.js response-streaming proxy for global and regional inference routes
│   ├── kubectl-applier-simple/          # Applies K8s manifests during deployment
│   │   └── manifests/                   # Kubernetes manifests (nodepools, RBAC, services, storage)
│   ├── proxy-shared/                    # Shared utilities for proxy Lambdas
│   ├── regional-api-proxy/              # Regional API Gateway → internal ALB proxy
│   ├── secret-rotation/                 # Daily secret rotation
│   ├── tls-certificate-manager/         # Private TLS root, regional ACM leaves, and trust publication
│   ├── tls-shared/                      # Strict private-root TLS client shared by the proxy Lambdas
│   ├── traffic-dial-controller/         # Health-driven Global Accelerator traffic dials
│   └── vector-ingest/                   # S3-triggered chunking + Bedrock embeddings for the vector store
│
├── gco_mcp/                             # MCP server for LLM interaction (142 tools default, up to 200 with feature flags)
├── images/                              # Screenshots and visual assets for docs and the wiki
├── scripts/                             # Utility scripts (version bump, cluster access setup)
├── tests/                               # PyTest + BATS test suites
├── wiki/                                # Orientation wiki sources, built with Zensical and published to GitHub Pages (wiki/README.md explains how)
└── zensical.toml                        # Zensical configuration for the GitHub Pages wiki (sources in wiki/, staged by scripts/build_wiki.py)

Contributing

See CONTRIBUTING.md for development setup, testing, the GitHub Actions CI/CD layout, the release process, and dependency scanning schedules. The whole test-suite runs from the same dev container scripts/setup-dev-alias.sh builds:

docker run --rm -v $(pwd):/workspace -w /workspace gco-dev pytest tests/ -v

Support

  • Check Troubleshooting for common issues
  • Review CloudWatch logs for Lambda and EKS errors
  • Open an issue on GitHub

Security

GCO implements defense in depth across the six controls of its Security Model, checks its infrastructure against five cdk-nag rule packs, and scans container images and dependencies in CI.

Reporting a vulnerability: do not open a public GitHub issue; follow the responsible disclosure process in .github/SECURITY.md.

Security controls in depth

Authentication and Authorization:

  • All API requests require AWS IAM (SigV4) authentication at the API Gateway
  • The trusted proxy Lambda adds a request-bound HMAC envelope using a rotating Secrets Manager key; the reusable key is never sent downstream
  • IRSA (IAM Roles for Service Accounts) provides pod-level AWS access with no static credentials
  • EKS access entries with explicit policy bindings (no aws-auth ConfigMap)

Network Security:

  • Regional platform ALBs are internal; the EKS API endpoint defaults to PRIVATE
  • EKS clusters run in private subnets with configurable endpoint access (PRIVATE or PUBLIC_AND_PRIVATE)
  • VPC gateway endpoints for S3 and DynamoDB keep that traffic off the public internet by default; opt-in interface endpoints (vpc_endpoints.interface in cdk.json) do the same for ECR, STS, SSM, CloudWatch, and SQS
  • VPC Flow Logs (30-day retention) capture all network traffic for audit
  • Kubernetes Network Policies enforce default-deny with explicit allow rules

Encryption:

  • Data at rest: S3 (KMS), EFS (KMS), EBS (KMS), DynamoDB (AWS-managed), and Secrets Manager (KMS)
  • Client-to-global/regional API traffic uses AWS-managed TLS and IAM SigV4. Cross-region aggregation also uses AWS-managed TLS plus SigV4 from the aggregator to each regional API bridge.
  • The normal global proxy → Global Accelerator → ALB path and each regional VPC proxy → ALB path use authenticated private-root TLS on TCP/HTTPS 443. Global Accelerator is a Layer 4 pass-through and does not terminate TLS.
  • Every ALB leaf is issued for backend.<project>.gco.internal; clients send that identity through SNI and assert it while connecting to dynamic accelerator or ALB DNS names. The ALB then re-encrypts to TLS-only pod proxy sidecars on port 8443; decrypted bytes travel only over pod loopback to the application process.
  • The root private key exists only in a customer-managed-KMS-encrypted Secrets Manager secret readable by the certificate-manager role. Proxy roles read only the public SSM trust bundle; rotating leaves are reimported into stable regional ACM ARNs.
  • The request-bound HMAC envelope adds integrity, freshness, and replay defense; it is not encryption and is independent of TLS confidentiality and server authentication.
  • EFS mount encryption in transit is enabled by the deployed storage configuration.
  • Kubernetes secrets are encrypted in etcd (EKS-managed encryption).

Infrastructure Policy Validation:

Supply Chain Security:

  • Container images scanned with Trivy on every push (CVE detection)
  • Python dependencies audited with pip-audit (GHSA/CVE detection)
  • Both repository-owned npm graphs are exact-pinned with committed lockfiles (see package.json and package-lock.json), audited on every PR, and updated by Dependabot
  • Production JavaScript is scanned by CodeQL and Semgrep; the inference-streaming Lambda has a separate Node.js 24 test workflow with exact 100% line/function/branch gates (see ./tests/inference-streaming-proxy/)
  • Every Python dependency, direct and transitive, pinned to an exact version in requirements-lock.txt (see requirements-lock.txt); CI fails when the lock drifts from pyproject.toml
  • Dependabot and CodeQL enabled for automated vulnerability alerts
  • Strict KICS and Checkov infrastructure scans
  • SBOM generation via Trivy for all container images

License

See the LICENSE file for details.

Reviews (0)

No results found