global-capacity-orchestrator-on-aws
Health Gecti
- License — License: MIT-0
- Description — Repository has a description
- Active repo — Last push 0 days ago
- Community trust — 49 GitHub stars
Code Basarisiz
- rm -rf — Recursive force deletion command in .github/actions/apt-install-with-retry/action.yml
- rm -rf — Recursive force deletion command in .github/actions/free-disk-space/action.yml
Permissions Gecti
- Permissions — No dangerous permissions requested
Bu listing icin henuz AI raporu yok.
AI/ML and HPC at planetary scale. One API. Every Accelerator. Any Region. MIT-0 licensed
Global Capacity Orchestrator
One API. Every Accelerator. Any Region.
Quick Start · Documentation · Wiki · Examples · CLI Reference · Contributing
Global Capacity Orchestrator (GCO) runs accelerated workloads — LLM training and inference, batch ML, HPC — on EKS Auto Mode clusters in as many AWS Regions as you configure, behind one IAM-authenticated API, CLI and MCP server. It finds where NVIDIA GPU, Trainium, Inferentia and CPU capacity actually is, places jobs there, and serves inference endpoints with automatic cross-Region failover.

A real gco session, reproducible from demo/live_demo.sh: fleet status, capacity discovery, four schedulers plus KEDA, shared storage, a globally replicated vector store and live LLM inference. More recordings are under See it running.
Why GCO?
Running accelerated workloads across Regions means finding capacity, standing up clusters, wiring authentication, handling failover, and keeping job outputs after the pods exit. GCO packages all of it as one platform you deploy with a single command:
- Capacity-aware placement.
gco capacityreads spot placement scores, spot price history, On-Demand Capacity Reservations and Capacity Blocks for ML, and auto-Region submission builds on them. - One authenticated front door. A SigV4 API, the
gcoCLI and an MCP server reach every Region with the AWS credentials you already have, with no kubeconfig distribution. - Accelerators on demand. EKS Auto Mode NodePools for NVIDIA GPU (x86 and Arm), EFA, Trainium, Inferentia and CPU launch nodes only when a workload needs them.
- Inference in every Region. One command deploys vLLM, SGLang or Triton endpoints everywhere, with automatic failover through Global Accelerator in the commercial
awspartition. - A curated ecosystem. KEDA, Volcano, Kueue, KubeRay and Kubeflow Trainer are on by default and Slurm, YuniKorn, Argo CD and Crossplane are opt-in; EFS ships with every cluster, and FSx for Lustre, Valkey and Aurora pgvector are a toggle away; Prometheus, Grafana, OpenCost and MLflow are on by default for observability and cost tracking.
- Agent-ready.
gco autopilotopens Claude Code, OpenAI Codex or OpenCode on Amazon Bedrock with the GCO MCP server already wired in.
| Challenge | Traditional Approach | With GCO |
|---|---|---|
| GPU availability | Manually check each region | Capacity tools and auto-region workflows compare configured regions |
| Node provisioning | Pre-provision or wait for scaling | EKS Auto Mode provisions on-demand |
| Multi-region ops | Manage clusters separately | One platform across unlimited SDK-known Regions in one partition |
| Authentication | Configure per-cluster access | IAM-based, uses existing AWS credentials |
| Job outputs | Lost unless persisted | EFS/FSx and per-region S3 available to every job that mounts or writes them |
| Inference serving | Deploy and manage per-region | Deploy once across selected Regions; global failover in aws |
| Failover | Manual intervention required | Automatic via Global Accelerator in aws; explicit regional selection elsewhere |
When to use GCO:
- You need to run GPU workloads (training, inference, batch processing)
- You want to deploy inference endpoints across multiple regions with a single command
- You want multi-region redundancy without managing multiple clusters
- You prefer IAM authentication over kubeconfig management
- You need job outputs to persist after completion
What it does. Spins up EKS Auto Mode clusters across any number of SDK-known CloudFormation Regions in one AWS partition. In commercial aws, Global Accelerator provides latency-aware anycast routing and automatic failover behind the global workload API; other partitions (aws-cn and aws-us-gov) use IAM-authenticated regional workload APIs while retaining the aggregate global API. Capacity tools and auto-region queue/CLI workflows select a target Region, EKS Auto Mode provisions matching nodes from the built-in system and general-purpose NodePools plus project-managed GPU x86, GPU ARM, inference, EFA, Mooncake EFA, Neuron, and CPU NodePools, and shared storage persists workload outputs. Network routing never substitutes for live GPU-capacity placement.
Why it's different. Capacity-aware placement tools and auto-region workflows, partition-aware authenticated routing, full-stack observability (CloudWatch dashboards, alarms, SNS), and a CDK app validated across the full curated configuration matrix in CI. Read Core Concepts for the ideas behind it and the Learning Path if Kubernetes is new to you.
Get started
Prerequisites
- An AWS account you can create infrastructure in, with credentials the AWS CLI can use
- Git and a container runtime: Docker, Finch or Podman (Colima also works)
The dev container brings everything else — Python 3.14, Node.js 24, CDK, kubectl, the AWS CLI, and Docker CLI + Buildx — at pinned versions. A host install is possible but advanced; see More ways to run GCO below.
Deploy
git clone https://github.com/aws-solutions-library-samples/global-capacity-orchestrator-on-aws.git
cd global-capacity-orchestrator-on-aws
./scripts/setup-dev-alias.sh # builds the dev container and installs the `gco` shell function
source ~/.zshrc # or ~/.bashrc — the script prints which file it updated
gco stacks deploy-all -y # stands up every Region in cdk.json; billing starts here
deploy-all bootstraps CDK where needed, deploys the global control plane and one regional stack per Region, and then installs the cluster add-ons in the background.
Run your first workload
gco capacity check --instance-type g4dn.xlarge --region us-east-1 # where is capacity right now?
gco jobs submit-sqs examples/simple-job.yaml --region us-east-1 # submit a job over SQS
gco jobs list --all-regions # watch it run
gco jobs logs hello-gco -n gco-jobs -r us-east-1 # read its output
Serve a model from every Region with one command:
gco inference deploy my-llm -i vllm/vllm-openai:v0.31.0 --gpu-count 1 # OpenAI-compatible endpoint
gco inference status my-llm # rollout state in every Region
gco inference scale my-llm --replicas 3 # scale it out
Let an agent drive
gco autopilot # default Claude Code session
gco autopilot --engine codex # or an OpenAI Codex session
gco autopilot --engine opencode # or an OpenCode session (Kimi K3)
Every engine runs on Amazon Bedrock with your AWS credentials and the GCO MCP server plus the recommended companion MCP servers wired in, so you can simply ask: "deploy everything", "where is p5 capacity cheapest right now?" The Autopilot Guide covers sessions, opt-in tool groups and dry runs.
Clean up
gco stacks destroy-all -y # destroys every stack, then sweeps known leftovers
Teardown is not an account-wide emptiness guarantee: retained ECR repositories, resources configured for retention, and unexpected resources can remain. See gco stacks destroy-all for the exact cleanup scope.
The dev container. GCO pins exact versions of a lot of Python packages (CDK, AWS SDKs, FastAPI, mypy, Ruff, etc.), and installing them on top of an existing Python environment is the most common source of "it doesn't install" reports. The dev container ships a fully resolved environment so you skip the whole problem, and scripts/setup-dev-alias.sh means you never hand-write a docker run … or live inside an interactive container shell:
./scripts/setup-dev-alias.sh # builds gco-dev from Dockerfile.dev + installs the `gco` shell function
source ~/.zshrc # or ~/.bashrc — the script prints which file it updated
gco --help # every command now runs inside the container, against your checkout
The script detects your container runtime, builds the gco-dev image from Dockerfile.dev (re-running rebuilds it, so a stale image is refreshed automatically — pass --no-build to skip), wires up the correct socket pass-through, and installs an idempotent gco shell function (not a bare alias, so arguments and pipes forward correctly and a TTY is attached only when one is present) into your profile. Re-run it whenever you switch runtimes, or use --print to preview the function, --runtime <name> to force one, and --rc <path> to target a specific profile. To move an existing deployment to a newer release later, gco upgrade checks out the release tag, rebuilds this image, and cycles the stacks in one pass.
The function shares your host Docker socket with every gco call because gco stacks deploy-all needs it: through your host daemon it builds and bundles the Lambda function assets (as linux/amd64, cross-built via Buildx so this works on Apple Silicon / arm64) and — when the Volcano image mirror is enabled in cdk.json — mirrors third-party images from Docker Hub into your ECR before the Helm install runs. The same socket backs gco images build / push. This is host-socket pass-through, not Docker-in-Docker: anyone with access to the container has root-equivalent access to the host Docker daemon, so keep the container on a trusted host. Finch users: Finch runs in its own VM with no host Docker socket to share, so the function omits the socket mount — everyday commands work as-is, while build-heavy ones like deploy-all run on the host with Finch as the CDK builder.
An interactive container shell. Build the image yourself and drop into it, running gco from inside — handy for ad-hoc tools and exploration. The -v /var/run/docker.sock:/var/run/docker.sock mount gives the container's Docker CLI access to your host daemon for the asset builds and image mirroring described above (Colima users: see the header of Dockerfile.dev for the socket path):
docker build -f Dockerfile.dev -t gco-dev .
docker run -it --rm \
-v ~/.aws:/root/.aws:ro \
-v $(pwd):/workspace \
-v /var/run/docker.sock:/var/run/docker.sock \
-w /workspace \
gco-dev
A host install (advanced). Host installs are the advanced, non-recommended path: GCO's exact pins frequently fail to resolve on top of an existing Python environment (ResolutionImpossible). If you still want one, use a clean virtual environment or pipx — the Quick Start has both recipes:
pipx install -e .
A host install additionally needs:
- Python 3.14+ and Node.js 24 (use
.nvmrc) - npm 12.2.0 and the repository's locked tooling graph: run
npm ci --ignore-scripts --no-audit --no-fundat the repository root;gcoprefers its localnode_modules/.bin/cdkover a global CLI - A clean Python virtual environment or pipx — GCO pins exact versions of many packages, so installing into an existing environment commonly fails with dependency-resolver errors. If you hit
ResolutionImpossible, switch to the dev container instead of debugging your local environment.
Heads up — Helm charts finish installing in the background. When
deploy-allreports the clusterCREATE_COMPLETE, the scheduler/operator Helm charts (KEDA, Volcano, KubeRay, cert-manager, Kueue, …) have only been kicked off; they converge asynchronously and can take 10–30+ minutes to all become ready. This is intentional — a slow chart never rolls back the cluster. Track progress withgco stacks addons status -r <region>and re-converge any failures withgco stacks addons install -r <region>. See docs/CUSTOMIZATION.md.
Every way to submit a job: SQS, the global queue, the REST API, or kubectlOptional: kubectl access. The EKS API endpoint is
PRIVATEby default and most users never need kubectl — submit jobs through SQS or API Gateway instead. When you do,gco cluster tunnel --via-ssm autoreaches the private endpoint from your laptop over SSM andgco stacks access -r <region>grants your IAM principal an EKS access entry;gco cluster doctorchecks both. See docs/CUSTOMIZATION.md.
Submit a job using whichever path fits your setup — via SQS (recommended), via the global DynamoDB queue, via API Gateway, or directly through kubectl:
gco jobs submit-sqs examples/simple-job.yaml --region us-east-1
gco queue submit examples/simple-job.yaml --region us-east-1
gco jobs submit examples/simple-job.yaml -n gco-jobs
gco jobs submit-direct examples/simple-job.yaml -r us-east-1
The Quick Start Guide walks through every step with the output to expect, and the CLI Reference lists every command.
See it running
Real terminal recordings, not mock-ups. Each one is reproducible from the script linked beneath it; the live demo itself plays at the top of this page.
📦 Deploy recording
Fresh gco stacks deploy-all -y --enable fsx_lustre,valkey,aurora_pgvector,vector_store,slurm,yunikorn from a clean account — the six add-ons ship disabled in cdk.json because each bills continuously, so the recording enables them for one run through run-scoped overrides (re-record)

Full teardown with gco stacks destroy-all -y --enable fsx_lustre,valkey,aurora_pgvector,vector_store,slurm,yunikorn, repeating the deploy's overrides so it evaluates the same app (re-record)

A real session: gco autopilot launches Claude Code on Amazon Bedrock (GCO's default Claude Opus profile) with the GCO MCP server + companion MCPs wired in, and it answers from the project's own MCP tools (docs · re-record).

A real Bedrock-backed gco autopilot --engine codex --no-companions session using a least-privilege recording profile: required GCO MCP, only find_docs/read_resource, no built-in shell, and no trust or approval prompts. A normal gco autopilot --engine codex launch includes the recommended companions (docs · re-record with DEMO_ENGINE=codex DEMO_MODE=live).

A real Bedrock-backed gco autopilot --engine opencode --no-companions session (Moonshot AI's Kimi K3) using a least-privilege recording profile: only the GCO MCP find_docs/read_resource tools, every built-in OpenCode tool denied, and no permission prompts. A normal gco autopilot --engine opencode launch includes the recommended companions and OpenCode's ordinary tool set behind an ask-first permission floor (docs · re-record with DEMO_ENGINE=opencode DEMO_MODE=live).
Install the MCP server
GCO ships with the GCO MCP server — an MCP server exposing 142 tools by default (up to 200 with feature flags) that index the whole project: docs, examples, source code, K8s manifests, and scripts. Connect it to an AI-powered IDE with MCP support and explore GCO conversationally — "How does region recommendation work?", "Walk me through the inference deployment flow" — or let it drive real operations. One click adds it (pinned to the latest release, no clone needed) to your client; uv and AWS credentials are the only prerequisites. Other clients and options: setup guide.
| Kiro | Cursor | VS Code |
|---|---|---|
Architecture Overview
Figure 1: GCO multi-region reference architecture — global control plane and workload entry
One CDK app deploys a global control plane (the API Gateway entry point, Global Accelerator in aws, shared DynamoDB state, the model bucket and monitoring) and one regional stack per Region: an EKS Auto Mode cluster with GCO's platform services, NodePools, storage and job queue. Architecture Details is the full deep dive, and the AWS Solutions Guidance presents the reference architecture in the AWS Solutions Library.
The generated reference architecture shows the commercial aws workload path. Other partitions retain the global aggregate API but route workload control and inference through each Region's IAM-authenticated bridge.
- DevOps / Platform engineers own the deployment. They configure the platform through cdk.json and drive everything from the
gcoCLI. - The AWS CDK app synthesises and deploys the GCO stacks with a single
gco stacks deploy-all, provisioning the global control plane and one regional stack per target region. - Users submit jobs and inference requests through the
gcoCLI, which signs every call with AWS SigV4 credentials. - In commercial
aws, Amazon API Gateway is edge-optimized and is the global workload and aggregate entry point. In other partitions it is regional and aggregate-only. Every exposed method enforces IAM (SigV4) authentication before integration. - In
aws, route-specific AWS Lambda proxies sign workload requests with a short-lived HMAC envelope derived from a rotating AWS Secrets Manager key;/api/v1/*stays buffered while/inference/*streams. Other partitions omit these global workload proxies and use equivalent VPC proxies behind the regional APIs. - In
aws, AWS Global Accelerator routes workload requests over the AWS backbone to a healthy registered Region. Other partitions create no accelerator resources. - A regional internal AWS Application Load Balancer terminates deployment-local private-root TLS from either Global Accelerator (
aws) or the regional VPC proxy, then re-encrypts to TLS-only proxy sidecars on the platform API pods. Each sidecar hot-reloads its projected certificate and forwards decrypted traffic only over pod loopback. - Each region runs an Amazon EKS Auto Mode cluster with built-in
systemandgeneral-purposeNodePools plus project-managed GPU, inference, EFA, Mooncake EFA, Neuron, and CPU NodePools. Platform services include the Cost Monitor, Health Monitor, Manifest Processor, Queue Processor, Inference Monitor, and dedicated Inference Proxy.
Figure 2: GCO regional reference architecture — EKS Auto Mode data plane and regional services
- An internal Application Load Balancer created from the shared
gco-system/gco-gatewayGateway API resources accepts only HTTPS/443 with a rotating regional ACM leaf, then re-encrypts target traffic to cert-manager-backed HTTPS listeners on the cluster API services. ALB target TLS provides confidentiality; HMAC proves trusted-proxy key possession and request integrity on protected paths, while API Gateway IAM authenticates the original caller. - The Amazon EKS Auto Mode cluster is the heart of the regional stack, hosting platform services and user workloads with a private API endpoint by default.
- NodePools provision capacity on demand: built-in
systemandgeneral-purpose, plusgpu-x86-pool,gpu-arm-pool,gpu-inference-pool,gpu-efa-pool,mooncake-efa-pool,neuron-pool, andcpu-general-pool. - Workloads and platform services run across namespaces:
gco-system(Health Monitor, Manifest Processor, Queue Processor, Inference Monitor, Inference Proxy) andgco-jobs/gco-inference(training and batch jobs, inference endpoints, and job DAG pipelines). - Storage and data services back workloads: Amazon EFS, optional FSx for Lustre, optional Valkey, optional Aurora pgvector, and Amazon S3 for KMS-encrypted model weights.
- An always-deployed Regional API Gateway bridge gives the aggregator a SigV4-authenticated path to the VPC Lambda and internal ALB. Direct same-account access is optional through
regional_api_enabledinawsand enabled automatically as the required workload ingress elsewhere. - Regional AWS services complete the stack: Amazon SQS for job ingestion, DynamoDB-backed state where applicable, and Amazon CloudWatch metrics and logs.

Figure 3: Generated CDK architecture for the global control plane and regional EKS data planes
Regenerate the full architecture and every per-stack view with python diagrams/infra_diagrams/generate.py. The generator synthesizes the current CDK app through cdk-dia so the committed diagrams track the deployed resource graph. See diagrams/infra_diagrams/README.md for per-stack flags (--stack global|api-gateway|regional|regional-api|monitoring|analytics|all).
Flowcharts of Lambda handlers, CLI commands, stack constructors, and MCP control paths live under diagrams/code_diagrams/. Regenerate them through the canonical two-commit workflow, which records an exact source commit without creating a self-referential SHA. Add newly charted functions to diagrams/code_diagrams/_targets.py.
The HTTP API surface has its own catalogue: diagrams/api_specs/ holds one spec sheet per surface — the two AWS API Gateways, the in-cluster Gateway (the internal ALB) and the four FastAPI services — with endpoint tables, parameters, request bodies, responses, component schemas and, for the gateways, each route's Lambda backend and the hops to the Service that answers; all rendered from generated OpenAPI documents (docs/openapi/: the services' own exports, the API Gateway stacks read at CDK synthesis, and the HTTPRoute composed with the service documents). The catalogue index embeds an interaction diagram of how the gateways, the Lambda proxies and the services fit together, drawn from the same documents. The sheets are also the wiki's API reference, and a Swagger UI console for each document is published at /swagger/. Regenerate with python diagrams/generate.py --api-only after refreshing the documents (python scripts/generate_openapi.py, python scripts/generate_api_gateway_openapi.py, python scripts/generate_cluster_gateway_openapi.py); python diagrams/generate.py --check fails when a route, rule or model changed without the sheets.
Security Model
Six complementary controls protect every backend request: IAM authentication at API Gateway, TLS trust separation, a request-bound HMAC, private backend exposure, freshness and integrity validation in the backend middleware, and IRSA / EKS Pod Identity for pod-level AWS access.
The six controls in detailSix complementary controls protect backend requests:
- IAM authentication — API Gateway validates AWS credentials with SigV4.
- TLS trust separation — API Gateway uses AWS-managed TLS; proxy-to-ALB traffic uses a deployment-local private root and explicit
backend.<project>.gco.internalSNI/hostname verification; the in-cluster hops behind the ALB (cost monitor, OpenCost, model endpoints, metrics scrapes) use a cluster-local cert-manager CA that every client verifies against (In-cluster TLS). - Request-bound HMAC — a trusted Lambda signs the version, timestamp, nonce, method, exact target, and body digest with a rotating key that is never transmitted. HMAC provides integrity/freshness/replay defense, not encryption.
- Private backend exposure — regional ALBs are internal and the EKS API endpoint is private by default.
- Freshness and integrity validation — backend middleware rejects stale, altered, or replayed envelopes.
- IRSA / EKS Pod Identity — pods receive scoped AWS permissions without static workload credentials.
Figure 4: GCO security model — layered controls and the authenticated request flow
Commercial `aws` request flow:
User → API Gateway (AWS TLS + SigV4) → Lambda (HMAC)
→ Global Accelerator (TCP/443 pass-through) → internal ALB (private-root TLS)
→ pod TLS proxy (re-encrypted HTTPS) → AuthenticationMiddleware (pod loopback)
Other partitions:
User → Regional API Gateway (AWS TLS + SigV4) → VPC Lambda (HMAC)
→ internal ALB (private-root TLS)
→ pod TLS proxy (re-encrypted HTTPS) → AuthenticationMiddleware (pod loopback)
Every Region's API bridge is required for aggregator fan-out. Direct regional
access
is optional for same-account callers in aws and enabled automatically as the
supported workload ingress in other partitions.
Key Features
Everything GCO does, grouped by area; expand one for the details and its guide.
Compute and orchestration: NodePools for every accelerator, four submission paths, distributed training, pipelines and Helm-managed schedulers- EKS Auto Mode with automatic node provisioning — no pre-scaling needed
- GPU and accelerator support through
gpu-x86-pool,gpu-arm-pool,gpu-inference-pool,gpu-efa-pool,mooncake-efa-pool, andneuron-pool, plus built-in and project-scoped CPU pools. Only families on the EKS Auto Mode supported instance list can launch at any one time; the pools deliberately list newer families (such asg7/g7e) ahead of that support so they come online without a GCO release — see which instance types can actually launch - Multiple submission methods: API Gateway, SQS queues, DynamoDB job queue, or direct kubectl
- Distributed training via Kubeflow Trainer v2 (on by default): multi-node PyTorch through the
TrainJobAPI against platform-shipped runtimes, validated end to end by the same security pipeline as every other submission, with optional Kueue gang admission — see the Distributed Training Guide - Job pipelines (DAGs): Multi-step ML pipelines with dependency ordering and failure handling
- Helm-managed ecosystem: mandatory KEDA; EFA and Neuron device plugins; Volcano, KubeRay, Kubeflow Trainer, cert-manager, kube-prometheus-stack (on by default with cluster observability), and Kueue; opt-in Slurm/Slinky, YuniKorn, Argo CD and Crossplane with its Crossview dashboard
- Multi-region inference: Deploy endpoints across regions with a single command. Each supported framework ships with a ready-to-run example manifest: vLLM (example), SGLang (example), and Triton (example)
- Canary deployments: A/B test new model versions with weighted traffic routing
- Model weight management: Central S3 bucket with KMS encryption, automatic sync to each region
- Spot instance support: Run inference on spot GPUs for significant cost savings
- Autoscaling: HPA-based scaling with CPU/memory metrics
- Global Accelerator: Single anycast endpoint with automatic failover
- IAM authentication: SigV4 at the API Gateway — no kubeconfig distribution
- Infrastructure policy validation: cdk-nag v3 rule packs for AWS Solutions, HIPAA, NIST 800-53, PCI DSS, and Serverless findings (these checks are not certifications)
- Network policies: Default-deny with explicit allow rules for all service communication
- Verified in-cluster HTTPS: every hop GCO owns inside a cluster (cost monitor, OpenCost, model endpoints, prefill/decode, Grafana's admin API, metrics scrapes) is TLS from a sidecar with a cert-manager leaf, verified by the client against one cluster-local CA — see In-cluster TLS
- EFA support: Optional Elastic Fabric Adapter for high-bandwidth distributed training and NIXL-based inference (toggle on/off)
- EFS: Shared elastic storage for job outputs that persist after pod termination
- FSx for Lustre: Optional high-performance parallel file system for ML training (toggle on/off)
- Valkey cache: Optional serverless key-value cache for prompt caching and session state
- Aurora pgvector: Optional serverless vector database for RAG, semantic search, and embedding storage
- Vector store: Optional globally replicated DynamoDB vector index over an S3-ingested document corpus — drop files in, search from every region (
gco vector)
- Multi-engine Autopilot: launch Claude Code by default with
gco autopilot, OpenAI Codex withgco autopilot --engine codex, or OpenCode withgco autopilot --engine opencode; every engine uses Amazon Bedrock with the GCO MCP server and recommended companion MCPs preconfigured - Cost visibility: Track spend by service, region, and workload via Cost Explorer integration
- Cost monitoring & analytics (on by default): per-cluster OpenCost with a Grafana cost dashboard, scheduled Parquet cost reports to a central S3 bucket, and cross-region Athena analytics via
gco costs k8s— see Cost Monitoring Guide - Spot price-aware scheduling: central-queue jobs can set a max spot price per instance type and dispatch only when the market clears it
- MLflow experiment tracking (on by default with observability): an in-cluster MLflow tracking server per region — run artifacts to S3 via a prefix-scoped IAM role, metadata on EBS, reached with
gco monitoring open --service mlflow— see MONITORING.md - Auto-bootstrap: CDK bootstrap runs automatically for new regions during deploy
- Multi-region monitoring: the
gco-monitoringstack's cross-region CloudWatch dashboards, alarms, and SNS alerts, complemented by per-cluster Prometheus/Grafana cluster observability - Distributed tracing (on by default, 5% sampled): the four API services export OpenTelemetry spans straight to AWS X-Ray with no collector, searchable in CloudWatch Transaction Search, and every service log line carries the matching trace id — see Distributed tracing
- GitOps with Argo CD (off by default): a self-managed Argo CD per regional cluster, installed from the upstream chart in namespaced mode and fenced to the job namespaces by a
gco-tenantsAppProjectand matching RBAC; point it at a repository path per cluster fromcdk.json, let the repo server autoscale, and open the UI over the private endpoint withgco gitops open— see GitOps with Argo CD - Crossplane (off by default): a self-managed Crossplane v2 and the Crossview dashboard, whose namespaced composite resources compose tenant workloads in the job namespaces only (
gco crossplane open) — see Crossplane - EKS Capabilities (off by default): attach the AWS-managed ACK and kro capabilities per regional cluster from
cdk.json, with IAM roles that carry only the permissions you configure and the tenant RBAC kro composes with;gco stacks capabilities statusreports drift — see EKS Capabilities
- ML & Analytics Environment: Optional SageMaker Studio domain + EMR Serverless + Cognito user pool for interactive notebook analytics, with an always-on
Cluster_Shared_Bucketthat all cluster jobs can read and write. Off by default — enable withgco analytics enable. See Analytics Guide.
Goal-directed iteration loop for orchestrated workflows. The operator declares a natural-language directive plus machine-checkable success criteria, a tool allowlist, and a budget; Mission runs five-phase iterations (propose → execute → observe → evaluate → decide) until a verdict is reached. Off by default — enable with GCO_ENABLE_MISSION=true. See Mission Guide.
- Deterministic verdict cascade with optional advisory LLM sampling (MCP host or Amazon Bedrock). Sampling shapes only the next strategy; it never moves the verdict.
- Budget caps on iterations and wall clock — the engine terminates cleanly when any cap fires. Cost guardrails live out-of-band via AWS Budgets and Cost Anomaly Detection at the account level.
- Scripted strategies opt-in: an AST-validated Python sandbox with bounded duration and memory limits.
- CLI + MCP surface: the
gco missioncommand group (including the chainedgco mission runthat scaffolds criteria and drives a session to completion in one call, andgco mission memoryfor the session-memory index) with matching MCP tools, plus threemission://sessions/{id}resource templates.
AWS Services in this Guidance
GCO is published in the AWS Solutions Library as the Guidance for EKS AutoMode Clusters with Global Capacity Orchestrator on AWS. It is built from 31 AWS services.
Every AWS service GCO uses, and what for| AWS Service | Usage |
|---|---|
| Amazon API Gateway | IAM-authenticated (SigV4) REST entry point for job submission and inference |
| Amazon Athena | Cross-region cost analytics — a KMS-enforced workgroup queried by gco costs k8s |
| Amazon Aurora | Optional Serverless v2 PostgreSQL with pgvector for RAG and semantic search |
| Amazon Bedrock | Multi-engine Autopilot (gco autopilot for Claude Code, --engine codex for OpenAI Codex, or --engine opencode for OpenCode), the optional AI capacity advisor (gco capacity ai-recommend / predict), and Mission strategy sampling |
| Amazon CloudWatch | Metrics, logs, alarms, dashboards, Container Insights for GPU utilization, and Transaction Search for the API services' trace spans |
| Amazon Cognito | Optional user pool authenticating analytics users to presigned Studio sessions |
| Amazon DynamoDB | Inference endpoint desired-state store, job queue state, and template storage |
| Amazon EC2 | Accelerated instance fleet plus the capacity APIs behind gco capacity — spot placement scores, spot price history, On-Demand Capacity Reservations, and Capacity Blocks for ML |
| Amazon ECR | Container image registry with cross-region replication for platform and user images |
| Amazon EFS | Shared elastic storage for job outputs, model weights, and inter-pod data sharing |
| Amazon EKS | Kubernetes control plane and Auto Mode compute (GPU, Trainium, Inferentia, CPU nodepools); optional EKS Capabilities — the AWS-managed ACK and kro — attached per cluster |
| Amazon ElastiCache (Valkey) | Optional serverless key-value cache for prompt caching and session state |
| Amazon EMR Serverless | Optional Spark application paired with the Studio domain for large-scale notebook analytics |
| Amazon FSx for Lustre | Optional high-performance parallel file system for ML training workloads |
| Amazon S3 | Model weight storage (KMS-encrypted), cluster shared bucket, CDK asset staging |
| Amazon SageMaker AI | Optional Studio domain for interactive notebook analytics (gco analytics enable) |
| Amazon SNS | Alert notifications for drift detection, health issues, and capacity events |
| Amazon SQS | Regional job ingestion queue with dead-letter queue and KEDA-driven scale-to-zero consumer |
| Amazon VPC | Network isolation with public/private subnets, NAT Gateways, and VPC endpoints |
| AWS CDK | Infrastructure as code — synthesizes, validates (cdk-nag), and deploys all stacks |
| AWS Certificate Manager | Stable regional certificate ARNs; rotating deployment-local ALB leaf certificates are reimported into them |
| AWS Cost Explorer | Cost tracking by service, region, and workload via the gco costs commands |
| AWS Global Accelerator | Anycast endpoint with health-based cross-region routing and automatic failover |
| AWS Glue | Data Catalog database and table (partition projection) over the Parquet cost reports — no crawlers or scheduled repair jobs |
| AWS IAM | IRSA roles for pod-level AWS access, service roles, the per-type roles the optional EKS Capabilities assume, and SigV4 authentication |
| AWS KMS | Encryption keys for S3 model buckets, EFS, application secrets, and the backend TLS root secret |
| AWS Lambda | HMAC-signing proxy functions, Global Accelerator registration, manifest application, and Helm chart installation orchestration |
| AWS Secrets Manager | Rotating HMAC signing key plus the KMS-encrypted deployment-local TLS root state |
| AWS Step Functions | Orchestrates Helm chart installs — one state per chart with per-chart retry and backoff |
| AWS X-Ray | OpenTelemetry traces of the four API services (OTLP endpoint, stored through CloudWatch Transaction Search) and active tracing for the Lambda functions |
| Elastic Load Balancing | Internal Application Load Balancers provisioned from the shared Gateway API resources; terminate deployment-local private-root TLS |
Sample Cost Table
A single-Region deployment with default settings carries roughly $210 a month of fixed platform cost before any workload runs. GPU instances dominate real spend and scale with the hours they run (one on-demand g5.xlarge around the clock is about $734 a month in us-east-1), and multi-Region deployments scale linearly. gco costs summary tracks what you actually spend.
The following estimates are for a single-region deployment with default settings. Multi-region deployments scale linearly. Costs vary by region, instance type, and utilization.
| Resource | Configuration | Estimated Monthly Cost (USD) |
|---|---|---|
| EKS cluster | 1 cluster (Auto Mode) | ~$73 |
| NAT Gateways | 2 (high availability) | ~$65 |
| Application Load Balancer | 1 (shared by all services) | ~$22 |
| Global Accelerator | 1 accelerator + data transfer | ~$18 + transfer |
| Lambda functions | 17 functions, minimal invocations | < $1 (often $0 within free tier) |
| Step Functions | ~10 state transitions per deploy | < $1 |
| DynamoDB | On-demand, low throughput | ~$5 |
| SQS | Standard queue, low message volume | < $1 |
| S3 | Model storage (varies with model size) | ~$2 (10 GB + API requests) |
| EFS | Elastic storage (varies with usage) | ~$3 (10 GB stored) |
| CloudWatch | Logs, metrics, Container Insights | ~$15 |
| ECR | Image storage + replication | ~$5 |
| Secrets Manager | 2 secrets with managed lifecycle (HMAC key and backend TLS root) | < $2 |
| Subtotal (platform, no GPU workloads) | ~$210/month | |
| GPU instances (example) | 1× g5.xlarge on-demand, 24/7 (us-east-1) | ~$734 |
| GPU instances (spot) | 1× g5.xlarge spot, 24/7 (us-east-1) | ~$250 |
Notes:
- Platform costs (~$210/month) are fixed regardless of workload volume.
- GPU costs dominate and scale with the number of instances and hours run. Use
gco costs summaryto track actual spend. - GPU estimates assume an on-demand g5.xlarge in us-east-1 at ~$1.006/hr (~$734/month over 730 hours); rates vary by region and instance type.
- Optional services (FSx, Valkey, Aurora, EKS Capabilities — billed per capability-hour) add additional cost depending on configuration.
- The cost table above uses US East (N. Virginia) pricing as of June 2025.
Supported AWS Regions
GCO deploys to any AWS Region in the aws, aws-cn or GovCloud partitions, as many as you list in cdk.json under deployment_regions.regional. GPU availability varies by Region, so gco capacity recommend-region --gpu is a good first stop.
// cdk.json
{
"context": {
"deployment_regions": {
"regional": ["us-east-1", "eu-west-1", "ap-northeast-1"]
}
}
}
Then redeploy: gco stacks deploy-all -y. CDK bootstrap runs automatically for new regions.
GPU instance availability varies by region. Use gco capacity check -i <instance-type> -r <region> or gco capacity recommend-region --gpu to find regions with available GPU capacity before deploying workloads.
A regional stack can be deployed to any CloudFormation Region known to the installed AWS SDK. Add or remove Regions in
deployment_regions.regional; all configured Regions must belong to one AWS partition, and GCO imposes no count limit.
Documentation
| Your Goal | Read This |
|---|---|
| Deploy GCO and run a first job | Quick Start Guide |
| Understand what GCO does and the ideas behind it | Core Concepts |
| Learn Kubernetes and GCO along a guided path | Learning Path |
| See how the system fits together | Architecture Details |
| Read the official AWS Solutions Library overview and reference architecture | AWS Solutions Guidance |
| Let Claude Code, OpenAI Codex, or OpenCode drive GCO from your terminal | Autopilot Guide |
| Understand the project's north star and decision priorities | Project Tenets |
| Browse every guide in one place | Documentation Index |
Prefer a website? The project wiki
is a short orientation site — what GCO is, how it works, what you can run, and
where to go deeper — published from this repository with the live coverage
reports for
Python,
Bash and
Node.js
embedded.
| Your Goal | Read This |
|---|---|
| CLI commands and usage | CLI Reference |
| Deploy inference endpoints | Inference Guide |
| Run multi-node distributed training | Distributed Training Guide |
| Use the REST API directly | API Reference |
| Fix issues | Troubleshooting |
| Respond to incidents | Operational Runbooks |
| Track and analyze workload cost | Cost Monitoring Guide |
| Run interactive notebook analytics | Analytics Guide |
| Drive a goal-directed iteration loop | Mission Guide |
| Perform routine maintenance & upgrades | Maintenance Guide |
| Your Goal | Read This |
|---|---|
| Add regions, tune nodepools, enable FSx | Customization Guide |
| Reconcile tenant workloads from Git with Argo CD | GitOps with Argo CD |
| Publish your own platform APIs with Crossplane | Crossplane |
| Attach AWS-managed ACK or kro | EKS Capabilities |
| Choose a scheduler for your workload | Schedulers & Orchestrators |
| Mirror Docker Hub images into ECR (Volcano) | Image Mirror |
| Configure the SQS queue processor | Queue Processor Config |
| Contribute to the project | Contributing |
| Take GCO into your own repository | Forking Guide |
| API client examples (Python, curl, AWS CLI) | Client Examples |
| IAM policy templates | IAM Policies |
| Demo walkthroughs, recordings, and re-record scripts | Demo Starter Kit |
Project Tenets
GCO is guided by the prioritized project tenets, beginning with workload,
data, and account safety and anchored by the north star One API. Every Accelerator.
Any Region. The tenets define how the project resolves trade-offs across truthful
state, security, regional behavior, accelerator policy, automation, recovery,
operations, cost, and maintainability. Earlier tenets outrank later ones; durable
exceptions require an Architecture Decision Record.
Project Structure
Repository layout.
├── app.py # CDK app entry point
├── TENETS.md # Prioritized project principles and north-star guidance
├── cdk.json # CDK configuration (regions, features, thresholds)
├── pyproject.toml # Project metadata, dependencies, and CLI installation
│
├── cli/ # GCO CLI (jobs, stacks, capacity, inference, costs, DAGs)
├── demo/ # Recorded CLI demos (GIFs + asciinema sources) with walkthroughs and re-record scripts
├── diagrams/ # Auto-generated architecture diagrams (infra_diagrams/), code flowcharts (code_diagrams/), API spec sheets (api_specs/)
├── dockerfiles/ # Distroless container images for the in-cluster GCO services
├── docs/ # Documentation (architecture, CLI, API, inference, customization, analytics)
├── examples/ # Example manifests (jobs, inference, Ray, Volcano, Kueue, Slurm, YuniKorn)
├── gco/
│ ├── config/ # Configuration loader with validation
│ ├── models/ # Data models for k8s clusters, health monitor, inference monitor and manifest processor
│ ├── services/ # K8s services (health/inference monitors, inference proxy, manifest/queue processors)
│ └── stacks/ # CDK stacks (global, API gateway, regional, regional API gateway, monitoring, analytics)
│ └── constants.py # Pinned versions: EKS addons, Lambda runtime, Aurora engine
│
├── lambda/ # Lambda functions
│ ├── analytics-cleanup/ # Custom resource that deletes Studio user profiles + EFS access points on stack destroy
│ ├── analytics-presigned-url/ # Generates presigned SageMaker Studio URLs for Cognito-authenticated users
│ ├── api-gateway-proxy/ # API Gateway → Global Accelerator proxy
│ ├── capacity-poller/ # Scheduled EC2 capacity snapshots for the historical-capacity surface
│ ├── cross-region-aggregator/ # Cross-region job/health aggregation
│ ├── drift-detection/ # Scheduled drift checks against deployed CDK stacks
│ ├── ga-registration/ # Global Accelerator endpoint registration
│ ├── helm-installer/ # Installs Helm charts (schedulers, cert-manager)
│ │ └── charts.yaml # Helm chart configuration (schedulers, cert-manager)
│ ├── helm-orchestrator/ # Custom-resource provider that starts and polls the Helm-install state machine
│ ├── image-lookup/ # Adopt-or-create custom resource for the project's gco/* ECR repositories
│ ├── inference-streaming-proxy/ # Node.js response-streaming proxy for global and regional inference routes
│ ├── kubectl-applier-simple/ # Applies K8s manifests during deployment
│ │ └── manifests/ # Kubernetes manifests (nodepools, RBAC, services, storage)
│ ├── proxy-shared/ # Shared utilities for proxy Lambdas
│ ├── regional-api-proxy/ # Regional API Gateway → internal ALB proxy
│ ├── secret-rotation/ # Daily secret rotation
│ ├── tls-certificate-manager/ # Private TLS root, regional ACM leaves, and trust publication
│ ├── tls-shared/ # Strict private-root TLS client shared by the proxy Lambdas
│ ├── traffic-dial-controller/ # Health-driven Global Accelerator traffic dials
│ └── vector-ingest/ # S3-triggered chunking + Bedrock embeddings for the vector store
│
├── gco_mcp/ # MCP server for LLM interaction (142 tools default, up to 200 with feature flags)
├── images/ # Screenshots and visual assets for docs and the wiki
├── scripts/ # Utility scripts (version bump, cluster access setup)
├── tests/ # PyTest + BATS test suites
├── wiki/ # Orientation wiki sources, built with Zensical and published to GitHub Pages (wiki/README.md explains how)
└── zensical.toml # Zensical configuration for the GitHub Pages wiki (sources in wiki/, staged by scripts/build_wiki.py)
Contributing
See CONTRIBUTING.md for development setup, testing, the GitHub Actions CI/CD layout, the release process, and dependency scanning schedules. The whole test-suite runs from the same dev container scripts/setup-dev-alias.sh builds:
docker run --rm -v $(pwd):/workspace -w /workspace gco-dev pytest tests/ -v
Support
- Check Troubleshooting for common issues
- Review CloudWatch logs for Lambda and EKS errors
- Open an issue on GitHub
Security
GCO implements defense in depth across the six controls of its Security Model, checks its infrastructure against five cdk-nag rule packs, and scans container images and dependencies in CI.
Reporting a vulnerability: do not open a public GitHub issue; follow the responsible disclosure process in .github/SECURITY.md.
Authentication and Authorization:
- All API requests require AWS IAM (SigV4) authentication at the API Gateway
- The trusted proxy Lambda adds a request-bound HMAC envelope using a rotating Secrets Manager key; the reusable key is never sent downstream
- IRSA (IAM Roles for Service Accounts) provides pod-level AWS access with no static credentials
- EKS access entries with explicit policy bindings (no aws-auth ConfigMap)
Network Security:
- Regional platform ALBs are internal; the EKS API endpoint defaults to
PRIVATE - EKS clusters run in private subnets with configurable endpoint access (PRIVATE or PUBLIC_AND_PRIVATE)
- VPC gateway endpoints for S3 and DynamoDB keep that traffic off the public internet by default; opt-in interface endpoints (
vpc_endpoints.interfaceincdk.json) do the same for ECR, STS, SSM, CloudWatch, and SQS - VPC Flow Logs (30-day retention) capture all network traffic for audit
- Kubernetes Network Policies enforce default-deny with explicit allow rules
Encryption:
- Data at rest: S3 (KMS), EFS (KMS), EBS (KMS), DynamoDB (AWS-managed), and Secrets Manager (KMS)
- Client-to-global/regional API traffic uses AWS-managed TLS and IAM SigV4. Cross-region aggregation also uses AWS-managed TLS plus SigV4 from the aggregator to each regional API bridge.
- The normal global proxy → Global Accelerator → ALB path and each regional VPC proxy → ALB path use authenticated private-root TLS on TCP/HTTPS 443. Global Accelerator is a Layer 4 pass-through and does not terminate TLS.
- Every ALB leaf is issued for
backend.<project>.gco.internal; clients send that identity through SNI and assert it while connecting to dynamic accelerator or ALB DNS names. The ALB then re-encrypts to TLS-only pod proxy sidecars on port 8443; decrypted bytes travel only over pod loopback to the application process. - The root private key exists only in a customer-managed-KMS-encrypted Secrets Manager secret readable by the certificate-manager role. Proxy roles read only the public SSM trust bundle; rotating leaves are reimported into stable regional ACM ARNs.
- The request-bound HMAC envelope adds integrity, freshness, and replay defense; it is not encryption and is independent of TLS confidentiality and server authentication.
- EFS mount encryption in transit is enabled by the deployed storage configuration.
- Kubernetes secrets are encrypted in etcd (EKS-managed encryption).
Infrastructure Policy Validation:
- Five cdk-nag v3 rule packs run during CDK policy validation:
- Findings are either fixed or explicitly acknowledged with justification in
gco/stacks/nag_suppressions.py. - These automated checks are not certifications and do not by themselves establish compliance.
Supply Chain Security:
- Container images scanned with Trivy on every push (CVE detection)
- Python dependencies audited with pip-audit (GHSA/CVE detection)
- Both repository-owned npm graphs are exact-pinned with committed lockfiles (see package.json and package-lock.json), audited on every PR, and updated by Dependabot
- Production JavaScript is scanned by CodeQL and Semgrep; the inference-streaming Lambda has a separate Node.js 24 test workflow with exact 100% line/function/branch gates (see ./tests/inference-streaming-proxy/)
- Every Python dependency, direct and transitive, pinned to an exact version in
requirements-lock.txt(see requirements-lock.txt); CI fails when the lock drifts frompyproject.toml - Dependabot and CodeQL enabled for automated vulnerability alerts
- Strict KICS and Checkov infrastructure scans
- SBOM generation via Trivy for all container images
License
See the LICENSE file for details.
Yorumlar (0)
Yorum birakmak icin giris yap.
Yorum birakSonuc bulunamadi


