Problem Statement
Organizations that run Large Language Models (LLMs) in air-gapped environments
quickly accumulate model endpoints: every open weight model served on
Google Distributed Cloud (GDC) air-gapped (see the companion guide set Open Weight Models
on GDC air-gapped) exposes its own Kubernetes Service,
its own address, its own credentials and its own operational lifecycle.
Application and agent developers end up in architectural silos, hard-coding one
endpoint per model and rewriting integrations with every model iteration. This
fragmentation creates technical debt and slows down the deployment of
mission-critical, multi-step agentic workloads.
Standard networking layers are also "AI-blind": a conventional load balancer
treats inference traffic like generic HTTP and can't read the model field of a
request, prioritize latency-sensitive workloads, account for tokens instead of
requests, or steer a request to the model server replica that already holds the
relevant KV cache. In high-cost air-gapped GPU clusters this leads to hardware
underutilization and a poor return on investment. While Gemini on
GDC
provides managed frontier models, customers who combine them with self-hosted
open weight models need an inference layer that offers cloud-like agility
without compromising data sovereignty: no request, prompt or model weight may
leave the GDC boundary.
Proposed solution
This document outlines the reference architecture for a self-managed AI gateway
on GDC air-gapped: a secure, centralized inference layer
that gives applications one OpenAI-compatible endpoint for all models and routes
each request to the right backend. The reference implementation builds on the
Kubernetes Gateway API: Envoy Gateway is the Gateway API implementation (control
plane and Envoy proxy data plane) and Envoy Agent Router (formerly Envoy AI
Gateway) extends it with AI-specific request processing: model extraction from
the request body, request and response translation, upstream authentication,
token accounting and, optionally, token-based rate limiting. The Gateway API
Inference Extension adds model-server-aware endpoint selection (InferencePool
and Endpoint Picker) so that GPU capacity is used efficiently. All artifacts are
staged in the Harbor registry of the GDC organization
and the whole stack runs in a standard cluster without any external
connectivity.
What is included: This architecture covers the deployment of Envoy Gateway
and Envoy Agent Router in a standard cluster, the Gateway API resources
(GatewayClass, Gateway) and Envoy Agent Router resources (AIGatewayRoute,
AIServiceBackend, BackendSecurityPolicy, GatewayConfig) that expose a
unified OpenAI-compatible endpoint, body-based routing on the model field, the
Gateway API Inference Extension (InferencePool, InferenceObjective, Endpoint
Picker) for model-aware load balancing, optional token-based rate limiting with
Redis, TLS termination at the gateway, network policies (NetworkPolicy /
ProjectNetworkPolicy) and the air-gapped artifact management in Harbor. The
model serving backends behind the gateway (Ollama, vLLM) are deployed with the
companion guide set Open Weight Models on GDC
air-gapped.
What isn't included: This architecture doesn't cover a managed AI gateway
product, the deployment and tuning of the models themselves (companion guide
set), model fine-tuning, non-LLM GenAI modalities, agent-to-tool traffic
(MCPRoute) or commercially supported gateway distributions. Gemini on
GDC is a managed product capability; its provisioning is
out of scope and this version of the guide set validates the gateway with
self-hosted open weight models.
Audience: The primary audience for this document includes Platform Administrators responsible for deploying and operating the gateway in a standard cluster and technical architects designing AI platforms within the air-gapped environment. Application developers who consume the unified endpoint are a secondary audience who benefit from understanding how their requests are routed.
Architecture
The solution runs entirely in a standard cluster of the
GDC air-gapped organization. In the Services layer the
Kubernetes Gateway API standardizes the configuration model; the AI gateway
control plane (the Envoy Gateway controller and the Envoy Agent Router
controller) reads the Gateway API and Envoy Agent Router resources and
configures the data plane: an Envoy proxy Deployment that is the request entry
point (exposed through a Service of type LoadBalancer) and the Envoy Agent
Router external processor that runs as a sidecar in the proxy Pod and executes
the AI-specific logic. In the Application layer, customer workloads
(applications and agents) send an AI request to the proxy (1), the proxy hands
the request to the external processor, which extracts the model name, translates
the request and records token usage (2), the proxy forwards the balanced traffic
to the inference Pods of the selected model (3), and the model runs on the GPU
nodes of the Infrastructure layer (4). Model serving backends are attached to
the gateway either as an AIServiceBackend (any OpenAI-compatible Service) or
as an InferencePool whose Endpoint Picker selects the best replica per
request. Container images and Helm charts are pulled from Harbor.

Solution components
The solution uses the following components:
| Solution Component | Responsibilities | Data Flow | Interfaces |
|---|---|---|---|
| Kubernetes Gateway API | Standard, role-oriented resource model for ingress and routing: GatewayClass (owned by the platform), Gateway (listeners, TLS) and route resources (owned by application teams). Successor of the Ingress API without vendor-specific annotations. |
Declarative configuration only; the resources are watched by the controllers and carry no request traffic. | Kubernetes API (gateway.networking.k8s.io/v1); the CRDs are installed from the Envoy Gateway CRD chart. |
| Envoy Gateway controller (control plane) | Implements the Gateway API: provisions one Envoy proxy Deployment and LoadBalancer Service per Gateway, translates routes and policies (BackendTrafficPolicy, SecurityPolicy, EnvoyExtensionPolicy, Backend, EnvoyProxy) into xDS configuration, runs the global rate limit service when configured, and calls the Envoy Agent Router extension server through its extension manager so that InferencePool backends are added to the translation. |
Watches the Kubernetes API and pushes xDS configuration to the proxies; no request traffic. | Kubernetes API (watch), xDS gRPC towards the proxies, gRPC towards the extension server (port 1063). |
| Envoy Agent Router controller (control plane) | Reconciles the aigateway.envoyproxy.io/v1beta1 resources, AIGatewayRoute (routing rules on the extracted model name), AIServiceBackend (a backend with its API schema), BackendSecurityPolicy (upstream credentials) and GatewayConfig (external processor settings per Gateway), into Envoy Gateway resources, injects the external processor sidecar into the Envoy proxy Pods and serves the extension server used by Envoy Gateway. |
Configuration only. | Kubernetes API, gRPC extension server (port 1063), mutating admission webhook for the sidecar injection. |
| Envoy proxy (data plane) | Request entry point: terminates TLS, matches listeners and routes, load-balances across backend endpoints, enforces rate limits and timeouts, retries and falls back to alternative backends, and emits access logs and metrics. | Receives client requests on the LoadBalancer VIP (80/443), consults the external processor for every AI request, forwards the request to the selected backend endpoint (a Service or the Pod chosen by the Endpoint Picker) and streams the response back to the client. |
HTTP/HTTPS towards clients (OpenAI-compatible API), ext_proc gRPC towards the sidecar and the Endpoint Picker, HTTP towards the backends, gRPC towards the rate limit service. |
| Envoy Agent Router external processor (ExtProc) | Parses the OpenAI-compatible request body, extracts the model field into the routing header x-ai-eg-model, applies model name overrides, translates request and response payloads between the client API schema and the backend API schema, injects the upstream credentials defined by BackendSecurityPolicy, and extracts token usage (input, output, total) from the response into dynamic metadata used for token-based rate limiting and metrics. |
Sidecar container in the Envoy proxy Pod; processes request and response bodies, including streamed responses. |
Envoy ext_proc gRPC; Prometheus metrics endpoint. |
| Gateway API Inference Extension (InferencePool + Endpoint Picker) | InferencePool (inference.networking.k8s.io/v1) groups the Pods of one model server deployment (label selector, target port, for example 8000) and references its Endpoint Picker (EPP). InferenceObjective (v1alpha2) assigns a priority per model or objective. The EPP scrapes the model server metrics (request queue depth, KV-cache utilization, loaded LoRA adapters) and selects the endpoint for every request. |
The proxy calls the EPP (ext_proc, port 9002) with the request headers; the EPP answers with the endpoint to use; the proxy forwards the request to that Pod. |
Kubernetes API (CRDs), gRPC ext_proc (9002), Prometheus scrape of the model servers. |
| Redis (optional) | Backing store of the Envoy rate limit service: shared counters for request-based and token-based limits across all proxy replicas. Only needed when global rate limiting is configured. | The rate limit service reads and increments counters on port 6379 for every limited request. | Redis protocol (6379). Redis 8.10 deployed in the cluster, or an existing Redis instance. |
| Model serving backends (Ollama / vLLM) | Serve the open weight models behind the gateway; deployed with the companion guide set Open Weight Models on GDC air-gapped. Attached as an AIServiceBackend (any OpenAI-compatible Service, for example Ollama) or as an InferencePool (model servers such as vLLM that expose the metrics the Endpoint Picker scores on). |
Receive the routed requests from the proxy, run inference on the GPU nodes (or on CPU for functional evaluation with small models) and return the OpenAI-compatible response. | OpenAI-compatible REST API: vLLM on port 8000, Ollama on port 11434; Prometheus metrics (vLLM) for the Endpoint Picker. |
| Harbor | Stores the container images (Envoy proxy, Envoy Gateway, Envoy Agent Router controller and external processor, rate limit service, Endpoint Picker, Redis) and the Helm charts (OCI artifacts) of the solution. | Artifacts are copied from the public registries into Harbor with seed_registry.sh (crane) from a workstation with internet access; the cluster pulls images with a robot account through image pull secrets and the charts are rendered from Harbor with helm template. |
OCI Distribution API; Harbor UI and robot accounts. |
| Standard cluster and GPU nodes | Hosts all gateway components (CPU nodes are sufficient) and the inference Pods (GPU nodes with the accelerator resources discovered from the nodes). Provides LoadBalancer VIPs, persistent storage for the model weights and the network policy enforcement. |
Foundational platform; requests enter through the LoadBalancer VIP of the Gateway and stay inside the cluster network. |
Kubernetes API, GDC networking APIs (networking.gdc.goog/v1 ProjectNetworkPolicy), GDC console and gdcloud for the Platform Administrator. |
Features and capabilities
This reference architecture provides the following core capabilities:
- Unified OpenAI-compatible endpoint: Applications call one address
(
/v1/chat/completions,/v1/models, …) for every model served behind the gateway.- Components:
Gateway,AIGatewayRoute, Envoy proxy, external processor. - Responsibilities: Exposing a stable virtual IP, terminating TLS, accepting the OpenAI-compatible request format and translating it to the backend schema when needed.
- Interfaces: HTTPS REST (JSON, including server-sent events for streaming).
- Components:
- Body-based (model-aware) routing: The gateway reads the
modelfield of the request body and routes the request to the matching backend without any client-side endpoint logic.- Components: External processor,
AIGatewayRouterules matching onx-ai-eg-model,AIServiceBackend,InferencePool. - Responsibilities: Extracting the model name, matching route rules,
renaming models (
modelNameOverride), splitting traffic by weight and falling back to alternative backends by priority. - Interfaces: Kubernetes API (
aigateway.envoyproxy.io/v1beta1).
- Components: External processor,
- Model-server-aware load balancing: Requests to an
InferencePoolare sent to the replica with the shortest queue and the most useful KV cache instead of round robin.- Components: Gateway API Inference Extension (
InferencePool,InferenceObjective, Endpoint Picker). - Responsibilities: Scraping model server metrics, scoring endpoints (queue depth, KV-cache utilization, prefix-cache affinity, LoRA adapters), honoring per-objective priorities under load.
- Interfaces: Kubernetes API (
inference.networking.k8s.io/v1,inference.networking.x-k8s.io/v1alpha2), ext_proc gRPC (9002).
- Components: Gateway API Inference Extension (
- Upstream authentication and security policies: Credentials for backends
that require them are injected by the gateway and never distributed to the
applications.
- Components:
BackendSecurityPolicy, Envoy GatewaySecurityPolicy, TLS listeners. - Responsibilities: Injecting API keys towards the backends, terminating
client TLS at the
Gateway, enforcing client authentication and authorization policies when required. - Interfaces: Kubernetes
Secrets referenced by the policies.
- Components:
- Token-based rate limiting (optional): Usage is measured in tokens, not
only in requests.
- Components: External processor (
llmRequestCosts), Envoy GatewayBackendTrafficPolicy(global rate limit), rate limit service, Redis. - Responsibilities: Extracting input, output and total tokens from every response, attaching them as request costs, enforcing per-consumer or per-model budgets across all proxy replicas.
- Interfaces: Kubernetes API, Redis (6379).
- Components: External processor (
- Traffic management: Standard Gateway API capabilities apply to AI
traffic.
- Components:
Gateway,HTTPRoute,AIGatewayRoute,BackendTrafficPolicy. - Responsibilities: Host-, path- and header-based matching, timeouts suited to long generations, retries, weighted traffic splitting and priority-based fallback between backends.
- Interfaces: Kubernetes API.
- Components:
- Observability: Request, latency and token metrics are available for
every model.
- Components: Envoy proxy access logs and statistics, external processor metrics (token usage, request duration, time to first token), Endpoint Picker metrics.
- Responsibilities: Exposing Prometheus-format metrics and structured access logs that the GDC observability stack or a customer-managed Prometheus can scrape.
- Interfaces: Prometheus metrics endpoints, standard output logs.
- Air-gapped artifact management: Every image and chart is staged in
Harbor before installation.
- Components: Harbor,
seed_registry.sh, robot accounts, image pull secrets,EnvoyProxyresource (proxy image and pull secret). - Responsibilities: Seeding images and OCI Helm charts from a connected
workstation, pinning versions, rendering charts offline with
helm template. - Interfaces: OCI Distribution API, Kubernetes API.
- Components: Harbor,
Architectural principles
- Standards First: The gateway is configured through the Kubernetes Gateway API and its Inference Extension and speaks the OpenAI-compatible API towards applications, so applications, routes and backends stay portable across gateway implementations.
- Security First: All components and data reside within the
GDC air-gapped boundary. TLS terminates at the
Gateway, backend credentials are injected by the gateway, and network access is restricted withNetworkPolicyandProjectNetworkPolicyobjects. - Data Sovereignty: Prompts, responses, model weights and usage data never leave the customer's GDC organization; no external connectivity is required for operation.
- Separation of Concerns: The role-oriented Gateway API lets the Platform
Administrator own the
GatewayClass,Gateway, TLS and network policies while application teams own theirAIGatewayRoute,AIServiceBackendandInferencePoolresources in their own namespaces. - Decoupled Control and Data Plane: Controllers only translate configuration; the request path consists of the Envoy proxy, its external processor sidecar and the Endpoint Picker, which keep serving traffic if a controller restarts.
- Open & Flexible: The architecture uses open-source software throughout (Envoy Gateway, Envoy Agent Router, Gateway API Inference Extension, Ollama, vLLM) and the same Gateway API resources can be served by alternative implementations.
- Scalability & Performance: The data plane scales horizontally by adding proxy replicas and vertically by adjusting CPU and memory; model-server-aware load balancing keeps the GPUs saturated without overloading individual replicas.
- Simplified End User Experience: One endpoint, one API format and one credential for every model shorten the path from model deployment to consumption.
- Resilience: Kubernetes health checks and replication, Envoy retries and priority-based fallback between backends keep the endpoint available when individual model replicas fail.
Considerations
- Scalability and Performance:
- The gateway components are CPU-only workloads; the Envoy proxy
Deploymentof aGatewayscales horizontally by increasing its replica count (EnvoyProxyresource) and vertically by raising CPU and memory requests. Body inspection by the external processor costs CPU proportional to request and response size; streamed responses are processed incrementally. - Model-server-aware load balancing only applies to backends attached as
an
InferencePool; backends attached as anAIServiceBackendare load-balanced by the proxy across the endpoints of theirService. - The gateway doesn't add GPU capacity: throughput and latency are bounded by the model servers. Size the model deployments with the companion guide set (vLLM for high throughput, Ollama for ease of use and CPU-only evaluation).
- Long generations need longer request and idle timeouts than typical HTTP APIs; configure them on the routes and policies rather than on the clients.
- The gateway components are CPU-only workloads; the Envoy proxy
- Resource Management:
- All gateway components run on CPU nodes and need no accelerator; the following table gives indicative sizing for a functional deployment. The reference implementation sets the exact requests and limits; production values depend on request rate and payload sizes.
- The port table that follows lists every port the network policies must admit.
- GPU planning belongs to the model deployments: the accelerator resource names are discovered from the nodes as described in the companion guide set.
Indicative sizing of the gateway components (CPU only)
| Component | Replicas | CPU (request) | Memory (request) | Notes |
|---|---|---|---|---|
| Envoy Gateway controller | 1 | 100m–500m | 256Mi–1Gi | Control plane; not in the request path |
| Envoy Agent Router controller | 1 | 100m–500m | 256Mi–512Mi | Control plane and extension server |
Envoy proxy Pod (proxy + external processor sidecar) |
1–2 per Gateway |
500m–2 | 512Mi–2Gi | Scale with request rate and payload size |
| Endpoint Picker | 1 per InferencePool |
100m–500m | 256Mi–512Mi | Scrapes the model server metrics |
| Rate limit service (optional) | 1 | 100m–500m | 256Mi–512Mi | Only with global rate limiting |
| Redis (optional) | 1 | 100m–500m | 256Mi–1Gi | Only with global rate limiting |
Network ports
| Port | Component | Purpose |
|---|---|---|
| 80 / 443 | Envoy proxy (Gateway listeners) |
HTTP / HTTPS entry point on the LoadBalancer VIP |
| 1063 | Envoy Agent Router controller (extension server) | gRPC calls from the Envoy Gateway extension manager |
| 9002 | Endpoint Picker | ext_proc gRPC calls from the Envoy proxy |
| 6379 | Redis | Rate limit counters |
| 8000 | vLLM | OpenAI-compatible API and metrics of the model server |
| 11434 | Ollama | Ollama and OpenAI-compatible API of the model server |
- Availability and Reliability:
- The proxy and the external processor run in the same
Pod; multiple replicas behind theLoadBalancerServicemask individualPodfailures, and the controllers can restart without interrupting traffic. AIGatewayRouterules can list several backends with priorities so that the gateway falls back to another replica set or model server when a backend is unhealthy.- The Endpoint Picker is a per-pool
Deployment; if it is unavailable, requests to thatInferencePoolfail until it is back. Run it with resource requests and readiness probes and monitor it like a data plane component. - Redis holds only rate limit counters; a Redis outage affects rate limiting, not routing.
- The proxy and the external processor run in the same
- Operational Complexity:
- Version alignment between Envoy Gateway, Envoy Agent Router and the Gateway API Inference Extension: upgrade the three together with the tested combination (Envoy Gateway 1.8.4, Envoy Agent Router 1.1.0, Gateway API Inference Extension 1.5.0).
- Air-gapped lifecycle: every image and chart is re-seeded into Harbor for
each upgrade; the
EnvoyProxyresource and the chart values pin the registry and the pull secret. - Two configuration layers: platform-owned
GatewayClass/Gateway/EnvoyProxyand team-ownedAIGatewayRoute/AIServiceBackend/InferencePool; onboarding a new model means adding a route rule and a backend, not a new gateway. - Monitoring: request rate, latency (including time to first token), token usage per model and per consumer, Endpoint Picker scoring, and GPU utilization of the model servers.
InferenceObjectiveis still an alpha API (v1alpha2); expect schema changes between releases of the Inference Extension.
- Security and Compliance:
- The entire solution operates within the GDC air-gapped security perimeter.
- Network policies are mandatory: a namespace
NetworkPolicyadmits traffic from the proxy to the model server ports (8000, 11434) and to the Endpoint Picker (9002), and aProjectNetworkPolicyadmits client traffic to theLoadBalancerVIP when consumers reside outside the project. - TLS terminates at the
Gatewaywith a certificate stored in a KubernetesSecret; clients only ever see the gateway endpoint. - Upstream credentials (for example an API key required by a model server)
are stored in
Secrets referenced by aBackendSecurityPolicyand injected by the gateway, so applications never hold backend credentials. - Token-based rate limiting bounds the consumption of expensive GPU capacity per consumer or per model.
- Harbor robot accounts and Kubernetes image pull secrets ensure secure access to container images; role-oriented Gateway API RBAC separates platform and application responsibilities.
Design decisions
The following key design decisions have been made for this reference architecture:
Decision 1: AI gateway implementation: open-source Envoy Agent Router on Envoy Gateway versus alternative distributions
| Context | Key Decision | Rationale |
|---|---|---|
| The AI gateway is the primary architectural choice with viable alternatives, from self-managed open-source options to commercially supported offerings: (a) Envoy Agent Router on Envoy Gateway (open source), (b) kgateway and agentgateway (open-source Gateway API implementations that also handle LLM provider traffic, MCP tools and agent-to-agent communication in one data plane, optionally with vendor support from solo.io), (c) Solo Enterprise for agentgateway as the commercially supported premium option. | The reference implementation uses the open-source Envoy Agent Router on Envoy Gateway. The alternatives remain valid choices behind the same Gateway API resources; a commercially supported distribution is recommended for production environments that require dedicated vendor support. | Open-source advantages: lightweight and standards-conformant, high-performance L7 routing, no licensing fees, ideal to evaluate the solution before committing to a paid offering; Envoy Agent Router is the option validated in this guide set. Open-source disadvantages: support is best effort by the Solutions team and community-driven, which may not satisfy core production environments. Premium advantages: enterprise support agreements for the gateway lifecycle and, often, advanced agentic capabilities. Premium disadvantages: paid licensing or support agreement and an additional vendor relationship. |
Decision 2: Kubernetes Gateway API with the Inference Extension versus Ingress or a proprietary gateway API
| Context | Key Decision | Rationale |
|---|---|---|
| Inference endpoints can be exposed with the legacy Ingress API, a gateway-specific configuration format, or the Kubernetes Gateway API. Model-aware load balancing additionally needs a way to describe model server pools and their priorities. | The solution standardizes on the Kubernetes Gateway API (GatewayClass, Gateway, routes) implemented by Envoy Gateway, and on the Gateway API Inference Extension (InferencePool, InferenceObjective, Endpoint Picker) for model-aware routing. |
The Gateway API is the Kubernetes standard that succeeds Ingress, is role-oriented (platform versus application teams) and portable across implementations. The Inference Extension is the community standard for model-server-aware routing (queue depth, KV cache, priorities) and is supported by Envoy Agent Router out of the box; the alternative gateway distributions implement the same APIs, which keeps Decision 1 reversible. |
Decision 3: Body-based routing on the model field of OpenAI-compatible
requests as the unified endpoint
| Context | Key Decision | Rationale |
|---|---|---|
Applications can address models through one endpoint per model (host- or path-based), through a custom header set by the client, or through the model field that every OpenAI-compatible request already carries. |
The gateway exposes one OpenAI-compatible endpoint and routes on the model field: the external processor extracts it into the x-ai-eg-model header and AIGatewayRoute rules match on that header. |
Existing OpenAI-compatible clients and SDKs work unchanged and switch models by changing one string; endpoints, credentials and model placement can change without touching applications; the same rules can map a model name to several backends for traffic splitting, fallback or model name overrides. |
Decision 4: Air-gapped artifact seeding into Harbor versus on-demand pulls
| Context | Key Decision | Rationale |
|---|---|---|
| The upstream installation pulls container images from public registries and Helm charts as OCI artifacts on demand, which is impossible in an air-gapped environment. | All container images (Envoy proxy, Envoy Gateway, Envoy Agent Router controller and external processor, rate limit service, Endpoint Picker, Redis) and the Envoy Gateway and Envoy Agent Router Helm charts are seeded into a Harbor project with seed_registry.sh (crane), charts are rendered offline with helm template from Harbor, and the EnvoyProxy resource and the chart values pin the Harbor registry and the pull secret. |
One repeatable, auditable seeding step per version; the cluster never needs external connectivity; version pinning makes upgrades explicit and reversible. |
Decision 5: Optional token-based rate limiting with Redis versus request-based limits only
| Context | Key Decision | Rationale |
|---|---|---|
| GPU capacity is consumed per token, not per request; a request-count limit can't protect a shared model from a few very large prompts or long generations. Global rate limiting in Envoy Gateway needs a shared counter store. | Token-based rate limiting is an optional feature of the architecture: the external processor records input, output and total tokens per request (llmRequestCosts), a BackendTrafficPolicy enforces the budget, and a Redis instance (Redis 8.10, deployed in the cluster or an existing instance) stores the counters. Deployments without rate limiting don't need Redis. |
Usage is measured in the unit that reflects cost, budgets are enforced consistently across all proxy replicas, and the dependency is only paid for by customers who need quotas. |
Decision 6: Topology: gateway control and data plane in a dedicated namespace of a standard cluster, models in application namespaces
| Context | Key Decision | Rationale |
|---|---|---|
| The gateway could run in a dedicated cluster, in every application namespace, or once per standard cluster next to the models it fronts. Open weight models on GDC air-gapped are deployed in standard clusters (companion guide set). | The Envoy Gateway and Envoy Agent Router control plane and the Envoy proxy data plane run in a dedicated namespace of the standard cluster that hosts the models; AIGatewayRoute, AIServiceBackend, InferencePool and the model deployments live in the application namespaces and are referenced across namespaces. |
Requests stay inside the cluster network between the proxy and the model Pods (lowest latency, no additional load balancers), the Platform Administrator owns the gateway namespace while application teams own their routes and models, and the Endpoint Picker can scrape the model servers directly. A per-project client gateway that forwards to a central server gateway remains an extension of this topology for multi-project consumption. |
Assumptions and limitations
- GDC air-gapped 1.16.2-hf1 or later environment is available with a standard cluster that runs Kubernetes v1.32.13-gke.400 or later.
- Sufficient hardware resources (especially GPUs) are present for the model serving backends; the gateway components themselves run on CPU nodes. Real models need GPUs; small models on Ollama can run on CPU for functional evaluation only.
- Harbor instance is set up and accessible, and a workstation with internet access and connectivity to the GDC environment (Harbor and Kubernetes API) is available for seeding images and charts.
- Users have the necessary permissions: project roles Harbor Instance Viewer
and Standard Cluster Admin plus a
StandardClusterRoleBindingto thecluster-adminrole on the standard cluster for the cluster-scoped installation (CRDs,GatewayClass, namespaces), see the reference implementation. - The model serving backends expose an OpenAI-compatible API; backends
attached as an
InferencePoolalso expose the vLLM-compatible metrics the Endpoint Picker relies on. - Limitation: support for the open-source components (Envoy Gateway, Envoy Agent Router, Gateway API Inference Extension) is best effort by the Solutions team; dedicated vendor support requires a commercially supported distribution (Decision 1).
- Limitation:
InferenceObjectiveis an alpha API (inference.networking.x-k8s.io/v1alpha2) and may change between releases of the Inference Extension. - Limitation: a
LoadBalancerVIP for theGatewayand aProjectNetworkPolicyare required for consumers outside the project; inside the project, in-cluster clients orkubectl port-forwardcan reach the gatewayServicedirectly. - Limitation: Gemini on GDC, model fine-tuning, non-LLM GenAI models and multi-node model serving are out of scope for this version of the solution.
User flow
The following is a comprehensive user flow for the AI gateway on GDC air-gapped: first the request path that an application developer relies on, then the setup path the Platform Administrator follows.
1. Developer request path
An application or agent needs nothing but the gateway endpoint, a client credential if the gateway enforces one, and the model name.
- Send the request: The application calls
POST /v1/chat/completionson the gateway endpoint with an OpenAI-compatible JSON body that names the model (for example"model": "gemma"), exactly as it would call any OpenAI-compatible service. - Enter the gateway: The request reaches the
LoadBalancerVIP of theGateway; the Envoy proxy terminates TLS and matches the listener and route. - Process the body: The proxy hands the request to the Envoy Agent Router
external processor, which parses the body, extracts the model name into the
x-ai-eg-modelheader and translates the payload to the backend schema if it differs. - Select the backend: The
AIGatewayRouterule that matches the model name selects anAIServiceBackend(for example the OllamaService) or anInferencePool(for example a vLLM replica set); for anInferencePool, the Endpoint Picker chooses the replica with the shortest queue and the best KV-cache match. - Forward and respond: The proxy injects upstream credentials when a
BackendSecurityPolicyapplies, forwards the request to the chosenPod, streams the response back and lets the external processor record the token usage, which feeds the metrics and, when configured, the token-based rate limit. - Switch models: To use another model the developer changes the
modelstring; no new endpoint, credential or client library is needed.
2. Platform Administrator setup path
- Seed artifacts into Harbor: Create the Harbor robot accounts, then copy
the container images and the Envoy Gateway and Envoy Agent Router Helm
charts into Harbor with
seed_registry.shfrom a workstation with internet access. - Install Envoy Gateway: Render the CRD and controller charts from Harbor
with
helm templateand apply them in the gateway namespace; create theEnvoyProxyresource that pins the proxy image and pull secret and verify the controller with a basicGatewayClass,GatewayandHTTPRoute. - Install Envoy Agent Router: Install the Envoy Agent Router CRDs and the
Gateway API Inference Extension CRDs, deploy the Envoy Agent Router
controller, and reconfigure Envoy Gateway to call its extension server (port
1063) and to accept
InferencePoolbackends; optionally deploy Redis and enable the global rate limit service. - Expose the gateway: Create the
GatewayClassand theGatewaywith HTTP and HTTPS listeners (TLS certificateSecret), wait for theLoadBalancerVIP, and apply theProjectNetworkPolicyfor consumers outside the project. - Attach models: Deploy the model serving backends with the companion
guide set, then create
AIServiceBackendresources for OpenAI-compatibleServices andInferencePoolplus Endpoint Picker resources for vLLM-compatible replica sets, together with the namespaceNetworkPolicyobjects. - Publish the route: Create the
AIGatewayRoutewith one rule per model name (body-based routing) and, where needed,BackendSecurityPolicyandBackendTrafficPolicyobjects for upstream credentials and token budgets. - Validate and operate: Call
/v1/modelsand/v1/chat/completionsthrough the VIP (or akubectl port-forwardfallback), check the proxy, external processor and Endpoint Picker metrics, and hand the endpoint and model names to the application teams. Scale proxy replicas and model deployments independently; upgrade Envoy Gateway, Envoy Agent Router and the Inference Extension together.
The detailed steps are provided in the Envoy Agent Router reference implementation and Body-based routing with Envoy Agent Router user guide.
Additional materials
- Envoy Agent Router reference implementation
- Body-based routing with Envoy Agent Router user guide
- Open-weight LLMs reference architecture
- Envoy Agent Router documentation
- Envoy Gateway
- Kubernetes Gateway API
- Gateway API Inference Extension
- Google Distributed Cloud air-gapped documentation