Featured image of post Google Open-Sources AX: A Kubernetes-Like Orchestrator for Billion-Scale Autonomous Agents

Google Open-Sources AX: A Kubernetes-Like Orchestrator for Billion-Scale Autonomous Agents

Google open-sources AX, a high-throughput declarative orchestrator for billion-scale agent workloads.

AX Open-Sourced: A Kubernetes-Like Orchestrator for Agent Workloads

On March 2026, Google open-sourced AX (github.com/google/ax), a high-throughput, declarative orchestrator runtime designed specifically for autonomous agent workloads. AX aims to schedule billions of agent tasks within a single cluster. Key facts:

  • Apache-2.0 licensed, led by Google
  • Current status: Alpha, with core concepts rapidly evolving
  • 11,600+ GitHub stars, 560+ forks
  • Kubernetes-compatible CLI: ax apply, get, describe, watch, delete
  • Built on Agent Substrate: AX handles orchestration only; sandbox isolation is delegated

The Agent Workload Gap

Traditional schedulers (e.g., Kubernetes Pod model) assume workloads are either stateless services or batch jobs. Agents differ fundamentally:

  • Accumulate state over time, not transient execution
  • Require strict sandbox isolation for security
  • Frequently call model APIs and tool servers
  • Risk burning token budgets without oversight
  • Need native support for pause-and-resume and interactive debugging

The AX team states bluntly: Agents are a new kind of workload, neither microservices nor batch. To address this, AX introduces three orthogonal declarative primitives:

  • Task: Isolated execution unit with suspend/resume capability
  • Workspace: Pre-warmed environment supporting Git clones and goal-driven setup
  • Model: Cluster-scoped model configuration for centralized credential rotation

A surprising architectural choice is storage backend: instead of Kubernetes etcd, AX uses Redis + Streams. The explanation: storing millions of short-lived tasks in etcd would hit its「single-digit GB storage limit」and write throughput bottlenecks. This shows AX deliberately abandoned the「everything is a CRD」K8s convention while retaining user-facing CLI familiarity.

Architecture and Notable Features

AX layered architecture:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
User Input (ax apply)
     │
     ▼
ax-server (stateless gRPC API)
     │
     ▼
Redis (Task Hash + Streams + PubSub)
     │
     ▼
ax-controller (horizontally scaled worker pool)
     │
     ▼
Agent Substrate (sandbox execution)

Key capabilities:

  • kubectl-style CLI: ax apply/get/describe/watch/delete plus ax ssh/suspend/resume
  • Minimalist Task design: Does not model agent internals (planning/delegation/retry); provides廉价, composable units
  • Workspace goal-driven setup: Natural language goal descriptions trigger automatic tool installation
  • Persistent Runner: ax-task-runner runs as PID 1; even after command exits, ax ssh remains accessible
  • Cluster-scoped Model resource: Centralized model credentials and parameters, avoiding duplication across agents
FeatureAX ApproachTraditional K8s Approach
StorageRedis + Streamsetcd
IsolationAgent SubstrateContainers/VMs
State ManagementControl plane / execution separationFull CRD persistence
Debuggingax ssh always-on serviceRequires debug container

Who Should Adopt Early?

AX is Alpha-stage with explicit warnings of breaking changes pre-stable-release. Appropriate for:

  • Teams running large-scale autonomous agents: e.g., mass code repair across repositories, multi-repo operations agents
  • Kubernetes-experienced teams: kubectl操作习惯 zero learning cost
  • Long-running tasks requiring suspension: ax suspend/resume enables checkpointing and resume
  • Tech-savvy contributors willing to accept Alpha risks: Can help shape future evolution (Actor API migration, idle auto-suspension)

Final Thoughts

AX’s value lies not in prescribing how agents should「think」, but in decomposing the infrastructure challenge of「simply running stateful workloads safely and at scale」into clean declarative primitives and a layered architecture. As agents transition from experiments to production, standardized orchestration at the base layer will define the ceiling for upper-layer capabilities.