Featured image of post DeepSeek.open-sources DSec: Agent Training Sandbox Platform Serving 3M Sandboxes Daily, Scaling to 32K Concurrent Instances

DeepSeek.open-sources DSec: Agent Training Sandbox Platform Serving 3M Sandboxes Daily, Scaling to 32K Concurrent Instances

DeepSeek releases first paper on Agent training infrastructure, detailing DSec's architecture and large-scale operations.

DeepSeek open-sources DSec: Agent training sandbox platform scaling to 30K+ concurrent instances

Core announcement: DSec (DeepSeek Elastic Compute), the infrastructure powering DeepSeek V3.2 to V4.1 RL training and evaluation, has been formally documented in a paper signed by founder Liang Wenfeng and submitted to arXiv on September 19, 2026 (ID: 2609.22978). Over 130 authors contributed. The platform had appeared previously only in the DeepSeek V4 technical report; this paper discloses its full architecture.

Key facts:

  • Paper submission date: September 19, 2026
  • Author count: 130+
  • Platform scale: ~300K sandboxes/day, peak concurrency >380K, creation rate >5,000/sec
  • Single-task capacity: up to 32,000 sandboxes launched simultaneously

Why so many sandboxes? The unique bottleneck of Agent training

Why so many sandboxes? The unique bottleneck of Agent training
Why so many sandboxes? The unique bottleneck of Agent training|News screenshot

Unlike traditional LLM RL, which uses static input-output-reward pairs, Agent training requires actual execution: installing dependencies, running software, modifying files, calling tools—each step alters environment state, requiring clean restoration for subsequent rollouts.

This creates scale challenges: DSec has launched 32,000 sandboxes in a single task. Counterintuitively, CPU usage during execution is intermittent and low, yet memory and writable state must persist throughout—rendering the traditional “start-container-run-task” model unsustainable.

DSec addresses four core needs: batch sandbox creation, resource scheduling, environment replication, and state preservation/restoration with security isolation.

Unified control plane: functions, containers, microVMs, and full VMs

Unified control plane: functions, containers, microVMs, and full VMs
Unified control plane: functions, containers, microVMs, and full VMs|News screenshot

For diverse Agent workloads, DSec provides four backends:

  • FnCall: short-lived function execution (e.g., code evaluation)
  • Container: full Linux user-space (software engineering tasks)
  • Firecracker microVM: high-isolation tasks
  • Full VM: near-physical-machine environment (e.g., commercial software)

The training framework uses a unified Python SDK (libdsec), oblivious to underlying backend type. After permission checks, the platform dispatches node-local Edge to create sandboxes; Aether and Chronus manage execution, while 3FS provides images on-demand.

Layered images and on-demand loading: decoupling maintenance from deployment

Key stat: A production week recorded 11,266 base images, 102,171 workspaces, and 103 toolkits across container backends; 67.8% of sandboxes layered workspaces/toolkits atop bases.

Monolithic image builds make any layer change trigger full rebuild and redistribution. DSec splits layers into read-only EROFS; overlayfs composes them at startup. Updates touch only modified layers.

Image distribution leverages on-demand loading: actual reads constitute only 4.2%–13.3% of full images. Images reside in 3FS; metadata prefetches locally, writes remain on-node, avoiding small random I/O overhead.

Results: Deploying 8,192 containers concurrently took 35 minutes with DSec vs. >60 minutes for Docker cold pull; per-node disk writes dropped from ~1.6TB to ~700GB.

ClassLoader also supports agent-driven environment construction: pack_diff generates delta snapshots, enabling new sandbox recovery.

Rollout offloaded from GPUs: uninterrupted execution during training

Rollout offloaded from GPUs: uninterrupted execution during training
Rollout offloaded from GPUs: uninterrupted execution during training|News screenshot

Early designs shared GPU pods for rollout and training; GPU preemption would interrupt execution. Since V4.1, rollout runs independently on DSec—Agent sandbox and worker container use CPU resources only, eliminating GPU dependency. Even when GPU training is preempted, rollout state persists.

For capacity bursts, DSec supports cloud bursting: when cluster utilization exceeds 80%, eligible sandboxes migrate to cloud VMs. A pre-staged ~30TB deduplicated image set (70% files accessed by container tasks) mitigates cold-pull latency. 200 cloud VMs can absorb ~30% of peak load.

Security isolation: AppArmor and eBPF with layered defense

Real environments introduce new risks: paper documents agents reading logs, forging RPC requests, modifying /bin/bash, scanning services, pulling external code, recursion-induced kernel crashes, and simple yes command causing log bloat to dozens of GB.

DSec employs:

  • AppArmor: restricts file and socket access
  • eBPF: dynamically limits network access per task phase

Paper admits: measures cover only partial risks; kernel-layer vulnerabilities remain imperfectly mitigated.

Adoption guidance

Adoption guidance
Adoption guidance|News screenshot

  • Build this if: Your team runs large-scale Agent training and needs infrastructure inspiration; security evaluation scenarios requiring high isolation benefit from DSec’s micro-isolation design.
  • Wait if: You lack PB-scale storage or 10K+ node clusters—sync on whether libdsec becomes open-source to assess migration feasibility.

Final thoughts

DSec signals a shift from ad-hoc to engineered infrastructure for Agent training. As model capabilities accelerate, execution platforms managing interactive, long-running tasks will become the next bottleneck—a frontier worth watching.