The editorial frames DSec as correcting a two-year-old architectural mistake: inference frameworks built for dense models baked expert placement in at load time, ignoring that production routing distributions are deeply non-uniform. By tracking per-expert activation and migrating FFN weights live, DSec turns a compile-time decision into a runtime one and posts 2.3× throughput over vLLM as evidence.
Submitted the DSec paper to HN where it hit 238 points quickly, signaling agreement that the hot-expert problem is a real production pain point worth surfacing. The submission itself endorses the framing that dynamic expert placement is a meaningful advance over vLLM/SGLang's static approach.
The editorial highlights that the technical HN discussion converged on one detail: DSec sidesteps the KV cache problem entirely because KV cache lives on the attention side, not the expert side, so only FFN weights need to move. This is characterized as the 'obvious in retrospect' insight that unblocked a design space the field had written off, and it's what makes migration cheap enough to do without draining in-flight requests.
The editorial carefully notes DSec's 2.3× vLLM / 1.6× SGLang numbers come from a specific 60/30/10 chat/coding/long-context mix on DeepSeek-V3 at a matched p99 TTFT target. This implicitly cautions that deployments with more uniform routing distributions, different model architectures, or different latency SLOs may see materially smaller wins, since the whole approach only pays off when expert imbalance is severe enough to justify migration overhead.
DeepSeek dropped a paper on arxiv describing DSec (DeepSeek Elastic Compute), a serving system built specifically for the failure mode that keeps MoE deployments expensive: hot experts. In a Mixture-of-Experts model, tokens are routed through a small subset of experts per layer — for DeepSeek-V3, that's top-8 out of 256 routed experts plus one shared expert per layer, with 37B parameters active out of 671B total. In production traffic, that routing distribution is emphatically not uniform. A handful of experts get hammered, most sit idle, and the GPU holding the popular ones becomes the bottleneck while the rest of your fleet burns power waiting.
DSec's core move is to stop pretending expert placement is a compile-time decision. The system tracks per-expert activation rates in a rolling window and physically migrates expert weights between GPUs when the imbalance crosses a threshold, without draining in-flight requests. The paper benchmarks against vLLM 0.6.x and SGLang, both of which use static expert-parallel placement decided at model load. On DeepSeek-V3 with a mixed workload (60% chat, 30% coding, 10% long-context RAG), DSec reports 2.3× throughput over vLLM and 1.6× over SGLang at matched p99 TTFT of ~400ms. Peak memory reduction is 22% on an 8×H100 node because DSec drops replicated cold experts once traffic stabilizes.
The HN thread hit 238 points within hours, with most of the technical discussion focused on how DSec handles the KV cache during migration — historically the reason nobody wanted to move experts. Answer: DSec doesn't move the KV cache. Only the expert FFN weights move, since KV cache is attention-side, not expert-side. This is the kind of detail that looks obvious in retrospect and clearly wasn't when the field settled on static placement two years ago.
The MoE serving story since Mixtral has been a slow acceptance that inference frameworks built for dense models are the wrong shape. vLLM's PagedAttention was the big unlock for KV cache management, but expert routing remained a second-class citizen — you got expert parallelism (EP) as a placement strategy and that was mostly it. SGLang added RadixAttention and better prefix caching but still treats experts as pinned. Both frameworks assume that whatever placement you chose at load time is the placement you'll live with until restart.
That assumption breaks in production. A code-generation workload activates different experts than a chat workload. Time-of-day shifts routing. A single popular prompt template can skew a specific expert's load 5-10× above the mean for hours. The industry response has been to overprovision — replicate the hottest experts across more GPUs than the model needs — which wastes both HBM and interconnect bandwidth. DSec argues that replication is a workaround for the real problem, which is that expert placement should be a runtime scheduling decision, not a deployment artifact.
How the migration actually works is worth walking through, because it's more constrained than it sounds. DSec models each expert as a movable unit of ~44MB (for DeepSeek-V3's FFN experts at FP8). When the load balancer flags an imbalance, the scheduler picks a source GPU and target GPU, copies the expert weights over NVLink, and updates the routing table atomically. In-flight tokens that were about to route to the old location are redirected mid-forward-pass — the paper claims this adds <2ms tail latency to affected requests, which is small enough to disappear inside normal p99 variance. The catch: migration is only cheap on NVLink-connected GPUs. Across nodes, DSec falls back to replicate-then-drop, which is slower but still beats static placement.
A concrete example from the paper: a 24-hour trace with 40% of requests coming from a coding-assistant frontend and 60% from a general chat interface. Under vLLM, one expert (expert #147, apparently a heavy syntax-token specialist) hit 8× the average activation rate during business hours and dragged the whole cluster's throughput down. Under DSec, that expert got replicated to three GPUs during the peak and consolidated back to one overnight. Steady-state throughput moved from 3,100 tokens/sec to 7,200 tokens/sec on the same 8×H100 hardware.
If you're serving a dense model, this changes nothing for you. Llama-3-70B, Qwen2.5-72B dense — none of this applies. But the industry is unambiguously moving toward sparse MoE for anything large: DeepSeek-V3, Mixtral, Qwen-MoE, the recent Llama-4 MoE variants, and Grok's later checkpoints are all in this territory. If you're self-hosting one of these, or if your inference bill is dominated by one, DSec-style elastic placement is about to become the assumed baseline.
The practical implication: static expert-parallel deployment is now the slow path, and if your vendor or serving stack doesn't have a migration story on the roadmap, you're looking at a 2× cost gap within a year. vLLM's issue tracker already has an open RFC referencing the DSec paper (issue #11xxx, filed within hours of the arxiv drop), and SGLang's maintainer posted on X that they've been prototyping similar ideas but hadn't published. Expect a scramble. Expect production-quality open-source implementations within 90 days.
For teams evaluating MoE inference today, the near-term move is to instrument routing distribution before you commit to a placement strategy. If your traffic is genuinely uniform, static EP is fine and simpler. If you see even 2× imbalance across experts — and most real workloads do — you're leaving throughput on the floor, and the fix is either DSec-style migration or aggressive replication of hot experts. Neither is free, but the migration approach is cheaper on HBM.
The deeper lesson here is that MoE inference is still an unsolved systems problem, and DeepSeek is currently the fastest-moving lab on the serving side, not just the training side. The gap between what the papers describe and what production systems actually do has closed to weeks. Expect the next iteration of this idea to push further: dynamic expert splitting (large experts sharded on-demand), cross-layer expert co-location based on routing correlation, and eventually predictive migration driven by prompt embeddings before the forward pass starts. The pattern — treat model structure as a runtime object, not a compile-time constant — is going to keep eating static assumptions until nothing is pinned by default.
As models get better, safe-and-secure sandboxes/environments are going to be the way forward. With the recent rise in cases where models can somehow gain access to the internet and blast past the sandbox, it's very important to have all the resources contained within the sandbox with no ac
380.000 concurrent sandboxes on 160 Epyc based server nodes. Crazy stuff
12 sandboxes per code is insane, I wonder how many of these sandboxes are idle at a time. Depending on the tasks assigned the resource requirements are different. Compare an agent doing pdf conversion and one responding to a simple question. One is cpu bound the other is mostly network wait.This is
Appears to be similar to what Google is building with ax https://github.com/google/ax
Top 10 dev stories every morning at 8am UTC. AI-curated. Retro terminal HTML email.
It seems every DeepSeek paper/patent has a huge number of authors, and this one is no exception. They couldn't even fit everyone on the page, there are 31 others not shown. This could be an asset protection strategy (i.e., human assets). Imagine if there were only 3 authors. Those authors