A new fine-tuning method cuts the memory traffic that slows down mixture-of-experts models during inference.
Researchers call the approach MaskCoFT, short for masked co-adaptive fine-tuning. It trains a model's router and its experts together, using a learnable binary mask that limits each layer's top-K routing to a smaller subset of experts during training. The experts then adapt to the specific tokens the new routing sends their way. At inference, the mask becomes a soft prior that reranks experts without locking any of them out. Tested on Mixtral-8x7B and DeepSeek-V2-Lite with simulated GPU caches of 4 and 12 experts per layer, MaskCoFT cut expert fetches per token by 23.7% and 10.1%, and lowered time per output token by up to 16.4% and 5.5% in real offloading systems. Accuracy across nine benchmarks stayed 0.92 and 0.53 points above the unmodified base models.
The real bottleneck for MoE models like Mixtral is memory, not compute. They are too big for a single GPU, so unused experts sit in host memory and get pulled in on demand, which is slow when routing is scattershot. Earlier fixes retrained only the router to make routing more cache-friendly, but left the experts frozen, so they were stuck handling whatever tokens the new routing threw at them. Training both pieces together closes that mismatch.
The DeepSeek-V2-Lite numbers are a reminder that gains shrink as models get better at expert specialization to begin with, so don't expect a flat 20% speedup on every MoE model you throw this at.