1. What this article is
FreeToken is other people's software. This article is our reading of the memory problem it makes visible. FreeToken does not run on Stallion today.
Fairview did not write FreeToken, does not ship it, and is not a partner on that work. The paper is arXiv:2608.16157, FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution (Yang et al., 2026). It runs Mixture-of-Experts models on personal computers using host RAM, a VRAM expert cache, and either PCIe fills or CPU execution. It does not need Fairview silicon.
PULSE, as released, is a host runtime with A1/B1 ISA packers. It is not a compiler from MLIR to GDSII. It is not CUDA. The rest of this note is commentary: what FreeToken actually does, what still hurts for on-prem and regulated inference, and what Stallion S100 is intended to be.
2. FreeToken, as published
The paper treats a PC as one elastic inference platform rather than a small GPU with an offload afterthought. The full expert pool lives in host memory. Non-expert weights stay on the GPU. Remaining VRAM is a shared expert cache. Three mechanisms do the work:
On each decode step, cache hits run on the GPU. Residual misses are split between PCIe cache-fill and in-place CPU execution using two measured bandwidths on that machine (host processing vs. pinned PCIe transfer). The closed-form split is q* ≈ m · BP / BH. The model is not approximated; GPU and CPU partials merge.
GPU memory is not a load-time constant. At scheduler safe points the engine rebuilds the expert cache under a new VRAM budget without restarting or reloading the host expert pool. The same pool is split between KV pages and expert slots as context grows and other apps steal VRAM.
Prefill double-buffers full layers over PCIe while the GPU computes the current layer. Recurrent and hybrid-attention state is checkpointed at semantic boundaries (thinking blocks, tool calls, turns) so an agent edit resumes from the nearest surviving anchor instead of re-prefilling the whole prefix.
That is a software answer to a hardware imbalance: sparse activation makes the FLOPs fit; the full expert pool still does not. FreeToken's own evaluation is on consumer and workstation GPUs, not on Stallion.
3. The residual problem for on-prem and regulated inference
Edge-native MoE on a PC is a different job from serving the same class of model inside a regulated on-prem envelope. Four residues remain after the software scheduler has done its work. The first is a capacity problem; the rest are bandwidth and flexibility problems.
At frontier expert counts the full MoE expert pool is larger than any single accelerator can hold. FreeToken's own headline result serves a 753B model, which at FP8 is roughly three quarters of a terabyte of weights. Widening an interface does not change how much memory is attached to it. This residue is capacity, not bandwidth, and no scheduler removes it.
q* can hide a miss better than a static placement. It cannot delete the miss. Expert weights still cross a host bus or run on DRAM-bound CPU kernels. On-prem serving that cannot spill to a gaming desktop still pays that movement on every cold expert.
Agentic sessions grow KV. The expert working set does not shrink to match. Elastic software can re-split VRAM at a safe point; both pools still compete for the same device memory. A long regulated context makes the split worse, not better.
Routing graphs, expert counts, and hybrid-attention mixes move on a software clock. A model-frozen attention ASIC does not. A general GPU can run the next graph, but the expert pool still does not fit, so the host bus remains the scheduler's ceiling.
That is the gap we read in the paper, and it splits in two.
Where the working set fits and the link is the ceiling — the cold expert, the streamed KV — software can schedule around a PCIe-class bus but cannot turn that bus into on-system memory. On-prem operators who cannot send tokens off-site need a programmable matrix path whose memory controller lives next to the compute, not a host-orchestrated spill.
Where the working set does not fit at all, bandwidth is not the first constraint; capacity is. A wider interface does not make a larger pool fit, and no amount of elastic residency creates memory that was never attached. That regime needs tiered memory that keeps cold experts adjacent to compute, which is the case for the asymmetric NVM tier rather than for the HBM4 interface alone.
Stallion and Gallium are intended to answer both regimes. We are explicit below about which parts are measured and which are targets, and we would rather be corrected on the analysis than have it read as a result.
4. Intended S100 answer
Packaged bring-up die is targeted for 1H 2027, and is not the S100 SKU. Product Stallion S100 (FV-ST-S100) is 2H 2027, with Gallium H4 (FV-GL-H4). Nothing in this section is a measured result on silicon.
The intended hardware answer is a programmable matrix ISA plus an on-system memory controller. The ISA is meant to take changing MoE routing graphs without freezing a transformer topology into the die. The memory controller is meant to keep expert weights and KV traffic on-package so the scheduler is not bounded by host PCIe or CPU DRAM on every miss.
PULSE today does not implement that stack. The public SDK is a host runtime and A1/B1 packer, with GRID=4 goldens. pulse-opt MLIR lowering is not in this release. There is no path from a FreeToken graph to Stallion GDSII.
| Family target | Caption |
|---|---|
| 16.0 TB/s | TARGET / NOT MEASURED · product family |
| 2 nm-class | TARGET / NOT MEASURED · product family |
| 4.72 tensor PFLOPS / die | TARGET / NOT MEASURED · product family |
| 128k-entry KV cache | TARGET / NOT MEASURED · product family |
| 512 GB HBM4 | TARGET / NOT MEASURED · product family · Gallium H4 with S100 |
Cite FreeToken as Yang et al., arXiv:2608.16157. Do not cite this page as a FreeToken result, a Fairview benchmark, or a Berkeley collaboration.
