Stallion Series
On-package AI MPU. HBM bandwidth is mirrored from Gallium — Stallion does not invent a second TB/s.
Topology is non-negotiable: Host → command processor → 2D mesh NoC → SM → L1 → partitioned L2 → Gallium → cubes. HBM never feeds L1.
Editorial Stallion die with copper halo. Not a metal layer plot.
Flagship
Training-class flagship. Consumes 8-stack Gallium-H4 (16,384-bit bus). 350W air-cooled envelope, up to 100B MoE capacity, PyTorch-Triton compiler native execution.
Inference
Inference SKU. Fatter MMA, same Gallium 8-stack geometry.
S100 peak FP32 is 0.06 PFLOPS (0.15 PFLOPS vector boost). At Gallium-H4 16 TB/s sustained across the 16,384-bit bus, the arithmetic intensity threshold to leave the memory roofline is easily satisfied for both dense tensor compute (10–15 PFLOPS) and vector activation pipelines. A narrow 2048-bit lock would invert every roofline sentence.