Stallion Series

Compute that cannot invent a second HBM.

On-package AI MPU. HBM bandwidth is mirrored from Gallium — Stallion does not invent a second TB/s.

Topology is non-negotiable: Host → command processor → 2D mesh NoC → SM → L1 → partitioned L2 → Gallium → cubes. HBM never feeds L1.

Editorial Stallion die with copper halo. Not a metal layer plot.

Roofline is why the width matters

S100 peak FP32 is 0.06 PFLOPS (0.15 PFLOPS vector boost). At Gallium-H4 16 TB/s sustained across the 16,384-bit bus, the arithmetic intensity threshold to leave the memory roofline is easily satisfied for both dense tensor compute (10–15 PFLOPS) and vector activation pipelines. A narrow 2048-bit lock would invert every roofline sentence.

Stallion MPU Series