Stallion Series

Compute that cannot invent a second HBM.

On-package AI GPU. HBM bandwidth is mirrored from Gallium — Stallion does not invent a second TB/s.

Topology is non-negotiable: Host → command processor → 2D mesh NoC → SM → L1 → partitioned L2 → Gallium → cubes. HBM never feeds L1.

Roofline is why the width matters

S100 peak FP32 is 58.9824 TFLOPS, displayed 59.0. At Gallium-H4 8.192 TB/s the intensity to leave the memory roof is 7.20 FLOP/B. A 1024-bit lock at the same clock would invert every roofline sentence.