Stallion · FV-ST-S100

Stallion S100

Training-class flagship. Consumes 8-stack Gallium-H4 (16,384-bit bus). 350W air-cooled envelope, up to 100B MoE capacity, PyTorch-Triton compiler native execution.

Flagship

Stallion compute die, editorial 3D view with copper halo
Peak FP32
0.15 PFLOPS58.9824 exact
Tensor path
4.72 PFLOPSdesign-target MMA rate, not a memory identity
I min (FP32)
7.20 FLOP/Bat 16 TB/s Gallium-H4
Threads in flight
262,144128 × 64 × 32
MPU lock. HBM geometry is mirrored from the bound Gallium card.
ParameterValue
SKUFV-ST-S100
SMs128
Warps / SM64 (dual-issue)
SM clock1.80 GHz
FP32 ALUs / SM128
MMA elements / clk / SM2048
L1 + shared / SM128 KiB
L2 (partitioned)96 MiB
NoC2D mesh
HostPCIe Gen6 x16 + CXL 3.2
UCIe-E32 GT/s die-to-die (sidecar / I/O, not HBM)
Package / node3Dx3D Heterogeneous Packaging (TSMC-SoIC + Glass Core Substrate) · 2 nm GAAFET nanosheet
TDP island350 W
Bound Gallium cubes8 × 16-hi × 2048-bit (16384-bit bus)
Mirrored HBM bandwidth16 TB/s
Package capacity512 GB Unified Memory

Stallion does not own B_agg

The MPU may print Gallium’s compiled aggregate for roofline. A different stack count, width, or attach mode is an error, not a richer-TB/s win.

All Stallion SKUsDesign-in
FV-ST-S100 Stallion S100