Stallion · FV-ST-S100
Stallion S100
Training-class flagship. Consumes 8-stack Gallium-H4 (16,384-bit bus). 350W air-cooled envelope, up to 100B MoE capacity, PyTorch-Triton compiler native execution.
Flagship

- Peak FP32
- 0.15 PFLOPS58.9824 exact
- Tensor path
- 4.72 PFLOPSdesign-target MMA rate, not a memory identity
- I min (FP32)
- 7.20 FLOP/Bat 16 TB/s Gallium-H4
- Threads in flight
- 262,144128 × 64 × 32
| Parameter | Value |
|---|---|
| SKU | FV-ST-S100 |
| SMs | 128 |
| Warps / SM | 64 (dual-issue) |
| SM clock | 1.80 GHz |
| FP32 ALUs / SM | 128 |
| MMA elements / clk / SM | 2048 |
| L1 + shared / SM | 128 KiB |
| L2 (partitioned) | 96 MiB |
| NoC | 2D mesh |
| Host | PCIe Gen6 x16 + CXL 3.2 |
| UCIe-E | 32 GT/s die-to-die (sidecar / I/O, not HBM) |
| Package / node | 3Dx3D Heterogeneous Packaging (TSMC-SoIC + Glass Core Substrate) · 2 nm GAAFET nanosheet |
| TDP island | 350 W |
| Bound Gallium cubes | 8 × 16-hi × 2048-bit (16384-bit bus) |
| Mirrored HBM bandwidth | 16 TB/s |
| Package capacity | 512 GB Unified Memory |
Stallion does not own B_agg
The MPU may print Gallium’s compiled aggregate for roofline. A different stack count, width, or attach mode is an error, not a richer-TB/s win.