Skip to content
Fairview Semiconductor
FairView Semi›Stallion Series
2NM GAAFET · 4.72 Dense PFLOPS (FP8 / FP4) · EARLY ACCESS

StallionSeries AI MPU

4.72 Dense PFLOPS (FP8 / FP4) of FP8 tensor compute across 185 billion transistors on 2nm GAAFET silicon. Fairview increases active compute duty cycles from 30% to >82%, driving a 65.5% reduction in TCO per token. Unlike model-frozen ASICs that hardwire transformer attention gates into brittle silicon, Fairview achieves 0.05 pJ/bit interconnect efficiency via 3Dx3D Glass Core packaging and 16.0 TB/s Gallium HBM4 memory while retaining 100% software agility.

Request Datasheet ›Explore Use Cases ›View Benchmarks
4.72PFLOPS
FP8 Tensor Compute
2nmGAAFET
Silicon Process Node
185BTransistors
Die Integration
Technical Specifications

Engineered for Hyperscale Dominance.

Complete engineering specifications for the Stallion Series AI MPU — built for ML researchers, MPU architects, and hyperscaler datacenter leads.

Compute Architecture
Process Node2nm-Class GAAFET (TSMC A16/N2P) + BSPDN SuperPower Rail
Power DeliveryBackside Power Delivery Network (Zero IR Drop)
Transistor Count185 Billion Integrated Transistors
Die ArchitectureDual-Compute Chiplet Array (2x 410 mm²)
Matrix Execution Units144 MEUs (Neural Compute Tiles)
Systolic Tensor Engines576 Matrix Processing Units
Clock Speed (Boost)2.4 GHz Sustained
Precision & Throughput
FP32 Vector Compute0.15 PFLOPS
FP16 / BF16 Tensor1.18 PFLOPS (FP16/BF16)
FP8 Tensor (Dense)10–15 Sustained PFLOPS (Dense FP8/FP4)
FP8 Tensor (Sparsity)20–30 PFLOPS (2:4 Sparsity)
FP4 / INT4 Tensor9.44 PFLOPS (9,437 TOPs)
Precision EngineAdaptive FP8 / FP4 Dynamic Scaler (2:4 Sparsity)
Memory & Interconnect
Memory AttachedGallium HBM4 (512 GB Unified Pool)
Memory Bandwidth16.0 TB/s Saturated Stream
Memory Bus Width16,384-bit Ultra-Wide Parallel Bus
Host InterfacePCIe Gen 6.0 x16 (128 GB/s Full-Duplex)
CXL SupportCXL 3.2 Direct & Disaggregated Memory Pool
Die-to-Memory AttachTSMC-SoIC Direct Cu-Cu Hybrid Bonding (<1µm)
Clustering & Advanced Packaging
Scale-Up InterconnectFV-Link 4.0 Co-Packaged Optics (CPO) 1.8 TB/s
Optical DomainUp to 512 MPUs Coherent (256 TB Direct / 1.5 PB CXL Fabric)
Package Technology3Dx3D Heterogeneous Glass Core Substrate
Substrate Line/Space< 2 µm Lithography (Zero Warpage @ 350W Air-Cooled)
TDP350 W (Air-Cooled, 1U-Ready Thermal Envelope)
Form FactorOAM 2.0 / SXM Enterprise Blade Compatible
Memory Architecture

16.0 TB/s.
Memory Redefined.

The Stallion MPU pairs with Gallium HBM4 via 3Dx3D Heterogeneous Packaging (TSMC-SoIC + Glass Core Substrate) — 8 stacks, 16,384-bit bus, direct die-to-memory attach. No DRAM bottleneck. No PCIe latency. Pure bandwidth for attention layers, KV-cache, and weight streaming.

Stallion + Gallium HBM416.0 TB/s
Competitor A (Legacy 8-Stack System)8.000 TB/s
Competitor B (Legacy 6-Stack System)4.800 TB/s
Competitor C (Standard Legacy Memory)3.350 TB/s

* Competitor figures based on publicly available specifications. Actual performance may vary.

3DX3D HETEROGENEOUS PACKAGING (TSMC-SOIC + GLASS CORE SUBSTRATE)
STALLION DUAL-COMPUTE ARRAY
2nm-Class GAAFET (TSMC A16) · Dual-Compute Array (2x 410 mm²) · 185B Integrated Transistors
Systolic Tensor Engines
MEU Compute Tiles
NoC Mesh
L2 Cache
HBM4Stack 1
64 GB
HBM4Stack 2
64 GB
HBM4Stack 3
64 GB
HBM4Stack 4
64 GB
HBM4Stack 5
64 GB
HBM4Stack 6
64 GB
HBM4Stack 7
64 GB
HBM4Stack 8
64 GB
Silicon Architecture

Four Breakthroughs. One Compute Package.

How 2nm-Class GAAFET (TSMC A16), 4th-gen Sparse Matrix Execution Units, FV-Link 4.0 CPO, and 3Dx3D Heterogeneous Packaging (TSMC-SoIC + Glass Core Substrate) combine to make Stallion the world's most powerful AI accelerator.

01
185B
integrated transistors across dual-compute array

2nm-Class GAAFET (TSMC A16) + BSPDN

Backside Power Delivery Network & Nanosheets

Stallion decouples signal and power routing using Backside Power Delivery (BSPDN). Buried power rails and nano-TSVs deliver pristine current from the wafer backside, reducing IR voltage drop by 25% and leaving the entire frontside metal stack dedicated to 16,384-bit memory and logic interconnect.

2nm-Class GAAFET (A16/N2P)BSPDNSuperPower RailZero IR Drop
02
4.72
Dense PFLOPS (FP8 / FP4)

4th-Gen Sparse Tensor Engine

Native FP8, FP4 & Micro-Scaling Precision

576 Matrix Processing Units with dynamic precision switching. The hardware transformer engine automatically executes FP8 (E4M3/E5M2) and FP4 micro-scaling representations with hardware-accelerated 2:4 structured sparsity doubling throughput to 4.72 Dense PFLOPS (FP8 / FP4).

4.72 Dense PFLOPS (FP8 / FP4)Micro-Scaling
03
1.8
TB/s on-package optical engine

FV-Link 4.0 Co-Packaged Optics (CPO)

Embedded Photonic Waveguides in Glass (0.05 dB/cm @ 1550nm)

Direct 1550nm optical routing laser-inscribed inside the 3.2 ppm/K glass core substrate (0.05 dB/cm propagation loss), bypassing copper SerDes heat and enabling <20ns inter-chassis optical clustering across up to 512 MPUs in a unified 256 TB direct HBM4 / 1.5 PB disaggregated CXL memory domain.

Glass PhotonicsEmbedded Waveguides (0.05 dB/cm)1.8 TB/s CPO512 MPU Coherent
04
16.0
TB/s memory-to-compute bus

3Dx3D Heterogeneous Glass Core Substrate

Sub-2µm Line/Space & Direct Cu-Cu Hybrid Bonding

Stallion and Gallium HBM4 mate over an ultra-flat Glass Core Substrate with TSMC-SoIC direct Cu-Cu hybrid bonding (<1µm pitch). Glass eliminates package warpage at 350W+ air-cooled thermal loads, maintaining sub-8ns latency at 0.05 pJ/bit.

Glass SubstrateDirect Cu-Cu (<1µm)TSMC-SoICSub-8ns Latency
05
16,000
GB/s direct expert swap bandwidth

Dynamic MoE & Elastic Sparsity Engine

Zero-Latency Expert Offloading & On-System Memory Placement

Native microarchitectural support for dynamic Mixture-of-Experts (MoE) routing. Eliminates the pipeline bubbles inherent in PCIe-based offloading frameworks (such as the elastic expert residency in Yang et al., arXiv:2608.16157) by feeding the matrix array directly via vertical Cu-Cu interconnects at 16.0 TB/s without host CPU arbitration. Scales to 100B-parameter MoE capacity with PyTorch-Triton compiler native execution mapping dynamic attention and expert-dispatch kernels directly onto the MEUs.

Dynamic MoEFreeToken Co-DesignZero PCIe Bubble16.0 TB/s Swap100B MoEPyTorch-Triton
Measured Performance

Real Workloads. Verified Gains.

Benchmarks across LLM inference, training time-to-accuracy, and sparse matrix operations. All numbers measured on pre-production Stallion silicon.

LLaMA 3 70B FP8 Inference Throughput
tokens / sec / MPU (normalized)
97/ 100
Mixture of Experts (MoE) Routing Latency
% Faster vs. Hopper
93/ 100
BERT-Large Training Time to Accuracy
Hours (lower = better)
95/ 100
Attention Layer Execution Efficiency
% Compute Peak Achieved
99/ 100
Sparse Tensor Core GEMM Peak
% of Theoretical FLOPS
96/ 100
Multi-MPU AllReduce Bandwidth
% of FV-Link Peak
98/ 100

Pre-production Stallion Series silicon measured at FairView Semiconductor benchmarking lab, Q3 2026. LLaMA 3 70B measured at batch size 32, FP8 precision with Flash Attention 3. Full benchmark reproducibility scripts available under NDA with early access agreement.

Early Access Program

Deploy Stallion in Your
AI Datacenter.

Early access silicon, software SDK (FV-CUDA compatible), and engineering sample boards are available for qualifying hyperscalers and AI research institutions.

FairView Semiconductor — Stallion AI MPU & Gallium HBM4