
Intel Gaudi 3 AI Accelerator
Intel's third-gen AI accelerator optimized for LLM training and inference. Cost-effective alternative with Ethernet-based scaling. Contact us for evaluation units and pricing.
Price Available Upon Request
Custom configuration and pricing
What it's for
Intel Gaudi 3 AI Accelerator is built for training and inference of large language models and generative AI. Built on TSMC's 5nm process, Gaudi 3 combines 64 tensor processor cores, 8 matrix multiplication engines, and 128GB HBM2e at 3.7TB/s. Its real differentiator: it scales on standard 200GbE Ethernet instead of proprietary interconnects, which changes the networking bill of materials. Contact us for evaluation units and a TCO comparison.
Key Features
- ✓64 tensor processor cores (TPCs) with 8 matrix multiplication engines (MMEs)
- ✓128GB HBM2e memory with 3.7TB/s bandwidth
- ✓1.8 petaFLOPS FP8 and BF16 compute performance
- ✓Integrated 24x 200 GbE networking for scale-out
- ✓96MB on-die SRAM cache for low-latency operations
- ✓14 integrated media engines (H.265, H.264, JPEG, VP9)
- ✓TSMC 5nm process node technology
- ✓Available in PCIe (600W) and OAM (900W) form factors
Use Cases
- →Large language model training (LLaMA, GPT-style models)
- →Generative AI inference at scale
- →Computer vision and video processing
- →Recommendation systems and embeddings
- →Multi-modal AI workloads
- →Cost-optimized AI infrastructure deployment
Technical Specifications
| Architecture | Gaudi 3 (TSMC 5nm) |
| GPU Memory | 128 GB HBM2e |
| Memory Bandwidth | 3.7 TB/s |
| Tensor Processor Cores | 64 TPCs |
| Matrix Multiply Engines | 8 MMEs |
| FP8 Performance | 1.8 petaFLOPS |
| BF16 Performance | 1.8 petaFLOPS |
| On-die SRAM | 96 MB |
| Network Interfaces | 24x 200 GbE |
| Media Engines | 14 (H.265/H.264/JPEG/VP9) |
| Max TDP (PCIe) | 600W |
| Max TDP (OAM) | 900W |
| Thermal Solution | Air cooling (PCIe) / Liquid (OAM) |
| Form Factor | PCIe Gen5 AIC |
| PCIe Interface | PCIe Gen5 x16 |
Related Products

NVIDIA H100 80GB PCIe Gen5
NVIDIA H100 Tensor Core GPU in PCIe form factor with 80GB HBM3 memory. Ideal for deploying AI inference and training in standard servers without NVLink clustering requirements. Contact our sales team for volume pricing and current lead time.
From
$29,999

NVIDIA H100 80GB SXM5
NVIDIA H100 Tensor Core GPU in SXM5 form factor with NVLink for multi-GPU scaling. Designed for HGX server platforms and large-scale AI training clusters. Enterprise volume discounts available - contact sales for custom configurations.
From
$32,999

NVIDIA H200 141GB HBM3e SXM5
Hopper GPU with 141GB HBM3e and 4.8TB/s bandwidth — the memory-bandwidth pick for LLM inference. Contact us for current lead time.
From
$39,999

NVIDIA B200 192GB Blackwell
Blackwell architecture with 192GB HBM3e and FP4 precision for next-gen AI. Contact us for availability and pricing.
Request pricing
Recommended Reading
The adjacent decisions buyers usually miss
Relevant guides tied to this product's real deployment questions: VRAM, power, cooling, financing, and market timing.

How to Size a GPU Cluster for LLM Training
Learn how to calculate the right number of GPUs, memory, and networking for training large language models efficiently.

AMD MI300X: The Cost-Effective Alternative for GenAI
AMD's MI300X offers 192GB of HBM3 memory at $14,999 — less than half the price of NVIDIA's H100. Here's what you need to know.

AI Cluster Networking Guide: InfiniBand vs Ethernet for GPU Training
Network design is often the bottleneck in distributed training. Learn when InfiniBand is worth it, when Ethernet is enough, and how to think about oversubscription, storage traffic, and scale-out planning.