| NVIDIA | Rubin (Vera Rubin NVL72) | 288 GB HBM4 | 50 PFLOPS NVFP4 inference; 35 training | TSMC | Not published; rentable from about $5 per H100 GPU-hour (AWS Capacity Blocks) |
| AMD | Instinct MI455X (Helios rack) | 432 GB HBM4 | 40 PFLOPS FP4; 20 FP8 (dense/sparse not stated) | TSMC | Not published; rentable from under $2 per MI300X GPU-hour (Vultr, pre-emptible) |
| Intel | Gaudi 3 | 128 GB HBM2e | 1,678 TFLOPS BF16/FP8 (white paper); 1,835 in Intel's Hot Chips slides | TSMC (5 nm) | $125,000 list for an 8-chip kit (Intel, 2024); IndiaAI ₹153/hour per card before GST |
| Google TPU | Ironwood (TPU7x); TPU 8t/8i coming | 192 GiB HBM | 4,614 TFLOPS FP8 (Google, "peak") | Google with Broadcom; foundry not disclosed | $12.00 per chip-hour on demand (Google Cloud list price) |
| Amazon Trainium | Trainium3 (Trn3 UltraServer) | 144 GB HBM3e | 2.52 PFLOPS FP8 (AWS) | Annapurna Labs designs; foundry not disclosed | Trn3 not published; Trainium2 $2.235 per chip-hour (AWS Capacity Block) |
| Microsoft Maia | Maia 200 | 216 GB HBM3e | Over 10 PFLOPS FP4; over 5 PFLOPS FP8 (Microsoft) | TSMC (3 nm) | Not for sale or rent (used for Microsoft services) |
| Meta MTIA | MTIA 300 (MTIA 400 next) | 216 GB HBM3E | 1.2 PFLOPS FP8 (as reported from Meta's roadmap) | Meta with Broadcom; TSMC for earlier chips | Not for sale or rent (Meta internal use only) |
| Apple | M5 Ultra (Mac Studio) | Up to 512 GB unified memory, 1.2 TB/s | Not published as TOPS (last figure: M4, 38 TOPS) | Apple designs; TSMC makes | Mac Studio with M5 Ultra from $5,499 |
| Cerebras | WSE-3 (CS-3); CS-4 shipping | 44 GB SRAM on the wafer, 21 PB/s | 125 PFLOPS per WSE-3; 250 per WSE-3 Turbo | Cerebras designs; TSMC 5 nm | Cloud: gpt-oss-120b $0.35 in / $0.75 out per million tokens |
| Groq | LPU (gen 1); NVIDIA Groq 3 LPX under licence | 230 MB SRAM per chip, 80 TB/s | 750 TOPS INT8 / 188 TFLOPS FP16 per chip | Groq design; 14 nm (GlobalFoundries per press) | Cloud: GPT OSS 120B $0.15 in / $0.60 out per million tokens |
| Huawei Ascend | Ascend 910C; 950 series from 2026 | 950DT: 144 GB own HBM, 4 TB/s | 950: 1 PFLOPS FP8 / 2 PFLOPS MXFP4 | HiSilicon; factory not disclosed (press: SMIC) | Not published by the company |