How to read this page. Each section starts Simple, then goes to a Deep dive: stop wherever you have what you need. The small numbers are sources: click one to open the original document. Where the maker has not published a figure -- a price, a die size, a factory address -- this page says so rather than estimate it.
1.At a glance
SimpleStart here
The V100 is a data-centre graphics chip NVIDIA announced on 10 May 2017 1. It was the first NVIDIA chip built around Tensor Cores — circuits designed specifically to speed up deep learning — and it is the chip that started NVIDIA’s run as the default hardware for training AI models.
Deep diveThe technical detail
Codename GV100, architecture Volta. Announced at Jensen Huang’s GTC 2017 keynote 1; Volta-based DGX-1, DGX Station and HGX-1 systems were slated to ship in Q3 2017 2. It shipped in 16GB and 32GB HBM2 variants, in both PCIe and SXM2 form factors 3.
2.Launch and history
SimpleStart here
NVIDIA framed Volta’s launch around a single new idea: a GPU core built just for the matrix math behind neural networks. Before V100, NVIDIA’s data-centre chips were general-purpose number-crunchers first and AI accelerators second. V100 flipped that.
Deep diveThe technical detail
NVIDIA’s own release frames Volta’s generational gains as a “5x improvement over Pascal architecture” and “15x over Maxwell” 1, naming P100 (Pascal) as the direct predecessor generation. DGX-1 (8× V100 via NVLink) and DGX Station (4× V100) were announced alongside the chip as the systems meant to put it to work immediately 2.
3.What’s inside it
SimpleStart here
V100’s headline feature was Tensor Cores: dedicated circuits for the specific kind of matrix multiplication that deep learning training and inference run millions of times over. NVIDIA said they gave “up to 12x higher peak TFLOPS for training and 6x higher peak TFLOPS for inference” compared with the previous generation 4.
Deep diveThe technical detail
Three architectural claims from NVIDIA’s own whitepaper 4:
- Tensor Cores — “new Tensor Cores designed specifically for deep learning deliver up to 12x higher peak TFLOPS for training and 6x higher peak TFLOPS for inference” versus the prior generation.
- Independent Thread Scheduling — “Volta’s new independent thread scheduling capability enables finer-grain synchronization and cooperation between parallel threads.” This broke the lock-step warp execution model NVIDIA GPUs had used since the start of CUDA, and every architecture since (Turing, Ampere, Hopper, Blackwell) has kept it.
- 2nd-generation NVLink — “delivers higher bandwidth, more links, and improved scalability for multi-GPU and multi-GPU/CPU system configurations.”
4.Full spec table
SimpleStart here
640 Tensor Cores, 5,120 CUDA cores, and either 16GB or 32GB of HBM2 memory, depending on the variant 3.
Deep diveThe technical detail
| Spec | V100 PCIe | V100 SXM2 | V100S PCIe (2019 refresh) |
|---|
| Process node | TSMC 12nm FFN (custom) 4 |
|---|
| Transistors | 21.1 billion 3 |
|---|
| Die size | 815 mm² 3 |
|---|
| CUDA cores | 5,120 3 |
|---|
| Tensor cores | 640 3 |
|---|
| Memory | 16GB or 32GB HBM2 | 16GB or 32GB HBM2 | 32GB HBM2 |
|---|
| Memory bandwidth | 900 GB/s | 1,134 GB/s | 1,134 GB/s |
|---|
| TDP | 250 W | 300 W | 250 W |
|---|
| FP64 | 7 TFLOPS | 7.8 TFLOPS | 8.2 TFLOPS |
|---|
| FP32 | 14 TFLOPS | 15.7 TFLOPS | 16.4 TFLOPS |
|---|
| Tensor (FP16) | 112 TFLOPS | 125 TFLOPS | 130 TFLOPS |
|---|
| NVLink | Not present (PCIe Gen3, 32 GB/s) | 2nd gen, 300 GB/s | Not present |
|---|
| Form factor | PCIe (FHFL) | SXM2 | PCIe |
|---|
All figures per NVIDIA’s own datasheet and whitepaper 34. FP8 was not supported — the format did not exist on Volta. An official INT8 TOPS figure was not found in the reviewed datasheet.
5.Where it’s made
SimpleStart here
V100 was made by TSMC, on a custom process NVIDIA calls 12nm FFN 4. NVIDIA designs its chips but does not operate the factories that manufacture them (a “fabless” company) — this is true of every NVIDIA chip on this site.
Deep diveThe technical detail
Process: TSMC 12nm FFN (FinFET NVIDIA, a custom high-performance variant of TSMC’s 12nm node) 4. NVIDIA’s public materials name the foundry and the process name but do not publish which of TSMC’s individual fabs produced V100 wafers — that level of detail is not published by either company.
6.Which systems use it
SimpleStart here
NVIDIA’s own systems built around V100 were DGX-1 (8 chips) and DGX Station (4 chips) 2. NVIDIA also named AWS, Google Cloud and Microsoft Azure as launch cloud partners offering V100 instances 1.
Deep diveThe technical detail
DGX-1 combines 8× V100 GPUs connected by NVLink 2. Oak Ridge National Laboratory’s Summit supercomputer, completed in 2018, was cited by NVIDIA at launch as an adopter of the Volta platform 1. Cloud availability: Amazon Web Services (P3 instances), Google Cloud, and Microsoft Azure N-series, all named in NVIDIA’s own launch release 1.
7.Official pricing
SimpleStart here
NVIDIA has not published a standalone price for the V100 chip itself. Its May 2017 release does not state a system price either — it directs readers to nvidia.com for DGX-1 and DGX Station pricing 2.
Deep diveThe technical detail
NVIDIA’s own DGX-1 announcement says pricing is available at nvidia.com/dgx-1 and nvidia.com/dgx-station, rather than stating a figure in the release itself 2. Numbers circulated by press outlets at the time (commonly cited around $149,000 for DGX-1) could not be traced to an NVIDIA-owned document in this page’s research pass, so this page does not state them as fact. Cloud providers publish their own hourly rates for V100 instances on their own pricing pages, which change over time and are not reproduced here.
9.What came before, what came next
SimpleStart here
Came before: P100 (Pascal architecture). Came after: A100 (Ampere architecture), announced May 2020 1.
Deep diveThe technical detail
NVIDIA’s own release frames Volta’s gains against “Pascal architecture” and “Maxwell” by name, making P100 the direct predecessor generation 1. V100’s direct successor is the A100 (Ampere), the next data-centre GPU generation NVIDIA announced.
10.Hidden in plain sight
SimpleStart here
V100’s quiet feature outlived its headline one. Tensor Cores got the announcement; independent thread scheduling — barely mentioned at launch — changed how every NVIDIA GPU since has executed code.
Deep diveThe technical detail
NVIDIA’s architecture whitepaper describes independent thread scheduling as enabling “finer-grain synchronization and cooperation between parallel threads” 4. It broke a rigid lock-step execution model NVIDIA GPUs had used since the beginning of CUDA — letting threads inside the same warp diverge and interleave independently. Every NVIDIA architecture released since (Turing, Ampere, Hopper, Blackwell) has kept this model, even though it was not what V100 was marketed for.