How to read this page. Each section starts Simple, then goes to a Deep dive: stop wherever you have what you need. The small numbers are sources: click one to open the original document. Where the maker has not published a figure -- a price, a die size, a factory address -- this page says so rather than estimate it.
1.At a glance
SimpleStart here
The V100 is a data-centre graphics chip NVIDIA announced on 10 May 2017 1. It was the first NVIDIA chip built around Tensor Cores — circuits designed specifically to speed up deep learning — and it is the chip that started NVIDIA’s run as the default hardware for training AI models.
Deep diveThe technical detail
Codename GV100, architecture Volta. Announced at Jensen Huang’s GTC 2017 keynote 1; Volta-based DGX-1, DGX Station and HGX-1 systems were slated to ship in Q3 2017 2. It shipped in 16GB and 32GB HBM2 variants, in both PCIe and SXM2 form factors 3.
2.Launch and history
SimpleStart here
NVIDIA framed Volta’s launch around a single new idea: a GPU core built just for the matrix math behind neural networks. Before V100, NVIDIA’s data-centre chips were general-purpose number-crunchers first and AI accelerators second. V100 flipped that.
Deep diveThe technical detail
NVIDIA’s own release puts the generational gain narrowly, in peak teraflops rather than overall: Volta “provides a 5x improvement over Pascal™, the current-generation NVIDIA GPU architecture, in peak teraflops, and 15x over the Maxwell™ architecture” 1 — naming P100 (Pascal) as the direct predecessor generation. Volta-based DGX-1 and DGX Station systems were announced alongside the chip as the machines meant to put it to work immediately; that release quantifies them by CPU-equivalence rather than GPU count, at “the computing capacity of 800 CPUs” and 400 CPUs respectively 2.
3.What’s inside it
SimpleStart here
V100’s headline feature was Tensor Cores: dedicated circuits for the specific kind of matrix multiplication that deep learning training and inference run millions of times over. NVIDIA said they gave “up to 12x higher peak TFLOPS for training and 6x higher peak TFLOPS for inference” compared with the previous generation 4.
Deep diveThe technical detail
Three architectural claims from NVIDIA’s own whitepaper 4:
- Tensor Cores — “new Tensor Cores designed specifically for deep learning deliver up to 12x higher peak TFLOPS for training and 6x higher peak TFLOPS for inference” versus the prior generation.
- Independent Thread Scheduling — “Volta’s new independent thread scheduling capability enables finer-grain synchronization and cooperation between parallel threads.” This broke the lock-step warp execution model NVIDIA GPUs had used since the start of CUDA, and every architecture since (Turing, Ampere, Hopper, Blackwell) has kept it.
- 2nd-generation NVLink — “delivers higher bandwidth, more links, and improved scalability for multi-GPU and multi-GPU/CPU system configurations.”
4.Full spec table
SimpleStart here
640 Tensor Cores, 5,120 CUDA cores, and either 16GB or 32GB of HBM2 memory, depending on the variant — NVIDIA’s own datasheet lists both capacities for the PCIe card and 32GB for the SXM2 module 3.
Deep diveThe technical detail
| Spec | V100 PCIe | V100 SXM2 | V100S PCIe (2019 refresh) |
|---|
| Process node | TSMC 12nm FFN (custom) 4 |
|---|
| Transistors | 21.1 billion 4 |
|---|
| Die size | 815 mm² 4 |
|---|
| CUDA cores | 5,120 3 |
|---|
| Tensor cores | 640 3 |
|---|
| Memory | 16GB or 32GB HBM2 | 32GB HBM2 (a 16GB SXM2 also shipped; this datasheet lists only 32GB) | 32GB HBM2 |
|---|
| Memory bandwidth | 900 GB/s | 900 GB/s | 1,134 GB/s |
|---|
| TDP | 250 W | 300 W | 250 W |
|---|
| FP64 | 7 TFLOPS | 7.8 TFLOPS | 8.2 TFLOPS |
|---|
| FP32 | 14 TFLOPS | 15.7 TFLOPS | 16.4 TFLOPS |
|---|
| Tensor (FP16) | 112 TFLOPS | 125 TFLOPS | 130 TFLOPS |
|---|
| NVLink | Not present (PCIe Gen3, 32 GB/s) | 2nd gen, 300 GB/s | Not present |
|---|
| Form factor | PCIe (FHFL) | SXM2 | PCIe |
|---|
All figures per NVIDIA’s own datasheet and whitepaper 34. FP8 was not supported — the format did not exist on Volta. An official INT8 TOPS figure was not found in the reviewed datasheet.
5.Where it’s made
SimpleStart here
V100 was made by TSMC, on a custom process NVIDIA calls 12nm FFN 4. NVIDIA designs its chips but does not operate the factories that manufacture them (a “fabless” company) — this is true of every NVIDIA chip on this site.
Deep diveThe technical detail
Process: TSMC 12nm FFN (FinFET NVIDIA, a custom high-performance variant of TSMC’s 12nm node) 4. NVIDIA’s public materials name the foundry and the process name but do not publish which of TSMC’s individual fabs produced V100 wafers — that level of detail is not published by either company.
6.Which systems use it
SimpleStart here
NVIDIA’s own systems built around V100 were DGX-1 and DGX Station, announced with the chip and expected to ship in the third quarter of 2017 2. NVIDIA also named AWS, Google Cloud and Microsoft Azure in its launch release — as partners intending to offer Volta, not as providers with instances already running: Google said it “plan[ned] to offer” Volta GPUs, and AWS spoke of its next GPU instance family “when Volta becomes available later in the year” 1.
Deep diveThe technical detail
Oak Ridge National Laboratory’s Summit supercomputer, completed in 2018, was cited by NVIDIA at launch as an adopter of the Volta platform 1. The DGX-1 and DGX Station configurations most often quoted for this chip — eight V100s joined by NVLink, and four V100s respectively — are accurate, but NVIDIA’s own launch releases state neither the GPU counts nor NVLink, so this page does not cite them for those figures 2.
7.Official pricing
SimpleStart here
NVIDIA has not published a standalone price for the V100 chip itself. Its May 2017 release does not state a system price either — it directs readers to nvidia.com for DGX-1 and DGX Station pricing 2.
Deep diveThe technical detail
NVIDIA’s own DGX-1 announcement says pricing is available at nvidia.com/dgx-1 and nvidia.com/dgx-station, rather than stating a figure in the release itself 2. Numbers circulated by press outlets at the time (commonly cited around $149,000 for DGX-1) could not be traced to an NVIDIA-owned document in this page’s research pass, so this page does not state them as fact. Cloud providers publish their own hourly rates for V100 instances on their own pricing pages, which change over time and are not reproduced here.
9.What came before, what came next
SimpleStart here
Came before: P100 (Pascal architecture). Came after: A100 (Ampere architecture), covered on its own page on this site.
Deep diveThe technical detail
NVIDIA’s own launch release names Pascal and Maxwell as the architectures Volta is measured against, making P100 the direct predecessor generation 1. V100’s successor is the A100 (Ampere), announced three years later and covered on its own page here — the 2017 Volta release naturally says nothing about it, so this page does not cite it for that.
10.Hidden in plain sight
SimpleStart here
V100’s quiet feature outlived its headline one. Tensor Cores got the announcement; independent thread scheduling — barely mentioned at launch — changed how every NVIDIA GPU since has executed code.
Deep diveThe technical detail
NVIDIA’s architecture whitepaper describes independent thread scheduling as enabling “finer-grain synchronization and cooperation between parallel threads” 4. It broke a rigid lock-step execution model NVIDIA GPUs had used since the beginning of CUDA — letting threads inside the same warp diverge and interleave independently. Every NVIDIA architecture released since (Turing, Ampere, Hopper, Blackwell) has kept this model, even though it was not what V100 was marketed for.