AI Chips · Product page

NVIDIA V100 (Volta): The Complete Guide

The chip that put Tensor Cores into a data-centre GPU for the first time, in 2017. Simple to expert, every number links to its source.

Stuck at any point? Ask an AI about this page →
Ask an AI about this chip:CAGPX
NVIDIA V100 10 sections · 2 levels 4 linked sources Checked September 2026
How to read this page. Each section starts Simple, then goes to a Deep dive: stop wherever you have what you need. The small numbers are sources: click one to open the original document. Where the maker has not published a figure -- a price, a die size, a factory address -- this page says so rather than estimate it.

1.At a glance

SimpleStart here

The V100 is a data-centre graphics chip NVIDIA announced on 10 May 2017 1. It was the first NVIDIA chip built around Tensor Cores — circuits designed specifically to speed up deep learning — and it is the chip that started NVIDIA’s run as the default hardware for training AI models.

Deep diveThe technical detail

Codename GV100, architecture Volta. Announced at Jensen Huang’s GTC 2017 keynote 1; Volta-based DGX-1, DGX Station and HGX-1 systems were slated to ship in Q3 2017 2. It shipped in 16GB and 32GB HBM2 variants, in both PCIe and SXM2 form factors 3.

2.Launch and history

SimpleStart here

NVIDIA framed Volta’s launch around a single new idea: a GPU core built just for the matrix math behind neural networks. Before V100, NVIDIA’s data-centre chips were general-purpose number-crunchers first and AI accelerators second. V100 flipped that.

Deep diveThe technical detail

NVIDIA’s own release puts the generational gain narrowly, in peak teraflops rather than overall: Volta “provides a 5x improvement over Pascal™, the current-generation NVIDIA GPU architecture, in peak teraflops, and 15x over the Maxwell™ architecture” 1 — naming P100 (Pascal) as the direct predecessor generation. Volta-based DGX-1 and DGX Station systems were announced alongside the chip as the machines meant to put it to work immediately; that release quantifies them by CPU-equivalence rather than GPU count, at “the computing capacity of 800 CPUs” and 400 CPUs respectively 2.

3.What’s inside it

SimpleStart here

V100’s headline feature was Tensor Cores: dedicated circuits for the specific kind of matrix multiplication that deep learning training and inference run millions of times over. NVIDIA said they gave “up to 12x higher peak TFLOPS for training and 6x higher peak TFLOPS for inference” compared with the previous generation 4.

Deep diveThe technical detail

Three architectural claims from NVIDIA’s own whitepaper 4:

  • Tensor Cores — “new Tensor Cores designed specifically for deep learning deliver up to 12x higher peak TFLOPS for training and 6x higher peak TFLOPS for inference” versus the prior generation.
  • Independent Thread Scheduling — “Volta’s new independent thread scheduling capability enables finer-grain synchronization and cooperation between parallel threads.” This broke the lock-step warp execution model NVIDIA GPUs had used since the start of CUDA, and every architecture since (Turing, Ampere, Hopper, Blackwell) has kept it.
  • 2nd-generation NVLink — “delivers higher bandwidth, more links, and improved scalability for multi-GPU and multi-GPU/CPU system configurations.”

4.Full spec table

SimpleStart here

640 Tensor Cores, 5,120 CUDA cores, and either 16GB or 32GB of HBM2 memory, depending on the variant — NVIDIA’s own datasheet lists both capacities for the PCIe card and 32GB for the SXM2 module 3.

Deep diveThe technical detail
SpecV100 PCIeV100 SXM2V100S PCIe (2019 refresh)
Process nodeTSMC 12nm FFN (custom) 4
Transistors21.1 billion 4
Die size815 mm² 4
CUDA cores5,120 3
Tensor cores640 3
Memory16GB or 32GB HBM232GB HBM2 (a 16GB SXM2 also shipped; this datasheet lists only 32GB)32GB HBM2
Memory bandwidth900 GB/s900 GB/s1,134 GB/s
TDP250 W300 W250 W
FP647 TFLOPS7.8 TFLOPS8.2 TFLOPS
FP3214 TFLOPS15.7 TFLOPS16.4 TFLOPS
Tensor (FP16)112 TFLOPS125 TFLOPS130 TFLOPS
NVLinkNot present (PCIe Gen3, 32 GB/s)2nd gen, 300 GB/sNot present
Form factorPCIe (FHFL)SXM2PCIe

All figures per NVIDIA’s own datasheet and whitepaper 34. FP8 was not supported — the format did not exist on Volta. An official INT8 TOPS figure was not found in the reviewed datasheet.

5.Where it’s made

SimpleStart here

V100 was made by TSMC, on a custom process NVIDIA calls 12nm FFN 4. NVIDIA designs its chips but does not operate the factories that manufacture them (a “fabless” company) — this is true of every NVIDIA chip on this site.

Deep diveThe technical detail

Process: TSMC 12nm FFN (FinFET NVIDIA, a custom high-performance variant of TSMC’s 12nm node) 4. NVIDIA’s public materials name the foundry and the process name but do not publish which of TSMC’s individual fabs produced V100 wafers — that level of detail is not published by either company.

6.Which systems use it

SimpleStart here

NVIDIA’s own systems built around V100 were DGX-1 and DGX Station, announced with the chip and expected to ship in the third quarter of 2017 2. NVIDIA also named AWS, Google Cloud and Microsoft Azure in its launch release — as partners intending to offer Volta, not as providers with instances already running: Google said it “plan[ned] to offer” Volta GPUs, and AWS spoke of its next GPU instance family “when Volta becomes available later in the year” 1.

Deep diveThe technical detail

Oak Ridge National Laboratory’s Summit supercomputer, completed in 2018, was cited by NVIDIA at launch as an adopter of the Volta platform 1. The DGX-1 and DGX Station configurations most often quoted for this chip — eight V100s joined by NVLink, and four V100s respectively — are accurate, but NVIDIA’s own launch releases state neither the GPU counts nor NVLink, so this page does not cite them for those figures 2.

7.Official pricing

SimpleStart here

NVIDIA has not published a standalone price for the V100 chip itself. Its May 2017 release does not state a system price either — it directs readers to nvidia.com for DGX-1 and DGX Station pricing 2.

Deep diveThe technical detail

NVIDIA’s own DGX-1 announcement says pricing is available at nvidia.com/dgx-1 and nvidia.com/dgx-station, rather than stating a figure in the release itself 2. Numbers circulated by press outlets at the time (commonly cited around $149,000 for DGX-1) could not be traced to an NVIDIA-owned document in this page’s research pass, so this page does not state them as fact. Cloud providers publish their own hourly rates for V100 instances on their own pricing pages, which change over time and are not reproduced here.

8.Real-world performance

SimpleStart here

NVIDIA’s own whitepaper claims “up to 12x higher peak TFLOPS for training and 6x higher peak TFLOPS for inference” over the prior generation 4 — a company-stated figure, not an independent benchmark.

Deep diveThe technical detail

All performance figures on this page (112–130 TFLOPS Tensor FP16 depending on variant) are NVIDIA’s own specification numbers 3, not third-party benchmark results. No independent MLPerf-style benchmark citation for V100 was confirmed in this research pass — MLPerf’s first round post-dates V100’s launch and V100 appears in later rounds as a reference point rather than this page’s primary performance source.

9.What came before, what came next

SimpleStart here

Came before: P100 (Pascal architecture). Came after: A100 (Ampere architecture), covered on its own page on this site.

Deep diveThe technical detail

NVIDIA’s own launch release names Pascal and Maxwell as the architectures Volta is measured against, making P100 the direct predecessor generation 1. V100’s successor is the A100 (Ampere), announced three years later and covered on its own page here — the 2017 Volta release naturally says nothing about it, so this page does not cite it for that.

10.Hidden in plain sight

SimpleStart here

V100’s quiet feature outlived its headline one. Tensor Cores got the announcement; independent thread scheduling — barely mentioned at launch — changed how every NVIDIA GPU since has executed code.

Deep diveThe technical detail

NVIDIA’s architecture whitepaper describes independent thread scheduling as enabling “finer-grain synchronization and cooperation between parallel threads” 4. It broke a rigid lock-step execution model NVIDIA GPUs had used since the beginning of CUDA — letting threads inside the same warp diverge and interleave independently. Every NVIDIA architecture released since (Turing, Ampere, Hopper, Blackwell) has kept this model, even though it was not what V100 was marketed for.

11.Sources

4 sources, checked September 2026. Where NVIDIA or another maker has not published a figure, this page says so rather than estimate it.

  1. NVIDIA Newsroom: NVIDIA Launches Revolutionary Volta GPU Platform, Fueling Next Era of AI and High Performance ComputingOfficial
  2. NVIDIA Newsroom: NVIDIA Advances AI Computing Revolution with New Volta-Based DGX SystemsOfficial
  3. NVIDIA: Tesla V100 GPU DatasheetOfficial
  4. NVIDIA: NVIDIA Tesla V100 GPU Architecture Whitepaper (WP-08608-001_v1.1)Official

Ask an AI about this page

Opens your assistant with this page as the source, and a question rather than a summary. It will ask what you are building before it answers.

ChatGPTClaudeGeminiPerplexityGrok

Nothing is sent from here. The link carries only this page’s title and address.