Skip to content
— CH. 1 · INTRODUCTION —

Tensor Processing Unit

16 min listen · Ch. 1 of 8
8 sections
  • The Tensor Processing Unit helped power AlphaGo during its 2016 series against Lee Sedol, one of the strongest Go players alive. Google did not reveal the chip's existence until later that year. In May, at the Google I/O conference, it admitted the TPU had already been running inside company data centers for over a year. A processor nobody outside Google had heard of had already changed how the company handled some of its hardest computing problems. What makes this chip different from the graphics processors and general-purpose chips already on the market? Who built it, and what happened when a rival company later accused Google of patent infringement over its design?

  • Eight bits of precision, in many cases, is all a Tensor Processing Unit needs to complete a calculation, far less than a graphics processor typically uses. Compared with a GPU, the TPU trades precision for volume, packing in more input and output operations for every joule of energy it spends. It also skips hardware built for rasterization and texture mapping, features a graphics card needs for rendering images but a neural network never touches.

    According to Norman Jouppi, the TPU's chips are mounted inside a heatsink assembly compact enough to slide into a hard drive slot within a data center rack. That physical detail says as much about the chip's purpose as any spec sheet: it was built to be racked by the thousands, not plugged into a single desktop tower.

    Convolutional neural networks run best on a TPU, while some fully connected neural networks favor a GPU instead. Recurrent neural networks, meanwhile, can play to the strengths of an ordinary CPU. No single processor wins every job.

    Other companies were racing to build their own accelerators too, though many aimed theirs at embedded devices and robotics rather than data centers.

  • In 2013, Google recruited Amir Salek to build the company's first custom silicon development capability for its data centers. Salek became founder and head of Custom Silicon for Google Technical Infrastructure and Google Cloud. He went on to lead development of the original TPU, Google's first production chip, along with TPUv2, which became the industry's first production deep-learning training chip. His team followed with TPUv3 and TPUv4, and also built the Edge-TPU and other silicon projects, including the VCU, the IPU, and OpenTitan.

    Jonathan Ross was one of the original TPU engineers before he left to found the chip company Groq. According to Ross, three separate teams inside Google were developing AI accelerators at the same time. The systolic array design that became the TPU was the one Google ultimately chose.

    Norman P. Jouppi served as tech lead and principal architect for the TPU, leading its design, verification, and deployment to production in just 15 months. He was lead author of the 2017 paper In-Datacenter Performance Analysis of a Tensor Processing Unit, presented at the 44th International Symposium on Computer Architecture, known as ISCA 2017. The paper showed the TPU reaching 15-30 times the performance of contemporary CPUs and GPUs, and 30-80 times their performance per watt. That finding established the TPU as a foundational platform for neural network inference across Google's production services.

    The chip was built specifically for TensorFlow, Google's symbolic math library for machine learning applications such as neural networks. Google's 2017 paper traced the TPU's lineage back to systolic matrix multipliers of similar design built in the 1990s. Even so, as of 2017 Google still relied on ordinary CPUs and GPUs for other kinds of machine learning work.

    None of that hardware would have reached a data center rack without a manufacturing partner willing to turn Google's blueprints into silicon.

  • Every calculation in the first-generation chip runs at 8-bit precision, using CISC instructions sent from the host processor over a PCIe 3.0 bus. It was built on a 28-nanometer process with a die no larger than 331 square millimeters, running at a 700 megahertz clock speed. Its thermal design power sits between 28 and 40 watts. Inside the package are 28 mebibytes of on-chip memory and 4 mebibytes of 32-bit accumulators, fed by a 256-by-256 systolic array of 8-bit multipliers. Alongside that sits 8 gibibytes of dual-channel 2133 megahertz DDR3 memory, offering 34 gigabytes per second of bandwidth. Instructions move data to and from the host, run matrix multiplications or convolutions, and apply activation functions.

    The second-generation TPU was announced in May 2017. Google said memory bandwidth had limited the first generation, and 16 gibibytes of new High Bandwidth Memory raised that figure to 600 gigabytes per second. Performance rose to 45 teraFLOPS, and four of the chips combined into a module reaching 180 teraFLOPS. Sixty-four of those modules were then assembled into 256-chip pods delivering 11.5 petaFLOPS. Unlike the integer-only first generation, the second generation could also calculate in floating point, using a new bfloat16 format invented by Google Brain. That shift made the chip useful for both training and running models, and Google offered it on Google Compute Engine for TensorFlow applications.

    The third-generation TPU was announced on the 8th of May 2018. Google said each processor was twice as powerful as the second generation and would ship in pods holding four times as many chips. That combination produced an eightfold jump in performance per pod, with as many as 1,024 chips in a single pod.

    A pod that once topped out at 1,024 chips was, within a few years, dwarfed by designs Google had not yet built.

  • On the 18th of May 2021, Google's CEO Sundar Pichai used his keynote at the Google I/O virtual conference to introduce TPU v4. Pichai said v4 improved performance by more than twice that of TPU v3. He put it this way: "A single v4 pod contains 4,096 v4 chips, and each pod has 10x the interconnect bandwidth per chip at scale, compared to any other networking technology." An April 2023 paper from Google claimed TPU v4 ran 5-87% faster than a Nvidia A100 on machine learning benchmarks. A separate inference-focused version called v4i skipped the liquid cooling the training chip required.

    In 2021, Google revealed it was designing the physical layout of TPU v5 with help from a novel application of deep reinforcement learning. Google claimed v5 was nearly twice as fast as v4, leading some observers to speculate it could match or beat an Nvidia H100. A lighter, cost-efficient version called v5e followed, similar in spirit to v4i. In December 2023, Google announced TPU v5p, which it claimed was competitive with the Nvidia H100.

    In May 2024, at the Google I/O conference, Google announced a chip called Trillium. It reached preview availability in October 2024. Google claimed a 4.7 times performance increase over TPU v5e, achieved through larger matrix multiplication units and a higher clock speed. High bandwidth memory capacity and bandwidth both doubled, and a single pod could now hold up to 256 Trillium units.

    In April 2025, at the Google Cloud Next conference, Google unveiled TPU v7, named Ironwood. Ironwood comes in two configurations, one built from a 256-chip cluster and another from a 9,216-chip cluster. Its peak computational performance reaches 4,614 teraflops per second.

    On the 22nd of April 2026, Google announced two specialized eighth-generation chips, the TPU 8t and the TPU 8i. It was the first time Google had split its TPU architecture into separate training and inference-optimized designs. Both chips run on Google's custom Arm-based Axion CPUs and use a fourth generation of liquid cooling. The TPU 8t, built for large-scale pre-training and embedding-heavy workloads, delivers 12.6 FP4 petaflops of peak performance. It carries 216 gigabytes of HBM3e memory with 6,528 gigabytes per second of bandwidth. Its Virgo Network fabric scales up to 9,600 chips in a single superpod, reaching 121 FP4 exaflops of compute. The TPU 8i, aimed at high-speed serving, AI agents, and long-context reasoning, delivers 10.1 FP4 petaflops of peak performance. It packs 288 gigabytes of HBM3e memory at 8,601 gigabytes per second, plus 384 megabytes of on-chip SRAM, three times the previous generation's amount. Its Boardfly network topology and Collectives Acceleration Engine cut synchronization latency by five times.

    None of that hardware, running at petaflop scale inside a superpod, resembles the far smaller chip Google had already built for phones years earlier.

  • In July 2018, Google announced the Edge TPU, a purpose-built chip for running machine learning models on edge devices rather than inside a data center. It is far smaller and consumes far less power than the Cloud TPUs running in Google's data centers. In January 2019, Google opened the Edge TPU to developers under a new product line called Coral. The chip can perform 4 trillion operations per second while using just 2 watts of power.

    Coral's lineup includes a single-board computer, a system on module, a USB accessory, a mini PCI-e card, and an M.2 card. The single-board computer and system on module both run Mendel Linux OS, a derivative of Debian. The USB, PCI-e, and M.2 versions act as add-ons for existing computers, supporting Debian-based Linux on x86-64 and ARM64 hosts, including the Raspberry Pi. Its runtime is built on TensorFlow Lite, and the chip only accelerates forward-pass operations. That makes it best suited for inference, though lightweight transfer learning is still possible. The Edge TPU also only supports 8-bit math, so a network needs quantization-aware training, or, since late 2019, post-training quantization, to run on it.

    On the 12th of November 2019, Asus announced two single-board computers built around the Edge TPU, the Tinker Edge T and the Tinker Edge R. Both boards, aimed at IoT and edge AI uses, officially support Android and Debian operating systems. Asus also demonstrated a mini PC called the Asus PN60T that likewise featured the Edge TPU. On the 2nd of January 2020, Google announced the Coral Accelerator Module and the Coral Dev Board Mini ahead of that year's CES. The Accelerator Module is a multi-chip design with PCIe and USB interfaces, while the Dev Board Mini pairs it with a MediaTek 8167s chip.

    On the 15th of October 2019, Google announced the Pixel 4 smartphone, which contained an Edge TPU called the Pixel Neural Core. Google described the chip as "customized to meet the requirements of key camera features in Pixel 4." It relied on a neural network search that traded away some accuracy in exchange for lower latency and reduced power use.

    Google carried that approach further with Google Tensor, a custom system-on-chip that folded an Edge TPU into the Pixel 6 line when it released in 2021. The Google Tensor SoC demonstrated "extremely large performance advantages over the competition" in machine-learning-focused benchmarks. Instantaneous power draw ran relatively high, but shorter bursts at peak performance meant the chip used less energy overall.

    None of these smaller chips faced the courtroom battle that was already underway over how the larger ones handled numbers.

  • In 2019, Singular Computing filed suit against Google, alleging that Google's TPU chips infringed on its patents. The company was founded in 2009 by Joseph Bates, a visiting professor at MIT. By 2020, Google had persuaded the court to narrow the case to just two claims: claim 53, filed in 2012, and claim 7, filed in 2013. Both claims covered a dynamic range running from 10 to the negative sixth up to 10 to the sixth for floating point numbers. The standard float16 format cannot reach that range without resorting to subnormal numbers, because it allots only five bits to the exponent.

    In a 2023 court filing, Singular Computing pointed directly at Google's use of bfloat16, arguing it exceeds the dynamic range of float16. Singular Computing argued that non-standard floating point formats were non-obvious back in 2009. Google countered that a configurable-exponent format called VFLOAT already existed as prior art in 2002.

    By January 2024, additional lawsuits from Singular Computing had raised the number of patents being litigated to eight. Toward the end of that month's trial, Google agreed to a settlement with Singular Computing, though the terms were never disclosed. The settlement resolved a dispute that had grown to cover eight separate patents, but questions about how much numerical precision a machine learning chip truly needs did not end with it.

Common questions

What is a Tensor Processing Unit used for?

A Tensor Processing Unit is an application-specific integrated circuit that Google built for neural network machine learning, supporting frameworks including TensorFlow, Jax, and PyTorch. Google has used TPUs for tasks such as AlphaGo, AlphaZero, Street View text recognition, Google Photos, and the RankBrain search system.

Who invented the Tensor Processing Unit at Google?

Norman P. Jouppi served as tech lead and principal architect, leading the TPU's design, verification, and deployment to production in 15 months. Amir Salek was recruited in 2013 to build Google's custom silicon development capability, and Jonathan Ross was one of the original TPU engineers before founding the chip company Groq.

How much faster is a Tensor Processing Unit than a CPU or GPU?

Google's 2017 paper reported the first TPU reached 15-30 times the performance of contemporary CPUs and GPUs, and 30-80 times their performance per watt. An April 2023 paper claimed TPU v4 ran 5-87% faster than a Nvidia A100 on machine learning benchmarks.

What is the difference between a Tensor Processing Unit and a GPU?

A Tensor Processing Unit is designed for high-volume, low-precision computation, sometimes as little as 8-bit, without the rasterization or texture-mapping hardware graphics processors carry. TPUs work best for convolutional neural networks, while GPUs suit some fully connected neural networks and CPUs suit recurrent neural networks.

What is the Edge TPU used for?

The Edge TPU is Google's smaller, lower-power chip for running machine learning models outside the data center, capable of 4 trillion operations per second on 2 watts. Google made it available to developers in January 2019 under the Coral brand, and it also underlies the Pixel Neural Core in the Pixel 4 and the Google Tensor chip in the Pixel 6.

What lawsuit did Google face over the Tensor Processing Unit?

In 2019, Singular Computing, founded in 2009 by Joseph Bates, sued Google alleging its TPU chips infringed on patents covering a wide dynamic range for floating point numbers. Google agreed to a settlement with undisclosed terms in January 2024, after the number of patents in dispute had grown to eight.

All sources

78 references cited across the entry

  1. 1In-Datacenter Performance Analysis of a Tensor Processing UnitNorman Jouppi et al. — Association for Computing Machinery — 2017
  2. 6Benchmarking TPU, GPU, and CPU Platforms for Deep LearningYu Emma Wang et al. — 2019-07-01
  3. 9Amir Salek – LeadershipCerberus Capital Management
  4. 10Jonathan Ross' PostJonathan Ross — LinkedIn
  5. 12In-Datacenter Performance Analysis of a Tensor Processing UnitNorman P. Jouppi — 2017-04-15
  6. 17DeepMind's AlphaZero crushes chessColin McGourty — 6 December 2017
  7. 25Google offers its TPUs to AI cloud providers - reportGeorgia Butler Have your say — 2025-09-08
  8. 27Ten lessons from three generations that shaped Google's TPUv4iJouppi, Norman P. et al. — June 14, 2021
  9. 29NewsCase Study on the Google TPU and GDDR5 from Hot Chips 29Patrick Kennedy — Serve The Home — 22 August 2017
  10. 31TPU v4: An Optically Reconfigurable Supercomputer for Machine Learning with Hardware Support for EmbeddingsNorman P. Jouppi et al. — 2023
  11. 39In-Datacenter Performance Analysis of a Tensor Processing Unit™Norman P. Jouppi et al. — June 26, 2017
  12. 40NewsGoogle brings 45 teraflops tensor flow processors to its compute cloudPeter Bright — Ars Technica — 17 May 2017
  13. 41NewsGoogle Cloud TPU Details RevealedPatrick Kennedy — Serve The Home — 17 May 2017
  14. 42NewsGoogle I/O Opening Keynote Live-BlogAndre Frumusanu — 8 May 2018
  15. 43NewsGoogle Offers Glimpse of Third-Generation TPU ProcessorMichael Feldman — Top 500 — 11 May 2018
  16. 44NewsTearing Apart Google's TPU 3.0 AI CoprocessorPaul Teich — The Next Platform — 10 May 2018
  17. 46TPU v4: An Optically Reconfigurable Supercomputer for Machine Learning with Hardware Support for EmbeddingsNorman Jouppi — 2023-04-20
  18. 49JournalA graph placement methodology for fast chip designAzalia Mirhoseini et al. — 2021-06-01
  19. 73The surprising usefulness of sloppy arithmeticLarry Hardesty — MIT — 2011-01-03
  20. 78Google Settles AI-Chip Suit That Had Sought Over $5 BillionLaurel Brubaker Calkins — Bloomberg Law — January 24, 2024
  21. 79Google settles AI-related chip patent lawsuit that sought $1.67 blnBlake Brittain et al. — Reuters — January 24, 2024