Google Launches New Tensor Ironwood Processor at Google Cloud Next 25

Google officially announced its presence Tensor Processing Unit (TPU) Ironwood, seventh generation artificial intelligence accelerator Special for large-scale inference, in the event Google Cloud Next which will take place on 25 April 2025.
Amin Vahdat, Vice President and General Manager of ML, Systems, and AI Cloud on Google, calls Ironwood The fastest TPU ever made by Google, at the same time become The first TPU design developed exclusively for the inference process Artificial intelligence model.
“This TPU is designed to process the thought of an inferential AI model on a large scale,” Amin explained in an official post on Google Blog, Wednesday, April 9, 2025.
Enter the AI Inference Era
Ironwood is called the main driver “Inference Era” by Google—a new stage in the development of AI where the system not only collects data, but also Actively capture, process, and present real-time information-based solutions.
“This is not just data collection, but the delivery of thoughts from large-scale AI models such as LLM, MOE, and other complex inference tasks,” explained Amin.
Gahar specifications for massive scale
Ironwood TPU consists of 9,216 liquid cooling processor, connected via network system Inter-Chip Interconnect (ICI) low latency and high bandwidth, with power consumption up to 10 megawatts per pod.
Some significant improvements compared to the previous generation (TPU trilium) include:
💾 HBM memory: 192 GB per chip (up 6x fold)
⚡ Memory bandwidth: 7.2 tbs per chip (up 3.5x)
🔗 ICI Bandwidth: 1.2 Tbps bidirectional (up 1.5x)
This combination is designed to meet the needs of generative AI inference and heavy computing workload efficiently and energy-efficient.
faster than its predecessor
CEO of Google and Alphabet, Sundar Pichai, also confirms that Ironwood will be launched This year. He claims that Ironwood’s performance Up to 3,600 times higher compared to the previous version.
“This is the most powerful AI processor we have ever developed. Ironwood will be the backbone of the next generation of generative and inferential AI models processing,” said Sundar.
From TPU V1 to Ironwood
Ironwood continues the evolution of Google’s TPU from the first generation (TPU V1) to Trilium. When trilium is widely used for training, Ironwood is focused on Fast and efficient inference, making it ideal for products and services based AI Real-time Such as virtual assistants, large language models, and cloud processing.
We’ll wait how Ironwood will change the competitive map of AI accelerators in the market, especially against NVIDIA and AWS Trainium. Its high performance promises major transformations in the inference and cloud computing ecosystem.























