AMD Announces World’s Fastest HPC Accelerator for Scientific Research¹
ꟷ AMD Instinct™ MI100 accelerators revolutionize high-Производительность computing (HPC) and AI with industry-leading compute Производительность ꟷ
ꟷ First GPU accelerator with new AMD CDNA Архитектура engineered for the exascale era ꟷ
SANTA CLARA, Calif., Nov. 16, 2020 (GLOBE NEWSWIRE) — AMD (NASDAQ: AMD) today announced the new AMD Instinct™ MI100 accelerator – the world’s fastest HPC GPU and the first x86 server GPU to surpass the 10 teraflops (FP64) Производительность barrier.1 Supported by new accelerated compute platforms from Dell, Gigabyte, HPE, and Supermicro, the MI100, combined with AMD EPYCTM CPUs and the ROCm™ 4.0 open software platform, is designed to propel new discoveries ahead of the exascale era.
Built on the new AMD CDNA Архитектура, the AMD Instinct MI100 GPU enables a new class of accelerated systems for HPC and AI when paired with 2nd Gen AMD EPYC processors. The MI100 offers up to 11.5 TFLOPS of peak FP64 Производительность for HPC and up to 46.1 TFLOPS Макс. производительность Matrix Производительность for AI and machine learning workloads2. With new AMD Matrix Core technology, the MI100 also delivers a nearly 7x boost in FP16 theoretical peak floating point Производительность for AI training workloads compared to AMD’s prior generation accelerators.3
“Today AMD takes a major step forward in the journey toward exascale computing as we unveil the AMD Instinct MI100 – the world’s fastest HPC GPU,” said Brad McCredie, corporate vice president, Data Center GPU and Accelerated Processing, AMD. “Squarely targeted toward the workloads that matter in scientific computing, our latest accelerator, when combined with the AMD ROCm open software platform, is designed to provide scientists and researchers a superior foundation for their work in HPC.”
Open Software Platform for the Exascale Era
The AMD ROCm developer software provides the foundation for exascale computing. As an open source toolset consisting of compilers, programming APIs and libraries, ROCm is used by exascale software developers to create high Производительность applications. ROCm 4.0 has been optimized to deliver Производительность at scale for MI100-based systems. ROCm 4.0 has upgraded the compiler to be open source and unified to support both OpenMP® 5.0 and HIP. PyTorch and Tensorflow frameworks, which have been optimized with ROCm 4.0, can now achieve higher Производительность with MI1007,8. ROCm 4.0 is the latest offering for HPC, ML and AI application developers which allows them to create Производительность portable software.
“We’ve received early access to the MI100 accelerator, and the preliminary results are very encouraging. We’ve typically seen significant Производительность boosts, up to 2-3x compared to other GPUs,” said Bronson Messer, director of science, Oak Ridge Leadership Computing Facility. “What’s also important to recognize is the impact software has on Производительность. The fact that the ROCm open software platform and HIP developer tool are open source and work on a variety of platforms, it is something that we have been absolutely almost obsessed with since we fielded the very first hybrid CPU/GPU system.”
Key capabilities and features of the AMD Instinct MI100 accelerator include:
- All-New AMD CDNA Архитектура- Engineered to Питание AMD GPUs for the exascale era and at the heart of the MI100 accelerator, the AMD CDNA Архитектура offers exceptional Производительность and Питание efficiency
- Leading FP64 and FP32 Производительность for HPC Workloads – Delivers industry leading 11.5 TFLOPS peak FP64 Производительность and 23.1 TFLOPS Макс. производительность Производительность, enabling scientists and researchers across the globe to accelerate discoveries in industries including life sciences, energy, finance, academics, government, defense and more.1
- All-New Matrix Core Technology for HPC and AI – Supercharged Производительность for a full range of single and mixed precision matrix operations, such as FP32, FP16, bFloat16, Int8 and Int4, engineered to boost the convergence of HPC and AI.
- 2nd Gen AMD Infinity Fabric™ Technology – Instinct MI100 provides ~2x the peer-to-peer (P2P) peak I/O bandwidth over PCIe® 4.0 with up to 340 GB/s of aggregate bandwidth per card with three AMD Infinity Fabric™ Links.4 In a server, MI100 GPUs can be configured with up to two fully-connected quad GPU hives, each providing up to 552 GB/s of P2P I/O bandwidth for fast data sharing.4
- Ultra-Fast HBM2 Память– Features 32GB High-bandwidth HBM2 Память at a clock rate of 1.2 GHz and delivers an ultra-high 1.23 TB/s of Пропускная способность to support large data sets and help eliminate bottlenecks in moving data in and out of Память.5
- Support for Industry’s Latest PCIe® Gen 4.0 – Designed with the latest PCIe Gen 4.0 technology support providing up to 64GB/s peak theoretical transport data bandwidth from CPU to GPU.6
Available Server Solutions
The AMD Instinct MI100 accelerators are expected by end of the year in systems from major OEM and ODM partners in the enterprise markets, including:
Dell
“Dell EMC ПитаниеEdge servers will support the new AMD Instinct MI100, which will enable faster insights from data. This would help our customers achieve more robust and efficient HPC and AI results rapidly,” said Ravi Pendekanti, senior vice president, ПитаниеEdge Servers, Dell Technologies. “AMD has been a valued partner in our support for advancing innovation in the data center. The high-Производительность capabilities of AMD Instinct accelerators are a natural fit for our ПитаниеEdge server AI & HPC portfolio.”
Gigabyte
“We’re pleased to again work with AMD as a strategic partner offering customers server hardware for high Производительность computing,” said Alan Chen, assistant vice president in NCBU, GIGABYTE. “AMD Instinct MI100 accelerators represent the next level of high-Производительность computing in the data center, bringing greater connectivity and data bandwidth for energy research, molecular dynamics, and deep learning training. As a new accelerator in the GIGABYTE portfolio, our customers can look to benefit from improved Производительность across a range of scientific and industrial HPC workloads.”
Hewlett Packard Enterprise (HPE)
“Customers use HPE Apollo systems for purpose-built capabilities and Производительность to tackle a range of complex, data-intensive workloads across high-Производительность computing (HPC), deep learning and analytics,” said Bill Mannel, vice president and general manager, HPC at HPE. “With the introduction of the new HPE Apollo 6500 Gen10 Plus system, we are further advancing our portfolio to improve workload Производительность by supporting the new AMD Instinct MI100 accelerator, which enables greater connectivity and data processing, alongside the 2nd Gen AMD EPYC™ processor. We look forward to continuing our collaboration with AMD to expand our offerings with its latest CPUs and accelerators.”
Supermicro
“We’re excited that AMD is making a big impact in high-Производительность computing with AMD Instinct MI100 GPU accelerators,” said Vik Malyala, senior vice president, field application engineering and business development, Supermicro. “With the combination of the compute Питание gained with the new CDNA Архитектура, along with the high Память and GPU peer-to-peer bandwidth the MI100 brings, our customers will get access to great solutions that will meet their accelerated compute requirements and critical enterprise workloads. The AMD Instinct MI100 will be a great addition for our multi-GPU servers and our extensive portfolio of high-Производительность systems and server building block solutions.”
MI100 Specifications
Compute Units |
Stream Processors |
FP64 TFLOPS (Peak) |
FP32 TFLOPS (Peak) |
FP32 Matrix TFLOPS (Peak) |
FP16/FP16 Matrix TFLOPS (Peak) |
INT4 | INT8 TOPS (Peak) |
bFloat16 TFLOPS (Peak) |
HBM2 ECC Память |
Память Bandwidth |
120 | 7680 | Up to 11.5 |
Up to 23.1 | Up to 46.1 |
Up to 184.6 |
Up to 184.6 |
Up to 92.3 TFLOPS |
32GB | Up to 1.23 TB/s |