Accelerate open source AI on IBM Z and LinuxONE with optimized performance and trusted support
AI Toolkit for IBM Z® and LinuxONE is a family of supported open source AI frameworks optimized for the Telum processor. Adopt AI with certified containers, integrated accelerators and expert support. These frameworks use on-chip AI acceleration in z16®, LinuxONE 4, z17® and LinuxONE 5.
Deploy open source AI with IBM Elite Support and IBM-vetted containers for compliance, security and nonwarranted software confidence.
IBM z17’s Telum II on-chip AI accelerator delivers inference performance comparable to a 13-core x86 server within the same system managing online transaction processing (OLTP) workloads.1
Deploy ML, DL and large language models (LLMs) with up to 3.5x faster inference for predictions.4 Seamlessly integrate with PyTorch, TensorFlow, Snap ML, Open Neural Network Exchange (ONNX) and more.
Seamlessly develop and deploy machine learning (ML) models with optimized TensorFlow and PyTorch frameworks tailored for IBM Z. Use integrated acceleration for improved neural network inference performance.
The AI Toolkit consists of IBM Elite Support (within IBM Selected Support) and IBM Secure Engineering. These tools vet and scan open source AI serving frameworks and IBM-certified containers for security vulnerabilities and validate compliance with industry regulations.
Use on-chip AI inferencing to analyze large volumes of unstructured data on IBM Z and LinuxONE. Deliver faster, more accurate predictions for chatbots, content classification and language understanding.
With up to 450 billion inferences per day and 99.9 percentile response under 1 ms, detect and act on fraudulent activity instantly by using composite AI models and Telum acceleration.5
Identify suspicious patterns in financial transactions by using Snap ML and Scikit-learn. With data compression, encryption and on-platform AI, improve AML response without sacrificing performance or security.
1 Using a single Integrated Accelerator for AI on an OLTP workload on IBM z17 matches the throughput of running inferencing on a compared remote x86 server with 13 cores.
DISCLAIMER: Performance results are based on IBM® internal tests running on IBM Systems Hardware of machine type 9175. The OLTP application and PostgreSQL was deployed on the IBM Systems Hardware. The Credit Card Fraud Detection (CCFD) ensemble AI setup consists of two models (LSTM, TabFormer). On IBM Systems Hardware, running the OLTP application with IBM Z Deep Learning Compiler (zDLC) compiled jar and IBM Z Accelerated for NVIDIA® Triton™ Inference Server locally and processing the AI inference operations on IFLs and the Integrated Accelerator for AI versus running the OLTP application locally and processing remote AI inference operations on a x86 server running NVIDIA Triton Inference Server with OpenVINO™ runtime backend on CPU (with AMX). Each scenario was driven from Apache JMeter™ 5.6.3 with 64 parallel users. IBM Systems Hardware configuration: 1 LPAR running Ubuntu 24.04 with 7 dedicated IFLs (SMT), 256 GB memory, and IBM FlashSystem® 9500 storage. The Network adapters were dedicated for NETH on Linux. x86 server configuration: 1 x86 server running Ubuntu 24.04 with 28 Emerald Rapids Intel® Xeon® Gold CPUs @ 2.20 GHz with hyper-threading turned on, 1 TB memory, local SSDs, UEFI with maximum performance profile enabled, CPU P-State Control and C-States disabled. Results may vary.
2 The IBM z17 Telum II processor supports INT8 quantization, designed to reduce inference latency when compared to the non-quantized models.
DISCLAIMER: INT8 quantization support in the IBM z17 Telum II processor reduces and stores the weights and activations from 32-bit floating point numbers to 8-bit integers. This reduction in precision allows for faster computations which can lead to lower inference times compared to the non-quantized models
3,5 With IBM z17, process up to 450 billion inference operations per day using multiple AI models for credit card fraud detection.
DISCLAIMER: Performance result is extrapolated from IBM® internal tests running on IBM Systems Hardware of machine type 9175. The benchmark was executed with 64 threads performing local inference operations using a synthetic credit card fraud detection (CCFD) model based on an LSTM and a TabFormer model. The benchmark exploited the Integrated Accelerator for AI using IBM Z Deep Learning Compiler (zDLC) and IBM Z Accelerated for PyTorch. The setup consists of 64 threads pinned in groups of 8 to each chip (1 for zDLC, 7 for PyTorch). The TabFormer (tabular transformer) model evaluated 0.035% of the inference requests. A batch size of 160 was used for the LSTM based model. IBM Systems Hardware configuration: 1 LPAR running Ubuntu 24.04 with 45 IFLs (SMT), 128 GB memory. Results may vary.
4 DISCLAIMER: Performance results based on IBM internal tests doing inferencing using a Random Forest model with Snap ML v1.12.0 backend which uses the Integrated Accelerator for AI on IBM Machine Type 3931 versus the NVIDIA Forest Inference Library backend on compared x86 server. The model was trained on the following public dataset and NVIDIA Triton™ was used on both platforms as model serving framework. The workload was driven via the http benchmarking tool Hey. IBM Machine Type 3931 configuration: Ubuntu 22.04 in an LPAR with 6 dedicated IFLs, 256 GB memory. x86 configuration: Ubuntu 22.04 on 6 Ice Lake Intel® Xeon® Gold CPU @ 2.80GHz with hyper-threading turned on, 1 TB memory.