AI Toolkit for IBM Z and LinuxONE

Accelerate open source AI on IBM Z and LinuxONE with optimized performance and trusted support

Illustration of the AI toolkit for IBM Z and IBM LinuxONE workflow

Deploy AI with speed and confidence

AI Toolkit for IBM Z® and LinuxONE is a family of supported open source AI frameworks optimized for the Telum processor. Adopt AI with certified containers, integrated accelerators and expert support. These frameworks use on-chip AI acceleration in z16®LinuxONE 4z17® and LinuxONE 5.

Confident AI deployment at scale

Deploy open source AI with IBM Elite Support and IBM-vetted containers for compliance, security and nonwarranted software confidence.

Accelerated real-time AI

IBM z17’s Telum II on-chip AI accelerator delivers inference performance comparable to a 13-core x86 server within the same system managing online transaction processing (OLTP) workloads.1

Inferencing at scale

IBM z17 and LinuxONE 5 enable INT8-optimized AI2, powering multiple models predictive scoring, while delivering up to 450 billion daily inferences with less than 1 ms response time. These tools manage such results because they use a deep learning model for credit card fraud detection.3

Support for multiple AI models

Deploy ML, DL and large language models (LLMs) with up to 3.5x faster inference for predictions.4 Seamlessly integrate with PyTorch, TensorFlow, Snap ML, Open Neural Network Exchange (ONNX) and more.

Features

Seamlessly develop and deploy machine learning (ML) models with optimized TensorFlow and PyTorch frameworks tailored for IBM Z. Use integrated acceleration for improved neural network inference performance.

A person on a laptop showing interaction with AI
PyTorch compatible

Accelerate seamless integration of PyTorch with IBM Z Accelerated for PyTorch to develop and deploy ML models on neural networks.

Explore PyTorch inference
IT engineer interacting with server module in sever room
TensorFlow compatible

Accelerate seamless integration of TensorFlow with IBM Z Accelerated for TensorFlow to develop and deploy ML models on neural networks.

Explore TensorFlow inference
A person using a laptop in a server room
ML models with TensorFlow Serving

Harness the benefits of TensorFlow Serving, a flexible and high-performance service system, with IBM Z Accelerated for TensorFlow Serving to help the deployment of ML models in production.

Explore TensorFlow Serving
Two people looking at an interactive screen with ''AI" logo on it
NVIDIA Triton Inference Server

Optimized for IBM Telum processors and Linux on Z, IBM Z Accelerated for NVIDIA Triton Inference Server enables high-performance AI inference. The tool offers support for dynamic batching, multiple frameworks and custom backends across CPUs and GPUs.

Discover Triton Inference Server
A person looking at desktop screen with laptop next to it
Run Snap ML

Use IBM Z Accelerated for Snap ML to build and deploy ML models with Snap ML, an IBM nonwarranted program that optimizes the training and scoring of popular ML models.

Explore IBM Snap machine learning
A person looking at three desktop screens.
Compile ML ONNX models with IBM zDLC

Use Telum and Telum II on-chip accelerated inference capabilities with ONNX models that use the IBM Z® Deep Learning Compiler (IBM zDLC) on IBM z/OS®, zCX and LinuxONE. IBM zDLC, an AI model compiler, provides capabilities such as auto-quantization for ML models with reduced latency and reduced energy consumption.

Explore IBM Deep Learning Compiler Using the IBM zDLC Container Images
A person on a laptop showing interaction with AI
PyTorch compatible

Accelerate seamless integration of PyTorch with IBM Z Accelerated for PyTorch to develop and deploy ML models on neural networks.

Explore PyTorch inference
IT engineer interacting with server module in sever room
TensorFlow compatible

Accelerate seamless integration of TensorFlow with IBM Z Accelerated for TensorFlow to develop and deploy ML models on neural networks.

Explore TensorFlow inference
A person using a laptop in a server room
ML models with TensorFlow Serving

Harness the benefits of TensorFlow Serving, a flexible and high-performance service system, with IBM Z Accelerated for TensorFlow Serving to help the deployment of ML models in production.

Explore TensorFlow Serving
Two people looking at an interactive screen with ''AI" logo on it
NVIDIA Triton Inference Server

Optimized for IBM Telum processors and Linux on Z, IBM Z Accelerated for NVIDIA Triton Inference Server enables high-performance AI inference. The tool offers support for dynamic batching, multiple frameworks and custom backends across CPUs and GPUs.

Discover Triton Inference Server
A person looking at desktop screen with laptop next to it
Run Snap ML

Use IBM Z Accelerated for Snap ML to build and deploy ML models with Snap ML, an IBM nonwarranted program that optimizes the training and scoring of popular ML models.

Explore IBM Snap machine learning
A person looking at three desktop screens.
Compile ML ONNX models with IBM zDLC

Use Telum and Telum II on-chip accelerated inference capabilities with ONNX models that use the IBM Z® Deep Learning Compiler (IBM zDLC) on IBM z/OS®, zCX and LinuxONE. IBM zDLC, an AI model compiler, provides capabilities such as auto-quantization for ML models with reduced latency and reduced energy consumption.

Explore IBM Deep Learning Compiler Using the IBM zDLC Container Images

Secure, compliant containers by IBM

Containers found in the AI Toolkit for IBM Z and LinuxONE

The AI Toolkit consists of IBM Elite Support (within IBM Selected Support) and IBM Secure Engineering. These tools vet and scan open source AI serving frameworks and IBM-certified containers for security vulnerabilities and validate compliance with industry regulations.

Access through IBM Container Registry
Use cases
A person holding a tech chip
Real-time natural language processing

Use on-chip AI inferencing to analyze large volumes of unstructured data on IBM Z and LinuxONE. Deliver faster, more accurate predictions for chatbots, content classification and language understanding.

A person holding a credit card
Credit card fraud detection in milliseconds

With up to 450 billion inferences per day and 99.9 percentile response under 1 ms, detect and act on fraudulent activity instantly by using composite AI models and Telum acceleration.5

A person tapping a credit card
Anti-money laundering at scale

Identify suspicious patterns in financial transactions by using Snap ML and Scikit-learn. With data compression, encryption and on-platform AI, improve AML response without sacrificing performance or security.

Take the next step

Discover how AI Toolkit for IBM Z and LinuxONE accelerate open source AI with optimized performance and trusted support.

  1. Access through IBM Container Registry
Footnotes

Using a single Integrated Accelerator for AI on an OLTP workload on IBM z17 matches the throughput of running inferencing on a compared remote x86 server with 13 cores.

DISCLAIMER: Performance results are based on IBM® internal tests running on IBM Systems Hardware of machine type 9175. The OLTP application and PostgreSQL was deployed on the IBM Systems Hardware. The Credit Card Fraud Detection (CCFD) ensemble AI setup consists of two models (LSTM, TabFormer). On IBM Systems Hardware, running the OLTP application with IBM Z Deep Learning Compiler (zDLC) compiled jar and IBM Z Accelerated for NVIDIA® Triton™ Inference Server locally and processing the AI inference operations on IFLs and the Integrated Accelerator for AI versus running the OLTP application locally and processing remote AI inference operations on a x86 server running NVIDIA Triton Inference Server with OpenVINO™ runtime backend on CPU (with AMX). Each scenario was driven from Apache JMeter™ 5.6.3 with 64 parallel users. IBM Systems Hardware configuration: 1 LPAR running Ubuntu 24.04 with 7 dedicated IFLs (SMT), 256 GB memory, and IBM FlashSystem® 9500 storage. The Network adapters were dedicated for NETH on Linux. x86 server configuration: 1 x86 server running Ubuntu 24.04 with 28 Emerald Rapids Intel® Xeon® Gold CPUs @ 2.20 GHz with hyper-threading turned on, 1 TB memory, local SSDs, UEFI with maximum performance profile enabled, CPU P-State Control and C-States disabled. Results may vary.

The IBM z17 Telum II processor supports INT8 quantization, designed to reduce inference latency when compared to the non-quantized models.

DISCLAIMER: INT8 quantization support in the IBM z17 Telum II processor reduces and stores the weights and activations from 32-bit floating point numbers to 8-bit integers. This reduction in precision allows for faster computations which can lead to lower inference times compared to the non-quantized models

3,5 With IBM z17, process up to 450 billion inference operations per day using multiple AI models for credit card fraud detection.

DISCLAIMER: Performance result is extrapolated from IBM® internal tests running on IBM Systems Hardware of machine type 9175. The benchmark was executed with 64 threads performing local inference operations using a synthetic credit card fraud detection (CCFD) model based on an LSTM and a TabFormer model. The benchmark exploited the Integrated Accelerator for AI using IBM Z Deep Learning Compiler (zDLC) and IBM Z Accelerated for PyTorch. The setup consists of 64 threads pinned in groups of 8 to each chip (1 for zDLC, 7 for PyTorch). The TabFormer (tabular transformer) model evaluated 0.035% of the inference requests. A batch size of 160 was used for the LSTM based model. IBM Systems Hardware configuration: 1 LPAR running Ubuntu 24.04 with 45 IFLs (SMT), 128 GB memory. Results may vary.

4 DISCLAIMER: Performance results based on IBM internal tests doing inferencing using a Random Forest model with Snap ML v1.12.0 backend which uses the Integrated Accelerator for AI on IBM Machine Type 3931 versus the NVIDIA Forest Inference Library backend on compared x86 server. The model was trained on the following public dataset and NVIDIA Triton™ was used on both platforms as model serving framework. The workload was driven via the http benchmarking tool Hey. IBM Machine Type 3931 configuration: Ubuntu 22.04 in an LPAR with 6 dedicated IFLs, 256 GB memory. x86 configuration: Ubuntu 22.04 on 6 Ice Lake Intel® Xeon® Gold CPU @ 2.80GHz with hyper-threading turned on, 1 TB memory.