Overview of AI Optimizer for IBM Z and IBM LinuxONE
IBM not only modernizes the mainframe systems with a full range of cutting-edge AI capabilities, but also offers a complete stack of AI-enabled solutions to help you leverage those capabilities. AI Optimizer for Z and LinuxONE is one of the solutions that empower you to optimize your mainframe AI workloads.
What is AI Optimizer for Z and LinuxONE?
AI Optimizer for Z and LinuxONE, previously known as IBM AI Optimizer for Z, is an enterprise-grade appliance that serves as a central AI inference gateway for large language model (LLM) workloads on IBM Z and LinuxONE systems. Built for the IBM Secure Service Container (SSC) platform, the appliance provides a unified control plane for deploying LLMs, routing inference requests, and executing routed requests across locally deployed and remotely hosted LLMs.
The AI Optimizer for Z and LinuxONE appliance is designed to fully leverage IBM Spyre Accelerator, an advanced hardware acceleration technology purpose‑built for AI inference workloads. Enabled by the accelerator, the appliance intelligently optimizes inference request routing across your mission‑critical AI workloads while preserving data integrity, compliance, and security already provided by the mainframe platforms.
As shown in the following diagram, each AI Optimizer for Z and LinuxONE instance is composed of integrated services that work together to provide a complete AI inference routing optimization and management solution.
At the core of the comprehensive solution are the multi-functional appliance manager, the intelligent inference router, the data-rich inference routing monitor, the intuitive web user interfaces, and the developer-friendly REST API.
-
Appliance manager. AI Optimizer for Z and LinuxONE Appliance Manager provide centralized tools for appliance administration, access control, system health monitoring, and diagnostic information retrieval.
-
Inference router. Powered by the Red Hat® AI Inference Server (RHAIIS), the AI Optimizer for Z and LinuxONE inference router acts as the central gateway for all incoming inference requests and dynamically routes them to different LLMs. RHAIIS provides the runtime engine and serves a range of LLMs that can be deployed locally in AI Optimizer for Z and LinuxONE. It also seamlessly integrates with Spyre accelerator cards, enabling the models to automatically take advantage of hardware acceleration and the inference router to optimize inference routing.
The inference router enables request routing by exposing OpenAI‑compatible API endpoints. You can interact with the router through those endpoints and send your inference queries to both locally deployed and remotely hosted LLMs. Using tag‑based routing, the router dynamically directs your queries to the most appropriate models based on the characteristics that you specify in the REST API call. The router also manages authentication, rate limiting, and request validation, ensuring that all inference operations are secure and well governed. By serving as a single point of entry for all inference routing operations, the router simplifies your interaction while providing sophisticated routing behind the scenes.
-
Inference routing monitor. Built on the open-source Prometheus and Grafana technologies, the AI Optimizer for Z and LinuxONE inference routing monitor gains deep visibility into your AI inference and provides comprehensive metrics of inference routing performance and system resource consumption. The Prometheus metrics collection service scrapes information, including inference latency, throughput, resource utilization, and error rates. Taking input from Prometheus, the Grafana dashboard visualizes the metrics in real time. The analytics from the metrics help you not only quickly identify potential bottlenecks and anomalies in LLM performance but also effectively guide your future planning for cost-effectiveness in infrastructure and resource allocation.
-
Web user interfaces. You can interact with the AI Optimizer for Z and LinuxONE appliance through the following user interfaces (UIs), each of which serves a different set of purposes:
-
The appliance manager UI is the main interface for you to interact with the Appliance Manager for configuring the appliance, managing users and permissions, and accessing logs and system dumps.
-
The inference router UI is the main interface for you to interact with the inference router for deploying LLMs, managing model deployments, routing inference requests, and executing inference operations.
-
The inference routing monitoring dashboard is a visually intuitive dashboard that collects, categorizes, and visualizes real-time data of LLM inference routing and system resource utilization.
-
-
REST API. The AI Optimizer for Z and LinuxONE REST API exposes endpoints that you can use to access the inference router and deployed LLMs, manage the remote model gateway, handle external server certificates, and generate API keys. By providing OpenAI-compatible REST API, AI Optimizer for Z and LinuxONE enables you to use existing code, tools, and libraries without modification to interact with the inference router on IBM Z and LinuxONE.
What are the key features of AI Optimizer for Z and LinuxONE?
Enabled by the core components of the solution, AI Optimizer for Z and LinuxONE delivers the following key features and capabilities:
- Optimized routing of LLM inference requests
-
AI Optimizer for Z and LinuxONE intelligently routes inference requests across multiple LLMs based on tags and other specified metadata. Tag-based routing represents an advanced capability that goes beyond simple load balancing. During deployment, you can tag the models based on use cases, application domains, and other characteristics, including latency, data sensitivity, computational intensity, GPU availability, and cost efficiency. Then, you specify one or more of these tags in your API inference calls to route your requests. This tag-based routing ensures that your requests are quickly and efficiently directed to the best suitable models, which not only reduces response time but also optimizes resource utilization across your AI infrastructure.
- Remote model gateway Tech Preview
-
AI Optimizer for Z and LinuxONE includes a model gateway for LLMs that are hosted remotely. AI Optimizer for Z and LinuxONE enables the registration, tagging, and grouping of remote models for unified management and inference routing. Wherever supported, it includes key performance metrics of external models in the inference monitoring dashboard, providing a holistic view of your entire AI landscape.
Tech Preview: The integration of remote models feature is currently available for technology preview only. - Integration with IBM Spyre Accelerator and native support of IBM Granite™ 3.3-8b-instruct
-
AI Optimizer for Z and LinuxONE natively supports the deployment of Granite 3.3-8b-instruct, a foundation LLM that is specifically optimized for IBM Spyre accelerators. The model provides state-of-the-art natural language understanding and generation capabilities while maintaining the performance and efficiency characteristics required for enterprise deployment. This tight integration with the Spyre accelerator and the native deployment of IBM Granite family of LLMs ensure optimal performance and resource utilization of the inference router.
- Real-time monitoring of inference request routing
- Powered by open-source technologies Prometheus and Grafana, the AI Optimizer for Z and LinuxONE inference routing monitor gains deep observability and visibility into your AI inference workloads in real time through the following capabilities:
-
Metrics collection: The Prometheus-based metrics collection service captures detailed performance data from every specified component of the inference router, including inference latency, throughput, error rates, and resource utilization.
-
Visualization: Grafana dashboard provides real-time visualization of the metrics from the metrics collection service.
-
Alerting: The observability stack includes alerting capabilities, notifying you when metrics exceed defined thresholds or when anomalous behavior is detected. This proactive monitoring helps prevent issues before they impact users.
-
Performance analysis: You can use the observability tools to analyze LLM performance, identify bottlenecks, and make data-driven decisions about model optimization and deployment strategies.
-
- Role-based access control (RBAC)
-
AI Optimizer for Z and LinuxONE requires authorized access to both the web UIs and the REST API. User authorization is centrally managed through the Appliance Manager. As an administrator, you can create users and assign them roles. The roles determine what tasks the users can do and what system resources they can access.
How does AI Optimizer for Z and LinuxONE work?
The operational flow of the AI Optimizer for Z and LinuxONE inference router follows a well-defined process that ensures secure, efficient, and intelligent handling of AI inference requests. The workflow is simple and straightforward. You deploy LLMs in the inference router UI and then you send inference queries by using the REST API. Your inference requests are dynamically routed to the most appropriate models.
The inference router has sophisticated model lifecycle management capabilities that allow you to deploy, tag, and select a set of Spyre accelerator cards for an LLM. If needed, you can update the deployment configuration at any time. Once deployed, the LLM usage can be scaled up or down based on demand because the inference router can automatically distribute inference workload across multiple instances of the same model to ensure consistent performance under varying conditions.
You can send inference requests to the deployed models by using the AI Optimizer for Z and LinuxONE REST API. Upon successful authentication, the inference router examines the request to determine the appropriate routing target. In addition to other information, such as model type, the router uses the specified tags to intelligently select the most appropriate target from available models. After the most optimal routing is determined, the router forwards the request to RHAIIS. For LLMs that are deployed on the Spyre accelerator cards, RHAIIS automatically leverages the hardware acceleration capabilities while processing the requests.
Throughout this entire inference routing and execution process, the inference routing monitor collects key metrics and logs that you can use for inference monitoring, troubleshooting, and further optimization.