Designed to Scale and Optimize GenAI Inferencing
IBM AI Optimizer for IBM Z and IBM LinuxONE is the unified AI inferencing stack for the IBM Spyre Accelerator - bringing together model on-boarding, routing, and monitoring into a single, integrated solution on IBM Z and IBM LinuxONE.
An integrated software appliance delivered as a single LPAR image - including the operating system, curated AI models, container runtime, observability tooling, and a management UI. This unified approach reduces operational complexity, accelerates deployment timelines, and ensures consistency across environments.
Gain full visibility into GenAI inferencing across IBM Z with enterprise‑grade observability. Built‑in Prometheus and Grafana dashboards provide deep insights into:
This transparency helps eliminate over‑provisioning, streamline capacity planning, and drive smarter infrastructure investment.
AI Optimizer registers models running on Spyre for optimization. Users can configure their own routing strategies or rely on the built‑in intelligent router, which considers performance, availability, and usage patterns. Semantic tagging allows grouping of models for use‑case‑aligned routing thus providing more flexibility on inferencing requests.
Models deployed outside IBM Z or LinuxONE can be registered, tagged, grouped, and monitored along with on‑platform models. This provides a unified operational view of GenAI inferencing across hybrid environments, ensuring consistency in governance and performance tracking.
AI Optimizer for Z automates installation and configuration of key IBM Z Gen AI components and products, such as IBM watsonx Assistant for Z, ensuring fast and reliable setup. It validates infrastructure and provides a health dashboard for easy monitoring. This reduces complexity and accelerates time to production.
Simplify and transform how your users interact and manage mainframe with AI.
Accelerate AI innovation at scale with IBM infrastructure.