Overview of IBM watsonx.data Premium

IBM watsonx.data Premium is a hybrid, gen AI data lakehouse designed to power your AI and analytics solutions across complex and distributed data environments. With watsonx.data Premium, you can unlock insights from both structured and unstructured data from diverse sources. You can run various workloads and integrate seamlessly with the IBM AI products so that you can scale your AI and analytics initiatives.

The use cases for watsonx.data Premium span data management, analytics, and the AI lifecycle:

  • Data engineers can store, query, and analyze data.
  • Data scientists can extract insights from the data to inform business decisions.
  • Data stewards can ensure that all data governance and quality requirements are addressed.
  • AI app developers can build AI solutions with curated and enriched data.

The following illustration shows the IBM watsonx.data Premium components.

IBM watsonx.data Premium components

watsonx.data Premium

IBM watsonx.data Premium includes the following key capabilities to handle structured and unstructured data at scale:

  • Hybrid deployment options – Software, SaaS, and Customer VPC (Bring your own cloud (BYOC)), and Developer edition
  • AI-ready architecture – with native support for unstructured and structured data processing, open source vector database integration (Milvus), first-class integration with watsonx.ai, and automated data ingestion and curation capabilities.
  • Multi-engine support – Presto (Java, C++), Spark (Java, C++), Milvus, integration with Db2wh, and Netezza.
  • Data source connectivity – Multiple connectors support for both structured and unstructured data
  • Unified governance and metadata – Integration with watsonx.data intelligence and Apache Ranger, end-to-end access control, and access control lists for unstructured data.
  • Open and flexible – Support for open source formats (Iceberg), separation of compute, metadata, and storage, and co-existing open source and proprietary tools.
  • Optimized for performance and cost – Object storage across hybrid and multi-cloud environments, reduced data duplication and storage costs, and high-performance query engines for large-scale analytics.
  • Model Context Protocol (MCP) servers - Integrate with AI agents and assistants through MCP servers to add your RAG pipeline to watsonx Orchestrate or other AI agents or query structured data in the lakehouse with an AI agent on your local computer.

Learn more