Overview of IBM watsonx.data
IBM watsonx.data is a hybrid, gen AI data lakehouse designed to power your AI and analytics solutions across complex and distributed data environments. With watsonx.data, you can unlock insights from both structured and unstructured data from diverse sources. You can run various workloads and integrate seamlessly with the IBM AI products so that you can scale your AI and analytics initiatives.
The use cases for watsonx.data span data management, analytics, and the AI lifecycle:
- Data engineers can store, query, and analyze data.
- Data scientists can extract insights from the data to inform business decisions.
- Data stewards can ensure that all data governance and quality requirements are addressed.
- AI app developers can build AI solutions with curated and enriched data.
The following illustration shows the IBM watsonx.data as a Service components.

watsonx.data
IBM watsonx.data includes the following key capabilities to handle structured and unstructured data at scale:
- Hybrid deployment options – Software, SaaS, and Customer VPC (Bring your own cloud (BYOC)), and Developer edition
- AI-ready architecture – with native support for unstructured and structured data processing, open source vector database integration (Milvus), first-class integration with watsonx.ai, and automated data ingestion and curation capabilities.
- Multi-engine support – Presto (Java, C++), Spark (Java, C++), Milvus, integration with Db2wh, and Netezza.
- Data source connectivity – Multiple connectors support for both structured and unstructured data
- Unified governance and metadata – Integration with watsonx.data intelligence and Apache Ranger, end-to-end access control, and access control lists for unstructured data.
- Open and flexible – Support for open source formats (Iceberg), separation of compute, metadata, and storage, and co-existing open source and proprietary tools.
- Optimized for performance and cost – Object storage across hybrid and multi-cloud environments, reduced data duplication and storage costs, and high-performance query engines for large-scale analytics.
- Model Context Protocol (MCP) servers - Integrate with AI agents and assistants through MCP servers to add your RAG pipeline to watsonx Orchestrate or other AI agents or query structured data in the lakehouse with an AI agent on your local computer.
Usage resources
Depending on your service plans, you might have a set amount of usage resources per month, or you might be billed for the resources that you consume. When you run tools on watsonx.data as a Service, you consume the following types of resources:
- Compute usage
- When you run jobs, or deployments, your compute resource usage is calculated based on the rate for the runtime environment and its active duration. Compute resources include the appropriate hardware and software that are specific to the workload. Compute usage is measured in capacity unit hours.
- Text extraction
- When you use text extraction to convert document files into an AI model-friendly JSON file format, you are charged per page.
Storage
Your object storage service is IBM Cloud Object Storage.
On IBM Cloud, watsonx.data is offered under Enterprise plan. You must have a pay-as-you-go or subscription IBM Cloud account to avail the Enterprise plan. It is available on IBM Cloud and AWS environments. The paid plan has Presto, Spark (Serverless Spark), Milvus, Apache Gluten accelerated Spark engines available on IBM Cloud. AWS environment includes Cassandra as well. For information about the usage, see Billing details for watsonx.data as a Service.