Integrating with DataStage

You can integrate IBM® watsonx.data with IBM DataStage on Cloud Pak for Data to ingest and to read data from watsonx.data.

watsonx.data on Red Hat® OpenShift®

watsonx.data Developer edition

Requirements
Following are the requirements to integrate watsonx.data with DataStage:
  • Cloud Pak for Data with DataStage (version 5.0.1 or later)
  • IBM watsonx.data (version 2.0.0 or later)
    Note: Ingestion is not supported by watsonx.data developer edition in DataStage. The developer edition supports only reading data.
  • Use watsonx.data connector in DataStage to read or ingest data into watsonx.data.
  • Make sure that the Data Access Service (DAS) is enabled in watsonx.data to enable data ingestion in DataStage.
  • Presto engine must be provisioned in watsonx.data to read and ingest data in DataStage.
  • Amazon S3 or IBM Cloud Object storage must be connected to watsonx.data to enable data ingestion in DataStage.
  • The Amazon S3 or IBM Cloud Object storage must be connected to the Iceberg catalog, which must be associated to the Presto engine to enable data ingestion in DataStage.

For more information about connecting to a data source in DataStage, see Connecting to a data source in DataStage.

For more information about creating a connection asset in watsonx.data, see IBM watsonx.data connection.

Limitation

Datastage ignores data rollback leading to data inconsistency in watsonx.data

The Datastage fails to consider the rollback actions in the watsonx.data tables when writing data. This results in data inconsistency in the watsonx.data when multiple transactions are involved. Specifically, if a rollback is performed on one of the transactions, the rollback operation fails, retaining the data in watsonx.data leading to incorrect data.