DataStage Integration Requirements
The following are the prerequisites necessary for IBM Automatic Data Lineage to connect to this third-party system, which you may choose to do at your sole discretion. Note that while these are usually sufficient to connect to this third-party system, we cannot guarantee the success of the connection or integration since we have no control, liability, or responsibility for third-party products or services, including their performance.
-
If using standalone DataStage:
-
IBM InfoSphere DataStage version 11.7 (other versions may work but haven't been tested)
-
The property
datastage.editionmust be set toStandalone DataStage. -
Currently, only parallel jobs are supported and can be imported to Automatic Data Lineage.
-
DataStage parallel job metadata must be exported from the DataStage Designer, by using the "xml" option in order to be processed by Automatic Data Lineage. Exports using the dscmdexport or dsexport tools are also supported. However, it is advised to use the DataStage Designer for optimal selection of Jobs, Parameter Sets and other related metadata. InfoSphere Information Server
.isxstyle exports are not supported.
-
-
If using DataStage for Cloud Pak for Data:
-
The property
datastage.editionmust be set toDataStage for Cloud Pak for Data. -
DataStage project metadata must be exported using the "Export to Desktop" option in the project.
-
The export should contain all flows you want to analyze together with all subflows, connections, and parameter sets used in these flows. You can also export jobs, which will provide additional configuration to the flows and can improve the lineage in some cases. If the export contains a job, it should also contain the corresponding flow with its dependencies, otherwise the job will be ignored.
-
The simplest way is to export everything in the project.
-
-
The appropriate
*.zipDataStage export files must be placed in the folder specified by the configurabledatastage.input.dirproperty, which by default is set to<MANTAFLOW_INSTALLATION_DIRECTORY>/cli/input/datastage/<CONNECTION_NAME>.
-
-
If you want to add new parameters to your project or override old ones, create a
datastageParameterOverride.txtfile in the location defined by thedatastage.parameter.override.fileproperty and fill it in according to the format defined below in the section describing the parameter override file.- After each analysis, when any unresolved parameters are found, a "helper" file
datastageParameterOverride.txtis generated in the location defined by thedatastage.parameter.override.helper.fileproperty. The hint the helper file gives can also be found in the log file and in logs inside Admin UI Log Viewer relevant to the current DataStage dataflow analysis. You can replace the original file with this helper file and complete the missing parameter values. Detailed instructions can be found in the section describing the parameter override file.
- After each analysis, when any unresolved parameters are found, a "helper" file
-
For unsupported stages, which use the wise method (find out more about this method below), you can link input and output columns manually instead. For manual mapping, create a CSV file defined by the
datastage.manual.mapping.fileproperty and fill it in according to the format defined below in the paragraph about manual mapping files. -
If you are loading SQL queries from files in your database connectors (Oracle, Db2, etc.) and you have Automatic Data Lineage installed on a server other than Manta Server, Automatic Data Lineage won't be able to find the SQL files, so you need to provide Automatic Data Lineage with the SQL files for correct analysis. The SQL files should be placed in the directory defined by the
datastage.sql.files.directoryproperty.
Supported Stages
Here is a list of supported stages. (Unlisted stages may be supported somehow, but no guarantees are provided. Also, if Automatic Data Lineage does not provide support for a third-party technology listed below, such as iWay or Classic Federation, then the flow analysis will end inside DataStage.)
Processing Stage
|
Processing Stage |
Standalone DS |
DSNG |
|---|---|---|
|
Aggregator |
|
|
|
Bloom Filter |
|
|
|
Change Apply |
|
|
|
Change Capture |
|
|
|
Checksum |
|
|
|
Compare |
|
|
|
Compress |
|
|
|
Copy |
|
|
|
Data Masking |
|
|
|
Decode |
|
|
|
Difference |
|
|
|
Encode |
|
|
|
Expand |
|
|
|
External Filter |
An external filter can be anything, so we provide a wise column connection of input and output columns. Find out more about the wise approach below. |
An external filter can be anything, so we provide a wise column connection of input and output columns. Find out more about the wise approach below. |
|
Filter |
|
|
|
FTP Enterprise |
|
|
|
FTP Plugin |
|
|
|
Funnel |
|
|
|
Generic |
A generic operator can be any orchestrate operator, so we provide a wise column connection of input and output columns. Find out more about the wise approach below. |
A generic operator can be any orchestrate operator, so we provide a wise column connection of input and output columns. Find out more about the wise approach below. |
|
Head |
|
|
|
Join |
|
|
|
Lookup |
|
|
|
Merge |
|
|
|
Modify |
|
|
|
Peek |
|
|
|
Pivot |
|
|
|
Pivot Enterprise |
|
|
|
Remove Duplicates |
|
|
|
Row Generator |
|
|
|
Slowly Changing Dimension |
|
|
|
Sort |
|
|
|
Surrogate Key Generator |
|
|
|
Switch |
|
|
|
Transformer |
DS Routines and functions in expressions are resolved as direct flows. |
DS Routines and functions in expressions are resolved as direct flows. |
|
Wave Generator |
|
|
API Stage
| API Stage | Standalone DS | DSNG |
|---|---|---|
| REST |
Quality stage
| Quality stage | Standalone DS | DSNG |
|---|---|---|
| Data Rule |
File Stage
| File Stage | Standalone DS | DSNG |
|---|---|---|
| Amazon S3 | ||
| Azure Storage | ||
| Azure Blob Storage | ||
| Azure File Storage | ||
| Cloud Object Storage | ||
| Complex Flat File | ||
| Data Set | ||
| External Source | ||
| External Target | ||
| File Connector | ||
| File Set | ||
| Lookup File Set | ||
| Sequential File | ||
| Unstructured Data | ||
| zOS File |
Container stage
| Container stage | Standalone DS | DSNG |
|---|---|---|
| Local Container | ||
| Shared Container |
Database stage
|
Database stage |
Standalone DS |
DSNG |
|---|---|---|
|
Classic Federation |
The flow will end in DataStage because Automatic Data Lineage doesn't support the Classic Federation database. |
The flow will end in DataStage because Automatic Data Lineage doesn't support the Classic Federation database. |
|
Cognos TM1 |
The flow will end in DataStage because Automatic Data Lineage doesn't support the Cognos TM1 tool. |
The flow will end in DataStage because Automatic Data Lineage doesn't support the Cognos TM1 tool. |
|
Db2 |
|
|
|
Distributed Transaction |
|
|
|
DRS |
|
|
|
Greenplum |
|
|
|
HBase |
The flow will end in DataStage because Automatic Data Lineage doesn't support the HBase database. |
The flow will end in DataStage because Automatic Data Lineage doesn't support the HBase database. |
|
Hive |
|
|
|
Informix Bulk Load |
The flow will end in DataStage because Automatic Data Lineage doesn't support the Informix tool. |
The flow will end in DataStage because Automatic Data Lineage doesn't support the Informix tool. |
|
Informix CLI |
The flow will end in DataStage because Automatic Data Lineage doesn't support the Informix tool. |
The flow will end in DataStage because Automatic Data Lineage doesn't support the Informix tool. |
|
Informix Enterprise |
The flow will end in DataStage because Automatic Data Lineage doesn't support the Informix tool. |
The flow will end in DataStage because Automatic Data Lineage doesn't support the Informix tool. |
|
iWay |
The flow will end in DataStage because Automatic Data Lineage doesn't support the iWay database. |
The flow will end in DataStage because Automatic Data Lineage doesn't support the iWay database. |
|
JDBC |
|
|
|
MS SQL |
|
|
|
Netezza |
|
|
|
ODBC |
|
|
|
Oracle |
|
|
|
PX Oracle (this stage does not seem to exist in Standalone DS in 11.7, but exists in 9.0) |
|
|
|
PostgreSQL |
|
|
|
Stored Procedure |
Automatic Data Lineage doesn't support stored procedures for Sybase and Teradata. |
Automatic Data Lineage doesn't support stored procedures for Sybase and Teradata. |
|
Sybase Bulk Load |
|
|
|
Sybase Enterprise |
|
|
|
Sybase OC |
|
|
|
Teradata |
|
|
|
BigQuery |
|
|
Supported Features in DataStage for Cloud Pak for Data
Automatic Data Lineage supports the analysis of jobs, parameter sets, connection assets, subflows, and local subflows.
Analysis of Job Runs
Runtime parameters from job runs can be used to analyze corresponding jobs. To use parameters from job runs, set the property datastage.analyze.job.runs to true. If you only want to consider runs that are more recent than a certain
date, configure the date in datastage.analyze.job.runs.since. These are only applicable if the job runs metadata is present in the export.
Since R42.10, the feature is replaced with runtime column propagation. For more information, see Lineage with runtime column propagation.
Wise Column Connection
By default, Automatic Data Lineage will generate lineage to connect 1-to-1 columns with matching names. Input columns that don't have corresponding matching output columns to be connected to all output columns except those that have the same name for the input and output. The behavior of the output columns is similar.
Set common property datastage.unsupported.stages.connect.all to true to have all input column without a corresponding output column connected to all output columns; all output columns that don’t have a corresponding input column connected
to all input columns.
Local connections
As of R42.8 embedded (stage-level) connections for Datastage Next Gen are supported.
Known Unsupported Features
Automatic Data Lineage does not support the following DataStage features. This list includes all of the features that IBM is aware are unsupported, but it might not be comprehensive.
-
Common data sorting — https://www.ibm.com/support/knowledgecenter/en/SSZJPZ_11.7.0/com.ibm.swg.im.iis.ds.parjob.dev.doc/topics/sortingdata.html
-
Infosphere (Standalone) Datastage:
-
.dsxexport format -
.xmlexport using istool and exports based on Information Server .isx files are not supported -
Runtime Column Propagation (RCP)
-
Server jobs
-
-
RCP support for Datastage Next Generation / Datastage on Cloudpak for Data as of R42.10:
-
Lineage for job runs can be generated only for projects that are automatically extracted. It cannot be generated from the information in a .zip file provided during metadata import.
-
Lineage cannot be generated for job runs with subflows.
-
Some column node attributes might not be included in the job run lineage.
-
Lineage for job runs cannot be created for DataStage as a Service and DataStage on Cloud Pak for Data as a Service.
-