Apache Hive lineage configuration
To import lineage metadata from Apache Hive, create a connection, data source definition, and metadata import job.
This information applies to IBM Manta Data Lineage service.
Apache Hive is a data warehouse software project that provides data query and analysis and is built on top of Apache Hadoop.
Supported Apache Hive versions
- Apache Hive 2.3 or later
Processed metadata
The following Apache Hive metadata is processed and displayed on lineage:
- Hive metastore
- Tables
- Views
- HiveQL scripts
Limitations
The following limitations apply:
- You can't import lineage in the instances that operate in a FIPS‑compliant environment, because the JDBC drivers are not FIPS‑compliant.
- Assets that use dynamically executed code through the EXECUTE IMMEDIATE statement are not processed.
Prerequisite configuration
Before you import lineage metadata, ensure that the following prerequisites are met:
- You must have
SELECTprivileges withGRANT OPTIONon all tables in the Hive databases to enable execution ofSHOW CREATE TABLEduring metadata extraction. - You must be authorized to run these statements:
SHOW DATABASES,SHOW CREATE TABLE,SHOW TABLES IN,SHOW FUNCTIONS LIKE, andSHOW VIEWS IN.
Creating a metadata import asset
Data source connection
To connect to the data source from which you want to import lineage metadata, you need to select a data source definition and a connection. You can create them before you start creating the metadata import, or you can create them when you create and configure the metadata import asset.
Data source definition
Select Apache Hive as the data source type.
Connection
For connection details, see Apache Hive connection.
Connection mode
You can connect to Apache Hive by using one of the following connection modes:
- Direct connection
- Remote connection with a Manta agent. When an agent is configured, select it from the list. For more information, see Configuring agents for lineage metadata import.
Include and exclude lists
You can include or exclude databases. Provide databases by their names separated by a comma. Each part is evaluated as a regular expression. Assets which are added later in the data source will also be included or excluded if they match the conditions specified in the lists. Example values:
myDB/: all assets inmyDB,myDB2/.*: all assets inmyDB2.
External inputs
You can add external Apache Hive QL scripts that work with the configured database in a .zip file as external inputs. You can organize the structure of a .zip file as subfolders that represent databases. After the scripts are scanned, they are added under respective databases in the selected catalog or project. The .zip file can have the following structure:
<database_name>
<script_name.sql>
<script_name.sql>
replace.csv
linkedServerConnectionsConfiguration.prm
The replace.csv file contains placeholder replacements for the scripts that are added in the .zip file. For more information about the format, see Placeholder replacements.
The linkedServerConnectionsConfiguration.prm file contains linked server connection definitions.
Advanced import options
- Performance profile
- For selected data sources you can choose a performance profile. Depending on your current needs, the lineage metadata import might be faster or more complete. You can choose between the following profiles:
- Fast: Low time and memory consumption are the priorities in this profile. If your input is large, lineage might not be complete.
- Balanced: Both performance and lineage completness are important. It is a compromise between the lineage completness and time and memory that is spent on lineage import.
- Complete: The completness for lineage is the priority in this profile. If your input is large, the lineage import might take a significant amount of resources and time.
- Custom profile: You can create your own performance profile by providing values for the following properties:
- Dataflow Analysis Timeout Limit: Specifies the maximum estimated time (in seconds) after which the dataflow analysis of a single input is stopped. The time is checked when each node is added, or in some cases when
edges are created. Therefore, in some cases, the timeout might slightly exceed the specified limit. If you set the value to 0, the analysis is not stopped. Example value:
60. - Dataflow Analysis Edge Limit: Specifies the maximum number of edges that are allowed for a single input during the dataflow analysis. If this limit is exceeded, all filter edges are removed and no more filter edges
are added. If the limit is still exceeded even after that, the analysis is stopped and the input fails. To disable the limit, set the value to 0. Example value:
2500.
- Dataflow Analysis Timeout Limit: Specifies the maximum estimated time (in seconds) after which the dataflow analysis of a single input is stopped. The time is checked when each node is added, or in some cases when
edges are created. Therefore, in some cases, the timeout might slightly exceed the specified limit. If you set the value to 0, the analysis is not stopped. Example value: