Apache Hive lineage configuration

To import lineage metadata from Apache Hive, create a connection, data source definition and metadata import job.

Apache Hive is a data warehouse software project that provides data query and analysis and is built on top of Apache Hadoop.

Supported Apache Hive versions

  • Apache Hive 2.3 or later

Processed metadata

The following Apache Hive metadata is processed and displayed on lineage:

  • Hive metastore
  • Tables
  • Views
  • HiveQL scripts

Limitations

The following limitations apply:

  • You can't import lineage in the instances that operate in a FIPS‑compliant environment, because the JDBC drivers are not FIPS‑compliant.
  • Assets that use dynamically executed code through the EXECUTE IMMEDIATE statement are not processed.

Prerequisite configuration

Before you import lineage metadata, ensure that the following prerequisites are met:

  • You must have SELECT privileges with GRANT OPTION on all tables in the Hive databases to enable execution of SHOW CREATE TABLE during metadata extraction.
  • You must be authorized to run these statements: SHOW DATABASES, SHOW CREATE TABLE, SHOW TABLES IN, SHOW FUNCTIONS LIKE, and SHOW VIEWS IN.

Creating a metadata import asset

Data source connection

To connect to the data source from which you want to import lineage metadata, you need to select a data source definition and a connection. You can create them before you start creating the metadata import, or you can create them when you create and configure the metadata import asset.

Data source definition

Select Apache Hive as the data source type.

Connection

For connection details, see Apache Hive connection.

Connection mode

You can connect to Apache Hive by using one of the following connection modes:

Include and exclude lists

You can include or exclude databases. Provide databases by their names separated by a comma. Each part is evaluated as a regular expression. Assets which are added later in the data source will also be included or excluded if they match the conditions specified in the lists. Example values:

  • myDB/: all assets in myDB,
  • myDB2/.*: all assets in myDB2.

External inputs

You can add external Apache Hive QL scripts that work with the configured database in a .zip file as external inputs. You can organize the structure of a .zip file as subfolders that represent databases. After the scripts are scanned, they are added under respective databases in the selected catalog or project. The .zip file can have the following structure:

<database_name>
   <script_name.sql>
<script_name.sql>
replace.csv
linkedServerConnectionsConfiguration.prm

The replace.csv file contains placeholder replacements for the scripts that are added in the .zip file. For more information about the format, see Placeholder replacements.

The linkedServerConnectionsConfiguration.prm file contains linked server connection definitions.

Advanced import options

Learn more