Overview of the Data Fabric experience
The Data Fabric experience is a set of services on IBM Software Hub that provides a secure and collaborative environment where you can provide business users with trusted data of high quality while managing governance and compliance demands. The watsonx.data intelligence services provide a data governance framework for curating structured and unstructured data, sharing data products, and exploring data lineage. The watsonx.data integration services provide a set of integrated tools to transform, integrate, and observe your data.
A data fabric architecture implements active metadata management, which uses machine learning to automate metadata processing. The outcomes of the metadata analysis facilitate automated data discovery, improve confidence in data, and enable data protection and data governance at scale.
Data engineers, data stewards, and compliance officers can accomplish the following goals in the Data Fabric experience:
- Prepare data for AI
- Import and enrich structured and unstructured data with your business vocabulary.
- Govern data
- Track and control data assets based on assigned metadata. Protect sensitive data and ensure data quality compliance.
- Monitor data
- View complete data lineage.
- Share data across your organization
- Share data assets with platform users in a catalog or share data products with anyone in your organization.
IBM watsonx.data intelligence
The IBM watsonx.data intelligence service includes APIs and tools for curating, governing, and sharing data.
Tools for curating, governing, and sharing data
Your watsonx.data intelligence tools are in collaborative workspaces called projects.
You can use watsonx.data intelligence tools to prepare and govern data in the following ways:
- Curate structured data
-
- Refine and visualize your data files or data tables in remote data sources with Data Refinery.
-
- Import metadata that represents data tables in remote data sources with Metadata import.
-
- Enrich structured data assets with metadata that provides information about the data with Metadata enrichment.
-
- Define data quality with data quality definitions.
-
- Measure the quality of your data with data quality rules.
-
- Monitor data quality SLA violations.
- See Curating structured data.
- Curate unstructured data
-
- Import, analyze, transform, and enrich documents from unstructured data sources with Unstructured data curation.
-
- Ingest, cleanse, transform, and enrich unstructured data for RAG processing with Unstructured Data Integration.
- See Curating and integrating unstructured data.
- Govern data
-
- Establish a business vocabulary with governance artifacts, such as, business terms, data classes, reference datasets, and classifications.
-
- Import industry-specific vocabularies with Knowledge Accelerators.
-
- Protect data with policies and rules, such as data protection rules, governance rules, and data quality SLAs.
-
- Customize the workflow, organization, properties, and relationships of governance artifacts with workflows and categories.
-
- Analyze governance and data usage trends with reports.
- See Data governance.
- Share data
-
- Publish high-quality data to governed catalogs.
-
- Publish curated data products in Data Product Hub.
- See Catalogs and Publishing data products.
- Track data
-
- Track data as it is moved and used by different software tools with Data lineage.
- See Data lineage.
IBM watsonx.data integration
Use IBM watsonx.data integration to deliver actionable data at speed and scale.
Watsonx.data integration provides unified tools that you can use to transform, integrate, and observe your data. You can use a range of diverse data integration styles, such as streaming, replication, observability, and bulk or batch processing.
With watsonx.data integration, data engineers can access and connect data across various data sources, including databases, file systems, real-time web services and messaging systems, and other enterprise applications.
Data engineers use an intuitive graphical design interface to build data flows that can transform and integrate the following types of data:
- Structured data, such as data stored in relational databases or CSV files
- Semi-structured data, such as JSON or XML files
- Unstructured data, such as PDF, HTML, or markdown files
Data engineers can create alerts to track the health of the end-to-end data integration process, allowing immediate investigations of data incidents.
Tools for transforming, integrating, and observing data
Your watsonx.data integration tools are in collaborative workspaces called projects. You use the tools to transform, integrate, and observe your data.
You can transform, integrate, and observe data in the following ways:
- Transform and integrate data
-
- Transform batch data with DataStage to create batch data flows that extract data from multiple source systems, transform the data as required, and deliver the data to target systems. With batch data flows, you can use both ETL (Extract, Transform, Load) and ELT (Extract, Load, Transform) patterns. See Transforming data with DataStage.
-
- Stream real-time data with StreamSets to create streaming data flows that act on time-sensitive data. A streaming data flow runs continuously to read, process, and write data as soon as the data becomes available. You can add processors to streaming data flows to transform the data as it moves from source to target systems. See Streaming real-time data.
-
- Replicate data with Data Replication to build a replication pipeline that synchronizes data between a source and target data store. Use the Data Replication tool for near-real-time data delivery with low impact to source data stores. See Replicating data.
-
- Prepare unstructured data with Unstructured Data Integration to ingest, transform, and enrich unstructured data from diverse sources. See Integrating unstructured data documents.
- Observe data
-
- Create alerts with Data Observability that notify you when a data integration process encounters errors or behaves differently than you expect. For example, you can set up alerts that monitor DataStage job status.
-
- Investigate data incidents with Data Observability to solve any problems or issues that occur in data quality, integrity, and access. See Observing data.