IBM Data Cataloging and Content-Aware Storage insights

IBM Data Cataloging and Content-Aware Storage (CAS) introduce data insights dashboards that provide immediate visibility into data once sources are connected and scanned. This experience is enabled when both IBM Data Cataloging and CAS are integrated. The provided insights surface key indicators of data composition on each established data source.

The UI dashboard surfaces foundational indicators of the associated data sources from the storage. No additional configuration is required to start using the dashboard.

Prerequisites

Overview

When a data source is added on CAS, the dashboards display insights that are scoped to that source.

You can go to Fusion UI > Content-Aware Storage > Data Sources > Details View.

The experience is organized into the following sections:

  • Data source summary: High-level file and storage metrics.
  • Data composition: Breakdown of file types and file size distribution.

Data source summary

At the top of the Data source view, four summary cards display key metrics for the selected data source.

Metric Description
Total files Total number of files that are indexed by IBM Data Cataloging across all data sources and prefixes for the data source.
Total storage Aggregate storage capacity consumed, reported in GiB.
Ready to process Number and percentage of files that match the CAS ingestion criteria defined in the active ruleset. For more information, see Managing IBM Data Cataloging rules when integrated with CAS. When IBM Data Cataloging and CAS are integrated, this value reflects the result of the rules that are specified in the custom resource.
Sensitive data Count of files flagged as containing sensitive content. See the Sensitive data section for details on how this count is determined.

Data composition

The Data composition section provides a structural breakdown of the cataloged dataset across two dimensions.

File types and formats

Files are grouped by category (Documents, Images, Videos, Audio, Other). Each category shows the total file count and storage size, and distinguishes between formats that CAS can process and those it cannot. Expanding a category reveals the individual extensions within each group.

The following file formats are supported for CAS ingestion:

pdf, doc, docx, txt, html, markdown, md, pptx, jpeg, png, bmp, xlsx, asciidoc, adoc, xhtml, csv, webp, wav, ogg, opus, mp3

For more information, see Managing IBM Data Cataloging rules when integrated with CAS. All other extensions appear under the Not Supported label within their category.

File size distribution

A horizontal bar chart shows the number of files that are grouped by size range. The buckets used are: < 1 MB, 1-10 MB, 10-100 MB, 100 MB-1 GB, and > 1 GB.

This metric is derived from a single aggregated API query against the IBM Data Cataloging catalog.