IBM Data Cataloging and Content-Aware Storage insights
IBM Data Cataloging and Content-Aware Storage (CAS) introduce data insights dashboards that provide immediate visibility into data once sources are connected and scanned. This experience is enabled when both IBM Data Cataloging and CAS are integrated. The provided insights surface key indicators of data composition on each established data source.
The UI dashboard surfaces foundational indicators of the associated data sources from the storage. No additional configuration is required to start using the dashboard.
Prerequisites
- Content-Aware Storage v1.1.5
- IBM Data Cataloging v2.5.3
- IBM Storage Scale v6.0.0.2 or later
- Ensure that both services are integrated. For more information, see Integrating IBM Data Cataloging for file filtering.
Overview
When a data source is added on CAS, the dashboards display insights that are scoped to that source.
You can go to .
The experience is organized into the following sections:
- Data source summary: High-level file and storage metrics.
- Data composition: Breakdown of file types and file size distribution.
Data source summary
At the top of the Data source view, four summary cards display key metrics for the selected data source.
| Metric | Description |
|---|---|
| Total files | Total number of files that are indexed by IBM Data Cataloging across all data sources and prefixes for the data source. |
| Total storage | Aggregate storage capacity consumed, reported in GiB. |
| Ready to process | Number and percentage of files that match the CAS ingestion criteria defined in the active ruleset. For more information, see Managing IBM Data Cataloging rules when integrated with CAS. When IBM Data Cataloging and CAS are integrated, this value reflects the result of the rules that are specified in the custom resource. |
| Sensitive data | Count of files flagged as containing sensitive content. See the Sensitive data section for details on how this count is determined. |
Data composition
The Data composition section provides a structural breakdown of the cataloged dataset across two dimensions.
- File types and formats
-
Files are grouped by category (Documents, Images, Videos, Audio, Other). Each category shows the total file count and storage size, and distinguishes between formats that CAS can process and those it cannot. Expanding a category reveals the individual extensions within each group.
The following file formats are supported for CAS ingestion:
pdf,doc,docx,txt,html,markdown,md,pptx,jpeg,png,bmp,xlsx,asciidoc,adoc,xhtml,csv,webp,wav,ogg,opus,mp3For more information, see Managing IBM Data Cataloging rules when integrated with CAS. All other extensions appear under the Not Supported label within their category.
- File size distribution
-
A horizontal bar chart shows the number of files that are grouped by size range. The buckets used are: < 1 MB, 1-10 MB, 10-100 MB, 100 MB-1 GB, and > 1 GB.
This metric is derived from a single aggregated API query against the IBM Data Cataloging catalog.