Triggering conditions for virtual machine alerts

You can set up IBM Storage Insights Pro so that it examines the attributes, capacity, and performance of a virtual machine and notifies you when changes or violations are detected.

Alerts can notify you of general changes, capacity, and performance issues on the following resources:
Important: Not all the attributes upon which you can alert are listed here. To view a complete list of attributes upon which you can alert, you can either edit alert definitions for a host with no alert policy assigned, or you can create a new custom alert policy and define new alerts in the policy. To create a custom alert policy, go to Configuration > Alert Policies and click Create Policy.

In the tables, default alerts are marked with an asterisk (*).

Virtual Machines (general changes)

Table 1. Pre-defined alerts for Virtual machines
General Attributes Defining Conditions for Attributes
Status
One of the following statuses is detected on a virtual machine:
Not Normal
An error or warning status was detected on the virtual machine.
Warning
A warning status was detected on the virtual machine.
Error (default)
An error status was detected on the virtual machine.
Unreachable
One or more of the monitored resources for a virtual machine are not responding. This status might be caused by a problem in the network.
Removed Virtual Machine A previously monitored virtual machine can no longer be found. Historical data about the virtual machine is retained, but no current data is being collected. Use this alert to be notified if an virtual machine is removed or becomes unavailable.
New Virtual Machine A virtual machine is detected for the first time.

Virtual Machines (capacity changes)

Table 2. Pre-defined alerts for Virtual Machines
General Attributes Defining Conditions for Attributes
SAN Capacity The total amount of storage capacity assigned to a virtual machine hosted by the ESXi host.
Used SAN Capacity The amount of used storage capacity on a virtual machine hosted by the ESXi host.

Virtual Machine (performance changes)

Define alerts that notify you when the performance of a virtual machine falls outside a specified threshold. In alerts, you can specify conditions based on metrics that measure the performance of virtual machine, including Response Time, I/O rates, and Data rates. By creating alerts with performance conditions, you can be informed about potential bottlenecks in your network infrastructure.

For example, you can define an alert to be notified when the Total Response Time for a disk is greater than or equal to a specified threshold. Total Response Time represents the average number of milliseconds for the back-end storage resources to respond to a read or a write operation. Use this alert to help identify disk conditions that might slow the performance of the virtual machine.

You can also be notified when a metric is less than a specified threshold, such as when you want to identify disks that are having low Response Time.

For a complete list of virtual machine metrics that can be alerted upon, see Performance metrics for virtual machines.
Tips for performance conditions:
  • Data Collection must run against a resource for a time before IBM Storage Insights Pro can determine whether a threshold is violated and an alert is generated for a performance condition.
Best practice: When you set thresholds for performance conditions, try to determine the best value so you can derive the maximum benefit without generating too many false alerts. Because suitable thresholds are highly dependent on the type of workload that is being run, hardware configuration, the number of physical disks, exact model numbers, and other factors, there are no easy or standard default rules.

A recommended approach is to monitor the performance of resources for a number of weeks and by using this historical data, determine reasonable threshold values for each performance condition. After that is done, you can fine-tune the condition settings to minimize the number of false alerts.

Click Edit alert Definitions, and for each performance alert definition click View History to see the history of host performance and set the threshold that you want relative to that data.

HBAs (general changes)

Table 3. Pre-defined alerts for HBAs
General Attributes Defining Conditions for Attributes
Removed HBA A previously monitored HBA can no longer be found. Historical data about the HBA is retained, but no current data is being collected. Use this alert to be notified if an HBA is removed or becomes unavailable.
New HBA An HBA is detected for the first time.

Drives (general changes)

Table 4. Pre-defined alerts for Drives
General Attributes Defining Conditions for Attributes
Removed Drive A previously monitored drive can no longer be found. Historical data about the drive is retained, but no current data is being collected. Use this alert to be notified if a drive is removed or becomes unavailable.
New Drive A drive is detected for the first time. Use this alert to be notified of hardware changes on hosts.
Paths The number of access paths that are associated with the drive falls outside a specified range, or is equal to or not equal to a specified value.

Drives (capacity changes)

Table 5. Pre-defined alerts for Drives
General Attributes Defining Conditions for Attributes
Capacity The total amount of storage capacity that is assigned to a virtual machine drive.
Available Drive Capacity The unused capacity on a virtual machine drive.
Used Capacity The amount of used storage capacity on a virtual machine drive.

Drives (performance changes)

Table 6. Pre-defined alerts for Drives
General Attributes Defining Conditions for Attributes
I/O Rate (Read) The average number of read operations per second that are issued to the back-end storage resources.
I/O Rate (Write) The average number of write operations per second that are issued to the back-end storage resources.
I/O Rate (Total) The average number of I/O operations per second that are transmitted between the back-end storage resources and the component. This value includes both read and write operations.
Data Rate (Read) The average number of MiB per second that are read from the back-end storage resources.
Data Rate (Write) The average number of MiB per second that are written to the back-end storage resources.
Data Rate (Total) The average rate at which data is transmitted between the back-end storage resources and the component. The rate is measured in MiB per second and includes both read and write operations.
Response Time (Read) The average number of milliseconds for the back-end storage resources to respond to a read operation.
Response Time (Write) The average number of milliseconds for the back-end storage resources to respond to a write operation.
Response Time (Total) The average number of milliseconds for the back-end storage resources to respond to a read or a write operation.

Paths (general changes)

Table 7. Pre-defined alerts for Paths
General Attributes Defining Conditions for Attributes
Status
One of the following statuses is detected on a path:
Not Normal
An error or warning status was detected on the path.
Warning
A warning status was detected on the path.
Error (default)
An error status was detected on the path.
Deleted Path A previously monitored access path for a virtual machine disk can no longer be found. This change might or might not affect the availability of the disk because there might be more than one path available.
New Path An access path for a disk is detected for the first time.