Start of change

Recoverygroup events

The following table lists the events that are created for the Recoverygroup component.

Table 1. Events for the Recoverygroup component
Event Event
Type
Severity Call Home Details
gnr_da_rebuild_failed INFO_EXTERNAL WARNING no Message: The fault tolerance on node {0} for recovery group {1} and declustered array {2} has remaining redundancy {3}.
Description: A declustered array rebuild failed.
Cause: The daRebuildFailed callback is invoked.
User Action: A recovery process should be active. Use the mmlsrecoverygroup command or the 'IBM Storage Scale RAID callbacks' documentation in 'IBM Storage Scale Erasure Code Edition' for more information. A possible problem could be exhausted spare space in a declustered array. If the problem is not solved within a couple of minutes, contact your IBM support.
gnr_log_groups_failed STATE_CHANGE
DEGRADED
WARNING no Message: Recovery group {id} has unassigned log groups: {0}.
Description: One or more log groups have no server assigned. Data access might be affected.
Cause: The mmvdisk rg list --log-group --rg <rg-name> command shows a log group without a server assignment.
User Action: Use the mmvdisk command to check the log group server. Check /var/adm/ras/mmfs.log.latest for more information.
gnr_log_groups_ok STATE_CHANGE
HEALTHY
INFO no Message: All log groups of recovery group {id} have a server assigned.
Description: Log Groups have a recovery group server.
Cause: N/A
User Action: N/A
gnr_rg_failed STATE_CHANGE
FAILED
ERROR no Message: GNR recovery group {0} is not active.
Description: A configured recovery group is not listed as active.
Cause: The mmlsrecoverygroup command reports that a recovery group is configured but not listed.
User Action: Examine the health of the recovery group server node and resolve any found health issues. Issue the mmlsrecoverygroup command to verify your modifications.
gnr_rg_found ADD_ENTITY INFO no Message: GNR recovery group {0} is found.
Description: A GNR recovery group, which is listed in the IBM Storage Scale configuration, was detected.
Cause: N/A
User Action: N/A
gnr_rg_is_primary STATE_CHANGE
HEALTHY
INFO no Message: The recovery group server {id} is the primary server.
Description: The recovery group server is the primary server.
Cause: N/A
User Action: N/A
gnr_rg_not_primary STATE_CHANGE
DEGRADED
WARNING no Message: Do not upgrade the system because the recovery group {id} active server {0} is not the primary server. Server list: {1}.
Description: The recovery group server is not the primary server.
Cause: The mmlsrecoverygroup <gr-group> -Y command shows that the ActiveRecoveryGroupServer is not the first server in the server list.
User Action: The 'tsrecgroupserver' command should be used to set the recovery group server as the primary server to avoid problems during the next upgrade.
gnr_rg_ok STATE_CHANGE
HEALTHY
INFO no Message: GNR recovery group {0} is OK.
Description: The recovery group is OK.
Cause: N/A
User Action: N/A
gnr_rg_relinquish: INFO_EXTERNAL INFO no Message: GNR recovery group server on node {0} has relinquished recovery group {1} for reason {2}.
Description: A recovery group server relinquished it's recovery groups.
Cause: This event is triggered through 'postRGRelinquish' callback.
User Action: N/A
gnr_rg_relinquish_warn INFO_EXTERNAL WARNING no Message: GNR recovery group server on node {0} relinquished recovery group {3} with error {1} and reports {2}.
Description: A recovery group server has problems to relinquish a recovery group.
Cause: The postRGRelinquish callback indicates a problem.
User Action: For debugging information see the Elastic Storage Server, section Recovery Group Issues.
gnr_rg_server_down STATE_CHANGE
DEGRADED
WARNING no Message: GNR recovery group server {0} in resource group {1} is unresponsive.
Description: The server node in this recovery group is reported as unresponsive.
Cause: The server node in this recovery group is down or unresponsive.
User Action: Examine the health of the recovery group server node and resolve any found health issues.
gnr_rg_server_panic INFO_EXTERNAL WARNING no Message: The GNR recovery group server on node {0} is no longer able to continue serving recovery group {3} with error {1} for reason {2}.
Description: A recovery group server stopped serving a recovery group.
Cause: The rgPanic callback is invoked.
User Action: A recovery process should be active. Use the mmlsrecoverygroup command or the 'IBM Storage Scale RAID callbacks' documentation in 'IBM Storage Scale Erasure Code Edition' for more information. A possible problem could be a faulty drive or cable. If the problem is not solved within a couple of minutes, contact your IBM support.
gnr_rg_server_up STATE_CHANGE
HEALTHY
INFO no Message: GNR recovery group server {0} in resource group {1} is active.
Description: The recovery group server node, which was previously reported as unresponsive, is now active.
Cause: N/A
User Action: N/A
gnr_rg_takeover INFO_EXTERNAL INFO no Message: The take over of recovery group {1} was done by GNR recovery group server on node {0} for reason {2}.
Description: The take over of a recovery group was successful.
Cause: This event is triggered through 'postRGTakeover' callback.
User Action: N/A
gnr_rg_takeover_warn INFO_EXTERNAL WARNING no Message: GNR recovery group server on node {0} has error {1} and reports {2} for recovery group {3}.
Description: A recovery group server is not working as expected.
Cause: The postRGTakeover callback indicates a problem.
User Action: For debugging information see the Elastic Storage Server, section Recovery Group Issues.
gnr_rg_vanished DELETE_ENTITY INFO no Message: GNR recovery group {0} has vanished.
Description: A GNR recovery group, which was previously listed in the IBM Storage Scale configuration, was not detected.
Cause: A GNR recovery group, which was previously listed in the IBM Storage Scale configuration, is no longer found. This can be a valid situation.
User Action: Run the mmlsrecoverygroup command to verify that all expected GNR recovery groups exist.
perfmon_gpfsfcm_active TIP INFO no Message: The GPFSFCM perfmon sensor {0} is active.
Description: The GPFSFCM perfmon sensors are active. This event's monitor is running only once an hour.
Cause: The GPFSFCM perfmon sensors' period attribute is greater than 0.
User Action: N/A
perfmon_gpfsfcm_inactive TIP TIP no Message: The GPFSFCM perfmon sensor {0} is inactive.
Description: The GPFSFCM perfmon sensors are inactive. This event's monitor is running only once an hour.
Cause: The GPFSFCM perfmon sensors' period attribute is 0.
User Action: Set the period attribute of the GPFSFCM sensors to a value greater than 0. For more information, use the mmperfmon config update SensorName.period=N command where 'SensorName' is the name of a specific GPFSFCM sensor and 'N' is a natural number greater than 0. Consider that this TIP monitor is running only once per hour and it might take up to one hour to detect the changes in the configuration.
perfmon_gpfsfcm_not_configured TIP TIP no Message: The GPFSFCM perfmon sensor {0} is not configured.
Description: The GPFSFCM perfmon sensor does not exist in the mmperfmon config show command.
Cause: The GPFSFCM perfmon sensor is not configured in the sensors' configuration file.
User Action: Include the sensors into the perfmon configuration by using the mmperfmon config add --sensors /opt/IBM/zimon/defaults/ZIMonSensors_GPFSFCM.cfg command. An example for the configuration file can be found in the mmperfmon command page in the Command Reference Guide.
perfmon_gpfsfcm_not_needed TIP INFO no Message: The GPFSFCM perfmon sensor not monitored as FCM drives are not present or have been removed from cluster.
Description: The GPFSFCM perfmon sensor status not monitored as FCM drives are not present or have been removed from cluster.
Cause: N/A
User Action: N/A
End of change