Recoverygroup events
The following table lists the events that are created for the Recoverygroup component.
| Event | Event Type |
Severity | Call Home | Details |
|---|---|---|---|---|
| gnr_da_rebuild_failed | INFO_EXTERNAL | WARNING | no | Message: The fault tolerance on node {0} for recovery group {1} and declustered array {2} has remaining redundancy {3}. |
| Description: A declustered array rebuild failed. | ||||
| Cause: The daRebuildFailed callback is invoked. | ||||
| User Action: A recovery process should be active. Use the mmlsrecoverygroup command or the 'IBM Storage Scale RAID callbacks' documentation in 'IBM Storage Scale Erasure Code Edition' for more information. A possible problem could be exhausted spare space in a declustered array. If the problem is not solved within a couple of minutes, contact your IBM support. | ||||
| gnr_log_groups_failed | STATE_CHANGE DEGRADED |
WARNING | no | Message: Recovery group {id} has unassigned log groups: {0}. |
| Description: One or more log groups have no server assigned. Data access might be affected. | ||||
| Cause: The mmvdisk rg list --log-group --rg <rg-name> command shows a log group without a server assignment. | ||||
| User Action: Use the mmvdisk command to check the log group server. Check /var/adm/ras/mmfs.log.latest for more information. | ||||
| gnr_log_groups_ok | STATE_CHANGE HEALTHY |
INFO | no | Message: All log groups of recovery group {id} have a server assigned. |
| Description: Log Groups have a recovery group server. | ||||
| Cause: N/A | ||||
| User Action: N/A | ||||
| gnr_rg_failed | STATE_CHANGE FAILED |
ERROR | no | Message: GNR recovery group {0} is not active. |
| Description: A configured recovery group is not listed as active. | ||||
| Cause: The mmlsrecoverygroup command reports that a recovery group is configured but not listed. | ||||
| User Action: Examine the health of the recovery group server node and resolve any found health issues. Issue the mmlsrecoverygroup command to verify your modifications. | ||||
| gnr_rg_found | ADD_ENTITY | INFO | no | Message: GNR recovery group {0} is found. |
| Description: A GNR recovery group, which is listed in the IBM Storage Scale configuration, was detected. | ||||
| Cause: N/A | ||||
| User Action: N/A | ||||
| gnr_rg_is_primary | STATE_CHANGE HEALTHY |
INFO | no | Message: The recovery group server {id} is the primary server. |
| Description: The recovery group server is the primary server. | ||||
| Cause: N/A | ||||
| User Action: N/A | ||||
| gnr_rg_not_primary | STATE_CHANGE DEGRADED |
WARNING | no | Message: Do not upgrade the system because the recovery group {id} active server {0} is not the primary server. Server list: {1}. |
| Description: The recovery group server is not the primary server. | ||||
| Cause: The mmlsrecoverygroup <gr-group> -Y command shows that the ActiveRecoveryGroupServer is not the first server in the server list. | ||||
| User Action: The 'tsrecgroupserver' command should be used to set the recovery group server as the primary server to avoid problems during the next upgrade. | ||||
| gnr_rg_ok | STATE_CHANGE HEALTHY |
INFO | no | Message: GNR recovery group {0} is OK. |
| Description: The recovery group is OK. | ||||
| Cause: N/A | ||||
| User Action: N/A | ||||
| gnr_rg_relinquish: | INFO_EXTERNAL | INFO | no | Message: GNR recovery group server on node {0} has relinquished recovery group {1} for reason {2}. |
| Description: A recovery group server relinquished it's recovery groups. | ||||
| Cause: This event is triggered through 'postRGRelinquish' callback. | ||||
| User Action: N/A | ||||
| gnr_rg_relinquish_warn | INFO_EXTERNAL | WARNING | no | Message: GNR recovery group server on node {0} relinquished recovery group {3} with error {1} and reports {2}. |
| Description: A recovery group server has problems to relinquish a recovery group. | ||||
| Cause: The postRGRelinquish callback indicates a problem. | ||||
| User Action: For debugging information see the Elastic Storage Server, section Recovery Group Issues. | ||||
| gnr_rg_server_down | STATE_CHANGE DEGRADED |
WARNING | no | Message: GNR recovery group server {0} in resource group {1} is unresponsive. |
| Description: The server node in this recovery group is reported as unresponsive. | ||||
| Cause: The server node in this recovery group is down or unresponsive. | ||||
| User Action: Examine the health of the recovery group server node and resolve any found health issues. | ||||
| gnr_rg_server_panic | INFO_EXTERNAL | WARNING | no | Message: The GNR recovery group server on node {0} is no longer able to continue serving recovery group {3} with error {1} for reason {2}. |
| Description: A recovery group server stopped serving a recovery group. | ||||
| Cause: The rgPanic callback is invoked. | ||||
| User Action: A recovery process should be active. Use the mmlsrecoverygroup command or the 'IBM Storage Scale RAID callbacks' documentation in 'IBM Storage Scale Erasure Code Edition' for more information. A possible problem could be a faulty drive or cable. If the problem is not solved within a couple of minutes, contact your IBM support. | ||||
| gnr_rg_server_up | STATE_CHANGE HEALTHY |
INFO | no | Message: GNR recovery group server {0} in resource group {1} is active. |
| Description: The recovery group server node, which was previously reported as unresponsive, is now active. | ||||
| Cause: N/A | ||||
| User Action: N/A | ||||
| gnr_rg_takeover | INFO_EXTERNAL | INFO | no | Message: The take over of recovery group {1} was done by GNR recovery group server on node {0} for reason {2}. |
| Description: The take over of a recovery group was successful. | ||||
| Cause: This event is triggered through 'postRGTakeover' callback. | ||||
| User Action: N/A | ||||
| gnr_rg_takeover_warn | INFO_EXTERNAL | WARNING | no | Message: GNR recovery group server on node {0} has error {1} and reports {2} for recovery group {3}. |
| Description: A recovery group server is not working as expected. | ||||
| Cause: The postRGTakeover callback indicates a problem. | ||||
| User Action: For debugging information see the Elastic Storage Server, section Recovery Group Issues. | ||||
| gnr_rg_vanished | DELETE_ENTITY | INFO | no | Message: GNR recovery group {0} has vanished. |
| Description: A GNR recovery group, which was previously listed in the IBM Storage Scale configuration, was not detected. | ||||
| Cause: A GNR recovery group, which was previously listed in the IBM Storage Scale configuration, is no longer found. This can be a valid situation. | ||||
| User Action: Run the mmlsrecoverygroup command to verify that all expected GNR recovery groups exist. | ||||
| perfmon_gpfsfcm_active | TIP | INFO | no | Message: The GPFSFCM perfmon sensor {0} is active. |
| Description: The GPFSFCM perfmon sensors are active. This event's monitor is running only once an hour. | ||||
| Cause: The GPFSFCM perfmon sensors' period attribute is greater than 0. | ||||
| User Action: N/A | ||||
| perfmon_gpfsfcm_inactive | TIP | TIP | no | Message: The GPFSFCM perfmon sensor {0} is inactive. |
| Description: The GPFSFCM perfmon sensors are inactive. This event's monitor is running only once an hour. | ||||
| Cause: The GPFSFCM perfmon sensors' period attribute is 0. | ||||
| User Action: Set the period attribute of the GPFSFCM sensors to a value greater than 0. For more information, use the mmperfmon config update SensorName.period=N command where 'SensorName' is the name of a specific GPFSFCM sensor and 'N' is a natural number greater than 0. Consider that this TIP monitor is running only once per hour and it might take up to one hour to detect the changes in the configuration. | ||||
| perfmon_gpfsfcm_not_configured | TIP | TIP | no | Message: The GPFSFCM perfmon sensor {0} is not configured. |
| Description: The GPFSFCM perfmon sensor does not exist in the mmperfmon config show command. | ||||
| Cause: The GPFSFCM perfmon sensor is not configured in the sensors' configuration file. | ||||
| User Action: Include the sensors into the perfmon configuration by using the mmperfmon config add --sensors /opt/IBM/zimon/defaults/ZIMonSensors_GPFSFCM.cfg command. An example for the configuration file can be found in the mmperfmon command page in the Command Reference Guide. | ||||
| perfmon_gpfsfcm_not_needed | TIP | INFO | no | Message: The GPFSFCM perfmon sensor not monitored as FCM drives are not present or have been removed from cluster. |
| Description: The GPFSFCM perfmon sensor status not monitored as FCM drives are not present or have been removed from cluster. | ||||
| Cause: N/A | ||||
| User Action: N/A |
