Troubleshooting clusters and nodes
Learn how to isolate and resolve problems that involve cluster and node failures.
- Temporal grouping policies not created after training historic events
- Illegal reflective access operation in Jobmanager pod log
- Concert Operate console not accessible after shutdown of backup cluster
- Log anomaly detection Kafka intermittent Issue, as Log anomaly detector abruptly stops making predictions
- Postgres database crashes
- Add additional network policies
- CouchDB pod reports a permission error when draining a node
- Kafka broker pods are unhealthy because Kafka PVCs are full
- aimanager is not in Ready state after cluster restart, integrations are missing from the console
- Deployment on Linux: unable to access Concert Operate console
- Deployment on Linux: API server down after cluster restart
- After restarting the Red Hat OpenShift cluster one or more Cassandra pods are missing or in Pending state
- After restarting the Red Hat OpenShift cluster one or more couchdb pods are missing or in Pending state
- OpenSearch pods restarting continuously due to insufficient entropy availability (deployments on Linux only)
Temporal grouping policies not created after training historic events
Groups of related alerts that occur together (temporal groups) are not correlated automatically by the system. This issue can occur when temporal grouping policies are not created automatically due to communication issues between the aiops-ir-core-archiving pod and some Kafka topics.
-
Run the following command, while logged in as an admin user in the project (namespace) where IBM Concert Operate is running:
kubectl logs $(kubectl get pods|grep aiops-ir-core-archiving|grep -v setup|awk '{print $1}') | grep "lib-rdkafka status"| tail -1Example command output:
{ "name": "client.kafka", "hostname": "aiops-ir-core-archiving-5dfd8fdcc5-l52g8", "pid": 20, "level": 30, "brokerStates": { "0": "UP", "-1": "UP" }, "partitionStates": { "cp4waiops-cartridge.irdatalayer.replay.alerts.0": "UP", "cp4waiops-cartridge.irdatalayer.replay.alerts.1": "UP", "cp4waiops-cartridge.irdatalayer.replay.alerts.2": "UP", "cp4waiops-cartridge.irdatalayer.replay.alerts.3": "UP", "cp4waiops-cartridge.irdatalayer.replay.alerts.4": "UP", "cp4waiops-cartridge.irdatalayer.replay.alerts.5": "UP" }, "msg": "lib-rdkafka status", "time": "2021-10-19T11:00:49.852Z", "v": 0 } -
If any of the broker or partition states are not up, restart the aiops-ir-core-archiving pod with the following command:
kubectl delete pod $(kubectl get pods|grep aiops-ir-core-archiving|grep -v setup|awk '{print $1}')
Illegal reflective access operation in Jobmanager pod log
Reports of illegal reflective access operation.
WARNING: An illegal reflective access operation has occurred
WARNING: Illegal reflective access by org.apache.flink.shaded.akka.org.jboss.netty.util.internal.ByteBufferUtil (file:/opt/flink/lib/flink-dist_2.11-1.13.2.jar) to method java.nio.DirectByteBuffer.cleaner()
WARNING: Please consider reporting this to the maintainers of org.apache.flink.shaded.akka.org.jboss.netty.util.internal.ByteBufferUtil
WARNING: Use --illegal-access=warn to enable warnings of further illegal reflective access operations
WARNING: All illegal access operations will be denied in a future release
This warning is reported when the Flink library is run by Java 11. For more information, see https://issues.apache.org/jira/browse/FLINK-17524
Solution: No action required. This warning message can be ignored and does not impact functionality.
Concert Operate console not accessible after shutdown of backup cluster
Unable to access the Concert Operate console on the backup cluster after a systematic shut down and start up of the backup cluster. The following error message is displayed:
"Document was encrypted with unknown key '<namespace>-<release name>-<timestamp>'"
Solution: Find and then delete the required databases through the API.
-
First, run the following command to retrieve the hostname:
oc get route couchdb-georedundancyFor this scenario, use the hostname
couchdb-georedundancy-geonoi.apps.geo01primary.myibm.com. -
Run the two
SECRET_NAMEcommands to retrieve theusernameandpasswordthat are required to retrieve the database names.For Username:
SECRET_NAME=$(oc get secret | grep couchdb-secret | awk '{ print $1 }'); kubectl get secret ${SECRET_NAME} -o json | grep "username" | cut -d : -f2 | cut -d '"' -f2 | base64 -d usernameFor Password:
SECRET_NAME=$(oc get secret | grep couchdb-secret | awk '{ print $1 }'); kubectl get secret ${SECRET_NAME} -o json | grep "password" | cut -d : -f2 | cut -d '"' -f2 | base64 -d password -
Use the
hostname,username, andpasswordobtained in the previous steps to run the command to retrieve the database names:curl -sk -u "username:password" https://couchdb-georedundancy-geonoi.apps.geo01primary.myibm.com/_all_dbs | jq "."Example output:
[ "_global_changes", "_replicator", "_users", "collabopsuser", "emailrecipients", "genericproperties", "icp-63666439356237652d336263372d343030362d613461382d613733613739633731323535-rba-as", "icp-63666439356237652d336263372d343030362d613461382d613733613739633731323535-rba-rbs", "integration", "noi-drdb", "noi-osregdb", "normalizercfd95b7e-3bc7-4006-a4a8-a73a79c71255", "osb-map", "otc_omaas_broker", "rba-pdoc", "schedule", "tenant", "trainingjob" ]The relevant database name here is collabopsuser.
-
Run the following command to delete the database:
curl -sk -u "root:netcool" -X DELETE "https://couchdb-georedundancy-geonoi.apps.geo01primary.myibm.com/${DATABASE_NAME}"Where
DATABASE_NAME=collabopsuser -
Restart the
cem-userspod to recreate the database:oc get pods |grep cem-users <cem_users_pod_name>oc delete pod <cem_users_pod_name> -
You can confirm the database is created by running the curl command from step 3:
curl -sk -u "username:password" https://couchdb-georedundancy-geonoi.apps.geo01primary.myibm.com/_all_dbs | jq "."
LAD Kafka intermittent Issue, as Log anomaly detector abruptly stops making predictions
The log anomaly detector abruptly stops making predictions.
On checking the anomaly pod logs, you see error messages related to Kafka, and due to these intermittent issues, log anomaly stops consuming data from the windowed logs topic, even when there are continuous windows generated from the data preprocessing component.
An example of such an error in the logs is as follows:
[2022-03-19 14:28:34,233] [logprophet.logs.watson_aiops] [ERROR] An error has occurred while processing message in batch mode. Error: 'NoneType' object is not iterable
[2022-03-19 14:28:34,234] [app.anomaly_detector.service_controller] [ERROR] Consumer failed with unhandled exception : 'NoneType' object is not iterable
Solution: There is a two-step process to resume the flow:
-
Clear the windowed logs topic content. This step is required to clear any old data that is not yet consumed by the log anomaly detector. If you do not clear this data, when LAD is ready to make predictions, it will try to generate events for old data showing a lag. So, you need to disable the data flow, and clear data from windowed logs.
-
Disable data flow from data integration
-
Clear the windowed logs topic using following command
oc edit kt cp4waiops-cartridge-windowed-logs-1000-1000 -
Set
retention.msto1. -
Wait for a few minutes, and once the data is cleared and no more windows are generated, reset the 'retention.ms' to the original value
-
Enable the data flow in the data integration
-
-
Once the data is cleared using the above steps, restart the anomaly pod logs with the following command:
oc get po | grep anomaly oc delete pod <anomaly-pod>
Postgres database crashes
After a failure of one of the Postgres replicas, a new replica is unable to come up, and the failed replica remains.
Solution: Delete the failed replica and its persistent volume claim (PVC) if the replica failure is due to a synchronization problem.
-
Run the following command to find all Postgres instances that are not running:
export PROJECT_CP4AIOPS=<project> oc get pod -n $PROJECT_CP4AIOPS -l "pg.ibm.com/podRole=instance" | grep -vE "Running"Where
<project>is the project (namespace) where Concert Operate is deployed.Example output where
zen-metastore-1has failed:oc get pod -n $PROJECT_CP4AIOPS -l "pg.ibm.com/podRole=instance" | grep -vE "Running" NAMESPACE NAME READY STATUS RESTARTS AGE aiops zen-metastore-1 0/1 CrashLoopBackOff 904 (36s ago) 7d1h -
Examine the logs for the failed replica to see whether the non-primary replicas are not synchronized with the WAL logs in the primary replica.
oc logs -n "${PROJECT_CP4AIOPS}" <failed_replica>Where
<failed_replica>is the failed replica returned in step 1, for examplezen-metastore-1.Look for entries similar to the following examples:
pg_rewind: servers diverged at WAL location 5/6606A4E0 on timeline 1 pg_rewind: error: could not open file \"/var/lib/postgresql/data/pgdata/pg_wal/000000010000000500000068\": No such file or directory pg_rewind: fatal: could not read WAL record at 5/68000028{"level":"info","ts":1677595436.8679705,"logger":"postgres","msg":"record","logging_pod":"ibm-cp-aiops-edb-postgres-1","record":{"log_time":"2023-02-28 14:43:56.867 UTC","user_name":"streaming_replica","process_id":"90988","connection_from":"10.254.24.22:32852","session_id":"63fe132c.1636c","session_line_num":"1","command_tag":"idle","session_start_time":"2023-02-28 14:43:56 UTC","virtual_transaction_id":"13/0","transaction_id":"0","error_severity":"ERROR","sql_state_code":"58P01","message":"requested WAL segment 0000000300000002000000A0 has already been removed","query":"START_REPLICATION 2/A0000000 TIMELINE 2","application_name":"ibm-cp-aiops-edb-postgres-2","backend_type":"walsender"}} -
Check that the failed replica is not the primary replica.
Run the following command, and check that the failed replica identified in step 1 is not listed in the PRIMARY column for any of the Postgres clusters.
oc get clusters.pg.ibm.com -n $PROJECT_CP4AIOPSExample output where
zen-metastore-1is not the primary:oc get clusters.pg.ibm.com -n $PROJECT_CP4AIOPS NAME AGE INSTANCES READY STATUS PRIMARY aiops-ir-analytics-postgres 125m 1 1 Cluster in healthy state aiops-ir-analytics-postgres-1 aiops-ir-core-postgres 125m 1 1 Cluster in healthy state aiops-ir-core-postgres-1 aiops-orchestrator-postgres 131m 1 1 Cluster in healthy state aiops-orchestrator-postgres-1 aiops-topology-postgres 129m 1 1 Cluster in healthy state aiops-topology-postgres-1 common-service-db 132m 2 2 Cluster in healthy state common-service-db-1 zen-metastore 117m 2 1 Instance Status Extraction Error: HTTP communication issue zen-metastore-2 -
Destroy the failed replica.
If it is not already installed, run the following command to install thekubectlplug-in.curl -fsSL https://raw.githubusercontent.com/IBM/kubectl-ibmpg/main/install.sh | shRun the following command if the failed replica is not the primary and the logs for the failed replica showed errors similar to the example output in step 2. Otherwise, contact IBM Support.
oc ibmpg destroy <failed_replica> <replica ID> -n "${PROJECT_CP4AIOPS}"Where
<failed_replica>is the failed replica returned in step 1, for examplezen-metastore-1.Example command:
oc ibmpg destroy zen-metastore-1 -n "${PROJECT_CP4AIOPS}"
Add additional network policies
Add additional network policies to allow network traffic from the kube-system namespace to the cluster.
Note that such traffic is prohibited by default in order to isolate the clusters for security. When not required, the policies should be deleted in order to restrict traffic according to the normal rules.
Use the kubectl plugin to check the aiops-ir-analytics-postgres or aiops-ir-core-postgres clusters for errors, and then add an additional network policy for each.
-
If it is not already installed, install the
kubectlplug-in.Run the following command:
curl -fsSL https://raw.githubusercontent.com/IBM/kubectl-ibmpg/main/install.sh | sh -
Check the postgres status for
aiops-ir-core-postgres.oc ibmpg status aiops-ir-core-postgresA typical error message would contain text such as, for example,
the server is currently unable to handle the request (get pods https:aiops-ir-core-postgres-1:8000). -
Add a network policy using the following commands, replacing the
NAMESPACEplaceholder with the namespace in which your instance of IBM Concert Operate is running.For cluster aiops-ir-core-postgres:
apiVersion: networking.k8s.io/v1 kind: NetworkPolicy metadata: name: aiops-ir-core-postgres-kube-system namespace: NAMESPACE spec: ingress: - fromEndpoints: - matchLabels: io.kubernetes.pod.namespace: kube-system toPorts: - ports: - port: "8000" protocol: TCP podSelector: matchLabels: pg.ibm.com/cluster: aiops-ir-core-postgres policyTypes: - IngressFor cluster aiops-ir-analytics-postgres:
apiVersion: networking.k8s.io/v1 kind: NetworkPolicy metadata: name: aiops-ir-analytics-postgres-kube-system namespace: NAMESPACE spec: ingress: - fromEndpoints: - matchLabels: io.kubernetes.pod.namespace: kube-system toPorts: - ports: - port: "8000" protocol: TCP podSelector: matchLabels: pg.ibm.com/cluster: aiops-ir-analytics-postgres policyTypes: - Ingress
CouchDB pod reports a permission error when draining a node
The CouchDB pod's log reports a permission error similar to the following example:
[error] 2022-12-02T01:24:20.726024Z couchdb@c-example-couchdbcluster-m-0.c-example-couchdbcluster-m <0.13913.7> -------- Could not open file /data/db/shards/00000000-7fffffff/_users.1669801403.couch: permission denied
After a restore into a new namespace, similar symptoms can be seen. For more information about the solution for this scenario, see Restore fails with CouchDB pod in CrashLoopBackOff.
The way that Kubernetes mounts persistent volumes can cause this problem.
Solution: Delete the CouchDB pods that are not ready. When the replacement pods are created, the persistent volumes are mounted correctly.
-
Find the names of your
couchdbpods.oc get pods -l app.kubernetes.io/name=IssueResolutionCoreCouchDB -n <namespace>Where
<namespace>is the namespace that Concert Operate is deployed in. -
Delete all the
couchdbpods that the previous step showed as not ready.oc delete pod <couchdb-pod> -n <namespace>Where:
-
<namespace>is the namespace that Concert Operate is deployed in. -
<couchdb-pod>is a couchdb pod that is not ready.
-
Kafka broker pods are unhealthy because Kafka PVCs are full
When the Kafka broker pods are unhealthy due to insufficient storage, the Events operator is unable to expand the PVCs.
Solution: Check whether the Kafka broker pods have storage available, and expand the storage available to Kafka if they do not.
-
Check the state of the Kafka broker pods:
export AIOPS_NAMESPACE=<project> oc get pod -n "${AIOPS_NAMESPACE}" -l app.kubernetes.io/name=kafkaWhere
<project>is the namespace (project) that IBM Concert Operate is deployed.Example output:
iaf-system-kafka-0 1/1 Running 0 5d21h iaf-system-kafka-1 0/1 CreateContainerError 0 95m iaf-system-kafka-2 0/1 CreateContainerError 1 (2d18h ago) 5d21h -
If the broker pods are in a CreateContainerError state, check the logs for errors:
oc logs -n "${AIOPS_NAMESPACE}" iaf-system-kafka-1Example output:
Warning Failed 5m40s (x13 over 89m) kubelet Error: relabel failed /var/lib/kubelet/pods/37e6519a-b372-429d-b5f5-9aaf5b12fd37/volumes/kubernetes.io~csi/pvc-ff53e966-fee1-41e2-9448-b0eeac2e49c6/mount: lsetxattr /var/lib/kubelet/pods/37e6519a-b372-429d-b5f5-9aaf5b12fd37/volumes/kubernetes.io~csi/pvc-ff53e966-fee1-41e2-9448-b0eeac2e49c6/mount/kafka-log1/.lock: no space left on deviceDo not proceed with the rest of these troubleshooting steps if the error is not related to disk space.
-
Adjust the Kafka storage size in the CustomResource:
oc patch kafka/iaf-system -n "${AIOPS_NAMESPACE}" --type merge -p '{"spec":{"kafka":{"storage":{"size":"<size>Gi"}}}}'Where
<size>is the increased storage size that you require for the Kafka PVCs. -
Increase the Kafka PVCs:
for p in $(oc get pvc -l app.kubernetes.io/name=kafka -n "${AIOPS_NAMESPACE}" -o jsonpath='{.items[*].metadata.name}'); do oc patch pvc $p -n "${AIOPS_NAMESPACE}" --type merge -p '{"spec":{"resources":{"requests":{"storage":"<size>Gi"}}}}' ; doneWhere
<size>is the same value that was used in the previous command.Note : The time required for this command to run depends on the number of PVCs and the amount of data that is stored in each one.
aimanager is not in Ready state after cluster restart, integrations are missing from the console
After a cluster restart, aimanager does not have a ready state when queried with oc get aimanager -o yaml, and some previously configured integrations such as log integrations, ChatOps, runbooks, kafka, or PagerDuty are missing from the console. This problem can be caused by a timing issue where the aimanager pod attempts to start before the Postgres database is ready.
Solution
-
Run the following command to check if the
aimanager-aio-controllerpod is running.oc get pods | grep aimanager-aio-controllerExample output for a pod that is running correctly:
aimanager-aio-controller-67d6857cbb-n9skp 1/1 Running 0 7d3h -
If the
aimanager-aio-controlleris not running, then check the pod's logs for errors.oc logs <aimanager-aio-controller-pod-name>Where
<aimanager-aio-controller-pod-name>is the name of theaimanager-aio-controller-pod-namefrom step 1.Example error output:
2024-12-03 04:51:14 ERROR ServletListener:801 - Failed to get edb connection from SQL query java.sql.SQLException: Cannot create PoolableConnectionFactory (Connection to ibm-cp-aiops-edb-postgres-rw:5432 refused. Check that the hostname and port are correct and that the postmaster is accepting TCP/IP connections.) at org.apache.commons.dbcp2.BasicDataSource.createPoolableConnectionFactory(BasicDataSource.java:653) ~[commons-dbcp2-2.9.0.jar:2.9.0]If you have an error similar to the preceding example error output, then note down the name of the
aimanager-aio-controllerpod and use the following steps to resolve the problem. -
Check that the
Postgresdatabase is up and running, and do not continue until it is.oc get pods | grep postgresExample output if PostgreSQL is running:
aiops-installation-edb-postgres-1 1/1 Running 0 6d2h -
Run the following command to restart the
aimanager-aio-controllerpod.oc delete pod <aimanager-aio-controller-pod-name>Where
<aimanager-aio-controller-pod-name>is the name of theaimanager-aio-controller-pod-namefrom step 1.
Deployment on Linux: unable to access Concert Operate console
The pods in the IBM Concert Operate cluster are in a Running state, but the user interface is not accessible.
Solution: When you installed IBM Concert Operate, you used an existing load balancer or configured a new one of your choice. For more information, see Load balancing. The load balancer is the entry point for accessing deployments of IBM Concert Operate on Linux. Problems with your load balancer can prevent access to the Concert Operate console. Check the load balancer status and logs for problems, and restart the load balancer if it is not running.
Deployment on Linux: API server is down after a cluster restart
After restarting the Linux cluster, all requests to the Kubernetes API server such as oc get pod and aiopsctl status fail, as in the following example output:
$ oc get pod -n aiops
E0903 08:55:39.586644 1287 memcache.go:265] couldn't get current server API group list: the server is currently unable to handle the request
E0903 08:55:39.589138 1287 memcache.go:265] couldn't get current server API group list: the server is currently unable to handle the request
E0903 08:55:39.592133 1287 memcache.go:265] couldn't get current server API group list: the server is currently unable to handle the request
E0903 08:55:39.594610 1287 memcache.go:265] couldn't get current server API group list: the server is currently unable to handle the request
E0903 08:55:39.597172 1287 memcache.go:265] couldn't get current server API group list: the server is currently unable to handle the request
Error from server (ServiceUnavailable): the server is currently unable to handle the request
$ aiopsctl status
o- [03 Sep 24 08:58 PDT] Getting cluster status
[ERROR] Failed to get cluster status apiserver not ready
Solution: Run the following steps to confirm that the problem is due to the Kubernetes API server being down, and to rectify this.
-
Check the control plane nodes.
-
On each control plane node, run the following command to assess the platform state. Affected deployments have a state of activating.
systemctl status k3s.serviceExample output for a deployment affected by this issue:
$ systemctl status k3s.service k3s.service - Lightweight Kubernetes Loaded: loaded (/etc/systemd/system/k3s.service; enabled; vendor preset: enabled) Active: activating (start) since Mon 2024-09-09 05:10:30 PDT; 6s ago -
On each control plane node, run the following command to check the platform logs:
journalctl -xeu k3s.serviceExample output for a deployment affected by this issue:
Sep 09 05:11:28 control-plane-2.acme.com k3s[50701]: {"level":"warn","ts":"2024-09-09T05:11:28.735466-0700","caller":"rafthttp/http.go:413","msg":"failed to find remote peer in cluster","local-member-id":"8c10b42dfa5ffc54","rem> Sep 09 05:11:28 control-plane-2.acme.com k3s[50701]: {"level":"warn","ts":"2024-09-09T05:11:28.828015-0700","caller":"rafthttp/http.go:413","msg":"failed to find remote peer in cluster","local-member-id":"8c10b42dfa5ffc54","rem> Sep 09 05:11:28 control-plane-2.acme.com k3s[50701]: {"level":"warn","ts":"2024-09-09T05:11:28.834347-0700","caller":"rafthttp/http.go:145","msg":"failed to process Raft message","local-member-id":"8c10b42dfa5ffc54","error":"ra>
If your deployment is affected, continue with the following steps.
-
-
Stop
k3son the primary control plane node:systemctl stop k3sThe primary control plane is the first control plane attached to the cluster. This is the value of CONTROL_PLANE_NODE in your aiops_var.sh environment variables file. The non-primary control plane nodes are the nodes in the ADDITIONAL_CONTROL_PLANE_NODES array in your aiops_var.sh environment variables file.
-
Run a cluster reset on the primary control plane node to soft reset
etcd.k3s server --cluster-reset -
Stop k3s on the non-primary control plane nodes.
systemctl stop k3s -
Back up the
etcddata directory on each non-primary control plane node.cp -r <data_directory> <backup_directory>Where
-
<data_directory>is:/var/lib/rancher/k3s/server/dbif you are not using a custom platform directory, orPLATFORM_STORAGE_PATHin aiops_var.sh if you are using a custom platform directory -
<backup_directory>is a directory where you want to store the backup.
-
-
Delete the
etcddata directory on each non-primary control plane node.Do not delete the platform data directory on the primary control plane node.
rm -rf <data_directory>Where
<data_directory>is/var/lib/rancher/k3s/server/dbif you are not using a custom platform directory, orPLATFORM_STORAGE_PATHin aiops_var.sh if you are using a custom platform directory. -
Run the following command on the primary control plane node, and then on each of the non-primary control plane nodes:
systemctl start k3s
After restarting the Red Hat OpenShift cluster one or more Cassandra pods are missing or in Pending state.
When the Red Hat OpenShift cluster restarted, the Cassandra pods might have tried to restart before Red Hat OpenShift was fully operational. Run the following command, and then check if the output has a warning that the Cassandra pods were not created because they could not be validated against security context constraints.
oc describe statefulset aiops-topology-cassandra
Example output:
Warning FailedCreate 32s (x61 over 5h4m) statefulset-controller create Pod aiops-topology-cassandra-0 in StatefulSet aiops-topology-cassandra failed error: pods "aiops-topology-cassandra-0" is forbidden: unable to validate against any security context constraint: [provider "anyuid": Forbidden: not usable by user or serviceaccount, provider restricted-v2: .spec.securityContext.fsGroup: Invalid value: []int64{1001}: 1001 is not an allowed group, provider restricted-v2: .initContainers[0].runAsUser: Invalid value: 1001: must be in the ranges: [1000750000, 1000759999], provider restricted-v2: .containers[0].runAsUser: Invalid value: 1001: must be in the ranges: [1000750000, 1000759999], provider "restricted": Forbidden: not usable by user or serviceaccount, provider "nonroot-v2": Forbidden: not usable by user or serviceaccount, provider "nonroot": Forbidden: not usable by user or serviceaccount, ...
Solution: Wait for the pods in the Red Hat OpenShift namespaces to have a status of Running, especially the pods in the openshift-apiserver namespace, and then restart the ASM operator.
-
Run the following command to check the status of the pods in the Red Hat OpenShift namespaces:
oc get po -A | grep ^openshift-Example output:
oc get po -A | grep ^openshift- openshift-apiserver-operator openshift-apiserver-operator-7bb4b546d-wq86x 1/1 Running 0 5h4m openshift-apiserver apiserver-9c76484c6-hbxdr 2/2 Running 0 5h9m openshift-apiserver apiserver-9c76484c6-hprjc 2/2 Running 0 5h15m openshift-apiserver apiserver-9c76484c6-wbg9h 2/2 Running 0 5h4m <...> -
Restart the ASM operator.
When all of the pods that are returned by the previous command are
Running, run the following command to restart theasm-operator:oc delete po -l name=asm-operator -n <namespace>Where
<namespace>is the namespace that IBM Concert Operate is deployed in.
After restarting the Red Hat OpenShift cluster one or more couchdb pods are missing or in Pending state.
When the Red Hat OpenShift cluster restarted, the couchdb pods might have tried to restart before Red Hat OpenShift was fully operational. Run the following command, and then check if the output has a warning that the couchdb pods were not created because they could not be validated against security context constraints.
oc describe statefulset c-example-couchdbcluster-m
Example output:
Warning FailedCreate 3m57s (x750 over 3d4h) statefulset-controller create Pod c-example-couchdbcluster-m-0 in StatefulSet c-example-couchdbcluster-m failed error: pods "c-example-couchdbcluster-m-0" is forbidden: unable to validate against any security context constraint: [provider "anyuid": Forbidden: not usable by user or serviceaccount, provider restricted-v2: .spec.securityContext.fsGroup: Invalid value: []int64{0}: 0 is not an allowed group, provider restricted-v2: .initContainers[0].runAsUser: Invalid value: 1000: must be in the ranges: [1000750000, 1000759999], provider restricted-v2: .containers[0].runAsUser: Invalid value: 1000: must be in the ranges: [1000750000, 1000759999], provider restricted-v2: .containers[1].runAsUser: Invalid value: 1000: must be in the ranges: [1000750000, 1000759999], provider "restricted": Forbidden: not usable by user or serviceaccount, provider "nonroot-v2": Forbidden: not usable by user or serviceaccount, provider "nonroot"
Solution: Wait for the pods in the Red Hat OpenShift namespaces to have a status of Running, especially the pods in the openshift-apiserver namespace, and then restart the issue resolution core operator.
-
Run the following command to check the status of the pods in the Red Hat OpenShift namespaces:
oc get po -A | grep ^openshift-Example output:
oc get po -A | grep ^openshift- openshift-apiserver-operator openshift-apiserver-operator-7bb4b546d-wq86x 1/1 Running 0 5h4m openshift-apiserver apiserver-9c76484c6-hbxdr 2/2 Running 0 5h9m openshift-apiserver apiserver-9c76484c6-hprjc 2/2 Running 0 5h15m openshift-apiserver apiserver-9c76484c6-wbg9h 2/2 Running 0 5h4m <...> -
Restart the issue resolution core operator.
When all of the pods that are returned by the previous command are
Running, run the following command to restart their-core-operator:oc delete po -l app.kubernetes.io/name=ir-core-operator -n <namespace>Where
<namespace>is the namespace that IBM Concert Operate is deployed in.
OpenSearch pods restarting continuously due to insufficient entropy availability (deployments on Linux only)
OpenSearch has high requirements for entropy availability at /dev/random. If there is insufficient entropy availability, the OpenSearch pods might restart continuously as the process hangs during startup.
Symptoms:
-
OpenSearch pods are restarting continuously:
oc get pod -l cluster.opensearch.cloudpackopen.ibm.com=aiops-opensearch -n aiopsExample output:
aiops-opensearch-all-000 0/1 Running 1 (114s ago) 4m54s aiops-opensearch-all-001 0/1 Running 12 (6m10s ago) 51m aiops-opensearch-all-002 0/1 Running 12 (6m17s ago) 51m -
The OpenSearch logs show the process hanging at startup:
oc logs aiops-opensearch-all-000 -pExample output:
Defaulted container "opensearch" out of: opensearch, opensearch-setup (init), opensearch-plugin-load-0 (init) ... Starting Opensearch... warning: no-jdk distributions that do not bundle a JDK are deprecated and will be removed in a future release warning: no-jdk distributions that do not bundle a JDK are deprecated and will be removed in a future release WARNING: Using incubator modules: jdk.incubator.vector WARNING: Unknown module: org.apache.arrow.memory.core specified to --add-opens
Solution: Verify and increase entropy availability on your cluster nodes.
-
Check the current entropy availability on each node. The result should be greater than 1000.
cat /proc/sys/kernel/random/entropy_availExample output showing insufficient entropy:
9 -
If entropy availability is too low, install and enable
rng-tools. For example, on Red Hat Enterprise Linux:yum install -y rng-tools systemctl start rngd systemctl enable rngd -
Verify that entropy availability is now greater than 1000.
cat /proc/sys/kernel/random/entropy_availExample output showing sufficient entropy:
3906The Opensearch pods recover automatically.