Preparing your cluster for fault tolerance
High availability configuration
For high availability and fault tolerance to be effective, it is critical that you set the appropriate number of replicas for the management service and for Business Performance Center. For Kafka and OpenSearch, the replicas are handled by the IBM Cloud Pak® for Business Automation environment. Refer to Configuration for high availability and fault tolerance.
Failure of Apache Flink job manager and task managers
When they are deployed, the job manager and task managers are configured to restart their failed pod, and jobs are configured for checkpointing.
Job failure
If a recoverable error occurs, such as temporary network outage that prevents connection to Kafka or OpenSearch, Flink jobs automatically restart. When a job cannot restart, see Restarting from a checkpoint or savepoint.
Known issues
To learn how to resolve known issues, see Troubleshooting IBM Business Automation Insights on Kubernetes.