IBM Support

InfoSphere BigInsights high availability configurations

Product Documentation


Abstract

InfoSphere BigInsights supports three high availability configurations that differ by the number of dedicated high availability nodes.

Content

High availability is important for clusters that are larger than a single rack. InfoSphere BigInsights offers IBM specific and open source high availability options. InfoSphere BigInsights supports three high availability configurations that differ by the number of dedicated high availability nodes.

The following chart shows the three high availability configurations available for InfoSphere BigInsights.

Tab navigation

When using the General Parallel File System (GPFS) as your distributed file system, adaptive MapReduce is used to manage high availability. The GPFS cluster requires three system pool machines that share the same disks for the high availability manager component, which provide high availability for adaptive MapReduce. At a minimum, the ZooKeeper service must have a dedicated local disk to provide high I/O transactions of ZooKeeper snapshots.
Another high availability topology uses HDFS as the distributed file system and a Network File System (NFS) for high availability. The high availability manager and the NameNode service share resources on the same machine, and these components are replicated across each rack in the cluster.
In a typical Hadoop cluster, high availability becomes important when the cluster size grows beyond a single rack. In these topologies, one HTTPFS server should exist on each rack for bandwidth load balancing, and three ZooKeeper servers should also exist on each management rack to support high availability. HBase master nodes are spread over three racks of machines for redundancy. The NameNode is spread across two racks for failover, and the JobTracker exists on its own machine. Data is spread across 80 data nodes to manage data in a large cluster.

[{"Product":{"code":"SSCRJT","label":"IBM Db2 Big SQL"},"Business Unit":{"code":"BU059","label":"IBM Software w\/o TPS"},"Component":"--","Platform":[{"code":"PF016","label":"Linux"}],"Version":"2.1.2","Edition":"","Line of Business":{"code":"LOB10","label":"Data and AI"}}]

Document Information

Modified date:
18 July 2020

UID

swg27041635