The Cisco HyperFlex Information Platform (HXDP) is a distributed hyperconverged infrastructure system that has been constructed from inception to deal with particular person part failures throughout the spectrum of {hardware} parts with out interruption in companies. Consequently, the system is very accessible and able to in depth failure dealing with. On this quick dialogue, we’ll outline the kinds of failures, briefly clarify why distributed techniques are the popular system mannequin to deal with these, how knowledge redundancy impacts availability, and what’s concerned in a web based knowledge rebuild within the occasion of the lack of knowledge parts.
You will need to notice that HX is available in 4 distinct varieties. They’re Customary Information Middle, Information Middle@ No-Material Interconnect (DC No-FI), Stretched Cluster, and Edge clusters. Listed here are the important thing variations:
Customary DC
- Has Material Interconnects (FI)
- May be scaled to very massive techniques
- Designed for infrastructure and VDI in enterprise environments and knowledge facilities
DC No-FI
- Just like customary DC HX however with out FIs
- Has scale limits
- Lowered configuration calls for
- Designed for infrastructure and VDI in enterprise environments and knowledge facilities
Edge Cluster
- Utilized in ROBO deployments
- Is available in numerous node counts from 2 nodes to eight nodes
- Designed for smaller environments the place preserving the functions or infrastructure near the customers is required
- No Material Interconnects – redundant switches as an alternative
Stretched Cluster
- Has 2 units of FIs
- Used for extremely accessible DR/BC deployments with geographically synchronous redundancy
- Deployed for each infrastructure and software VMs with extraordinarily low outage tolerance
The HX node itself consists of the software program parts required to create the storage infrastructure for the system’s hypervisor. That is completed through the HX Information Platform (HXDP) that’s deployed at set up on the node. The HX Information Platform makes use of PCI pass-through which removes storage ({hardware}) operations from the hypervisor making the system extremely performant. The HX nodes use particular plug-ins for VMware known as VIBs which can be used for redirection of NFS datastore visitors to the proper distributed useful resource, and for {hardware} offload of complicated operations like snapshots and cloning.

These nodes are integrated right into a distributed Zookeeper based mostly cluster as proven under. ZooKeeper is actually a centralized service for distributed techniques to a hierarchical key-value retailer. It’s used to supply a distributed configuration service, synchronization service, and naming registry for big distributed techniques.

To being, let’s take a look at all of the doable the kinds of failures that may occur and what they imply to availability. Then we are able to focus on how HX handles these failures.
- Node loss. There are numerous the reason why a node might go down. Motherboard, rack energy failure,
- Disk loss. Information drives and cache drives.
- Lack of community interface (NIC) playing cards or ports. Multi-port VIC and help for add on NICs.
- Material Interconnect (FI) No all HX techniques have FIs.
- Energy provide
- Upstream connectivity interruption
Node Community Connectivity (NIC) Failure
Every node is redundantly linked to both the FI pair or the change, relying on which deployment structure you might have chosen. The digital NICs (vNICs) on the VIC in every node are in an energetic standby mode and break up between the 2 FIs or upstream switches. The bodily ports on the VIC are unfold between every upstream machine as properly and you might have extra VICs for additional redundancy if wanted.

Let’s observe up with a easy resiliency resolution earlier than analyzing want and disk failures. A conventional Cisco HyperFlex single-cluster deployment consists of HX-Sequence nodes in Cisco UCS linked to one another and the upstream change by a pair of material interconnects. A material interconnect pair might embody a number of clusters.
On this situation, the material interconnects are in a redundant active-passive major pair. Within the occasion of an FI failure, the associate will take over. This is identical for upstream change pairs whether or not they’re immediately linked to the VICs or by the FIs as proven above. Energy provides, in fact, are in redundant pairs within the system chassis.
Cluster State with Variety of Failed Nodes and Disks
How the variety of node failures impacts the storage cluster depends upon:
- Variety of nodes within the cluster—As a result of nature of Zookeeper, the response by the storage cluster is completely different for clusters with 3 to 4 nodes and 5 or larger nodes.
- Information Replication Issue—Set throughout HX Information Platform set up and can’t be modified. The choices are 2 or 3 redundant replicas of your knowledge throughout the storage cluster.
- Entry Coverage—May be modified from the default setting after the storage cluster is created. The choices are strict for shielding towards knowledge loss, or lenient, to help longer storage cluster availability.
- The kind
The desk under exhibits how the storage cluster performance modifications with the listed variety of simultaneous node failures in a cluster with 5 or extra nodes working HX 4.5(x) or larger. The case with 3 or 4 nodes has particular issues and you may examine the admin information for this info or discuss to your Cisco consultant.
The identical desk can be utilized with the variety of nodes which have a number of failed disks. Utilizing the desk for disks, notice that the node itself has not failed however disk(s) inside the node have failed. For instance: 2 signifies that there are 2 nodes that every have no less than one failed disk.
There are two doable kinds of disks on the servers: SSDs and HDDs. After we discuss a number of disk failures within the desk under, it’s referring to the disks used for storage capability. For instance: If a cache SSD fails on one node and a capability SSD or HDD fails on one other node the storage cluster stays extremely accessible, even with an Entry Coverage strict setting.
The desk under lists the worst-case situation with the listed variety of failed disks. This is applicable to any storage cluster 3 or extra nodes. For instance: A 3 node cluster with Replication Issue 3, whereas self-healing is in progress, solely shuts down if there’s a complete of three simultaneous disk failures on 3 separate nodes.
3+ Node Cluster with Variety of Nodes with Failed Disks
A storage cluster therapeutic timeout is the size of time the cluster waits earlier than robotically therapeutic. If a disk fails, the therapeutic timeout is 1 minute. If a node fails, the therapeutic timeout is 2 hours. A node failure timeout takes precedence if a disk and a node fail at identical time or if a disk fails after node failure, however earlier than the therapeutic is completed.
You probably have deployed an HX Stretched Cluster, the efficient replication issue is 4 since every geographically separated location has an area RF 2 for web site resilience. The tolerated failure situations for a Stretched Cluster are out of scope for this weblog, however all the small print are lined in my white paper right here.
In Conclusion
Cisco HyperFlex techniques include all of the redundant options one would possibly anticipate, like failover parts. Nonetheless, additionally they include replication components for the information as defined above that supply redundancy and resilience for a number of node and disk failure. These are necessities for correctly designed enterprise deployments, and all components are addressed by HX.
Share:



