There are many things worth learning about high-end router equipment. Here we mainly introduce equipment-Level Reliability Technology, including power fan redundancy and uninterrupted forwarding technology NSF/GR. With the rapid development of IP technology, various value-added services have been widely used on the Internet. Emerging important telecommunication services such as NGN/3G, IPTV streaming media, major customer leased lines, and VPN interconnection have high requirements on the reliability of IP Telecom networks. The reliability of telecommunication services for IP networks mainly includes three layers: device reliability, link reliability, and network reliability. In the bearer network, the availability requirement of network equipment reaches 99.999%, which is roughly equivalent to less than 5 minutes of downtime maintenance due to various possible reasons during the continuous operation of the equipment for one year. High reliability is the basic requirement of telecommunication equipment and the basic starting point for telecom operators to build networks.
Equipment-level reliability technologies mainly include:
Hot swapping technology refers to directly plugging components or boards without affecting the business of other components or boards when the device is not shut down and running. Hot swapping features include: adding or removing a single board to the frame without affecting the use of the Board; online board replacement, that is, removing the board for a new board or the original Single Board re-insertion, the new board can inherit the original configuration without affecting the work of other boards. For distributed devices, when you add or remove a board, the FIBForwarding Information Base table can be synchronized to the board. All components of the Huawei NE series high-end router equipment support hot swapping, including the main control board, switching network board, power supply, Fan and various business boards. With the hot swapping function, you can maintain and update components without affecting services, expand more services, add more users, and provide more functions.
Power fan Redundancy
Power supply is the foundation of equipment operation assurance. Once a power supply problem occurs, the device cannot start normally, so power redundancy is necessary. Power supply redundancy includes power supply input redundancy and device Power Supply Module Redundancy. To ensure the stability of power input, high-end router devices generally provide dual or multi-channel power input. When a power input fails, it can automatically switch to another power input without affecting the normal operation of the device. In addition, the high-end router device also uses multiple power supply modules for power supply, and adopts the N + 1 backup mode. One power supply module works with N other power supply modules at the same time and provides backup for them, when a power supply module fails, other power supplies immediately share the load of the faulty power supply, so as to ensure that sufficient power is always provided to ensure normal operation of the equipment. As an important means of heat dissipation, the fan has a direct impact on the stable operation of the equipment. If the fan fails and the heat dissipation fails in time, the device may experience high temperature and high heat, and the chip and board may burn down. Fan redundancy is also very important. High-end router devices generally provide multiple fan boxes, which can be replaced online without affecting the device functions.
Master Control Redundancy
The main control board MPUMain Processing Unit) is the core of the entire router and undertakes global functions such as route Processing, resource management, status monitoring, and network management proxy of the entire system. Generally, it also integrates third-level clock, CFCompact Flash) card and other functional modules. Some devices even include the MPU of the switching network module to provide the switching plane for the entire high-end router device. The Master Control Board redundancy means that clock redundancy, storage device redundancy, and switching network redundancy are also realized. The Master Control Redundancy technology is introduced here.
If the device only has a single master, if the master Board fails, you need to load the image file, initialize the configuration, and re-register the Business Board to restart the master board, and then recreate the control plane and forwarding plane table items, the entire process takes several minutes. This time is intolerable for the telecom network, especially for nodes with single points of failure in the network, because the business will be completely interrupted in this process, this will cause huge losses. Therefore, in order to shorten the master restart time and reduce the loss caused by service interruption, high-end router devices must adopt master redundancy technology. Master Control Redundancy means that the device provides two master control boards for mutual backup. One Master is in the active state, the other is in the active state, and the other is in the STANDBY state. During the operation of the main control board, all static configuration information and some dynamic information are backed up to the backup control board, so that the backup control board has the same configuration information as the master control board. When the Main Control Board fails due to hardware or software failure, the Standby Control Board takes over the failure of the main control board, restart the control plane and management plane, to ensure that the router can recover to normal in a short period of time. You can use either the hardware heartbeat or the IPC channel or other methods to detect switching between the Master and the Slave Master.
Compared with single-master, dual-master has much better convergence performance. In the case of dual-master, Slave has completed loading and configuration initialization of the image file in advance, and the Business Board does not need to be re-registered during master-Slave switchover, the interfaces on the second and third layers do not show up/down either. In addition, the Slave has also backed up a forwarding table item, which can take on the forwarding task immediately to avoid service interruption to a certain extent.
However, because the new Master node does not take part in the control plane processing before the Master/Slave switchover, you need to re-negotiate sessions with the neighbors after the switchover. Therefore, although the complete forwarding table item is saved, however, this can only prevent some traffic from being interrupted. For example, the second-layer service and the traffic sent from the device will not be interrupted. In addition, if Static Routing or static LSP is configured with the neighbor, the traffic will not be interrupted. However, if there is a dynamic routing protocol or a dynamic Label Distribution Protocol between the two neighbors, the traffic between the two neighbors will be interrupted because the plane session is reset, the control plane of the neighbor is re-computed, and the path that it considers appropriate is selected. Taking the OSPF protocol as an example, the new Master does not have the RID of the original neighbor in the Hello message, which will cause the neighbor to reset the OSPF session status, the LSA associated with the high-end router device that has been switched is deleted, resulting in route re-calculation. If there are other optional paths, the traffic bypasses the device with master-slave switchover. If there is no optional path, you need to wait for OSPF to converge again. Before convergence, the neighbor will not send traffic to the high-end router device with master-slave switchover.
Uninterrupted forwarding technology NSF/GR
From the above analysis, we can see that when a Router performs a master-slave switchover, it will fluctuate with its neighbors at the routing protocol layer. This kind of neighbor relationship fluctuation will eventually lead to the appearance of Route fluctuation, so that the winner's slave switchover router may encounter a routing black hole within a period of time or cause the neighbor to bypass the data service, in this case, the business may be temporarily interrupted. Continuous Forwarding NSFNone Stop Forwarding is an important high-reliability technology. It can ensure that data Forwarding works continuously when there is a fault at the control layer of the router, such as a fault restart or route fluctuation, thus, the network traffic is almost unaffected. The router must have a distributed architecture, data forwarding and control separation, and support dual-controller design. In the case of master-slave switchover, the slave board must be able to successfully Save the IP/MPLS Forwarding Table entry forwarding plane ).
Secondly, you may need to save the Protocol's status control plane as needed ). For complicated protocols such as OSPF, IS-IS, BGP, and LDP, the complex state of the control plane IS completely backed up, which IS too costly or impossible to implement. On the contrary, by extending the current protocol to a certain extent while keeping forward compatibility as far as possible, you can simply use partial backup or no backup at all) protocol status, with the help of the neighboring high-end router devices, the session connection of the control plane is not reset and the forwarding is not interrupted when the master/Slave switchover occurs.
These technologies that do not reset the control layer are collectively referred to as the Graceful Restart extension of the routing protocol, or GR for short. The GR technique prevents the neighbor relationship from fluctuating flap when the master-slave switchover is restarted. Once the master-slave switchover is restarted, the router restarts as soon as possible to synchronize the route information with the neighbor router, then update the local route information. Currently, the GR implementation usually requires the assistance of the neighbor router Helper), requiring Helper to be able to perceive the occurrence of GR in the neighbor and assist the neighbor to complete the GR, which also puts forward high requirements for the Helper in the network. Currently, the GR-capable routing protocols include OSPF, IS-IS, BGP, and LDP. Although each protocol has its own unique implementation, its basic principles are similar.