Reprint: Original link
Golden Partner for deep learning: The GPU is reinventing computing
Ofweek Electronic Engineering Network with the deepening of neural networks and deep learning-in particular, speech recognition and natural language processing, image and pattern recognition, text and data analysis, and other complex areas-researchers are constantly looking for new and better ways to extend and expand computing power.
For decades, the gold standard in this area has been a high performance computing (HCP) cluster, which solves a lot of processing power problems, albeit at a high cost. But this approach has helped drive progress in a number of areas, including weather forecasts, financial services, and energy exploration.
In 2012, however, a new approach emerged. Researchers at the University of Illinois have previously studied the possibility of using GPUs in desktop supercomputers to speed up processing tasks (like rebuilding), and now a group of computer scientists and engineers at the University of Toronto have demonstrated a way to drive computer vision significantly with a deep neural network running on GPUs. After inserting the GPUs (previously used primarily in graphics), the performance of the computational neural network is immediately greatly improved, which is reflected in the noticeable improvement of the computer visual effect.
It was a revolutionary advance.
"Just a few years later, GPUs has been at the heart of deep learning," said Kurt Keutzer, professor of electronic engineering and computer science at the University of California, Berkeley. "The use of GPUs is becoming mainstream, and GUP is fundamentally changing computing by using dozens of to hundreds of processors in an application. ”
"GPU is a great throughput computing device," said Wen-mei W. Hwu, honorary chairman of Sanders iii–advanced Micro device, the University of Illinois at Urbana-Champaign, electronic and computer engineering. If you have only one task, there is no need to use the GPUs, because the speed is not going anywhere. However, if you have a lot of independent tasks between each other, use GPUs to be right. ”
A deep perspective
GPU architectures originate from underlying graphics rendering operations, such as shading graphics. In 1999, Nvida launched the GeForce 256, the world's first GPU. Simply put, this dedicated circuit ——-can be built into a video card or motherboard-leading and optimizing the computer's memory to speed up display rendering. Today, GPUs is used in a wider range of devices, including personal computers, tablets, mobile phones, workstations, electronic signage, gaming consoles, and embedded systems.
However, "the memory of many new applications in computer vision and deep learning is limited bandwidth," Keutzer explains, "In these applications, the speed of an application often ultimately depends on how much time it takes to extract data from memory and how long it will take to flow through the processor. ”
One of the most often overlooked advantages of deploying GPUs is the super-bandwidth of their processor-to-memory. The result, Keutzer points points out, is that "in applications with limited bandwidth, the relative advantage of this processor-to-memory bandwidth translates directly into super-application performance." "The key is that GPUs uses less power to provide faster floating-point operations (FLOPs, floating-point operations per second) that extend the energy efficiency advantage by supporting 16-bit floating-point numbers, which is more energy efficient than single-precision (32-bit) or double-precision (64-bit) floating-point numbers.
Multicore GPUs rely on a larger number of deployments of 32-bit to 64-bit, simpler processor cores. By contrast, using smaller, traditional microprocessors, typically 2-bit to 4-bit to 8-bit, how does it work?
"The use of microprocessor-GPUs achieves superior performance and provides better architecture support for deep neural networks." The performance advantages shown by GPUs on the deep neural network are gradually being transformed into more kinds of applications. "Keutzer said.
Today, a typical GPU cluster contains 8 to 16 GPUs, and researchers like Keutzer are trying to use hundreds of GPUs to train multiple deep neural networks simultaneously on very large datasets, otherwise it will take weeks of training time. This training requires a large amount of data to be run through the system to get it to a problem-solving state. At that point, it might be able to run in a GPU or a hybrid processor. "This is not an academic exercise. "Keutzer pointed out. "We're going to need this speed when we train neural networks for new applications like auto-driving cars." ”
The use of GPUs is becoming mainstream, and the calculation can be fundamentally changed by using multiple processors in a single application.
GPU technology is now progressing much faster than a traditional CPU, with powerful floating-point horsepower and low power consumption, the GPU's scalable performance accelerates the efficiency of deep learning and machine learning tasks in a way that is comparable to a turbocharged engine on a car, says Baidu senior researcher Bryan Catanzaro. "Deep learning is not something new. GPUs is not. But this area has been greatly improved in computing power and has a wealth of data to use before it starts to really sail. ”
Most of the progress comes from Nvidia, which continues to introduce more complex GPUs, including the new Pascal architecture, specifically designed to solve special tasks such as training and reasoning. In this latest GPU system, the Tesla P100 chip enables the encapsulation of 15 billion transistors on a piece of silicon, twice times the number of previous processors.
Another example is that Baidu is pushing the new frontiers of language recognition research. Its "Deep Speech" project, which relies on an end-to-end neural network, makes speech recognition accurate to human levels in short audio clips in both English and Chinese. The company is also exploring GPU technology in autonomous cars, which has been developing autonomous vehicles that automatically navigate the streets of Beijing, and has done maneuvers to change lanes, overtake, stop, and start.
At the same time, Microsoft's researchers in Asia use GPUs and a variant of deep neural networks-deep residual networks-to achieve higher accuracy in the task of classifying and identifying objects in computer vision.
Google, too, is using these technologies to continuously improve image recognition algorithms. Former Google AI researcher, now Open AI Research director Ilya Sutskever said: "The neural network is reviving." The core concepts of neural networks and deep learning have been discussed and thought over for years, but it is the development of the universal GPU that is the key to the success of neural networks and deep learning. ”
One Step Beyond
"While GPU technology is advancing to new frontiers in deep learning, many computational challenges persist. First, the efficient implementation of a standalone programmatic multi-core device such as the GPU is still difficult, and this difficulty can deteriorate as multiple GPU parallelism intensifies. "Keutzer said.
Unfortunately, he added, "Many of the highly efficient programmatic expertise of these devices is confined to the company, and many of the technical details that have been developed are still not widely used." ”
Similarly, keutzer that the design of deep neural networks is still widely described as "black technology" and that building a new deep neural network architecture is as complex as building a new microprocessor architecture. To make things worse, once this deep neural network architecture is built, "there are many knobs that are similar to hyper-parameters that, when applied in training, produce the accuracy they deserve only when these knobs are properly set." All of this creates a knowledge gap between these known and unknown "
"In the field of deep neural networks or GPU programming, individuals with expertise are scarce, and those who know both are rarer. ”
Another challenge is to understand how to use the GPU most effectively. For example, Baidu needs 8-16 GPUs to train a model to reach a 40%-50% floating point peak across the application. "This means that the performance is very limited. "The reality is that we need to use gpu,8 or 16 more on a larger scale than enough, and what we might need is 128 GPUs in parallel," Catanzaro said. "This requires better connectivity and the ability to support 32-bit floating-point support to 16-bit floating-point support." Nvidia's next generation gpu--pascal are likely to solve these problems.
In addition, there is a big hurdle for the GPU to better integrate with other GPUs and CPUs. Hwu points out that these two types of processors are not often integrated, and they rarely have enough bandwidth. This ultimately translates into a limited number of applications and the ability to run the system.
"You really need your GPU to have the ability to run big data tasks, and your GPU can pause to make the uninstallation process more cost-effective." "Catanzaro explains.
The Nvidia GPUs now exist on different chips, and they are usually connected to the CPU via an I/O bus (PCIe). That's one reason they can send a lot of tasks to the GPU. The future system integrates the GPU and CPU into a single package, and it can take on higher bandwidth and less risk, as well as maintaining shared consistency through the GPU and CPU.
Keutzer hopes that with the passage of time, the CPU and GPU can be better integrated, the more consistency and synchronization between the two can be achieved. In fact, Nvidia and Intel are also looking at this area. Keutzer noted that a new Intel chip called Knight's Landing (KNL) provides unprecedented computing power in the Xeon Phi 72-core super-computing processor, and integrates both CPU and GPU of the feature. At the same time, the chip also provides a bandwidth requirement of up to GB processor-to-memory per second, which will erode the GPU's advantage in this area.
Hwu noted that KNL's 72 cores can execute "a broad vector instruction (512 bytes) with each other." When converting to double precision (8 bytes) and single precision (4 bytes), Vector broadband will be 64 and 128. At this level, it has a similar execution model to the GPU. ”
Keutzer hopes that with the passage of time, the CPU and GPU can be better integrated, the more consistency and synchronization between the two can be achieved.
KNL Chip's programming model is a traditional x86 model, so HWU that programmers "need to write code through Intel C Compiler to make the chip quantifiable, or to use the intrinsic library functions of Intel AVX vectors." He added that the GPU programming model needed to be based on a core programming model.
In addition, the X86 kernel has cache consistency for the cache hierarchy, Hwu says, "but the GPU's first-tier cache is not clearly consistent, and it is accompanied by some reduced storage bandwidth." "However, he added," cache consistency is less important for the first-tier caching of most algorithms for deep learning applications. ”
In the next decade, all this versatility lies in how the big industry environment develops. Hwu says he believes Moore's law will continue to work for more than three generations, and designers and programmers can transition from almost discrete CPUs and GPU systems to integrated designs.
"If Moore's law stops working, it will also significantly affect the future of these systems, as well as the way people use hardware and software in deep learning and other tasks." "However, even if we solve the problem at the hardware level, the task of specific deep learning still requires a lot of tabbed data," Hwu said. At some level, we need to make breakthroughs in labeling data so that we have the ability to train in the necessary areas, especially in the field of autonomous driving. ”
Over the next few years, Sutskever says, machine learning will be widely applied to the GPU. "As machine learning approaches evolve, they are applied to areas far beyond today's scope of use and affect everything from healthcare to robotics to financial services and user experience." These advancements rely on the development of Faster GPUs, which will also enable machine learning to have research capabilities. ”
Adds Catanzaro says: The GPU is the gateway to future computing. Deep learning is exciting because when you add more data, it can scale. At this point, we will never be satisfied with the pursuit of more data and computational resources to solve complex problems. GPU technology is a very important part of expanding computing limits.
Golden Partner in deep learning: The GPU is Reinventing Computing (reproduced)