CONNECT WITH US
Sign out

Nvidia's CEO talks about trends towards ChatGPT and large language models

Monica Chen and Jessie Shen, DIGIMES Asia, Taipei
0

Credit: DIGITIMES

Nvidia CEO Jensen Huang at this year's GPU Technology Conference (GTC) discussed a variety of subjects and technological development trends, particularly those related to accelerated computing and AI. In a Q&A call hosted by Nvidia for the press based in Asia on March 22, Huang revealed his vision for the era of accelerated computing and AI that he believes has come, and trends towards applications like ChatGPT and massive language models.

Huang discussed ChatGPT's outstanding AI advancement before the Q&A began. Nvidia's DGX AI supercomputer was responsible for the enormous language model breakthrough. The inflection point in AI has been brought about by generative AI, said Huang, adding that this is the AI equivalent of the iPhone moment. The following is a summary of the Q&A.

Q: As you may be aware, the trend of AI toward larger AI models and data science raises the requirement for computational capacity. So, how is Nvidia using its technology to alleviate energy consumption issues and promote sustainable growth through the usage of green energy?

A: This is a critical question, and it was the first thing I mentioned during my talk regarding sustainability. Keep in mind that artificial intelligence accounts for a very small portion of the world's computing today. In truth, the vast bulk of the world's computing today has been essentially driven in the previous 40 years by a very powerful force known as Moore's Law. Moore's Law lasted roughly 30 years and has begun to decelerate considerably in the last five years since we've reached the boundaries of physics. We can downsize the transistor but not the atom.

One issue is to imagine what would happen if Moore's Law failed yet more computing was still required. When we enhance performance by ten times or throughput by ten times, does the power increase by ten times and the cost increase by ten times? Of course, this is unsustainable.

The first thing we need to do is accelerate every potential workload. Accelerated computing is great since it is a whole stack problem. You have to modify and invent new algorithms, new software algorithms, new chips, and new systems. That is what Nvidia does for a living. Domain by domain, domain by domain. This cannot be done once for each application. This must be done once for each domain of applications. However, if you do, you may boost performance by 50 times, as we recently accomplished with TSMC and semiconductor manufacturing, computational lithography, the most computationally intensive application in all of EDA. As a result, we can enhance performance by 50 times. We can cut power and costs by nearly tenfold. That is what we must accomplish, task by job, domain by domain by domain. First and foremost, accelerate the workload.

Second, we can reclaim the electricity because, of course, the data center already has it. If we speed the workload, we will recapture power, cut power, and redirect it to new growth. Imagine how much growth opportunity we have if we can recapture 10 times the power through acceleration. So, first and foremost, we must accelerate application processing. There isn't another good example. We don't have any other good ideas for long-term sustainability.

Artificial intelligence also has the ability to reduce the amount of computation by maybe 10,000 or 100,000 times. And the reason for that is that you can teach an AI the laws of physics. After that, then the AI can use the knowledge, use the skill, to predict physics instead of computing physics. Let me give you an example. If I were to throw a ball, my dogs, even when they were puppies, can jump into the air and predict where the ball is going to fly to and grab it from the air. The reason why is because they... Of course, you are aware that Newtonian physics was not computed. They don't grasp gravity's laws. They've just noticed that when I throw a ball, they can guess where it'll land. We can teach an AI the laws of physics and many forms of physics, and we have now proved that we can cut the amount of computing by 1,000, 10,000, 100,000 times for thermal dynamics, fluid dynamics, and quantum chemistry. This will save energy.

Accelerated computing and AI will both save energy.

Q: Many Chinese enterprises, particularly small and medium-size ones building large language models, are concerned about computer capacity scarcity. Nvidia is now providing cloud services, which is fantastic news. Could you please explain to us when and how this cloud service will be available to Chinese users, as well as whether it will compete with existing cloud service providers such as Microsoft Azure?

A: We will engage with cloud service providers in China in the same way that we do in the West. In the West, we launched a number of cloud services in which we collaborated with CSPs to deploy the Nvidia DGX AI supercomputer in their cloud.

Ampere and Hopper will be available in China, as previously reported. The processors will strictly adhere to all export regulations and laws. They will be used by Chinese cloud companies. Alibaba, Tencent, and Baidu are all excellent partners, and I am confident that they will have the most powerful AI computer systems accessible. Okay, so I believe that Alibaba, Tencent, and Baidu will have fantastic cloud capabilities within NVIDIA's AI for the young enterprises that are now building enormous language models and jumping on the generative AI revolution.

Q: Can you elaborate more on the future demands of the AI supercomputer, particularly in 2023? Another question is that you've contributed the DGX to OpenAI in 2016, but did you ever expect OpenAI to have such a "iPhone moment of AI" in 2016?

A: Well, let me start backward. The reason why I delivered the world's first DGX to OpenAI is because I had so much confidence in the team. This is an extraordinary team. Ilya was there, Greg was there, Sam's there. Simply put, this is a top-notch crew. At DeepMind, Denis also oversees a top-notch AI research group. These two teams are therefore among the best in the world at what they do. And I'm quite proud of the work they accomplish.

And inflection point. There are two things that are happening simultaneously. In the near term, of course, we are seeing accelerating demand for AI supercomputers, this DGX AI supercomputer, and we're seeing accelerating demand for the Hopper inference system. So this is for training. This is for inference. We're seeing accelerating demand for both of these things. The accelerating demand is coming broad-based from CSPs all over the world, startups all over the world. And the reason everybody now realized that the inflection point is here and they know what is now possible to achieve with AI.

What is going to happen and why is this an inflection point? The first thing that's going to happen is that these AI supercomputers, the DGX AI supercomputer that's here is going to move from, extend from research, which is where it's really been for the last six to eight years. There's a few areas of industrial work, computer vision for self-driving cars, computer vision for robotics, but really narrowly focused on computer vision. Now all of a sudden, this AI is going to move into every single industry because one of the things that I said earlier, and it's a very, very important concept, is that AI has now learned the language of many domains. It learned the language of 3D graphics, it learned the language, of course, human language. It learned the language of motion. It learned the language of images. And once you can learn the language of something, you can of course understand it and produce it.

I picked up Mandarin when I first moved to the United States, and now that I know it, I can both understand it and produce or generate information in it. The same concept also applies to proteins, chemicals, and a wide variety of different domains. And this AI supercomputer, the transformer engine, will be used by every single business and industry to learn their languages. It might be a language unique to the industry, such as the language of banking, the language of law, the language of your own company, the language of farming, the language of astronomy, or any number of other languages. Therefore, this is the moment when this computer transitions from a research computer to now essentially an AI factory.

Keep in mind that the previous industrial revolution involved factories producing mechanical items. The most valuable thing we now know will be produced by the next factory, the next industrial revolution, which will be a digital factory that produces intelligence. Because this is a new industry, the first thing you will see is the development of AI factories. The second is that in the future, generative AI will be integrated into every single application. Browsers will connect to the internet. Naturally, Microsoft has already stated that the entire Office suite will be linked. Teams will be able to access generative AI. Generic AI will be linked to Google Docs. Adobe also announced they would connect to generative AI. In the future, generative AI will be included into every application. And the GPU will need to perform the inference when they are coupled and produce some information, knowledge, or content. And as a result, the biggest market window for GPU inferencing has opened. These are the two inflection points.

Q: In comparison to the previous three AI booms, how significant do you believe the focus on generative AI will be this time around? Can you again be more precise about how Nvidia is getting ready for this demand?

A: The AI boom really already has several chapters. The first chapter was 10-12 years ago when Nvidia recognized that deep learning was going to change how computing is done, and that deep learning and machine learning will revolutionize the way software will be developed and deployed. And we went to reinvent the modern computer, and this is our latest generation. This computer looks completely different than previous generation computers. The reason for that is that in the past, a computer programmer would sit and type and create. Now, the computer programmer works with this computer to type and create. Together, they write software that was impossible before. So phase one was building AI infrastructure.

Phase two was AI learning perception. Before you can make a prediction or do something helpful, you have to understand the environment. Human perception includes our eyes, our ears, our senses, and previous knowledge. All of that enhances our perception capability. In the last decade, as we were developing the computer, we were also inventing perception. Now, of course, the breakthrough is very clear. You can use it for computer vision, for self-driving cars, computer vision for automated robotics, computer vision for automatic checkout, and so many different things. There's a robotics company in Japan that has robots in 7-Eleven to automate stocking and inventory management.

So there are so many things that you can do with computer vision already, but the big breakthrough isn't perception only. All of us have perceptions, but the reason why we have value to society is that we generate information. We generate stories, we generate music, we generate movies, we generate designs, we generate strategies, we generate things. This is now the era of generative AI. This is chapter three. Chapter three is going to now allow AI to help us and be our co-creators, our co-pilots, our co-producers. They will be our partners in almost everything that we do.

Now, what is the benefit of having a partner that is able to help us create the first draft? Maybe help us start with an initial design to inspire our imagination. Sometimes we have a mental block and we need a partner to open our minds. Maybe it's to simplify just a mountain of work that we have so that we can add the ultimate value. So it will inspire us. It will accelerate our work. It will make us much more productive. Our software engineers around the world are already using Copilot to help write software. In just the last six months, we already experienced that Copilot has improved our productivity almost by a factor of two. Now, remember, software engineers are some of the most expensive engineers in the world. If we can improve their productivity by a factor of two, incredible value has been created.

The next chapter is generative AI. Our company is preparing for all of this in several different ways. The first way is building the computer and developing the algorithms and software. The second thing that we're doing is putting everything into the cloud, either cloud of our partners like Alibaba, Tencent and Baidu, or Azure and GCP and Amazon and OCI, and so on. We are also standing up with the CSPs to bring up Nvidia AI supercomputers in their cloud. So we have many different ways that we're going to make Nvidia AI infrastructure as accessible and as quickly as possible.

The demand for both training and inference has already grown and accelerated because of the demand for generative AI. Thankfully, we are in excellent hands when it comes to our supply issue. Although we have a sufficient supply for the entire year, we are going to work as quickly as we can to meet our clients' delivery requests. As you are aware, neither the economy nor the business sector are particularly vibrant. In order to access a large supply of systems, assembly, packaging, memory, and cutting-edge chips, we have this capability. So, there is a plentiful supply.

Q: At the same time that ChatGPT and AI are becoming increasingly popular, the gaming business is declining. Meanwhile, the Chinese electric vehicle market is faltering. In China, EV sales are slowing. So, what is Nvidia's current plan? Are you shifting your emphasis away from gaming and toward data centers, or how do you prioritize different businesses?

A: These are three highly major segments that are also three quite different businesses. And, at its core, the computer engine is the same. Nvidia's foundation is accelerated computing. So, while the businesses are different, the engine is the same, which is why improving computer graphics helps our AI. When we develop AI, we improve computer graphics. When we improve computer graphics, we can utilize it to simulate the Omniverse, which aids self-driving cars. AI advancements benefit self-driving automobiles. So I chose all of these businesses for a reason: they are all similar, and if we are good at one, we can be good at others. When we invest in one, we are simultaneously investing in the others.

Frankly, I think that the gaming industry is recovering, it's not slowing. It had a difficult last year, and we've been slowly recovering, and from our perspective, Ada is doing extremely well. And now the China market is back, it's open, and games are being certified again, and the West was always quite vibrant.

So I'm quite enthusiastic about the gaming industry. And then with respect to electric vehicles, remember that we serve the electric vehicle industry in both the data center and the car. In the data center, we serve in AI for training the model, as well as Omniverse for simulation of the cars. And so sometimes, the growth of the EV itself is quite good, and around the world is still quite good. Sometimes, even if the growth rate is a little bit different, slowed, it's okay because everybody's still investing in AI for their future software. And so I'm not at all concerned with AV. In fact, AV is one of our fastest-growing markets. It has grown over 100% year over year, and I expect the next year to also be a very big year.

Q: Will Nvidia continue to build larger scale foundation models on its own, or will it really focus on providing a platform for other players' foundation models?

A: If I can describe the customers in several ways, first of all, there's some excellent laboratories like OpenAI, DeepMind, Google Brain, Meta Research. They have just extraordinary capabilities in developing AI models. And we partner very closely with them. We dedicate a lot of engineering capability towards them and help them advance their work. And there, we are very much an AI infrastructure, AI computing partner.

There are customers on the other extreme, which would like not to develop the AI models at all, but would like to utilize the amazing models of the companies that I described. And so they would use the AI models as a service.

There's also a group of customers in the middle. They can't use the open models. And the reason for that is because their usage pattern, their usage case, is too domain specific. Maybe it's highly related to biology, maybe it's biomolecules, maybe it's synthetic molecules, maybe it's synthetic proteins. And so maybe it's related to drug design, maybe it's related to physics. Physics has its own foundation model that will likely occur. There are many different areas, many domains where foundation models can be created.

And the data to train that foundation model is completely proprietary. So if a customer has proprietary data and they would like to create and operate a custom model to perform skills very specific to their company and their industry, then we have the skills to create the AI with them. They have the domain knowledge, they have the data, we have the expertise, and we have the computing. So that is what we call NVIDIA AI Foundations. And we started with three foundation models, one for human language, one for visual language, and one for biology. And we'll have other languages in the future. The purpose of it is singular, not to offer it as a service. It's not B2C. It is also not going to be offered as a service that you can just use. It's designed to help companies build custom models. And we will continue to advance our model development. We have many big projects coming, and NVIDIA has six supercomputers, we have some of the largest in the world. We designed them so that we can advance large language model development.

Q: Do you think cuLitho will extend the life of Moore's Law?

A: Moore's Law was defined as twice performance at the same price, at the same power, every year and a half. So every five years you can increase your performance at the same power and the same price, which is the reason why the personal computer has improved performance by so much in the 40 years that I've been in the industry.

And yet the price of a personal computer is still about the same price, about $1,000. However, that trend has now ended. If you would like to have 10 times the performance every five years, the power will increase substantially and the price will increase, and that's the reason why Moore's law has ended. We will continue to add more transistors and TSMC and other foundries are, and other manufacturers are, really pushing the limits.

So I think that we're going to get more and more transistors. We'll have all kinds of architecture. As you know, Nvidia created the super chip. This is essentially two dies that are connected together into one giant die across a chip-to-chip interconnect. And so this is really one chip, and so as we call it a superchip. Notice, this is eight GPUs into one giant GPU, okay, so this is eight chips and this is two chips, and these are superchips.

There are other ways to do it. We can create chiplets, which is essentially the same as superchips. And there's a lot of different ways to configure, and we'll use more and more and more transistors. However, the fundamental CPU law of more performance at the same power and same price is largely over.

And so the answer is really two things. We want to, number one, use a different way of doing computing. So we add an accelerator, accelerated computing, so that we can take the workload of the software that is parallelizable and accelerated on the GPU. And by acceleration, we can bring many orders of magnitude of more performance and yet still reduce the power and reduce the cost. So that's the number one way - accelerated computing.

The second way is to completely change the application. Instead of using just human engineered software, which is not very efficient, to use a computer to write the software, use the GPU to write the software. We call it AI. That AI software is very efficient and also has the benefit of being very easy to accelerate. And so two fundamental approaches to the future of computing so that we can continue to gain more performance at the same price, or even lower price, and lower power. And that is accelerated computing and artificial intelligence.

Jensen Huang

Jensen Huang, co-founder and CEO of Nvidia
Photo: Nvidia