Free - Ai Labs Talks

Tech Talk | LLM Inference Energy Efficiency

Dubai Exhibition Centre (DEC)
Wednesday, Feb 5, 2025
1:00 PM - 1:15 PM | Asia/Dubai

Hall 1, Ai Labs

Hall 1 | Ai Labs
English
About
With the widespread use of large language models (LLMs) across industries, and the rise of LLM based-AI Agentic-AI Framework Applications, the inference serving for these models is ever expanding. Given the high compute and memory requirements of modern LLMs, massive quantities of top-of-the line GPUs ( and other bare-metal silicon ) are being deployed to serve these models. Energy availability has come to the forefront as the biggest challenge for data center expansion to serve these models. We would present the trade-offs brought up by making energy efficiency the primary goal of LLM serving under performance SLOs. We show that depending on the inputs, the model, and the service-level agreements, there are several tweaks available to the LLM inference provider to use for being energy efficient.We will discuss the impact of these changes on the latency, throughput, as well as the energy. By exploring these tradeoffs, we offer valuable insights into optimizing energy usage without compromising on performance, thereby paving the way for sustainable and cost-effective LLM deployment in data center environments.

Speakers

Roger Habr

Chief Data And Ai OfficerGulf Data Hub