57. Eldar Kurtic - Efficient Inference through sparsity and quantization - Part 2/2 - Austrian Artificial Intelligence Podcast | Transcription & Insights

Description

Hello and welcome back to the AAIP This is the second part of my interview with Eldar Kurtic and his research on how to optimiz inference of deep neural networks. In the first part of the interview, we focused on sparsity and how high unstructured sparsity can be achieved without loosing model accuracy on CPU's and in part on GPU's. In this second part of the interview, we are going to focus on quantization. Quantization tries to reduce model size by finding ways to represent the model in numeric representations with less precision while retaining model performance. This means that a model that for example has been trained in a standard 32bit floating point representation is during post training quantization converted to a representation that is only using 8 bits. Reducing the model size to one forth. We will discuss how current quantization method can be applied to quantize model weights down to 4 bits while retaining most of the models performance and why doing so with the models activation is much more tricky. Eldar will explain how current GPU architectures, create two different type of bottlenecks. Memory bound and compute bound scenarios. Where in the case of memory bound situations, the model size causes most of the inference time to be spend in transferring model weights. Exactly in these situations, quantization has its biggest impact and reducing the models size can accelerate inference. Enjoy. ## AAIP Community Join our discord server and ask guest directly or discuss related topics with the community. https://discord.gg/5Pj446VKNU ### References Eldar Kurtic: https://www.linkedin.com/in/eldar-kurti%C4%87-77963b160/ Neural Magic: https://neuralmagic.com/ IST Austria Alistarh Group: https://ist.ac.at/en/research/alistarh-group/

Audio

Featured in this Episode

No persons identified in this episode.

Transcription

This episode hasn't been transcribed yet

Help us prioritize this episode for transcription by upvoting it.

0 upvotes

🗳️ Sign in to Upvote

Popular episodes get transcribed faster

Other recent transcribed episodes

Transcribed and ready to explore now

3ª PARTE | 17 DIC 2025 | EL PARTIDAZO DE COPE

01 Jan 1970

El Partidazo de COPE

Buchladen: Tipps für Weihnachten

20 Dec 2025

eat.READ.sleep. Bücher für dich

BOJ alza 25pb decennale sopra 2%, Oracle vola con accordo Tik Tok, 90 mld eurobond per Ucraina | Morning Finance 

19 Dec 2025

Black Box - La scatola nera della finanza

365. The BEST advice for managing ADHD in your 20s ft. Chris Wang

19 Dec 2025

The Psychology of your 20s

LVST 19 de diciembre de 2025

19 Dec 2025

La Venganza Será Terrible (oficial)

Cuando la Ciencia Ficción Explicó el Mundo que Hoy Vivimos

19 Dec 2025

El Podcast de Marc Vidal

Comments

There are no comments yet.

Please log in to write the first comment.

Austrian Artificial Intelligence Podcast

57. Eldar Kurtic - Efficient Inference through sparsity and quantization - Part 2/2

This episode hasn't been transcribed yet

Other recent transcribed episodes

3ª PARTE | 17 DIC 2025 | EL PARTIDAZO DE COPE

Buchladen: Tipps für Weihnachten

BOJ alza 25pb decennale sopra 2%, Oracle vola con accordo Tik Tok, 90 mld eurobond per Ucraina | Morning Finance

365. The BEST advice for managing ADHD in your 20s ft. Chris Wang

LVST 19 de diciembre de 2025

Cuando la Ciencia Ficción Explicó el Mundo que Hoy Vivimos

Sign in to Audioscrape

Share this moment