Quantizing Transformers by Helping Attention Heads Do Nothing with Markus Nagel - #663

Quantizing Transformers by Helping Attention Heads Do Nothing with Markus Nagel - #663

Today we’re joined by Markus Nagel, research scientist at Qualcomm AI Research, who helps us kick off our coverage of NeurIPS 2023. In our conversation with Markus, we cover his accepted papers at the conference, along with other work presented by Qualcomm AI Research scientists. Markus’ first paper, Quantizable Transformers: Removing Outliers by Helping Attention Heads Do Nothing, focuses on tackling activation quantization issues introduced by the attention mechanism and how to solve them. We also discuss Pruning vs Quantization: Which is Better?, which focuses on comparing the effectiveness of these two methods in achieving model weight compression. Additional papers discussed focus on topics like using scalarization in multitask and multidomain learning to improve training and inference, using diffusion models for a sequence of state models and actions, applying geometric algebra with equivariance to transformers, and applying a deductive verification of chain of thought reasoning performed by LLMs. The complete show notes for this episode can be found at twimlai.com/go/663.

Tämä jakso on lisätty Podme-palveluun avoimen RSS-syötteen kautta eikä se ole Podmen omaa tuotantoa. Siksi jakso saattaa sisältää mainontaa.

Jaksot(793)

World Models and the Future of Spatial AI with Justin Johnson - #775

World Models and the Future of Spatial AI with Justin Johnson - #775

In this episode, Justin Johnson, co-founder of World Labs, joins us to discuss world models and the emerging field of spatial AI. We explore why many researchers see capabilities beyond language as an...

1 Syys 1h 6min

Why the Next AI Breakthrough May Come from Physics with Max Welling - #774

Why the Next AI Breakthrough May Come from Physics with Max Welling - #774

The conventional wisdom in AI is that the next breakthrough will come from more compute, more data, and larger models. But what if the next leap comes from somewhere else? In this episode, Max Wellin...

26 Elo 58min

Why Image Generation Needs More Than Bigger Models with Fatih Porikli - #773

Why Image Generation Needs More Than Bigger Models with Fatih Porikli - #773

Text-to-image models have become remarkably good at producing realistic images. But realism isn’t the same as correctness. Ask for several distinct people, a specific composition, or a high-resolution...

12 Elo 56min

Why Models Are AI’s Next Training Dataset with Damian Borth - #772

Why Models Are AI’s Next Training Dataset with Damian Borth - #772

For more than a decade, AI has advanced by training ever-larger models on ever-larger datasets. But as high-quality training data becomes harder to find and pretraining grows increasingly expensive, r...

27 Heinä 47min

How AI Learns to Smell with Alex Wiltschko - #771

How AI Learns to Smell with Alex Wiltschko - #771

In this episode, Alex Wiltschko, founder and CEO of Osmo, joins the show to discuss his goal of giving computers a sense of smell and what it takes to build olfactory intelligence. We explore the sci...

8 Heinä 59min

Why AI Agents Break the GenAI Security Model with Devvret Rishi - #770

Why AI Agents Break the GenAI Security Model with Devvret Rishi - #770

In this episode, Sam talks with Dev Rishi, GM of AI at Rubrik, about what happens when agents move beyond answering questions and start taking action across tools, systems, and business processes. We...

16 Kesä 56min

Is RAG Dead? Lessons from Building AI for Tax Law with Alex Bowcut - #769

Is RAG Dead? Lessons from Building AI for Tax Law with Alex Bowcut - #769

As context windows grow into the millions of tokens, many AI practitioners are questioning whether retrieval-augmented generation (RAG) is still necessary. If modern models can ingest entire libraries...

9 Kesä 51min

Relational Foundation Models for Enterprise Data with Jure Leskovec - #768

Relational Foundation Models for Enterprise Data with Jure Leskovec - #768

In this episode, Jure Leskovec, co-founder and chief scientist at Kumo and professor of computer science at Stanford, joins us to explore two fronts of his work: AI for science and relational deep lea...

21 Touko 1h 6min

Suosittua kategoriassa Politiikka ja uutiset

uutiscast
aikalisa
politiikan-puskaradio
rss-viihde-media
rss-ootsa-kuullut-tasta
ootsa-kuullut-tasta-2
vallattomat
otetaan-yhdet
rss-voi-venaja
rss-vaalirankkurit-podcast
rss-raha-talous-ja-politiikka
rss-girls-finish-f1rst
rss-asiastudio
rss-seksicast
rss-podme-livebox
rss-ulkopoditiikkaa
et-sa-noin-voi-sanoo-esittaa
linda-maria
rss-pallo-keskelle-2
rss-kaikki-uusiksi