Transformers Need Glasses! - Federico Barbero

Transformers Need Glasses! - Federico Barbero

Federico Barbero (DeepMind/Oxford) is the lead author of "Transformers Need Glasses!".


Have you ever wondered why LLMs struggle with seemingly simple tasks like counting or copying long strings of text? We break down the theoretical reasons behind these failures, revealing architectural bottlenecks and the challenges of maintaining information fidelity across extended contexts.


Federico explains how these issues are rooted in the transformer's design, drawing parallels to over-squashing in graph neural networks and detailing how the softmax function limits sharp decision-making.


But it's not all bad news! Discover practical "glasses" that can help transformers see more clearly, from simple input modifications to architectural tweaks.


SPONSOR MESSAGES:

***

CentML offers competitive pricing for GenAI model deployment, with flexible options to suit a wide range of models, from small to large-scale deployments. Check out their super fast DeepSeek R1 hosting!

https://centml.ai/pricing/


Tufa AI Labs is a brand new research lab in Zurich started by Benjamin Crouzier focussed on o-series style reasoning and AGI. They are hiring a Chief Engineer and ML engineers. Events in Zurich.


Goto https://tufalabs.ai/

***


https://federicobarbero.com/


TRANSCRIPT + RESEARCH:

https://www.dropbox.com/s/h7ys83ztwktqjje/Federico.pdf?dl=0


TOC:

1. Transformer Limitations: Token Detection & Representation

[00:00:00] 1.1 Transformers fail at single token detection

[00:02:45] 1.2 Representation collapse in transformers

[00:03:21] 1.3 Experiment: LLMs fail at copying last tokens

[00:18:00] 1.4 Attention sharpness limitations in transformers


2. Transformer Limitations: Information Flow & Quantization

[00:18:50] 2.1 Unidirectional information mixing

[00:18:50] 2.2 Unidirectional information flow towards sequence beginning in transformers

[00:21:50] 2.3 Diagonal attention heads as expensive no-ops in LAMA/Gemma

[00:27:14] 2.4 Sequence entropy affects transformer model distinguishability

[00:30:36] 2.5 Quantization limitations lead to information loss & representational collapse

[00:38:34] 2.6 LLMs use subitizing as opposed to counting algorithms


3. Transformers and the Nature of Reasoning

[00:40:30] 3.1 Turing completeness conditions in transformers

[00:43:23] 3.2 Transformers struggle with sequential tasks

[00:45:50] 3.3 Windowed attention as solution to information compression

[00:51:04] 3.4 Chess engines: mechanical computation vs creative reasoning

[01:00:35] 3.5 Epistemic foraging introduced


REFS:

[00:01:05] Transformers Need Glasses!, Barbero et al.

https://proceedings.neurips.cc/paper_files/paper/2024/file/b1d35561c4a4a0e0b6012b2af531e149-Paper-Conference.pdf


[00:05:30] Softmax is Not Enough, Veličković et al.

https://arxiv.org/abs/2410.01104


[00:11:30] Adv Alg Lecture 15, Chawla

https://pages.cs.wisc.edu/~shuchi/courses/787-F09/scribe-notes/lec15.pdf


[00:15:05] Graph Attention Networks, Veličković

https://arxiv.org/abs/1710.10903


[00:19:15] Extract Training Data, Carlini et al.

https://arxiv.org/pdf/2311.17035


[00:31:30] 1-bit LLMs, Ma et al.

https://arxiv.org/abs/2402.17764


[00:38:35] LLMs Solve Math, Nikankin et al.

https://arxiv.org/html/2410.21272v1


[00:38:45] Subitizing, Railo

https://link.springer.com/10.1007/978-1-4419-1428-6_578


[00:43:25] NN & Chomsky Hierarchy, Delétang et al.

https://arxiv.org/abs/2207.02098


[00:51:05] Measure of Intelligence, Chollet

https://arxiv.org/abs/1911.01547


[00:52:10] AlphaZero, Silver et al.

https://pubmed.ncbi.nlm.nih.gov/30523106/


[00:55:10] Golden Gate Claude, Anthropic

https://www.anthropic.com/news/golden-gate-claude


[00:56:40] Chess Positions, Chase & Simon

https://www.sciencedirect.com/science/article/abs/pii/0010028573900042


[01:00:35] Epistemic Foraging, Friston

https://www.frontiersin.org/journals/computational-neuroscience/articles/10.3389/fncom.2016.00056/full

Det här avsnittet är hämtat från ett öppet RSS-flöde och publiceras inte av Podme. Det kan innehålla reklam.

Avsnitt(260)

Stealing Reasoning Traces from Proprietary LLM APIs — Ilia Shumailov & Alexander Panfilov

Stealing Reasoning Traces from Proprietary LLM APIs — Ilia Shumailov & Alexander Panfilov

Tim Scarfe speaks with Ilia Shumailov and Alexander Panfilov about their paper, Stealing Reasoning Traces from Proprietary LLM APIs.The core bug sounds deceptively simple: providers return encrypted r...

22 Aug 49min

Every Exponential Ends — Silicon Valley Forgot — Adam Becker

Every Exponential Ends — Silicon Valley Forgot — Adam Becker

Astrophysicist Adam Becker, author of "What Is Real?", joins Tim Scarfe to take apart the futures Silicon Valley keeps selling: the 2045 singularity, mind uploading, Mars colonies, and the AI apocalyp...

20 Aug 1h 18min

AI Is Learning at the Wrong Level of Abstraction — Matthieu Wyart

AI Is Learning at the Wrong Level of Abstraction — Matthieu Wyart

This episode is sponsored by Notion. Learn more about Notion's Developer Platform today at https://notion.com/mlstWhy can deep networks discover abstractions that shallow models miss? Statistical phys...

10 Aug 1h 18min

How Researchers Test AI for Hidden Goals — Apollo Research

How Researchers Test AI for Hidden Goals — Apollo Research

Can an AI do the right thing for the wrong reason? Tim Scarfe speaks with Apollo Research’s Alexander Meinke, Axel Højmark and Jérémy Scheurer about Measuring Reward-Seeking via Contrastive Belief Upd...

31 Juli 1h 18min

Why a Nation Can't Outsource Its Frontier AI - Alistair Pullen (Cosine AI)

Why a Nation Can't Outsource Its Frontier AI - Alistair Pullen (Cosine AI)

This episode is sponsored by Notion. Learn more about Notion's Developer Platform today at https://notion.com/mlstBritain's most capable coding model can't be exported, and that ban is the whole reaso...

13 Juli 55min

 The Benchmark With No Instructions — ARC-AGI-3 (winning team!)

The Benchmark With No Instructions — ARC-AGI-3 (winning team!)

Tim Scarfe travels to Zurich to sit down with the Tufa Labs ARC-AGI-3 team — founder Benjamin Crouzier, with Jeroen Cottaar, Dries Smit, Stefano Viel and Michal Tesnar — to work out what their leaderb...

1 Juli 1h 24min

The Thermodynamic AI Computing Chip - Thomas Ahle

The Thermodynamic AI Computing Chip - Thomas Ahle

Thomas Ahle wants Normal Computing to be the Lovable for chip design: type your intent, and a swarm of agents carries it from design through optimisation, formalisation and verification to tape-out. T...

28 Juni 1h 2min

He won a Nobel here for AlphaFold. Then he left. - John Jumper

He won a Nobel here for AlphaFold. Then he left. - John Jumper

This episode is sponsored by Notion. Learn more about Notion's Developer Platform today at https://notion.com/mlstProtein folding stalled biology for fifty years. A sequence of amino acids dictates a ...

22 Juni 53min

Populärt inom Teknik

uppgang-och-fall
market-makers
rss-elektrikerpodden
 och-bilen-gar-bra
rss-laddstationen-med-elbilen-i-sverige
bli-saker-podden
rss-technokratin
rss-en-ai-till-kaffet
rss-uppgang-och-fall
skogsforum-podcast
hej-bruksbil
gubbar-som-tjotar-om-bilar
bilar-med-sladd
rss-milpodden
natets-morka-sida
klocksnack-tillsammans-med-nymans-ur-1851
rss-fabriken-2
rss-veckans-ai
elbilsveckan
developers-mer-an-bara-kod