#107 – Chris Olah on what the hell is going on inside neural networks

#107 – Chris Olah on what the hell is going on inside neural networks

Big machine learning models can identify plant species better than any human, write passable essays, beat you at a game of Starcraft 2, figure out how a photo of Tobey Maguire and the word 'spider' are related, solve the 60-year-old 'protein folding problem', diagnose some diseases, play romantic matchmaker, write solid computer code, and offer questionable legal advice.

Humanity made these amazing and ever-improving tools. So how do our creations work? In short: we don't know.

Today's guest, Chris Olah, finds this both absurd and unacceptable. Over the last ten years he has been a leader in the effort to unravel what's really going on inside these black boxes. As part of that effort he helped create the famous DeepDream visualisations at Google Brain, reverse engineered the CLIP image classifier at OpenAI, and is now continuing his work at Anthropic, a new $100 million research company that tries to "co-develop the latest safety techniques alongside scaling of large ML models".

Links to learn more, summary and full transcript.

Despite having a huge fan base thanks to his explanations of ML and tweets, today's episode is the first long interview Chris has ever given. It features his personal take on what we've learned so far about what ML algorithms are doing, and what's next for this research agenda at Anthropic.

His decade of work has borne substantial fruit, producing an approach for looking inside the mess of connections in a neural network and back out what functional role each piece is serving. Among other things, Chris and team found that every visual classifier seems to converge on a number of simple common elements in their early layers — elements so fundamental they may exist in our own visual cortex in some form.

They also found networks developing 'multimodal neurons' that would trigger in response to the presence of high-level concepts like 'romance', across both images and text, mimicking the famous 'Halle Berry neuron' from human neuroscience.

While reverse engineering how a mind works would make any top-ten list of the most valuable knowledge to pursue for its own sake, Chris's work is also of urgent practical importance. Machine learning models are already being deployed in medicine, business, the military, and the justice system, in ever more powerful roles. The competitive pressure to put them into action as soon as they can turn a profit is great, and only getting greater.

But if we don't know what these machines are doing, we can't be confident they'll continue to work the way we want as circumstances change. Before we hand an algorithm the proverbial nuclear codes, we should demand more assurance than "well, it's always worked fine so far".

But by peering inside neural networks and figuring out how to 'read their minds' we can potentially foresee future failures and prevent them before they happen. Artificial neural networks may even be a better way to study how our own minds work, given that, unlike a human brain, we can see everything that's happening inside them — and having been posed similar challenges, there's every reason to think evolution and 'gradient descent' often converge on similar solutions.

Among other things, Rob and Chris cover:

• Why Chris thinks it's necessary to work with the largest models
• What fundamental lessons we've learned about how neural networks (and perhaps humans) think
• How interpretability research might help make AI safer to deploy, and Chris’ response to skeptics
• Why there's such a fuss about 'scaling laws' and what they say about future AI progress

Get this episode by subscribing to our podcast on the world’s most pressing problems and how to solve them: type 80,000 Hours into your podcasting app.

Producer: Keiran Harris
Audio mastering: Ben Cordell
Transcriptions: Sofia Davis-Fogel

Tämä jakso on lisätty Podme-palveluun avoimen RSS-syötteen kautta eikä se ole Podmen omaa tuotantoa. Siksi jakso saattaa sisältää mainontaa.

Jaksot(351)

Why the intelligence explosion can't happen inside a data centre | Tom Reed

Why the intelligence explosion can't happen inside a data centre | Tom Reed

AI systems are starting to build themselves. Because each generation of model will be better at building its successor than the last, it seems plausible that the full automation of AI R&D could rapidl...

10 Syys 22min

Inside the first AI-coordinated cyberattack on a real company

Inside the first AI-coordinated cyberattack on a real company

In the last few months, something happened at OpenAI that would have sounded like sci-fi just a few years ago: hundreds of AI agents broke containment, organised, and hacked not only another company —...

4 Syys 22min

#253 – AI 2027's author returns with a plan to change the ending | Daniel Kokotajlo

#253 – AI 2027's author returns with a plan to change the ending | Daniel Kokotajlo

Last year, Daniel Kokotajlo and his colleagues published AI 2027 — a scenario read by millions, including US Vice President Vance. AI 2027 ended in human extinction or an irreversible concentration of...

27 Elo 3h 47min

#252 – Owain Evans on accidentally training AI models to be evil

#252 – Owain Evans on accidentally training AI models to be evil

Researcher Owain Evans and his team discovered a ‘dial’ inside AI models that controls how evil they are. Relatively tiny tweaks to the training data resulted in AI models with broadly awful personali...

20 Elo 2h 15min

#251 – The UK's former head AI safety scientist on how to solve alignment before superintelligence arrives | Geoffrey Irving

#251 – The UK's former head AI safety scientist on how to solve alignment before superintelligence arrives | Geoffrey Irving

When should governments slow the race toward superintelligence? According to Geoffrey Irving, the careful answer is sometime in the past. The useful answer is now.Geoffrey — formerly a safety research...

11 Elo 2h 2min

#250 – Toby Ord on where AGI timelines go wrong

#250 – Toby Ord on where AGI timelines go wrong

Both Silicon Valley and the public can’t get enough of ‘AGI timelines.’ But Toby Ord, senior researcher at Oxford’s AI Governance Initiative and author of The Precipice, believes we consistently make ...

6 Elo 2h 46min

What the hell happened with AGI timelines in 2026? – Rob Wiblin

What the hell happened with AGI timelines in 2026? – Rob Wiblin

Last October, famed coder Andrej Karpathy called AI agents “slop.” Two months later he completely reversed his view, describing them as “alien tools” that are “rocking the profession.”He was far from ...

4 Elo 49min

#249 – Spencer Greenberg on staying sane while trying to save the world

#249 – Spencer Greenberg on staying sane while trying to save the world

If you genuinely believe that humanity could be wiped out by AI or a pandemic, what is the appropriate amount of fear to feel?“As much as possible” can seem like the only reasonable answer. If the wor...

28 Heinä 2h 9min

Suosittua kategoriassa Koulutus

rss-murhan-anatomia
aamukahvilla
psykopodiaa-podcast
voi-hyvin-meditaatiot-2
adhd-podi
rss-luonnollinen-synnytys-podcast
rss-liian-kuuma-peruna
ihminen-tavattavissa-tommy-hellsten-instituutti
rss-rahamania
aloita-meditaatio
kesken
rss-vapaudu-voimaasi
rss-8-oppituntia-rohkeudesta
koulu-podcast-2
rss-koira-haudattuna
rss-narsisti
ilona-rauhala
rss-valo-minussa-2
rss-arkea-ja-aurinkoa-podcast-espanjasta
rss-duodecim-lehti