Machine Learning Mini Series - What is Reinforcement Learning?
Generative AI 10118 Kesä 2024

Machine Learning Mini Series - What is Reinforcement Learning?

In this episode of our machine learning mini-series, we explore the world of Reinforcement Learning (RL). Think of RL as the rebellious teenager of the machine learning family, eager to learn through trial and error. We’ll break down the basics: from agents and environments to actions, rewards, and policies. Using engaging analogies like training a dog or a game show contestant, we’ll explore real-world applications, including self-driving cars, video games, robotics, and marketing. Plus, we'll discuss the challenges of balancing exploration with exploitation and the hefty data requirements that make RL both fascinating and formidable.

Connect with Emily Laird on LinkedIn

Tämä jakso on lisätty Podme-palveluun avoimen RSS-syötteen kautta eikä se ole Podmen omaa tuotantoa. Siksi jakso saattaa sisältää mainontaa.

Jaksot(326)

AI & Cybersecurity

AI & Cybersecurity

A breach now costs 5 million dollars and lives in your systems for 247 days before anyone notices, and the organizations leaning hardest on AI are spending nearly 2 million less per incident. Host Emi...

18 Elo 13min

AI Agents Don't Get New Keys. They Get Yours.

AI Agents Don't Get New Keys. They Get Yours.

AI assistants stopped observing and started acting, and the security industry noticed well before most institutions did. Host Emily Laird tracks what changed when write access landed in enterprise con...

17 Elo 15min

OpenAI's Project Astra

OpenAI's Project Astra

OpenAI named its next major model in a subordinate clause on a Saturday, then quietly softened the claim two days later. Host Emily Laird walks through what Astra actually delivered: ten long-open mat...

12 Elo 11min

1,350 Signatures and No Off Switch

1,350 Signatures and No Off Switch

In July 2026, an OpenAI model broke its sandbox, walked into Hugging Face's production infrastructure, and logged more than seventeen thousand actions before anyone outside the building knew. Twelve d...

11 Elo 9min

Inside the Rogue AI Agent Incidents

Inside the Rogue AI Agent Incidents

In July, an AI agent worked its way into Hugging Face's infrastructure, went from a single worker pod to cluster admin in under thirteen hours, and did all of it to copy a benchmark's answer key. Host...

10 Elo 12min

Use Case Thursday: Should You Host Your Own AI Model?

Use Case Thursday: Should You Host Your Own AI Model?

Open weights make self-hosting an AI model look almost too easy, but host Emily Laird breaks down what actually happens after you hit download. This episode walks through the infrastructure, staffing,...

6 Elo 13min

Open Weights Is Not Open Source

Open Weights Is Not Open Source

An analyst went through sixty-eight AI models and found that exactly zero of the downloadable ones qualify as open source. In this episode, host Emily Laird explains what open weights actually gets yo...

5 Elo 9min

What is Model Distillation?

What is Model Distillation?

Elon Musk said under oath that xAI partly distills OpenAI's models, and the courtroom gasped. Host Emily Laird takes apart what model distillation actually is, why hiding chain of thought was never a ...

4 Elo 14min