2-1-5. The Industrial Pipeline — Training Large Models
LLM Primer17 Feb

2-1-5. The Industrial Pipeline — Training Large Models

In this episode, we move from the theoretical blueprint of the Transformer to the operational reality of building a Large Language Model. We explore how an empty mathematical shell is transformed into a capable system through a massive, coordinated engineering process known as training.

Join us as we:

Curate the Curriculum: We discuss why "more data" isn't always better, explaining the critical steps of deduplication, filtering, and balancing diverse sources like web text, books, and code.

Minimize the Surprise: We break down the mathematical objective of Cross-Entropy Loss and the optimization algorithm Gradient Descent, revealing how billions of parameters are nudged iteratively to improve prediction accuracy.

Distribute the Load: We examine the physical infrastructure required for training, detailing how strategies like Data Parallelism and Model Parallelism allow engineers to split massive models across thousands of GPUs.

Balance the Learning: We analyze the risks of Overfitting (memorizing data) versus Underfitting (failing to learn patterns), and how regularization ensures a model can generalize to new, unseen text.

This episode reveals that training an LLM is not just a math problem, but a large-scale systems engineering challenge.

Denne episoden er hentet fra en åpen RSS-feed og er ikke publisert av Podme. Den kan derfor inneholde annonser.

Episoder(23)

Prompt Injection and Jailbreaks

Prompt Injection and Jailbreaks

This chapter examines prompt injection and jailbreak attacks, which exploit a language model's inherent inability to distinguish between authoritative developer instructions and untrusted user data. I...

7 Jul 45min

Data Security and Privacy

Data Security and Privacy

This chapter examines data security and privacy throughout the LLM lifecycle. It explores the inherent risks of training data, such as copyright issues, personal information (PII) contamination, and d...

7 Jul 32min

Threat Modeling for LLM Systems

Threat Modeling for LLM Systems

This chapter adapts traditional threat modeling frameworks (such as STRIDE, PASTA, and attack trees) specifically for the unique vulnerabilities of LLM systems. It guides defenders through identifying...

6 Jul 54min

Why AI Security Is Different

Why AI Security Is Different

This chapter explains that AI security fundamentally and structurally differs from traditional software security. Instead of finding and patching clear bugs in readable source code, defenders must sec...

6 Jul 49min

2-7-7. Hallucinations and Reliability: Managing Confident Errors

2-7-7. Hallucinations and Reliability: Managing Confident Errors

This episode covers Chapter 7, examining why Large Language Models confidently generate false information. We discuss the probabilistic nature of "hallucinations," the dangerous gap between fluency an...

19 Feb 16min

2-7-6. Retrieval-Augmented Generation Risks: Securing the Knowledge Pipeline

2-7-6. Retrieval-Augmented Generation Risks: Securing the Knowledge Pipeline

This episode covers Chapter 6, focusing on the security implications of connecting models to external data (RAG). We discuss how this introduces new trust boundaries, the dangers of malicious document...

19 Feb 34min

2-7-5. Input Validation and Output Filtering: The Defense Pipeline

2-7-5. Input Validation and Output Filtering: The Defense Pipeline

This episode covers Chapter 5, detailing how to build disciplined pipelines around an AI model. We discuss strategies for sanitizing user inputs to catch attacks early, the importance of structured pr...

18 Feb 29min

2-7-4. Prompt Injection and Jailbreaks: Defending the Interpreter

2-7-4. Prompt Injection and Jailbreaks: Defending the Interpreter

This episode explores Chapter 4, detailing how attackers manipulate model behavior through crafted inputs like instruction overrides. We discuss why prompt injection is an inherent property of instruc...

18 Feb 37min

Populært innen Teknologi

lydartikler-fra-aftenposten
tomprat-med-gunnar-tjomlid
teknisk-sett
smart-forklart
elektropodden
nasjonal-sikkerhetsmyndighet-nsm
rss-ki-praten
fornybaren
shifter
rss-kunstig-intelligens-med-elisabeth-maren-og-morten
teknologi-og-mennesker
rss-alt-som-gar-pa-strom
rss-snakk-om-sikkerhet
rss-ai-forklart
rss-alt-vi-kan
rss-fisketimen
digital-forretningsforstaelse
rss-bouvet-bobler
pedagogisk-intelligens
rss-polypod