DeepSeek-OCR: Contexts Optical Compression
GenAI Level UP24 Okt 2025

DeepSeek-OCR: Contexts Optical Compression

The single biggest bottleneck for Large Language Models isn't intelligence—it's cost. The quadratic scaling of self-attention makes processing truly long documents prohibitively expensive, a fundamental barrier that has stalled progress. But what if the solution wasn't more compute, but a radically simpler, more elegant idea?

In this episode, we dissect a groundbreaking paper from DeepSeek-AI that presents a counterintuitive yet insanely great solution: Contexts Optical Compression. We explore the astonishing feasibility of converting thousands of text tokens into a handful of vision tokens—effectively compressing text into a picture—to achieve unprecedented efficiency.

This isn't just theory. We go deep on the novel DeepEncoder architecture that makes this possible, revealing the specific engineering trick that allows it to achieve near-lossless compression at a 10:1 ratio while outperforming models that use 9x more tokens. If you're wrestling with context length, memory limits, or soaring GPU bills, this is the paradigm shift you've been waiting for.

In this episode, you will discover:

    • (02:10) The Quadratic Tyranny: Why long context is the most expensive problem in AI today and the physical limits it imposes.

    • (06:45) The Counterintuitive Leap: Unpacking the "Big Idea"—compressing text by turning it back into an image, and why it's a game-changer.

    • (11:20) Inside the DeepEncoder: A breakdown of the brilliant architecture that serially combines local and global attention with a 16x compressor to achieve maximum efficiency.

    • (17:05) The 10x Proof: We analyze the staggering benchmark results: achieving over 96% accuracy at 10x compression and still retaining 60% at a mind-bending 20x.

    • (23:50) Beyond Simple Text: How this method enables "deep parsing"—extracting structured data from charts, chemical formulas, and complex layouts automatically.

    • (28:15) A Glimpse of the Future: The visionary concept of mimicking human memory decay to unlock a path toward theoretically unlimited context.


Denne episoden er hentet fra en åpen RSS-feed og er ikke publisert av Podme. Den kan derfor inneholde annonser.

Episoder(45)

Recursive Self Improvement

Recursive Self Improvement

Imagine holding a wrench on an assembly line. Suddenly, it leaps from your hand, sprouts its own mechanical arms, and begins forging a faster, lighter wrench without you. You are no longer the creator...

7 Jun 1h

Master the New Physics of AI with Context Graphs & GraphRAG

Master the New Physics of AI with Context Graphs & GraphRAG

Stop trying to find the "magic words" to hack your LLM. The era of the Prompt Engineer—tweaking adjectives and hoping for the best—is officially over. We are entering the age of the Context Engineer, ...

1 Feb 17min

Context Graph

Context Graph

Stop feeding your AI static facts in a dynamic world.Most RAG systems and Knowledge Graphs rely on a fundamental unit called the "Triple" (Subject, Verb, Object). It’s efficient, but it’s brittle. It ...

25 Jan 19min

Nested Learning: The Illusion of Deep Learning Architectures

Nested Learning: The Illusion of Deep Learning Architectures

Why do today's most powerful Large Language Models feel... frozen in time? Despite their vast knowledge, they suffer from a fundamental flaw: a form of digital amnesia that prevents them from truly le...

14 Nov 202513min

Memento: Fine-tuning LLM Agents without Fine-tuning LLMs

Memento: Fine-tuning LLM Agents without Fine-tuning LLMs

What if you could build AI agents that get smarter with every task, learning from successes and failures in real-time—without the astronomical cost and complexity of constant fine-tuning? This isn't a...

1 Nov 202518min

MemGPT: Towards LLMs as Operating Systems

MemGPT: Towards LLMs as Operating Systems

Have you ever felt the frustration of an LLM losing the plot mid-conversation, its brilliant insights vanishing like a dream? This "goldfish memory"—the limited context window—is the Achilles' heel of...

1 Nov 202518min

A Definition of AGI

A Definition of AGI

For decades, Artificial General Intelligence has been a moving target, a nebulous concept that shifts every time a new AI masters a complex task. This ambiguity fuels unproductive debates and obscures...

23 Okt 202519min

Populært innen Teknologi

lydartikler-fra-aftenposten
teknisk-sett
tomprat-med-gunnar-tjomlid
shifter
rss-ki-praten
elektropodden
rss-ai-forklart
fornybaren
nasjonal-sikkerhetsmyndighet-nsm
rss-bouvet-bobler
hans-petter-og-co
pedagogisk-intelligens
rss-polypod
rss-alt-som-gar-pa-strom
rss-fish-ships
smart-forklart
rss-kunstig-intelligens-med-elisabeth-maren-og-morten
rss-fisketimen
rss-nkom-innsikt
energi-og-klima