Titans: Learning to Memorize at Test Time
GenAI Level UP18 Jan 2025

Titans: Learning to Memorize at Test Time

Are current AI models hitting a memory wall? Join us as we delve into the fascinating research behind "Titans: Learning to Memorize at Test Time," an innovative approach to AI learning.

The podcast covers key concepts from the paper, including:

  • The challenges of long-term memory in AI, noting that models like Transformers are good at understanding immediate relationships but struggle with retaining information from the past.
  • How the Titan model addresses these limitations by equipping AI with both short-term and long-term memory.
  • The concept of "learning to memorize at test time", where the model figures out what is important to remember as it encounters new information.
  • The use of a surprise-based approach, where the model prioritizes information that is most surprising or unexpected.
  • The combination of surprise-based long-term memory with a more traditional short-term memory.
  • The way long-term memory is stored, which is within the parameters of a deep neural network.
  • The use of a technique similar to gradient descent with momentum for efficient memory formation.
  • The model's built-in forgetting mechanism to manage memory capacity and prioritize important information.
  • The use of attention to guide the search for relevant information in long-term memory.
  • The ability of Titans to handle longer sequences of information by using long-term memory to free up short-term memory.
  • The advantages of Titans in real-world applications such as language modeling, common sense reasoning, and the needle in a haystack problem.
  • The three variants of the Titan architecture: Memory as a Context (MAC), Memory as a Gate (MAG), and Memory as a Layer (MAL). Each variant uses long-term memory differently.



  • Denne episoden er hentet fra en åpen RSS-feed og er ikke publisert av Podme. Den kan derfor inneholde annonser.

    Episoder(45)

    Recursive Self Improvement

    Recursive Self Improvement

    Imagine holding a wrench on an assembly line. Suddenly, it leaps from your hand, sprouts its own mechanical arms, and begins forging a faster, lighter wrench without you. You are no longer the creator...

    7 Jun 1h

    Master the New Physics of AI with Context Graphs & GraphRAG

    Master the New Physics of AI with Context Graphs & GraphRAG

    Stop trying to find the "magic words" to hack your LLM. The era of the Prompt Engineer—tweaking adjectives and hoping for the best—is officially over. We are entering the age of the Context Engineer, ...

    1 Feb 17min

    Context Graph

    Context Graph

    Stop feeding your AI static facts in a dynamic world.Most RAG systems and Knowledge Graphs rely on a fundamental unit called the "Triple" (Subject, Verb, Object). It’s efficient, but it’s brittle. It ...

    25 Jan 19min

    Nested Learning: The Illusion of Deep Learning Architectures

    Nested Learning: The Illusion of Deep Learning Architectures

    Why do today's most powerful Large Language Models feel... frozen in time? Despite their vast knowledge, they suffer from a fundamental flaw: a form of digital amnesia that prevents them from truly le...

    14 Nov 202513min

    Memento: Fine-tuning LLM Agents without Fine-tuning LLMs

    Memento: Fine-tuning LLM Agents without Fine-tuning LLMs

    What if you could build AI agents that get smarter with every task, learning from successes and failures in real-time—without the astronomical cost and complexity of constant fine-tuning? This isn't a...

    1 Nov 202518min

    MemGPT: Towards LLMs as Operating Systems

    MemGPT: Towards LLMs as Operating Systems

    Have you ever felt the frustration of an LLM losing the plot mid-conversation, its brilliant insights vanishing like a dream? This "goldfish memory"—the limited context window—is the Achilles' heel of...

    1 Nov 202518min

    DeepSeek-OCR: Contexts Optical Compression

    DeepSeek-OCR: Contexts Optical Compression

    The single biggest bottleneck for Large Language Models isn't intelligence—it's cost. The quadratic scaling of self-attention makes processing truly long documents prohibitively expensive, a fundament...

    24 Okt 202513min

    A Definition of AGI

    A Definition of AGI

    For decades, Artificial General Intelligence has been a moving target, a nebulous concept that shifts every time a new AI masters a complex task. This ambiguity fuels unproductive debates and obscures...

    23 Okt 202519min

    Populært innen Teknologi

    lydartikler-fra-aftenposten
    teknisk-sett
    tomprat-med-gunnar-tjomlid
    rss-ki-praten
    shifter
    elektropodden
    rss-ai-forklart
    rss-alt-som-gar-pa-strom
    pedagogisk-intelligens
    nasjonal-sikkerhetsmyndighet-nsm
    smart-forklart
    rss-polypod
    rss-fish-ships
    fornybaren
    rss-bouvet-bobler
    rss-kunstig-intelligens-med-elisabeth-maren-og-morten
    hans-petter-og-co
    rss-fisketimen
    teknologi-og-mennesker
    rss-digitaliseringspadden