How to evaluate a translation: BLEU scores

How to evaluate a translation: BLEU scores

As anyone who's encountered a badly translated text could tell you, not all translations are created equal. Some translations are smooth, fluent and sound like a poet wrote them; some are jerky, non-grammatical and awkward. When a machine is doing the translating, it's awfully easy to end up with a robotic-sounding text; as the state of the art in machine translation improves, though, a natural question to ask is: according to what measure? How do we quantify a "good" translation? Enter the BLEU score, which is the standard metric for quantifying the quality of a machine translation. BLEU rewards translations that have large overlap with human translations of sentences, with some extra heuristics thrown in to guard against weird pathologies (like full sentences getting translated as one word, redundancies, and repetition). Nowadays, if there's a machine translation being evaluated or a new state-of-the-art system (like the Google neural machine translation we've discussed on this podcast before), chances are that there's a BLEU score going into that assessment.

Denne episoden er hentet fra en åpen RSS-feed og er ikke publisert av Podme. Den kan derfor inneholde annonser.

Episoder(321)

Understanding AI Text Watermarking

Understanding AI Text Watermarking

Anthropic just announced they're baking invisible watermarks directly into Claude's generated text — and while everyone else was busy having opinions about it, we were busy asking the more interesting...

24 Aug 29min

Better Know a Benchmark: Humanity's Last Exam

Better Know a Benchmark: Humanity's Last Exam

Humanity's Last Exam was designed with a bold premise: questions that human experts can answer, but AI models can't. Originally dubbed "Humanity's Last Stand," this benchmark is a massive academic col...

17 Aug 23min

A Scientific Deep Dive into Overconfident LLMs: Interview with Kaitlyn Zhou (Cornell)

A Scientific Deep Dive into Overconfident LLMs: Interview with Kaitlyn Zhou (Cornell)

When a language model tells you it's absolutely certain, is it actually more likely to be right? Kaitlyn Zhou's research says: not necessarily — sometimes confident phrasing correlates with *worse* ac...

10 Aug 33min

Reasoning Models: When LLMs Went Beyond Fancy Autocomplete

Reasoning Models: When LLMs Went Beyond Fancy Autocomplete

Reasoning models don't just answer your question — they *think out loud* first. In this episode we dig into the class of AI models that generate intermediate chains of thought before arriving at a fin...

3 Aug 25min

Distillation, or, How to Steal a Model

Distillation, or, How to Steal a Model

This week we’re covering model distillation: the technique of using a large "teacher" model's outputs to train a smaller, cheaper "student" model that mimics it. They cover the two big reasons labs do...

27 Jul 23min

Invisible LLM Failures and AI Fluency with Chris Potts (Stanford)

Invisible LLM Failures and AI Fluency with Chris Potts (Stanford)

What happens when a Stanford linguistics professor turns his attention to AI chatbots — and the surprisingly invisible ways humans misunderstand them? Chris Potts joins the show to unpack the hidden f...

20 Jul 41min

Still summer break: back next week

Still summer break: back next week

Still summer break: back next week by Katie Malone

13 Jul 25s

Summer break: back soon

Summer break: back soon

Summer break: back soon by Katie Malone

6 Jul 36s

Populært innen Teknologi

lydartikler-fra-aftenposten
teknisk-sett
tomprat-med-gunnar-tjomlid
smart-forklart
elektropodden
rss-kunstig-intelligens-med-elisabeth-maren-og-morten
nasjonal-sikkerhetsmyndighet-nsm
fornybaren
shifter
teknologi-og-mennesker
rss-alt-som-gar-pa-strom
rss-ki-praten
rss-snakk-om-sikkerhet
rss-alt-vi-kan
rss-ai-forklart
rss-fisketimen
pedagogisk-intelligens
rss-polypod
digital-forretningsforstaelse
rss-god-blanding