Evaluating LLMs with Chatbot Arena and Joseph E. Gonzalez

Evaluating LLMs with Chatbot Arena and Joseph E. Gonzalez

In this episode of Gradient Dissent, Joseph E. Gonzalez, EECS Professor at UC Berkeley and Co-Founder at RunLLM, joins host Lukas Biewald to explore innovative approaches to evaluating LLMs.

They discuss the concept of vibes-based evaluation, which examines not just accuracy but also the style and tone of model responses, and how Chatbot Arena has become a community-driven benchmark for open-source and commercial LLMs. Joseph shares insights on democratizing model evaluation, refining AI-human interactions, and leveraging human preferences to improve model performance. This episode provides a deep dive into the evolving landscape of LLM evaluation and its impact on AI development.

🎙 Get our podcasts on these platforms:

Apple Podcasts: http://wandb.me/apple-podcasts

Spotify: http://wandb.me/spotify

Google: http://wandb.me/gd_google

YouTube: http://wandb.me/youtube

Follow Weights & Biases:

https://twitter.com/weights_biases

https://www.linkedin.com/company/wandb

Join the Weights & Biases Discord Server:

https://discord.gg/CkZKRNnaf3

Det här avsnittet är hämtat från ett öppet RSS-flöde och publiceras inte av Podme. Det kan innehålla reklam.

Avsnitt(136)

Aaron Colak — ML and NLP in Experience Management

Aaron Colak — ML and NLP in Experience Management

Aaron Colak is the Leader of Core Machine Learning at Qualtrics, an experiment management company that takes large language models and applies them to real-world, B2B use cases.In this episode, Aaron ...

26 Aug 202250min

Jordan Fisher — Skipping the Line with Autonomous Checkout

Jordan Fisher — Skipping the Line with Autonomous Checkout

Jordan Fisher is the CEO and co-founder of Standard AI, an autonomous checkout company that’s pushing the boundaries of computer vision.In this episode, Jordan discusses “the Wild West” of the MLOps s...

4 Aug 202257min

Drago Anguelov — Robustness, Safety, and Scalability at Waymo

Drago Anguelov — Robustness, Safety, and Scalability at Waymo

Drago Anguelov is a Distinguished Scientist and Head of Research at Waymo, an autonomous driving technology company and subsidiary of Alphabet Inc.We begin by discussing Drago's work on the original I...

14 Juli 20221h 9min

James Cham — Investing in the Intersection of Business and Technology

James Cham — Investing in the Intersection of Business and Technology

James Cham is a co-founder and partner at Bloomberg Beta, an early-stage venture firm that invests in machine learning and the future of work, the intersection between business and technology.James ex...

7 Juli 20221h 6min

Boris Dayma — The Story Behind DALL·E mini, the Viral Phenomenon

Boris Dayma — The Story Behind DALL·E mini, the Viral Phenomenon

Check out this report by Boris about DALL-E mini:https://wandb.ai/dalle-mini/dalle-mini/reports/DALL-E-mini-Generate-images-from-any-text-prompt--VmlldzoyMDE4NDAyhttps://wandb.ai/_scott/wandb_example/...

17 Juni 202235min

Tristan Handy — The Work Behind the Data Work

Tristan Handy — The Work Behind the Data Work

Tristan Handy is CEO and founder of dbt Labs. dbt (data build tool) simplifies the data transformation workflow and helps organizations make better decisions.Lukas and Tristan dive into the history of...

9 Juni 20221h

Johannes Otterbach — Unlocking ML for Traditional Companies

Johannes Otterbach — Unlocking ML for Traditional Companies

Johannes Otterbach is VP of Machine Learning Research at Merantix Momentum, an ML consulting studio that helps their clients build AI solutions.Johannes and Lukas talk about Johannes' background in ph...

12 Maj 202244min

Mircea Neagovici — Robotic Process Automation (RPA) and ML

Mircea Neagovici — Robotic Process Automation (RPA) and ML

Mircea Neagovici is VP, AI and Research at UiPath, where his team works on task mining and other ways of combining robotic process automation (RPA) with machine learning for their B2B products.Mircea ...

21 Apr 202246min

Populärt inom Business & ekonomi

framgangspodden
varvet
rss-jossan-nina
badfluence
rss-svart-marknad
rss-borsens-finest
svd-tech-brief
avanzapodden
uppgang-och-fall
bathina-en-podcast
fill-or-kill
kapitalet-en-podd-om-ekonomi
dynastin
tabberaset
rss-dagen-med-di
lastbilspodden
rss-inga-dumma-fragor-om-pengar
rikatillsammans-om-privatekonomi-rikedom-i-livet
rss-kort-lang-analyspodden-fran-di
rss-dr-bjorklund