$75k Contest Launch! ChinaTalk Hiring + Evals for the Situation Room
ChinaTalk10 Aug

$75k Contest Launch! ChinaTalk Hiring + Evals for the Situation Room

Contest links: General $50k https://www.chinatalk.media/p/50k-chinatalk-submission-hiring-contest we're looking at submissions on a rolling basis! AI evals: Due Over the past few years, we’ve seen hints of policymakers and national leaders using AI models in their actual policy decision-making. The Prime Minister of Sweden said he uses it for second opinions on policy; the German Chancellor is testing it to draft legislation; even Trump said he had it re-write a speech at some point. It’s fair to assume that senior leadership across the world, and in Washington as well, have started using AI not just for tactical or operational tasks, but increasingly for broad strategic decision-making. While that’s exciting — there’s the promise of uplift and smarter calls on some of the most consequential decisions leaders face in foreign policy and national security — we’re also flying blind. Enormous effort and energy goes into benchmarking and evaluation for things like coding. The experiments you can run to make models better at software development are much easier to execute and much lower-stakes than running a real experiment when you’re deciding whether to invade a country or sign a treaty. That’s why we at ChinaTalk are trying to kickstart a field aimed at helping researchers and policymakers understand exactly what they’re working with when they ask these models to support some of the most consequential decisions they may make in their lifetimes. We’re launching an evals/essay project contest to explore this theme. I’ve brought on two expert AI eval creators to discuss why the field is important, what interesting work has already been done on how models approach broad national-security and strategic questions, and how you — as an eval professional, semi-professional, or just a concerned person — can contribute new ways to poke and prod at these models and see what they can really do. Joining us today: Florian Brand, research engineer at Prime Intellect, and John Chen, professor at the University of Arizona, who’s done some pretty wild things getting models to start nuclear wars with each other in Civ V. Our conversation covers: How frontier labs are hitting eval-building limits — they’ve gone from undergrads to PhDs to field experts, and the models are now catching the experts’ mistakes. How Civilization V exposes AI’s strategic blind spots (terrible second-order reasoning) and models’ distinct strategic personalities (Claude’s really into science!). Why ethical prompting in Civilization V still doesn’t stop models from launching nukes. Jordan's PresidentBench eval, where a Chinese model was nonchalant about a Taiwan invasion, while Claude wanted to keep Taiwan free and independent. Advice for designing better AI evals and ChinaTalk’s new essay/evals contest! Learn more about your ad choices. Visit megaphone.fm/adchoices

Det här avsnittet är hämtat från ett öppet RSS-flöde och publiceras inte av Podme. Det kan innehålla reklam.

Avsnitt(573)

Data is the Hard Part | Building AI Tools for the Military

Data is the Hard Part | Building AI Tools for the Military

Everyone wants AI that wins wars. Very few people want to collect, label, store, and test the data it runs on. Bharat Patel is currently Accenture’s AI and data lead for its defense portfolio, a form...

6 Okt 42min

WarTalk: AutoWarCom + Project Meridian

WarTalk: AutoWarCom + Project Meridian

It's a new fiscal year with no new budget, and the Pentagon has a fresh batch of proposals. Two stand out: a four-star Autonomous Warfare Command, built to buy and field autonomous weapons fast, and P...

5 Okt 55min

Section 301 and the Spice Man who Sued Trump Over Them

Section 301 and the Spice Man who Sued Trump Over Them

In February, the Supreme Court ruled Trump's IEEPA tariffs illegal. So why is everyone still paying them? Because the administration rebuilt them: first under Section 122, then under Section 301, and ...

2 Okt 1h 5min

China's Economy is Broken | Logan Wright

China's Economy is Broken | Logan Wright

After the global financial crisis, China added a third of global GDP to its bank assets in just eight years... the largest single-country credit expansion in at least a century. Today investment is ...

29 Sep 1h 28min

WarTalk: Potemkin Pacific, Midterm Budget Woes

WarTalk: Potemkin Pacific, Midterm Budget Woes

WarTalk regulars Bryan Clark of the Hudson Institute, Justin McIntosh, and Tony Stark join Jordan. We discuss… The Potemkin Pacific — why Paparo has half the planet on his map and nothing to figh...

25 Sep 1h 12min

Julian Gewirtz on Trump-Xi Summit, AI, and Mao Dead Fifty Years Ago!

Julian Gewirtz on Trump-Xi Summit, AI, and Mao Dead Fifty Years Ago!

Julian Gewirtz, author and former Biden NSC China director, joins ChinaTalk days before Xi Jinping’s first White House visit since 2015. With little substantive agenda on the table, what is the US get...

22 Sep 1h 6min

Can AI Beat the Pentagon’s Bureaucracy?

Can AI Beat the Pentagon’s Bureaucracy?

Garrett Berntsen was last seen on ChinaTalk in his role as Deputy CDAO at the State Department. He then went to the Department of Defense for a little bit, and now makes a return in his new role as Ch...

22 Sep 49min

WarTalk: PLA, 2027, Nukes, and the Pivot to Asia Pirouette

WarTalk: PLA, 2027, Nukes, and the Pivot to Asia Pirouette

In the deserts of western China sit 320 brand-new missile silos, and for a while, by the Pentagon's account, some of them simply didn't work. The gap between the PLA on paper and the PLA in practice i...

19 Sep 1h 17min

Populärt inom Politik & nyheter

svenska-fall
fordomspodden
aftonbladet-krim
rss-krimstad
p3-krim
aftonbladet-daily
flashback-forever
spar
svd-dokumentara-berattelser-2
motiv
rss-sanning-konsekvens
de-fyras-gang
rss-vad-fan-hande
rss-krimreportrarna
svd-ledarredaktionen
rss-flodet
omni-podd
en-runda-till
rss-frandfors-horna
rss-politikrummet