
Investigating the Role of Prompting and External Tools in Hallucination Rates of LLMs
🔎 Investigating the Role of Prompting and External Tools in Hallucination Rates of Large Language ModelsThis paper examines the effectiveness of different prompting techniques and frameworks for miti...
3 Nov 202416min

Mind Your Step (by Step)
🌀 Mind Your Step (by Step): Chain-of-Thought can Reduce Performance on Tasks where Thinking Makes Humans WorseThis research paper examines how chain-of-thought (CoT) prompting—encouraging models to r...
2 Nov 202416min

SimpleQA
❓Measuring short-form factuality in large language modelsThis document introduces SimpleQA, a new benchmark for evaluating the factuality of large language models. The benchmark consists of over 4,000...
31 Okt 202417min

GPT-4o System Card
📜 GPT-4o System CardThis technical document is the System Card for OpenAI's GPT-4o, a multimodal, autoregressive language model that can process and generate text, audio, images, and video. The card ...
30 Okt 202424min

Mixture of Parrots
🦜 Mixture of Parrots: Experts improve memorization more than reasoningThis research paper investigates the effectiveness of Mixture-of-Experts (MoE) architectures in deep learning, particularly compa...
29 Okt 202410min

Improve Vision Language Model Chain-of-thought Reasoning
🖼 Improve Vision Language Model Chain-of-thought ReasoningThis research paper investigates how to improve the chain-of-thought (CoT) reasoning capabilities of vision language models (VLMs). The autho...
28 Okt 202415min

Breaking the Memory Barrier
🧠 Breaking the Memory Barrier: Near Infinite Batch Size Scaling for Contrastive LossThis research paper introduces Inf-CL, a novel approach for contrastive learning that dramatically reduces GPU memo...
27 Okt 202415min

LLMs Reflect the Ideology of their Creators
⚖️ Large Language Models Reflect the Ideology of their CreatorsThis study examines the ideological stances of large language models (LLMs) by analyzing their responses to prompts about a vast set of h...
26 Okt 202411min



















