What do we actually gain from using an LLM?Preprint: https://arxiv.org/abs/2609.16793 In our new preprint, 535 people solved 40 reasoning problems, alone or with GPT-5.6-Luna, Claude Opus 4.8, Gemini 3.6 Flash, or Kimi K3. We also tested each model on every problem 100 times to see how reliably it answered. People benefited more on problems the… Continue reading Better AI is only part of the story