norgitov/ trends
Technology · people · ideas
Back to discovery/LessWrong41 minutes ago

Why Do Models Really Fail on HLE Tasks?

The study investigates why large language models fail on the Humanity's Last Exam (HLE) tasks by analyzing failure modes using the Inspect Scout framework. Researchers tested three GPT models on a subset of HLE tasks, identifying failure reasons and examining how modifying the task harness affects failure prevalence.

Open original
SIGNAL FROM THE SOURCE
7
source points
Tracking sinceSeptember 28, 20267 source points
Momentum0/hover 1.02 h
PublishedSeptember 28, 2026Ana Leonescu
BEHIND THE NUMBERS

How interest changes

History starts here

The chart will appear after repeat observations. The current metric comes from the source.

7 source points

Real observations only. History before source connection is not reconstructed.

WHY IT MAY MATTER

This research provides insights into failure modes of language models on complex tasks, highlighting the importance of task design and evaluation frameworks.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic