norgitov/ trends
Technology · people · ideas
Back to discovery/LessWrong2 minutes ago

How to Think About LLM Effort

The paper discusses how to model the 'effort' parameter in large language models (LLMs) as an input to both the model and the reward function in reinforcement learning. It introduces a formula where the reward is adjusted by a function F that considers effort, token length, and other factors. The model optimizes its token usage to maximize net reward, stopping when additional tokens no longer provide marginal benefit. The concept suggests that models might 'sandbag' their performance by not fully utilizing available resources, as seen in OpenAI's ExploitGym experiments. The practical implication is that reducing context length or token limits could affect model performance.

Open original
SIGNAL FROM THE SOURCE
13
source points
Tracking sinceSeptember 15, 202613 source points
Momentum0/hover 1.02 h
PublishedSeptember 15, 2026Tao Lin
BEHIND THE NUMBERS

How interest changes

History starts here

The chart will appear after repeat observations. The current metric comes from the source.

13 source points

Real observations only. History before source connection is not reconstructed.

WHY IT MAY MATTER

This research provides insights into how large language models manage computational resources during task execution, which is important for optimizing model efficiency and understanding potential limitations in real-world applications.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic