norgitov/ trends
Technology · people · ideas
Back to discovery/Daily Papers1 hour ago

WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents

WhatWorkedBench is a benchmark designed to evaluate how well AI agents understand the impact of experimental changes. It measures predictions about component changes after budgeted experimentation, using a response surface that predicts scores for different configuration settings. The benchmark includes data from 36 tasks, 30 data sources, and 8 workflow types, with 1,248 configuration records. Evaluation involves 4,206 numerical-control records and 108 agent episodes, showing improvements in effect recovery when using a Gaussian process (GP) model.

Open original
SIGNAL FROM THE SOURCE
6
source votes
Tracking sinceSeptember 24, 20266 source votes
MomentumMore observations needed
DiscussionRead comments ↗
PublishedSeptember 23, 2026Jingjie Ning, Xueqi Li, Yibo Kong
BEHIND THE NUMBERS

How interest changes

History starts here

The chart will appear after repeat observations. The current metric comes from the source.

6 source votes

Real observations only. History before source connection is not reconstructed.

WHY IT MAY MATTER

This benchmark helps researchers evaluate how well AI agents can predict the outcomes of experimental changes, improving their ability to design and interpret experiments effectively.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic