norgitov/ trends
Technology · people · ideas
Back to discovery/LessWrong17 minutes ago

CommentBench: Can Models Match Human Comments on AI Safety Posts?

The study evaluates how well AI models generate comments that match human comments on AI safety documents. It introduces a pipeline to measure the alignment between model-generated and human comments, finding that Fable 5 and Fable 5.1 perform best, matching 8.3% and 7.5% of human points, respectively. Performance across models is highly correlated across different document types.

Open original
SIGNAL FROM THE SOURCE
13
source points
Tracking sinceSeptember 19, 202613 source points
Momentum+2.95/hover 1.02 h
PublishedSeptember 19, 2026OscarGilg
BEHIND THE NUMBERS

How interest changes

History starts here

The chart will appear after repeat observations. The current metric comes from the source.

13 source points

Real observations only. History before source connection is not reconstructed.

WHY IT MAY MATTER

This work helps assess AI models' ability to generate relevant and aligned comments on safety-related content, which is crucial for ensuring effective collaboration between AI and human experts in safety research.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic