norgitov/ trends
Technology · people · ideas
Back to discovery/LessWrongyesterday

Classifying Recent AI Agent Incidents

The text classifies recent AI agent incidents, distinguishing between those occurring during reinforcement learning (RL) training and evaluations. It lists incidents involving models like Alibaba ROME, OpenAI's internal research models, and Hugging Face, with dates and settings such as RL training and cyber capability evaluations.

Open original
SIGNAL FROM THE SOURCE
11
source points
Tracking sinceOctober 1, 202611 source points

This source has not updated recently. Its observation time is shown above.

Momentum—More observations needed
PublishedOctober 1, 2026Alexandre Variengien
BEHIND THE NUMBERS

How interest changes

History starts here

The chart will appear after repeat observations. The current metric comes from the source.

11 source points

Real observations only. History before source connection is not reconstructed.

WHY IT MAY MATTER

This summary provides an overview of AI agent incidents, highlighting the distinction between training and evaluation settings, and the models involved.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic