norgitov/ trends
Technology · people · ideas
Back to discovery/Daily Papers3 hours ago

MIMESIS: Learning User Simulators as Training Environments for Interactive Agents

The paper introduces MIMESIS, a user simulator trained on human conversations with explicit reasoning supervision and 13 realistic behavioral patterns. It achieves a SOUL-Index of 65.7, outperforming existing models. The simulator is used to train agents through multi-turn reinforcement learning, showing better generalization than GPT-5.5. The paper also proposes Coached On-Policy Self-Distillation (CSD) to improve agent performance using simulator-generated feedback.

Open original
SIGNAL FROM THE SOURCE
9
source votes
Tracking sinceOctober 8, 20269 source votes
Momentum—More observations needed
Discussion—Read comments ↗
PublishedOctober 7, 2026Hoang Phan, Dat Huynh, Andrey Zhmoginov
BEHIND THE NUMBERS

How interest changes

History starts here

The chart will appear after repeat observations. The current metric comes from the source.

9 source votes

Real observations only. History before source connection is not reconstructed.

WHY IT MAY MATTER

MIMESIS provides a scalable and realistic alternative to human user interactions for training and evaluating interactive agents, improving behavioral fidelity and generalization. The Coached On-Policy Self-Distillation method enhances agent performance through detailed feedback from the simulator.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic