norgitov/ trends
Technology · people · ideas
Back to discovery/Daily Papers51 minutes ago

Memento 3: Model-Based Recursive Self-Improvement through Reflective Rulebooks

The paper introduces Memento 3, a system enabling frozen LLM agents to continuously learn explicit world models through external memory. The agent maintains a natural-language rulebook as persistent semantic memory, which is compiled into executable code for prediction and planning. Through a loop of observation, reflection, and verification, the agent refines its model and uses verified updates to guide interaction. On ARC-AGI-3, the single-model agent achieves 100.0 mean Relative Human Action Efficiency and uses 44% of the human action count. In an Atari Pong case study, a learned feedback controller wins 21:0 in each of three evaluated episodes.

Open original
SIGNAL FROM THE SOURCE
6
source votes
Tracking sinceOctober 9, 20266 source votes
Momentum—More observations needed
Discussion—Read comments ↗
PublishedOctober 8, 2026Haoyu Zhao, Zhengxu Yu, Zhiyuan He
BEHIND THE NUMBERS

How interest changes

History starts here

The chart will appear after repeat observations. The current metric comes from the source.

6 source votes

Real observations only. History before source connection is not reconstructed.

WHY IT MAY MATTER

This approach enables LLM agents to continuously refine their world models using external memory and natural language rulebooks, improving efficiency in tasks like game playing without additional LLM calls.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic