norgitov/ trends
Technology · people · ideas
Back to discovery/Daily Papers2 hours ago

A Zeroth-Order Paradigm for LLM Preference Alignment

This paper introduces Comparison-based Preference Optimization (ComPO), a zeroth-order method for aligning large language models with human preferences by using comparison oracles to extract directional information from preference pairs without direct optimization of a differentiable loss. The method includes offline and online variants with convergence and performance guarantees, showing improvements over existing alignment methods in experiments with various models.

Open original
SIGNAL FROM THE SOURCE
21
source votes
Tracking sinceSeptember 17, 202621 source votes
MomentumMore observations needed
DiscussionRead comments ↗
PublishedSeptember 16, 2026Peter Chen, Xi Chen, Wotao Yin
BEHIND THE NUMBERS

How interest changes

History starts here

The chart will appear after repeat observations. The current metric comes from the source.

21 source votes

Real observations only. History before source connection is not reconstructed.

WHY IT MAY MATTER

This method offers a new approach to aligning language models with human preferences without relying on direct optimization of a differentiable loss function, which could be useful in scenarios where such optimization is challenging or impractical.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic