norgitov/ trends
Technology · people · ideas
Back to discovery/Daily Papers53 minutes ago

One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts

This paper introduces reViT, a recurrent vision transformer architecture that uses a single block applied multiple times to achieve performance comparable to deeper models with similar computational costs. The method employs a shared expert bank for feed-forward networks, with depth-specific transformations controlled by a normalized-depth coordinate. It demonstrates effectiveness in both supervised training and distillation from a DINOv2 teacher, achieving high accuracy with fewer parameters and flexible depth adaptation.

Open original
SIGNAL FROM THE SOURCE
2
source votes
Tracking sinceOctober 9, 20262 source votes
Momentum—More observations needed
Discussion—Read comments ↗
PublishedOctober 8, 2026Adrian Bulat, Yassine Ouali, Georgios Tzimiropoulos
BEHIND THE NUMBERS

How interest changes

History starts here

The chart will appear after repeat observations. The current metric comes from the source.

2 source votes

Real observations only. History before source connection is not reconstructed.

WHY IT MAY MATTER

This approach enables efficient model deployment by allowing a single model to operate at multiple depths without retraining, reducing storage requirements and computational overhead.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic