norgitov/ trends
Technology · people · ideas
Back to discovery/Daily Papers2 hours ago

FLAT: Resampling Image and Text into 1D Flexible-Length Aligned Transmodal Tokens for Retrieval and Generation

FLAT is a pre-training framework that jointly optimizes a shared multimodal encoder with text-to-image and image-to-text decoders, enabling flexible-length aligned transmodal representations for retrieval and generation. It maps visual and textual inputs into a unified 1D sequence space and demonstrates strong performance on various tasks.

Open original
SIGNAL FROM THE SOURCE
12
source votes
Tracking sinceSeptember 16, 202612 source votes
MomentumMore observations needed
DiscussionRead comments ↗
PublishedSeptember 15, 2026Guangyu Sun, Shlok Kumar Mishra, Wentao Bao
BEHIND THE NUMBERS

How interest changes

History starts here

The chart will appear after repeat observations. The current metric comes from the source.

12 source votes

Real observations only. History before source connection is not reconstructed.

WHY IT MAY MATTER

FLAT enables flexible-length cross-modal generation and retrieval by aligning visual and textual representations in a unified 1D sequence space, supporting tasks like linear interpolation and zero-shot composed retrieval.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic