norgitov/ trends
Technology · people · ideas
Back to discovery/Daily Papers2 hours ago

Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs

The paper introduces Flash-dLLM, a framework for accelerating diffusion large language models (dLLMs) by addressing I/O bottlenecks in key-value (KV) caching and enabling efficient parallel decoding. It proposes an I/O-aware fused KV-cache kernel and a draft-and-verify decoding strategy that uses the dLLM itself for both tasks, improving speed and memory efficiency.

Open original
SIGNAL FROM THE SOURCE
3
source votes
Tracking sinceSeptember 23, 20263 source votes
MomentumMore observations needed
DiscussionRead comments ↗
PublishedSeptember 22, 2026Quan Nguyen-Tri, Mukul Ranjan, Zhiqiang Shen
BEHIND THE NUMBERS

How interest changes

History starts here

The chart will appear after repeat observations. The current metric comes from the source.

3 source votes

Real observations only. History before source connection is not reconstructed.

WHY IT MAY MATTER

Flash-dLLM improves inference speed and memory efficiency for diffusion LLMs by optimizing I/O in KV caching and enabling parallel decoding without additional models.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic