norgitov/ trends
Technology · people · ideas
Back to discovery/arXiv52 minutes ago

FD-VAD: Semantic Endpoint Detection for Streaming Full-Duplex Speech

The paper introduces FD-VAD, an ASR-free streaming endpointer that uses causal audio-language reasoning to determine if a speech pause indicates hesitation or completion. It combines a frozen speech encoder with a lightweight modality adapter and a parameter-efficient language model, achieving high EOT recall in zero-shot settings.

Open original
SIGNAL FROM THE SOURCE
30
September 2026
Tracking sinceSeptember 30, 2026
Momentum—More observations needed
Discussion—No comment count provided
PublishedSeptember 30, 2026Puneet Mathur, Dinesh Manocha
BEHIND THE NUMBERS

How interest changes

The publication is the signal

This official source does not publish popularity metrics. The story is refreshed from its RSS feed.

Real observations only. History before source connection is not reconstructed.

WHY IT MAY MATTER

This approach enables real-time semantic endpoint detection without relying on ASR, improving accuracy in conversational settings.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic