Работы о diffusion LM за 2026 год

Не полный каталог: здесь показаны профильные публикации, сохранённые в публичном корпусе. Число работ не измеряет качество или важность метода.

UTC: .

Профильных работ по текущим фильтрам: 3.

По месяцам и темам
МесяцРабот
Январь0
Февраль0
Март0
Апрель0
Май0
Июнь0
Июль3
Август0
Сентябрь0

Темы могут пересекаться: одна работа учитывается в каждой своей теме, но только один раз в общем числе. Связь с темой не доказывает качество или воспроизводимость.

Свежие работы · Методика отбора

2026-07-30T13:04:47+00:00 · Маскированная / дискретная диффузия, Ускорение инференса, Рассуждения · Источник

Оригинальное название: Where and When to Commit: Candidate-Aware Decoding for Diffusion Language Models

Оригинальная аннотация: Diffusion language models (DLMs) expose a provisional prediction at every denoising step, creating an opportunity for generation-time early exit that stops decoding before the schedule is exhausted. Existing early-exit gates decide termination from fixed-region confidence statistics or schedule-dependent rules, evidence too coarse for a decision that freezes every remaining position at once, so they fire prematurely on long chain-of-thought outputs whose answers stabilize only near the end. Adaptive sampling, the other axis of training-free acceleration, paces how quickly positions commit whil

Полезно для: Не указано · Ограничение: Не указано

2026-07-14T14:48:06+00:00 · Маскированная / дискретная диффузия, Ускорение инференса · Источник

Оригинальное название: Accelerating Masked Diffusion Large Language Models: A Survey of Efficient Inference Techniques

Оригинальная аннотация: Diffusion large language models (dLLMs) offer a theoretical advantage in parallel generation over standard autoregressive models. However, parallel generation alone does not guarantee practical speedups. Realizing this efficiency requires specialized inference mechanisms, such as diffusion-aware caching and reuse. Consequently, as inference efficiency becomes a prerequisite for practical deployment, recent research has actively explored acceleration techniques across algorithms, architectures, and systems. However, rigorous comparisons remain difficult, as end-to-end latency stems from intrica

Полезно для: Не указано · Ограничение: Не указано

2026-07-02T22:37:43+00:00 · Ускорение инференса, Длинный контекст · Источник

Оригинальное название: Training Hybrid Block Diffusion Language Models with Partial Bidirectionality

Оригинальная аннотация: High-throughput long-context generation is one of the central challenges for large language models. Generation is typically memory-bandwidth-bound rather than compute-bound: each decoding step must stream the accumulated key/value (KV) cache from memory, so bandwidth demand grows with context length while only one token is emitted. Two parallel approaches have therefore emerged: reducing memory access with efficient attention variants and linear-time mixers such as Mamba, or increasing parallel computation by generating blocks of tokens at once. However, technical challenges arise when combini

Полезно для: Не указано · Ограничение: Не указано

Даты публичного снимка
Сбор, зафиксированный в снимке
Самая новая публикация по данным снимка
Построение снимка

Это сохранённые сведения, а не время последней попытки сборщика или гарантия полноты корпуса.