Работы о diffusion LM за 2026 год

Не полный каталог: здесь показаны профильные публикации, сохранённые в публичном корпусе. Число работ не измеряет качество или важность метода.

UTC: .

Профильных работ по текущим фильтрам: 10.

По месяцам и темам
МесяцРабот
Январь0
Февраль0
Март0
Апрель0
Май0
Июнь0
Июль0
Август10
Сентябрь0

Темы могут пересекаться: одна работа учитывается в каждой своей теме, но только один раз в общем числе. Связь с темой не доказывает качество или воспроизводимость.

Свежие работы · Методика отбора

2026-08-31T23:50:31+00:00 · Маскированная / дискретная диффузия · Источник

Оригинальное название: Beyond Token Positions: Safety Alignment Across Denoising Steps in Diffusion Language Models

Оригинальная аннотация: Diffusion large language models (dLLMs) generate text through iterative denoising rather than left-to-right decoding. This generation paradigm introduces two axes that can influence safety alignment: when tokens are generated during denoising and where they appear in the response. In this paper, we measure dLLM safety behavior under harmful prompts by tracing intermediate token distributions and commitment decisions throughout denoising. Our analysis shows that refusal signals are concentrated in early denoising steps and leading response positions, and the tokens committed early can strongly

Полезно для: Не указано · Ограничение: Не указано

2026-08-31T15:00:30+00:00 · Маскированная / дискретная диффузия, Рассуждения, Генерация кода · Источник

Оригинальное название: CARVE: Verified Expansion for Variable-Length Generation in Diffusion Language Models

Оригинальная аннотация: Masked diffusion language models predict tokens from a partially observed response canvas, enabling bidirectional conditioning and parallel token refinement. Yet standard masked-diffusion decoders use a rigid inference interface: the number of masked positions allocated to the answer is fixed before generation begins. Choosing this length is difficult. A short canvas can truncate reasoning or code, while a long canvas wastes computation and can perturb denoising. We introduce CARVE (Counterfactual-Aware Reveal with Verified Expansion), a training-free variable-length algorithm for masked diffu

Полезно для: Не указано · Ограничение: Не указано

2026-08-20T14:52:17+00:00 · Маскированная / дискретная диффузия · Источник

Оригинальное название: Discrete Diffusion Inference-Time Control with Nested Sequential Monte Carlo

Оригинальная аннотация: We study inference-time control for text generation in discrete diffusion language models, where the goal is to steer sampling toward sequence-level rewards without retraining. Prior work in this domain has focused on particle-based methods such as best-of-$n$ sampling and bootstrap sequential Monte Carlo, which may suffer from overoptimism and weight degeneracy, respectively. We address these limitations using \emph{nested} sequential Monte Carlo methods. We formulate nested SMC (NSMC) and fully-adapted nested SMC (FA-NSMC) for Feynman--Kac steering, identifying and correcting errors in prior

Полезно для: Не указано · Ограничение: Не указано

2026-08-10T10:52:20+00:00 · Маскированная / дискретная диффузия · Источник

Оригинальное название: Reducing Pretraining-Generation Mismatch in Diffusion Language Models

Оригинальная аннотация: Autoregressive language models align training and use: generation conditions on a clean prompt, and training predicts future tokens from clean left context. Diffusion language models offer parallel denoising, but native dLLM pretraining can randomly corrupt prompt and continuation tokens together, weakening the clean-prefix interface needed for prompt-conditioned generation. We identify this mismatch for prompt continuation and propose PCD (Prefix-Conditioned Diffusion), a pretraining objective that combines AR prefix supervision with no-shift suffix denoising. At the training-objective level,

Полезно для: Не указано · Ограничение: Не указано

2026-08-08T11:59:38+00:00 · Маскированная / дискретная диффузия, Управляемость · Источник

Оригинальное название: Commitment Before Realization: When Classifier-Free Guidance Becomes Unnecessary in Masked Diffusion Language Models

Оригинальная аннотация: Classifier-free guidance (CFG) is usually kept on throughout masked diffusion language model decoding, although its benefit varies across prompts and over time. We study when CFG is actually needed by comparing, from any partial output, the probability of eventual constraint satisfaction under continued CFG and under base-only continuation. Their difference defines the remaining value of guidance. Guidance dependence is highly prompt-specific. Many prompts already succeed without CFG, while for others it provides no measurable benefit or can be harmful. For prompts that do benefit, the gain is

Полезно для: Не указано · Ограничение: Не указано

2026-08-06T22:34:20+00:00 · Маскированная / дискретная диффузия, Ускорение инференса, Дообучение · Источник

Оригинальное название: Retrofitting Linear Attention into Diffusion Language Models

Оригинальная аннотация: Diffusion language models (dLLMs) offer a promising alternative to autoregressive models by accelerating inference through parallel decoding. Recent dLLMs commonly use blockwise semi-autoregressive decoding, generating blocks autoregressively while denoising tokens within each active block in parallel. However, despite KV caching, each denoising step still attends to all previous blocks, repeatedly incurring prefix-attention cost. Motivated by this bottleneck, we ask whether dLLM inference can be further accelerated by linearizing attention over previous blocks. We introduce block-hybrid atten

Полезно для: Не указано · Ограничение: Не указано

2026-08-06T19:23:32+00:00 · Маскированная / дискретная диффузия · Источник

Оригинальное название: Lost in Interpolation: Why Predictive Feedback Fails in Diffusion Language Models

Оригинальная аннотация: Soft-masking accelerates the convergence of Masked Diffusion Language Models (MDLMs). Existing formulations build this blend with linear interpolation (LERP) in the raw embedding space, which implicitly treats that space as Euclidean. We analyze the embedding space of MDLMs and find that the mask and predicted-token embeddings maintain a near-constant angle of (\approx 73^\circ) throughout training, while embedding norms remain essentially flat across vocabulary-frequency rank. These indicate a hyperspherical geometry, for which LERP is the wrong interpolation primitive. We introduce Spherical

Полезно для: Не указано · Ограничение: Не указано

2026-08-04T14:54:14+00:00 · Маскированная / дискретная диффузия, Дообучение · Источник

Оригинальное название: MDLMPE: Distribution Aware Positional Encoding for Masked Diffusion Language Models

Оригинальная аннотация: Masked diffusion language models (MDLMs) enable parallel generation and bidirectional context modeling, but their positional context differs fundamentally from that of autoregressive (AR) models. Whereas AR decoding exposes a contiguous prefix, MDLM denoising produces dynamic, non-contiguous configurations of revealed and masked tokens. Conventional positional encodings such as RoPE capture sequence order and pairwise displacement but remain insensitive to this evolving token-availability structure. To address this limitation, we propose MDLMPE, a positional encoding designed specifically for

Полезно для: Не указано · Ограничение: Не указано

2026-08-04T10:53:02+00:00 · Маскированная / дискретная диффузия, Рассуждения, Дообучение · Источник

Оригинальное название: LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models

Оригинальная аннотация: Diffusion language models (dLLMs) offer an alternative to autoregressive (AR) language modeling, yet the scaling behavior of Mixture-of-Experts (MoE) dLLMs remains poorly understood. We systematically characterize how optimization hyperparameters, compute allocation, and architecture scale for MoE dLLMs, identifying quantitative differences from scaling trends previously reported for AR models. Specifically, for optimization, the optimal nominal batch size grows faster, while the optimal learning rate decays more rapidly with compute. For model--data allocation, IsoFLOP analysis reveals a slig

Полезно для: Не указано · Ограничение: Не указано

2026-08-03T06:59:06+00:00 · Маскированная / дискретная диффузия · Источник

Оригинальное название: REFLEX: Rethinking MoE Inference as Refinement-Aware Compute Allocation in Diffusion Language Models

Оригинальная аннотация: Mixture-of-experts (MoE) models increase parameter capacity by activating only a small subset of experts for each token. This conditional-computation paradigm has enabled autoregressive language models to scale model capacity without a proportional increase in per-token computation. In diffusion language models (DLMs), however, each denoising forward jointly revisits all token positions despite their sharply different refinement demands, while the default fixed token-choice routing assigns them a uniform expert budget, creating a mismatch between expert computation and refinement demand. We ar

Полезно для: Не указано · Ограничение: Не указано

Даты публичного снимка
Сбор, зафиксированный в снимке
Самая новая публикация по данным снимка
Построение снимка

Это сохранённые сведения, а не время последней попытки сборщика или гарантия полноты корпуса.