developer.nvidia.com
2026-06-23
NVIDIA has introduced DFlash speculative decoding on its Blackwell platform, achieving up to 15x performance improvements in large language model (LLM) inference. By replacing traditional autoregressive draft models with a lightweight block-diffusion drafter, DFlash enables parallel block-level toke