Semiconductor News & Analysis Feed

2 articles
2026-09-03
developer.nvidia.com 2026-09-03
NVIDIA's latest developer article explores the use of speculative decoding to accelerate large language model (LLM) inference. This technique leverages a small draft model to predict multiple tokens, which are then verified in parallel by a larger target model, reducing total decoding iterations. Th
2026-06-23
developer.nvidia.com 2026-06-23
NVIDIA has introduced DFlash speculative decoding on its Blackwell platform, achieving up to 15x performance improvements in large language model (LLM) inference. By replacing traditional autoregressive draft models with a lightweight block-diffusion drafter, DFlash enables parallel block-level toke