Published event
Hardware
ModelRelease
1 source(s)
v0.26.0
Summary
v0.26.0 vllm-project / vllm Public Uh oh! There was an error while loading.
Why it matters
This ModelRelease is relevant to the technology intelligence record because it involves DeepSeek, AMD, Mistral AI, OpenAI. The source article should remain the factual reference for follow-up coverage.
Key facts
- vllm-project / vllm Public Uh oh!
- There was an error while loading.
- Notifications You must be signed in to change notification settings Fork 22.7k Star 92.7k v0.26.0 khluu released this 27 Jul 01:06 · 3112 commits to main since this release v0.26.0 568afb3 vLLM v0.26.0 Release Notes Highlights This release features 411 commits from 212 contributors (61 new)!
- New Inkling model family with a full support stack: base modeling ( #48799 ), piecewise CUDA graph support ( #48822 ), Hopper FA4 relative attention ( #48858 ), MTP=1 speculative decoding ( #48869 ), LoRA ( #48884 ), and standard ModelOpt NVFP4 quantization ( #48990 ).
- DeepSeek-V4 performance push across vendors: a specialized routing kernel (2.94% E2E TPOT, #48660 ), fused_topk_bias (1.5–2x kernel, #47463 ), and redundant repeat/copy removal (1.8% E2E TPOT, #48137 ), plus ROCm two-stage compressor for HCA prefill ( #47718 ), sparse decode/prefill optimizations ( #48519 , #48788 , #46275 ), and DSpark speculative decoding on AMD ( #47419 ) and XPU ( #47677 ).
- fp32 lm_head for generation models via head_dtype ( #48390 ), extended to the LoRA path ( #48525 ) and given a ROCm torch.mm fast path ( #48688 ), improving accuracy for generation heads.
Entities in this story
Related events