Published event
ArtificialIntelligence
Other
1 source(s)
v0.20.2
Summary
v0.20.2 vllm-project / vllm Public Uh oh! There was an error while loading.
Why it matters
This Other is relevant to the technology intelligence record because it involves DeepSeek, DeepSeek v4, gpt-oss, GPT. The source article should remain the factual reference for follow-up coverage.
Key facts
- vllm-project / vllm Public Uh oh!
- There was an error while loading.
- Notifications You must be signed in to change notification settings Fork 22.7k Star 92.7k v0.20.2 khluu released this 10 May 07:37 · 5903 commits to main since this release v0.20.2 bc150f5 vLLM v0.20.2 Highlights This release features 6 commits from 6 contributors (0 new)!
- This is a small patch release with bug fixes for DeepSeek V4, gpt-oss, and Qwen3-VL Bug Fixes DeepSeek V4 sparse attention : Re-enable the persistent topk path on Hopper and ensure the memset kernel runs at CUDA graph capture time regardless of max_seq_len , fixing the MTP=1 hang on DeepSeek V4 ( #41665 , revert of #41605 ).
- DeepSeek V4 KV cache : Fixed a "failure to allocate KV blocks" error in the V1 engine KV cache manager ( #41282 ).
- gpt-oss MXFP4 + torch.compile : Plumbed hidden_dim_unpadded through the moe_forward fake op so MXFP4 works under torch.compile on v0.20.x ( #42002 , backport of #41646 ).
Entities in this story
Related events