Published event
Hardware
ModelRelease
1 source(s)
v0.12.0
Summary
v0.12.0 vllm-project / vllm Public Uh oh! There was an error while loading.
Why it matters
This ModelRelease is relevant to the technology intelligence record because it involves AMD, DeepSeek, Mistral AI, NVIDIA. The source article should remain the factual reference for follow-up coverage.
Key facts
- vllm-project / vllm Public Uh oh!
- There was an error while loading.
- Notifications You must be signed in to change notification settings Fork 22.7k Star 92.7k v0.12.0 khluu released this 03 Dec 09:36 · 10157 commits to main since this release v0.12.0 4fd9d6a vLLM v0.12.0 Release Notes Highlights Highlights This release features 474 commits from 213 contributors (57 new)! Breaking Changes : This release includes PyTorch 2.9.0 upgrade (CUDA 12.9), V0 deprecations including xformers backend, and scheduled removals - please review the changelog carefully.
- Major Features : EAGLE Speculative Decoding Improvements : Multi-step CUDA graph support ( #29559 ), DP>1 support ( #26086 ), and multimodal support with Qwen3VL ( #29594 ).
- Significant Performance Optimizations : 18.1% throughput improvement from batch invariant BMM ( #29345 ), 2.2% throughput improvement from shared experts overlap ( #28879 ).
- AMD ROCm Expansion : DeepSeek v3.2 + SparseMLA support ( #26670 ), FP8 MLA decode ( #28032 ), AITER attention backend ( #28701 ).
Entities in this story
Related events