Published event
ArtificialIntelligence Funding 1 source(s)

Release 5.8.0

Updated September 26, 2026 · 2:47 PM · source date May 5, 2026

Summary

Release 5.8.0 huggingface / transformers Public Notifications You must be signed in to change notification settings Fork 34.7k Star 167k Release 5.8.0 vasqu released this 05 May 16:52 · 1333 commits to main since this release v5.8.0 049d2bf Release v5.8.0 New Model additions DeepSeek-V4 DeepSeek-V4 is the next-generation MoE (Mixture of Experts) language model from DeepSeek that introduces several architectural innovations over DeepSeek-V3. The architecture replaces Multi-head Latent Attention (MLA) with a hybrid local + long-range attention design, swaps residual connections for Manifold-Constrained Hyper-Connections (mHC), and bootstraps the first few MoE layers with a static token-id → expert-id hash table.

Why it matters

This Funding is relevant to the technology intelligence record because it involves DeepSeek, GitHub, DeepSeek that, DeepSeek v4. The source article should remain the factual reference for follow-up coverage.

Key facts
  • huggingface / transformers Public Notifications You must be signed in to change notification settings Fork 34.7k Star 167k Release 5.8.0 vasqu released this 05 May 16:52 · 1333 commits to main since this release v5.8.0 049d2bf Release v5.8.0 New Model additions DeepSeek-V4 DeepSeek-V4 is the next-generation MoE (Mixture of Experts) language model from DeepSeek that introduces several architectural innovations over DeepSeek-V3.
  • The architecture replaces Multi-head Latent Attention (MLA) with a hybrid local + long-range attention design, swaps residual connections for Manifold-Constrained Hyper-Connections (mHC), and bootstraps the first few MoE layers with a static token-id → expert-id hash table.
  • This implementation covers DeepSeek-V4-Flash, DeepSeek-V4-Pro, and their -Base pretrained variants, which share the same architecture but differ in width, depth, expert count and weights.
  • Links: Documentation | Paper Add DeepSeek V4 ( #45643 ) by @ArthurZucker in #45643 Gemma 4 Assistant Gemma 4 Assistant is a small, text-only model that enables speculative decoding for Gemma 4 models using the Multi-Token Prediction (MTP) method and associated candidate generator.
  • The model shares the same Gemma4TextModel backbone as other Gemma 4 models but uses KV sharing throughout the entire model, allowing it to reuse the KV cache populated by the target model and skip the pre-fill phase entirely.
  • This architecture includes cross-attention to make the most of the target model's context, allowing the assistant to accurately predict more drafted tokens per drafting round.
Entities in this story
Related events