Published event
ArtificialIntelligence ModelRelease 1 source(s)

Release v5.9.0

Updated September 26, 2026 · 2:47 PM · source date May 20, 2026

Summary

Release v5.9.0 huggingface / transformers Public Notifications You must be signed in to change notification settings Fork 34.7k Star 167k Release v5.9.0 Cyrilvallez released this 20 May 14:12 · 1222 commits to main since this release v5.9.0 0a2757d This commit was signed with the committer’s verified signature . Cyrilvallez Cyril Vallez SSH Key Fingerprint: OjK+mdCRyLrQ4vIh2+8FffCzuDZs1WByt+LD1Z6WNw4 Verified Learn about vigilant mode .

Why it matters

This ModelRelease is relevant to the technology intelligence record because it involves Cohere, GitHub, Qwen. The source article should remain the factual reference for follow-up coverage.

Key facts
  • huggingface / transformers Public Notifications You must be signed in to change notification settings Fork 34.7k Star 167k Release v5.9.0 Cyrilvallez released this 20 May 14:12 · 1222 commits to main since this release v5.9.0 0a2757d This commit was signed with the committer’s verified signature .
  • Cyrilvallez Cyril Vallez SSH Key Fingerprint: OjK+mdCRyLrQ4vIh2+8FffCzuDZs1WByt+LD1Z6WNw4 Verified Learn about vigilant mode .
  • Release v5.9.0 New Model additions Cohere2Moe Command A+ is a Mixture-of-Experts (MoE) language model from Cohere that features a hybrid attention pattern combining sliding window and full attention layers.
  • The model incorporates both shared and routed experts and supports a very large context window for processing extensive text sequences.
  • Links: Documentation Add new cohere2_moe model ( #46115 ) by @Cyrilvallez in #46115 Parakeet tdt ( #44171 ) Parakeet tdt ( #44171 ) by @lmaksym HRM-Text HRM-Text is an improved autoregressive language-modeling variant of the Hierarchical Reasoning Model (HRM) that uses a hierarchical recurrent forward pass with two transformer stacks - one for slow, abstract planning (H) and one for fast, detailed computation (L) - reused inside a nested recurrence.
  • It features PrefixLM attention where instruction tokens attend bidirectionally while response tokens attend causally, per-head sigmoid output gates, and parameterless RMSNorm.
Entities in this story
Related events