Published event
ArtificialIntelligence Research 1 source(s)

Show HN: DynamicTune – Closed-form trajectory weight surgery from 4B into 0.8B

Updated October 3, 2026 · 12:04 AM · 4 · source date October 2, 2026

Summary

Show HN: DynamicTune – Closed-form trajectory weight surgery from 4B into 0.8B dsadawq3 / DynamicTune Public Notifications You must be signed in to change notification settings Fork 0 Star 0 Branches Tags Open more actions menu Latest commit History 6 Commits 6 Commits Folders and files Name Name Last commit message Last commit date benchmarks benchmarks configs configs data data docs docs faytuna_flow faytuna_flow scripts scripts tests tests .gitignore .gitignore README.md README.md pyproject.toml pyproject.toml uv.lock uv.lock Repository files navigation DynamicTune Challenging the trillion-token orthodoxy: cross-model hidden trajectory transport and closed-form weight surgery across architectures and model widths. Tested across radically different model families: modern hybrid Qwen3.5 (4B with d=2560 -> 0.8B with d=1024 ) and notoriously fragile GPT-2 (XL with d=1600 -> small with d=768 ).

Why it matters

This Research is relevant to the technology intelligence record because it involves Anthropic, Cohere, Perplexity, AMD. The source article should remain the factual reference for follow-up coverage.

Key facts
  • dsadawq3 / DynamicTune Public Notifications You must be signed in to change notification settings Fork 0 Star 0 Branches Tags Open more actions menu Latest commit History 6 Commits 6 Commits Folders and files Name Name Last commit message Last commit date benchmarks benchmarks configs configs data data docs docs faytuna_flow faytuna_flow scripts scripts tests tests .gitignore .gitignore README.md README.md pyproject.toml pyproject.toml uv.lock uv.lock Repository files navigation DynamicTune Challenging the trillion-token orthodoxy: cross-model hidden trajectory transport and closed-form weight surgery across architectures and model widths.
  • Tested across radically different model families: modern hybrid Qwen3.5 (4B with d=2560 -> 0.8B with d=1024 ) and notoriously fragile GPT-2 (XL with d=1600 -> small with d=768 ).
  • Breaking the Trillion-Token Orthodoxy The prevailing consensus in deep learning is that transferring capabilities from a larger teacher model to a smaller student demands billions or trillions of tokens, massive synthetic dataset pipelines, and weeks of GPU cluster compute running token-level cross-entropy or KL divergence minimization.
  • DynamicTune challenges this dogma: A transformer stack is fundamentally a discrete dynamical system over depth: $h_{l+1} = h_l + f_l(h_l)$ .
  • A capable teacher traces an informational velocity field through representation space.
  • By aligning these trajectories through a local orthogonal Procrustes atlas and solving for closed-form weight updates in the student's MLP blocks, we can physically transfer teacher trajectory dynamics into the student without backpropagation or training runs.
Entities in this story
Related events