BLOGS

Distilling Kimi-K3 into Qwen3-8B: a ~13x Agentic-SWE Gain, and Where It Plateaus

The Corpus Entropy Profile: How Hard vs. How Unevenly Hard

Multi-Head Attention Residuals

Delta Attention Residuals: It's the Routing, Not the Sources

Open Attention Residuals: Replacing Additive Residuals with Learned Cross-Layer Attention

Fine-tuning Large Language Models with Mini-Sequence Technology and Distributed Training

Extending LLAMA Training Context with Mini-Sequence Technology

Extending Mistral Training Context with Mini-Sequence Technology

Extending Qwen Training Context with Mini-Sequence Technology

Extending gemma2 Training with Mini-Sequence Technology

Revolutionizing LLM Training with Mini-Sequence Technology