Distilling Agentic Software Engineering into Qwen3-8B: Two Trace Pipelines, One End-to-End Study
From Black-Box Agent Traces to Distillation Data: Fable Collection and Causal Kimi Rationale Synthesis
Distilling Kimi-K3 into Qwen3-8B: a ~13x Agentic-SWE Gain, and Where It Plateaus
K3 → Qwen3-8B Distillation: The Complete Runbook
The Corpus Entropy Profile: How Hard vs. How Unevenly Hard
Fine-tuning Large Language Models with Mini-Sequence Technology and Distributed Training
Extending LLAMA Training Context with Mini-Sequence Technology
Extending Mistral Training Context with Mini-Sequence Technology
Extending Qwen Training Context with Mini-Sequence Technology
Extending gemma2 Training with Mini-Sequence Technology
Revolutionizing LLM Training with Mini-Sequence Technology