ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation
Summary
arXiv:2607.13124v1 Announce Type: new Abstract: Structured pruning is a hardware-friendly way to compress LLMs, but it is mostly validated on multiple-choice recognition tasks, while the same compressed checkpoints can collapse on the free-form generation that deployment actually requires. Two observations trace this gap. First, greedy \textsc{pass}@$1$ nearly vanishes after compression, yet \textsc{pass}@$k$ recovers substantially under repeated sampling: useful generations are demoted, not erased.
Why It Matters
This Robotics & Manufacturing development accelerates factory automation, industrial AI and precision manufacturing across the region. For Asia, it is a signal worth tracking: it shapes who supplies, who scales, and who sets the standard over the next five years.
Key Facts
- SectorRobotics & Manufacturing
- Market—
- ImpactMedium (58/100)
- SignalResearch