The Distillation Paradox: Is Training One AI on Another Learning or Theft?
As open-weight models challenge frontier AI, knowledge distillation has become the ultimate battlefield. Is using a teacher model's output to train a student a legitimate optimization, or is it algorithmic plagiarism?