Model Compression Techniques such as pruning, quantization, and distillation that shrink models for faster, cheaper deployment.