Dynamic Quantization
reduce resource usage
Popularity:
Ease-of-Use:
Description
Post-training conversion of neural network weights to lower precision formats like INT8 to reduce memory use and computational cost during inference, enhancing speed and hardware efficiency.