Dynamic Quantization

reduce resource usage
Popularity:
Ease-of-Use:

Description

Post-training conversion of neural network weights to lower precision formats like INT8 to reduce memory use and computational cost during inference, enhancing speed and hardware efficiency.