Pruning Binarized Neural Networks: A Dedicated Framework and Globally Weighted Algorithms
Proposes a framework for pruning binarized neural networks to improve efficiency on edge hardware like FPGAs.
Extreme compression of deep neural networks, up to full binarization, dramatically reduces memory footprint and arithmetic complexity, facilitating deployment on constrained edge hardware with field-programmable gate arrays (FPGAs) and microcontrollers. Although combining binarization with pruning promises additional efficiency gains, existing pruning strategies are ill-suited to binarized representations and rarely…