TY - GEN
T1 - Precision and Performance-Aware Voltage Scaling in DNN Accelerators
AU - Rathore, Mallika
AU - Milder, Peter
AU - Salman, Emre
N1 - Publisher Copyright:
© 2023 ACM.
PY - 2023/6/5
Y1 - 2023/6/5
N2 - A methodology is proposed to enhance the energy efficiency of systolic array based deep neural network (DNN) accelerators by enabling precision- and performance-aware voltage scaling. The proposed framework consists of three primary steps. In the first step, the voltage-dependent timing error probability for each output bit within the processing elements is analytically estimated. Next, these timing errors are injected into DNN models, helping us understand how inference accuracy is affected by lower operating voltages. In the last step, we apply error detection and correction to only select bits within the network, thereby improving inference accuracy while minimizing circuit overhead. For a 256X256 array operating at 0.7GHz and evaluating MobileNetV2 on ImageNet, we can reduce the nominal supply voltage from 0.9V to 0.5V with negligible (0.001%) latency overhead. This reduction in supply voltage reduces the inference energy by 79.4% while degrading inference accuracy by only 0.29%.
AB - A methodology is proposed to enhance the energy efficiency of systolic array based deep neural network (DNN) accelerators by enabling precision- and performance-aware voltage scaling. The proposed framework consists of three primary steps. In the first step, the voltage-dependent timing error probability for each output bit within the processing elements is analytically estimated. Next, these timing errors are injected into DNN models, helping us understand how inference accuracy is affected by lower operating voltages. In the last step, we apply error detection and correction to only select bits within the network, thereby improving inference accuracy while minimizing circuit overhead. For a 256X256 array operating at 0.7GHz and evaluating MobileNetV2 on ImageNet, we can reduce the nominal supply voltage from 0.9V to 0.5V with negligible (0.001%) latency overhead. This reduction in supply voltage reduces the inference energy by 79.4% while degrading inference accuracy by only 0.29%.
KW - dnn accelerator
KW - energy efficiency
KW - voltage scaling
UR - https://www.scopus.com/pages/publications/85163143658
U2 - 10.1145/3583781.3590202
DO - 10.1145/3583781.3590202
M3 - Conference contribution
AN - SCOPUS:85163143658
T3 - Proceedings of the ACM Great Lakes Symposium on VLSI, GLSVLSI
SP - 237
EP - 242
BT - GLSVLSI 2023 - Proceedings of the Great Lakes Symposium on VLSI 2023
PB - Association for Computing Machinery
T2 - 33rd Great Lakes Symposium on VLSI, GLSVLSI 2023
Y2 - 5 June 2023 through 7 June 2023
ER -