Resolving Latency-Accuracy Trade-offs in Real-Time Edge-Deployed Convolutional Neural Networks via Adaptive Quantization-Aware Pruning Schedules
Keywords:
quantization-aware training, structured neural network pruning, edge AI inference optimization, mixed-precision quantization, convolutional neural network compression, Pareto-optimal model efficiency, latency-accuracy trade-off, resource-constrained deployment, adaptive sparsity schedulingAbstract
The deployment of deep convolutional neural networks (CNNs) on resource-constrained edge hardware presents a persistent conflict between inference latency and predictive accuracy, particularly under dynamic workload conditions in industrial AI systems. This paper proposes a unified Adaptive Quantization-Aware Pruning (AQAP) framework that iteratively recalibrates structured sparsity masks and mixed-precision quantization policies during training, guided by a multi-objective Pareto-optimization criterion. Empirical evaluations conducted on NVIDIA Jetson AGX Orin and ARM Cortex-M85 platforms demonstrate that AQAP achieves up to 3.8× latency reduction while retaining 97.4% of baseline top-1 accuracy on ImageNet-1K and CIFAR-100 benchmarks. The proposed schedule outperforms static post-training quantization and magnitude-based pruning baselines across all tested compression ratios, establishing a reproducible methodology for latency-constrained AI engineering deployments.
References
Semeniuk, V. V. (2025). OPTIMIZATION OF LOCAL DEVELOPMENT PROCESS USING DOCKER PHP IMAGE THAT COMES WITH A FULL SET OF TOOLS OUT OF THE BOX: DATABASE AND INTERNATIONALIZATION EXTENSIONS. ВЧЕНІ ЗАПИСКИ, 12025226.