vPALs: Towards Verified Performance-aware Learning System For Resource Management
Abstract: Accurately predicting task performance at runtime in a cluster is advantageous for a resource management system to determine whether a task should be migrated due to performance degradation caused by interference. This is beneficial for both cluster operators and service owners. However, deploying performance prediction systems with learning methods requires sophisticated safeguard mechanisms due to the inherent stochastic and black-box natures of these models, such as Deep Neural Networks (DNNs). Vanilla Neural Networks (NNs) can be vulnerable to out-of-distribution data samples that can lead to sub-optimal decisions. To take a step towards a safe learning system in performance prediction, We propose vPALs that leverage well-correlated system metrics, and verification to produce safe performance prediction at runtime, providing an extra layer of safety to integrate learning techniques to cluster resource management systems. Our experiments show that vPALs can outperform vanilla NNs across our benchmark workload.
- Dynamic allocation of computational resources for deep learning-enabled cellular image analysis with Kubernetes. BioRxiv, 505032.
- Virtual Machine Allocation with Lifetime Predictions. Proceedings of Machine Learning and Systems, 5.
- Take it to the limit: peak prediction-driven resource overcommitment in datacenters. In Proceedings of the Sixteenth European Conference on Computer Systems, 556–573.
- Pearson correlation coefficient. In Noise reduction in speech processing, 37–40. Springer.
- Deep neural network approximation theory for high-dimensional functions. arXiv:2112.14523.
- Cilantro: Performance-Aware Resource Allocation for General Objectives via Online Feedback. In 17th USENIX Symposium on Operating Systems Design and Implementation (OSDI 23), 623–643. Boston, MA: USENIX Association. ISBN 978-1-939133-34-2.
- Pattern recognition and machine learning, volume 4. Springer.
- Borg, Omega, and Kubernetes: Lessons learned from three container-management systems over a decade. Queue, 14(1): 70–93.
- Parties: Qos-aware resource partitioning for multiple interactive services. In Proceedings of the Twenty-Fourth International Conference on Architectural Support for Programming Languages and Operating Systems, 107–120.
- Benchmarking cloud serving systems with YCSB. In Proceedings of the 1st ACM symposium on Cloud computing, 143–154.
- Paragon: QoS-aware scheduling for heterogeneous datacenters. ACM SIGPLAN Notices, 48(4): 77–88.
- Caladan: Mitigating interference at microsecond timescales. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20), 281–297.
- Dropout as a bayesian approximation: Representing model uncertainty in deep learning. In international conference on machine learning, 1050–1059. PMLR.
- Google. 2023. cAdvisor: Analyzes resource usage and performance characteristics of running containers. Accessed: 2023-11-20.
- Grafana. 2023. Promtail: An agent which ships the contents of local logs to a private Grafana Loki instance or Grafana Cloud. Accessed: 2023-11-20.
- Monitorless: Predicting performance degradation in cloud applications with machine learning. In Proceedings of the 20th international middleware conference, 149–162.
- Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, 770–778.
- Profiling a warehouse-scale computer. In Proceedings of the 42nd Annual International Symposium on Computer Architecture, 158–169.
- Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980.
- ImageNet Classification with Deep Convolutional Neural Networks. Advances in neural information processing systems, 25.
- GraphCast: Learning skillful medium-range global weather forecasting.
- Pond: CXL-based memory pooling systems for cloud platforms. In Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2, 574–587.
- Algorithms for verifying deep neural networks. Foundations and Trends® in Optimization, 4(3-4): 244–404.
- Nnlqp: A multi-platform neural network latency query and prediction system with an evolving database. In Proceedings of the 51st International Conference on Parallel Processing, 1–14.
- Understanding and Optimizing Workloads for Unified Resource Management in Large Cloud Platforms. In Proceedings of the Eighteenth European Conference on Computer Systems, 416–432.
- Interference-aware scheduling for inference serving. In Proceedings of the 1st Workshop on Machine Learning and Systems, 80–88.
- Prometheus. 2023a. Node Exporter: Prometheus exporter for hardware and OS metrics exposed by *NIX kernels. Accessed: 2023-11-20.
- Prometheus. 2023b. Prometheus: An open-source systems monitoring and alerting toolkit. Accessed: 2023-11-20.
- Google-Wide Profiling: A Continuous Profiling Infrastructure for Data Centers. IEEE Micro, 65–79.
- Mage: Online and interference-aware scheduling for multi-scale heterogeneous systems. In Proceedings of the 27th International Conference on Parallel Architectures and Compilation Techniques, 1–13.
- Skynet: Performance-driven resource management for dynamic workloads. In 2021 IEEE 14th International Conference on Cloud Computing (CLOUD), 527–539. IEEE.
- Provisioning Differentiated Last-Level Cache Allocations to VMs in Public Clouds. In Proceedings of the ACM Symposium on Cloud Computing, SoCC ’21, 319–334. New York, NY, USA: Association for Computing Machinery. ISBN 9781450386388.
- Building Verified Neural Networks for Computer Systems with Ouroboros. Proceedings of Machine Learning and Systems, 5.
- Building Verified Neural Networks with Specifications for Systems. In Proceedings of the 12th ACM SIGOPS Asia-Pacific Workshop on Systems, APSys ’21, 42–47. New York, NY, USA: Association for Computing Machinery. ISBN 9781450386982.
- Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
- Evaluating job packing in warehouse-scale computing. In 2014 IEEE International Conference on Cluster Computing (CLUSTER), 48–56.
- VictoriaMetrics. 2023. VictoriaLogs: An open source user-friendly database for logs from VictoriaMetrics. Accessed: 2023-11-20.
- FaaSNet: Scalable and Fast Provisioning of Custom Serverless Container Runtimes at Alibaba Cloud Function Compute. In 2021 USENIX Annual Technical Conference (USENIX ATC 21), 443–457. USENIX Association. ISBN 978-1-939133-23-6.
- Characterizing Job Microarchitectural Profiles at Scale: Dataset and Analysis. In Proceedings of the 51st International Conference on Parallel Processing, 1–11.
- Efficient Formal Safety Analysis of Neural Networks. In Proceedings of the 32nd International Conference on Neural Information Processing Systems, NIPS’18, 6369–6379. Red Hook, NY, USA: Curran Associates Inc.
- Smartharvest: Harvesting idle cpus safely and efficiently in the cloud. In Proceedings of the Sixteenth European Conference on Computer Systems, 1–16.
- Weiner, J. 2018. PSI - Pressure Stall Information. https://www.kernel.org/doc/html/latest/accounting/psi.html. Accessed: 19-11-2023.
- Tmo: transparent memory offloading in datacenters. In Proceedings of the 27th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, 609–621.
- Pythia: Improving datacenter utilization via precise contention prediction for multiple co-located workloads. In Proceedings of the 19th International Middleware Conference, 146–160.
- Horus: Interference-Aware and Prediction-Based Scheduling in Deep Learning Systems. IEEE Transactions on Parallel and Distributed Systems, 33(1): 88–100.
- Delay scheduling: a simple technique for achieving locality and fairness in cluster scheduling. In Proceedings of the 5th European conference on Computer systems, 265–278.
- Nn-meter: Towards accurate latency prediction of deep-learning model inference on diverse edge devices. In Proceedings of the 19th Annual International Conference on Mobile Systems, Applications, and Services, 81–93.
- URSA: Precise Capacity Planning and Fair Scheduling Based on Low-Level Statistics for Public Clouds. In Proceedings of the 49th International Conference on Parallel Processing, ICPP ’20. New York, NY, USA: Association for Computing Machinery. ISBN 9781450388160.
- CPI2: CPU performance isolation for shared compute clusters. In Proceedings of the 8th ACM European Conference on Computer Systems, 379–391.
- Sinan: ML-based and QoS-aware resource management for cloud microservices. In Proceedings of the 26th ACM international conference on architectural support for programming languages and operating systems, 167–181.
Paper Prompts
Sign up for free to create and run prompts on this paper.