2019 IEEE International Symposium on Information Theory (ISIT) | 2019

Timely Coded Computing

Abstract

In modern distributed computing systems, unpredictable and unreliable infrastructures result in high variability of computing resources. Meanwhile, there is significantly increasing demand for timely and event-driven services with deadline constraints. Motivated by measurements over Amazon EC2 clusters, we consider a two-state Markov model for variability of computing speed in cloud networks. In this model, each worker can be either in a good state or a bad state in terms of the computation speed, and the transition between these states is modeled as a Markov chain which is unknown to the scheduler. We then consider a Coded Computing framework, in which the data is possibly encoded and stored at the worker nodes in order to provide robustness against nodes that may be in a bad state. Our goal is to design the optimal computation-load allocation strategy that maximizes the timely computation throughput (i.e, the average number of computation tasks accomplished before their deadline). Our main result is the development of a dynamic computation strategy called Estimate-and-Allocate (EA) strategy, which achieves the optimal timely computation throughput. Compared with the static allocation strategy, EA improves the timely computation throughput by 1.44 ×4.6 in experiments over Amazon EC2 clusters.

Volume None

2019 IEEE International Symposium on Information Theory (ISIT) | 2019

Timely Coded Computing

Abstract

Volume None

Pages 2798-2802

DOI 10.1109/ISIT.2019.8849235

Language English

Journal 2019 IEEE International Symposium on Information Theory (ISIT)

Full Text