Sciweavers

SC
2009
ACM

Supporting fault-tolerance for time-critical events in distributed environments

13 years 11 months ago
Supporting fault-tolerance for time-critical events in distributed environments
In this paper, we consider the problem of supporting fault tolerance for adaptive and time-critical applications in heterogeneous and unreliable grid computing environments. Our goal for this class of applications is to optimize a user-specified benefit function while meeting the time deadline. Our first contribution in this paper is a multi-objective optimization algorithm for scheduling the application onto the most efficient and reliable resources. In this way, the processing can achieve the maximum benefit while also maximizing the success rate, which is the probability of finishing execution without failures. However, for the cases where failures do occur, we have developed a hybrid failure-recovery scheme to ensure that the application can complete within the pre-specified time interval. Our experimental results show that our scheduling algorithm can achieve better benefit when compared to several heuristics-based greedy scheduling algorithms, while still having a neglig...
Qian Zhu, Gagan Agrawal
Added 19 May 2010
Updated 19 May 2010
Type Conference
Year 2009
Where SC
Authors Qian Zhu, Gagan Agrawal
Comments (0)