跳到主要导航 跳到搜索 跳到主要内容

Checkpointing Workflows for Fail-Stop Errors

  • Li Han*
  • , Louis Claude Canon
  • , Henri Casanova
  • , Yves Robert
  • , Frederic Vivien
  • *此作品的通讯作者
  • CNRS
  • Universite de Bourgogne Franche-Comte
  • University of Hawai'i at Mānoa
  • The University of Tennessee, Knoxville

科研成果: 期刊稿件文章同行评审

摘要

We consider the problem of orchestrating the execution of workflow applications structured as Directed Acyclic Graphs (DAGs) on parallel computing platforms that are subject to fail-stop failures. The objective is to minimize expected overall execution time, or makespan. A solution to this problem consists of a schedule of the workflow tasks on the available processors and of a decision of which application data to checkpoint to stable storage, so as to mitigate the impact of processor failures. To address this challenge, we consider a restricted class of graphs, Minimal Series-Parallel Graphs (M-SPGs), which is relevant to many real-world workflow applications. For this class of graphs, we propose a recursive list-scheduling algorithm that exploits the M-SPG structure to assign sub-graphs to individual processors, and uses dynamic programming to decide how to checkpoint these sub-graphs. We assess the performance of our algorithm for production workflow configurations, comparing it to an approach in which all application data is checkpointed and an approach in which no application data is checkpointed. Results demonstrate that our algorithm outperforms both the former approach, because of lower checkpointing overhead, and the latter approach, because of better resilience to failures.

源语言英语
页(从-至)1105-1120
页数16
期刊IEEE Transactions on Computers
67
8
DOI
出版状态已出版 - 1 8月 2018

学术指纹

探究 'Checkpointing Workflows for Fail-Stop Errors' 的科研主题。它们共同构成独一无二的学术指纹。

引用此