Test environment running 7.6.6

Cultural advice

The Australian National University acknowledges, celebrates and pays our respects to the Ngunnawal and Ngambri people of the Canberra region and to all First Nations Australians on whose traditional lands we meet and work, and whose cultures are among the oldest continuing cultures in human history.

Aboriginal and Torres Strait Islander peoples are advised that ANU Library collections may include images, names, voices, and other representations of deceased persons.

Material in the collection may contain terms, language or views that reflect the period in which the item was created and may be considered inappropriate today.

Application level fault recovery: Using fault-tolerant open MPI in a PDE solver

dc.contributor.authorAli, Muhammad
dc.contributor.authorSouthern, J.
dc.contributor.authorStrazdins, Peter
dc.contributor.authorHarding, Brendan
dc.coverage.spatialArizona, USA
dc.date.accessioned2015-12-10T22:39:40Z
dc.date.createdMay 19-23 2014
dc.date.issued2014
dc.date.updated2015-12-09T10:53:22Z
dc.description.abstractA fault-tolerant version of Open Message Passing Interface (Open MPI), based on the draft User Level Failure Mitigation (ULFM) proposal of the MPI Forum's Fault Tolerance Working Group, is used to create fault-tolerant applications. This allows applications and libraries to design their own recovery methods and control them at the user level. However, only a limited amount of research work on user level failure recovery (including the implementation and performance evaluation of this prototype) has been carried out. This paper contributes a fault-tolerant implementation of an application solving 2D partial differential equations (PDEs) by means of a sparse grid combination technique which is capable of surviving multiple process failures caused by the faults. Our fault recovery involves reconstructing the faulty communicators without shrinking the global size by re-spawning failed MPI processes on the same physical processors where they were before the failure (for load balancing). It also involves restoring lost data from either exact check pointed data on disk, approximated data in memory (via an alternate sparse grid combination technique) or a near-exact copy of replicated data in memory. The experimental results show that the faulty communicator reconstruction time is currently large in the draft ULFM, especially for multiple process failures. They also show that the alternate combination technique has the lowest data recovery overhead, except on a system with very low disk write latency for which checkpointing has the lowest overhead. Furthermore, the errors due to the recovery of approximated data are within a factor of 10 in all cases, with the surprising result that the alternate combination technique being more accurate than the near-exact replication method. The contributed implementation details, including the analysis of the experimental results, of this paper will help application developers to resolve different issues of design and implementation of fault-tolerant applications by means of the Open MPI ULFM standard.
dc.identifier.isbn9780769552088
dc.identifier.urihttp://hdl.handle.net/1885/57279
dc.publisherIEEE
dc.relation.ispartofseries28th IEEE International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2014
dc.sourceProceedings of the International Parallel and Distributed Processing Symposium, IPDPS
dc.titleApplication level fault recovery: Using fault-tolerant open MPI in a PDE solver
dc.typeConference paper
local.bibliographicCitation.lastpage1178
local.bibliographicCitation.startpage1169
local.contributor.affiliationAli, Muhammad, College of Engineering and Computer Science, ANU
local.contributor.affiliationSouthern, J., Fujitsu Laboratories of Europe
local.contributor.affiliationStrazdins, Peter, College of Engineering and Computer Science, ANU
local.contributor.affiliationHarding, Brendan, College of Physical and Mathematical Sciences, ANU
local.contributor.authoruidAli, Muhammad, u4616239
local.contributor.authoruidStrazdins, Peter, u8914893
local.contributor.authoruidHarding, Brendan, u4409191
local.description.embargo2037-12-31
local.description.notesImported from ARIES
local.description.refereedYes
local.identifier.absfor010204 - Dynamical Systems in Applications
local.identifier.absfor080501 - Distributed and Grid Systems
local.identifier.absfor080304 - Concurrent Programming
local.identifier.absseo970108 - Expanding Knowledge in the Information and Computing Sciences
local.identifier.ariespublicationa383154xPUB394
local.identifier.doi10.1109/IPDPSW.2014.132
local.identifier.scopusID2-s2.0-84918798613
local.type.statusPublished Version

Downloads

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
01_Ali_Application_level_fault_2014.pdf
Size:
545.39 KB
Format:
Adobe Portable Document Format