Sciweavers

ICDCS
2007
IEEE

Fault Tolerance in Multiprocessor Systems Via Application Cloning

13 years 11 months ago
Fault Tolerance in Multiprocessor Systems Via Application Cloning
Record and Replay (RR) is a software based state replication solution designed to support recording and subsequent replay of the execution of unmodified applications running on multiprocessor systems for fault-tolerance. Multiple instances of the application are simultaneously executed in separate virtualized environments called Containers. Containers facilitate state replication between the application instances by resolving the resource conflicts and providing a uniform view of the underlying operating system across all clones. The virtualization layer that creates ainer abstraction actively monitors the primary instance of the application and synchronizes its state with that of the clones by transferring the necessary information to enforce identical state among them. In particular, we address the replication of relevant operating system state, such as network state to preserve network connections across failures, and the state that results from nondeterministic interleaved accesse...
Philippe Bergheaud, Dinesh Subhraveti, Marc Vertes
Added 03 Jun 2010
Updated 03 Jun 2010
Type Conference
Year 2007
Where ICDCS
Authors Philippe Bergheaud, Dinesh Subhraveti, Marc Vertes
Comments (0)