BlobSeer: Bringing high throughput under heavy concurrency to Hadoop Map-Reduce applications

15 years 2 months ago

Download hal.inria.fr

Hadoop is a software framework supporting the Map/Reduce programming model. It relies on the Hadoop Distributed File System (HDFS) as its primary storage system. The efficiency of HDFS is crucial for the performance of Map/Reduce applications. We substitute the original HDFS layer of Hadoop with a new, concurrency-optimized data storage layer based on the BlobSeer data management service. Thereby, the efficiency of Hadoop is significantly improved for data-intensive Map/Reduce applications, which naturally exhibit a high degree of data access concurrency. Moreover, BlobSeer's features (built-in versioning, its support for concurrent append operations) open the possibility for Hadoop to further extend its functionalities. We report on extensive experiments conducted on the Grid'5000 testbed. The results illustrate the benefits of our approach over the original HDFS-based implementation of Hadoop. Keywords-Large-scale distributed computing; Data-intensive; Map/Reduce-based appl...

Bogdan Nicolae, Diana Moise, Gabriel Antoniu, Luc

Real-time Traffic

Data Storage Layer | Distributed And Parallel Computing | File System | Hadoop | IPPS 2010 |

claim paper

» Enabling High Data Throughput in Desktop Grids through Decentralized Data and Metadata Man...

» High Throughput DataCompression for Cloud Storage

» Using Global Behavior Modeling to Improve QoS in Cloud Data Storage Services

Post Info
More Details (n/a)

Added	13 Feb 2011
Updated	13 Feb 2011
Type	Journal
Year	2010
Where	IPPS
Authors	Bogdan Nicolae, Diana Moise, Gabriel Antoniu, Luc Bougé, Matthieu Dorier

Comments (0)

Sciweavers

BlobSeer: Bringing high throughput under heavy concurrency to Hadoop Map-Reduce applications

Data Storage Layer | Distributed And Parallel Computing | File System | Hadoop | IPPS 2010 |

Explore & Download

Productivity Tools

Sciweavers