Hadoop Distributed File System(HDFS) Content of this Lecture:In this lecture, we will discuss design goals of HDFS, theread/write process to HDFS, the main configurationtuning parameters to control HDFS performance androbustness.Big Data Computing Hadoop Distributed File System (HDFS)Vu PhamIntroductionHadoop provides a distributed file system and a framework forthe analysis and transformation
...[Show More]
Hadoop Distributed File System
(HDFS)
Content of this Lecture:
In this lecture, we will discuss design goals of HDFS, the
read/write process to HDFS, the main configuration
tuning parameters to control HDFS performance and
robustness.
Big Data Computing Hadoop Distributed File System (HDFS)
Vu Pham
Introduction
Hadoop provides a distributed file system and a framework for
the analysis and transformation of very large data sets using
the MapReduce paradigm.
An important characteristic of Hadoop is the partitioning of
data and computation across many (thousands) of hosts, and
executing application computations in parallel close to their
data.
A Hadoop cluster scales computation capacity, storage capacity
and IO bandwidth by simply adding commodity servers.
Hadoop clusters at Yahoo! span 25,000 servers, and store 25
petabytes of application data, with the largest cluster being
3500 servers. One hundred other organizations worldwide
report using Hadoop.
Big Data Computing Hadoop Distributed File System (HDFS)
Vu Pham
Introduction
Hadoop is an Apache project; all components are available via
the Apache open source license.
Yahoo! has developed and contributed to 80% of the core of
Hadoop (HDFS and MapReduce).
HBase was originally developed at Powerset, now a department
at Microsoft.
Hive was originated and developed at Facebook.
Pig, ZooKeeper, and Chukwa were originated and developed at
Yahoo!
Avro was originated at Yahoo! and is being co-developed with
Cloudera.
Big Data Computing Hadoop Distributed File System (HDFS
[Show Less]