Bookbot

Big Data Made Easy

A Working Guide to the Complete Hadoop Toolset

Autor*innen

  • Autorenkollektiv

Parameter

  • 392 Seiten
  • 14 Lesestunden

Mehr zum Buch

Many corporations are struggling as their data sets exceed the capabilities of traditional systems for storage and processing. The solution lies in implementing a big data system. Apache Hadoop provides a scalable, fault-tolerant framework for parallel data storage and processing. It features a comprehensive toolset for various functions: storage (Hadoop), configuration (YARN and ZooKeeper), data collection (Nutch and Solr), processing (Storm, Pig, and MapReduce), scheduling (Oozie), data movement (Sqoop and Avro), monitoring (Chukwa, Ambari, and Hue), testing (Big Top), and analysis (Hive). This resource approaches the management of massive data sets from a systems perspective, detailing the roles of each project and how to utilize the Hadoop toolset effectively at each stage. With clear explanations and numerous examples, it guides users on employing each tool. The book also outlines a sliding scale of tools based on data size, instructing developers, architects, testers, and project managers on storing, configuring, processing, scheduling, moving, monitoring, analyzing, reporting, and testing big data systems. It caters to developers, architects, IT project managers, database administrators, and anyone interested in Hadoop or big data, as well as those looking to enhance their careers with big data skills.

Buchkauf

Big Data Made Easy, Autorenkollektiv

Sprache
Erscheinungsdatum
2014
product-detail.submit-box.info.binding
(Paperback)
Wir benachrichtigen dich per E-Mail.

Lieferung

  • Gratis Versand ab 16,99 € in ganz Österreich! Mehr Infos.

Zahlungsmethoden

Keiner hat bisher bewertet.Abgeben

Titel
Big Data Made Easy
Untertitel
A Working Guide to the Complete Hadoop Toolset
Sprache
Englisch
Autor*innen
Autorenkollektiv
Erscheinungsdatum
2014
Einband
Paperback
Seitenzahl
392
ISBN10
1484200950
ISBN13
9781484200957
Reihe
Beschreibung
Many corporations are struggling as their data sets exceed the capabilities of traditional systems for storage and processing. The solution lies in implementing a big data system. Apache Hadoop provides a scalable, fault-tolerant framework for parallel data storage and processing. It features a comprehensive toolset for various functions: storage (Hadoop), configuration (YARN and ZooKeeper), data collection (Nutch and Solr), processing (Storm, Pig, and MapReduce), scheduling (Oozie), data movement (Sqoop and Avro), monitoring (Chukwa, Ambari, and Hue), testing (Big Top), and analysis (Hive). This resource approaches the management of massive data sets from a systems perspective, detailing the roles of each project and how to utilize the Hadoop toolset effectively at each stage. With clear explanations and numerous examples, it guides users on employing each tool. The book also outlines a sliding scale of tools based on data size, instructing developers, architects, testers, and project managers on storing, configuring, processing, scheduling, moving, monitoring, analyzing, reporting, and testing big data systems. It caters to developers, architects, IT project managers, database administrators, and anyone interested in Hadoop or big data, as well as those looking to enhance their careers with big data skills.