Pro Apache Hadoop

  • Authors
  • Sameer Wadkar
  • Madhu Siddalingaiah

Table of contents

  1. Front Matter
    Pages i-xxvi
  2. Sameer Wadkar, Madhu Siddalingaiah
    Pages 1-10
  3. Sameer Wadkar, Madhu Siddalingaiah
    Pages 11-30
  4. Sameer Wadkar, Madhu Siddalingaiah
    Pages 31-46
  5. Sameer Wadkar, Madhu Siddalingaiah
    Pages 47-72
  6. Sameer Wadkar, Madhu Siddalingaiah
    Pages 73-106
  7. Sameer Wadkar, Madhu Siddalingaiah
    Pages 107-150
  8. Sameer Wadkar, Madhu Siddalingaiah
    Pages 151-183
  9. Sameer Wadkar, Madhu Siddalingaiah
    Pages 185-202
  10. Sameer Wadkar, Madhu Siddalingaiah
    Pages 203-215
  11. Sameer Wadkar, Madhu Siddalingaiah
    Pages 217-239
  12. Sameer Wadkar, Madhu Siddalingaiah
    Pages 241-269
  13. Sameer Wadkar, Madhu Siddalingaiah
    Pages 271-282
  14. Sameer Wadkar, Madhu Siddalingaiah
    Pages 283-291
  15. Sameer Wadkar, Madhu Siddalingaiah
    Pages 293-323
  16. Sameer Wadkar, Madhu Siddalingaiah
    Pages 325-342
  17. Sameer Wadkar, Madhu Siddalingaiah
    Pages 343-356
  18. Sameer Wadkar, Madhu Siddalingaiah
    Pages 357-379
  19. Sameer Wadkar, Madhu Siddalingaiah
    Pages 381-390
  20. Sameer Wadkar, Madhu Siddalingaiah
    Pages 391-398

About this book

Introduction

Pro Apache Hadoop, Second Edition brings you up to speed on Hadoop – the framework of big data. Revised to cover Hadoop 2.0, the book covers the very latest developments such as YARN (aka MapReduce 2.0), new HDFS high-availability features, and increased scalability in the form of HDFS Federations. All the old content has been revised too, giving the latest on the ins and outs of MapReduce, cluster design, the Hadoop Distributed File System, and more.

This book covers everything you need to build your first Hadoop cluster and begin analyzing and deriving value from your business and scientific data. Learn to solve big-data problems the MapReduce way, by breaking a big problem into chunks and creating small-scale solutions that can be flung across thousands upon thousands of nodes to analyze large data volumes in a short amount of wall-clock time. Learn how to let Hadoop take care of distributing and parallelizing your software—you just focus on the code; Hadoop takes care of the rest.

  • Covers all that is new in Hadoop 2.0
  • Written by a professional involved in Hadoop since day one
  • Takes you quickly to the seasoned pro level on the hottest cloud-computing framework  

Bibliographic information

Industry Sectors
Pharma
Automotive
Biotechnology
Electronics
Telecommunications
Consumer Packaged Goods
Energy, Utilities & Environment
Aerospace