Guide to High Performance Distributed Computing Case Studies with Hadoop, Scalding and Spark

You must be logged in to access this title.

Sign up now

Already a member? Log in

Synopsis

This timely text/reference describes the development and implementation of large-scale distributed processing systems using open source tools and technologies. Comprehensive in scope, the book presents state-of-the-art material on building high performance distributed computing systems, providing practical guidance and best practices as well as describing theoretical software frameworks. Features: describes the fundamentals of building scalable software systems for large-scale data processing in the new paradigm of high performance distributed computing; presents an overview of the Hadoop ecosystem, followed by step-by-step instruction on its installation, programming and execution; Reviews the basics of Spark, including resilient distributed datasets, and examines Hadoop streaming and working with Scalding; Provides detailed case studies on approaches to clustering, data classification and regression analysis; Explains the process of creating a working recommender system using Scalding and Spark.

Book details

Edition:
2015
Series:
Computer Communications and Networks
Author:
K.G. Srinivasa, Anil Kumar Muppalla
ISBN:
9783319134970
Related ISBNs:
9783319134963
Publisher:
Springer International Publishing
Pages:
N/A
Reading age:
Not specified
Includes images:
No
Date of addition:
2019-09-05
Usage restrictions:
Copyright
Copyright date:
2015
Copyright by:
N/A 
Adult content:
No
Language:
English
Categories:
Art and Architecture, Computers and Internet, Nonfiction