Guide to High Performance Distributed Computing Case Studies with Hadoop, Scalding and Spark
Synopsis
This timely text/reference describes the development and implementation of large-scale distributed processing systems using open source tools and technologies. Comprehensive in scope, the book presents state-of-the-art material on building high performance distributed computing systems, providing practical guidance and best practices as well as describing theoretical software frameworks. Features: describes the fundamentals of building scalable software systems for large-scale data processing in the new paradigm of high performance distributed computing; presents an overview of the Hadoop ecosystem, followed by step-by-step instruction on its installation, programming and execution; Reviews the basics of Spark, including resilient distributed datasets, and examines Hadoop streaming and working with Scalding; Provides detailed case studies on approaches to clustering, data classification and regression analysis; Explains the process of creating a working recommender system using Scalding and Spark.
Book details
- Edition:
- 2015
- Series:
- Computer Communications and Networks
- Author:
- K.G. Srinivasa, Anil Kumar Muppalla
- ISBN:
- 9783319134970
- Related ISBNs:
- 9783319134963
- Publisher:
- Springer International Publishing
- Pages:
- N/A
- Reading age:
- Not specified
- Includes images:
- No
- Date of addition:
- 2019-09-05
- Usage restrictions:
- Copyright
- Copyright date:
- 2015
- Copyright by:
- N/A
- Adult content:
- No
- Language:
-
English
- Categories:
-
Art and Architecture, Computers and Internet, Nonfiction