Highly Parallel Programming with Apache Spark

Tutorials – Apache Spark

Article from Issue 202/2017

Author(s): Ben Everard

Churn through lots of data with cluster computing on Apache's Spark platform.

As a society, we're creating more data than ever before. We're monitoring everything from the planet's weather to the performance of our computers, and we're storing all this information. But how do you process all this data? On a single machine, you can get a few terabytes of disk space and a few hundred gigabytes of memory (at least, you can if your pockets are deep enough), but how do you churn through a petabyte of raw ones and zeros? Basically, you're going to need more than one computer, and you're going to look for a method of running your programs on many machines at the same time: Apache Spark [1].

Before you run off and buy a rack of servers, slow down! We're going to start by introducing Spark on a single machine. Once you've mastered the basics, you can scale up.

Spark is a data processing engine that is often used with Hadoop for managing large amounts of data in a highly distributed manner. If you move forward with Spark, you're probably going to end up with a complete Hadoop setup; however, that's also getting ahead of ourselves. We can start Spark as a standalone service on a single computer.

[...]

Use Express-Checkout link below to read the full article (PDF).

Buy this article as PDF

Express-Checkout as PDF

Price $2.95
(incl. VAT)

Buy Linux Magazine

SINGLE ISSUES

Print Issues

Digital Issues

SUBSCRIPTIONS

Print Subs

Digisubs

TABLET & SMARTPHONE APPS

US / Canada

UK / Australia

Support Our Work

Linux Magazine content is made possible with support from readers like you. Please consider contributing when you’ve found an article to be beneficial.

News

Linux Mint 22.2 Beta Available for Testing

Linux mint , Operating Systems , Wayland

Some interesting new additions and improvements are coming to Linux Mint. Check out the Linux Mint 22.2 Beta to give it a test run.
Debian 13.0 Officially Released

DEBIAN , Linux , Operating Systems

After two years of development, the latest iteration of Debian is now available with plenty of under-the-hood improvements.
Upcoming Changes for MXLinux

MXLinux , Plasma , Wayland

MXLinux 25 has plenty in store to please all types of users.
A New Linux AI Assistant in Town

Artificial Inte... , Linux , LLM

Newelle, a Linux AI assistant, works with different LLMs and includes document parsing and profiles.
Linux Kernel 6.16 Released with Minor Fixes

Kernel , Linux , Security

The latest Linux kernel doesn't really include any big-ticket features, just a lot of lines of code.
EU Sovereign Tech Fund Gains Traction

funding , open source , Security

OpenForum Europe recently released a report regarding a sovereign tech fund with backing from several significant entities.
FreeBSD Promises a Full Desktop Installer

Desktop , FreeBSD , open source

FreeBSD has lacked an option to include a full desktop environment during installation.
Linux Hits an Important Milestone

Linux , open source

If you pay attention to the news in the Linux-sphere, you've probably heard that the open source operating system recently crashed through a ceiling no one thought possible.
Plasma Bigscreen Returns

KDE , open source , Plasma

A developer discovered that the Plasma Bigscreen feature had been sitting untouched, so he decided to do something about it.
CachyOS Now Lets Users Choose Their Shell

CachyOS , shell , Wayland

Imagine getting the opportunity to select which shell you want during the installation of your favorite Linux distribution. That's now a thing.

Highly Parallel Programming with Apache Spark

Tutorials – Apache Spark

Buy this article as PDF

Buy Linux Magazine

Related content

Subscribe to our Linux Newsletters
Find Linux and Open Source Jobs
Subscribe to our ADMIN Newsletters

Support Our Work

News

Linux Mint 22.2 Beta Available for Testing

Debian 13.0 Officially Released

Upcoming Changes for MXLinux

A New Linux AI Assistant in Town

Linux Kernel 6.16 Released with Minor Fixes

EU Sovereign Tech Fund Gains Traction

FreeBSD Promises a Full Desktop Installer

Linux Hits an Important Milestone

Plasma Bigscreen Returns

CachyOS Now Lets Users Choose Their Shell

Highly Parallel Programming with Apache Spark

Tutorials – Apache Spark

Buy this article as PDF

Buy Linux Magazine

Related content

Subscribe to our Linux Newsletters Find Linux and Open Source Jobs Subscribe to our ADMIN Newsletters

Support Our Work

News

Tag Cloud

Subscribe to our Linux Newsletters
Find Linux and Open Source Jobs
Subscribe to our ADMIN Newsletters