Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Spark

  • A fast and general compute engine (originally for Hadoop data).

    • Often paired with Hadoop for its distributed filesystem (HDFS), cluster resource management and parallel processing.

  • Spark provides a simple and expressive programming model that supports a wide range of applications, including ETL (extract, transform, load), Machine Learning, stream processing, and graph computation.

  • Also communicates well with some databases and other resources.

  • Installation of Spark and its dependencies is explained in the Installation chapter.

Spark and Cassandra

  • Cassandra is one of the databases that work well with Spark.

    • Same type of distributed processing.

    • Same way of replicating for fault tolerance.

  • Spark can be deployed on the same nodes as Cassandra for:

    • local (short traveled) data manipulation, and

    • combination of results to a central hub (MapReduce).

  • Requires drivers from Datastax

    • Automatically downloaded and applied with the following configuration.

    • The combination of Java version, Spark version and datax driver is very sensitive.

  • A SparkSession instantiates Spark, applies configurations and connects to a data source.

Accessing tables

Note: The following sets of commands assume that the Cassandra notebook has been run first to set up the relevant keyspace and tables.

Database views

  • Useful for “setting the scene” before a more simplified data extraction.

  • The below example simply attaches to the correct keyspace and table.

    • The view could also be a selection into that table to query further.

Spark DataFrame

  • Related to a Pandas data frame, but can be distributed over compute nodes.

  • Various functions like filters, statistical calculations, groupBy, Pandas functions (mapInPandas), joins, etc.

  • Export to Pandas and JSON.

  • Reads many formats, including SQL, JSON, Excel, ...

Aggregation, grouping and filtering

  • These can be combined in many ways.

  • Starting from the left.

  • Order is important.

Write data to Cassandra

  • One can append or overwrite data in existing database tables.

  • PySpark is picky regarding data formats.

    • Reading data from the existing table and extracting formatting is possible.

  • PySpark is case sensitive, while Cassandra is not by default.

Exercise

  • Create a Cassandra table matching the structure of the planets data.

  • Insert the planets data into the table using Spark.

  • Read the data back from Cassandra through Spark and print it.