ML & Data Science

Big data · Graphs

Big Data Graph Processing

The same graph problem on two big data engines, plus a native graph database for relationship queries.

Role
Data Engineer
Year
2025
Focus
Big data
Stack
4 technologies
  • PySpark
  • Hadoop
  • Neo4j
  • Cypher
  • 68M+

    Edges (LiveJournal)

  • 3

    SNAP graphs

  • 2

    Engines benchmarked

In-degree distributions are computed for three SNAP datasets: email-EuAll, web-BerkStan, and soc-LiveJournal1 (68M+ edges). One version uses Apache Spark (PySpark), the other Hadoop MapReduce streaming, so their runtime and developer experience can be compared directly.

A Neo4j project loads a social network in Docker and answers relationship questions in Cypher: connections, the most-connected people, and shortest paths.

Experiment

  1. 01

    Edge lists

    SNAP datasets, with comment lines skipped.

  2. 02

    Spark job

    PySpark

    Map destinations, then reduceByKey for in-degree.

  3. 03

    Hadoop job

    MapReduce streaming

    Mapper and reducer scripts, with a local fallback.

  4. 04

    Distribution

    Aggregate in-degree frequencies and compare runtimes.