Big data · Graphs
Big Data Graph Processing
The same graph problem on two big data engines, plus a native graph database for relationship queries.
- Role
- Data Engineer
- Year
- 2025
- Focus
- Big data
- Stack
- 4 technologies
- PySpark
- Hadoop
- Neo4j
- Cypher
68M+
Edges (LiveJournal)
3
SNAP graphs
2
Engines benchmarked
Overview
In-degree distributions are computed for three SNAP datasets: email-EuAll, web-BerkStan, and soc-LiveJournal1 (68M+ edges). One version uses Apache Spark (PySpark), the other Hadoop MapReduce streaming, so their runtime and developer experience can be compared directly.
A Neo4j project loads a social network in Docker and answers relationship questions in Cypher: connections, the most-connected people, and shortest paths.
Architecture
Experiment
- 01
Edge lists
SNAP datasets, with comment lines skipped.
- 02
Spark job
PySparkMap destinations, then reduceByKey for in-degree.
- 03
Hadoop job
MapReduce streamingMapper and reducer scripts, with a local fallback.
- 04
Distribution
Aggregate in-degree frequencies and compare runtimes.