ML & Data Science

Graph database · Cypher

Neo4j Social Graph

MSc Data Science · Big Data coursework (Task 3) · Coventry University

When the relationships matter more than the records, a graph database answers in one hop what a relational database needs a stack of JOINs for.

Role
Data Engineer
Year
2025
Focus
Graph database
Stack
4 technologies
  • Neo4j 5
  • Cypher
  • Docker
  • cypher-shell
  • 15

    User nodes

  • 23

    FOLLOWS relationships

  • 12 / 14

    Users within 3 hops of Alice

  • 3 hops

    Shortest path, Alice → Mallory

In a relational database, questions like 'friends of friends within 3 hops' or 'the shortest path between two people' need repeated self-JOINs that get slower as the network grows. Graph databases store each relationship directly on its nodes, so traversals stay fast no matter how connected the data is.

Neo4j 5 runs in Docker, either as a single container or through Docker Compose with persistent volumes for data, logs, and imports. The Neo4j Browser is on port 7474 and the Bolt protocol on 7687.

The dataset is a small social network of 15 users across 7 Sri Lankan cities (Colombo, Galle, Kandy, Jaffna, Negombo, Matara, Anuradhapura), each with a city and a join year from 2020 to 2024, linked by 23 directed FOLLOWS relationships. The seed script uses UNWIND and MERGE so it is idempotent: running it twice never creates duplicates. A cleanup script resets the graph.

Three parameterised Cypher queries answer real social-network questions: who is connected to a user within 3 hops, who has the most connections, and the shortest path between two people. Three further queries validate the model by listing every user, city, and join year.

Workflow

  1. 01

    Run Neo4j

    Docker

    docker run neo4j or docker compose up with the neo4j:5.23-community image, ports 7474 and 7687.

  2. 02

    Reset

    cleanup.cypher

    Remove all nodes and relationships so every run starts from a clean graph.

  3. 03

    Seed

    seed.cypher

    UNWIND a list of users and MERGE them with city and joinYear, then MERGE the FOLLOWS edges.

  4. 04

    Verify

    count(u) and count(r) in cypher-shell confirm 15 users and 23 relationships.

  5. 05

    Query

    cypher-shell --param

    Run the three parameterised queries from the shell or in the Browser with :param.

Screenshots from the Neo4j Browser after seeding and running the queries. Tap any figure to zoom.

  1. Fig. 01 · Graph view

    The social network as a graph

    Each circle is a User node and each arrow a FOLLOWS relationship from follower to followed. The network is accurately modelled as a directed graph.

    • A dense community sits in the middle (Alice, Bob, Carol, Eve, Frank, Dave), where many users follow each other.
    • Peripheral users such as Trent, Peggy, Mallory, and Niaj connect only through longer chains.
  2. Fig. 02 · Cypher · UNION ALL

    Validate the model: all users

    Returns every distinct name property on nodes (and, via UNION ALL, on relationships if any exist). All 15 expected users are present.

  3. Fig. 03 · Node properties

    Where users live

    Seven distinct cities. Properties stored on nodes support location-based queries such as 'find all users from Colombo' or regional recommendations.

  4. Fig. 04 · Temporal properties

    When users joined

    Join years from 2020 to 2024 show how the network grew over time and allow comparing early adopters with newer users.

  • Q1 · connected users

    variable-length path

    MATCH p = (target)-[*..3]-(other) and return min(length(p)) per user: everyone reachable within 3 hops and how far away they are.

  • Q2 · most connected

    degree

    Two OPTIONAL MATCHes count outgoing and incoming FOLLOWS per user, then rank by total degree (in + out).

  • Q3 · shortest path

    shortestPath()

    shortestPath((a)-[:FOLLOWS*..15]-(b)) between two named users, bounded to avoid expansion blow-up.

  • Recommendation query

    collaborative filtering

    Find items liked by users who share likes with User 501, exclude items they already like, and rank by how many similar users liked them.

Most connected users

Q2: total degree = followers + following

  • Alice (4 in · 2 out)6
  • Bob (4 in · 2 out)6
  • Carol (3 in · 2 out)5
  • Eve (2 in · 2 out)4
  • Frank (2 in · 2 out)4

Users reachable from Alice

Q1: closest distance, within 3 hops

  • 1 hop6
  • 2 hops4
  • 3 hops2
  • Q1 (Alice, 3 hops): Bob, Carol, Dave, Eve, Frank and Judy at 1 hop; Grace, Heidi, Ivan and Olivia at 2; Mallory and Niaj at 3. Only Peggy and Trent sit further out.
  • Q2: Alice and Bob tie as the most connected (degree 6, each followed by 4 users), followed by Carol (5), Eve (4) and Frank (4).
  • Q3 (Alice → Mallory): Alice → Eve ← Heidi ← Mallory, a 3-hop path found by treating FOLLOWS as undirected.
  • Social networks: 'friends of friends within 3 hops', 'people you may know', and finding influencers or communities are direct traversals instead of repeated JOINs.
  • Fraud detection: accounts, devices, IPs, cards, and transactions become nodes, so rings of accounts sharing one device or IP show up as connected clusters within a few hops.
  • Recommendations: users, items, and genres linked by WATCHED, LIKED, PURCHASED, and SIMILAR_TO make 'users who liked this also liked that' a short walk through the graph.
  • Nodes: users, items (movies, products, songs), and categories or genres.
  • Relationships: [:VIEWED], [:PURCHASED], [:RATED] and [:LIKED] for interactions; [:BELONGS_TO_GENRE] and [:IN_CATEGORY] for content; [:SIMILAR_TO] from shared features or co-purchases.
  • Collaborative filtering becomes one query: start at the user, walk to items they liked, out to other users who liked those items, then to new items, and rank by how many similar users liked each one.