Neo4j: Building a Movie-Director Graph
This notebook demonstrates a complete end-to-end application that builds a graph database of movies and directors from the Netflix dataset, then visualizes the relationships.
Workflow:
- Load the Netflix CSV dataset with Pandas
- Clean and preprocess the data
- Populate a Neo4j graph database using the
Neo4jAPIwrapper - Visualize the director-movie graph with NetworkX
Why graphs for this? Graphs are an intuitive way to represent relationships. Each movie is connected to its director, forming a natural web. Using Neo4j, we can query and explore this data efficiently, and by visualizing it we reveal connections that might otherwise be hard to see in rows and columns.
%load_ext autoreload
%autoreload 2
%matplotlib inlineimport logging
import pandas as pd
import helpers.hdbg as hdbg
import helpers.hnotebook as hnotebo
import tutorials.tutorial_neo4j.neo4j_utils as ttneouti
hdbg.init_logger(verbosity=logging.INFO)
_LOG = logging.getLogger(__name__)
hnotebo.config_notebook()Part 1: Data Loading and Inspection¶
Load the Netflix dataset from artifacts/data/netflix.csv. The dataset
contains movie and TV show metadata including title, director, cast, country,
and release year.
Sample structure:
| title | release_year | director | cast | country |
|--------------|--------------|---------------|---------------|---------|
| Movie Title | 2021 | John Doe | Actor A, ... | USA |# Load the dataset into a Pandas DataFrame.
csv_file = "artifacts/data/netflix.csv"
data = pd.read_csv(csv_file)
_LOG.info("Dataset shape: %s", data.shape)
data.head()Part 2: Data Cleaning and Preprocessing¶
Before loading into Neo4j we need clean, unambiguous data:
- Fill empty
castfields with empty strings (cast is optional) - Drop rows with missing
directororcountry(required for graph edges) - Strip whitespace from
titleto avoid duplicate nodes
# Fill missing cast values.
data["cast"] = data["cast"].fillna("")
# Remove rows with missing director or country.
data = data.dropna(subset=["director", "country"])
# Strip whitespace from title to avoid duplicates.
data["title"] = data["title"].str.strip()
_LOG.info("Cleaned dataset shape: %s", data.shape)
data.head()Part 3: Graph Construction with Neo4j¶
The Neo4jAPI wrapper class handles connection management and data loading.
It uses MERGE instead of CREATE to prevent duplicate nodes:
MERGE (movie:Movie {title: $title, release_year: $release_year})
MERGE (director:Director {name: $director})
MERGE (director)-[:DIRECTED]->(movie)This creates:
- Movie nodes with
titleandrelease_yearproperties - Director nodes with a
nameproperty - [:DIRECTED] relationships from director to movie
# Start the Neo4j server.
!sudo neo4j start# Initialize the Neo4j API wrapper.
neo4j_api = ttneouti.Neo4jAPI(
uri="neo4j://localhost:7687",
user="neo4j",
password="new_password",
)# Load the first 40 rows into Neo4j (lightweight demo).
neo4j_api.load_data(data[:40])
_LOG.info("Data loaded into Neo4j.")Part 4: Graph Visualization¶
Query the database and render the Director → Movie graph using NetworkX:
MATCH (d:Director)-[r:DIRECTED]->(m:Movie)
WHERE d.name <> 'Unknown'
RETURN d.name AS director, m.title AS movie, m.release_year AS yearVisualization style:
- Blue nodes: Directors
- Green nodes: Movies (with release year)
- Arrows: DIRECTED relationships
- Layout: Spring layout for optimal spacing
# Generate the visualization.
neo4j_api.visualize_graph()Clean Up¶
# Close the Neo4j connection.
neo4j_api.close()
_LOG.info("Connection closed.")