Skip to content

Repository files navigation

🎬 Movie Recommendation System

A production-ready, AI-powered movie recommendation system built with Django and advanced machine learning. Scalable from thousands to millions of movies.

Python Django License


Logo Image


πŸ“‘ Table of Contents


🎯 Overview

The Movie Recommendation System provides intelligent movie suggestions using content-based filtering with TF-IDF and SVD dimensionality reduction. It features a modern web interface, RESTful API, and supports datasets from 2K to 1M+ movies.

Header Image

Why This Project?

  • βœ… Production Ready - Security hardened, optimized, well-documented
  • βœ… Scalable Architecture - Handles millions of movies efficiently
  • βœ… Modern Tech Stack - Django 6.0, Python 3.11+, advanced ML
  • βœ… Easy to Use - Simple installation, clear documentation
  • βœ… Flexible - Train on 10K or 1M+ movies from any TMDB-shaped CSV

Key Technologies

  • Backend: Django 6.0, Python 3.11+
  • ML/Data: scikit-learn, pandas, numpy, scipy
  • Storage: Parquet (efficient data format)
  • Deployment: Render, Heroku, Docker compatible

πŸ“Έ Screenshots & Demo

Demo Video

Application Demo

Model Loading

Model Loading

Home Page

Home Page

Movie Search Recommendations

Movie Recommendations


✨ Features

User Features

  • πŸ” Smart Search - Real-time autocomplete with fuzzy matching
  • 🎬 AI Recommendations - Content-based filtering with 15+ suggestions
  • ⭐ Rich Metadata - Ratings, votes, genres, production companies
  • πŸ”— External Links - Google Search and IMDb integration
  • πŸ“± Responsive Design - Works seamlessly on all devices
  • ⚑ Fast Performance - Sub-50ms recommendation generation

Technical Features

  • πŸ€– Advanced ML - TF-IDF + SVD dimensionality reduction
  • πŸ“Š Scalable - Handles 2K to 1M+ movies
  • πŸ’Ύ Efficient Storage - Parquet format with compression
  • πŸ”§ Configurable - Easy model switching via MODEL_DIR
  • πŸ“‘ REST API - JSON endpoints for integration
  • πŸ”’ Secure - Production-ready security settings
  • πŸ“ Logging - Comprehensive error tracking
  • πŸš€ Deployment Ready - render.yaml, Procfile and Dockerfile included

πŸš€ Quick Start

Prerequisites

  • Python 3.11 or higher (numpy 2.3+ requires it)
  • pip package manager
  • 8GB RAM (recommended for training)
  • Git

Installation

# 1. Clone the repository
git clone https://github.com/yourusername/movie-recommendation-system.git
cd movie-recommendation-system

# 2. Create virtual environment
python -m venv venv

# 3. Activate virtual environment
# Windows:
venv\Scripts\activate
# macOS/Linux:
source venv/bin/activate

# 4. Install dependencies
pip install -r requirements.txt

# 5. Run database migrations
python manage.py migrate

# 6. Start the development server
python manage.py runserver

Access the Application

Open your browser and navigate to:

http://localhost:8000

That's it β€” a 6,248-movie demo model ships in demo_model/ and is found automatically, so search works immediately with no configuration. πŸŽ‰

Want a bigger catalogue, or your own dataset? See Model Training β€” it is entirely optional.


πŸ“ Project Structure

movie-recommendation-system/
β”‚
β”œβ”€β”€ πŸ“š Documentation
β”‚   β”œβ”€β”€ README.md                  # This file - overview and quick start
β”‚   β”œβ”€β”€ PROJECT_GUIDE.md           # Complete technical guide
β”‚   └── CHANGELOG.md               # Version history and changes
β”‚
β”œβ”€β”€ βš™οΈ Django Application
β”‚   β”œβ”€β”€ movie_recommendation/      # Django project settings
β”‚   β”‚   β”œβ”€β”€ settings.py           # Configuration
β”‚   β”‚   β”œβ”€β”€ urls.py               # URL routing
β”‚   β”‚   └── wsgi.py               # WSGI entry point
β”‚   β”‚
β”‚   β”œβ”€β”€ recommender/              # Main application
β”‚   β”‚   β”œβ”€β”€ views.py              # Recommendation logic
β”‚   β”‚   β”œβ”€β”€ urls.py               # App URLs
β”‚   β”‚   └── templates/            # HTML templates
β”‚   β”‚       └── recommender/
β”‚   β”‚           β”œβ”€β”€ index.html    # Search page
β”‚   β”‚           β”œβ”€β”€ result.html   # Results page
β”‚   β”‚           └── error.html    # Error page
β”‚   β”‚
β”‚   β”œβ”€β”€ manage.py                 # Django management script
β”‚   └── requirements.txt          # Python dependencies
β”‚
β”œβ”€β”€ πŸŽ“ Model Training
β”‚   └── training/
β”‚       β”œβ”€β”€ train.py              # Training pipeline
β”‚       β”œβ”€β”€ infer.py              # Inference examples
β”‚       └── guide.md              # Training documentation
β”‚
β”œβ”€β”€ 🎯 Models (Created after training)
β”‚   β”œβ”€β”€ demo_model/               # Committed: 6,248 movies, ~3.8 MB
β”‚   β”‚   β”œβ”€β”€ movie_metadata.parquet    # Movie information
β”‚   β”‚   β”œβ”€β”€ neighbors_idx.npy         # Top-K neighbour indices
β”‚   β”‚   β”œβ”€β”€ neighbors_scores.npy      # Top-K neighbour scores
β”‚   β”‚   β”œβ”€β”€ title_to_idx.json         # Title mappings
β”‚   β”‚   └── config.json               # Model configuration
β”‚   β”‚
β”‚   └── models/                   # git-ignored; your own trained models
β”‚       └── ... same layout, plus *.pkl retraining artifacts
β”‚
β”œβ”€β”€ πŸ“¦ Static Files
β”‚   └── static/
β”‚       └── logo.ico                  # Application icon
β”‚
β”œβ”€β”€ πŸš€ Deployment
β”‚   β”œβ”€β”€ Procfile                  # Heroku configuration
β”‚   β”œβ”€β”€ render.yaml               # Render configuration
β”‚   β”œβ”€β”€ Dockerfile                # Container image
β”‚   β”œβ”€β”€ .dockerignore             # Build context exclusions
β”‚   └── .gitignore                # Git ignore rules
β”‚
└── βš™οΈ Project config
    β”œβ”€β”€ requirements.txt          # Runtime dependencies
    β”œβ”€β”€ requirements-train.txt    # Extra dependencies for training
    β”œβ”€β”€ .env.example              # Documented environment variables
    └── .github/workflows/ci.yml  # Checks and tests on every push

πŸ’‘ Usage

Web Interface

  1. Search for a Movie

    • Go to http://localhost:8000
    • Start typing a movie name in the search box
    • Select from autocomplete suggestions or type the full name
  2. View Recommendations

    • Click "Get Recommendations"
    • Browse 15 similar movie suggestions
    • Each card shows: rating, release date, genres, production company
  3. Explore Movies

    • Click "Google" to search for the movie
    • Click "IMDb" to view on IMDb (if available)

API Usage

Search Movies (Autocomplete)

GET /api/search/?q=matrix

Response:
{
  "movies": ["The Matrix", "The Matrix Reloaded", "The Matrix Revolutions"],
  "count": 3
}

Health Check

GET /api/health/

Response:
{
  "status": "healthy",
  "movies_loaded": 100000,
  "model_dir": "./models",
  "model_loaded": true
}

πŸŽ“ Model Training

Training a Model

Training is optional. The committed demo_model/ (6,248 movies with 500+ votes) covers most well-known films. Train your own for a larger catalogue, a different language, or your own dataset. Download the TMDB Movies Dataset (or any CSV with the same columns), then:

pip install -r requirements-train.txt

# Top 50,000 movies by quality score, 50 neighbours each (~20 MB output)
python training/train.py ./TMDB_movie_dataset_v11.csv -o models

# Smaller and faster: only movies with 500+ votes
python training/train.py ./TMDB_movie_dataset_v11.csv -o models -q high -m 10000

python training/train.py --help   # all options

The trainer stores the top K most similar movies per title rather than a full N x N similarity matrix. At 50,000 movies that is roughly 20 MB instead of 10 GB, which is what makes the free tier of most hosts viable.

Training Options

Want to train on more movies or your own dataset? See the Training Guide for:

  • πŸ“– Complete training documentation
  • 🎯 Configuration options (10K to 1M+ movies)
  • βš™οΈ Performance tuning guidelines
  • πŸ“Š Dataset requirements
  • πŸ”§ Advanced features

Using the trainer from Python:

from training.train import MovieRecommenderTrainer

trainer = MovieRecommenderTrainer(
    output_dir='./models',
    use_dimensionality_reduction=True,
    n_components=500,
    top_k=50,            # neighbours stored per movie
)

df, neighbors = trainer.train(
    'path/to/your/dataset.csv',
    quality_threshold='medium',  # low/medium/high
    max_movies=50000,
)

For detailed training instructions, see:


πŸ“‘ API Reference

Endpoints

Endpoint Method Description
/ GET Home page with search interface
/ POST Submit movie search and get recommendations
/api/search/ GET Search movies (autocomplete)
/api/model-status/ GET Model loading progress
/api/health/ GET Health check endpoint

Search Movies

Request:

GET /api/search/?q=inception

Response:

{
  "movies": ["Inception", "Inception: The Cobol Job"],
  "count": 2
}

Health Check

Request:

GET /api/health/

Response:

{
  "status": "healthy",
  "movies_loaded": 100000,
  "model_dir": "./models",
  "model_loaded": true
}

For complete API documentation, see PROJECT_GUIDE.md - API Reference


βš™οΈ Configuration

Environment Variables

Create a .env file (optional for development):

# Django Settings
SECRET_KEY=your-secret-key-here
DEBUG=True
ALLOWED_HOSTS=localhost,127.0.0.1

# Model Configuration
MODEL_DIR=./models

# Database (optional - defaults to SQLite)
# DATABASE_URL=postgresql://user:password@localhost/dbname

# Deployment
# RENDER_EXTERNAL_HOSTNAME=your-app.onrender.com

Using Different Models

To switch between models, set the MODEL_DIR environment variable:

# Use your trained model (the default)
export MODEL_DIR=models

# Use absolute path
export MODEL_DIR=/path/to/your/models

For detailed configuration options, see PROJECT_GUIDE.md - Configuration


πŸ“š Documentation

Main Documentation

  • README.md (this file) - Overview, quick start, basic usage
  • PROJECT_GUIDE.md - Complete technical guide
    • Installation
    • Model training
    • Configuration
    • Development
    • Deployment
    • API reference
    • Troubleshooting
  • CHANGELOG.md - Version history and changes

Training Documentation

  • training/guide.md - Complete model training guide
    • Dataset requirements
    • Training configurations
    • Performance tuning
    • Advanced features

Quick Links

Topic Documentation
Installation Quick Start or PROJECT_GUIDE.md
Model Training training/guide.md
Deployment PROJECT_GUIDE.md - Deployment
API Reference API Reference or PROJECT_GUIDE.md
Troubleshooting PROJECT_GUIDE.md - Troubleshooting
Configuration Configuration or PROJECT_GUIDE.md

πŸš€ Deployment

Deployments work as-is. The committed demo_model/ is found automatically, so a deploy from a fresh clone serves recommendations without any model configuration. The options below are only for shipping a larger model than the demo.

Shipping a bigger model (optional)

Option A β€” replace the demo model (simplest). Train into demo_model/ (or any directory that is not git-ignored) and commit:

pip install -r requirements-train.txt
python training/train.py ./TMDB_movie_dataset_v11.csv -o demo_model -q high -m 10000
git add demo_model && git commit -m "Add demo model"

That is all β€” demo_model/ is one of the locations checked automatically, so you can leave MODEL_DIR unset (or set it explicitly if you prefer).

Two things keep this small: models/ is git-ignored on purpose so full-size models never enter history, while demo_model/ is not; and *.pkl is ignored repository-wide, so the retraining artifacts (an SVD pickle is easily 10Γ— the rest of the model combined) are skipped by git add automatically. Nothing loaded at serving time is a pickle.

Option B β€” fetch it at build time (keeps the repo small). Train a model, package it, and upload model.tar.gz as a GitHub Release asset:

python training/train.py ./TMDB_movie_dataset_v11.csv -o models
tar czf model.tar.gz -C models --exclude='*.pkl' .   # paths relative; no pickles

Set MODEL_URL to the asset URL. render.yaml's build step downloads and unpacks it into MODEL_DIR. This is build-time only β€” the app never fetches anything at runtime.

Render

  1. Push to GitHub and connect the repository to Render
  2. Render picks up render.yaml automatically
  3. Set ALLOWED_HOSTS (your custom domain, if any) and MODEL_URL (option B)
  4. Deploy. SECRET_KEY is generated for you.

Heroku

Uses the Procfile. Commit a model (option A) and set config vars:

heroku config:set DEBUG=False
heroku config:set MODEL_DIR=demo_model
heroku config:set ALLOWED_HOSTS=your-app.herokuapp.com
heroku config:set SECRET_KEY="$(python -c 'import secrets; print(secrets.token_urlsafe(50))')"

Docker

The image bakes in whatever is in MODEL_DIR at build time, so train first:

python training/train.py ./TMDB_movie_dataset_v11.csv -o models
docker build -t movie-recommender .
export SECRET_KEY="$(python -c 'import secrets; print(secrets.token_urlsafe(50))')"
docker run --rm -p 8000:8000 \
  -e SECRET_KEY="$SECRET_KEY" \
  -e ALLOWED_HOSTS=localhost,127.0.0.1 \
  -e SECURE_SSL_REDIRECT=False \
  movie-recommender

SECURE_SSL_REDIRECT=False is required when nothing is terminating TLS in front of the container; otherwise every request is redirected to https:// and loops.

Deployment checklist

  • SECRET_KEY set (the app refuses to start with the dev key when DEBUG=False)
  • DEBUG=False
  • ALLOWED_HOSTS includes your domain
  • A model is reachable β€” curl https://your-app/api/health/ returns "status": "healthy"

For more detail, see PROJECT_GUIDE.md - Deployment


🀝 Contributing

Contributions are welcome. See CONTRIBUTING.md for setup, the layout of the code, and the handful of non-obvious things worth knowing before changing the model pipeline.

The short version:

pip install -r requirements.txt
python manage.py migrate
python manage.py runserver        # a demo model ships, so this just works

# before opening a PR -- CI runs the same three
python manage.py check
python manage.py test
ruff check .

Style is configured in pyproject.toml; ruff check . --fix handles the mechanical parts. Please add a test with a bug fix and note the change under [Unreleased] in CHANGELOG.md.


🎞️ Data & attribution

Movie metadata and poster images come from TMDB.

This product uses data from TMDB but is not endorsed or certified by TMDB.

  • The committed demo_model/ is derived from the TMDB Movies Dataset and contains titles, overviews, ratings and poster paths originating from TMDB.
  • Poster images are loaded directly from image.tmdb.org at render time; none are redistributed in this repository.
  • If you extend this project to call the TMDB API directly, their terms require the wording "This product uses the TMDB API but is not endorsed or certified by TMDB" together with the TMDB logo. See TMDB's terms of use.

The MIT licence below covers this project's own source code, not the movie data.


πŸ“„ License

This project is licensed under the MIT License - see the LICENSE file for details.


πŸ†˜ Support

Need help? Here are your options:


🎯 Roadmap

Version 2.1 (Planned)

  • User authentication system
  • Personal watchlists
  • Movie rating system
  • Advanced filtering (multiple genres, year ranges)
  • Recommendation history

Version 2.2 (Planned)

  • Collaborative filtering
  • Social features (sharing, comments)
  • Movie reviews
  • Advanced analytics dashboard

Version 3.0 (Long-term)

  • Mobile applications (iOS/Android)
  • Real-time recommendations
  • Streaming service integration
  • Enhanced ML models (hybrid recommendations)

πŸ“Š Performance

Metric Value
Recommendation Time < 50ms
Search Response < 100ms
Page Load < 200ms
Memory Usage ~200MB (100K movies)
Concurrent Users 1000+
Model Size 180MB (100K movies)

πŸ™ Acknowledgments

  • Movie data from TMDB and IMDb
  • Built with Django, scikit-learn, pandas
  • UI inspired by modern design principles
  • Community contributions and feedback

Made with ❀️ for movie lovers and developers

⭐ Star this repo β€’ πŸ› Report Bug β€’ πŸ’‘ Request Feature