A production-ready, AI-powered movie recommendation system built with Django and advanced machine learning. Scalable from thousands to millions of movies.
- Overview
- Screenshots
- Features
- Quick Start
- Project Structure
- Usage
- Model Training
- API Reference
- Configuration
- Documentation
- Contributing
- License
The Movie Recommendation System provides intelligent movie suggestions using content-based filtering with TF-IDF and SVD dimensionality reduction. It features a modern web interface, RESTful API, and supports datasets from 2K to 1M+ movies.
- β Production Ready - Security hardened, optimized, well-documented
- β Scalable Architecture - Handles millions of movies efficiently
- β Modern Tech Stack - Django 6.0, Python 3.11+, advanced ML
- β Easy to Use - Simple installation, clear documentation
- β Flexible - Train on 10K or 1M+ movies from any TMDB-shaped CSV
- Backend: Django 6.0, Python 3.11+
- ML/Data: scikit-learn, pandas, numpy, scipy
- Storage: Parquet (efficient data format)
- Deployment: Render, Heroku, Docker compatible
- π Smart Search - Real-time autocomplete with fuzzy matching
- π¬ AI Recommendations - Content-based filtering with 15+ suggestions
- β Rich Metadata - Ratings, votes, genres, production companies
- π External Links - Google Search and IMDb integration
- π± Responsive Design - Works seamlessly on all devices
- β‘ Fast Performance - Sub-50ms recommendation generation
- π€ Advanced ML - TF-IDF + SVD dimensionality reduction
- π Scalable - Handles 2K to 1M+ movies
- πΎ Efficient Storage - Parquet format with compression
- π§ Configurable - Easy model switching via
MODEL_DIR - π‘ REST API - JSON endpoints for integration
- π Secure - Production-ready security settings
- π Logging - Comprehensive error tracking
- π Deployment Ready -
render.yaml,ProcfileandDockerfileincluded
- Python 3.11 or higher (numpy 2.3+ requires it)
- pip package manager
- 8GB RAM (recommended for training)
- Git
# 1. Clone the repository
git clone https://github.com/yourusername/movie-recommendation-system.git
cd movie-recommendation-system
# 2. Create virtual environment
python -m venv venv
# 3. Activate virtual environment
# Windows:
venv\Scripts\activate
# macOS/Linux:
source venv/bin/activate
# 4. Install dependencies
pip install -r requirements.txt
# 5. Run database migrations
python manage.py migrate
# 6. Start the development server
python manage.py runserverOpen your browser and navigate to:
http://localhost:8000
That's it β a 6,248-movie demo model ships in demo_model/ and is found
automatically, so search works immediately with no configuration. π
Want a bigger catalogue, or your own dataset? See Model Training β it is entirely optional.
movie-recommendation-system/
β
βββ π Documentation
β βββ README.md # This file - overview and quick start
β βββ PROJECT_GUIDE.md # Complete technical guide
β βββ CHANGELOG.md # Version history and changes
β
βββ βοΈ Django Application
β βββ movie_recommendation/ # Django project settings
β β βββ settings.py # Configuration
β β βββ urls.py # URL routing
β β βββ wsgi.py # WSGI entry point
β β
β βββ recommender/ # Main application
β β βββ views.py # Recommendation logic
β β βββ urls.py # App URLs
β β βββ templates/ # HTML templates
β β βββ recommender/
β β βββ index.html # Search page
β β βββ result.html # Results page
β β βββ error.html # Error page
β β
β βββ manage.py # Django management script
β βββ requirements.txt # Python dependencies
β
βββ π Model Training
β βββ training/
β βββ train.py # Training pipeline
β βββ infer.py # Inference examples
β βββ guide.md # Training documentation
β
βββ π― Models (Created after training)
β βββ demo_model/ # Committed: 6,248 movies, ~3.8 MB
β β βββ movie_metadata.parquet # Movie information
β β βββ neighbors_idx.npy # Top-K neighbour indices
β β βββ neighbors_scores.npy # Top-K neighbour scores
β β βββ title_to_idx.json # Title mappings
β β βββ config.json # Model configuration
β β
β βββ models/ # git-ignored; your own trained models
β βββ ... same layout, plus *.pkl retraining artifacts
β
βββ π¦ Static Files
β βββ static/
β βββ logo.ico # Application icon
β
βββ π Deployment
β βββ Procfile # Heroku configuration
β βββ render.yaml # Render configuration
β βββ Dockerfile # Container image
β βββ .dockerignore # Build context exclusions
β βββ .gitignore # Git ignore rules
β
βββ βοΈ Project config
βββ requirements.txt # Runtime dependencies
βββ requirements-train.txt # Extra dependencies for training
βββ .env.example # Documented environment variables
βββ .github/workflows/ci.yml # Checks and tests on every push
-
Search for a Movie
- Go to
http://localhost:8000 - Start typing a movie name in the search box
- Select from autocomplete suggestions or type the full name
- Go to
-
View Recommendations
- Click "Get Recommendations"
- Browse 15 similar movie suggestions
- Each card shows: rating, release date, genres, production company
-
Explore Movies
- Click "Google" to search for the movie
- Click "IMDb" to view on IMDb (if available)
GET /api/search/?q=matrix
Response:
{
"movies": ["The Matrix", "The Matrix Reloaded", "The Matrix Revolutions"],
"count": 3
}GET /api/health/
Response:
{
"status": "healthy",
"movies_loaded": 100000,
"model_dir": "./models",
"model_loaded": true
}Training is optional. The committed demo_model/ (6,248 movies with 500+
votes) covers most well-known films. Train your own for a larger catalogue, a
different language, or your own dataset. Download the
TMDB Movies Dataset
(or any CSV with the same columns), then:
pip install -r requirements-train.txt
# Top 50,000 movies by quality score, 50 neighbours each (~20 MB output)
python training/train.py ./TMDB_movie_dataset_v11.csv -o models
# Smaller and faster: only movies with 500+ votes
python training/train.py ./TMDB_movie_dataset_v11.csv -o models -q high -m 10000
python training/train.py --help # all optionsThe trainer stores the top K most similar movies per title rather than a full N x N similarity matrix. At 50,000 movies that is roughly 20 MB instead of 10 GB, which is what makes the free tier of most hosts viable.
Want to train on more movies or your own dataset? See the Training Guide for:
- π Complete training documentation
- π― Configuration options (10K to 1M+ movies)
- βοΈ Performance tuning guidelines
- π Dataset requirements
- π§ Advanced features
Using the trainer from Python:
from training.train import MovieRecommenderTrainer
trainer = MovieRecommenderTrainer(
output_dir='./models',
use_dimensionality_reduction=True,
n_components=500,
top_k=50, # neighbours stored per movie
)
df, neighbors = trainer.train(
'path/to/your/dataset.csv',
quality_threshold='medium', # low/medium/high
max_movies=50000,
)For detailed training instructions, see:
- π Training Guide - Complete training documentation
- π PROJECT_GUIDE.md - Training setup and configurations
| Endpoint | Method | Description |
|---|---|---|
/ |
GET | Home page with search interface |
/ |
POST | Submit movie search and get recommendations |
/api/search/ |
GET | Search movies (autocomplete) |
/api/model-status/ |
GET | Model loading progress |
/api/health/ |
GET | Health check endpoint |
Request:
GET /api/search/?q=inceptionResponse:
{
"movies": ["Inception", "Inception: The Cobol Job"],
"count": 2
}Request:
GET /api/health/Response:
{
"status": "healthy",
"movies_loaded": 100000,
"model_dir": "./models",
"model_loaded": true
}For complete API documentation, see PROJECT_GUIDE.md - API Reference
Create a .env file (optional for development):
# Django Settings
SECRET_KEY=your-secret-key-here
DEBUG=True
ALLOWED_HOSTS=localhost,127.0.0.1
# Model Configuration
MODEL_DIR=./models
# Database (optional - defaults to SQLite)
# DATABASE_URL=postgresql://user:password@localhost/dbname
# Deployment
# RENDER_EXTERNAL_HOSTNAME=your-app.onrender.comTo switch between models, set the MODEL_DIR environment variable:
# Use your trained model (the default)
export MODEL_DIR=models
# Use absolute path
export MODEL_DIR=/path/to/your/modelsFor detailed configuration options, see PROJECT_GUIDE.md - Configuration
- README.md (this file) - Overview, quick start, basic usage
- PROJECT_GUIDE.md - Complete technical guide
- Installation
- Model training
- Configuration
- Development
- Deployment
- API reference
- Troubleshooting
- CHANGELOG.md - Version history and changes
- training/guide.md - Complete model training guide
- Dataset requirements
- Training configurations
- Performance tuning
- Advanced features
| Topic | Documentation |
|---|---|
| Installation | Quick Start or PROJECT_GUIDE.md |
| Model Training | training/guide.md |
| Deployment | PROJECT_GUIDE.md - Deployment |
| API Reference | API Reference or PROJECT_GUIDE.md |
| Troubleshooting | PROJECT_GUIDE.md - Troubleshooting |
| Configuration | Configuration or PROJECT_GUIDE.md |
Deployments work as-is. The committed
demo_model/is found automatically, so a deploy from a fresh clone serves recommendations without any model configuration. The options below are only for shipping a larger model than the demo.
Option A β replace the demo model (simplest).
Train into demo_model/ (or any directory that is not git-ignored) and commit:
pip install -r requirements-train.txt
python training/train.py ./TMDB_movie_dataset_v11.csv -o demo_model -q high -m 10000
git add demo_model && git commit -m "Add demo model"That is all β demo_model/ is one of the locations checked automatically, so
you can leave MODEL_DIR unset (or set it explicitly if you prefer).
Two things keep this small: models/ is git-ignored on purpose so full-size
models never enter history, while demo_model/ is not; and *.pkl is ignored
repository-wide, so the retraining artifacts (an SVD pickle is easily 10Γ the
rest of the model combined) are skipped by git add automatically. Nothing
loaded at serving time is a pickle.
Option B β fetch it at build time (keeps the repo small).
Train a model, package it, and upload model.tar.gz as a GitHub Release asset:
python training/train.py ./TMDB_movie_dataset_v11.csv -o models
tar czf model.tar.gz -C models --exclude='*.pkl' . # paths relative; no picklesSet MODEL_URL to the asset URL. render.yaml's build step downloads and
unpacks it into MODEL_DIR. This is build-time only β the app never fetches
anything at runtime.
- Push to GitHub and connect the repository to Render
- Render picks up
render.yamlautomatically - Set
ALLOWED_HOSTS(your custom domain, if any) andMODEL_URL(option B) - Deploy.
SECRET_KEYis generated for you.
Uses the Procfile. Commit a model (option A) and set config vars:
heroku config:set DEBUG=False
heroku config:set MODEL_DIR=demo_model
heroku config:set ALLOWED_HOSTS=your-app.herokuapp.com
heroku config:set SECRET_KEY="$(python -c 'import secrets; print(secrets.token_urlsafe(50))')"The image bakes in whatever is in MODEL_DIR at build time, so train first:
python training/train.py ./TMDB_movie_dataset_v11.csv -o models
docker build -t movie-recommender .
export SECRET_KEY="$(python -c 'import secrets; print(secrets.token_urlsafe(50))')"
docker run --rm -p 8000:8000 \
-e SECRET_KEY="$SECRET_KEY" \
-e ALLOWED_HOSTS=localhost,127.0.0.1 \
-e SECURE_SSL_REDIRECT=False \
movie-recommenderSECURE_SSL_REDIRECT=False is required when nothing is terminating TLS in
front of the container; otherwise every request is redirected to https://
and loops.
-
SECRET_KEYset (the app refuses to start with the dev key whenDEBUG=False) -
DEBUG=False -
ALLOWED_HOSTSincludes your domain - A model is reachable β
curl https://your-app/api/health/returns"status": "healthy"
For more detail, see PROJECT_GUIDE.md - Deployment
Contributions are welcome. See CONTRIBUTING.md for setup, the layout of the code, and the handful of non-obvious things worth knowing before changing the model pipeline.
The short version:
pip install -r requirements.txt
python manage.py migrate
python manage.py runserver # a demo model ships, so this just works
# before opening a PR -- CI runs the same three
python manage.py check
python manage.py test
ruff check .Style is configured in pyproject.toml; ruff check . --fix handles the
mechanical parts. Please add a test with a bug fix and note the change under
[Unreleased] in CHANGELOG.md.
Movie metadata and poster images come from TMDB.
This product uses data from TMDB but is not endorsed or certified by TMDB.
- The committed
demo_model/is derived from the TMDB Movies Dataset and contains titles, overviews, ratings and poster paths originating from TMDB. - Poster images are loaded directly from
image.tmdb.orgat render time; none are redistributed in this repository. - If you extend this project to call the TMDB API directly, their terms require the wording "This product uses the TMDB API but is not endorsed or certified by TMDB" together with the TMDB logo. See TMDB's terms of use.
The MIT licence below covers this project's own source code, not the movie data.
This project is licensed under the MIT License - see the LICENSE file for details.
Need help? Here are your options:
- π Documentation: Check PROJECT_GUIDE.md for detailed guides
- π Training Help: See training/guide.md for model training
- π Issues: Open an issue on GitHub
- π¬ Discussions: GitHub Discussions
- User authentication system
- Personal watchlists
- Movie rating system
- Advanced filtering (multiple genres, year ranges)
- Recommendation history
- Collaborative filtering
- Social features (sharing, comments)
- Movie reviews
- Advanced analytics dashboard
- Mobile applications (iOS/Android)
- Real-time recommendations
- Streaming service integration
- Enhanced ML models (hybrid recommendations)
| Metric | Value |
|---|---|
| Recommendation Time | < 50ms |
| Search Response | < 100ms |
| Page Load | < 200ms |
| Memory Usage | ~200MB (100K movies) |
| Concurrent Users | 1000+ |
| Model Size | 180MB (100K movies) |
- Movie data from TMDB and IMDb
- Built with Django, scikit-learn, pandas
- UI inspired by modern design principles
- Community contributions and feedback
Made with β€οΈ for movie lovers and developers
β Star this repo β’ π Report Bug β’ π‘ Request Feature





