This repository contains a full-stack, secure AI platform deployed on Google Cloud via Terraform. It enables secure, RAG-based chat with enterprise-grade security and automated PII protection, accessible to public users via Firebase Authentication.
Figure 1: Google Cloud Platform Architecture
Figure 2: Correspondent AWS Architecture
- Location:
/frontend-nextjs - Tech: React 18, Tailwind CSS, Lucide Icons, Firebase.
- Resilience: Circuit Breaker
opossumfor fail-fast backend communication. - Security: Acts as a secure proxy to the Backend; Authentication handled via Firebase.
- Scalability: Configured with
min_instances = 1for zero-latency response. - Reliability: Implements circuit breakers (opossum) for external API calls.
- Testing: Comprehensive coverage with Jest (Unit) and Playwright (E2E).
The frontend is a modern Next.js application using the App Router with the following structure:
Core Components:
AuthProvider.tsx: The cornerstone of the application's authentication system. Uses Firebase Authentication to manage user authentication state and protect routes from unauthorized access.ChatInterface.tsx: The main chat interface providing a real-time, streaming chat experience. Tightly integrated with the backend API, it handles authentication and payment-related errors gracefully, redirecting users to the appropriate page when necessary.PaymentClient.tsx: Provides a seamless and secure payment experience using Stripe's Embedded Checkout. Guides users through the payment process with graceful error handling.
Routing Structure:
| Route | Description |
|---|---|
/ |
Landing page - main entry point, funnels users to payment |
/chat |
Main chat interface |
/login |
Login page |
/payment |
Payment page |
/payment-success |
Displayed after successful payment |
API Routes (Server-Side):
/api/chat: Acts as a secure and resilient proxy to the backend chat service. Uses a circuit breaker to prevent cascading failures and OIDC tokens for secure service-to-service authentication. Streams responses from the backend to the client for real-time chat./api/check-payment-status: Checks the status of a Stripe Checkout session and sets a cookie to persist payment status client-side./api/create-checkout-session: Creates a Stripe Checkout session and returns theclient_secretfor displaying the Stripe payment form.
Security Features:
- Secure handling of API keys and secrets
- OIDC tokens for service-to-service authentication
- User authentication tokens forwarded to backend for authorization
- Location:
/backend-agent - Neural Core: Orchestrates RAG using LangChain and Vertex AI.
- Resilience: Retries (
tenacity) for transient errors & OpenTelemetry tracing. - Vector DB: Cloud SQL for PostgreSQL 16 with
pgvectorandasyncpg. - Networking: Set to
INGRESS_TRAFFIC_INTERNAL_ONLYto ensure it is unreachable from the public internet. - Observability: Full OpenTelemetry instrumentation (Traces exported to Google Cloud Trace).
- Security: Rate limiting (slowapi) and integration with Google Cloud DLP (Data Loss Prevention) suggests a focus on enterprise compliance.
The backend is a robust FastAPI application with a strong focus on security, observability, and scalability.
API Endpoints:
| Endpoint | Description |
|---|---|
/health |
Standard health check endpoint |
/webhook |
Stripe webhook handler for managing user subscriptions |
/chat |
Primary, non-streaming chat endpoint |
/stream |
Streaming version of chat endpoint for real-time communication |
Security Implementation:
- Rate Limiting: API is rate-limited to 10 requests per minute per IP address to prevent abuse.
- Input Validation: Message size validation to prevent Denial-of-Service (DoS) attacks.
- Authentication: The
get_current_userdependency ensures all chat requests are authenticated. - IDOR Prevention: Session IDs are scoped to authenticated users, preventing cross-user session access.
Observability:
- Structured Logging: Provides clear and actionable log data.
- Distributed Tracing: OpenTelemetry integration for monitoring and debugging in microservices architecture.
Data Layer:
- Data Model: PostgreSQL database stores user information including subscription status and Stripe customer ID. The
Usermodel is defined using SQLAlchemy. - CRUD Operations: All database operations (Create, Read, Update, Delete) are encapsulated in
crud.pyfor maintainability and testability. - Subscription Management: Tight integration with Stripe; the
/webhookendpoint listens for Stripe events and updates user subscription status. - Secure Database Connection: Database URL fetched from Google Secret Manager.
Knowledge Base (RAG Pipeline):
The ingest.py script builds and maintains the knowledge base:
- Document Ingestion: Ingests PDF documents from the
./datadirectory usingDirectoryLoaderandPyPDFLoaderfrom LangChain. - Text Chunking: Uses
RecursiveCharacterTextSplitterto break documents into manageable chunks for efficient RAG retrieval. - Vector Embeddings: Uses
textembedding-gecko@003model from Vertex AI to generate semantic vector embeddings. - Vector Store: Uses
PGVectorto store document chunks and embeddings in PostgreSQL (AlloyDB) for scalable similarity searches. - Data Upsert: The
add_documentsfunction upserts data to keep the knowledge base up-to-date.
Core AI Capabilities:
- Intelligent Query Routing: Distinguishes between general conversation and knowledge-base queries
- Retrieval-Augmented Generation (RAG): Retrieves relevant information from PDF documents for accurate answers
- Conversational Memory: Remembers previous turns for context-aware responses
- Multi-Layered Security: Protection against common web vulnerabilities
- Data Loss Prevention (DLP): Prevents leakage of sensitive information
- Resilient Design: Handles transient errors and network failures
cicd: CI/CD pipeline.network: VPC, Subnets, Cloud NAT, and PSA.compute: Cloud Run services and granular IAM policies.database: Cloud SQL (PostgreSQL) and Firestore (Chat History).redis: Memorystore for semantic caching.ingress: Global Load Balancer, Cloud Armor, and SSL.billing_monitoring: Budgets, Alert Policies, and Notification Channels.function: Google Cloud Functions for PDF Ingestion.storage: Buckets and lifecycle policies.
The infrastructure is defined as code (IaC) using modular Terraform, adhering to Google Cloud best practices:
- Compute: Decoupled frontend and backend services (likely Cloud Run) and event-driven Cloud Functions for async processing.
- Data Layer: Primary DB: Cloud SQL (PostgreSQL) with pgvector for vector similarity search. Caching: Cloud Memorystore (Redis) for session/cache management. Storage: Cloud Storage for raw assets (PDFs).
- Networking: Custom VPC with private subnets and specific ingress controls.
- Security: IAM roles are granularly assigned (e.g., specific service accounts accessing specific secrets).
The Terraform configuration adheres to Google Cloud best practices for security, scalability, and maintainability with modular design, explicit dependency management, and a security-first approach.
Key Strengths:
Security-First Design:
- Network: Private VPC, private subnets, and Cloud NAT gateway ensure services are not exposed to the public internet.
- Database: Cloud SQL instance has no public IP and uses IAM authentication.
- Secrets Management: Google Secret Manager stores all sensitive information.
- Ingress: Cloud Armor with pre-configured OWASP Top 10 rules and rate limiting provides strong first-line defense.
Scalability and Resilience:
- Serverless: Cloud Run for frontend and backend enables automatic traffic-based scaling.
- Load Balancing: Global external HTTPS load balancer distributes traffic efficiently with a single entry point.
- CDN: Cloud CDN improves performance by caching static assets closer to users.
- Health Checks: Startup and liveness probes for the backend service improve reliability.
Automation:
- The CI/CD pipeline in the
cicdmodule automates build and deployment processes.
Frontend-Backend Flow Alignment:
- Service Communication: The
computemodule correctly configures Cloud Run services communication with secure environment variables and secrets. - Data Flow: The
database,redis, andstoragemodules provision necessary data stores; thefunctionmodule sets up the RAG data ingestion pipeline. - Ingress and Egress: The
ingressmodule routes traffic to the frontend; thenetworkmodule ensures outbound internet access through Cloud NAT.
- Authentication: Firebase Authentication vs. Google Identity (IAP)
I explicitly chose Firebase Authentication over Identity-Aware Proxy (IAP) for this architecture.
- Seamless Frontend Integration: Firebase provides a rich, client-side SDK that integrates natively with the Next.js application, offering a smoother, customizable user experience (login pages, social providers) compared to IAP's rigid, infrastructure-level interception.
- Public-Facing Scalability: Unlike IAP, which is optimized for internal enterprise tools (Google Workspace identities), Firebase Authentication is designed for consumer-scale applications (B2C), supporting millions of users with a generous free tier and ease of external sign-ups.
- Developer Experience: It allows for rapid prototyping and deployment without complex load balancer configurations, while still maintaining high security standards through JWT verification on the backend using the Firebase Admin SDK.
- Communication: Asyncio vs. Pub/Sub
While Pub/Sub is excellent for decoupled, asynchronous background tasks, I utilize Python's asyncio within FastAPI for the chat interface.
- Real-Time Requirement: Chat users expect immediate, streaming responses. Pub/Sub is a "fire-and-forget" mechanism designed for background processing, not for maintaining the open, bidirectional HTTP connections required for streaming LLM tokens to a user in real-time.
- Concurrency:
asyncioallows a single Cloud Run instance to handle thousands of concurrent waiting connections (e.g., waiting for Vertex AI to reply) without blocking, providing high throughput for chat without the architectural complexity of a message queue.
- Event-Driven Ingestion: Cloud Functions
I moved the document ingestion logic from a manual script to a Google Cloud Function triggered by Cloud Storage events.
- Automation: Uploading a PDF to the
data_bucketautomatically triggers the function to parse, chunk, embed, and upsert the document into the vector database. - Efficiency: This is a serverless, event-driven approach. Resources are only consumed when a file is uploaded, rather than having a long-running service waiting for input.
- Scalability: Each file upload triggers a separate function instance, allowing parallel processing of mass uploads without blocking the main chat application.
- Automation: Uploading a PDF to the
The Backend Agent is designed as a stateful, retrieval-augmented system that balances high-performance search with secure session management.
- Storage: Utilizes Google Cloud Firestore (Native Mode) for low-latency persistence of chat history.
-
Implementation: Leverages
FirestoreChatMessageHistorywithin the LangChain framework. -
Security & Isolation: Every session is cryptographically scoped to the authenticated user's email (
user_email:session_id). This ensures strict multi-tenancy where users can never access or "leak" into another's conversation history (IDOR protection). -
Context Injection: The system automatically retrieves the last
$N$ messages and injects them into the history placeholder of the RAG prompt, enabling multi-turn, context-aware dialogue.
-
Vector Database: Powered by PostgreSQL 16 (Cloud SQL) with the vector extension (
pgvector). -
Retrieval Logic: Employs semantic similarity search using
VertexAIEmbeddings(textembedding-gecko@003). For every query, the engine retrieves the top 5 most relevant chunks ($k=5$ ) to provide grounded context to the LLM. -
Semantic Caching: Integrated with Redis (Memorystore) using a
RedisSemanticCache. If a user asks a question semantically similar to a previously cached query (threshold: 0.05), the system returns the cached response instantly, bypassing the LLM to save cost and reduce latency.
- Ingestion Pipeline: A specialized
ingest.pyscript handles the transformation of raw data into "AI-ready" vectors. - Smart Chunking: Uses the
RecursiveCharacterTextSplitterto maintain semantic integrity:- Chunk Size: 1000 characters/tokens.
- Chunk Overlap: 200 characters (ensures no loss of context at the edges of chunks).
- Separators: Prioritizes splitting by double newlines (paragraphs), then single newlines, then spaces.
- Document Support: Includes a
DirectoryLoaderwithPyPDFLoaderto automatically parse and index complex PDF structures.
- Cost Savings: You pay for the system instruction tokens once per hour (cache creation) instead of every single request.
- Latency: The model doesn't need to re-process the large system prompt for every user query, leading to faster Time to First Token (TTFT).
- Implicit vs. Explicit: I relied on Implicit Caching for the short-term chat history (managed automatically by Gemini) and implemented Explicit Caching for the static, heavy system prompt.
This platform implements a robust, multi-layered security strategy. The codebase and infrastructure have been hardened against the following threats:
- SQL Injection (SQLi) Protection:
- Infrastructure Level: Google Cloud Armor is configured with pre-configured WAF rules (
sqli-v33-stable) to filter malicious SQL patterns at the edge. - Application Level: The backend uses
asyncpg(via LangChain's PGVector), which strictly employs parameterized queries, ensuring user input is never executed as raw SQL.
- Infrastructure Level: Google Cloud Armor is configured with pre-configured WAF rules (
- Cross-Site Scripting (XSS) Protection:
- Infrastructure Level: Cloud Armor WAF rules (
xss-v33-stable) detect and block malicious script injection attempts. - Framework Level: Next.js (Frontend) automatically sanitizes and escapes content by default, and the backend returns structured JSON to prevent direct script rendering.
- Infrastructure Level: Cloud Armor WAF rules (
- Broken Access Control & IDOR (Insecure Direct Object Reference):
- Verified Identity (Firebase): The frontend acts as a Secure Proxy. It captures the user's identity from Firebase Authentication tokens (
X-Firebase-Tokenheader) and propagates it to the backend for verification. - Session Isolation: Chat histories are cryptographically scoped to the authenticated user's identity (
user_email:session_id), preventing IDOR attacks where one user could access another's private history.
- Verified Identity (Firebase): The frontend acts as a Secure Proxy. It captures the user's identity from Firebase Authentication tokens (
- Edge Protection: Cloud Armor implements a global rate-limiting policy (500 requests/min per IP) and a "rate-based ban" to mitigate large-scale volumetric DDoS and brute-force attacks.
- Application Resilience: The backend core utilizes
slowapito enforce granular rate limiting (5 requests/min per user) specifically for expensive LLM operations, protecting against cost-based denial-of-service and resource exhaustion. - Input Validation: Pydantic models in the backend enforce a strict 10,000-character limit on user messages to prevent memory-exhaustion attacks.
- Prompt Injection Mitigation: The RAG prompt template uses strict structural delimiters (
----------) and prioritized system instructions to ensure the model adheres to its enterprise role and ignores adversarial overrides contained within documents or user queries. - Sensitive Data Leakage (PII): Google Cloud DLP (Data Loss Prevention) is integrated into the core pipeline with a Regex Fast-Path and Asynchronous Threading. This automatically detects and masks PII in real-time without blocking the main event loop, ensuring high performance while minimizing API costs.
- Knowledge Base Security: Data is stored in a private Cloud SQL (PostgreSQL) instance reachable only via a Serverless VPC Access connector, ensuring the "Brain" of the AI is never exposed to the public internet.
-
Preventing Wallet Exhaustion (Rate Limiting & Input Validation)
- Action: Reduced API rate limit: 10/minute for authenticated users.
- Action: Added a
validate_token_countfunction (using a lightweight 4-char/token heuristic) to strictly enforce input size limits before processing, rejecting requests that exceed the limit (2000 tokens) with a 400 error.
-
RAG Hardening (Prompt Injection Defense)
- Action: Prompt Templates with Cache Hit and Cache Miss scenarios.
- Action: "Sandwich Defense" using XML tagging (
<trusted_knowledge_base>) and explicit instructions to ignore external commands found within the retrieved context.
-
Guardrail Layer (Improved Security Judge)
- Action: Regex
SecurityBlockerfor low-latency filtering of obvious attacks. - Action: The
security_judge_llmuses a more specific system prompt acting as a specialized classifier ("SAFE" vs "BLOCKED"). - Action: Google's Native Content Safety Settings (
HarmBlockThreshold.BLOCK_LOW_AND_ABOVE) on the Judge model to leverage Vertex AI's built-in safety classifiers for Hate Speech, Dangerous Content, etc.
- Action: Regex
-
Model Safety (Generation Hardening)
- Action: Strict
SAFETY_SETTINGS(blocking low and above harm probability) to the main RAG generation models (ChatVertexAI). This acts as a final line of defense against generating harmful content, even if prompt injection succeeds.
- Action: Strict
Note on Streaming DLP:
The current protected_chain_stream sanitizes the input but streams the output directly from the LLM to the client to maintain responsiveness. By enforcing the strict SAFETY_SETTINGS on the generation model itself, we have mitigated the risk of the model generating harmful content, serving as an effective output guardrail for the streaming endpoint.
This platform has been upgraded for production-scale performance, cost efficiency, and sub-second perceived latency:
- Horizontal Autoscaling: Both Frontend and Backend services are configured for automatic horizontal scaling in Cloud Run. They can scale from zero to hundreds of concurrent instances to handle massive traffic spikes.
- Cold-Start Mitigation: The Frontend service maintains a minimum of 1 warm instance
min_instance_count = 1), ensuring immediate responsiveness and eliminating "cold start" latency for users. - Cloud SQL Read Pool: While currently using a single instance for cost efficiency, the architecture is ready for a dedicated Read Replica in Cloud SQL. This horizontally scales read capacity for the vector database, ensuring that heavy document retrieval and search operations do not bottleneck the primary write instance.
- Asynchronous I/O (Neural Core): The backend is built on FastAPI and uses
asyncpgfor non-blocking database connections. This allows a single instance to handle thousands of concurrent requests with minimal resource usage. - Server-Sent Events (SSE): Real-time token streaming from the LLM (Gemini 2.5 Flash) directly to the Next.js UI provides sub-second "Time-To-First-Token," creating a highly responsive user experience.
- Asynchronous Thread Pooling: Expensive operations like PII de-identification via Google Cloud DLP are offloaded to asynchronous background threads, preventing them from blocking the main request-response cycle.
- Gemini 2.5 Flash Integration: Utilizes the high-efficiency Flash model (
gemini-2.5-flash-preview-05-20) for a 10x reduction in token costs and significantly lower latency compared to larger models. - DLP Fast-Path Guardrails: Implemented a high-performance regex-based "pre-check" for PII. This intelligently bypasses expensive Google Cloud DLP API calls for clean content, invoking the API only when potential PII patterns are detected.
- Global CDN Caching: Google Cloud CDN is enabled at the Load Balancer level to cache static assets and common frontend resources globally, reducing origin server load and improving page load times.
- Smart Storage Versioning: Implemented Object Lifecycle Management on Cloud Storage buckets. Files are automatically transitioned to Nearline storage after 7 days, Archive storage after 30 days, and deleted after 90 days. This ensures disaster recovery capabilities (versioning is enabled) without indefinite storage costs.
The current infrastructure is designed for high efficiency and is benchmarked to handle approximately 2,500 users per hour with the standard provisioned resources.
To handle this load, you must change the architecture: Solution A: Offload Vector Search (Recommended) Use a specialized engine designed for high-throughput vector search.
- Use: Google Vertex AI Vector Search (formerly Matching Engine).
- Why: It is fully managed and designed to handle billions of vectors and thousands of QPS with <10ms latency.
- Architecture Change:
- Postgres: Only stores Chat History and User Metadata (cheap writes).
- Vertex AI: Handles the 2,800 QPS vector load.
The platform now enforces a strict "Login -> Pay -> Chat" workflow using Stripe and Cloud SQL.
- Source of Truth: The Cloud SQL (PostgreSQL) database is the single source of truth for user subscription status.
- Stripe Integration:
- Webhooks: A secure
/webhookendpoint listens forcheckout.session.completedandinvoice.payment_succeededevents from Stripe. - Automatic Activation: When a payment succeeds, the webhook updates the user's
is_activestatus in theuserstable.
- Webhooks: A secure
- Security Enforcement:
- Backend Middleware: The
get_current_userdependency checks the database for every request. Ifis_activeis false, it raises a403 Forbiddenerror. - Frontend Redirect: The frontend intercepts these 403 errors and automatically redirects the user to the
/paymentpage.
- Backend Middleware: The
The new users table tracks subscription state:
email(Primary Key): Linked to Firebase Identity.is_active(Boolean): Grants access to the chat.stripe_customer_id: Links to the Stripe Customer.subscription_status: Status string (e.g., 'active', 'past_due').
Before running the project locally or deploying to the cloud, ensure you have the following installed:
- Docker Desktop: Required for running the local database (Postgres/Vector) and Redis.
- Node.js (v18+): For the Frontend.
- Python (3.10+): For the Backend.
- Google Cloud CLI
gcloud: For authenticated access to GCP services (Vertex AI, Firestore, etc.). - Stripe CLI: For testing payments locally.
The following table details the Zero-Trust permission model enforced by the infrastructure:
| Source | Target | Role | Status |
|---|---|---|---|
| Frontend SA | Backend Service | roles/run.invoker |
✅ Present |
| Backend SA | Vertex AI | roles/aiplatform.user |
✅ Present |
| Backend SA | Cloud SQL | roles/cloudsql.client |
✅ Present |
| Backend SA | Secret Manager | roles/secretmanager.secretAccessor |
✅ Present |
| Backend SA | Firestore | roles/datastore.user |
✅ Present |
| Backend SA | Cloud DLP | roles/dlp.user |
✅ Present |
| Function SA | Storage | roles/storage.objectViewer |
✅ Present |
| Cloud Build SA | CI/CD | roles/run.admin |
✅ Present |
| Cloud Build SA | CI/CD | roles/iam.serviceAccountAdmin |
✅ Present |
| Cloud Build SA | CI/CD | roles/artifactregistry.writer |
✅ Present |
This guide details how to run the Enterprise AI Platform locally for development and testing.
- Docker Desktop (for running the database and cache locally).
- Python 3.10+ (for the Backend).
- Node.js 18+ (for the Frontend).
- Google Cloud Project with Firebase Authentication enabled.
- Stripe Account (for testing payments).
Since this is a cloud-native application, you need to connect to a few real external services even for local development.
- Go to the Firebase Console.
- Create a project (or use an existing one).
- Enable Authentication and set up the Email/Password provider.
- Go to Project Settings > General and scroll to "Your apps".
- Select "Web app", register it, and copy the
firebaseConfigobject. You will need these values for the Frontend. - Service Account Key (for Backend):
- Go to Project Settings > Service accounts.
- Click "Generate new private key".
- Save this JSON file as
service-account-key.jsonin the root of thebackend-agentdirectory. DO NOT COMMIT THIS FILE.
- Go to the Stripe Dashboard.
- Enable "Test Mode".
- Get your Publishable Key and Secret Key.
- Create a Webhook endpoint pointing to
http://localhost:8080/webhook(you may need a tool likengrokor Stripe CLI to forward local events, or just mock the payment flow in the DB manually).
We will use Docker Compose to run PostgreSQL (with pgvector) and Redis locally.
Use the docker-compose.yml file in the root of the project:
Start the infrastructure:
docker-compose up -dCreate a .env file in backend-agent/:
# Disable Secret Manager loading
PROJECT_ID=""
# Debug Mode
DEBUG=true
# Database (Matches docker-compose.yml)
DB_HOST=localhost
DB_USER=postgres
DB_PASSWORD=password
DB_NAME=postgres
# Redis
REDIS_HOST=localhost
REDIS_PASSWORD=""
# Stripe (From Step 1B)
STRIPE_API_KEY=sk_test_...
STRIPE_WEBHOOK_SECRET=whsec_...
# Google Cloud (Required for Vertex AI & Firestore)
# Ensure you are authenticated via 'gcloud auth application-default login'
# OR set GOOGLE_APPLICATION_CREDENTIALS to your key file path
GOOGLE_APPLICATION_CREDENTIALS=service-account-key.json
# Vertex AI Region
REGION=us-central1- Navigate to the backend directory:
cd backend-agent - Create a virtual environment:
python -m venv venv source venv/bin/activate # Windows: venv\Scripts\activate
- Install dependencies:
pip install -r requirements.txt
- Run the server:
uvicorn main:app --reload --host 0.0.0.0 --port 8080
Note: The backend will attempt to connect to Google Cloud services (Vertex AI for embeddings, Firestore for chat history). Ensure your service-account-key.json has permissions for:
roles/aiplatform.userroles/datastore.user
Create a .env.local file in frontend-nextjs/:
# Backend URL (Proxy or Direct)
BACKEND_URL=http://localhost:8080
# Firebase Config (From Step 1A)
NEXT_PUBLIC_FIREBASE_API_KEY=AIzaSy...
NEXT_PUBLIC_FIREBASE_AUTH_DOMAIN=your-project.firebaseapp.com
NEXT_PUBLIC_FIREBASE_PROJECT_ID=your-project-id
NEXT_PUBLIC_FIREBASE_STORAGE_BUCKET=your-project.appspot.com
NEXT_PUBLIC_FIREBASE_MESSAGING_SENDER_ID=...
NEXT_PUBLIC_FIREBASE_APP_ID=...
# Stripe Config (From Step 1B)
NEXT_PUBLIC_STRIPE_PUBLISHABLE_KEY=pk_test_...- Navigate to the frontend directory:
cd frontend-nextjs - Install dependencies:
npm install
- Run the development server:
npm run dev
- Open http://localhost:3000.
- Sign Up: Open the frontend and create an account via Firebase Auth.
- Payment (Manual Activation): Since local Stripe webhooks might be tricky without
ngrok, you can manually activate your user in the local database:- Connect to local Postgres:
psql postgres://postgres:password@localhost:5432/postgres
- Find your user (created after login attempt) and update status:
UPDATE users SET is_active = true WHERE email = 'your-email@example.com';
- Connect to local Postgres:
- Chat: You should now be able to access the chat interface. Messages will be:
- Embedded via Vertex AI (Cloud).
- Stored in Firestore (Cloud).
- Vector-searched in Postgres (Local).
This project contains automated tests for both the backend (FastAPI) and frontend (Next.js) applications. Follow the instructions below to run the tests.
The backend uses pytest for testing.
Ensure you have the Python dependencies installed:
cd backend-agent
pip install -r requirements.txtTo run all tests:
# From the backend-agent directory
pytesttests/conftest.py: Contains test fixtures (e.g.,clientfor API requests, mocks).tests/test_main.py: Contains API endpoint tests.
The frontend uses Jest and React Testing Library.
Ensure you have the Node.js dependencies installed:
cd frontend-nextjs
npm installTo run the test suite:
# From the frontend-nextjs directory
npm testTo run tests in watch mode (re-runs on file changes):
npm run test:watchnpx playwright install
npx playwright test │cd functions/pdf-ingest
pytest__tests__/: Contains the test files (e.g.,LandingPage.test.tsx).jest.config.ts: Jest configuration.jest.setup.ts: Global test setup (e.g., loadingjest-dommatchers).
This guide outlines the step-by-step process to deploy your AI Platform to Google Cloud.
Core Concept: Your infrastructure (Terraform) creates the "shell" services first (using a placeholder image). Then, your CI/CD pipelines (Cloud Build) build the actual code and "fill" those shells with your application.
Before starting, ensure you have the following CLI tools installed: gcloud, terraform, docker, npm, python.
Ensure you have a GCP project with billing enabled.
gcloud auth login
gcloud config set project [YOUR_PROJECT_ID]
gcloud config set compute/region us-central1- Go to the Firebase Console.
- Add a new project and select your existing Google Cloud Project.
- Enable Authentication (Google Provider, Email/Password, etc.).
- Enable Firestore (Create Database → Native Mode → Select same region as your GCP resources, e.g.,
us-central1). - Go to Project Settings → General → Your apps → Add app (Web).
- Copy the Firebase Config SDK values (apiKey, authDomain, projectId, storageBucket, messagingSenderId, appId). You will need these later.
gcloud services enable \
compute.googleapis.com \
iam.googleapis.com \
run.googleapis.com \
artifactregistry.googleapis.com \
cloudbuild.googleapis.com \
secretmanager.googleapis.com \
sqladmin.googleapis.com \
firestore.googleapis.com \
dlp.googleapis.com \
aiplatform.googleapis.com \
redis.googleapis.com \
cloudfunctions.googleapis.com \
storage.googleapis.comTerraform creates Cloud Build triggers that watch your repo. You must connect your repository to Google Cloud Build manually before running Terraform.
- Go to the Cloud Build Triggers page.
- Click Manage Repositories → Connect Repository.
- Select GitHub and follow the authorization flow.
- Select the repository containing this code.
Terraform needs permission to manage IAM policies, and Cloud Build needs permission to deploy.
PROJECT_NUMBER=$(gcloud projects describe $(gcloud config get-value project) --format="value(projectNumber)")
gcloud projects add-iam-policy-binding $(gcloud config get-value project) \
--member="serviceAccount:${PROJECT_NUMBER}@cloudbuild.gserviceaccount.com" \
--role="roles/run.admin"
gcloud projects add-iam-policy-binding $(gcloud config get-value project) \
--member="serviceAccount:${PROJECT_NUMBER}@cloudbuild.gserviceaccount.com" \
--role="roles/iam.serviceAccountAdmin"
gcloud projects add-iam-policy-binding $(gcloud config get-value project) \
--member="serviceAccount:${PROJECT_NUMBER}@cloudbuild.gserviceaccount.com" \
--role="roles/artifactregistry.writer"
gcloud projects add-iam-policy-binding $(gcloud config get-value project) \
--member="serviceAccount:${PROJECT_NUMBER}@cloudbuild.gserviceaccount.com" \
--role="roles/run.developer"
gcloud projects add-iam-policy-binding $(gcloud config get-value project) \
--member="serviceAccount:${PROJECT_NUMBER}@cloudbuild.gserviceaccount.com" \
--role="roles/iam.serviceAccountUser"Your Cloud SQL instance is protected by a private IP configuration, which means you cannot connect to it directly from your local machine unless you are connected to the VPC (e.g., via VPN) or use a proxy.
The init-db.sql script MUST be executed to enable the vector extension, otherwise the RAG functionality will fail.
- Upload
init-db.sqlto your Cloud Shell. - Connect to the database using the private IP (if Cloud Shell is configured for VPC access) or use the Auth Proxy.
-
Enable the API:
gcloud services enable sqladmin.googleapis.com -
Start the proxy (background):
./cloud_sql_proxy -instances=<PROJECT_ID>:<REGION>:<INSTANCE_NAME>=tcp:5432 &
-
Run the script:
psql "host=127.0.0.1 port=5432 sslmode=disable user=postgres dbname=postgres" -f init-db.sql(You will need the password generated by Terraform. Check Secret Manager:
projects/<PROJECT_ID>/secrets/<PROJECT_ID>-cloudsql-password)
- Create a small VM in the same VPC (
<project-id>-vpc) and Subnet as the database. - SSH into the VM.
- Install the postgres client:
sudo apt-get update && sudo apt-get install -y postgresql-client - Upload or copy-paste
init-db.sql. - Connect and run:
psql -h <DB_PRIVATE_IP> -U postgres -d postgres -f init-db.sql
To verify the extension is installed, log into the database and run:
SELECT * FROM pg_extension WHERE extname = 'vector';This step sets up the VPC, Database, Artifact Registry, Cloud Build Triggers, and the initial "Hello World" Cloud Run services.
Navigate to the terraform directory and create a terraform.tfvars file:
cd terraformproject_id = "your-project-id"
region = "us-central1"
domain_name = "ai.your-domain.com"
github_owner = "your-github-username"
github_repo_name = "your-repo-name"
billing_account = "000000-000000-000000"
notification_email = "admin@example.com"terraform init
terraform validate
terraform plan -out=after_review.tfplan
terraform apply after_review.tfplanType yes when prompted. This process will take 15-20 minutes.
In case of error:
terraform state list #see what Terraform thinks it successfully created
## IF state saved
# Fix errors
terraform plan
terraform apply
# ELSE
terraform import <resource_type>.<name> <existing_id>
terraform applyWhat gets created:
- VPC Network & Serverless Access
- Cloud SQL (PostgreSQL) & Redis
- Cloud Run Services (Frontend & Backend) with placeholder images
- Cloud Function (PDF Ingest)
- Cloud Build Triggers
- Secret Manager placeholders
The Frontend build requires your Firebase keys to be "baked" into the Docker image at build time.
- Go to Google Cloud Console → Cloud Build → Triggers.
- Locate the
frontend-nextjs-triggerand click Edit. - Scroll down to Substitution variables.
- Add the following variables using the Firebase values from Prerequisites:
| Variable Name | Description |
|---|---|
_NEXT_PUBLIC_FIREBASE_API_KEY |
Your Firebase API Key |
_NEXT_PUBLIC_FIREBASE_AUTH_DOMAIN |
Your Firebase Auth Domain (e.g., project.firebaseapp.com) |
_NEXT_PUBLIC_FIREBASE_PROJECT_ID |
Your Firebase Project ID |
_NEXT_PUBLIC_FIREBASE_STORAGE_BUCKET |
Your Firebase Storage Bucket (e.g., project.appspot.com) |
_NEXT_PUBLIC_FIREBASE_MESSAGING_SENDER_ID |
Your Firebase Messaging Sender ID |
_NEXT_PUBLIC_FIREBASE_APP_ID |
Your Firebase App ID |
- Click Save.
Security Note: Since these variables are prefixed with NEXT_PUBLIC_, Next.js will bundle them into the client-side JavaScript code. They are safe to be exposed to the browser (as they are required for the Firebase client SDK to work), but ensure your Firebase Security Rules are configured correctly.
The backend deployment is simpler because environment variables are injected at runtime via Cloud Run, not build time.
- Go to Cloud Build → Triggers.
- Click Create Trigger (or edit the existing
backend-agent-trigger). - Configure with:
- Name:
backend-deploy - Event: Push to a branch
- Source: Your repository and branch (e.g.,
^main$) - Configuration: Cloud Build configuration file
- Location:
cloudbuild-backend.yaml - Ignored Files (Optional):
frontend-nextjs/**
- Name:
- Click Save/Create.
It will: 1. Test (Run pytest/npm test) -> 2. Build -> 3. Deploy
Terraform created the containers for your secrets, but you need to add the values.
- Find the secret
STRIPE_SECRET_KEYin Secret Manager and add a new version with your Stripe Secret Key. - Find the secret
STRIPE_PUBLISHABLE_KEYand add your Stripe Publishable Key.
Create these secrets manually:
# Stripe Webhook Secret (from Stripe Dashboard)
gcloud secrets create STRIPE_WEBHOOK_SECRET --replication-policy="automatic"
echo -n "whsec_..." | gcloud secrets versions add STRIPE_WEBHOOK_SECRET --data-file=-
# Google API Key (Optional, if not using ADC)
gcloud secrets create GOOGLE_API_KEY --replication-policy="automatic"
echo -n "AIzaSy..." | gcloud secrets versions add GOOGLE_API_KEY --data-file=-
# DB_HOST (use the IP address output by Terraform)
gcloud secrets create DB_HOST --replication-policy="automatic"
echo -n "10.x.x.x" | gcloud secrets versions add DB_HOST --data-file=-Now that the infrastructure and triggers are configured, you can build the actual applications.
In the Cloud Build Triggers page:
- Click Run on
backend-agent-trigger. - Click Run on
frontend-nextjs-trigger.
Push a commit to your main branch to automatically fire both triggers:
git add .
git commit -m "Deploy applications"
git push origin mainWhat happens: Cloud Build builds the Docker images, pushes them to Artifact Registry, and updates the Cloud Run services with your actual application code.
If needed, you can also run builds directly:
gcloud builds submit --config cloudbuild-backend.yaml .
gcloud builds submit --config cloudbuild-frontend.yaml .Your Cloud SQL database is running but empty.
The easiest way is to use the Cloud SQL Auth Proxy from Cloud Shell:
# Download proxy
curl -o cloud-sql-proxy https://storage.googleapis.com/cloud-sql-connectors/cloud-sql-proxy/v2.8.0/cloud-sql-proxy.linux.amd64
chmod +x cloud-sql-proxy
# Start proxy (replace INSTANCE_CONNECTION_NAME from SQL Console)
./cloud-sql-proxy --address 0.0.0.0 --port 5432 INSTANCE_CONNECTION_NAME &Go to Secret Manager → [project-id]-cloudsql-password → View Secret Value.
psql "host=127.0.0.1 port=5432 sslmode=disable user=postgres dbname=postgres" -f init-db.sqlWithin the database, enable the vector extension:
CREATE EXTENSION IF NOT EXISTS vector;Unlike the frontend, backend secrets are set on the Cloud Run service itself:
- Go to Cloud Run.
- Select the
backend-agentservice. - Click Edit & Deploy New Revision.
- Go to the Variables & Secrets tab.
- Add the required environment variables (see
backend-agent/config.pyfor the list, e.g.,DB_HOST,DB_PASSWORD,STRIPE_API_KEY). - Click Deploy.
- Get the Load Balancer IP:
cd terraform && terraform output public_ip-
Update your DNS provider (e.g., GoDaddy, Cloudflare) to point your domain (e.g.,
ai.your-domain.com) to this IP. -
Wait 15-30 minutes for the managed SSL certificate to provision.
- Go to Cloud Run in the Console.
- Click on the
frontend-agentservice. - Click the URL provided at the top.
- You should see your Next.js application (not the Hello World page).
-
What is deployed:
- Budget: A budget alert for $100 (alerts you at 50%, 90%, 100%).
- Monitoring: Custom metric for error counts and an alert policy for high error rates.
- Notification: Email channel.
-
Cost: ~$0.00 / month.
- Google Cloud Budgets and standard Alerting are generally free.
- Custom metrics can incur costs if you send millions of data points, but for this scale, it's negligible.
-
What is deployed:
- Artifact Registry: A Docker repository (
cloud-run-source-deploy). - Cloud Build: 2 Triggers (Frontend & Backend) connected to GitHub.
- Artifact Registry: A Docker repository (
-
Cost: Pay-as-you-go (Low).
- Builds: You get 120 free build-minutes/day. Unless you commit code constantly, this is likely free.
- Storage: Artifact Registry charges ~$0.020 per GB/month for storing your Docker images.
-
What is deployed:
-
Cloud Run Job:
ingest-job(Runs only when triggered). -
Cloud Run Service:
backend-agent(2 vCPU, 4GB RAM). -
Cloud Run Service:
frontend-agent(1 vCPU, 2GB RAM).
-
Cloud Run Job:
-
Cost: $0.00 / month (If Idle).
-
Configuration: Both services currently have
min_instance_count = 0. This means they scale to zero when not in use. -
Active Cost: You only pay when requests are processed.
- Backend: ~$0.000048 per vCPU-second.
- Frontend: ~$0.000024 per vCPU-second.
-
Configuration: Both services currently have
-
What is deployed:
-
Cloud SQL (PostgreSQL):
- Tier:
db-g1-small(Shared core). - Edition: Enterprise.
- Disk: SSD (Autoscaling).
- Backups: Enabled (7 days retention).
- Tier:
- Firestore: Native mode database.
-
Cloud SQL (PostgreSQL):
-
Cost: ~$34 - $45 / month.
- The SQL instance charges an hourly rate 24/7 (
$0.041/hour) plus storage costs ($0.17/GB/month). - Firestore is pay-as-you-go (reads/writes) and has a generous free tier.
- The SQL instance charges an hourly rate 24/7 (
-
What is deployed:
- Cloud Function:
pdf-ingest-function(Python 3.11). - Trigger: Eventarc trigger watching a Storage Bucket for new files.
- Cloud Function:
-
Cost: Pay-as-you-go (Low).
- You only pay when a file is uploaded and the function runs. The first 2 million invocations per month are usually free.
-
What is deployed:
- Load Balancer: Global External Application Load Balancer.
-
SSL: Managed Google Certificate for
app.yourdomain.com. - Cloud Armor: Security Policy with WAF rules (SQLi, XSS, etc.) and Rate Limiting.
-
Cost: ~$33 - $40 / month.
-
Forwarding Rule: The Load Balancer charges
$0.025/hour ($18/month). - Cloud Armor: ~$5/month per policy + $1/month per rule.
-
Warning: The plan enables
layer_7_ddos_defense_config. We have confirmedgoogle_compute_project_cloud_armor_tieris NOT used, avoiding the $3,000/mo Enterprise subscription.
-
Forwarding Rule: The Load Balancer charges
-
What is deployed:
- VPC Network: Custom subnets.
-
Cloud NAT: A NAT Gateway (
your-actual-project-id-12345-nat).
-
Cost: ~$33 / month + Data Fees.
- The NAT Gateway charges
$0.045/hour ($33/month) just to exist, regardless of traffic. - You also pay $0.045 per GB for data processing through the NAT.
- The NAT Gateway charges
-
What is deployed:
- Memorystore for Redis: Basic Tier, 1 GB capacity.
-
Cost: ~$36 / month.
- This is a fixed instance charged hourly (~$0.049/hour).
-
What is deployed:
- Buckets:
data_bucket(with lifecycle rules) andsource_bucket.
- Buckets:
-
Cost: Pay-as-you-go (Low).
- Standard storage is ~$0.02 per GB. Unless you store Terabytes, this is negligible.
| Module | Resource | Est. Monthly Cost (Idle) |
|---|---|---|
| Compute | Cloud Run (Min 0 Instances) | ~$0.00 |
| Database | Cloud SQL (db-g1-small) | ~$34.00 |
| Redis | Memorystore (1GB Basic) | ~$36.00 |
| Network | Cloud NAT Gateway | ~$33.00 |
| Ingress | Load Balancer Rule | ~$18.00 |
| Ingress | Cloud Armor Policy | ~$15.00 |
| TOTAL | Baseline "Rent" | ~$136.00 / month |
Recommendation: To reduce costs to <$50/mo:
- Delete the Redis module and use a local container or smaller service if possible (~$36 savings).
- Delete the NAT Gateway if your Cloud Run services don't strictly need a static outgoing IP (~$33 savings). Note: This would require changing the Cloud Run VPC egress settings.
-
Downgrade Cloud SQL to
db-f1-micro(if available in your region/project type) or use Firestore only (~$34 savings).
This section outlines the procedures for recovering critical data stores.
Our Cloud SQL instance is configured with:
- Automated Backups: Retained for 7 days.
- Point-in-Time Recovery (PITR): Allows restoration to any second within the retention window.
- Deletion Protection: Prevents accidental deletion of the instance.
Restore the database to a state before the corruption occurred:
-
Identify the Timestamp: Determine the exact UTC time just before the error occurred.
-
Perform Restore (Clone):
gcloud sql instances clone <SOURCE_INSTANCE_ID> restored-db-instance \
--point-in-time "2023-10-27T13:00:00Z"-
Verify Data: Connect to
restored-db-instanceand verify the data integrity. -
Switch Traffic: Update the application secrets to point to the new instance IP/Host.
Restore from the last successful nightly backup:
- List Backups:
gcloud sql backups list --instance=<INSTANCE_ID>- Restore:
gcloud sql backups restore <BACKUP_ID> --restore-instance=<TARGET_INSTANCE_ID>Our Firestore database is configured with a Daily Backup Schedule retained for 7 days.
- List Available Backups:
gcloud firestore backups list --location=<REGION>Note the resource name of the backup you wish to restore.
- Restore to a New Database:
Firestore does not support in-place restores. You must restore to a new database ID.
gcloud firestore databases restore \
--source-backup=projects/<PROJECT_ID>/locations/<REGION>/backups/<BACKUP_ID> \
--destination-database=restored-firestore-db- Update Application: Update the backend configuration (
FIRESTORE_DATABASE_ID) to point torestored-firestore-db.
- Verify Connectivity: Ensure backend services can connect to the restored databases.
- Data Integrity Check: Run application-level smoke tests.
- Re-enable Backups: Ensure the new/restored instances have backup schedules re-applied (Terraform apply might be needed).
pre-commit install
pre-commit run --all-files
git status && git diff
git add ... ... && git commit -m "...."Acknowledgements ✨ Google ML Developer Programs and Google Developers Program supported this work by providing Google Cloud Credits (and awesome tutorials for the Google Developer Experts)✨