Welcome to the AI Studio documentation. This guide provides comprehensive information for developers, operators, and administrators.
Complete system architecture including technology stack, design patterns, database schema, and security model.
Step-by-step deployment guide for AWS infrastructure using CDK, including Google OAuth setup and first administrator configuration.
Complete API documentation for REST endpoints and server actions with request/response examples.
Error codes, handling patterns, and debugging workflow.
Common issues and solutions for development, deployment, and production.
Complete reference of all environment variables required for development and production environments.
Index of all architectural diagrams (9 comprehensive Mermaid.js diagrams with 10,000+ lines of documentation):
- 01. CDK Stack Dependencies - Deployment order and SSM parameter flows
- 02. VPC Network Topology - Multi-AZ subnets, security groups, and VPC endpoints
- 03. AWS Service Architecture - Complete service breakdown with cost analysis
- 04. Database ERD - All 54 PostgreSQL tables with relationships
- 05. Authentication Flow - OAuth 2.0 via Cognito + NextAuth v5
- 06. Request Flow Diagrams - Nexus Chat, Model Compare, Document Processing
- 07. Assistant Architect Execution - Multi-prompt chains with variable substitution
- 08. Document Processing Pipeline - Upload → S3 → Lambda → Textract → Embedding → pgvector
- 09. Streaming Architecture - SSE via ECS Fargate with circuit breaker pattern
Comprehensive logging patterns with examples for server actions, API routes, and error handling.
Testing strategies including unit tests, integration tests, and E2E testing with Playwright.
TypeScript best practices, conventions, and guidelines for maintaining type safety.
OAuth and auth security review checklist for PRs modifying auth flows or token handling.
Catalog of silent failure patterns (Drizzle ORM, AI SDK v6, prototype pollution, SessionProvider).
React pitfalls and patterns specific to this codebase (initialization guards, derived state, deferred loading).
Step-by-step provider integration guide for adding new AI providers.
AWS Secrets Manager integration and best practices.
Getting started with the AI Studio REST API v1.
OAuth2/OIDC integration for authenticating external applications.
Connecting AI tools (Claude Code, Cursor) via Model Context Protocol.
Complete guide for creating and managing database migrations with Drizzle ORM.
Common query patterns and best practices for Drizzle ORM.
Common issues and solutions when working with Drizzle ORM.
Quick reference for the hybrid Drizzle-kit to Lambda migration workflow.
Tracking the RDS Data API to Drizzle ORM migration progress.
Overview of library directory structure and common utilities.
Database access layer using Drizzle ORM with postgres.js driver.
Unified streaming service, provider adapters, and circuit breaker implementation.
Complete CDK infrastructure guide with deployment commands, environment configuration, and best practices.
VPC consolidation and network architecture optimization.
Aurora Serverless v2 cost optimization and monitoring strategies.
Lambda optimization framework with PowerTuning results (66% memory reduction).
Detailed PowerTuning results for all Lambda functions.
Multi-architecture Docker builds for ARM64/AMD64 support.
security/USING_IAM_SECURITY.md ⭐ START HERE
How to use the IAM security framework with examples and patterns.
Comprehensive IAM security architecture with least privilege and tag-based access control.
Step-by-step guide for migrating existing infrastructure to secure IAM constructs.
Operational procedures, monitoring, and maintenance guidelines.
Load testing and performance benchmarking procedures.
ECS streaming infrastructure operations and monitoring.
Digest-bound harness, prompt, and model comparison decisions from the first full PSD Agent eval run.
Reproduction evidence, prompt hardening, repeated trials, and the non-promotion decision for the GLM-5 stable-skill contract follow-up.
What changed to move the agent harness to Claude Sonnet 5.5, why no request-shaping change was needed, deploy order, the 10-prompt dev verification run, and rollback.
Annual NSPRA publication replacement, protected-source handling, repository configuration, retrieval validation, and prior-edition retirement.
Comprehensive checklist for deploying to production.
Deployment, stale-DLQ cleanup, explicit invocation, and verification steps for the owner-bound agent schedule record and Scheduler-target migration.
Inventory, backfill, reconciliation, cutover, rollback, and legacy-content retirement runbook.
Dev-first repository recovery, production drift reconciliation, guarded cutover order, and cross-surface acceptance gates.
Managing Assistant Architect tools and permissions.
OIDC signing-key architecture, provisioning, rotation, health checks, and incident response.
AI integration patterns using Vercel AI SDK v6, provider factory implementation, and streaming techniques.
features/nexus-conversation-architecture.md ⭐ CRITICAL
Complete Nexus conversation system architecture - Component hierarchy, state management (conversationId vs stableConversationId vs ref), message flows, runtime memoization, format conversions, common pitfalls, and troubleshooting. Read this before modifying conversation code.
Nexus model routing — Standard/Advanced UX, Nova Micro classification, family/tier resolution, automatic web search, image and PSD-data MCP dispatch, configuration, fallbacks, and rollout modes.
Assistant Architect model routing — Standard/Advanced authoring, shared capability-aware tier routing, legacy compatibility, execution surfaces, and independent rollout controls.
Unified repository product integration — Authoritative Repository Manager source/ACL/version UI, Assistant Architect repository-only knowledge, Nexus private ephemeral attachments, promotion, retention, and rollback boundaries.
Unified content agents and Projects — Repository catalog MCP/REST scopes, per-user OpenClaw PKCE OAuth, live skill repository bindings, and durable Nexus Project membership/repository/chat boundaries.
PSD Observances agent skill — Cited, bounded NSPRA calendar lookups, coverage limits, repository resolution, authorization behavior, and operational ownership.
Google Workspace content synchronization — PKCE OAuth and Picker selection across My Drive and user-accessible Shared Drives, durable cursor/version reconciliation, deployment-owned configuration, deletion grace, and operations.
Agent-owned Google Workspace integration — User and agent credential slots, DWD minting, automatic agnt_ provisioning, Phase 1 safety boundaries, deployment, automated checks, live smoke tests, and troubleshooting.
ClassLink OneRoster synchronization — Nightly full-collection roster ingestion, OAuth1 and static-bearer Proxy configuration, fail-safe reconciliation invariants, metrics, alarms, rollout, and rollback.
Dynamic navigation system with role-based menu items.
Document upload and processing system with S3 integration.
Vector embedding and semantic search with pgvector.
Server-Sent Events for Assistant Architect execution.
Tool integration for Assistant Architect prompts.
Tool & skill versioning contract — v1/v2 versions, identifier@version addressing, deprecation lifecycle (90-day grace period, replaced_by, removal), MCP/REST visibility, and SKILL.md version pinning.
Skill publishing & export — SKILL.md format, publish → scan → review pipeline, how skills are consumed (Nexus session binding vs skill.{slug} catalog tools), and the zip-export format + portability caveats for Claude Code / Desktop.
Email triage user guide — enabling triage, the three-label folder model, escalation modes (what pings vs what gets labeled), rule tuning, the correction → nightly-learning → suggestion loop, digest, and @psd/Task email-to-task. End-user companion to operations/email-triage.md.
Nexus workspace chat editing (§1087) — when a document/artifact is open beside the chat, the chat can read + edit it (live Yjs document edits via the agent bridge, artifact edits via createVersion); server-bound, canView/canEdit-gated, §28.3-screened.
Connecting agents to Atrium content — how a local MCP client (Claude Code etc.) or the PSD AI Agents (OpenClaw) read/write Atrium documents: API-key setup, the MCP content tools + scope table, the version-based vs live-document distinction, delegated tokens, and the loopback binding hazard for the live bridge.
Atrium collection management — district/shared and owner-bound private hierarchies, inherited view/create grants, lifecycle and conflict rules, explicit count/filter semantics, audit coverage, and the UI/REST/skill surfaces.
Atrium artifact data — session-authenticated AtriumData.submit() / list() persistence for sandboxed artifacts, the source-authenticated postMessage protocol, append-only records, agent broker reads, CSP invariants, and locked design decisions.
Complete JSON import specification for generating valid assistant import files. Includes schema reference, field types, variable substitution, execution patterns, and comprehensive examples.
Document processing infrastructure setup.
Testing strategy for document processing pipeline.
Client integration patterns for polling APIs.
S3 lifecycle policies and cost optimization.
OAuth2/OIDC provider architecture and implementation.
Model Context Protocol server for AI tool integrations.
Structured decision capture and context graph system.
K-12 content filtering with Amazon Bedrock Guardrails.
Real-time voice conversations via WebSocket with Gemini Live API, transcript persistence, permissions, and content safety guardrails.
Atrium content workspace (Epic #1059) — agent-native documents and sandboxed interactive artifacts: collaborative editing, visibility model and permission-aware retrieval, publishing connectors with the §26.4 public-publish approval gate, the anonymous /p/[slug] public reader, OKF export/import, and the MCP/REST content tool surfaces.
- Start with ARCHITECTURE.md to understand the system
- Review diagrams/README.md for visual architecture
- Check /infra/README.md for infrastructure details
- Review ENVIRONMENT_VARIABLES.md for setup
- Follow guides/TYPESCRIPT.md for code standards
- Reference guides/LOGGING.md for logging patterns
- Working on Nexus? Read features/nexus-conversation-architecture.md first
- Follow DEPLOYMENT.md for initial deployment
- Study /infra/README.md for CDK stack details
- Review diagrams/01-cdk-stack-dependencies.md for deployment order
- Check operations/OPERATIONS.md for maintenance
- Use TROUBLESHOOTING.md for common issues
- Read guides/TESTING.md for testing strategies
- Use Playwright MCP for E2E testing during development
- Add tests to
working-tests.spec.tsfor CI/CD
All server actions return a consistent response structure. See ARCHITECTURE.md#actionstate-pattern and API_REFERENCE.md#server-actions.
Unified interface for multiple AI providers. See API/AI_SDK_PATTERNS.md and /lib/streaming/README.md.
Every operation gets a unique request ID for end-to-end tracing. See guides/LOGGING.md and ERROR_REFERENCE.md#debugging-workflow.
Database-first configuration with environment fallback. See ARCHITECTURE.md#settings-management.
Direct ECS execution for real-time AI streaming with HTTP/2 support. See diagrams/09-streaming-architecture.md and /lib/streaming/README.md.
Drizzle ORM with postgres.js driver and connection pooling. See database/drizzle-patterns.md, /lib/db/README.md, and diagrams/04-database-erd.md.
- Design the database schema (see diagrams/04-database-erd.md)
- Create server actions with proper logging (see guides/LOGGING.md)
- Build UI components
- Add E2E tests (see guides/TESTING.md)
- Update documentation
- Follow guides/adding-ai-providers.md
- Create provider adapter in /lib/streaming/provider-adapters/
- Add to database models and configuration
- Test with real API and update monitoring
- Deploy and verify in staging environment
- Use request ID to trace through CloudWatch logs
- Check ERROR_REFERENCE.md for error codes
- Review TROUBLESHOOTING.md for common issues
- Follow operations/OPERATIONS.md for procedures
- Test locally with
bun run dev - Run
bun run lintandbun run typecheck(entire codebase) - Deploy with CDK:
bunx cdk deploy - Monitor CloudWatch for errors
- See /infra/README.md for detailed deployment commands
- Next.js Documentation
- AWS CDK Guide
- Vercel AI SDK
- Playwright Documentation
- PostgreSQL Documentation
- AWS Well-Architected Framework
- Always when adding new features
- Always when changing architecture
- Always when modifying deployment process
- When fixing complex bugs (document the solution)
- When discovering non-obvious patterns
- Keep documentation close to code
- Use clear, concise language
- Include code examples (tested and working)
- Include visual diagrams where helpful
- Update this README index
- Cross-reference related documents
- Current, active documentation in main folders
- Diagrams in
/docs/diagrams/(Mermaid.js format) - Feature docs in
/docs/features/ - Operations docs in
/docs/operations/ - Infrastructure docs in
/docs/infrastructure/ - Guides in
/docs/guides/
When contributing to documentation:
- Follow the existing structure
- Use proper markdown formatting
- Include practical, tested examples
- Cross-reference related documents
- Update this README index
- Add diagrams where helpful (Mermaid.js preferred)
Key architectural decisions documented:
- ADR-001: Authentication Optimization - NextAuth v5 with Cognito integration
- ADR-002: Streaming Architecture Migration - Amplify to ECS Fargate migration
- ADR-003: ECS Streaming Migration - Lambda workers to direct ECS execution
- ADR-004: Docker Container Optimization - Multi-stage builds and layer caching
- ADR-006: Centralized Secrets Management - AWS Secrets Manager integration
Last updated: April 2026 Status: Active - comprehensive documentation with 9 architectural diagrams Total Documentation: 10,000+ lines across 50+ files For AI assistant guidelines, see CLAUDE.md