Skip to content

About

Production-ready AI Business OS — 46 autonomous AI Employees, Graph RAG (Neo4j), multi-channel, BYOK across 17 LLM providers plus custom/local endpoints (Ollama), automatic fallback. 8 libs (Python+Java). Self-hosted, MIT licensed, guardrailed autonomy for regulated industries.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

17 stars

Watchers

2 watching

Forks

Latest commit

 

History

965 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

RagLeap Core logo

RagLeap Core

Autonomous AI Agents — Not Just RAG.

license MIT Core Release 46 AI Employees Autonomous Self--Hosted PyPI ragleap-rag Downloads PyPI ragleap-graph Downloads PyPI ragleap-vectorstores Downloads PyPI ragleap-tools Downloads PyPI ragleap-integrations Downloads PyPI ragleap-ops Downloads PyPI ragleap-app-chart Downloads PyPI ragleap-observability Downloads PyPI ragleap-terraform Downloads Maven Central ragleap-rag (Java) Python Discord Repo clones

English · Afrikaans · العربية · Български · বাংলা · Català · Čeština · Cymraeg · Dansk · Deutsch · Ελληνικά · Español · Eesti · فارسی · Suomi · Français · ગુજરાતી · עברית · हिन्दी · Hrvatski · Magyar · Bahasa Indonesia · Italiano · 日本語 · ಕನ್ನಡ · 한국어 · Lietuvių · Latviešu · Македонски · മലയാളം · मराठी · नेपाली · Nederlands · Norsk · ਪੰਜਾਬੀ · Polski · Português · Română · Русский · Slovenčina · Slovenščina · Soomaali · Shqip · Svenska · Kiswahili · தமிழ் · తెలుగు · ไทย · Tagalog · Türkçe · Українська · اردو · Tiếng Việt · 简体中文 · 繁體中文

RagLeap Core is the open-source engine behind RagLeap — a self-hosted, agentic system that runs your business from your own documents on your own server, with no vendor lock-in.

46 role-based AI Employees: 9 core generalist roles (AI Manager, Secretary, CEO, Sales, Support, HR, Finance, Marketing, Operations) plus 37 vertical-specific global roles (Recruiter, Real Estate Agent, Legal Intake, Healthcare Intake, Insurance Agent, and more) — full list in core/employees/defaults.py.

What it does:

  • Self-learning, outcome-weighted memory
  • Auto-trigger workflows and escalation
  • Autonomy modes: off (default, the AI acts only when you start it), semi (you approve each action) and full (opt-in; sensitive roles are never fully autonomous)
  • Think-act-decide loop, not just retrieval

Quickstart · Docs · Website · Packages


Open source, self-hosted. ragleap-core (this repo) is MIT-licensed, completely free, and never requires a license key. A managed cloud version of RagLeap is planned; for now, RagLeap is open source only.

Install

pip install ragleap-rag

⭐ If this helps you, please consider starring the repo — it genuinely helps more people find it.

Prefer Java? ragleap-rag is also on Maven Central:

<dependency>
    <groupId>io.github.antonyrag</groupId>
    <artifactId>ragleap-rag</artifactId>
    <version>0.5.0</version>
</dependency>

Independently versioned from the Python package above — see java/ragleap-rag for current scope.

Add ragleap-graph too if you want Neo4j-backed knowledge graph retrieval:

pip install ragleap-rag ragleap-graph

Add ragleap-vectorstores too if you want pluggable vector backends beyond ragleap-rag's built-in six (Chroma today):

pip install ragleap-rag ragleap-vectorstores[chroma]
# or, with uv
uv add ragleap-rag ragleap-vectorstores[chroma]

Add ragleap-tools too if you want ready-made tools for LLM tool-calling (calculator, sandboxed file ops, date/time, unit conversion, JSON/CSV parsing, text utilities):

pip install ragleap-rag ragleap-tools
# or, with uv
uv add ragleap-rag ragleap-tools

Deploying to Kubernetes? ragleap-ops ships live-tested manifests and a Helm chart for the full stack:

pip install ragleap-ops
# or, with uv
uv add ragleap-ops

Need a generic, reusable chart for your own app (not RagLeap-specific)? ragleap-app-chart takes an arbitrary services: list, not hardcoded names:

pip install ragleap-app-chart
# or, with uv
uv add ragleap-app-chart

Want metrics and logs for your ragleap-ops deployment? ragleap-observability ships Prometheus, Grafana, Loki + Promtail, and AlertManager, live-verified end-to-end against a real cluster:

pip install ragleap-observability
# or, with uv
uv add ragleap-observability

Or run the full self-hosted app (channels, web chat UI, Docker Compose) — see Quickstart below. Browse every package at packages.ragleap.com. Try it hands-on with the runnable scripts in examples/ — 01_ingest_and_query.py (upload a document, ask a question via the API) and 02_test_channel_directly.py (test channel answering logic without real bot credentials).

If a RAG chatbot answers questions, RagLeap runs your business

Most open-source RAG projects give you a toolkit — you still have to build the app, wire up a UI, add memory, and connect every channel yourself. RagLeap Core gives you one AI across WhatsApp, Telegram, Discord, and voice calls, instead of a different disconnected bot per channel. Memory persists across sessions and channels, so it isn't relearning who a customer is every time. Voice is a real, working inbound-call handler (Twilio Media Streams, Whisper STT, OpenAI TTS) — not just document Q&A with a phone number bolted on — and n8n workflow automation can be triggered directly from any conversation.

What this repo doesn't include: a multi-tenant admin dashboard (settings, analytics, team, billing) and an executive-assistant layer (Manager AI). RagLeap Core itself is single-tenant, self-hosted, and configured via .env.

What makes RagLeap Core specifically different

This repo isn't a general-purpose RAG framework you assemble into something — it's a real, working engine you can run today. The code here is honest about being early.

Why RagLeap exists

Open-source AI agent projects like OpenClaw took off for a specific reason: people wanted an assistant that runs on their own infrastructure, with their own keys, answering from the chat apps they already use — not a black box hosted by someone else. That same principle is what RagLeap Core is built on for business AI specifically.

Your keys, your infrastructure, your data. RagLeap Core never asks for a system API key. You bring your own Gemini key, you run your own PostgreSQL database, your documents are stored on your own server, and text is sent only to the AI providers you configure (your chat model, and Gemini for embeddings).

Chat is the interface, not a separate dashboard you have to learn. The same way OpenClaw meets people on WhatsApp, Telegram, and Slack, RagLeap's full platform meets business owners on the channels they already use — WhatsApp, Telegram, Discord, and real phone calls — not a new app they have to check.

A real, working system — not an abstract framework. This isn't a toolkit like LangChain where you assemble your own app from primitives. RagLeap Core is the actual chunking → embedding → retrieval → generation pipeline extracted from a production system that already answers real customer questions, at a company that already runs on it.

Built in public, honestly. This repo says clearly what's done and what isn't. No inflated claims, no vaporware Quickstart commands that don't work yet — the Roadmap reflects the real state of the code, updated as it progresses.

What RagLeap Core is

RagLeap Core is a document-grounded chat engine. Upload your documents, ask questions, get cited answers — self-hosted, on your own infrastructure, with your own API key.

WhatsApp, Telegram, and Discord bots are included in this repo too — single-tenant, .env-configured channel adapters that answer from the same document knowledge base.

RagLeap Core is right for you if

  • ✅ You want a self-hosted RAG chatbot with full control over your data
  • ✅ You want to understand exactly how document retrieval and citation works, not use a black box
  • ✅ You're comfortable running your own server and your own AI provider key
  • ✅ You want to contribute to or extend an open document-QA engine
  • ✅ You'd rather see the code than trust a vendor's word on data privacy

What RagLeap Core is not

It's not... It is...
A hosted product Self-hosted software you run yourself
Multi-tenant, with persistent cross-session memory Single-tenant — one bot, one document set, per deployment
A multi-tenant platform WhatsApp/Telegram/Discord/Voice channel adapters included, single-tenant
A no-code SaaS dashboard A codebase you deploy and configure
Feature-complete The foundational subset — see Roadmap

Features

📄 Document ingestion Upload PDFs, text, and common document formats
🔍 RAG retrieval Vector search over your documents via pgvector
💬 Chat with citations Answers reference the source document, not a black box
🔌 Bring your own AI key OpenAI, Gemini, Anthropic, or any OpenAI-compatible endpoint
🌐 Web chat widget Embed a chat widget on any website
🐳 Docker-based setup One-command local deployment
🕸️ Knowledge Graph (Neo4j) Entity extraction and graph-boosted retrieval alongside vector search
🌍 Language detection Auto-detects document and query language, applied across every channel
🔗 Integrations Connect MySQL, PostgreSQL, MongoDB, REST APIs, Salesforce, HubSpot, Shopify, Google Sheets, Stripe
🔀 Hybrid search Combines dense (vector) and sparse (full-text) retrieval via Reciprocal Rank Fusion
⚡ Streaming responses Answers stream token-by-token instead of waiting for the full response
🔁 Provider fallback Automatically retries with a backup LLM provider if the primary fails
💰 Token usage reporting Real per-call token counts from the provider, plus context-size budget trimming
🧑‍💼 AI Employees Role-based agents (46 default roles) with persistent business-context memory, wired into /chat via role=<role>
🛠️ Build your own AI Employee Define a fully custom role (any name, personality, channels, skill tags) via PATCH /employees/{role} - no fork, no code change. Verified live: create, retrieve, and list a new role end-to-end.
🔗 n8n workflow automation Fire a webhook after the AI replies on WhatsApp/Telegram/Discord — no-code automations triggered directly from a conversation

Architecture

How RagLeap Core is put together:

flowchart TD
    subgraph Core["RagLeap Core — this repo (open)"]
        WebUI["Web Chat UI"] --> ChatAPI["Chat API"]
        ChatAPI --> WA["WhatsApp"]
        ChatAPI --> TG["Telegram"]
        ChatAPI --> DC["Discord"]
        ChatAPI --> VC["Voice"]

        WA --> N8N["n8n Workflow Trigger (fires after AI reply)"]
        TG --> N8N
        DC --> N8N

        Employees["AI Employees (role context, learned memory)"] --> Provider

        WA --> Ingest["Document Ingest"]
        TG --> Ingest
        DC --> Ingest
        VC --> Ingest

        WA --> RAG["RAG Retrieve"]
        TG --> RAG
        DC --> RAG
        VC --> RAG

        WA --> Provider["AI Provider Adapter"]
        TG --> Provider
        DC --> Provider
        VC --> Provider

        Ingest --> PG[("PostgreSQL + pgvector")]
        RAG --> PG
        Provider --> PG
        PG --> Neo[("Neo4j (Knowledge Graph)")]
    end

Loading

Optional background job queue. The periodic integration sync job runs inline in the API process by default. Setting REDIS_URL switches it to a real Redis queue (RQ) instead, processed by separate worker process(es) that scale independently — horizontally via docker compose up -d --scale worker=N, or in Kubernetes via a queue-depth autoscaler like KEDA (see examples/keda-scaledobject-worker.yaml for a reference ScaledObject). Entirely optional — nothing changes if REDIS_URL is unset.

Repo structure

ragleap-core/
├── core/                  # RAG engine — chunking, embedding, retrieval, generation
│   ├── chunker.py
│   ├── embedding.py       # Embeddings: Gemini (default, 3072-dim) or Ollama/OpenAI/Mistral/...
│   ├── retrieval.py       # pgvector cosine search
│   ├── generation.py      # 19-provider BYOK generation (Gemini, OpenAI, Anthropic, etc.)
│   ├── ingest.py          # chunk -> embed -> store pipeline
│   ├── parsers.py         # PDF/DOCX/TXT text extraction
│   ├── employees/         # AI Employees — roles, business profile, learned memory
│   ├── workflows.py       # n8n workflow automation — webhook triggers
│   └── api.py             # FastAPI app — /health, /upload, /chat, /profile, /employees, /n8n-workflows, /webhook/*
├── channels/              # Messaging + voice channel adapters
│   ├── whatsapp/          # Twilio + Gupshup
│   ├── telegram/
│   ├── discord/
│   └── voice/             # Twilio Media Streams, WebSocket server
├── db/
│   └── schema.sql         # documents + chunks tables, pgvector index
├── examples/              # Runnable example scripts
├── .github/workflows/     # CI: compile check, Docker build, smoke tests
├── docker-compose.yml     # app + db + voice services
└── Dockerfile

Quickstart

✅ Status: core pipeline verified working. Ingest -> embed -> retrieve -> generate runs end-to-end via Docker Compose, including a clean fresh-clone test. See the Roadmap for what's next (PDF/DOCX support, alternative BYOK providers).

Fastest way to try it — one command checks Docker, clones the repo, and sets up .env for you:

curl -fsSL https://raw.githubusercontent.com/antonyrag/ragleap-core/main/install.sh | bash

(Windows users: run this in Git Bash, not Command Prompt or PowerShell.)

The script will pause after cloning and ask you to add your Gemini API key to .env — get a free one at aistudio.google.com/apikey, then re-run the same command.

Or, the manual way — better if you want to read the code before running anything:

git clone https://github.com/antonyrag/ragleap-core.git
cd ragleap-core
cp .env.example .env
# add your Gemini API key to .env
docker compose up --build -d

Requirements: Docker, Docker Compose, and a Gemini API key (embeddings use Gemini by default; set EMBEDDING_PROVIDER=ollama or another provider to run without a Gemini key). For chat you can use Gemini or any of the 19 providers listed under Supported LLM Providers, including a local Ollama model and any OpenAI-compatible endpoint.

Try it in 30 seconds — with the stack running, see examples/ for two verified, runnable scripts:

  • examples/01_ingest_and_query.py — upload a document and ask a question via the API
  • examples/02_test_channel_directly.py — test the WhatsApp/Telegram/Discord answering logic without real bot credentials

Just want the RAG engine as a Python library?

If you don't need the full Docker app — WhatsApp/Telegram/Discord/Voice adapters, the web chat UI, all of it — the core retrieval engine is also published as standalone, pip-installable Python packages:

pip install ragleap-rag
  • ragleap-rag — the chunking → embedding → retrieval → generation pipeline as a library. Pluggable embeddings (8 providers: Gemini, OpenAI, Mistral, Together, Ollama, Cohere, Voyage and a custom endpoint), 6 vector backends (FAISS, PgVector, Pinecone, Weaviate, Qdrant, Milvus), cross-encoder reranking, and more.
  • ragleap-graph — Neo4j-backed knowledge graph retrieval, usable standalone or alongside ragleap-rag.
  • ragleap-vectorstores — pluggable vector backends beyond ragleap-rag core's 6. Backends: Chroma and LanceDB (embedded/local), Redis (RediSearch/Redis Stack), Upstash Vector (managed serverless) and OpenSearch (k-NN). Install with pip install ragleap-vectorstores[chroma] or uv add ragleap-vectorstores[chroma].
  • ragleap-tools — standalone, dependency-light tools for LLM tool-calling. 12 stateless tools (calculator, date/time, unit conversion, JSON/CSV parsing, text utilities), sandboxed file ops (symlink-escape protected), search_documents (hybrid vector+keyword search), search_web (pluggable BYOK providers such as Tavily and Serper), and optional ragleap-rag-backed document ingestion. Does not own a tool-calling execution loop — provides Tool objects for your own loop or ragleap-agents once it ships. Install with pip install ragleap-tools or uv add ragleap-tools.
  • ragleap-ops — Kubernetes deployment manifests for RagLeap Core, live-tested end-to-end on a real cluster. Install with pip install ragleap-ops or uv add ragleap-ops.
  • ragleap-app-chart — generic, reusable Helm chart for deploying arbitrary services to Kubernetes, not RagLeap-specific. Point it at your own app via a services: list. Install with pip install ragleap-app-chart or uv add ragleap-app-chart.
  • ragleap-observability — Prometheus, Grafana, Loki + Promtail for ragleap-ops, live-verified end-to-end against a real cluster (real connection confirmed, real non-zero metrics returned; 21+ real log streams with correct namespace/pod/container labels). AlertManager is wired to Prometheus with a first real alert rule, plus optional Slack/email receivers (off by default; delivery verified against a stand-in, real send not yet confirmed). Install with pip install ragleap-observability or uv add ragleap-observability.
  • ragleap-terraform — Terraform module that creates a local kind cluster and installs the RagLeap Helm charts, applied and destroyed on a real cluster (no cloud modules yet). Install with pip install ragleap-terraform or uv add ragleap-terraform.

All eight are MIT licensed. Browse the full package index at packages.ragleap.com.

Supported LLM Providers (BYOK)

RagLeap Core is bring-your-own-key only there is no system-provided key for any provider. Set LLM_PROVIDER in .env to choose which one to use for the generation (chat) step. Embeddings currently always use Gemini (gemini-embedding-001), regardless of LLM_PROVIDER.

LLM_PROVIDER value Required env vars Notes
gemini (default) GEMINI_API_KEY Get a key at aistudio.google.com/apikey
anthropic ANTHROPIC_API_KEY, ANTHROPIC_MODEL (optional) Get a key at console.anthropic.com
openai OPENAI_API_KEY, OPENAI_MODEL
mistral MISTRAL_API_KEY, MISTRAL_MODEL
groq GROQ_API_KEY, GROQ_MODEL Free tier available
together TOGETHER_API_KEY, TOGETHER_MODEL
openrouter OPENROUTER_API_KEY, OPENROUTER_MODEL
ollama OLLAMA_MODEL (no API key needed) Self-hosted; requires Ollama running locally
deepseek DEEPSEEK_API_KEY, DEEPSEEK_MODEL
xai XAI_API_KEY, XAI_MODEL
cohere COHERE_API_KEY, COHERE_MODEL
perplexity PERPLEXITY_API_KEY, PERPLEXITY_MODEL
qwen QWEN_API_KEY, QWEN_MODEL
moonshot MOONSHOT_API_KEY, MOONSHOT_MODEL
zhipu ZHIPU_API_KEY, ZHIPU_MODEL
yi YI_API_KEY, YI_MODEL
baidu BAIDU_API_KEY, BAIDU_MODEL
minimax MINIMAX_API_KEY, MINIMAX_MODEL
custom CUSTOM_API_KEY, CUSTOM_MODEL, CUSTOM_BASE_URL Any OpenAI-compatible endpoint

Example, switching to Groq in .env:

LLM_PROVIDER=groq
GROQ_API_KEY=your-groq-key
GROQ_MODEL=llama-3.3-70b-versatile

Ollama in Docker: known gotchas

Running Ollama as a fallback/local provider from inside this app's Docker container (rather than bare metal) has real, verified setup steps beyond just setting OLLAMA_MODEL:

  1. host.docker.internal isn't reliable on plain Linux Docker Engine (unlike Docker Desktop) — it can resolve to the wrong bridge network's gateway. If Ollama connections fail with this alias, override OLLAMA_BASE_URL with the container's actual compose-network gateway IP directly, e.g. OLLAMA_BASE_URL=http://172.18.0.1:11434/v1 (find your real gateway with docker network inspect <network> | grep Gateway).
  2. Ollama binds to 127.0.0.1 by default, which a container can't reach even with the correct gateway IP. Override it to listen on all interfaces:
   mkdir -p /etc/systemd/system/ollama.service.d
   cat > /etc/systemd/system/ollama.service.d/override.conf << 'EOF'
   [Service]
   Environment="OLLAMA_HOST=0.0.0.0:11434"
   EOF
   systemctl daemon-reload && systemctl restart ollama
  1. The host firewall may silently drop container→host traffic even after the above — container-to-bridge-gateway traffic hits ufw's INPUT chain, not FORWARD/DOCKER-USER. Allow it explicitly: ufw allow from 172.18.0.0/16 to any port 11434 proto tcp (adjust the subnet to match your actual Docker network).

Voice Channel (Twilio)

RagLeap Core includes a real-time voice channel: Twilio Media Streams connects via WebSocket, your speech is transcribed with OpenAI Whisper, answered by the core RAG pipeline, and spoken back with OpenAI TTS. The voice-activity detection and echo-suppression logic is carried over from a production system tuned against real call traffic.

Runs as a separate service on port 8765 (see docker-compose.yml), since Twilio's real-time audio protocol needs a raw WebSocket server, not an HTTP route.

Setup:

  1. Set OPENAI_API_KEY in .env (used for both Whisper STT and TTS in v1)
  2. Optionally set VOICE_BOT_NAME, VOICE_GREETING, VOICE_TTS_VOICE
  3. Point a Twilio phone number's <Connect><Stream> TwiML at wss://your-domain.com:8765

Honest status: the WebSocket server, Twilio event protocol handling, and error handling are verified working. The full Whisper/TTS round-trip has not yet been live-tested end-to-end (requires OpenAI API credits). If you try it and hit issues, please open one — this is exactly the kind of real-world testing this project needs.

Known limitations, carried over from production and not yet fixed here:

  • Only OpenAI Whisper (STT) and OpenAI TTS are supported in v1 — Deepgram and ElevenLabs (multi-language support) are good-first-issue candidates
  • Non-English TTS quality varies since OpenAI's TTS voices are English-tuned
  • Typical round-trip latency in production was 6-8 seconds

Knowledge Graph (Neo4j)

RagLeap Core builds a lightweight entity co-occurrence graph alongside its vector index. When you ingest a document, entities (product names, acronyms, proper nouns) are extracted and linked in Neo4j. When you ask a question, the same extraction runs on your query, and any documents linked to matching entities get a similarity boost in retrieval — on top of, not instead of, normal pgvector search.

Runs as a fourth Docker Compose service on ports 7475/7688 (remapped from Neo4j's defaults to avoid colliding with another Neo4j instance on the same host). If Neo4j is unreachable or NEO4J_PASSWORD is unset, the graph degrades gracefully — retrieval falls back to pure vector search, ingestion is unaffected.

Setup:

  1. Set NEO4J_URI, NEO4J_USER, and NEO4J_PASSWORD in .env (matching the NEO4J_AUTH value in docker-compose.yml)
  2. Optionally set DOMAIN_TERMS — a comma-separated list of domain-specific terms to boost during extraction (e.g. DOMAIN_TERMS=API,SDK,RAG)

Honest status: entity extraction, document graph writes, entity-based document lookup, and graph-boosted chat retrieval are all verified working end-to-end, including in CI (fresh build, real ingest, real query, real graph lookup). The graph boost is currently a simple additive score bump, not a full weighted re-ranker — a richer hybrid ranking system is a good next step for anyone who wants to dig in.

Known limitations:

  • Entity extraction is regex-based (CamelCase, acronyms, capitalized phrases, plus optional domain terms) — not a trained NER model, so it will miss some entities and occasionally include noise
  • search_related_entities() (multi-hop graph traversal) is implemented but not yet wired into the retrieval pipeline — good-first-issue candidate for anyone wanting a project

Language Detection

RagLeap Core auto-detects language during document ingestion (per chunk) and during chat (per query), using the langdetect library plus script-based heuristics for CJK, Hangul, and Kana text. Since every channel (WhatsApp, Telegram, Discord, Voice, and the API directly) routes through the same core chat pipeline, detection applies consistently everywhere without per-channel wiring.

Setup: works out of the box with no configuration. Optionally set DEFAULT_LANGUAGE (fallback when detection fails or text is too short), LANGUAGE_DETECTION_CONFIDENCE_THRESHOLD (default 0.7), and LANGUAGE_DETECTION_SUPPORTED_LANGUAGES (comma-separated allowlist, blank = unrestricted).

Honest status: verified working end-to-end — document-level detection tested at high confidence (0.9999) on a real mixed-language document, and query-level detection confirmed working via both the API and CLI.

Known limitations:

  • langdetect covers roughly 55 languages — noticeably fewer than the hosted platform's 222+, which layers additional detection and per-user language preferences on top
  • Short queries in closely-related languages can be misdetected (in testing, a short French query was detected as Italian) — this is an inherent limitation of statistical detection on short text, not specific to this port. A good-first-issue candidate for anyone wanting to improve short-query accuracy
  • Detection is one-way only: RagLeap Core detects the query's language and surfaces it, but does not yet steer the AI's response language to match — that's a reasonable next step for a contributor

Integrations

flowchart TD
    subgraph Connectors["10 connectors, one shared interface"]
        C1[MySQL] --- C2[PostgreSQL] --- C3[MongoDB]
        C4[REST API] --- C5[Salesforce] --- C6[HubSpot]
        C7[Shopify] --- C8[Google Sheets] --- C9[Stripe]
        C10[CSV Upload]
    end
    Connectors --> Svc["RealTimeExternalDataService<br/>real-time query, 5-min cache, no separate sync step"]

    Owner["Owner configures an action:<br/>trigger phrase or auto-match on SQL/field names<br/>+ a query template (SELECT / UPDATE / INSERT / DELETE)"] --> Svc

    Svc --> Match["match_and_execute_action(workspace_id, user_message, user_identifier)<br/>extracts order_id / email / phone from the message,<br/>substitutes into the template, runs the query"]

    Match --> Chat["Chat channels (WhatsApp/Telegram/Discord/Web)<br/>api/personal_bot_views.py"]
    Match --> Voice["Voice channel<br/>memory/voice_views.py::twilio_voice_speech<br/>wired in 2026-08"]

    Chat --> RAGCtx["Result injected as context<br/>into the RAG prompt"]
    Voice --> RAGCtx
    RAGCtx --> Answer["AI answers with real account/order/appointment<br/>data, not just document knowledge"]
Loading

Verified live, read from api/addon_realtime.py's RealTimeExternalDataService. Two things worth being direct about: match_and_execute_action genuinely supports write queries (UPDATE/INSERT/DELETE), not just read-only lookups — owner-configured, so the safety boundary is whatever SQL the owner writes into the template, not something the framework restricts on its own. And until 2026-08, this action-matching step only ran on chat channels; voice calls had no equivalent, which is the gap closed in the Voice Channel Routing diagram above.

RagLeap Core connects to external databases and business tools, syncing per-user context to personalize RAG responses. Nine connectors are included: MySQL, PostgreSQL, MongoDB, generic REST APIs, Salesforce, HubSpot, Shopify, Google Sheets, and Stripe.

Every CRM/SaaS connector uses credentials you provide directly — a username/password, a private-app token, an admin API token, a service-account JSON file, or a secret key, depending on the service. None require registering an OAuth app; nothing here depends on RagLeap owning any third-party developer account.

Credentials are encrypted at rest (Fernet/AES-128) before being stored.

Setup:

  1. Generate an encryption key: python3 -c "from cryptography.fernet import Fernet; print(Fernet.generate_key().decode())"
  2. Set ADDON_ENCRYPTION_KEY in .env to that value
  3. Install the SDK for the connector(s) you want (each is optional — see requirements.txt)
  4. Create a data source: POST /integrations with name, source_type, and the relevant credential fields
  5. Test it: POST /integrations/{id}/test
  6. Sync it: POST /integrations/{id}/sync

Honest status: verified end-to-end against a real public API — connection testing, syncing, correct identifier-field matching, and credential encryption (checked as actual ciphertext in the database, not just assumed) all confirmed working.

Known limitations:

  • 9 of the 18 source types listed in the hosted platform's UI have real connectors here. CSV Upload, Snowflake, BigQuery, WooCommerce, Airtable, Notion, Razorpay, Slack, and Gmail are good-first-issue candidates for anyone wanting to add one
  • Sync is on-demand only (POST /integrations/{id}/sync) — no scheduled background sync yet, though the schema tracks sync_interval_minutes for a future Celery-beat-equivalent
  • Synced context isn't automatically injected into chat responses yet — each channel adapter would need to know its own user's identifier first, which is a reasonable next contribution

Retrieval, Generation & Reliability

Beyond the core RAG pipeline, /chat (and the underlying core.chat.ask()) support several controls aimed at production use: retrieval quality, response latency, provider reliability, and cost.

Hybrid search (dense + sparse). By default, retrieval combines pgvector cosine similarity with Postgres full-text search (tsvector/ GIN index), fused via Reciprocal Rank Fusion — catching both semantic matches and exact keyword/identifier matches a pure embedding search can miss. Pass hybrid=false to use dense-only retrieval instead (cheaper — one query instead of two).

Streaming. POST /chat/stream streams the answer as it's generated (text/plain, chunked transfer) instead of waiting for the full response. Implemented natively per provider (Gemini, Anthropic, and OpenAI-compatible each have different streaming APIs — all three are real, not one stubbed).

Provider fallback. Set LLM_FALLBACK_PROVIDERS (comma-separated) to automatically retry with backup providers if the primary fails — a rate limit, outage, or bad key on your primary provider doesn't have to mean a failed request. Each fallback needs its own API key configured normally. Streaming can only fall back before any text has been sent to the caller — a mid-stream failure surfaces as an error rather than silently switching providers and confusing the output.

Generation controls. temperature, system_prompt, and max_tokens are all real per-call parameters (not just env-var defaults) — build your own agent behavior on top of RagLeap's retrieval without forking the library.

Token usage & context budget. Every blocking /chat call returns real token usage (prompt_tokens, completion_tokens, total_tokens) pulled directly from the provider's response — not an estimate. Retrieved context is also trimmed to MAX_CONTEXT_CHARS (default 12000, roughly 4 characters per token for English text) before being sent, dropping the lowest-ranked chunks first, so you're not paying for more context than necessary. Set MAX_CONTEXT_CHARS=0 to disable trimming.

Honest status: hybrid search's RRF fusion math verified correct against hand calculation. Streaming verified working end-to-end for the default provider. Provider fallback verified with a real broken-primary test — deliberately invalid API key, confirmed fallback to a working secondary provider with a correct answer. Token usage and context trimming verified with real numbers: a 3-chunk retrieval trimmed to 1 chunk under a tight budget reduced actual prompt_tokens by 38% on the same live API.

Known limitations:

  • Token usage reporting is not available for streaming responses — each provider's streaming API surfaces usage differently, and doing all three correctly is separate, not-yet-done work
  • MAX_CONTEXT_CHARS is a character-count approximation (~4 chars/token for English), not an exact per-provider tokenizer count
  • Hybrid search hasn't been benchmarked for actual ranking-quality improvement on a multi-document corpus with genuinely conflicting dense vs. sparse rankings — only correctness (fusion math, tokenization of unusual identifiers) has been verified so far

Roadmap

  • Public repository created
  • Core RAG engine extracted and cleaned from production codebase
  • Standalone Docker Compose setup (no external Django project dependency)
  • Document ingestion module (28+ formats, not just PDF/TXT/DOCX)
  • Web chat widget
  • Bring-your-own-API-key support (19 providers)
  • WhatsApp, Telegram, Discord, and Voice channel adapters (single-tenant)
  • Knowledge Graph (Neo4j), language detection, database/CRM integrations
  • Contribution guide and good-first-issue labels — 8+ issues labeled, with a real external contributor active on #134
  • Community Discord

See ROADMAP.md for the full phase-by-phase history.

Contributing

RagLeap Core is working, tested, and open for contributions now. See CONTRIBUTING.md for how to get started, and check the good first issue label for scoped tasks.

Want to see exactly what's being worked on and what's open to claim? Check the Project board — issues are staged as Good First Issue, Ready (Scoped), or Needs Scoping, so you can pick something that matches how much design work you want to do versus just build.

Student, professor, or looking for a capstone/thesis project? See STUDENT_PROJECTS.md for scoped project ideas at starter, semester, and research-grade levels.

Community

License

MIT © 2026 RagLeap

Troubleshooting

"could not translate host name 'db'" error after a failed docker compose up: If your first docker compose up attempt fails (e.g. a port conflict on 5433 or 8000), a retry can sometimes leave the database container attached to a stale, orphaned Docker network. Fix:

docker compose down
docker network prune -f
docker compose up --build -d

Port 5433 or 8000 already in use: Another instance of this project (or something else) is using the port. Either stop it, or change the host-side port mapping in docker-compose.yml (the "5433:5432" and "8000:8000" lines) to something free.

Add AI Employees example scripts to examples/

Documentation reference and guidelines for #135.

About

Production-ready AI Business OS — 46 autonomous AI Employees, Graph RAG (Neo4j), multi-channel, BYOK across 17 LLM providers plus custom/local endpoints (Ollama), automatic fallback. 8 libs (Python+Java). Self-hosted, MIT licensed, guardrailed autonomy for regulated industries.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

17 stars

Watchers

2 watching

Forks

Releases

Sponsor this project

Used by

Contributors

Languages