Autonomous AI Agents — Not Just RAG.
English · Afrikaans · العربية · Български · বাংলা · Català · Čeština · Cymraeg · Dansk · Deutsch · Ελληνικά · Español · Eesti · فارسی · Suomi · Français · ગુજરાતી · עברית · हिन्दी · Hrvatski · Magyar · Bahasa Indonesia · Italiano · 日本語 · ಕನ್ನಡ · 한국어 · Lietuvių · Latviešu · Македонски · മലയാളം · मराठी · नेपाली · Nederlands · Norsk · ਪੰਜਾਬੀ · Polski · Português · Română · Русский · Slovenčina · Slovenščina · Soomaali · Shqip · Svenska · Kiswahili · தமிழ் · తెలుగు · ไทย · Tagalog · Türkçe · Українська · اردو · Tiếng Việt · 简体中文 · 繁體中文
RagLeap Core is the open-source engine behind RagLeap — a self-hosted, agentic system that runs your business from your own documents on your own server, with no vendor lock-in.
46 role-based AI Employees: 9 core generalist roles (AI Manager, Secretary, CEO, Sales, Support, HR, Finance, Marketing, Operations) plus 37 vertical-specific global roles (Recruiter, Real Estate Agent, Legal Intake, Healthcare Intake, Insurance Agent, and more) — full list in core/employees/defaults.py.
What it does:
- Self-learning, outcome-weighted memory
- Auto-trigger workflows and escalation
- Autonomy modes:
off(default, the AI acts only when you start it),semi(you approve each action) andfull(opt-in; sensitive roles are never fully autonomous) - Think-act-decide loop, not just retrieval
Quickstart · Docs · Website · Packages
Open source, self-hosted.
ragleap-core(this repo) is MIT-licensed, completely free, and never requires a license key. A managed cloud version of RagLeap is planned; for now, RagLeap is open source only.
pip install ragleap-rag⭐ If this helps you, please consider starring the repo — it genuinely helps more people find it.
Prefer Java? ragleap-rag is also on Maven Central:
<dependency>
<groupId>io.github.antonyrag</groupId>
<artifactId>ragleap-rag</artifactId>
<version>0.5.0</version>
</dependency>Independently versioned from the Python package above — see java/ragleap-rag for current scope.
Add ragleap-graph too if you want Neo4j-backed knowledge graph retrieval:
pip install ragleap-rag ragleap-graphAdd ragleap-vectorstores too if you want pluggable vector backends beyond ragleap-rag's built-in six (Chroma today):
pip install ragleap-rag ragleap-vectorstores[chroma]
# or, with uv
uv add ragleap-rag ragleap-vectorstores[chroma]Add ragleap-tools too if you want ready-made tools for LLM tool-calling (calculator, sandboxed file ops, date/time, unit conversion, JSON/CSV parsing, text utilities):
pip install ragleap-rag ragleap-tools
# or, with uv
uv add ragleap-rag ragleap-toolsDeploying to Kubernetes? ragleap-ops ships live-tested manifests and a Helm chart for the full stack:
pip install ragleap-ops
# or, with uv
uv add ragleap-opsNeed a generic, reusable chart for your own app (not RagLeap-specific)? ragleap-app-chart takes an arbitrary services: list, not hardcoded names:
pip install ragleap-app-chart
# or, with uv
uv add ragleap-app-chartWant metrics and logs for your ragleap-ops deployment? ragleap-observability ships Prometheus, Grafana, Loki + Promtail, and AlertManager, live-verified end-to-end against a real cluster:
pip install ragleap-observability
# or, with uv
uv add ragleap-observabilityOr run the full self-hosted app (channels, web chat UI, Docker Compose) — see Quickstart below. Browse every package at packages.ragleap.com. Try it hands-on with the runnable scripts in examples/ — 01_ingest_and_query.py (upload a document, ask a question via the API) and 02_test_channel_directly.py (test channel answering logic without real bot credentials).
Most open-source RAG projects give you a toolkit — you still have to build the app, wire up a UI, add memory, and connect every channel yourself. RagLeap Core gives you one AI across WhatsApp, Telegram, Discord, and voice calls, instead of a different disconnected bot per channel. Memory persists across sessions and channels, so it isn't relearning who a customer is every time. Voice is a real, working inbound-call handler (Twilio Media Streams, Whisper STT, OpenAI TTS) — not just document Q&A with a phone number bolted on — and n8n workflow automation can be triggered directly from any conversation.
What this repo doesn't include: a multi-tenant admin dashboard (settings, analytics, team, billing) and an executive-assistant layer (Manager AI). RagLeap Core itself is single-tenant, self-hosted, and configured via .env.
This repo isn't a general-purpose RAG framework you assemble into something — it's a real, working engine you can run today. The code here is honest about being early.
Open-source AI agent projects like OpenClaw took off for a specific reason: people wanted an assistant that runs on their own infrastructure, with their own keys, answering from the chat apps they already use — not a black box hosted by someone else. That same principle is what RagLeap Core is built on for business AI specifically.
Your keys, your infrastructure, your data. RagLeap Core never asks for a system API key. You bring your own Gemini key, you run your own PostgreSQL database, your documents are stored on your own server, and text is sent only to the AI providers you configure (your chat model, and Gemini for embeddings).
Chat is the interface, not a separate dashboard you have to learn. The same way OpenClaw meets people on WhatsApp, Telegram, and Slack, RagLeap's full platform meets business owners on the channels they already use — WhatsApp, Telegram, Discord, and real phone calls — not a new app they have to check.
A real, working system — not an abstract framework. This isn't a toolkit like LangChain where you assemble your own app from primitives. RagLeap Core is the actual chunking → embedding → retrieval → generation pipeline extracted from a production system that already answers real customer questions, at a company that already runs on it.
Built in public, honestly. This repo says clearly what's done and what isn't. No inflated claims, no vaporware Quickstart commands that don't work yet — the Roadmap reflects the real state of the code, updated as it progresses.
RagLeap Core is a document-grounded chat engine. Upload your documents, ask questions, get cited answers — self-hosted, on your own infrastructure, with your own API key.
WhatsApp, Telegram, and Discord bots are included in this repo too — single-tenant, .env-configured channel adapters that answer from the same document knowledge base.
- ✅ You want a self-hosted RAG chatbot with full control over your data
- ✅ You want to understand exactly how document retrieval and citation works, not use a black box
- ✅ You're comfortable running your own server and your own AI provider key
- ✅ You want to contribute to or extend an open document-QA engine
- ✅ You'd rather see the code than trust a vendor's word on data privacy
| It's not... | It is... |
|---|---|
| A hosted product | Self-hosted software you run yourself |
| Multi-tenant, with persistent cross-session memory | Single-tenant — one bot, one document set, per deployment |
| A multi-tenant platform | WhatsApp/Telegram/Discord/Voice channel adapters included, single-tenant |
| A no-code SaaS dashboard | A codebase you deploy and configure |
| Feature-complete | The foundational subset — see Roadmap |
| 📄 Document ingestion | Upload PDFs, text, and common document formats |
| 🔍 RAG retrieval | Vector search over your documents via pgvector |
| 💬 Chat with citations | Answers reference the source document, not a black box |
| 🔌 Bring your own AI key | OpenAI, Gemini, Anthropic, or any OpenAI-compatible endpoint |
| 🌐 Web chat widget | Embed a chat widget on any website |
| 🐳 Docker-based setup | One-command local deployment |
| 🕸️ Knowledge Graph (Neo4j) | Entity extraction and graph-boosted retrieval alongside vector search |
| 🌍 Language detection | Auto-detects document and query language, applied across every channel |
| 🔗 Integrations | Connect MySQL, PostgreSQL, MongoDB, REST APIs, Salesforce, HubSpot, Shopify, Google Sheets, Stripe |
| 🔀 Hybrid search | Combines dense (vector) and sparse (full-text) retrieval via Reciprocal Rank Fusion |
| ⚡ Streaming responses | Answers stream token-by-token instead of waiting for the full response |
| 🔁 Provider fallback | Automatically retries with a backup LLM provider if the primary fails |
| 💰 Token usage reporting | Real per-call token counts from the provider, plus context-size budget trimming |
| 🧑💼 AI Employees | Role-based agents (46 default roles) with persistent business-context memory, wired into /chat via role=<role> |
| 🛠️ Build your own AI Employee | Define a fully custom role (any name, personality, channels, skill tags) via PATCH /employees/{role} - no fork, no code change. Verified live: create, retrieve, and list a new role end-to-end. |
| 🔗 n8n workflow automation | Fire a webhook after the AI replies on WhatsApp/Telegram/Discord — no-code automations triggered directly from a conversation |
How RagLeap Core is put together:
flowchart TD
subgraph Core["RagLeap Core — this repo (open)"]
WebUI["Web Chat UI"] --> ChatAPI["Chat API"]
ChatAPI --> WA["WhatsApp"]
ChatAPI --> TG["Telegram"]
ChatAPI --> DC["Discord"]
ChatAPI --> VC["Voice"]
WA --> N8N["n8n Workflow Trigger (fires after AI reply)"]
TG --> N8N
DC --> N8N
Employees["AI Employees (role context, learned memory)"] --> Provider
WA --> Ingest["Document Ingest"]
TG --> Ingest
DC --> Ingest
VC --> Ingest
WA --> RAG["RAG Retrieve"]
TG --> RAG
DC --> RAG
VC --> RAG
WA --> Provider["AI Provider Adapter"]
TG --> Provider
DC --> Provider
VC --> Provider
Ingest --> PG[("PostgreSQL + pgvector")]
RAG --> PG
Provider --> PG
PG --> Neo[("Neo4j (Knowledge Graph)")]
end
Optional background job queue. The periodic integration sync job runs inline in the API process by default. Setting REDIS_URL switches it to a real Redis queue (RQ) instead, processed by separate worker process(es) that scale independently — horizontally via docker compose up -d --scale worker=N, or in Kubernetes via a queue-depth autoscaler like KEDA (see examples/keda-scaledobject-worker.yaml for a reference ScaledObject). Entirely optional — nothing changes if REDIS_URL is unset.
ragleap-core/
├── core/ # RAG engine — chunking, embedding, retrieval, generation
│ ├── chunker.py
│ ├── embedding.py # Embeddings: Gemini (default, 3072-dim) or Ollama/OpenAI/Mistral/...
│ ├── retrieval.py # pgvector cosine search
│ ├── generation.py # 19-provider BYOK generation (Gemini, OpenAI, Anthropic, etc.)
│ ├── ingest.py # chunk -> embed -> store pipeline
│ ├── parsers.py # PDF/DOCX/TXT text extraction
│ ├── employees/ # AI Employees — roles, business profile, learned memory
│ ├── workflows.py # n8n workflow automation — webhook triggers
│ └── api.py # FastAPI app — /health, /upload, /chat, /profile, /employees, /n8n-workflows, /webhook/*
├── channels/ # Messaging + voice channel adapters
│ ├── whatsapp/ # Twilio + Gupshup
│ ├── telegram/
│ ├── discord/
│ └── voice/ # Twilio Media Streams, WebSocket server
├── db/
│ └── schema.sql # documents + chunks tables, pgvector index
├── examples/ # Runnable example scripts
├── .github/workflows/ # CI: compile check, Docker build, smoke tests
├── docker-compose.yml # app + db + voice services
└── Dockerfile
✅ Status: core pipeline verified working. Ingest -> embed -> retrieve -> generate runs end-to-end via Docker Compose, including a clean fresh-clone test. See the Roadmap for what's next (PDF/DOCX support, alternative BYOK providers).
Fastest way to try it — one command checks Docker, clones the repo, and sets up .env for you:
curl -fsSL https://raw.githubusercontent.com/antonyrag/ragleap-core/main/install.sh | bash(Windows users: run this in Git Bash, not Command Prompt or PowerShell.)
The script will pause after cloning and ask you to add your Gemini API key to .env — get a free one at aistudio.google.com/apikey, then re-run the same command.
Or, the manual way — better if you want to read the code before running anything:
git clone https://github.com/antonyrag/ragleap-core.git
cd ragleap-core
cp .env.example .env
# add your Gemini API key to .env
docker compose up --build -dRequirements: Docker, Docker Compose, and a Gemini API key (embeddings use Gemini by default; set EMBEDDING_PROVIDER=ollama or another provider to run without a Gemini key). For chat you can use Gemini or any of the 19 providers listed under Supported LLM Providers, including a local Ollama model and any OpenAI-compatible endpoint.
Try it in 30 seconds — with the stack running, see examples/ for two verified, runnable scripts:
examples/01_ingest_and_query.py— upload a document and ask a question via the APIexamples/02_test_channel_directly.py— test the WhatsApp/Telegram/Discord answering logic without real bot credentials
If you don't need the full Docker app — WhatsApp/Telegram/Discord/Voice adapters, the web chat UI, all of it — the core retrieval engine is also published as standalone, pip-installable Python packages:
pip install ragleap-ragragleap-rag— the chunking → embedding → retrieval → generation pipeline as a library. Pluggable embeddings (8 providers: Gemini, OpenAI, Mistral, Together, Ollama, Cohere, Voyage and a custom endpoint), 6 vector backends (FAISS, PgVector, Pinecone, Weaviate, Qdrant, Milvus), cross-encoder reranking, and more.ragleap-graph— Neo4j-backed knowledge graph retrieval, usable standalone or alongsideragleap-rag.ragleap-vectorstores— pluggable vector backends beyondragleap-ragcore's 6. Backends: Chroma and LanceDB (embedded/local), Redis (RediSearch/Redis Stack), Upstash Vector (managed serverless) and OpenSearch (k-NN). Install withpip install ragleap-vectorstores[chroma]oruv add ragleap-vectorstores[chroma].ragleap-tools— standalone, dependency-light tools for LLM tool-calling. 12 stateless tools (calculator, date/time, unit conversion, JSON/CSV parsing, text utilities), sandboxed file ops (symlink-escape protected),search_documents(hybrid vector+keyword search),search_web(pluggable BYOK providers such as Tavily and Serper), and optionalragleap-rag-backed document ingestion. Does not own a tool-calling execution loop — providesToolobjects for your own loop orragleap-agentsonce it ships. Install withpip install ragleap-toolsoruv add ragleap-tools.ragleap-ops— Kubernetes deployment manifests for RagLeap Core, live-tested end-to-end on a real cluster. Install withpip install ragleap-opsoruv add ragleap-ops.ragleap-app-chart— generic, reusable Helm chart for deploying arbitrary services to Kubernetes, not RagLeap-specific. Point it at your own app via aservices:list. Install withpip install ragleap-app-chartoruv add ragleap-app-chart.ragleap-observability— Prometheus, Grafana, Loki + Promtail forragleap-ops, live-verified end-to-end against a real cluster (real connection confirmed, real non-zero metrics returned; 21+ real log streams with correct namespace/pod/container labels). AlertManager is wired to Prometheus with a first real alert rule, plus optional Slack/email receivers (off by default; delivery verified against a stand-in, real send not yet confirmed). Install withpip install ragleap-observabilityoruv add ragleap-observability.ragleap-terraform— Terraform module that creates a local kind cluster and installs the RagLeap Helm charts, applied and destroyed on a real cluster (no cloud modules yet). Install withpip install ragleap-terraformoruv add ragleap-terraform.
All eight are MIT licensed. Browse the full package index at packages.ragleap.com.
RagLeap Core is bring-your-own-key only there is no system-provided key for any provider. Set LLM_PROVIDER in .env to choose which one to use for the generation (chat) step. Embeddings currently always use Gemini (gemini-embedding-001), regardless of LLM_PROVIDER.
LLM_PROVIDER value |
Required env vars | Notes |
|---|---|---|
gemini (default) |
GEMINI_API_KEY |
Get a key at aistudio.google.com/apikey |
anthropic |
ANTHROPIC_API_KEY, ANTHROPIC_MODEL (optional) |
Get a key at console.anthropic.com |
openai |
OPENAI_API_KEY, OPENAI_MODEL |
|
mistral |
MISTRAL_API_KEY, MISTRAL_MODEL |
|
groq |
GROQ_API_KEY, GROQ_MODEL |
Free tier available |
together |
TOGETHER_API_KEY, TOGETHER_MODEL |
|
openrouter |
OPENROUTER_API_KEY, OPENROUTER_MODEL |
|
ollama |
OLLAMA_MODEL (no API key needed) |
Self-hosted; requires Ollama running locally |
deepseek |
DEEPSEEK_API_KEY, DEEPSEEK_MODEL |
|
xai |
XAI_API_KEY, XAI_MODEL |
|
cohere |
COHERE_API_KEY, COHERE_MODEL |
|
perplexity |
PERPLEXITY_API_KEY, PERPLEXITY_MODEL |
|
qwen |
QWEN_API_KEY, QWEN_MODEL |
|
moonshot |
MOONSHOT_API_KEY, MOONSHOT_MODEL |
|
zhipu |
ZHIPU_API_KEY, ZHIPU_MODEL |
|
yi |
YI_API_KEY, YI_MODEL |
|
baidu |
BAIDU_API_KEY, BAIDU_MODEL |
|
minimax |
MINIMAX_API_KEY, MINIMAX_MODEL |
|
custom |
CUSTOM_API_KEY, CUSTOM_MODEL, CUSTOM_BASE_URL |
Any OpenAI-compatible endpoint |
Example, switching to Groq in .env:
LLM_PROVIDER=groq
GROQ_API_KEY=your-groq-key
GROQ_MODEL=llama-3.3-70b-versatile
Running Ollama as a fallback/local provider from inside this app's Docker
container (rather than bare metal) has real, verified setup steps beyond
just setting OLLAMA_MODEL:
host.docker.internalisn't reliable on plain Linux Docker Engine (unlike Docker Desktop) — it can resolve to the wrong bridge network's gateway. If Ollama connections fail with this alias, overrideOLLAMA_BASE_URLwith the container's actual compose-network gateway IP directly, e.g.OLLAMA_BASE_URL=http://172.18.0.1:11434/v1(find your real gateway withdocker network inspect <network> | grep Gateway).- Ollama binds to
127.0.0.1by default, which a container can't reach even with the correct gateway IP. Override it to listen on all interfaces:
mkdir -p /etc/systemd/system/ollama.service.d
cat > /etc/systemd/system/ollama.service.d/override.conf << 'EOF'
[Service]
Environment="OLLAMA_HOST=0.0.0.0:11434"
EOF
systemctl daemon-reload && systemctl restart ollama- The host firewall may silently drop container→host traffic even
after the above — container-to-bridge-gateway traffic hits ufw's
INPUTchain, notFORWARD/DOCKER-USER. Allow it explicitly:ufw allow from 172.18.0.0/16 to any port 11434 proto tcp(adjust the subnet to match your actual Docker network).
RagLeap Core includes a real-time voice channel: Twilio Media Streams connects via WebSocket, your speech is transcribed with OpenAI Whisper, answered by the core RAG pipeline, and spoken back with OpenAI TTS. The voice-activity detection and echo-suppression logic is carried over from a production system tuned against real call traffic.
Runs as a separate service on port 8765 (see docker-compose.yml), since
Twilio's real-time audio protocol needs a raw WebSocket server, not an
HTTP route.
Setup:
- Set
OPENAI_API_KEYin.env(used for both Whisper STT and TTS in v1) - Optionally set
VOICE_BOT_NAME,VOICE_GREETING,VOICE_TTS_VOICE - Point a Twilio phone number's
<Connect><Stream>TwiML atwss://your-domain.com:8765
Honest status: the WebSocket server, Twilio event protocol handling, and error handling are verified working. The full Whisper/TTS round-trip has not yet been live-tested end-to-end (requires OpenAI API credits). If you try it and hit issues, please open one — this is exactly the kind of real-world testing this project needs.
Known limitations, carried over from production and not yet fixed here:
- Only OpenAI Whisper (STT) and OpenAI TTS are supported in v1 — Deepgram and ElevenLabs (multi-language support) are good-first-issue candidates
- Non-English TTS quality varies since OpenAI's TTS voices are English-tuned
- Typical round-trip latency in production was 6-8 seconds
RagLeap Core builds a lightweight entity co-occurrence graph alongside its vector index. When you ingest a document, entities (product names, acronyms, proper nouns) are extracted and linked in Neo4j. When you ask a question, the same extraction runs on your query, and any documents linked to matching entities get a similarity boost in retrieval — on top of, not instead of, normal pgvector search.
Runs as a fourth Docker Compose service on ports 7475/7688 (remapped
from Neo4j's defaults to avoid colliding with another Neo4j instance on the
same host). If Neo4j is unreachable or NEO4J_PASSWORD is unset, the graph
degrades gracefully — retrieval falls back to pure vector search, ingestion
is unaffected.
Setup:
- Set
NEO4J_URI,NEO4J_USER, andNEO4J_PASSWORDin.env(matching theNEO4J_AUTHvalue indocker-compose.yml) - Optionally set
DOMAIN_TERMS— a comma-separated list of domain-specific terms to boost during extraction (e.g.DOMAIN_TERMS=API,SDK,RAG)
Honest status: entity extraction, document graph writes, entity-based document lookup, and graph-boosted chat retrieval are all verified working end-to-end, including in CI (fresh build, real ingest, real query, real graph lookup). The graph boost is currently a simple additive score bump, not a full weighted re-ranker — a richer hybrid ranking system is a good next step for anyone who wants to dig in.
Known limitations:
- Entity extraction is regex-based (CamelCase, acronyms, capitalized phrases, plus optional domain terms) — not a trained NER model, so it will miss some entities and occasionally include noise
search_related_entities()(multi-hop graph traversal) is implemented but not yet wired into the retrieval pipeline — good-first-issue candidate for anyone wanting a project
RagLeap Core auto-detects language during document ingestion (per chunk)
and during chat (per query), using the langdetect library plus
script-based heuristics for CJK, Hangul, and Kana text. Since every
channel (WhatsApp, Telegram, Discord, Voice, and the API directly)
routes through the same core chat pipeline, detection applies
consistently everywhere without per-channel wiring.
Setup: works out of the box with no configuration. Optionally set
DEFAULT_LANGUAGE (fallback when detection fails or text is too short),
LANGUAGE_DETECTION_CONFIDENCE_THRESHOLD (default 0.7), and
LANGUAGE_DETECTION_SUPPORTED_LANGUAGES (comma-separated allowlist,
blank = unrestricted).
Honest status: verified working end-to-end — document-level detection tested at high confidence (0.9999) on a real mixed-language document, and query-level detection confirmed working via both the API and CLI.
Known limitations:
langdetectcovers roughly 55 languages — noticeably fewer than the hosted platform's 222+, which layers additional detection and per-user language preferences on top- Short queries in closely-related languages can be misdetected (in testing, a short French query was detected as Italian) — this is an inherent limitation of statistical detection on short text, not specific to this port. A good-first-issue candidate for anyone wanting to improve short-query accuracy
- Detection is one-way only: RagLeap Core detects the query's language and surfaces it, but does not yet steer the AI's response language to match — that's a reasonable next step for a contributor
flowchart TD
subgraph Connectors["10 connectors, one shared interface"]
C1[MySQL] --- C2[PostgreSQL] --- C3[MongoDB]
C4[REST API] --- C5[Salesforce] --- C6[HubSpot]
C7[Shopify] --- C8[Google Sheets] --- C9[Stripe]
C10[CSV Upload]
end
Connectors --> Svc["RealTimeExternalDataService<br/>real-time query, 5-min cache, no separate sync step"]
Owner["Owner configures an action:<br/>trigger phrase or auto-match on SQL/field names<br/>+ a query template (SELECT / UPDATE / INSERT / DELETE)"] --> Svc
Svc --> Match["match_and_execute_action(workspace_id, user_message, user_identifier)<br/>extracts order_id / email / phone from the message,<br/>substitutes into the template, runs the query"]
Match --> Chat["Chat channels (WhatsApp/Telegram/Discord/Web)<br/>api/personal_bot_views.py"]
Match --> Voice["Voice channel<br/>memory/voice_views.py::twilio_voice_speech<br/>wired in 2026-08"]
Chat --> RAGCtx["Result injected as context<br/>into the RAG prompt"]
Voice --> RAGCtx
RAGCtx --> Answer["AI answers with real account/order/appointment<br/>data, not just document knowledge"]
Verified live, read from api/addon_realtime.py's RealTimeExternalDataService. Two things worth being direct about: match_and_execute_action genuinely supports write queries (UPDATE/INSERT/DELETE), not just read-only lookups — owner-configured, so the safety boundary is whatever SQL the owner writes into the template, not something the framework restricts on its own. And until 2026-08, this action-matching step only ran on chat channels; voice calls had no equivalent, which is the gap closed in the Voice Channel Routing diagram above.
RagLeap Core connects to external databases and business tools, syncing per-user context to personalize RAG responses. Nine connectors are included: MySQL, PostgreSQL, MongoDB, generic REST APIs, Salesforce, HubSpot, Shopify, Google Sheets, and Stripe.
Every CRM/SaaS connector uses credentials you provide directly — a username/password, a private-app token, an admin API token, a service-account JSON file, or a secret key, depending on the service. None require registering an OAuth app; nothing here depends on RagLeap owning any third-party developer account.
Credentials are encrypted at rest (Fernet/AES-128) before being stored.
Setup:
- Generate an encryption key:
python3 -c "from cryptography.fernet import Fernet; print(Fernet.generate_key().decode())" - Set
ADDON_ENCRYPTION_KEYin.envto that value - Install the SDK for the connector(s) you want (each is optional — see
requirements.txt) - Create a data source:
POST /integrationswithname,source_type, and the relevant credential fields - Test it:
POST /integrations/{id}/test - Sync it:
POST /integrations/{id}/sync
Honest status: verified end-to-end against a real public API — connection testing, syncing, correct identifier-field matching, and credential encryption (checked as actual ciphertext in the database, not just assumed) all confirmed working.
Known limitations:
- 9 of the 18 source types listed in the hosted platform's UI have real connectors here. CSV Upload, Snowflake, BigQuery, WooCommerce, Airtable, Notion, Razorpay, Slack, and Gmail are good-first-issue candidates for anyone wanting to add one
- Sync is on-demand only (
POST /integrations/{id}/sync) — no scheduled background sync yet, though the schema trackssync_interval_minutesfor a future Celery-beat-equivalent - Synced context isn't automatically injected into chat responses yet — each channel adapter would need to know its own user's identifier first, which is a reasonable next contribution
Beyond the core RAG pipeline, /chat (and the underlying core.chat.ask())
support several controls aimed at production use: retrieval quality,
response latency, provider reliability, and cost.
Hybrid search (dense + sparse). By default, retrieval combines
pgvector cosine similarity with Postgres full-text search (tsvector/
GIN index), fused via Reciprocal Rank Fusion — catching both semantic
matches and exact keyword/identifier matches a pure embedding search can
miss. Pass hybrid=false to use dense-only retrieval instead (cheaper —
one query instead of two).
Streaming. POST /chat/stream streams the answer as it's generated
(text/plain, chunked transfer) instead of waiting for the full response.
Implemented natively per provider (Gemini, Anthropic, and OpenAI-compatible
each have different streaming APIs — all three are real, not one stubbed).
Provider fallback. Set LLM_FALLBACK_PROVIDERS (comma-separated) to
automatically retry with backup providers if the primary fails — a rate
limit, outage, or bad key on your primary provider doesn't have to mean a
failed request. Each fallback needs its own API key configured normally.
Streaming can only fall back before any text has been sent to the
caller — a mid-stream failure surfaces as an error rather than silently
switching providers and confusing the output.
Generation controls. temperature, system_prompt, and max_tokens
are all real per-call parameters (not just env-var defaults) — build your
own agent behavior on top of RagLeap's retrieval without forking the
library.
Token usage & context budget. Every blocking /chat call returns real
token usage (prompt_tokens, completion_tokens, total_tokens) pulled
directly from the provider's response — not an estimate. Retrieved
context is also trimmed to MAX_CONTEXT_CHARS (default 12000, roughly
4 characters per token for English text) before being sent, dropping the
lowest-ranked chunks first, so you're not paying for more context than
necessary. Set MAX_CONTEXT_CHARS=0 to disable trimming.
Honest status: hybrid search's RRF fusion math verified correct
against hand calculation. Streaming verified working end-to-end for the
default provider. Provider fallback verified with a real broken-primary
test — deliberately invalid API key, confirmed fallback to a working
secondary provider with a correct answer. Token usage and context
trimming verified with real numbers: a 3-chunk retrieval trimmed to 1
chunk under a tight budget reduced actual prompt_tokens by 38% on the
same live API.
Known limitations:
- Token usage reporting is not available for streaming responses — each provider's streaming API surfaces usage differently, and doing all three correctly is separate, not-yet-done work
MAX_CONTEXT_CHARSis a character-count approximation (~4 chars/token for English), not an exact per-provider tokenizer count- Hybrid search hasn't been benchmarked for actual ranking-quality improvement on a multi-document corpus with genuinely conflicting dense vs. sparse rankings — only correctness (fusion math, tokenization of unusual identifiers) has been verified so far
- Public repository created
- Core RAG engine extracted and cleaned from production codebase
- Standalone Docker Compose setup (no external Django project dependency)
- Document ingestion module (28+ formats, not just PDF/TXT/DOCX)
- Web chat widget
- Bring-your-own-API-key support (19 providers)
- WhatsApp, Telegram, Discord, and Voice channel adapters (single-tenant)
- Knowledge Graph (Neo4j), language detection, database/CRM integrations
- Contribution guide and good-first-issue labels — 8+ issues labeled, with a real external contributor active on #134
- Community Discord
See ROADMAP.md for the full phase-by-phase history.
RagLeap Core is working, tested, and open for contributions now. See CONTRIBUTING.md for how to get started, and check the good first issue label for scoped tasks.
Want to see exactly what's being worked on and what's open to claim? Check the Project board — issues are staged as Good First Issue, Ready (Scoped), or Needs Scoping, so you can pick something that matches how much design work you want to do versus just build.
Student, professor, or looking for a capstone/thesis project? See STUDENT_PROJECTS.md for scoped project ideas at starter, semester, and research-grade levels.
- GitHub Issues — bugs and feature requests
- GitHub Discussions — ideas and questions
- ragleap.com — the project website
MIT © 2026 RagLeap
"could not translate host name 'db'" error after a failed docker compose up:
If your first docker compose up attempt fails (e.g. a port conflict on 5433 or 8000), a retry can sometimes leave the database container attached to a stale, orphaned Docker network. Fix:
docker compose down
docker network prune -f
docker compose up --build -dPort 5433 or 8000 already in use:
Another instance of this project (or something else) is using the port. Either stop it, or change the host-side port mapping in docker-compose.yml (the "5433:5432" and "8000:8000" lines) to something free.
Documentation reference and guidelines for #135.