The command is ctxt — short for Contexting.
Contexting keeps a live map of your codebase so AI agents can reason about paths without hunting through the filesystem manually. It builds a recursive JSON tree of every folder and file, extracts code symbols (functions, classes, types, variables) using language-specific extractors, attaches LLM-generated synonyms, and exposes ranked search hints plus health tooling.
Read the visual getting started guide for the product story, architecture, installation paths, MCP setup, and privacy notes.
ctxt supports Linux, macOS, and Windows on amd64 and arm64. Download the
archive for your platform from GitHub Releases, verify it against
checksums.txt, and place the binary on your PATH. Go users can install
from source:
go install github.com/ktappdev/contexting/cmd/ctxt@latest
cd your-repo
ctxt init .
ctxt watch .Set an API key for synonym generation (optional but recommended):
export OPENROUTER_API_KEY="sk-or-v1-..."First run creates a .ctxt/ctx_config.toml config file. Edit it, then press Enter to continue. Subsequent runs use the saved config.
Synonyms are optional. Contexting works without an LLM — you get symbols and path matching, just no synonym expansion. With an LLM, search gets a significant boost since synonyms bridge the gap between how code is named and how developers talk about it.
The default uses OpenRouter with deepseek/deepseek-v4-flash — fast, nearly free (~$0.0004 per 60 names). But any OpenAI-compatible API works:
- OpenRouter (default) — access to dozens of models, free tier available
- Local — point
endpointat any local server (Ollama, llama.cpp, vLLM) - Other OpenAI-compatible APIs — set
endpoint,model, andapi_key_env - No LLM — skip the
[llm]section entirely or setapi_key = ""
Adjust batch_size, parallel_requests, and synonyms_min/synonyms_max to trade off speed vs. coverage for your model.
Remote LLM endpoints must use HTTPS; HTTP is accepted only on loopback.
ctxt init walks the filesystem tree and builds .ctxt/ctx_index.json — a single recursive JSON tree where every entry has full_path, type, optional symbols, optional synonyms, and nested children for directories.
Symbols: Every source file gets scanned with language-specific extractors that pull out named declarations, including Go receiver methods — functions, types, variables, classes, constants. These become a symbols array on the file node. The default extractor is auto (tree-sitter with regex fallback). Supported languages: Go (go/parser), Python/JavaScript/TypeScript/Rust/Svelte/Astro (tree-sitter), Vue/Ruby (regex fallback).
Synonyms: Each node (file or directory) gets 5–12 LLM-generated synonyms via batched API calls to an OpenAI-compatible endpoint. For example, Skeletons.tsx → ["skeletons", "loading", "placeholder", "animation"]. Names are smart-batched (up to 60 per request, 10 parallel by default). The prompt includes extracted symbols and, for JS/TS files, ESM imports to generate contextual terms. Use --offline to guarantee that no LLM requests are made.
Bootstrap diff: On subsequent runs, watch diffs the filesystem against the existing snapshot using file modification times and sizes. Timestamp-preserving edits with a different size are detected; edits preserving both size and timestamp require a full rebuild. Deleted files are removed, new files are added, modified files are re-extracted. Watch bootstrap reuses cached synonyms; run ctxt sync to fill gaps and refresh changed symbol/import context. Synonym cache keys use the full project-relative path (e.g., billing/routes/index.ts) so identically named files in separate features remain distinct.
When you run search-hints "product review rating", the query is lowercased and split into tokens. Each token is scored against every node in the tree using the matchers below. Each token receives points for its strongest path match, strongest synonym match, and up to three distinct symbol matches, so repeated mentions of one concept cannot inflate a file’s score:
| Matcher | Points | Description |
|---|---|---|
| exact | +12 | Token exactly matches a directory basename |
| basename | +7 | Token matches the file's basename |
| exact basename | +15 | Entire query exactly matches the full filename (with or without extension) |
| path segment | +4 | Token matches any segment in the full path |
| segment prefix | +5 | Token is a prefix of a path segment |
| syn-exact | +8 | Token exactly matches a synonym |
| syn overlap | +5 | Token partially overlaps a synonym |
| sym exact | +24 | Token exactly matches a symbol name |
| sym contains | +5 | Token is contained in a symbol name |
| sym token | +4 | Token matches a camelCase/PascalCase split token in a symbol |
The --explain flag reveals this breakdown:
syn exact +8: loading
sym contains +5: PageSkeleton
Matching additional distinct query terms adds +4 per term after the first; generated stems do not count as extra query terms. Exact symbol matches outweigh three partial symbol matches. Results are ranked by total score. Low-signal short/common words are filtered from query and synonym matching to reduce noise. Results are truncated at confidence gaps — if there's a 50%+ score drop between consecutive results, lower-scoring results are omitted even if --limit would include them. Exact filename/symbol lookups can return just one strong match; conceptual queries retain up to three alternatives. Zero-score candidates are never returned.
ctxt watch maintains the index in memory. It watches for filesystem changes via debounced events (750ms default), re-extracts symbols for modified files, and serves a loopback endpoint advertised by .ctxt/ctx_runtime.json. The on-disk snapshot is flushed on graceful shutdown.
Create a full snapshot in .ctxt/ctx_index.json with extracted symbols and optional synonyms.
ctxt init .
ctxt init . --output .ctxt/ctx_index.json --synonym-cache .ctxt/ctx_cache.jsonKey flags:
--no-config-prompt,--create-config— non-interactive automation--llm-model,--batch-size,--synonyms-min,--synonyms-max,--api-key,--ignore--symbol-extractor— symbol extraction mode:auto(default, tree-sitter with regex fallback),treesitter, orregex-v, --verbose— show symbol extraction progress and batch completion- Always rebuilds the entire tree; use when you need a clean snapshot
On first run, creates a starter .ctxt/ctx_config.toml and pauses so you can edit it.
Targeted synonym generation — generates synonyms for indexed names that are missing, below synonyms_min, or have changed symbols/imports. It refreshes symbol extraction from disk before comparing context fingerprints. Works on the existing index without rebuilding the file tree.
ctxt sync .
ctxt sync . --path command_sync.go
ctxt sync . --path docs --force--path selects an indexed file or directory; --force regenerates selected
entries even if their fingerprints are current. Use --force after changing
models or generation instructions. Body-only edits that leave symbols/imports
unchanged also require --force.
Successful LLM responses are saved even if another batch fails. The command
still exits with an error; rerun without --force to request only missing,
short, or stale entries. --batch-size explicitly controls batch size when set.
Only version 2 caches and indexes are supported. Older files are rejected; there is no migration or synonym-key fallback. For the default paths, rebuild:
rm .ctxt/ctx_cache.json
ctxt init .Use ctxt init . --offline to rebuild without LLM requests.
Use init or restart watch after adding files, then run sync to fill gaps.
Keep the index in memory with live filesystem updates.
ctxt watch . --debounce 750ms --verboseKey flags:
--llm-on-watch(default false) — opt into live remote synonym enrichment--search-log(default false) — log memory search requests; queries may be sensitive--search-log-query-max(default 120) — truncate logged queries--persist(default "shutdown") — snapshots flush on graceful shutdown- Starts a local memory-search endpoint and writes
.ctxt/ctx_runtime.json - Events applied via a single worker; logs show changed files per cycle
Query the index for ranked paths with explainable scores.
ctxt search-hints "update storage" --json
ctxt search-hints "routing auth" --dir-summary --dir-limit 5 --drill-limit 3Flags:
--limit,--min-score,--type files|dirs|all--dir-summary,--dir-limit,--drill-limit— top-down directory-first results--explain,--show-tokens,--json--memory(default true) — query live watch index first, fall back to snapshot--memory-only— fail if live memory unavailable--runtime-file— path to runtime discovery file (default.ctxt/ctx_runtime.json)--hybrid— augment index results by searching indexed files with ripgrep (default false); honors type and minimum-score filters--hybrid-score— score assigned to content-matched results (default 1)--hybrid-root— project root for content matching (defaults to index root)
Benchmark Hit@1/3/5 + MRR from manual query cases.
ctxt eval --cases ctx_cases.json --jsonInput format (version 2 with categories):
{
"version": 2,
"categories": {
"path-intent": {
"description": "Find file by describing path/location purpose",
"cases": [
{"query": "auth middleware", "expect_any": ["internal/auth/middleware.go"]}
]
}
}
}Bare arrays and other case versions are rejected.
Benchmark ctxt against find, grep, fd, rg, hybrid, and combined engines.
ctxt bench --cases docs/bench_cases.json --by-category
ctxt bench --cases docs/bench_cases.json --engines ctxt,find --jsonFlags:
--cases— path to case file (v2 format with categories)--engines— comma-separated list: ctxt,find,grep (default); also available: fd, rg, hybrid, combined--by-category— group results by category--json— output structured JSON--limit— max results per engine (default: 10)--min-score— minimum score threshold--root— project root--index— path to index file--grep-max-bytes— shared max file size for grep, rg, and hybrid content search--runs— measured runs per query and engine (default: 5)--warmup— unmeasured warmups per query and engine (default: 1)
Benchmark answers use exact, case-sensitive project-relative paths. All engines share the indexed file set and result limit; ctxt/hybrid rank by relevance, while other engines return alphabetical results. Reports distinguish query success from document recall and estimate path-only token size. JSON report version 2 includes timing samples, methodology, per-query results, and errors. Missing tools or failed commands cause a nonzero exit. See the benchmark guide for timing scope and validation cases.
Report index health, watch state, and path information.
ctxt status
ctxt status --jsonRemove the .ctxt/ directory.
ctxt clean
ctxt clean --dry-runHealth-check config, root, index, cache, and API key.
ctxt doctor --jsonCreate or overwrite .ctxt/ctx_config.toml:
ctxt config init --output .ctxt/ctx_config.toml.ctxt/ctx_config.toml drives all defaults. CLI flags override config, which overrides hard-coded defaults.
[common]
output = ".ctxt/ctx_index.json"
synonym_cache = ".ctxt/ctx_cache.json"
llm_model = "deepseek/deepseek-v4-flash"
batch_size = 8 # names per LLM request; 0 enables smart batching
synonyms_min = 5 # min synonyms per name
synonyms_max = 12 # max synonyms per name
ignore = [".git", ".venv", "site-packages", "__pycache__", "node_modules", "vendor", "dist", "migrations", "pb_migrations", "alembic", "flyway"]
dot_whitelist = [] # extra dot files to keep (merged with built-in defaults)
verbose = true
[llm]
provider = "openrouter"
endpoint = "https://openrouter.ai/api/v1/chat/completions"
model = "deepseek/deepseek-v4-flash"
api_key_env = "OPENROUTER_API_KEY" # reads key from env var
temperature = 0.9
parallel_requests = 10 # concurrent LLM batchesLLM config resolution: flag → config api_key → config api_key_env → LLM_API_KEY → OPENROUTER_API_KEY.
Supported providers: OpenRouter (default) and APIs that implement the
OpenAI chat-completions request/response shape. The provider value labels
the configuration; endpoint and model control the request.
Contexting loads common .gitignore patterns by default. Additional ignores come from:
- Built-in defaults —
.git,.venv,site-packages,__pycache__,node_modules,vendor,dist,migrations,pb_migrations,alembic,flyway .gitignorepatterns — loaded from the project rootignorein config — extra patterns merged with defaults
Root .gitignore rules support ordered negation (!path), rooted patterns,
directory-only patterns, and ** globs. Negation can undo earlier Git rules,
but cannot re-include an excluded parent directory or override built-in/config
exclusions. Symlink entries are not followed or indexed.
Dot files are skipped by default (any path segment starting with .). whitelisted dot files (.env.example, .prettierrc, .editorconfig, etc.) remain eligible unless explicitly ignored. Add more via dot_whitelist in config.
init → walk filesystem → extract symbols → build symbols map → (LLM: generate synonyms WITH symbols) → .ctxt/ctx_index.json
↓
watch → load .ctxt/ctx_index.json → keep in RAM → filesystem events → mutate in-memory
↓
search-hints → load .ctxt/ctx_index.json (or query live memory) → score tokens → ranked results
.ctxt/ctx_index.json— schema version, root path, timestamp, and indexed tree.ctxt/ctx_cache.json— versioned project-relative synonyms and symbol/import context fingerprints.ctxt/ctx_config.toml— config-driven defaults.ctxt/ctx_runtime.json— live watch discovery for memory search- Bench/eval case files — v2 format with categories (path-intent, symbol-lookup, concept-synonym, exact-file, narrow-scope, vague-intent); requires version 2
Initial indexing supports up to 10,000 files after ignores. Larger projects fail with an actionable error instead of producing a partial index. Add ignore patterns and retry. At 5,000 files, ctxt warns that narrower ignores may help.
go test ./...ctxt doctor --jsonfor diagnostics- If
.ctxt/ctx_index.jsonis stale, restart watch or runctxt init - If you changed ignore rules, run
ctxt initor restartwatchto rebuild - Synonym generation requires the configured key environment or
--api-key. Use--offlineto disable every LLM request regardless of configuration. - Watch mode must be stopped gracefully (Ctrl+C) to flush the snapshot
- If a writer crashes, verify no ctxt writer remains before removing the
empty
.ctxt-writerdirectory. - See SECURITY.md before enabling a remote LLM on private code.