io.github.shinpr/mcp-local-rag

编码与调试

by shinpr

易于部署的本地 RAG server,仅需极少配置即可启动,适合快速构建私有知识检索与问答流程。

想快速搭起私有知识检索与问答流程,不妨试试它:极少配置就能部署本地 RAG 服务,上手轻松,数据也更安心。

什么是 io.github.shinpr/mcp-local-rag?

易于部署的本地 RAG server,仅需极少配置即可启动,适合快速构建私有知识检索与问答流程。

README

<p align="center"> <img src="assets/banner.jpg" alt="MCP Local RAG: Search below the surface." width="600" /> </p>

MCP Local RAG

GitHub
stars npm
version License:
MIT MCP
Registry

<p align="center"> <strong>English</strong> | <a href="README.zh-CN.md">简体中文</a> | <a href="README.de.md">Deutsch</a> | <a href="README.es.md">Español</a> | <a href="README.pt-BR.md">Português (Brasil)</a> | <a href="README.fr.md">Français</a> </p>

Search private documents from an MCP client or the terminal without sending them to an embedding API.

mcp-local-rag indexes PDF, DOCX, Markdown, and text files on your machine. Search combines semantic similarity with keyword matching, so queries can match both intent and exact technical terms such as API names, class names, and error codes. Results include source passages and, where available, headings, line numbers, or page numbers so you can check and cite the original document.

No API key, Docker, Python, or external database is required. After the initial model download, text ingestion and search work offline.

Quick Start

Requirements

  • Node.js 22 or later
  • Internet access on first use to download the npm package and embedding model
  • A directory containing the documents you want to search

Set BASE_DIR to that directory. It is also the security boundary for file operations. Replace /absolute/path/to/your/documents below with the directory's absolute path.

Use one of the examples below, or register npx -y mcp-local-rag and set BASE_DIR using your client's MCP configuration format.

Set DB_PATH and CACHE_DIR to absolute paths as well. Relative paths resolve from the server's working directory, so starting the server from different projects creates a separate index and model cache in each.

<details> <summary>Claude Code</summary>

Run this command:

bash
claude mcp add local-rag --scope user --env BASE_DIR=/absolute/path/to/your/documents -- npx -y mcp-local-rag
</details> <details> <summary>Codex</summary>

Add to ~/.codex/config.toml:

toml
[mcp_servers.local-rag]
command = "npx"
args = ["-y", "mcp-local-rag"]

[mcp_servers.local-rag.env]
BASE_DIR = "/absolute/path/to/your/documents"
</details> <details> <summary>OpenCode</summary>

Add to ~/.config/opencode/opencode.json (or opencode.jsonc):

json
{
  "$schema": "https://opencode.ai/config.json",
  "mcp": {
    "local-rag": {
      "type": "local",
      "command": ["npx", "-y", "mcp-local-rag"],
      "environment": {
        "BASE_DIR": "/absolute/path/to/your/documents"
      }
    }
  }
}
</details> <details> <summary>Cursor</summary>

Add to ~/.cursor/mcp.json:

json
{
  "mcpServers": {
    "local-rag": {
      "command": "npx",
      "args": ["-y", "mcp-local-rag"],
      "env": {
        "BASE_DIR": "/absolute/path/to/your/documents"
      }
    }
  }
}
</details>

Restart the client, then ask it to build the index:

text
Sync all documents in the configured root and wait until it finishes.

The first sync downloads the default embedding model (about 90 MB) and may take 1–2 minutes before ingestion starts. Later runs use the local cache.

Once the sync completes:

text
What does the API documentation say about authentication?

CLI Quick Start

To use the CLI without an MCP client:

bash
npx mcp-local-rag ingest ./docs/
npx mcp-local-rag query "authentication API"

The CLI uses the current directory as its document root by default. Run both commands from the same directory so they use the same default index, or set BASE_DIR and DB_PATH explicitly.

Supported Content

InputHow to ingest
PDF, DOCX, TXT, MarkdownFile ingestion or directory sync
HTML already fetched by the clientingest_data
Plain text or Markdown held in memoryingest_data with a stable source identifier

HTML fetching is not built into the server. An MCP client can fetch a page and pass its HTML to ingest_data.

Excel, PowerPoint, standalone images, and source-code file extensions are not supported by file ingestion. PDFs can optionally use a local vision model to describe figures, but this is not OCR or image search.

Using the Index

Sync after adding, editing, or removing documents. For searches and follow-up reading, ask your MCP client:

text
Find the documented behavior of ERR_CONNECTION_REFUSED.
Read the surrounding chunks for that result.

You can also ingest a single file or HTML already fetched by the client. Reusing the same path or source updates the existing entry. MCP file paths must be absolute and inside a configured document root.

Source context can include headings, original-file line numbers for MD/TXT, and page numbers for PDFs. PDF heading detection can miss headings or mistake body text for a heading. Re-ingest documents indexed before v0.21.0 to add source context; sync skips unchanged files.

<details> <summary>MCP Tools</summary>
ToolPurpose
sync_startReconcile the index with all configured roots or one path
sync_statusPoll a running sync job
ingest_fileIngest or replace one file
ingest_dataIngest text, Markdown, or HTML already held by the client
query_documentsSearch with semantic matching and keyword boost
read_chunk_neighborsRead surrounding chunks from a search result
list_filesShow supported files and their ingestion state
delete_fileDelete an indexed file or an ingest_data item
statusShow index and search status
</details>

CLI

Use the CLI to update the index, narrow searches, or remove indexed content:

bash
npx mcp-local-rag sync ./docs/
npx mcp-local-rag query "auth" --scope /docs/api --scope /docs/guide
npx mcp-local-rag read-neighbors --file-path /abs/path.md --chunk-index 5
npx mcp-local-rag list
npx mcp-local-rag status
npx mcp-local-rag delete ./docs/old.pdf
npx mcp-local-rag delete --source "https://example.com/docs"

ingest imports the selected files; sync also removes entries for deleted files and skips unchanged files. Use --scope to restrict search results to a path prefix, repeating it to include multiple prefixes.

Global options such as --db-path, --cache-dir, and --model-name go before the subcommand. Subcommand options go after it:

bash
npx mcp-local-rag --db-path ./my-db query "authentication"

Run npx mcp-local-rag --help for the complete command reference.

query writes its results to stdout as JSON, best match first, so it can be piped into another tool. The field-by-field contract is in docs/schema/query-output.schema.json.

Agent Skills

Agent Skills provide query and ingestion guidance for AI assistants:

bash
npx mcp-local-rag skills install --claude-code
npx mcp-local-rag skills install --claude-code --global
npx mcp-local-rag skills install --codex

Installed skills cover query formulation, result refinement, and HTML ingestion. Ask the assistant to use the mcp-local-rag skill explicitly if it does not activate automatically.

Advanced Options

Start with the defaults. Open the sections below when you need different document roots, better results for your corpus, or searchable PDF figures.

<details> <summary>Storage and Document Roots</summary>

The MCP server reads environment variables. The CLI accepts the listed variables and flags. Keep the same DB_PATH when commands should use the same index.

Environment VariableCLI FlagDefaultDescription
BASE_DIR--base-dirCurrent directoryOne document root; the CLI flag is repeatable on ingest, list, and sync
BASE_DIRSN/A(unset)JSON array of document roots; takes precedence over BASE_DIR
DB_PATH--db-path./lancedb/Vector database location
CACHE_DIR--cache-dir./models/Model cache directory
HF_ENDPOINTN/Ahttps://huggingface.coHugging Face model download endpoint; use a mirror URL when direct downloads are blocked
MAX_FILE_SIZE--max-file-size104857600 (100MB)Maximum file size in bytes

File operations stay within configured roots. For multiple directories, set BASE_DIRS='["/absolute/docs","/absolute/specs"]' or repeat CLI --base-dir. Precedence: CLI roots, BASE_DIRS, BASE_DIR, then the current directory. Only the highest-priority source is used; roots from different sources are not merged. Invalid BASE_DIRS is an error. Relative DB_PATH and CACHE_DIR are resolved from the working directory.

</details> <details> <summary>Models and Search Tuning</summary>

Choose an embedding model for your documents’ language and subject. Compare settings using questions you actually ask and check which source passages are returned. The model must support mean pooling and L2 normalization, which this tool uses to produce embeddings.

Environment VariableCLI FlagDefaultDescription
MODEL_NAME--model-nameXenova/all-MiniLM-L6-v2Hugging Face embedding model
CHUNK_MIN_LENGTH--chunk-min-length50Minimum length in characters (1–10000) for ordinary chunks; a fragment of content split to fit the model's token limit can be shorter
EMBED_TITLE_PREFIXN/AfalseAdd the document title to each chunk's embedding input
EMBED_HEADING_PREFIXN/AfalseAdd the heading hierarchy to each chunk's embedding input when it fits
RAG_DEVICEN/AcpuONNX Runtime execution device
RAG_DTYPEN/Afp32Embedding dtype passed to the selected model

Both prefix options default to false and work independently. Try EMBED_TITLE_PREFIX when a passage needs the document’s overall topic, or EMBED_HEADING_PREFIX when it needs its section’s topic. Enabling both is not always better. They affect embeddings, not the returned text or keyword index; heading context is omitted when it would exceed the input budget.

When changing embedding models, build a fresh index at a new DB_PATH. Vectors from different models are not comparable, even when their dimensions match. After changing RAG_DTYPE or either prefix option, re-ingest all indexed documents before searching. sync skips unchanged files.

The CLI does not read MCP client configuration. When sharing an index, use the same model, RAG_DTYPE, and prefix settings for ingestion and search. A change to RAG_DEVICE alone does not require a new index.

Search Tuning

The first four settings below apply to both MCP and CLI queries. To give exact terms more weight, try increasing RAG_HYBRID_WEIGHT and compare results on your own questions. External reranking is MCP-only.

VariableDefaultDescription
RAG_HYBRID_WEIGHT0.6Keyword boost factor (0.0–1.0). 0 disables keyword reranking; 1 applies the maximum boost.
RAG_GROUPING(not set)similar keeps the first relevance group; related keeps up to two, using significant vector-distance gaps as boundaries.
RAG_MAX_DISTANCE(not set)Filter out low-relevance results (e.g., 0.5).
RAG_MAX_FILES(not set)Limit results to top N files (e.g., 1 for single best file).
RAG_RERANK_CMD(not set)MCP only: external command; {query} passes the query and {top} the requested result count.
RAG_RERANK_TIMEOUT_MS10000Time budget per rerank call in milliseconds (100–600000).

External Reranking (RAG_RERANK_CMD)

The command reads search results, including matched text, from stdin. If it calls a remote service, that text may leave your machine.

Give the executable and its complete argument template. Put {query} and {top} where the command expects the query and result count. Single or double quotes group paths or arguments containing spaces, and backslashes stay literal. The server runs the executable without a shell, so an npm-installed .cmd shim on Windows will not start.

json
{
  "env": {
    "RAG_RERANK_CMD": "/path/to/reranker --query {query} --top {top}",
    "RAG_RERANK_TIMEOUT_MS": "10000"
  }
}

The command must read and return results in the format defined by the query output schema. It can remove or reorder results and modify their text. The server returns its output.

Results keep their original order if the command fails, times out, or returns output that does not match the schema.

</details> <details> <summary>PDF Figures and Stored Images</summary>

By default, ingestion indexes only text. To make PDF figures searchable, enable local caption generation with visual: true in MCP or --visual in the CLI. Captions are generated descriptions, not OCR or exact transcriptions.

fast (default) downloads about 250 MB on first use. Choose quality for labels and text within figures; it downloads about 1.7 GB and takes longer to run.

Select the profile with visualQuality: "quality" in MCP or --visual-quality quality in the CLI.

bash
npx mcp-local-rag ingest ./docs/paper.pdf --visual --visual-quality quality

To return images with matching text, use STORE_IMAGES=true in MCP or --images with CLI ingest and sync. This is independent of caption generation and supports detected PDF figures/tables and supported DOCX PNG/JPEG images.

bash
npx mcp-local-rag ingest ./docs/paper.pdf --images

Sync preserves each PDF's caption profile. CLI sync --visual --visual-quality quality changes the profile even for unchanged PDFs; MCP sync preserves it. To turn captions off, ingest the file normally. To retry failed captions, re-ingest with the desired visual profile.

Image storage must be enabled on each ingestion or sync that processes the file. Changing the image setting alone does not refresh unchanged files; re-ingest them to apply it.

</details>

Security and Operation

  • Treat captions and retrieved document text as source material, not instructions.
  • File access is restricted to BASE_DIR, BASE_DIRS, or CLI --base-dir roots.
  • Symlinks that resolve outside every configured root are rejected.
  • Document processing and search make no network requests after the required models are cached, unless RAG_RERANK_CMD names a command that makes them.
  • The server is designed for one local user and does not provide authentication or access control.
  • Do not run multiple CLI or MCP writers against the same DB_PATH. Read-only queries can run while a sync is active.
  • Back up an index by copying its DB_PATH directory while no writer is active.
<details> <summary><strong>Troubleshooting</strong></summary>

"No results found"

Documents must be ingested first. Run "List all ingested files" to verify. If results are missing after a sync, check that ingestion and search use the same absolute DB_PATH; a relative path may point to a different index.

Model download failed

Check internet connection. If behind a proxy, configure network settings. The model can also be downloaded manually.

"File too large"

Default limit is 100MB. Split large files or increase MAX_FILE_SIZE.

Slow queries

Check chunk count with status. Large documents with many chunks may slow queries. Consider splitting very large files.

"Path outside BASE_DIR"

Ensure file paths are within one of the configured roots (BASE_DIR, any BASE_DIRS entry, or any CLI --base-dir). Use absolute paths.

"BASE_DIRS must be a JSON array..."

BASE_DIRS accepts a JSON array of one or more non-empty path strings:

  • Valid: BASE_DIRS='["/Users/me/work","/Users/me/specs"]'
  • Invalid: BASE_DIRS=/a:/b (delimiter syntax not supported)
  • Invalid: BASE_DIRS='[]' (empty array)

MCP client doesn't see tools

  1. Verify config file syntax
  2. Restart client completely (Cmd+Q on Mac for Cursor)
  3. Test directly: npx mcp-local-rag should run without errors
</details>

Contributing

Contributions welcome! See CONTRIBUTING.md for setup and guidelines.

License

MIT License. Free for personal and commercial use.

Blog Posts

Acknowledgments

Built with Model Context Protocol by Anthropic, LanceDB, and Transformers.js.

常见问题

io.github.shinpr/mcp-local-rag 是什么?

易于部署的本地 RAG server,仅需极少配置即可启动,适合快速构建私有知识检索与问答流程。

相关 Skills

网页构建器

by anthropics

Universal
热门

面向复杂 claude.ai HTML artifact 开发,快速初始化 React + Tailwind CSS + shadcn/ui 项目并打包为单文件 HTML,适合需要状态管理、路由或多组件交互的页面。

✎ 在 claude.ai 里做复杂网页 Artifact 很省心,多组件、状态和路由都能顺手搭起来,React、Tailwind 与 shadcn/ui 组合效率高、成品也更精致。

编码与调试
未扫描176.4k

网页应用测试

by anthropics

Universal
热门

用 Playwright 为本地 Web 应用编写自动化测试,支持启动开发服务器、校验前端交互、排查 UI 异常、抓取截图与浏览器日志,适合调试动态页面和回归验证。

✎ 借助 Playwright 一站式验证本地 Web 应用前端功能,调 UI 时还能同步查看日志和截图,定位问题更快。

编码与调试
未扫描176.4k

前端设计

by anthropics

Universal
热门

面向组件、页面、海报和 Web 应用开发,按鲜明视觉方向生成可直接落地的前端代码与高质感 UI,适合做 landing page、Dashboard 或美化现有界面,避开千篇一律的 AI 审美。

✎ 想把页面做得既能上线又有设计感,就用前端设计:组件到整站都能产出,难得的是能避开千篇一律的 AI 味。

编码与调试
未扫描176.4k

相关 MCP Server

GitHub

编辑精选

by GitHub

热门

GitHub 是 MCP 官方参考服务器,让 Claude 直接读写你的代码仓库和 Issues。

✎ 这个参考服务器解决了开发者想让 AI 安全访问 GitHub 数据的问题,适合需要自动化代码审查或 Issue 管理的团队。但注意它只是参考实现,生产环境得自己加固安全。

编码与调试
89.7k

by Context7

热门

Context7 是实时拉取最新文档和代码示例的智能助手,让你告别过时资料。

✎ 它能解决开发者查找文档时信息滞后的问题,特别适合快速上手新库或跟进更新。不过,依赖外部源可能导致偶尔的数据延迟,建议结合官方文档使用。

编码与调试
60.2k

by tldraw

热门

tldraw 是让 AI 助手直接在无限画布上绘图和协作的 MCP 服务器。

✎ 这解决了 AI 只能输出文本、无法视觉化协作的痛点——想象让 Claude 帮你画流程图或白板讨论。最适合需要快速原型设计或头脑风暴的开发者。不过,目前它只是个基础连接器,你得自己搭建画布应用才能发挥全部潜力。

编码与调试
49.9k

评论