io.github.ggozad/haiku-rag
AI 与智能体by ggozad
基于 LanceDB 的 Agentic Retrieval Augmented Generation(RAG)方案,支持更主动的检索与生成。
想把传统 RAG 做得更聪明,haiku-rag 借助 LanceDB 实现更主动的检索与生成,让知识问答更准也更灵活。
什么是 io.github.ggozad/haiku-rag?
基于 LanceDB 的 Agentic Retrieval Augmented Generation(RAG)方案,支持更主动的检索与生成。
README
haiku.rag
Agentic RAG that answers questions about your own documents with citations to page numbers and section headings. It runs on an embedded LanceDB database with open models through Ollama by default, so no server or API key is needed. Any provider Pydantic AI supports works in their place, and the same database can live on S3, GCS, Azure or LanceDB Cloud.
Built on LanceDB, Pydantic AI and Docling. Documentation: ggozad.github.io/haiku.rag.
Features
- Ingest PDFs, office documents, HTML, Markdown and images with Docling, in-process or on docling-serve. The stored DoclingDocument keeps headings, tables, pictures and page provenance.
- Search with hybrid vector and full-text retrieval, optional reranking (cross-encoders, Jina, Cohere, Zero Entropy, vLLM, OpenRouter), section-aware context expansion, and image search with a multimodal embedder (vLLM, OpenRouter, VoyageAI, Cohere). Across several named databases at once.
- Answer with the RAG capability: it searches, runs sandboxed Python over the documents for counting and aggregation, and cites page numbers and headings. Vision models receive the figures. Optional capabilities compact earlier evidence in long conversations and require every answer to declare its grounding.
- Check a citation by drawing its chunk on the page image, from the CLI, the chat TUI or Python.
- Integrate through the Python API, native Pydantic AI capabilities, an MCP server for Claude Code, Codex and Claude Desktop, and a reference web app.
- Operate with the
haiku-ingesterservice (filesystem, HTTP, S3 and WebDAV sources, a SQLite or Postgres job queue with retries, a control plane and dashboard), tags and rollback, vacuum, andhaiku-rag doctorhealth checks.
Installation
Python 3.12 or newer.
pip install haiku.rag # Docling, the VoyageAI and Cohere embedders, every reranker, the TUI
pip install haiku.rag-slim # the core, with extras chosen by you
The ingester, S3 access and model providers other than Ollama and OpenAI-compatible endpoints are extras. See Installation.
Quick start
The default configuration uses Ollama for embeddings and answers. The quickstart covers the models to pull and using OpenAI instead.
haiku-rag init # create the database
haiku-rag add-src paper.pdf # index a file, URL or directory
haiku-rag search "attention mechanism"
haiku-rag ask "What datasets were used for evaluation?"
haiku-rag ask "How many documents mention transformers?"
haiku-rag ask "Does this figure match the spec?" --image figure.png
haiku-rag chat # multi-turn chat in the terminal
Continuous ingestion from configured sources runs as a separate service, with the ingester extra (pip install 'haiku.rag[ingester]'):
haiku-ingester serve
Python API
from haiku.rag.client import HaikuRAG
async with HaikuRAG("knowledge.lancedb", create=True) as rag:
await rag.create_document_from_source("paper.pdf")
await rag.create_document_from_source("https://arxiv.org/pdf/1706.03762")
results = await rag.search("self-attention")
for result in results:
print(f"{result.score:.2f} | p.{result.page_numbers} | {result.content[:100]}")
answer, citations = await rag.ask("What is the complexity of self-attention?")
print(answer)
for cite in citations:
print(f" [{cite.chunk_id}] p.{cite.page_numbers}: {cite.content[:80]}")
To compose your own agent, see Capabilities.
MCP server
haiku-rag mcp --stdio
The server gives an assistant search, document reading and a Python sandbox over the documents. In Claude Code, the plugin registers it with a skill:
claude plugin marketplace add ggozad/haiku.rag
claude plugin install haiku-rag
Codex and Claude Desktop setup is in the MCP docs.
Examples
- Docker setup: docling-serve, the ingester and the MCP server
- Web application: conversational RAG over AG-UI with a CopilotKit frontend
Documentation
- Quickstart: install, index, chat
- Installation: packages and extras
- Architecture: how a document becomes a cited answer
- CLI and Chat and inspector
- Capabilities: native Pydantic AI capabilities
- Configuration: every setting, and tuning
- Ingester, MCP and remote processing
- Python API and custom pipelines
- Benchmarks and the changelog
License
MIT.
<!-- mcp-name is used by the MCP registry to identify this server -->mcp-name: io.github.ggozad/haiku-rag
常见问题
io.github.ggozad/haiku-rag 是什么?
基于 LanceDB 的 Agentic Retrieval Augmented Generation(RAG)方案,支持更主动的检索与生成。
相关 Skills
Claude接口
by anthropics
面向接入 Claude API、Anthropic SDK 或 Agent SDK 的开发场景,自动识别项目语言并给出对应示例与默认配置,快速搭建 LLM 应用。
✎ 想把Claude能力接进应用或智能体,用claude-api上手快、兼容Anthropic与Agent SDK,集成路径清晰又省心
多智能体架构
by alirezarezvani
聚焦多智能体系统架构设计,梳理 Supervisor、Swarm、分层和 Pipeline 等模式,覆盖角色定义、通信协作与性能评估,适合规划稳健可扩展的 AI agent 编排方案。
✎ 帮你系统解决多智能体应用的架构设计与协同编排难题,适合构建复杂 AI 工作流,成熟度高、社区认可也很亮眼。
RAG架构师
by alirezarezvani
聚焦生产级RAG系统设计与优化,覆盖文档切块、检索链路、索引构建、召回评估等关键环节,适合搭建可扩展、高准确率的知识库问答与检索增强应用。
✎ 面向RAG落地,把知识库、向量检索和生成链路系统串联起来,做架构设计时更清晰,也更少踩坑。
相关 MCP Server
知识图谱记忆
编辑精选by Anthropic
Memory 是一个基于本地知识图谱的持久化记忆系统,让 AI 记住长期上下文。
✎ 帮 AI 和智能体补上“记不住”的短板,用本地知识图谱沉淀长期上下文,连续对话更聪明,数据也更可控。
顺序思维
编辑精选by Anthropic
Sequential Thinking 是让 AI 通过动态思维链解决复杂问题的参考服务器。
✎ 这个服务器展示了如何让 Claude 像人类一样逐步推理,适合开发者学习 MCP 的思维链实现。但注意它只是个参考示例,别指望直接用在生产环境里。
by deusdata
持久化的代码库知识图谱,可跨会话保留上下文,在 session 重启或上下文压缩后仍能继续使用。
✎ 专治 AI 编程助手“会话失忆”,把代码库沉淀为持久知识图谱,重启或压缩上下文后也能无缝续上开发状态。