io.github.ggozad/haiku-rag

AI 与智能体

by ggozad

基于 LanceDB 的 Agentic Retrieval Augmented Generation(RAG)方案,支持更主动的检索与生成。

想把传统 RAG 做得更聪明,haiku-rag 借助 LanceDB 实现更主动的检索与生成,让知识问答更准也更灵活。

什么是 io.github.ggozad/haiku-rag?

基于 LanceDB 的 Agentic Retrieval Augmented Generation(RAG)方案,支持更主动的检索与生成。

README

haiku.rag

PyPI Python Downloads Docs Tests codecov

Agentic RAG that answers questions about your own documents with citations to page numbers and section headings. It runs on an embedded LanceDB database with open models through Ollama by default, so no server or API key is needed. Any provider Pydantic AI supports works in their place, and the same database can live on S3, GCS, Azure or LanceDB Cloud.

Built on LanceDB, Pydantic AI and Docling. Documentation: ggozad.github.io/haiku.rag.

Features

  • Ingest PDFs, office documents, HTML, Markdown and images with Docling, in-process or on docling-serve. The stored DoclingDocument keeps headings, tables, pictures and page provenance.
  • Search with hybrid vector and full-text retrieval, optional reranking (cross-encoders, Jina, Cohere, Zero Entropy, vLLM, OpenRouter), section-aware context expansion, and image search with a multimodal embedder (vLLM, OpenRouter, VoyageAI, Cohere). Across several named databases at once.
  • Answer with the RAG capability: it searches, runs sandboxed Python over the documents for counting and aggregation, and cites page numbers and headings. Vision models receive the figures. Optional capabilities compact earlier evidence in long conversations and require every answer to declare its grounding.
  • Check a citation by drawing its chunk on the page image, from the CLI, the chat TUI or Python.
  • Integrate through the Python API, native Pydantic AI capabilities, an MCP server for Claude Code, Codex and Claude Desktop, and a reference web app.
  • Operate with the haiku-ingester service (filesystem, HTTP, S3 and WebDAV sources, a SQLite or Postgres job queue with retries, a control plane and dashboard), tags and rollback, vacuum, and haiku-rag doctor health checks.

Installation

Python 3.12 or newer.

bash
pip install haiku.rag        # Docling, the VoyageAI and Cohere embedders, every reranker, the TUI
pip install haiku.rag-slim   # the core, with extras chosen by you

The ingester, S3 access and model providers other than Ollama and OpenAI-compatible endpoints are extras. See Installation.

Quick start

The default configuration uses Ollama for embeddings and answers. The quickstart covers the models to pull and using OpenAI instead.

bash
haiku-rag init                                  # create the database
haiku-rag add-src paper.pdf                     # index a file, URL or directory
haiku-rag search "attention mechanism"
haiku-rag ask "What datasets were used for evaluation?"
haiku-rag ask "How many documents mention transformers?"
haiku-rag ask "Does this figure match the spec?" --image figure.png
haiku-rag chat                                  # multi-turn chat in the terminal

Continuous ingestion from configured sources runs as a separate service, with the ingester extra (pip install 'haiku.rag[ingester]'):

bash
haiku-ingester serve

Python API

python
from haiku.rag.client import HaikuRAG

async with HaikuRAG("knowledge.lancedb", create=True) as rag:
    await rag.create_document_from_source("paper.pdf")
    await rag.create_document_from_source("https://arxiv.org/pdf/1706.03762")

    results = await rag.search("self-attention")
    for result in results:
        print(f"{result.score:.2f} | p.{result.page_numbers} | {result.content[:100]}")

    answer, citations = await rag.ask("What is the complexity of self-attention?")
    print(answer)
    for cite in citations:
        print(f"  [{cite.chunk_id}] p.{cite.page_numbers}: {cite.content[:80]}")

To compose your own agent, see Capabilities.

MCP server

bash
haiku-rag mcp --stdio

The server gives an assistant search, document reading and a Python sandbox over the documents. In Claude Code, the plugin registers it with a skill:

bash
claude plugin marketplace add ggozad/haiku.rag
claude plugin install haiku-rag

Codex and Claude Desktop setup is in the MCP docs.

Examples

  • Docker setup: docling-serve, the ingester and the MCP server
  • Web application: conversational RAG over AG-UI with a CopilotKit frontend

Documentation

License

MIT.

<!-- mcp-name is used by the MCP registry to identify this server -->

mcp-name: io.github.ggozad/haiku-rag

常见问题

io.github.ggozad/haiku-rag 是什么?

基于 LanceDB 的 Agentic Retrieval Augmented Generation(RAG)方案,支持更主动的检索与生成。

相关 Skills

Claude接口

by anthropics

Universal
热门

面向接入 Claude API、Anthropic SDK 或 Agent SDK 的开发场景,自动识别项目语言并给出对应示例与默认配置,快速搭建 LLM 应用。

✎ 想把Claude能力接进应用或智能体,用claude-api上手快、兼容Anthropic与Agent SDK,集成路径清晰又省心

AI 与智能体
未扫描176.4k

多智能体架构

by alirezarezvani

Universal
热门

聚焦多智能体系统架构设计,梳理 Supervisor、Swarm、分层和 Pipeline 等模式,覆盖角色定义、通信协作与性能评估,适合规划稳健可扩展的 AI agent 编排方案。

✎ 帮你系统解决多智能体应用的架构设计与协同编排难题,适合构建复杂 AI 工作流,成熟度高、社区认可也很亮眼。

AI 与智能体
未扫描26.0k

RAG架构师

by alirezarezvani

Universal
热门

聚焦生产级RAG系统设计与优化,覆盖文档切块、检索链路、索引构建、召回评估等关键环节,适合搭建可扩展、高准确率的知识库问答与检索增强应用。

✎ 面向RAG落地,把知识库、向量检索和生成链路系统串联起来,做架构设计时更清晰,也更少踩坑。

AI 与智能体
未扫描26.0k

相关 MCP Server

知识图谱记忆

编辑精选

by Anthropic

热门

Memory 是一个基于本地知识图谱的持久化记忆系统,让 AI 记住长期上下文。

✎ 帮 AI 和智能体补上“记不住”的短板,用本地知识图谱沉淀长期上下文,连续对话更聪明,数据也更可控。

AI 与智能体
89.7k

顺序思维

编辑精选

by Anthropic

热门

Sequential Thinking 是让 AI 通过动态思维链解决复杂问题的参考服务器。

✎ 这个服务器展示了如何让 Claude 像人类一样逐步推理,适合开发者学习 MCP 的思维链实现。但注意它只是个参考示例,别指望直接用在生产环境里。

AI 与智能体
89.2k

by deusdata

热门

持久化的代码库知识图谱,可跨会话保留上下文,在 session 重启或上下文压缩后仍能继续使用。

✎ 专治 AI 编程助手“会话失忆”,把代码库沉淀为持久知识图谱,重启或压缩上下文后也能无缝续上开发状态。

AI 与智能体
37.3k

评论