io.github.IgorGanapolsky/rlhf-feedback-loop
编码与调试by igorganapolsky
为AI agents提供RLHF feedback loop,捕获反馈信号、提升记忆、阻止错误,并导出DPO数据。
什么是 io.github.IgorGanapolsky/rlhf-feedback-loop?
为AI agents提供RLHF feedback loop,捕获反馈信号、提升记忆、阻止错误,并导出DPO数据。
README
ThumbGate 👍 👎
<p align="center"> <a href="https://thumbgate.ai"> <img src="docs/media/thumbgate-hero-banner.svg" alt="ThumbGate Infrastructure Firewall with Thumbs Up and Thumbs Down" width="100%" /> </a> </p> <p align="center"> <b>ThumbGate is the self-improving pre-action firewall for AI coding agents</b><br> AI coding agents repeat mistakes — and one wrong tool call can wipe a directory, leak a key, or push broken code. </p> <p align="center"> <a href="https://mcptoplist.com/server/glama%2FIgorGanapolsky%2FThumbGate"><img src="https://mcptoplist.com/badge/glama%2FIgorGanapolsky%2FThumbGate.svg" alt="MCP Toplist" /></a> <a href="https://github.com/IgorGanapolsky/ThumbGate/actions/workflows/ci.yml"><img src="https://github.com/IgorGanapolsky/ThumbGate/actions/workflows/ci.yml/badge.svg" alt="CI" /></a> <a href="https://www.npmjs.com/package/thumbgate"><img src="https://img.shields.io/npm/v/thumbgate" alt="npm version" /></a> <a href="https://www.npmjs.com/package/thumbgate"><img src="https://img.shields.io/npm/dw/thumbgate" alt="npm weekly downloads" /></a> <a href="https://github.com/IgorGanapolsky/ThumbGate"><img src="https://img.shields.io/github/stars/IgorGanapolsky/ThumbGate" alt="GitHub stars" /></a> <a href="https://github.com/marketplace/actions/thumbgate-agent-governance"><img src="https://img.shields.io/badge/GitHub_Marketplace-ThumbGate_Agent_Governance-0969da" alt="GitHub Marketplace: ThumbGate Agent Governance" /></a> <a href="LICENSE"><img src="https://img.shields.io/badge/License-MIT-green.svg" alt="License: MIT" /></a> </p> <p align="center"> <a href="#quick-start"><img src="https://img.shields.io/badge/⚡_Quick_Start-npx_thumbgate_init-22d3ee?style=for-the-badge" alt="Quick Start" /></a> <a href="https://thumbgate.ai/#demo?utm_source=github&utm_medium=readme"><img src="https://img.shields.io/badge/🎬_Watch-90s_Demo-ff647c?style=for-the-badge" alt="Watch Demo" /></a> <a href="https://thumbgate.ai/go/gpt?utm_source=github&utm_medium=readme"><img src="https://img.shields.io/badge/💬_Try-ThumbGate_GPT-56e39f?style=for-the-badge" alt="Try GPT" /></a> <a href="https://thumbgate.ai/checkout/pro?utm_source=github&utm_medium=readme"><img src="https://img.shields.io/badge/💼_Pro-$19/mo-ffd166?style=for-the-badge" alt="Pro Tier" /></a> </p>What it does
ThumbGate is the local-first Pre-Action Checks engine for AI coding agents. It runs in the PreToolUse hook to evaluate the proposed tool call before execution — so costly mistakes can be caught before they happen.
ThumbGate GitHub star growth is measured with GitHub's privacy-safe GET /repos/{owner}/{repo}/stargazers/history endpoint (weekly counts, no stargazer identities). Run npm run stars:history -- --fixture tests/fixtures/github-star-history.json --json for the local proof. Stars are not npm installs and not revenue. The live GitHub Marketplace Action is ThumbGate Agent Governance (uses: IgorGanapolsky/ThumbGate@v1).
Usage over star count. Evaluate ThumbGate from the install path and live usage badges above (npx thumbgate init, Marketplace uses:, npm weekly downloads, GitHub clones), not from whether the repo has twenty stars or twenty thousand. No pitch deck is required. We do not farm GitHub profile badges (no YOLO-merge of protected main, no 5-minute Issue close theater, no fake Co-authored-by). Galaxy Brain needs real accepted answers in Discussions Q&A. npm run github:achievements -- --fixture tests/fixtures/github-achievements.json --json inventories what is already earned vs what we refuse to farm.
Who it's for
ThumbGate is for operators whose AI coding agents can leak a secret or destroy a checkout before a human sees the tool call (Claude Code, Cursor, Codex, Gemini CLI, MCP). Discovery should reach those operators — not a star campaign.
ThumbGate is not a GitHub star package, not fake engagement, and not a substitute for npm installs or merged PRs. Real engagement is npx thumbgate init and a PreToolUse hook that actually fires.
Tech memes (shareable)
Lightweight visuals for how agents fail without a pre-action gate:
| Meme | Meaning |
|---|---|
| Unchecked tool calls ship destructive commands. | |
| A prompt is advice; a PreToolUse hook is enforcement. |
It hard-blocks detected secret leaks and two direct self-disable command classes by default — commands that terminate the ThumbGate gate process or enable its bypass environment override. Other high-risk classes (rm -rf, force-push, fetch-and-run, direct guardrail edits) warn and log by default. Set THUMBGATE_STRICT_ENFORCEMENT=1 for strict enforcement (warnings become hard denies).
| Verdict | Default behavior |
|---|---|
| ⛔ Hard-block | Detected secret leaks; process-kill/environment-override self-disable |
| 👎 Warn + log | rm -rf, git push --force, fetch-and-run, direct guardrail edits — warn by default |
| 👍 Allow | Everything else |
Accepted feedback is stored as local lessons. Repeated concrete failures can become prevention rules that promote from warnings to blocking gates. The firewall improves from operations without retraining the model. Prompt evaluation (npx thumbgate eval) turns accepted feedback into reusable eval cases and local proof reports.
Honest disclaimer: ThumbGate does not update model weights. It intercepts tool calls at runtime. Local-first — no cloud required for the enforcement path.
Works with Claude Code, Cursor, Codex, Gemini CLI, Amp, Cline, OpenCode, and other MCP agents.
Agent tries: rm -rf tests/
ThumbGate: 👎 WARN + LOG — "Never delete test directories"
Pattern matched: rm.*-rf.*tests
Source: your thumbs-down from last Tuesday
Strict mode: ⛔ DENY before tool execution
Agentic development cycle fit
Agentic development is becoming a loop: Guide → Generate → Verify → Solve. ThumbGate is the pre-action gate / pre-action boundary between generated intent and executed action.
Quick Start
Want a phased walkthrough with a verify step at every stage? Follow the Progressive Setup Guide.
Progressive wiring — prove the pipe before you turn matching on. Empty dashboard is success.
npx thumbgate init # Phase 1: hooks only
npx thumbgate doctor # verify: exits 0 only when PreToolUse hook is wired (hidden metric = hook install, not gate count)
npx thumbgate dashboard --open # Phase 2: open local HTML; empty stats are OK
npx thumbgate capture --feedback=down --context="Never run DROP on production tables" --what-went-wrong="agent proposed DROP" --what-to-change="require review for DROP"
Later DROP attempts in the same scope surface the check:
⚠️ Check fired: "Never run DROP on production tables"
Pattern: DROP.*production
Verdict: 👎 WARN + LOG (⛔ BLOCK when THUMBGATE_STRICT_ENFORCEMENT=1)
Numbered configs: config/progressive/. Guide: progressive wiring.
MCP / Glama / registry install (stdio)
Directories and clients that install ThumbGate as an MCP server must start stdio MCP, not the HTTP API:
npx -y thumbgate serve
- Equivalent:
npx -y thumbgate mcp - Do not use
npm startfor MCP — that launches the hosted HTTP API (src/api/server.js), not the agent-facing stdio server.
▶ 90-second demo · GIF walkthrough
Install for your agent
| Agent | Command | Enforcement |
|---|---|---|
| Claude Code | npx thumbgate init --agent claude-code | 🛡️ Hard — PreToolUse |
| Codex | npx thumbgate init --agent codex | 🛡️ Hard — pre_tool_use |
| Gemini CLI | npx thumbgate init --agent gemini | 🛡️ Hard — PreToolUse |
| ForgeCode | npx thumbgate init --agent forge | 🛡️ Hard — pre_tool_use |
| Cursor | npx thumbgate init --agent cursor | 💬 Advisory — MCP gate_check |
| Cline | npx thumbgate init --agent cline | 💬 Advisory — MCP + .clinerules |
| OpenCode | npx thumbgate init --agent opencode | 💬 Advisory — MCP gate_check |
| Any MCP agent | npx thumbgate serve | 💬 Advisory — MCP gate_check |
| Amp | npx thumbgate init --agent amp | 📝 Feedback capture |
| GitHub Actions | uses: IgorGanapolsky/ThumbGate@v1 | 🩺 Marketplace Action — doctor / AI inventory in CI |
Per-agent guides: Claude/Codex bridge · Codex profile · Cursor · MCP setup
Install scope: machine-wide vs per-project
| Scope | Command | Settings | Lessons | Best for |
|---|---|---|---|---|
| Machine-wide (default) | npx thumbgate init | ~/.claude/settings.json | ~/.claude/memory/feedback/ | Solo operators — same machine-local feedback store across repos |
| Per-project | npx thumbgate init --project | <repo>/.claude/settings.json | <repo>/.claude/memory/feedback/ | Client / compliance — separate dashboard / isolated lessons per repo |
Both scopes write mcpServers.thumbgate plus PreToolUse / UserPromptSubmit / PostToolUse / SessionStart hooks. Machine-wide is the right default for most developers. Cross-repo blocking is not automatic: a lesson learned in one project only applies elsewhere when you share the store (machine-wide) or export/import lessons.
MCP tools (surface): gate_check (read/evaluate proposed tool call), feedback capture + session tools (write), dashboard/stats (read). Destructive agent actions stay blocked/warned by PreToolUse — ThumbGate does not execute user shell commands for you.
Discoverable slash-commands — the guardrail layer for spec-driven agents
Spec-driven agent frameworks like GSD (get-shit-done) and GitHub Spec Kit plan and generate work. ThumbGate is the guardrail layer for spec-driven agents: it sits after the plan, on the boundary between a generated tool call and its execution — alongside GSD / Spec-Kit, not instead of them.
npx thumbgate init installs these into your agent palette:
| Command | What it does |
|---|---|
/thumbgate-dashboard | Open local project dashboard |
/thumbgate-guard | Turn last mistake into a hard prevention rule |
/thumbgate-rules | List active rules & lessons |
/thumbgate-blocked | Gate stats + enforcement matrix |
/thumbgate-protect | Branch governance + scoped approval |
/thumbgate-doctor | Health-check hooks, MCP, readiness |
Pricing & buyer paths
Free tier: 2 feedback captures/day (10 total) and up to 3 active auto-promoted prevention rules. Pro ($19/mo or $149/yr) is the individual tier for unlimited rules, history-aware lessons, linked feedback session flow, personal dashboard, and DPO export. Enterprise is custom and scoped after intake; hosted team lesson sync and a hosted org dashboard are not general availability.
| Free | Pro ($19/mo or $149/yr) | Enterprise | |
|---|---|---|---|
| Local CLI + PreToolUse | ✅ | ✅ | Scoped after intake |
| Feedback captures | 2 feedback captures/day (10 total) | Unlimited | Scoped after intake |
| Active auto-promoted rules | up to 3 active auto-promoted prevention rules | Unlimited | Scoped after intake |
| Personal dashboard + DPO export | — | ✅ | Reviewed during intake |
| Hosted team lesson sync | — | — | Not general availability |
| Hosted org dashboard | — | — | Not general availability |
Enterprise intake path: the Workflow Hardening Sprint scopes one repeated failure before any broader rollout commitment. Start intake →
Local technical path: install the CLI and use init plus the documented setup so Pre-Action Checks evaluate tool calls where the agent actually runs.
First-dollar activation path: open the ThumbGate GPT, paste the risky action, capture typed feedback (thumbs down: / thumbs up:). Native ChatGPT rating buttons are not the ThumbGate capture path. Ask: what repeated AI mistake would be worth catching before the tool executes?
Paid path for individual operators: ThumbGate Pro is the self-serve side lane for a personal dashboard and export-ready evidence.
Start free · Pro $19/mo · Live Dashboard · Team Sprint intake · Workflow Hardening Sprint · First Dollar Playbook
Popular buyer questions: AI search topical presence · Relational knowledge and AI recommendations · AI Mode ads for agent governance · MCP tool governance · AI agent pre-action approval gates · Background agent governance · GPT-5.5 model evaluation · Stop repeated AI agent mistakes · Browser automation safety · Native messaging host security · Autoresearch agent safety · Cursor guardrails · Codex CLI guardrails · Gemini CLI memory + enforcement · Google Cloud MCP guardrails · Roo Code alternative: migrate to Cline
How it works (short)
- Capture 👍/👎 feedback (CLI, MCP, linked feedback session flow /
open_feedback_session, or ThumbGate GPT) - Promote concrete lessons via history-aware lesson distillation into prevention rules
- Evaluate the next proposed tool call against active rules (literal/AST + local vectors)
- Allow / warn / deny before the tool runs
npx thumbgate brain --write # → .thumbgate/BRAIN.md (lessons + gates in one artifact)
Pro operators can invoke search_lessons through MCP and use npx thumbgate lessons from the CLI. History-aware feedback sessions and lesson search are Pro capabilities; Free does not include recall or search.
flowchart LR
A["Agent tool call"] --> B{"Rule match?"}
B -- exact --> D["On-device gate"]
B -- semantic --> C["Local LanceDB"]
C --> D
D -- secret/kill --> E["⛔ Hard-block"]
D -- known-bad --> G["👎 Warn + log"]
D -- safe --> F["👍 Allow"]
⛔ secret-exfiltration → hard-block (default)
⛔ self-protect-kill → hard-block (default)
⛔ self-protect-env → hard-block (default)
⚠️ force-push → warn; hard-block under strict
⚠️ protected-branch → warn; hard-block under strict
⚠️ unresolved-threads → warn; hard-block under strict
⚠️ package-lock-reset → warn; hard-block under strict
npx thumbgate init
npx thumbgate doctor
npx thumbgate capture up|down "<text>"
npx thumbgate lessons
npx thumbgate brain --write
npx thumbgate dashboard --open
npx thumbgate break-glass --reason="ThumbGate over-fired" # 5-min recovery
# Portable lessons
curl -X POST http://localhost:3456/v1/lessons/export \
-H "Authorization: Bearer $THUMBGATE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"outputPath": "./lessons-export.json"}'
# DPO pairs for fine-tuning
curl -X POST http://localhost:3456/v1/dpo/export \
-H "Authorization: Bearer $THUMBGATE_API_KEY" \
-o dpo-pairs.jsonl
Tech Stack
| Layer | Tech |
|---|---|
| Runtime | Node.js ≥18 |
| Interfaces | MCP stdio, HTTP API, CLI |
| Storage | SQLite + FTS5, LanceDB vectors, JSONL logs |
| Intelligence | MemAlign dual recall, Thompson Sampling, local embeddings |
| Billing / host | Stripe, Railway |
| Execution | Railway, Cloudflare Workers, Docker Sandboxes |
| Governance | Workflow Sentinel, control plane, Docker Sandboxes |
Every Changeset is tied to the exact main merge commit and generates Verification Evidence for Release Confidence.
Integrations (compact)
| Surface | Start here |
|---|---|
| Open ThumbGate GPT | thumbgate.ai/go/gpt — ThumbGate GPT: start here. Paste agent actions, get advice + checkpointing. No, users do not have to keep chatting inside the ThumbGate GPT to use ThumbGate — the hard enforcement layer still runs where the work happens. |
| Install Codex Plugin | Open the Codex plugin install page: thumbgate.ai/codex-plugin · zip: thumbgate-codex-plugin.zip · plugins/codex-profile/INSTALL.md |
Claude Desktop .mcpb | latest release |
| VS Code / Open VSX | plugins/vscode-extension/README.md |
| Antigravity-compatible | plugins/antigravity-extension/INSTALL.md |
| JetBrains | plugins/jetbrains-plugin/README.md · JetBrains Marketplace path for the same runtime |
| ChatGPT App / GPT Action | thumbgate.ai/chatgpt-app |
| ThumbGate-Core (staging) | https://github.com/IgorGanapolsky/ThumbGate-Core — pre-release staging + a few internal cache scripts; not the product moat |
Docs
Full index: docs/INDEX.md
| Need | Link |
|---|---|
| Agent workflow contract | WORKFLOW.md |
| Ready-for-agent intake | .github/ISSUE_TEMPLATE/ready-for-agent.yml |
| Verification Evidence | docs/VERIFICATION_EVIDENCE.md |
| Release Confidence | docs/RELEASE_CONFIDENCE.md |
| Changeset strategy | docs/CHANGESET_STRATEGY.md |
| First Dollar Playbook | docs/FIRST_DOLLAR_PLAYBOOK.md |
| Security policy | SECURITY.md |
| Threat model | THREAT_MODEL.md |
| Federal / regulated | docs/FEDERAL.md |
| Commercial Truth | docs/COMMERCIAL_TRUTH.md |
| Issues / PRs | GitHub Issues · PR template |
FAQ (one-liners): Not a fine-tuner (runtime intercept only). Different from CLAUDE.md / .cursorrules (those are context; ThumbGate is an external allow/warn/deny before tools run).
Who builds this
Igor Ganapolsky — payments (Stripe/Connect), AI agent guardrails/MCP, Android + backends. Small number of contract slots: $120–150/hr, 1099, remote US. LinkedIn · thumbgate.ai
License
MIT — see LICENSE. Project policy: SECURITY.md · THREAT_MODEL.md.
常见问题
io.github.IgorGanapolsky/rlhf-feedback-loop 是什么?
为AI agents提供RLHF feedback loop,捕获反馈信号、提升记忆、阻止错误,并导出DPO数据。
相关 Skills
网页构建器
by anthropics
面向复杂 claude.ai HTML artifact 开发,快速初始化 React + Tailwind CSS + shadcn/ui 项目并打包为单文件 HTML,适合需要状态管理、路由或多组件交互的页面。
✎ 在 claude.ai 里做复杂网页 Artifact 很省心,多组件、状态和路由都能顺手搭起来,React、Tailwind 与 shadcn/ui 组合效率高、成品也更精致。
网页应用测试
by anthropics
用 Playwright 为本地 Web 应用编写自动化测试,支持启动开发服务器、校验前端交互、排查 UI 异常、抓取截图与浏览器日志,适合调试动态页面和回归验证。
✎ 借助 Playwright 一站式验证本地 Web 应用前端功能,调 UI 时还能同步查看日志和截图,定位问题更快。
前端设计
by anthropics
面向组件、页面、海报和 Web 应用开发,按鲜明视觉方向生成可直接落地的前端代码与高质感 UI,适合做 landing page、Dashboard 或美化现有界面,避开千篇一律的 AI 审美。
✎ 想把页面做得既能上线又有设计感,就用前端设计:组件到整站都能产出,难得的是能避开千篇一律的 AI 味。
相关 MCP Server
GitHub
编辑精选by GitHub
GitHub 是 MCP 官方参考服务器,让 Claude 直接读写你的代码仓库和 Issues。
✎ 这个参考服务器解决了开发者想让 AI 安全访问 GitHub 数据的问题,适合需要自动化代码审查或 Issue 管理的团队。但注意它只是参考实现,生产环境得自己加固安全。
Context7 文档查询
编辑精选by Context7
Context7 是实时拉取最新文档和代码示例的智能助手,让你告别过时资料。
✎ 它能解决开发者查找文档时信息滞后的问题,特别适合快速上手新库或跟进更新。不过,依赖外部源可能导致偶尔的数据延迟,建议结合官方文档使用。
by tldraw
tldraw 是让 AI 助手直接在无限画布上绘图和协作的 MCP 服务器。
✎ 这解决了 AI 只能输出文本、无法视觉化协作的痛点——想象让 Claude 帮你画流程图或白板讨论。最适合需要快速原型设计或头脑风暴的开发者。不过,目前它只是个基础连接器,你得自己搭建画布应用才能发挥全部潜力。