io.github.IgorGanapolsky/rlhf-feedback-loop

编码与调试

by igorganapolsky

为AI agents提供RLHF feedback loop,捕获反馈信号、提升记忆、阻止错误,并导出DPO数据。

什么是 io.github.IgorGanapolsky/rlhf-feedback-loop

为AI agents提供RLHF feedback loop,捕获反馈信号、提升记忆、阻止错误,并导出DPO数据。

README

ThumbGate 👍 👎

<p align="center"> <a href="https://thumbgate.ai"> <img src="docs/media/thumbgate-hero-banner.svg" alt="ThumbGate Infrastructure Firewall with Thumbs Up and Thumbs Down" width="100%" /> </a> </p> <p align="center"> <b>ThumbGate is the self-improving pre-action firewall for AI coding agents</b><br> AI coding agents repeat mistakes — and one wrong tool call can wipe a directory, leak a key, or push broken code. </p> <p align="center"> <a href="https://mcptoplist.com/server/glama%2FIgorGanapolsky%2FThumbGate"><img src="https://mcptoplist.com/badge/glama%2FIgorGanapolsky%2FThumbGate.svg" alt="MCP Toplist" /></a> <a href="https://github.com/IgorGanapolsky/ThumbGate/actions/workflows/ci.yml"><img src="https://github.com/IgorGanapolsky/ThumbGate/actions/workflows/ci.yml/badge.svg" alt="CI" /></a> <a href="https://www.npmjs.com/package/thumbgate"><img src="https://img.shields.io/npm/v/thumbgate" alt="npm version" /></a> <a href="https://www.npmjs.com/package/thumbgate"><img src="https://img.shields.io/npm/dw/thumbgate" alt="npm weekly downloads" /></a> <a href="https://github.com/IgorGanapolsky/ThumbGate"><img src="https://img.shields.io/github/stars/IgorGanapolsky/ThumbGate" alt="GitHub stars" /></a> <a href="https://github.com/marketplace/actions/thumbgate-agent-governance"><img src="https://img.shields.io/badge/GitHub_Marketplace-ThumbGate_Agent_Governance-0969da" alt="GitHub Marketplace: ThumbGate Agent Governance" /></a> <a href="LICENSE"><img src="https://img.shields.io/badge/License-MIT-green.svg" alt="License: MIT" /></a> </p> <p align="center"> <a href="#quick-start"><img src="https://img.shields.io/badge/⚡_Quick_Start-npx_thumbgate_init-22d3ee?style=for-the-badge" alt="Quick Start" /></a> <a href="https://thumbgate.ai/#demo?utm_source=github&utm_medium=readme"><img src="https://img.shields.io/badge/🎬_Watch-90s_Demo-ff647c?style=for-the-badge" alt="Watch Demo" /></a> <a href="https://thumbgate.ai/go/gpt?utm_source=github&utm_medium=readme"><img src="https://img.shields.io/badge/💬_Try-ThumbGate_GPT-56e39f?style=for-the-badge" alt="Try GPT" /></a> <a href="https://thumbgate.ai/checkout/pro?utm_source=github&utm_medium=readme"><img src="https://img.shields.io/badge/💼_Pro-$19/mo-ffd166?style=for-the-badge" alt="Pro Tier" /></a> </p>

What it does

ThumbGate is the local-first Pre-Action Checks engine for AI coding agents. It runs in the PreToolUse hook to evaluate the proposed tool call before execution — so costly mistakes can be caught before they happen.

ThumbGate GitHub star growth is measured with GitHub's privacy-safe GET /repos/{owner}/{repo}/stargazers/history endpoint (weekly counts, no stargazer identities). Run npm run stars:history -- --fixture tests/fixtures/github-star-history.json --json for the local proof. Stars are not npm installs and not revenue. The live GitHub Marketplace Action is ThumbGate Agent Governance (uses: IgorGanapolsky/ThumbGate@v1).

Usage over star count. Evaluate ThumbGate from the install path and live usage badges above (npx thumbgate init, Marketplace uses:, npm weekly downloads, GitHub clones), not from whether the repo has twenty stars or twenty thousand. No pitch deck is required. We do not farm GitHub profile badges (no YOLO-merge of protected main, no 5-minute Issue close theater, no fake Co-authored-by). Galaxy Brain needs real accepted answers in Discussions Q&A. npm run github:achievements -- --fixture tests/fixtures/github-achievements.json --json inventories what is already earned vs what we refuse to farm.

Who it's for

ThumbGate is for operators whose AI coding agents can leak a secret or destroy a checkout before a human sees the tool call (Claude Code, Cursor, Codex, Gemini CLI, MCP). Discovery should reach those operators — not a star campaign.

ThumbGate is not a GitHub star package, not fake engagement, and not a substitute for npm installs or merged PRs. Real engagement is npx thumbgate init and a PreToolUse hook that actually fires.

Tech memes (shareable)

Lightweight visuals for how agents fail without a pre-action gate:

MemeMeaning
Agent destroys prod without a gateUnchecked tool calls ship destructive commands.
Prompt vs PreToolUse hookA prompt is advice; a PreToolUse hook is enforcement.

It hard-blocks detected secret leaks and two direct self-disable command classes by default — commands that terminate the ThumbGate gate process or enable its bypass environment override. Other high-risk classes (rm -rf, force-push, fetch-and-run, direct guardrail edits) warn and log by default. Set THUMBGATE_STRICT_ENFORCEMENT=1 for strict enforcement (warnings become hard denies).

VerdictDefault behavior
Hard-blockDetected secret leaks; process-kill/environment-override self-disable
👎 Warn + logrm -rf, git push --force, fetch-and-run, direct guardrail edits — warn by default
👍 AllowEverything else

Accepted feedback is stored as local lessons. Repeated concrete failures can become prevention rules that promote from warnings to blocking gates. The firewall improves from operations without retraining the model. Prompt evaluation (npx thumbgate eval) turns accepted feedback into reusable eval cases and local proof reports.

Honest disclaimer: ThumbGate does not update model weights. It intercepts tool calls at runtime. Local-first — no cloud required for the enforcement path.

Works with Claude Code, Cursor, Codex, Gemini CLI, Amp, Cline, OpenCode, and other MCP agents.

AI Agent without ThumbGate vs Agent guarded by ThumbGate

code
  Agent tries:   rm -rf tests/
  ThumbGate:     👎 WARN + LOG — "Never delete test directories"
                 Pattern matched: rm.*-rf.*tests
                 Source: your thumbs-down from last Tuesday
                 Strict mode: ⛔ DENY before tool execution

Agentic development cycle fit

Agentic development is becoming a loop: Guide → Generate → Verify → Solve. ThumbGate is the pre-action gate / pre-action boundary between generated intent and executed action.


Quick Start

Want a phased walkthrough with a verify step at every stage? Follow the Progressive Setup Guide.

Progressive wiring — prove the pipe before you turn matching on. Empty dashboard is success.

bash
npx thumbgate init          # Phase 1: hooks only
npx thumbgate doctor        # verify: exits 0 only when PreToolUse hook is wired (hidden metric = hook install, not gate count)
npx thumbgate dashboard --open  # Phase 2: open local HTML; empty stats are OK
npx thumbgate capture --feedback=down --context="Never run DROP on production tables" --what-went-wrong="agent proposed DROP" --what-to-change="require review for DROP"

Later DROP attempts in the same scope surface the check:

code
⚠️ Check fired: "Never run DROP on production tables"
   Pattern: DROP.*production
   Verdict: 👎 WARN + LOG   (⛔ BLOCK when THUMBGATE_STRICT_ENFORCEMENT=1)

Numbered configs: config/progressive/. Guide: progressive wiring.

MCP / Glama / registry install (stdio)

Directories and clients that install ThumbGate as an MCP server must start stdio MCP, not the HTTP API:

bash
npx -y thumbgate serve
  • Equivalent: npx -y thumbgate mcp
  • Do not use npm start for MCP — that launches the hosted HTTP API (src/api/server.js), not the agent-facing stdio server.

▶ 90-second demo · GIF walkthrough


Install for your agent

AgentCommandEnforcement
Claude Codenpx thumbgate init --agent claude-code🛡️ Hard — PreToolUse
Codexnpx thumbgate init --agent codex🛡️ Hard — pre_tool_use
Gemini CLInpx thumbgate init --agent gemini🛡️ Hard — PreToolUse
ForgeCodenpx thumbgate init --agent forge🛡️ Hard — pre_tool_use
Cursornpx thumbgate init --agent cursor💬 Advisory — MCP gate_check
Clinenpx thumbgate init --agent cline💬 Advisory — MCP + .clinerules
OpenCodenpx thumbgate init --agent opencode💬 Advisory — MCP gate_check
Any MCP agentnpx thumbgate serve💬 Advisory — MCP gate_check
Ampnpx thumbgate init --agent amp📝 Feedback capture
GitHub Actionsuses: IgorGanapolsky/ThumbGate@v1🩺 Marketplace Action — doctor / AI inventory in CI

Per-agent guides: Claude/Codex bridge · Codex profile · Cursor · MCP setup

Install scope: machine-wide vs per-project

ScopeCommandSettingsLessonsBest for
Machine-wide (default)npx thumbgate init~/.claude/settings.json~/.claude/memory/feedback/Solo operators — same machine-local feedback store across repos
Per-projectnpx thumbgate init --project<repo>/.claude/settings.json<repo>/.claude/memory/feedback/Client / compliance — separate dashboard / isolated lessons per repo

Both scopes write mcpServers.thumbgate plus PreToolUse / UserPromptSubmit / PostToolUse / SessionStart hooks. Machine-wide is the right default for most developers. Cross-repo blocking is not automatic: a lesson learned in one project only applies elsewhere when you share the store (machine-wide) or export/import lessons.

MCP tools (surface): gate_check (read/evaluate proposed tool call), feedback capture + session tools (write), dashboard/stats (read). Destructive agent actions stay blocked/warned by PreToolUse — ThumbGate does not execute user shell commands for you.


Discoverable slash-commands — the guardrail layer for spec-driven agents

Spec-driven agent frameworks like GSD (get-shit-done) and GitHub Spec Kit plan and generate work. ThumbGate is the guardrail layer for spec-driven agents: it sits after the plan, on the boundary between a generated tool call and its execution — alongside GSD / Spec-Kit, not instead of them.

npx thumbgate init installs these into your agent palette:

CommandWhat it does
/thumbgate-dashboardOpen local project dashboard
/thumbgate-guardTurn last mistake into a hard prevention rule
/thumbgate-rulesList active rules & lessons
/thumbgate-blockedGate stats + enforcement matrix
/thumbgate-protectBranch governance + scoped approval
/thumbgate-doctorHealth-check hooks, MCP, readiness

Pricing & buyer paths

Free tier: 2 feedback captures/day (10 total) and up to 3 active auto-promoted prevention rules. Pro ($19/mo or $149/yr) is the individual tier for unlimited rules, history-aware lessons, linked feedback session flow, personal dashboard, and DPO export. Enterprise is custom and scoped after intake; hosted team lesson sync and a hosted org dashboard are not general availability.

FreePro ($19/mo or $149/yr)Enterprise
Local CLI + PreToolUseScoped after intake
Feedback captures2 feedback captures/day (10 total)UnlimitedScoped after intake
Active auto-promoted rulesup to 3 active auto-promoted prevention rulesUnlimitedScoped after intake
Personal dashboard + DPO exportReviewed during intake
Hosted team lesson syncNot general availability
Hosted org dashboardNot general availability

Enterprise intake path: the Workflow Hardening Sprint scopes one repeated failure before any broader rollout commitment. Start intake →

Local technical path: install the CLI and use init plus the documented setup so Pre-Action Checks evaluate tool calls where the agent actually runs.

First-dollar activation path: open the ThumbGate GPT, paste the risky action, capture typed feedback (thumbs down: / thumbs up:). Native ChatGPT rating buttons are not the ThumbGate capture path. Ask: what repeated AI mistake would be worth catching before the tool executes?

Paid path for individual operators: ThumbGate Pro is the self-serve side lane for a personal dashboard and export-ready evidence.

Start free · Pro $19/mo · Live Dashboard · Team Sprint intake · Workflow Hardening Sprint · First Dollar Playbook

Popular buyer questions: AI search topical presence · Relational knowledge and AI recommendations · AI Mode ads for agent governance · MCP tool governance · AI agent pre-action approval gates · Background agent governance · GPT-5.5 model evaluation · Stop repeated AI agent mistakes · Browser automation safety · Native messaging host security · Autoresearch agent safety · Cursor guardrails · Codex CLI guardrails · Gemini CLI memory + enforcement · Google Cloud MCP guardrails · Roo Code alternative: migrate to Cline


How it works (short)

  1. Capture 👍/👎 feedback (CLI, MCP, linked feedback session flow / open_feedback_session, or ThumbGate GPT)
  2. Promote concrete lessons via history-aware lesson distillation into prevention rules
  3. Evaluate the next proposed tool call against active rules (literal/AST + local vectors)
  4. Allow / warn / deny before the tool runs
bash
npx thumbgate brain --write   # → .thumbgate/BRAIN.md (lessons + gates in one artifact)

Pro operators can invoke search_lessons through MCP and use npx thumbgate lessons from the CLI. History-aware feedback sessions and lesson search are Pro capabilities; Free does not include recall or search.

<details> <summary><b>Architecture diagram & stack</b></summary>

ThumbGate Architecture

mermaid
flowchart LR
    A["Agent tool call"] --> B{"Rule match?"}
    B -- exact --> D["On-device gate"]
    B -- semantic --> C["Local LanceDB"]
    C --> D
    D -- secret/kill --> E["⛔ Hard-block"]
    D -- known-bad --> G["👎 Warn + log"]
    D -- safe --> F["👍 Allow"]
</details> <details> <summary><b>Built-in checks</b></summary>
code
⛔ secret-exfiltration → hard-block (default)
⛔ self-protect-kill   → hard-block (default)
⛔ self-protect-env    → hard-block (default)
⚠️ force-push          → warn; hard-block under strict
⚠️ protected-branch    → warn; hard-block under strict
⚠️ unresolved-threads  → warn; hard-block under strict
⚠️ package-lock-reset  → warn; hard-block under strict
</details> <details> <summary><b>CLI cheatsheet</b></summary>
bash
npx thumbgate init
npx thumbgate doctor
npx thumbgate capture up|down "<text>"
npx thumbgate lessons
npx thumbgate brain --write
npx thumbgate dashboard --open
npx thumbgate break-glass --reason="ThumbGate over-fired"   # 5-min recovery
</details> <details> <summary><b>Pro: lesson + DPO export</b></summary>
bash
# Portable lessons
curl -X POST http://localhost:3456/v1/lessons/export \
  -H "Authorization: Bearer $THUMBGATE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"outputPath": "./lessons-export.json"}'

# DPO pairs for fine-tuning
curl -X POST http://localhost:3456/v1/dpo/export \
  -H "Authorization: Bearer $THUMBGATE_API_KEY" \
  -o dpo-pairs.jsonl
</details>

Tech Stack

LayerTech
RuntimeNode.js ≥18
InterfacesMCP stdio, HTTP API, CLI
StorageSQLite + FTS5, LanceDB vectors, JSONL logs
IntelligenceMemAlign dual recall, Thompson Sampling, local embeddings
Billing / hostStripe, Railway
ExecutionRailway, Cloudflare Workers, Docker Sandboxes
GovernanceWorkflow Sentinel, control plane, Docker Sandboxes

Every Changeset is tied to the exact main merge commit and generates Verification Evidence for Release Confidence.


Integrations (compact)

SurfaceStart here
Open ThumbGate GPTthumbgate.ai/go/gptThumbGate GPT: start here. Paste agent actions, get advice + checkpointing. No, users do not have to keep chatting inside the ThumbGate GPT to use ThumbGate — the hard enforcement layer still runs where the work happens.
Install Codex PluginOpen the Codex plugin install page: thumbgate.ai/codex-plugin · zip: thumbgate-codex-plugin.zip · plugins/codex-profile/INSTALL.md
Claude Desktop .mcpblatest release
VS Code / Open VSXplugins/vscode-extension/README.md
Antigravity-compatibleplugins/antigravity-extension/INSTALL.md
JetBrainsplugins/jetbrains-plugin/README.md · JetBrains Marketplace path for the same runtime
ChatGPT App / GPT Actionthumbgate.ai/chatgpt-app
ThumbGate-Core (staging)https://github.com/IgorGanapolsky/ThumbGate-Core — pre-release staging + a few internal cache scripts; not the product moat

Docs

Full index: docs/INDEX.md

NeedLink
Agent workflow contractWORKFLOW.md
Ready-for-agent intake.github/ISSUE_TEMPLATE/ready-for-agent.yml
Verification Evidencedocs/VERIFICATION_EVIDENCE.md
Release Confidencedocs/RELEASE_CONFIDENCE.md
Changeset strategydocs/CHANGESET_STRATEGY.md
First Dollar Playbookdocs/FIRST_DOLLAR_PLAYBOOK.md
Security policySECURITY.md
Threat modelTHREAT_MODEL.md
Federal / regulateddocs/FEDERAL.md
Commercial Truthdocs/COMMERCIAL_TRUTH.md
Issues / PRsGitHub Issues · PR template

FAQ (one-liners): Not a fine-tuner (runtime intercept only). Different from CLAUDE.md / .cursorrules (those are context; ThumbGate is an external allow/warn/deny before tools run).


Who builds this

Igor Ganapolsky — payments (Stripe/Connect), AI agent guardrails/MCP, Android + backends. Small number of contract slots: $120–150/hr, 1099, remote US. LinkedIn · thumbgate.ai

License

MIT — see LICENSE. Project policy: SECURITY.md · THREAT_MODEL.md.

常见问题

io.github.IgorGanapolsky/rlhf-feedback-loop 是什么?

为AI agents提供RLHF feedback loop,捕获反馈信号、提升记忆、阻止错误,并导出DPO数据。

相关 Skills

网页构建器

by anthropics

Universal
热门

面向复杂 claude.ai HTML artifact 开发,快速初始化 React + Tailwind CSS + shadcn/ui 项目并打包为单文件 HTML,适合需要状态管理、路由或多组件交互的页面。

在 claude.ai 里做复杂网页 Artifact 很省心,多组件、状态和路由都能顺手搭起来,React、Tailwind 与 shadcn/ui 组合效率高、成品也更精致。

编码与调试
未扫描176.4k

网页应用测试

by anthropics

Universal
热门

用 Playwright 为本地 Web 应用编写自动化测试,支持启动开发服务器、校验前端交互、排查 UI 异常、抓取截图与浏览器日志,适合调试动态页面和回归验证。

借助 Playwright 一站式验证本地 Web 应用前端功能,调 UI 时还能同步查看日志和截图,定位问题更快。

编码与调试
未扫描176.4k

前端设计

by anthropics

Universal
热门

面向组件、页面、海报和 Web 应用开发,按鲜明视觉方向生成可直接落地的前端代码与高质感 UI,适合做 landing page、Dashboard 或美化现有界面,避开千篇一律的 AI 审美。

想把页面做得既能上线又有设计感,就用前端设计:组件到整站都能产出,难得的是能避开千篇一律的 AI 味。

编码与调试
未扫描176.4k

相关 MCP Server

GitHub

编辑精选

by GitHub

热门

GitHub 是 MCP 官方参考服务器,让 Claude 直接读写你的代码仓库和 Issues。

这个参考服务器解决了开发者想让 AI 安全访问 GitHub 数据的问题,适合需要自动化代码审查或 Issue 管理的团队。但注意它只是参考实现,生产环境得自己加固安全。

编码与调试
89.7k

by Context7

热门

Context7 是实时拉取最新文档和代码示例的智能助手,让你告别过时资料。

它能解决开发者查找文档时信息滞后的问题,特别适合快速上手新库或跟进更新。不过,依赖外部源可能导致偶尔的数据延迟,建议结合官方文档使用。

编码与调试
60.2k

by tldraw

热门

tldraw 是让 AI 助手直接在无限画布上绘图和协作的 MCP 服务器。

这解决了 AI 只能输出文本、无法视觉化协作的痛点——想象让 Claude 帮你画流程图或白板讨论。最适合需要快速原型设计或头脑风暴的开发者。不过,目前它只是个基础连接器,你得自己搭建画布应用才能发挥全部潜力。

编码与调试
49.9k

评论