VMware AIops

DevOps

by zw008

基于 AI 的 VMware vCenter/ESXi 监控与运维工具集,内含 20 个 MCP 工具,帮助排障、巡检和自动化操作。

什么是 VMware AIops?

基于 AI 的 VMware vCenter/ESXi 监控与运维工具集,内含 20 个 MCP 工具,帮助排障、巡检和自动化操作。

README

<!-- mcp-name: io.github.vmware-skills/vmware-aiops -->

VMware AIops

Author: Wei Zhou, VMware by Broadcom — wei-wz.zhou@broadcom.com This is a community-driven project by a VMware engineer, not an official VMware product. For official VMware developer tools see developer.broadcom.com.

English | 中文

AI-powered VMware vCenter/ESXi VM lifecycle and deployment tool — 60 tools.

Companion skills handle everything else:

SkillScopeInstall
vmware-monitorRead-only: inventory, health, alarms, events, metricsuv tool install vmware-monitor
vmware-storageDatastores, iSCSI, vSAN managementuv tool install vmware-storage
vmware-vksTanzu Namespaces, TKC cluster lifecycleuv tool install vmware-vks

Need read-only monitoring only? Use VMware-Monitor — zero destructive code in the codebase.

ClawHub Skills.sh Claude Code Marketplace License: MIT

⚡ Quick Investigation Reports (read-only)

Triage → investigate → act, all in one conversation. Five opinionated read-only reports aggregate and correlate server-side and hand back a high-signal result (never raw inventory), so you can decide where to look before changing anything. Each renders a self-contained offline HTML snapshot with --html (no external assets; drill-down detail collapses in native <details>, zero JavaScript). All delegate to the vmware-monitor library using AIops's own vCenter connection.

QuestionCommandWhat it correlates
"What needs attention now?" across all vCentersvmware-aiops attentionEvery vCenter merged into one globally-ranked issue list; unreachable targets degrade gracefully
"Is anything on fire?" across all clustersvmware-aiops summaryEvery cluster's hosts + VM power + live CPU/mem + alarms → ranked top-N issues + per-cluster status
"What's happening around this VM?"vmware-aiops investigate vm <name>VM state + host + cluster + backing datastores + snapshots + alarms + performance + a merged event timeline
"What's happening around this host?"vmware-aiops investigate host <name>Host state + cluster + the VMs it runs + mounted datastores + alarms + performance + correlated timeline
"What's happening around this datastore?"vmware-aiops investigate datastore <name>Capacity/free + mounting hosts + VMs it backs + alarms + correlated timeline
bash
vmware-aiops attention                            # what needs attention now, all vCenters
vmware-aiops investigate vm web-01 --hours 72     # everything around a VM, then act on it
vmware-aiops investigate vm web-01 --html         # → offline snapshot in ~/vmware-health/

Via MCP these are the tools cluster_health_summary, cross_vcenter_attention, vm_investigation_bundle, host_investigation_bundle, datastore_investigation_bundle. (Requires vmware-monitor installed.)

Quick Install (Recommended)

Works with Claude Code, Cursor, Codex, Gemini CLI, Trae, and 30+ AI agents:

bash
# Via Skills.sh
npx skills add vmware-skills/VMware-AIops

# Via ClawHub
clawhub install @zw008/vmware-aiops

PyPI Install (No GitHub Access Required)

bash
# Install via uv (recommended)
uv tool install vmware-aiops

# Or via pip
pip install vmware-aiops

# China mainland mirror (faster)
pip install vmware-aiops -i https://pypi.tuna.tsinghua.edu.cn/simple

Offline / Air-Gapped Install (from source)

This project uses the modern PEP 517 build system (hatchling), so there is no setup.py by design — that is expected, not a missing file. If you cloned the source and hit ERROR: File "setup.py" or "setup.cfg" not found ... editable mode currently requires a setuptools-based build, your pip is older than 21.3 and cannot do an editable (-e) install with a non-setuptools backend. Editable mode is a developer convenience, not needed to run the tool — do one of:

bash
# From the source tree — a normal (non-editable) install builds a wheel:
pip install .              # NOT  pip install -e .

# ...or upgrade pip first, and editable works too:
pip install --upgrade pip && pip install -e .

For a truly air-gapped host, build the wheels on a connected machine and copy them over — the target then needs no network:

bash
# On a connected machine, collect this package + its dependencies as wheels:
pip wheel . -w dist        # → dist/*.whl   (or: uv build, for just this package)

# Copy dist/ to the air-gapped host, then install offline:
pip install --no-index --find-links dist vmware-aiops

Why this over other VMware MCP servers

Most open-source VMware MCP servers (e.g. bright8192/esxi-mcp-server, giuliolibrando/vmware-vsphere-mcp-server) are single-vCenter VM wrappers: list/power/snapshot a VM, basic monitoring, a confirm=True flag. They explicitly do not cover networking, storage, Kubernetes, ops analytics, load balancing, or compliance — and "logging is documented" is not an audit trail.

This is one skill in an 11-package family that covers the whole estate and runs every tool through a governed harness:

Other VMware MCP serversThis family
VM lifecycle + monitoring✅✅
NSX networking (segments/gateways/NAT/routing/IPAM)❌✅ vmware-nsx
NSX security (DFW/groups/IDS-IPS/traceflow)❌✅ vmware-nsx-security
Storage (datastore/iSCSI/vSAN)❌✅ vmware-storage
Tanzu Kubernetes (Supervisor/Namespace/TKC)❌✅ vmware-vks
Aria Operations (metrics/alerts/capacity)❌✅ vmware-aria
AVI / NSX ALB load balancing + AKO❌✅ vmware-avi
Compliance baselines + drift (CIS/SCG/等保/PCI)❌✅ vmware-harden
Governed harness (unified audit, policy engine, token budget + runaway breaker, graduated risk tiers, undo-token, prompt-injection sanitize)❌✅ vmware-policy on every tool

If you only ever power-cycle VMs in one vCenter, a single-file server is fine. If you run a real (regulated, NSX-segmented, multi-domain) VMware estate and need an AI operator an auditor can sign off on, that's what this family is for — see docs/compliance-ready.md.

Capabilities Overview

What This Skill Does

CategoryToolsCount
VM Lifecyclepower on/off, TTL auto-delete, clean slate6
DeploymentOVA, template, linked clone, batch clone/deploy8
Guest Opsexec commands, upload/download files, provision5
Plan/Applymulti-step planning with rollback4
Clustercreate, delete, HA/DRS config, add/remove hosts6
Datastorebrowse files, scan for images2
NetworkdvSwitch portgroup list/create, host VMkernel list/add/remove, DF-bit MTU-path ping6

CLI vs MCP: Which Mode to Use

ScenarioRecommendedWhy
Local/small models (Ollama, Qwen <32B)CLI~2K tokens context vs ~10K for MCP; small models struggle with many tool schemas
Token-sensitive workflowsCLISKILL.md + Bash tool = minimal overhead
Cloud models (Claude, GPT-4o)EitherBoth work; MCP gives structured JSON I/O
Automated pipelines / Agent chainingMCPType-safe parameters, structured output, no shell parsing
Monitoring / storage / K8sCompanion skillsSee vmware-monitor, vmware-storage, vmware-vks

Rule of thumb: Use CLI for cost efficiency and small models. Use MCP for structured automation with large models.

Architecture

code
User (Natural Language)
  ↓
AI CLI Tool (Claude Code / Gemini / Codex / Aider / Continue / Trae / Kimi)
  ↓ reads SKILL.md / AGENTS.md / rules
  ↓
vmware-aiops CLI
  ↓ pyVmomi (vSphere SOAP API)
  ↓
vCenter Server ──→ ESXi Cluster ──→ VM
    or
ESXi Standalone Host ──→ VM

Version Compatibility

vSphere / VCF VersionSupportNotes
VCF 9.1 / vSphere 9.1✅ FullReleased 2026-05-12. pyVmomi <10.0 resolves and connects via SOAP; new REST-only features (PATCH /deployment/size, IPv6-only GOSC) not yet wrapped — see VCF Python SDK for those.
VCF 9.0 / vSphere 9.0✅ FullpyVmomi 8.0.3+ connects against vSphere 9 SOAP API. From VCF 9, pyVmomi is also bundled inside the unified VCF Python SDK.
8.0 / 8.0U1-U3✅ FullCreateSnapshot_Task deprecated → use CreateSnapshotEx_Task
7.0 / 7.0U1-U3✅ FullAll APIs supported
6.7✅ CompatibleBackward-compatible, tested
6.5✅ CompatibleBackward-compatible, tested

pyVmomi auto-negotiates the API version during SOAP handshake — no manual configuration needed. The same codebase manages 7.0 / 8.0 / 9.0 / 9.1 environments seamlessly.

Official Broadcom References


Common Workflows

Deploy a Lab Environment

  1. Browse datastore for OVA images → vmware-aiops datastore browse <ds> --pattern "*.ova"
  2. Deploy VM from OVA → vmware-aiops deploy ova ./image.ova --name lab-vm --datastore ds1
  3. Install software inside VM → vmware-aiops vm guest-exec lab-vm --cmd /bin/bash --args "-c 'apt-get install -y nginx'" --user root
  4. Create baseline snapshot → vmware-aiops vm snapshot-create lab-vm --name baseline
  5. Set TTL for auto-cleanup → vmware-aiops vm set-ttl lab-vm --minutes 480

Batch Clone for Testing

  1. Create plan: vm_create_plan with multiple clone + reconfigure steps
  2. Review plan with user (shows affected VMs, irreversible warnings)
  3. Apply: vm_apply_plan executes sequentially, stops on failure
  4. If failed: vm_rollback_plan reverses executed steps
  5. Set TTL on all clones for auto-cleanup

Migrate VM to Another Host

  1. Check VM info via vmware-monitor → verify power state and current host
  2. Migrate: vmware-aiops vm migrate my-vm --to-host esxi-02
  3. Verify migration completed

VM Lifecycle

OperationCommandConfirmationvCenterESXi
Power Onvm power-on <name>—✅✅
Graceful Shutdownvm power-off <name>Double✅✅
Force Power Offvm power-off <name> --forceDouble✅✅
Resetplan action reset via vm_create_plan (MCP; no CLI command)—✅✅
Suspendplan action suspend via vm_create_plan (MCP; no CLI command)—✅✅
Create VMvm create <name> --cpu --memory --disk—✅✅
Delete VMvm delete <name>Double✅✅
Reconfigurevm reconfigure <name> --cpu --memoryDouble✅✅
Create Snapshotvm snapshot-create <name> --name <snap>—✅✅
List Snapshotsvm snapshot-list <name>—✅✅
Revert Snapshotvm snapshot-revert <name> --name <snap>Double✅✅
Delete Snapshotvm snapshot-delete <name> --name <snap> [--no-wait]Double✅✅
Task Statusvm task-status <task-id>—✅✅
Clone VMvm clone <name> --new-name <new>Double✅✅
vMotionvm migrate <name> --to-host <host>Double✅❌
Set TTLvm set-ttl <name> --minutes <n>Double✅✅
Cancel TTLvm cancel-ttl <name>—✅✅
List TTLsvm list-ttl—✅✅
Clean Slatevm clean-slate <name> [--snapshot baseline]Double✅✅
Guest Execvm guest-exec <name> --cmd /bin/bash --args "..." --user <account>Double✅✅
Guest Exec (with output)MCP only: vm_guest_exec_output (username required) — no CLI command—✅✅
Guest Uploadvm guest-upload <name> --local f.sh --guest /tmp/f.sh --user <account>Double✅✅
Guest Downloadvm guest-download <name> --guest /var/log/syslog --local ./syslog --user <account>—✅✅

Guest Operations require VMware Tools running inside the guest OS, and the guest account is always named explicitly — --user on the CLI, username over MCP. There is no default, so no call runs as root without choosing root. vm_guest_exec_output (MCP only; the CLI has no equivalent) auto-detects Linux/Windows shell and captures stdout/stderr.

Plan → Apply (Multi-step Operations)

For complex operations involving 2+ steps or 2+ VMs, use the plan/apply workflow instead of executing individually:

StepWhat Happens
1. Create PlanAI calls vm_create_plan — validates actions, checks targets in vSphere, generates plan with rollback info
2. ReviewAI shows plan to user: steps, affected VMs, irreversible warnings
3. Applyvm_apply_plan previews first; with confirm=True executes sequentially and stops on failure. Each gated step is checked as its own tool checks it; iSCSI/rescan steps are refused (use vmware-storage)
4. Rollback (if failed)Asks user whether to rollback, then vm_rollback_plan reverses executed steps (irreversible steps skipped); each destructive rollback step is checked first, and a refused check stops the rollback

Plans stored in ~/.vmware-aiops/plans/, auto-deleted on success, auto-cleaned after 24h.

VM Deployment & Provisioning

OperationCommandSpeedvCenterESXi
Deploy from OVAdeploy ova <path> --name <vm>Minutes✅✅
Deploy from Templatedeploy template <tmpl> --name <vm>Minutes✅✅
Linked Clonedeploy linked-clone --source <vm> --snapshot <snap> --name <new>Seconds✅✅
Attach ISOdeploy iso <vm> --iso "[ds] path/to.iso"Instant✅✅
Convert to Templatedeploy mark-template <vm>Instant✅✅
Batch Clonedeploy batch-clone --source <vm> --count <n>Minutes✅✅
Batch Deploy (YAML)deploy batch spec.yamlAuto✅✅

Cluster Management

OperationCommandConfirmationvCenterESXi
Cluster Infocluster info <name>—✅❌
Create Clustercluster create <name> [--ha] [--drs]—✅❌
Delete Clustercluster delete <name>Double✅❌
Add Hostcluster add-host <cluster> --host <host>Double✅❌
Remove Hostcluster remove-host <cluster> --host <host>Double✅❌
Configure HA/DRScluster configure <name> [--ha/--no-ha] [--drs/--no-drs]Double✅❌

remove-host requires the host to be in maintenance mode first; the host is moved out of the cluster into the datacenter's host folder as a standalone host.

Alarm Management

OperationCommandConfirmationvCenterESXi
List Triggered Alarmsalarm list [--target <t>]—✅❌
Acknowledge Alarmalarm acknowledge <entity> <alarm>—✅❌
Clear (Reset) Alarmsalarm reset <entity> <alarm>Double✅❌

Blast radius: vSphere has no per-alarm clear API. alarm reset uses AlarmManager.ClearTriggeredAlarms, which clears all triggered alarms matching the named alarm's entity type (host/VM/all) and current status (red/yellow) — not just the named one. The named alarm is looked up first (typos fail fast), and the output's scope field reports exactly what was cleared. Cleared alarms re-trigger automatically if their underlying condition persists.

Datastore Browser

FeaturevCenterESXiDetails
Browse Files✅✅List files/folders in any datastore path
Scan Images✅✅Discover ISO, OVA, OVF, VMDK across all datastores

Scheduled Scanning & Notifications

FeatureDetails
DaemonAPScheduler-based, configurable interval (default 15 min)
Multi-target ScanSequentially scan all configured vCenter/ESXi targets
Scan ContentEach cycle: triggered alarms, vCenter events from the last lookback_hours, and new lines in the ESXi host logs hostd, vmkernel, vpxa
Host LogsRead incrementally: each line is reported once per daemon run (a restarted daemon re-reads each log's last 500 lines once). A rotated log, or more than 500 new lines between cycles, adds an info row saying which lines were not scanned. Reading host logs needs the Global.Diagnostics privilege, which vCenter's Read-Only role does not include; a log that cannot be read becomes an info row with the reason, never a silent "all clear"
Log AnalysisHost-log lines matching error, fail, critical, panic, lost access, cannot, timeout, refused, corrupt — lines with critical/panic/corrupt are critical, the rest warning
Structured LogJSONL output to ~/.vmware-aiops/scan.log — every issue, info rows included
WebhookSlack, Discord, or any HTTP endpoint. Receives every critical issue and every alarm/event warning; host-log warnings go to the scan log only, and info rows are never sent
Cycle SummaryOne line per cycle in the daemon's log output: findings (and how many went to the webhook), unreadable host logs, logs with unscanned lines, failed passes. If any pass failed or a target could not be reached it reads Scan INCOMPLETE, never "all clear"
Daemon Managementdaemon start/stop/status, PID file, graceful shutdown

Safety Features

FeatureDetails
Dry-Run Mode (CLI only)--dry-run prints the exact API call without executing, on every CLI write except deploy iso, deploy mark-template, vm cancel-ttl and vm guest-download
Plan → Confirm → Execute → LogCLI workflow: show current state, confirm changes, execute, audit log
Double Confirmation (CLI only)Destructive and deploy CLI commands (vm power-off, delete, reconfigure, snapshot-revert/delete, clone, migrate, set-ttl, clean-slate, guest-exec, guest-upload; deploy ova, template, linked-clone, batch, batch-clone, mark-template; cluster delete, add-host, remove-host, configure, drs-rule-set/create/delete; alarm reset) require 2 sequential prompts and take no bypass flag
Only destructive MCP tools confirm22 of the 43 write tools an agent sees over MCP — every tool annotated destructive, plus vm_migrate and the network/DRS authoring tools — preview first and need confirm=True; the other 21 (creates, clones, deploys, power-on, reconfigure, snapshot create, host add, template and ISO operations, guest download, plan creation, alarms) act on the first call. There is no approval tier and no read-only switch. What decides whether a write lands is the privilege of the vCenter account, and what records it is the audit trail. See What protects you
Rejection LoggingDeclined CLI confirmations are recorded in the audit trail
Audit TrailAll operations logged to ~/.vmware-aiops/audit.log (JSONL) with before/after state
Input ValidationVM name, CPU (1-128), memory (128-1048576 MB), disk (1-65536 GB) validated
Password Protection.env file loading with permission check; never in shell history
SSL Self-signed Supportverify_ssl: false — only for ESXi with self-signed certs in isolated labs; production should use CA-signed certificates
Prompt Injection ProtectionvSphere event messages and host logs are truncated, stripped of control characters, and wrapped in boundary markers before output
Webhook Data ScopeDisabled by default. When configured, the daemon posts to your URL only: every critical issue (alarms, events, ESXi log lines matching critical/panic/corrupt, targets it could not connect to) and every alarm/event warning — host-log warnings stay in the scan log, and info rows are never sent. Each issue carries its entity name and message: sanitized alarm, event, or ESXi log text, or the connection error, which can include host names, IP addresses, and user names. No credentials from the skill's config or .env are sent
Task WaitingAll async operations wait for completion and report result
State ValidationPre-operation checks (VM exists, power state correct)

vCenter vs ESXi Comparison

CapabilityvCenterESXi Standalone
vMotion migration✅❌
Cross-host clone✅❌
Cluster management✅❌
All VM lifecycle ops✅✅
OVA/Template/Linked Clone deploy✅✅
Datastore browsing & image scan✅✅
Snapshots✅✅
Guest operations✅✅

Inventory, alarms, events, sensors, host services, and scanning are now in vmware-monitor.

What protects you

The table above lists two different surfaces and it is worth being blunt about which protections apply to which, because getting this wrong is worse than having no protection at all — a guardrail you believe in is one you stop compensating for.

On the CLI, a destructive command asks twice and takes no bypass flag, and --dry-run previews every write except deploy iso, deploy mark-template, vm cancel-ttl and vm guest-download. That defends a mistyped command typed by a human. It does not defend against an agent, which satisfies both prompts with yes |.

Over MCP, every tool annotated destructive, plus vm_migrate and the network/DRS authoring tools — 22 of the 43 write tools, including vm_power_off, vm_migrate, cluster_delete, the snapshot, guest, TTL, Clean Slate and plan tools — takes one argument, confirm, whose default is a no-write preview: a bare call returns its blast radius and changes nothing. confirm=True re-measures and is refused, with a teaching error audited as a failure, if the preview found a blocker (a VM without running VMware Tools, a target host in maintenance mode, a cluster that still has hosts, a missing or duplicated snapshot) or could not read something. vm_delete also takes the preview's acknowledge_with echoed back, and is refused if the VM changed since, is powered on or suspended. The other 21 write tools — creates, clones, deploys, power-on, reconfigure, snapshot create, host add, template and ISO operations, guest download, plan creation, alarms — act on the first call. A confirmation is not authorization, which is why the VMWARE_READ_ONLY switch stays removed (it was enforced on the MCP path only, and any agent with a shell walked around it via the CLI). What the preview buys is narrower: an agent does not destroy something it has not looked at.

What actually decides whether a write lands is the vCenter/ESXi service account. Give the skill an account with the privileges the work needs and no more; vCenter refuses the rest itself, on every surface, with no way around it from inside this skill. To run an agent read-only, give it a read-only vCenter role — one decision, enforced where it is made. Every call is then recorded in ~/.vmware/audit.db before the caller sees a result, which is how you find out what happened. Optional deny rules in ~/.vmware/rules.yaml, checked before every MCP call, can refuse operations — for example, writes to targets labelled environment: production. The shipped baseline denies nothing, and the rules run inside the same process: a guardrail on top of RBAC, not a replacement.

vm_guest_exec is the one to think hardest about. It runs a caller-supplied command inside the guest OS with the credentials handed to it — its username is required (no default account); nothing bounds what the command may be. The guest account is a separate authorization boundary from the vCenter one — a read-only vCenter role does not constrain what this tool does inside a VM. The skill stores no guest credentials: over MCP they are tool arguments the agent sees (the audit row redacts the password). Pass a least-privilege guest account, and do not hand the agent guest credentials it does not need. vm_guest_upload reads any local file the server process can read.

The full inventory of which tools are gated and which are not is in references/capabilities.md, where the numbers are checked against the live tool registry by the test suite rather than maintained by hand.


Troubleshooting

"VM not found" error

VM names are case-sensitive in vSphere. Use exact name from vmware-monitor inventory vms.

Guest exec returns empty output

Use vm_guest_exec_output instead of vm_guest_exec — it auto-captures stdout/stderr. Basic vm_guest_exec only returns exit code.

Deploy OVA times out

Large OVA files (>10GB) may exceed the default 120s timeout. The upload happens via HTTP NFC lease — ensure network between the machine running vmware-aiops and ESXi is stable.

Plan apply fails mid-way

Run vmware-aiops plan list to see failed plan status. Ask user if they want to rollback with vm_rollback_plan. Irreversible steps (delete_vm) are skipped during rollback.

Connection refused / SSL error

  1. Verify target is reachable: vmware-aiops doctor
  2. For self-signed certs: set verify_ssl: false in config.yaml (lab environments only)

Supported AI Platforms

PlatformStatusConfig FileAI Model
Claude Code✅ Native Skillskills/vmware-aiops/SKILL.mdAnthropic Claude
Gemini CLI✅ Context file + MCPskills/vmware-aiops/SKILL.mdGoogle Gemini
OpenAI Codex CLI✅ Skill + AGENTS.mdskills/vmware-aiops/SKILL.mdOpenAI GPT
Aider✅ Conventionsskills/vmware-aiops/SKILL.mdAny (cloud + local)
Continue CLI✅ Rulesskills/vmware-aiops/SKILL.mdAny (cloud + local)
Trae IDE✅ Rulesskills/vmware-aiops/SKILL.mdClaude/DeepSeek/GPT-4o/Doubao
Kimi Code CLI✅ Skillskills/vmware-aiops/SKILL.mdMoonshot Kimi
MCP Server✅ MCP Protocolvmware_aiops/mcp_server/Any MCP client
Python CLI✅ StandaloneN/AN/A

Platform Comparison

FeatureClaude CodeGemini CLICodex CLIAiderContinueTrae IDEKimi CLI
Cloud AIAnthropicGoogleOpenAIAnyAnyMultiMoonshot
Local models———OllamaOllama——
Skill systemSKILL.mdContext fileSKILL.md—RulesRulesSKILL.md
MCP supportNativeNativeVia SkillsThird-partyNative——
Free tier—60 req/min—Self-hostedSelf-hosted——

MCP Server Integrations

The vmware-aiops MCP server works with any MCP-compatible agent or tool. Ready-to-use configuration templates are in examples/mcp-configs/.

Agent / ToolLocal Model SupportConfig TemplateIntegration Guide
Xiaoguai (小怪)✅ Self-hosted, any LLMMCP setupGuide
Goose✅ Ollama, LM Studiogoose.jsonGuide
LocalCowork✅ Fully offlinelocalcowork.jsonGuide
mcp-agent✅ Ollama, vLLMmcp-agent.yamlGuide
VS Code Copilot—vscode-copilot.jsonGuide
Cursor—cursor.jsonGuide
Continue✅ Ollamacontinue.yamlGuide
Claude Code—claude-code.json—

Xiaoguai (小怪) — a self-hostable, audit-first agent platform (Rust, single binary + embedded SQLite) from the same maintainer. It runs the vmware-aiops MCP server as one of its toolboxes; being both an MCP consumer and an MCP server, its HMAC-chained audit log and human-on-the-loop approval gates line up with this skill's own audit + confirm design. See its MCP integration guide.

Fully local operation (no cloud API required):

bash
# Aider + Ollama + vmware-aiops (via SKILL.md)
aider --conventions skills/vmware-aiops/SKILL.md --model ollama/qwen2.5-coder:32b

# Any MCP agent + local model + vmware-aiops MCP server
# See examples/mcp-configs/ for your agent's config format

Installation

Step 0: Prerequisites

bash
# Python 3.10+ required
python3 --version

# Node.js 18+ required for Gemini CLI and Codex CLI
node --version

Step 1: Clone & Install Python Backend

All platforms share the same Python backend.

bash
git clone https://github.com/vmware-skills/VMware-AIops.git
cd VMware-AIops
python3 -m venv .venv
source .venv/bin/activate
pip install -e .

Step 2: Configure

bash
mkdir -p ~/.vmware-aiops
cp config.example.yaml ~/.vmware-aiops/config.yaml
# Edit config.yaml with your vCenter/ESXi targets

Set passwords via .env file (recommended):

bash
# Use the template
cp .env.example ~/.vmware-aiops/.env

# Edit and fill in your passwords, then lock permissions
chmod 600 ~/.vmware-aiops/.env

Security note: Prefer .env file over command-line export to avoid passwords appearing in shell history. The .env file should have chmod 600 (owner-only read/write).

Password environment variable naming convention:

code
VMWARE_{TARGET_NAME_UPPER}_PASSWORD
# Replace hyphens with underscores, UPPERCASE
# Example: target "home-esxi" → VMWARE_HOME_ESXI_PASSWORD
# Example: target "prod-vcenter" → VMWARE_PROD_VCENTER_PASSWORD

Security Best Practices

  • NEVER hardcode passwords in scripts or config files
  • NEVER pass passwords as command-line arguments (visible in ps)
  • ALWAYS use ~/.vmware-aiops/.env with chmod 600
  • ALWAYS configure connections via config.yaml — credentials are loaded from .env automatically
  • Config File Contents: config.yaml stores target hostnames, ports, and a reference to the .env file. It does not contain passwords or tokens. All secrets are stored exclusively in .env
  • TLS: Enabled by default. Disable only for ESXi hosts with self-signed certificates in isolated lab environments
  • Webhook: Disabled by default. When enabled, the daemon posts to your own configured URL only — no third-party service — every critical issue (alarms, events, ESXi log lines matching critical/panic/corrupt, targets it could not connect to) and every alarm/event warning — host-log warnings stay in the scan log, and info rows are never sent. Payloads carry the full issue text (entity names, alarm names, event messages, ESXi log excerpts, connection errors), which can include host names, IP addresses, and user names; they carry no credentials from your config or .env
  • Least Privilege: Use a dedicated vCenter service account with minimal permissions. For monitoring-only use cases, prefer the read-only VMware-Monitor
  • Prompt Injection Protection: All vSphere-sourced content is truncated, stripped of control characters, and wrapped in boundary markers before output
  • Code Review: We recommend reviewing the source code and commit history before deploying in production
  • Production Safety: For production environments, use the read-only VMware-Monitor instead. AI agents can misinterpret context and execute unintended destructive operations — real-world incidents have shown that AI-driven infrastructure tools without proper isolation can delete production databases and entire environments. VMware-Monitor eliminates this risk at the code level: no destructive functions exist in its codebase

Step 3: Connect Your AI Tool

Choose one (or more) of the following:


Option A: Claude Code

Method 1: Skills.sh or ClawHub (recommended)

Either installer places the skill in Claude Code's skills directory for you:

bash
npx skills add vmware-skills/VMware-AIops
# or
clawhub install @zw008/vmware-aiops

Method 2: Manual skill install

bash
git clone https://github.com/vmware-skills/VMware-AIops.git
cd VMware-AIops

# Copy the skill into Claude Code's personal skills directory
mkdir -p ~/.claude/skills/vmware-aiops
cp -r skills/vmware-aiops/. ~/.claude/skills/vmware-aiops/

For tool access (not just skill context), also register the MCP server:

bash
claude mcp add vmware-aiops -- vmware-aiops mcp

Restart Claude Code, then:

code
> Show me all VMs on esxi-lab.example.com

Submit to Official Marketplace

This plugin can also be submitted to the Anthropic official plugin directory for public discovery.


Option B: Gemini CLI

bash
# Install Gemini CLI
npm install -g @google/gemini-cli

# Load the skill as project context (Gemini CLI reads GEMINI.md on startup)
cp skills/vmware-aiops/SKILL.md ./GEMINI.md

For tool access (not just context), register the MCP server in ~/.gemini/settings.json:

json
{
  "mcpServers": {
    "vmware-aiops": {
      "command": "vmware-aiops",
      "args": ["mcp"],
      "env": { "VMWARE_AIOPS_CONFIG": "~/.vmware-aiops/config.yaml" }
    }
  }
}

Then start Gemini CLI:

code
gemini
> Show me all VMs on my ESXi host

Option C: OpenAI Codex CLI

bash
# Install Codex CLI
npm i -g @openai/codex
# Or on macOS:
# brew install --cask codex

# Copy skill to Codex skills directory
mkdir -p ~/.codex/skills/vmware-aiops
cp skills/vmware-aiops/SKILL.md ~/.codex/skills/vmware-aiops/SKILL.md

# Copy AGENTS.md to project root
cp skills/vmware-aiops/SKILL.md ./AGENTS.md

Then start Codex CLI:

bash
codex --enable skills
> List all VMs on my ESXi

Option D: Aider (supports local models)

bash
# Install Aider
pip install aider-chat

# Install Ollama for local models (optional)
# macOS:
brew install ollama
ollama pull qwen2.5-coder:32b

# Run with cloud API
aider --conventions skills/vmware-aiops/SKILL.md

# Or with local model via Ollama
aider --conventions skills/vmware-aiops/SKILL.md \
  --model ollama/qwen2.5-coder:32b

Option E: Continue CLI (supports local models)

bash
# Install Continue CLI
npm i -g @continuedev/cli

# Copy rules file
mkdir -p .continue/rules
cp skills/vmware-aiops/SKILL.md .continue/rules/vmware-aiops.md

Configure ~/.continue/config.yaml for local model:

yaml
models:
  - name: local-coder
    provider: ollama
    model: qwen2.5-coder:32b

Then:

bash
cn
> Check ESXi health and alarms

Option F: Trae IDE

Copy the rules file to your project's .trae/rules/ directory:

bash
mkdir -p .trae/rules
cp skills/vmware-aiops/SKILL.md .trae/rules/project_rules.md

Trae IDE's Builder Mode reads .trae/rules/ Markdown files at startup.

Note: You can also install Claude Code extension in Trae IDE and use .claude/skills/ format directly.


Option G: Kimi Code CLI

bash
# Copy skill file to Kimi skills directory
mkdir -p ~/.kimi/skills/vmware-aiops
cp skills/vmware-aiops/SKILL.md ~/.kimi/skills/vmware-aiops/SKILL.md

Option H: MCP Server (Glama / Claude Desktop)

The MCP server exposes VMware operations as tools via the Model Context Protocol. Works with any MCP-compatible client (Claude Desktop, Cursor, etc.).

After uv tool install vmware-aiops, start the MCP server with one command (v1.5.15+):

bash
# Recommended — single command, no network re-resolve
vmware-aiops mcp

# With a custom config path
VMWARE_AIOPS_CONFIG=/path/to/config.yaml vmware-aiops mcp

Claude Desktop config (claude_desktop_config.json):

json
{
  "mcpServers": {
    "vmware-aiops": {
      "command": "vmware-aiops",
      "args": ["mcp"],
      "env": {
        "VMWARE_AIOPS_CONFIG": "/path/to/config.yaml"
      }
    }
  }
}
<details> <summary>Alternative: uvx (no install) or legacy entry point</summary>
bash
# Run without installing (requires PyPI access each launch)
uvx --from vmware-aiops vmware-aiops mcp

# Legacy entry point (still works, kept for backward compatibility)
vmware-aiops-mcp

Behind a corporate TLS proxy? uvx may fail with invalid peer certificate: UnknownIssuer. Use the recommended vmware-aiops mcp form above (no network needed), or set UV_NATIVE_TLS=true.

</details>

Option I: Standalone CLI (no AI)

bash
# Already installed in Step 1
source .venv/bin/activate

vmware-aiops vm power-on my-vm --target home-esxi
vmware-aiops deploy ova ./ubuntu.ova --name my-vm --target home-esxi
vmware-aiops datastore browse datastore1 --target home-esxi

Update / Upgrade

Already installed? Re-run the install command for your channel to get the latest version:

Install ChannelUpdate Command
ClawHubclawhub install @zw008/vmware-aiops
Skills.shnpx skills add vmware-skills/VMware-AIops
Git clonecd VMware-AIops && git pull origin main && uv pip install --no-sources -e . (without --no-sources, uv looks for a sibling ../VMware-Monitor checkout)
uvuv tool install vmware-aiops --force

Check your current version: vmware-aiops --version


Chinese Cloud Models

For users in China who prefer domestic cloud APIs or have limited access to overseas services.

DeepSeek

Cost-effective, strong coding capability.

bash
# Set DeepSeek API key (get from https://platform.deepseek.com)
export DEEPSEEK_API_KEY="your-key"

# Run with Aider
aider --conventions skills/vmware-aiops/SKILL.md \
  --model deepseek/deepseek-coder

Persistent config ~/.aider.conf.yml:

yaml
model: deepseek/deepseek-coder
conventions: skills/vmware-aiops/SKILL.md

Qwen (Alibaba Cloud)

Alibaba Cloud's coding model, free tier available.

bash
# Set DashScope API key (get from https://dashscope.console.aliyun.com)
export DASHSCOPE_API_KEY="your-key"

aider --conventions skills/vmware-aiops/SKILL.md \
  --model qwen/qwen-coder-plus

Or via OpenAI-compatible endpoint:

bash
export OPENAI_API_BASE="https://dashscope.aliyuncs.com/compatible-mode/v1"
export OPENAI_API_KEY="your-dashscope-key"

aider --conventions skills/vmware-aiops/SKILL.md \
  --model qwen-coder-plus-latest

Doubao (ByteDance)

bash
export OPENAI_API_BASE="https://ark.cn-beijing.volces.com/api/v3"
export OPENAI_API_KEY="your-ark-key"

aider --conventions skills/vmware-aiops/SKILL.md \
  --model your-doubao-endpoint-id

With Continue CLI

Configure ~/.continue/config.yaml:

yaml
# DeepSeek
models:
  - name: deepseek-coder
    provider: openai-compatible
    apiBase: https://api.deepseek.com/v1
    apiKey: your-deepseek-key
    model: deepseek-coder

# Qwen
models:
  - name: qwen-coder
    provider: openai-compatible
    apiBase: https://dashscope.aliyuncs.com/compatible-mode/v1
    apiKey: your-dashscope-key
    model: qwen-coder-plus-latest

Local Models (Aider + Ollama)

For fully offline operation — no cloud API, no internet, full privacy.

Aider + Ollama + local Qwen/DeepSeek is ideal for air-gapped environments.

Step 1: Install Ollama

bash
# macOS
brew install ollama

# Linux — download from https://ollama.com/download and install manually
# See https://github.com/ollama/ollama for platform-specific instructions

Step 2: Pull a model

ModelCommandSizeNote
Qwen 2.5 Coder 32Bollama pull qwen2.5-coder:32b~20GBBest local coding model
Qwen 2.5 Coder 7Bollama pull qwen2.5-coder:7b~4.5GBLow-memory option
DeepSeek Coder V2ollama pull deepseek-coder-v2~8.9GBStrong reasoning
CodeLlama 34Bollama pull codellama:34b~19GBMeta coding model

Hardware: 32B → ~20GB VRAM (or 32GB RAM for CPU). 7B → 8GB RAM.

Step 3: Run with Aider

bash
pip install aider-chat
ollama serve

# Aider + local Qwen (recommended)
aider --conventions skills/vmware-aiops/SKILL.md \
  --model ollama/qwen2.5-coder:32b

# Aider + local DeepSeek
aider --conventions skills/vmware-aiops/SKILL.md \
  --model ollama/deepseek-coder-v2

# Low-memory option
aider --conventions skills/vmware-aiops/SKILL.md \
  --model ollama/qwen2.5-coder:7b

Persistent config ~/.aider.conf.yml:

yaml
model: ollama/qwen2.5-coder:32b
conventions: skills/vmware-aiops/SKILL.md

Local Architecture

code
User → Aider CLI → Ollama (localhost:11434) → Qwen / DeepSeek local model
  │                                                    ↓
  │                                          reads AGENTS.md instructions
  │                                                    ↓
  └──────────────────────────────→ vmware-aiops CLI ──→ ESXi / vCenter

Tip: Local models are fully offline — perfect for air-gapped environments or strict data compliance.


CLI Reference

bash
# Diagnostics
vmware-aiops doctor                   # Check environment, config, connectivity
vmware-aiops doctor --skip-auth       # Skip vSphere auth check (faster)

# MCP Config Generator
vmware-aiops mcp-config generate --agent goose        # Generate config for Goose
vmware-aiops mcp-config generate --agent claude-code  # Generate config for Claude Code
vmware-aiops mcp-config list                          # List all supported agents

# VM operations
vmware-aiops vm power-on my-vm                                 # Power on
vmware-aiops vm power-off my-vm                                # Graceful shutdown (2x confirm)
vmware-aiops vm power-off my-vm --force                        # Force power off (2x confirm)
vmware-aiops vm create my-new-vm --cpu 4 --memory 8192 --disk 100  # Create VM
vmware-aiops vm delete my-vm                                   # Delete VM (asks twice; --dry-run previews)
vmware-aiops vm reconfigure my-vm --cpu 4 --memory 8192        # Reconfigure (2x confirm)
vmware-aiops vm snapshot-create my-vm --name "before-upgrade"  # Create snapshot
vmware-aiops vm snapshot-list my-vm                            # List snapshots
vmware-aiops vm snapshot-revert my-vm --name "before-upgrade"  # Revert snapshot
vmware-aiops vm snapshot-delete my-vm --name "before-upgrade"  # Delete snapshot (waits ≤30 min for consolidation)
vmware-aiops vm snapshot-delete my-vm --name "old-big" --no-wait  # Fire async, return a task id
vmware-aiops vm task-status task-1234                          # Poll an async task by id
vmware-aiops vm clone my-vm --new-name my-vm-clone             # Clone VM
vmware-aiops vm migrate my-vm --to-host esxi-02                # vMotion
vmware-aiops vm set-ttl my-vm --minutes 60                     # Auto-delete in 60 min
vmware-aiops vm cancel-ttl my-vm                               # Cancel TTL
vmware-aiops vm list-ttl                                       # Show all TTLs
vmware-aiops vm clean-slate my-vm --snapshot baseline          # Revert to baseline (2x confirm)

# Guest Operations (requires VMware Tools in guest)
vmware-aiops vm guest-exec my-vm --cmd /bin/bash --args "-c 'whoami'" --user root
vmware-aiops vm guest-upload my-vm --local ./script.sh --guest /tmp/script.sh --user root
vmware-aiops vm guest-download my-vm --guest /var/log/syslog --local ./syslog.txt --user root

# Plan → Apply (multi-step operations)
vmware-aiops plan list                                        # List pending/failed plans

# Deploy
vmware-aiops deploy ova ./ubuntu.ova --name my-vm --datastore ds1      # Deploy from OVA
vmware-aiops deploy template golden-ubuntu --name new-vm               # Deploy from template
vmware-aiops deploy linked-clone --source base-vm --snapshot clean --name test-vm  # Linked clone (seconds)
vmware-aiops deploy iso my-vm --iso "[datastore1] iso/ubuntu-22.04.iso"  # Attach ISO
vmware-aiops deploy mark-template golden-vm                            # Convert VM to template
vmware-aiops deploy batch-clone --source base-vm --count 5 --prefix lab  # Batch clone
vmware-aiops deploy batch deploy.yaml                                  # Batch deploy from YAML spec

# Cluster
vmware-aiops cluster info my-cluster                                   # Cluster details (HA/DRS status)
vmware-aiops cluster create my-cluster --ha --drs                      # Create cluster with HA+DRS
vmware-aiops cluster delete my-cluster                                 # Delete cluster (2x confirm)
vmware-aiops cluster add-host my-cluster --host esxi-03                # Add host to cluster (2x confirm)
vmware-aiops cluster remove-host my-cluster --host esxi-03             # Remove host (2x confirm)
vmware-aiops cluster configure my-cluster --ha --drs                   # Configure HA/DRS (2x confirm)

# Alarm management
vmware-aiops alarm list                                                # List triggered alarms
vmware-aiops alarm acknowledge esxi-01 "Host memory usage"             # Acknowledge alarm
vmware-aiops alarm reset esxi-01 "Host memory usage"                   # Clear alarms (2x confirm; clears ALL matching entity type + status)

# Datastore (browse and scan only — iSCSI/vSAN moved to vmware-storage)
vmware-aiops datastore browse datastore1 --path "iso/"                 # Browse datastore
vmware-aiops datastore scan-images --target home-esxi                  # Scan all datastores for images

# Scan
vmware-aiops scan now              # One-time scan of alarms and events (host logs: daemon only)

# Daemon
vmware-aiops daemon start          # Start scanner
vmware-aiops daemon status         # Check status
vmware-aiops daemon stop           # Stop daemon

# Companion skills for other operations:
#   vmware-monitor: inventory, alarms, events, sensors
#   vmware-storage: datastores, iSCSI, vSAN
#   vmware-vks:     Tanzu/TKC cluster lifecycle

Configuration

See config.example.yaml for all options.

SectionKeyDefaultDescription
targetsname—Friendly name
targetshost—vCenter/ESXi hostname or IP
targetstypevcentervcenter or esxi
targetsport443Connection port
targetsverify_ssltrueVerify the target's TLS certificate (set false only for self-signed lab hosts)
scannerinterval_minutes15Scan frequency
scannerseverity_thresholdwarningMin severity: critical/warning/info
scannerlookback_hours1How far back to scan
scannerlog_types—Not read by any code — the daemon always reads the hostd, vmkernel and vpxa host logs. Setting it changes nothing
notifylog_file~/.vmware-aiops/scan.logJSONL log output
notifywebhook_url—Webhook endpoint (Slack, Discord, etc.)

Project Structure

code
VMware-AIops/
├── skills/                        # Skills index (npx skills add)
│   └── vmware-aiops/
│       ├── SKILL.md               # Slimmed-down skill (progressive disclosure)
│       └── references/            # Detailed docs loaded on-demand
│           ├── capabilities.md    # Full capabilities tables
│           ├── cli-reference.md   # Complete CLI reference
│           └── setup-guide.md     # Install, security, AI platforms
├── vmware_aiops/                  # Python backend
│   ├── config.py                  # YAML + .env config
│   ├── connection.py              # Multi-target pyVmomi
│   ├── cli/                       # Typer CLI (double confirm)
│   ├── ops/                       # Operations
│   │   ├── inventory.py           # VMs, hosts, datastores, clusters
│   │   ├── health.py              # Alarms, events, sensors
│   │   ├── vm_lifecycle.py        # VM CRUD, snapshots, clone, migrate
│   │   ├── vm_deploy.py           # OVA, template, linked clone, batch deploy
│   │   └── datastore_browser.py   # Datastore browsing, image discovery
│   ├── scanner/                   # Log scanning daemon
│   ├── notify/                    # Notifications (JSONL + webhook)
│   └── mcp_server/                # MCP server wrapper
│       ├── server.py              # FastMCP server with tools
│       └── __main__.py
├── examples/mcp-configs/          # MCP client config templates
├── tests/                         # Test suite
├── smithery.yaml                  # Smithery marketplace config
├── RELEASE_NOTES.md
├── config.example.yaml
└── pyproject.toml

API Coverage

Built on pyVmomi (vSphere Web Services API / SOAP).

API ObjectUsage
vim.VirtualMachineVM lifecycle, snapshots, clone, migrate
vim.HostSystemESXi host info, sensors, services
vim.DatastoreStorage capacity, type, accessibility
vim.host.DatastoreBrowserFile browsing, image discovery (ISO/OVA/VMDK)
vim.OvfManagerOVA import and deployment
vim.ClusterComputeResourceCluster, DRS, HA
vim.NetworkNetwork listing
vim.alarm.AlarmManagerActive alarm monitoring
vim.event.EventManagerEvent/log queries

Related Projects

SkillScopeToolsInstall
vmware-aiopsVM lifecycle, deployment, guest ops, cluster, datastore browse, triage49uv tool install vmware-aiops
vmware-monitorRead-only monitoring, alarms, events, investigation bundles27uv tool install vmware-monitor
vmware-storageDatastores, iSCSI, vSAN11uv tool install vmware-storage
vmware-vksTanzu Namespaces, TKC cluster lifecycle20uv tool install vmware-vks
vmware-nsxNSX networking: segments, gateways, NAT, routing, IPAM33uv tool install vmware-nsx-mgmt
vmware-nsx-securityDFW policies/rules, security groups, Traceflow, IDS/IPS21uv tool install vmware-nsx-security
vmware-ariaAria Operations metrics, alerts, capacity, anomalies28uv tool install vmware-aria
vmware-aviAVI (NSX ALB) load balancing, AKO Kubernetes ops28uv tool install vmware-avi
vmware-hardenCompliance baselines (CIS / vSphere SCG / 等保 / PCI-DSS), drift detection6uv tool install vmware-harden

Troubleshooting & Contributing

If you encounter any errors or issues, please send the error message, logs, or screenshots to zhouwei008@gmail.com. Contributions are welcome — feel free to join us in maintaining and improving this project!

License

MIT

常见问题

VMware AIops 是什么?

基于 AI 的 VMware vCenter/ESXi 监控与运维工具集,内含 20 个 MCP 工具,帮助排障、巡检和自动化操作。

相关 Skills

更新日志

by alirezarezvani

Universal
热门

基于 Conventional Commits 自动解析提交记录、判断语义化版本升级并生成规范 changelog,适合在 CI、发版前检查提交格式并批量输出可审计发布说明。

✎ 自动生成和管理更新日志与发布说明,帮团队把版本变更说清楚;聚焦版本化与流程自动化,省时又更规范。

DevOps
未扫描26.0k

可观测性设计

by alirezarezvani

Universal
热门

面向生产系统规划可落地的可观测性体系,串起指标、日志、链路追踪与 SLI/SLO、错误预算、告警和仪表盘设计,适合搭建监控平台与优化故障响应。

✎ 把监控、日志、链路追踪串起来,帮助团队从设计阶段构建可观测性,排障更快、系统演进更稳。

DevOps
未扫描26.0k

环境密钥管理

by alirezarezvani

Universal
热门

统一梳理dev/staging/prod的.env和密钥流程,自动生成.env.example、校验必填变量、扫描Git历史泄漏,并联动Vault、AWS SSM、1Password、Doppler完成轮换。

✎ 统一管理环境变量、密钥与配置,减少泄露和部署混乱,安全治理与团队协作一起做好,DevOps 场景很省心。

DevOps
未扫描26.0k

相关 MCP Server

kubefwd

编辑精选

by txn2

热门

kubefwd 是让 AI 帮你批量转发 Kubernetes 服务到本地的开发神器。

✎ 微服务开发者最头疼的本地调试问题,它一键搞定——自动分配 IP 避免端口冲突,还能用自然语言查询状态。但依赖 AI 工作流,纯命令行爱好者可能觉得不够直接。

DevOps
4.2k

Cloudflare

编辑精选

by Cloudflare

热门

Cloudflare MCP Server 是让你用自然语言管理 Workers、KV 和 R2 等云资源的工具。

✎ 这个工具解决了开发者频繁切换控制台和文档的痛点,特别适合那些在 Cloudflare 上部署无服务器应用、需要快速调试或管理配置的团队。不过,由于它依赖多个子服务器,初次设置可能有点繁琐,建议先从 Workers Bindings 这类核心功能入手。

DevOps
4.1k

Terraform

编辑精选

by hashicorp

热门

Terraform MCP Server 是让 AI 助手直接操作 Terraform Registry 和 HCP Terraform 的桥梁。

✎ 如果你经常在 Terraform 里翻文档找模块配置,这个服务器能省不少时间——直接问 Claude 就能生成准确的代码片段。最适合管理多云基础设施的团队,但注意它目前只适合本地使用,别在生产环境里暴露 HTTP 端点。

DevOps
1.5k

评论