DeskVNC DeskVNC

Registering DeskVNC with Gemini CLI

Gemini CLI is Google's terminal agent, and it speaks the Model Context Protocol. Pairing it with DeskVNC's dvv lets the agent open a saved machine, read the screen, and send clicks and keystrokes through the VNC, RDP, or SSH protocol the remote already speaks. No software is installed on the remote target, and the user can take the wheel back with one click. The DeskVNC integration notes document the exact shape that ~/.gemini/settings.json takes for the DeskVNC MCP server, including the httpUrl key the HTTP transport expects.

Prerequisite: the dvv binary

dvv is the executable that ships inside DeskVNCViewer. Install the viewer, open it once, and save a machine with its password so the agent has something concrete to reach. The binary sits next to the application:

The path you point Gemini CLI at is the path the viewer installed.

The registration step

Gemini CLI reads ~/.gemini/settings.json. The DeskVNC integration notes call out that the same mcpServers block used by other clients works here, with one difference for the HTTP transport: the key is httpUrl rather than url. On macOS:

{
  "mcpServers": {
    "deskvnc": {
      "command": "/Applications/DeskVNCViewer.app/Contents/MacOS/dvv",
      "args": ["mcp", "--stdio"]
    }
  }
}

On Linux, the same block with the system path:

{
  "mcpServers": {
    "deskvnc": {
      "command": "/usr/bin/dvv",
      "args": ["mcp", "--stdio"]
    }
  }
}

On Windows, use a JSON-safe path. Forward slashes keep the file portable:

{
  "mcpServers": {
    "deskvnc": {
      "command": "C:/Program Files/DeskVNCViewer/dvv.exe",
      "args": ["mcp", "--stdio"]
    }
  }
}

The top-level key is mcpServers. That is the key the DeskVNC integration notes name for ~/.gemini/settings.json, and it matches the standard MCP shape Gemini CLI reads.

For the HTTP transport, swap command and args for a block whose transport key is httpUrl (per the DeskVNC integration notes) and whose headers carry the bearer token:

{
  "mcpServers": {
    "deskvnc": {
      "httpUrl": "http://127.0.0.1:7333/mcp",
      "headers": { "Authorization": "Bearer <token>" }
    }
  }
}

The shell form is dvv setup. It writes the same MCP entry, copies the skill into the skills directory Gemini CLI reads, and points the agent at the right dvv. After the file is saved, restart Gemini CLI so it picks up the new server, then run /mcp in a session to confirm deskvnc is connected.

A first agent session

A typical first prompt names a saved machine and asks the agent to do something visible:

You: Open "qa-runner" and tell me the result of the last test run.

The agent drives the four-call loop. Behind the scenes, Gemini CLI forwards each MCP call to the dvv subprocess:

dvv_hosts   {}
dvv_open    {"hostId": "qa-runner", "perceive": true}
dvv_control {"limbId": "limb-13ac", "action": "acquire"}
dvv_screen  {"limbId": "limb-13ac", "form": "full", "scale": 0.25}

The screenshot comes back with a geometry_generation of, say, 12. The agent reads the picture, finds the test runner window, and acts:

dvv_click   {"limbId": "limb-13ac", "x": 240, "y": 96, "generation": 12}
dvv_type    {"limbId": "limb-13ac", "text": "pytest -q", "wpm": 3000}
dvv_key     {"limbId": "limb-13ac", "keys": "Return"}
dvv_screen  {"limbId": "limb-13ac", "form": "damage-crop"}

The follow-up screenshot confirms the test run completed, and the agent reports the pass or fail count. Each observe-then-act cycle is well under a second, so the agent keeps up with a moving desktop.

Machines are independent. Each saved machine is its own limb with its own lease, so an agent runs as many loops in parallel as it has machines. The same action across a set of machines is one dvv_group_run call that reports each member's outcome separately, and any single-limb tool accepts a groupId plus a member to address one of them.

The coordinate generation fence

Every dvv_screen carries two fences: a geometry_generation and a content_generation. Clicks must echo the geometry generation they were computed against. A click computed against a screen that has since resized is refused rather than landing in the wrong place. Typing and keys are fenced by the content generation, and dvv_type plus dvv_key are refused with SCREEN_CHANGED when something window-sized has repainted since the last dvv_screen. The fence is the safety net that stops an agent from typing into a dialog that appeared over the editor it was aiming at. One dvv_screen clears the fence, dvv_status does not (it reads no pixels), and there is no override. Terminal limbs are not fenced, because a PTY echoes what it is sent into a stream the agent reads back.

Human takeover

The person at the remote machine can take control back at any moment. The viewer shows an "agent driving" badge while the agent holds the lease, and one click in the viewer hands the desktop back to the person. The plane releases every held key and every held button the moment a person takes over, so a half-finished drag cannot strand the desktop. The recipient can revoke the agent's control or end the session outright. This is attended automation by design: the agent never has a stronger position than the person at the keyboard.

Troubleshooting

Symptom Likely cause Try this
LIMB_GONE from any tool The viewer closed the connection or the last dvv_close detached the limb Run dvv limbs to list live limbs, then dvv open <host> again
SCREEN_CHANGED on dvv_type or dvv_key A window-sized region repainted since the last dvv_screen Call dvv_screen once, then retry the action
Gemini CLI does not list deskvnc in /mcp JSON parse error, wrong file path, or wrong top-level key Validate the JSON, confirm the path is ~/.gemini/settings.json, and check the current Gemini CLI MCP docs for the expected key name
dvv exits immediately on spawn The path in command is wrong or the file is not executable Run the command value in a shell and confirm dvv mcp --stdio starts
HTTP transport returns 401 The bearer token changed Re-export DVV_MCP_TOKEN to the same value, then restart dvv
Click lands in the wrong place Stale geometry_generation after a resize Take a fresh dvv_screen, then send the new generation with the click