LLM agent quickstart: drive a real desktop in one sitting
This is the shortest path from a clean machine to an AI agent that opens a real desktop through DeskVNC. Every command below is meant to be copied verbatim. The expected output after each step is described, so an agent can self check before moving on.
The agent here is Claude Code. The dvv server is the same MCP server for OpenCode, Codex CLI, Pi, Gemini CLI, Cursor and the rest, so the only thing that changes for another client is the registration line.
Step 1: get the binary
Download the build for the agent host from the latest DeskVNC release.
| Platform | File |
|---|---|
| macOS 12 or newer | DeskVNCViewer_<version>_universal.dmg |
| Windows 10 or 11 | DeskVNCViewer_<version>_x64-setup.exe |
| Debian, Ubuntu | DeskVNCViewer_<version>_amd64.deb |
| Fedora, RHEL | DeskVNCViewer-<version>-1.x86_64.rpm |
| Any x64 Linux | DeskVNCViewer_<version>_amd64.AppImage |
On macOS, open the DMG and drag the app to Applications. The build is signed with a Developer ID and notarized by Apple, so it opens normally without a "damaged and can't be opened" dialog. On Windows, the installer is signed; if SmartScreen asks once, choose More info, then Run anyway. The publisher line should read Open Source Developer Godwin Josh.
Expected result: a working DeskVNCViewer icon in the launcher. The bundled dvv executable is what registers with the agent client in the next step. From v0.27.5 on macOS, and v0.27.6 on Windows and Linux, that executable is shipped inside the install. Earlier releases omitted it on some platforms; if the AI Agents panel reports no bundled executable, install a newer build or build dvv from source with cargo build -p dvv and use that absolute path in step 3.
Step 2: open the app once and save a machine
Start DeskVNCViewer. The first launch walks through two permission prompts on macOS: Local Network (for mDNS discovery) and Accessibility (only if global input capture is turned on). Both can be declined or accepted later without breaking the agent.
In the app, press New Host and save one machine with its password, or paste an address straight into the bar at the top and press Connect, which saves nothing. Double click a saved tile to confirm the connection works. Closing the viewer is fine: the agent will start it again.
Expected result: at least one tile in the host library, with a live thumbnail, and a credential in the operating system keychain.
Step 3: switch the agent plane on
Open the AI Agents panel in DeskVNCViewer and switch the plane on. It is off by default and stays off until a person turns it on, so this step is the gate. The app then scans the local machine for installed agents and points each one at the bundled dvv binary.
Expected result: a list of detected agents in the panel, each checked and ready.
Step 4: register the MCP server
For Claude Code, the app's AI Agents panel offers a one click Register with Claude Code button that runs the right claude mcp add line. For everything else, run the line by hand. The exact command for each platform is:
# macOS
claude mcp add --scope user deskvnc -- /Applications/DeskVNCViewer.app/Contents/MacOS/dvv mcp --stdio
# Windows (run in PowerShell or cmd)
claude mcp add --scope user deskvnc -- "%LOCALAPPDATA%\DeskVNCViewer\dvv.exe" mcp --stdio
# Linux
claude mcp add --scope user deskvnc -- /usr/bin/dvv mcp --stdioOpenCode, Codex CLI, Cursor, VS Code, Windsurf and Gemini CLI take the same shape: a command plus arguments, registered in their mcp config. The dvv setup command writes all of them at once and also drops the agent skill into each client's skills directory.
Expected result from claude mcp list after restart: a row named deskvnc with command /Applications/DeskVNCViewer.app/Contents/MacOS/dvv mcp --stdio (or the platform equivalent) and a green status. OpenCode reports the same in its MCP panel.
Step 5: the first loop
Restart the agent so it picks up the new MCP server, then ask it to do something on the saved machine by name. The first four calls of the loop, in order:
dvv_hosts {} // what there is to open
dvv_open {"hostId": "<id>", "perceive": true} // -> limbId, size, state
dvv_control {"limbId": "...", "action": "acquire"}
dvv_screen {"limbId": "...", "form": "full", "scale": 0.25}After the first screen read, the agent has a generation and a size. The next two calls are the observe then act pair the benchmark numbers in the README are measured against:
dvv_click {"limbId": "...", "x": 700, "y": 400, "generation": 1}
dvv_screen {"limbId": "...", "form": "damage-crop"} // look again
dvv_type {"limbId": "...", "text": "notepad", "wpm": 3000}
dvv_key {"limbId": "...", "keys": "meta+r"}A fresh dvv_screen on a primed mirror can answer with a request for one more full refresh. That is ordinary priming behaviour, not an error. Call dvv_screen again with form: "full" and carry on.
Expected result: a typed string, a key chord, and a settled screen, all inside one observe-then-act cycle. The figures the project publishes are 4 ms for dvv_open and attach, under 1 ms for dvv_control acquire, 25 ms for dvv_screen at scale: 0.25, and 19 ms for one full cycle, about 52 actions per second. wpm is a throttle in words per minute, so omitting it types as fast as the wire carries; dvv_type runs at 447 characters per second. A first call on a fresh mirror is slower by the priming handshake.
Step 6: read refusal codes as designed behaviour
The loop is fenced in two places, and both fences are features, not failures. The agent meets them on the first task; the right move is to recognise them and act.
GEOMETRY_CHANGED: a click carries a generation read from a prior dvv_screen. If the screen resized between the read and the click, the call is refused with GEOMETRY_CHANGED and nothing is delivered. The fix is one more dvv_screen and one more recomputation of the coordinate. A click never lands somewhere unintended, and the agent never has to ask whether it did.
SCREEN_CHANGED: dvv_type and dvv_key are refused with SCREEN_CHANGED when something window sized has repainted since the last dvv_screen, and refused outright on a limb that has never been read at all. A keystroke carries no coordinate, so it lands wherever focus is, and focus moves when a dialog or an application appears. One dvv_screen clears the fence, and dvv_status does not (it reads no pixels, deliberately). There is no override. The fix is to call dvv_screen, read what has focus and what is selected, and only then type.
LEASE_REVOKED: a person took the wheel, which is the intended behaviour. The right move is to stop, not to retry. Call dvv_control with action: "yield_status" and read lease.human_took_over. If it is true, a person is driving now and the agent should tell the user what it was doing, then stand down until the human hands control back. Held keys and buttons were already released when the lease went away, so a half finished drag cannot strand the desktop.
There is also a fourth ordinary result, not a refusal: a dvv_wait that times out. It returns settled: false and the observation, never an error. Call again or read dvv_status to see whether the limb is still connected.
The four codes an agent sees on a real loop are the only codes it needs to know by name. Every other code the manifest can raise arrives with a sentence that names the next step, so the agent reads the message rather than the code in those cases.
Step 7: close the limb
When the task is done, close the limb so the next call does not pay for an idle connection:
dvv_close {"limbId": "..."}dvv_close is idempotent. Closing a limb that is already gone returns an ordinary success. Every intent still in flight is withdrawn and settles first.
Shell equivalent
For agents without MCP tools, or for an operator working alongside the agent, the same loop is the dvv command line. After dvv setup has wired the binary onto PATH, the loop reads:
dvv hosts
dvv limbs
dvv open <name or hostId> --perceive
dvv wait <limbId> --until connected
dvv control acquire <limbId>
dvv screen <limbId> --scale 0.5 --out ./dvv-screen.png
dvv click <limbId> <x> <y>
dvv click <limbId> <x> <y> --action double
dvv type <limbId> "text to type"
dvv key <limbId> super+r
dvv wait <limbId> --until screen-stable
dvv reconnect <limbId>
dvv close <limbId>dvv screen writes the picture to disk and prints the path, which Pi's read tool opens as an image. The shell path runs the same dispatch table as the MCP server, so a call behaves identically regardless of how the agent reached it.