DeskVNC agent hub
Benchmarks

19 ms from seeing a screen to acting on it

These are the project's own measurements, taken on real hardware rather than a mock. The machine, the conditions and the meaning of each number are all below, so you can judge whether it fits the loop you intend to write.

Methodology

What was measured, and on what

What a LAN means for you

The wire between the machine running the agent and the target is an ordinary network connection, so the per call costs scale with the path between them. The shape of the loop is the same wherever the agent runs: a screen back, an action out, one settled result in.

The numbers

Every measured call

Agent plane measurements, 1920x1080 Windows desktop over a LAN
Call Time What it covers
dvv_open and attach 4 ms Spawning the session against a saved machine and attaching to it.
dvv_control acquire under 1 ms Taking the control lease. dvv_status measures about the same.
dvv_screen at scale: 0.25 25 ms A 112 KB image of the whole 1920x1080 framebuffer.
dvv_screen at full scale 70 ms A 1.5 MB image of the same frame.
One observe then act cycle 19 ms Look, act, settle. About 52 actions per second.
dvv_type at wpm: 12000 447 chars/s Typing throughput, one key event pair per Unicode code point.

The full table, with the surrounding notes, is in the project README.

In practice

Choosing the settings a loop should run at

Scale the screen, and the picture shrinks with it

Use scale: 0.25 unless you need to read small text. That is 25 ms and 112 KB instead of 70 ms and 1.5 MB, and the loop still has everything a decision needs on it. When you do need a region, ask for one at native resolution rather than asking for an image somebody will resize, because then the scale factor is one you did not choose and cannot invert.

MCP the loop
dvv_screen {"limbId": "...", "form": "full", "scale": 0.25}
dvv_click  {"limbId": "...", "x": 700, "y": 400, "generation": 1}
dvv_wait   {"limbId": "...", "until": "screen-stable", "quietMs": 750}
dvv_screen {"limbId": "...", "form": "damage-crop"}   # only what changed

damage-crop is the cheapest useful answer and the one to use after an action: it reads only what moved.

Raise wpm for throughput, not for speed

wpm is a correctness control rather than a politeness one. Neither wire carries an acknowledgement, so a machine that drops characters under a fast synthetic type does it silently. Raise it when you want throughput, and keep it moderate when the machine is remote enough that the input path is the one under test.

Ask for the mirror only where you need one

dvv_open with perceive: true attaches a framebuffer mirror, which is what dvv_screen reads. It is off by default, and for good arithmetic: eight 4K mirrors cost 264 MB before anything is decoded. Open mirrors on the limbs you will look at, and leave terminal limbs without one, because a terminal limb answers with text.

Scale

Ten machines are ten loops

Every open machine is its own limb with its own lease. There is no shared queue and no shared lock, so an agent runs as many independent loops as it has machines, and a group runs one action on every member at once, concurrently, rather than in a loop.

MCP one action, every member
dvv_group_open {"hostIds": ["<a>", "<b>", "<c>", "<d>", "<e>",
                               "<f>", "<g>", "<h>", "<i>", "<j>"],
                 "perceive": true}
dvv_group_run  {"groupId": "...",
                "action": "run",
                "arguments": {"command": "uptime", "timeoutMs": 5000}}
dvv_group_close {"groupId": "..."}
Two desktops at once

One agent holding two desktops and typing a different sum into a calculator on each finished both in 0.95 seconds. Every member starts before any finishes, and one member failing is reported for that member alone without stopping the others.

Guide 7 has the group calls in full, including growing and shrinking a set.

Read next

Put the numbers to work