How the DeskVNC agent benchmark was measured
A benchmark number without a method is a slogan. A number with a method is a fact. This page is the method behind the numbers the DeskVNC project publishes for the dvv agent control plane, and the contract a third party would follow to reproduce them. It is written so an infrastructure engineer can read it, disagree with it, and rerun it.
Two rules govern what follows. First, every number on this page is from the project's published benchmarks, measured against a real 1920x1080 Windows desktop over a LAN, reproduced from the project README. Second, the methodology is described at the level the project commits to: a description of the protocol, the environment that protocol was run in, and the variables a reproducer would control. Where the project has not published a specific figure, this page says so plainly. Omission is honest. Invention is not.
The numbers
| Call | Result | Unit | Measurement point |
|---|---|---|---|
dvv_open and attach |
4 | milliseconds | wall clock from dvv open to first screen ready |
dvv_control acquire |
under 1 | millisecond | wall clock of the acquire RPC |
dvv_screen at scale: 0.25 |
25 | milliseconds | end to end, MCP dispatch to bytes returned |
dvv_screen at full scale (1.5 MB) |
70 | milliseconds | end to end, MCP dispatch to bytes returned |
| one observe then act cycle | 19 | milliseconds | screen at 0.25, click, screen at damage crop |
| observe then act rate | about 52 | actions per second | reciprocal of the cycle time |
dvv_type at wpm: 12000 |
447 | characters per second | wall clock throughput of the type call |
Two things to notice in the table. The numbers are wall clock on the agent's side, not the protocol's wire time alone. They are the cost an agent actually pays per call, including the MCP dispatch, the framing, the protocol round trip, and the bytes crossing the IPC channel on the way back. A number that measured only the wire would be smaller. A number that measures wall clock is the one that fits inside an agent's budget.
What an observe then act cycle is
The agent's inner loop is four calls deep, but the unit that matters for throughput is the second pair: look, then act.
dvv_screen {"form": "full", "scale": 0.25} // look at the whole screen
dvv_click {"x": 700, "y": 400, "generation": 1} // click on what was seen
dvv_screen {"form": "damage-crop"} // look at what changedThe first dvv_screen is what an agent uses when it does not know which part of the screen matters. The second dvv_screen, with form: "damage-crop", asks the server for the rectangle that has changed since the previous frame. That is the path an agent takes on every subsequent step, and the path the 19 ms figure measures.
Why 19 ms matters is the question that decides whether the number is fast. An agent's outer loop is a language model call, which costs hundreds of milliseconds at the low end and seconds at the high end. An agent's inner loop is observe then act, and the cost of that inner loop is the tax the agent pays on every decision. At 19 ms the model can decide and act dozens of times per second. Slow the loop down by an order of magnitude and the model starts waiting on the screen the way a person waits on a slow session. At 19 ms the loop is fast enough to feel like a single thought. Slower, and it feels like a tool.
The comparison that matters is against the round trip a person perceives between clicking and seeing the next paint in an ordinary interactive session. The agent's inner loop at 19 ms is inside that budget. That is the point: the agent is not pretending to be a person. It runs a loop the person cannot run.
The measurement protocol
The numbers above were taken on a real 1920x1080 Windows desktop reached over a LAN. The release does not publish a fixed reference configuration, so the protocol below is the one a reproducer would follow, not a specific make and model.
Hardware
The agent host and the Windows desktop are two physical machines on the same LAN, not a single host with a virtual machine. The Windows desktop runs the stock Windows desktop at 1920x1080 with display scaling at 100 percent. The agent host runs DeskVNCViewer and the agent process. The project publishes per run numbers with each release so a reader can compare their own machine against the published distribution.
Network
The two machines are on a single LAN segment, with no proxy between the agent and the desktop. The connection goes over the protocol the operator chose, typically VNC, RDP or SSH, on the standard port. The 19 ms figure is the LAN figure, not the WAN figure.
Resolution and scale
The Windows desktop is 1920x1080. A dvv_screen at full scale returns a 1.5 MB picture, which is the 70 ms row. The agent's inner loop does not run at full scale. It runs at scale: 0.25, which returns about 112 KB, and that is the 25 ms row. The reason is in the "coordinate scaling" section below.
What was measured, how
Each call was timed from the moment the MCP call entered dvv to the moment the response left. For a dvv_screen, that is from dispatch to the last byte of the picture. For a dvv_click, that is from dispatch to the acknowledgement. The observe then act cycle is timed as a single end to end measurement: dispatch the first dvv_screen, dispatch the dvv_click against the generation the first call returned, and time the second dvv_screen to its last byte. The 19 ms figure is the median of that triple over many runs.
Sample counts and variance
The project has not published a specific sample count or variance figure for these benchmarks. The release does publish the raw per run numbers, which is what a reader should compare against. A median over many runs is the headline, and the raw distribution is what tells the reader whether their own environment will sit close to or far from the median. This page does not invent a confidence interval the project has not measured.
Warm versus cold
A first dvv_screen on a fresh session can answer "the mirror is priming, send a full refresh and read again". The benchmarks above are warm numbers. A cold start includes that priming handshake, and is slower for the first one or two calls. The 4 ms attach figure is the warm attach, after the first connection. The first connection on a fresh desktop is slower, by an amount the project has not separately published.
Coordinate scaling
dvv_screen prints an imageSpace line that describes the geometry of the picture the agent is looking at. With scale: 0.5, a point at (mx, my) on the picture is (mx times 2, my times 2) on the machine. With scale: 0.25, a point at (mx, my) on the picture is (mx times 4, my times 4) on the machine. The agent must convert before every click, and a click that does not carry the conversion is refused on a geometry mismatch.
Reduced scale is what makes the loop fast for two compounding reasons. First, the picture is smaller, so the wire transfer, the IPC framing and the decode on the agent host all cost less. The 1.5 MB full frame at 70 ms becomes 112 KB at 25 ms at a quarter scale, and most of the saving is in the picture's bytes. Second, the smaller picture is what most agent decisions need. An agent looking for a button does not need to read serif type at body size. It needs to know that a roughly button shaped region exists at roughly the right coordinates. A quarter scale picture is the right resolution for that decision, and the wrong resolution for reading a serial number off a label. The right reflex is to start at scale: 0.25, raise the scale when the agent needs to read small text, and drop back when the read is over.
What the release publishes
Every release ships the benchmark numbers, the protocol the numbers were taken on, the resolution of the target desktop, and the raw per run distribution. The headline numbers above are the medians of those distributions. A reader who wants to know whether their own environment will sit close to the median can pull the same release, run the same calls against their own desktop, and compare. The distribution is the contract. The median is the summary.
Why citable methodology matters
Numbers with a method get quoted. Numbers without a method get ignored, because the next reader cannot tell whether the number is portable, or honest. The discipline of publishing the method alongside the result, and the raw data alongside the median, is the difference between a benchmark that travels and a benchmark that stays in the project that produced it. The DeskVNC agent control plane publishes the performance of the no install path to the machines an enterprise actually runs, in enough detail to be reproduced. That detail is the reason the number travels.