- 19 ms One observe then act cycle, start to settled result
- 52 actions/s What that cycle rate works out to
- 4 ms Open a saved machine and attach to it
- 447 chars/s Typing throughput at wpm 12000
What was measured, and on what
- The machine: a real 1920x1080 Windows desktop, not a synthetic source and not a mock.
- The network: a LAN, so the figures are connection time rather than round trip time across a WAN.
- What a cycle is: the full loop: observe the screen, then act, and receive one settled result for the action. Not a single call, and not fire and forget.
- Where the figures come from: the project measures them as part of its own work, and publishes them in the README. They are reproduced here unchanged.
The wire between the machine running the agent and the target is an ordinary network connection, so the per call costs scale with the path between them. The shape of the loop is the same wherever the agent runs: a screen back, an action out, one settled result in.
Every measured call
| Call | Time | What it covers |
|---|---|---|
dvv_open and attach |
4 ms | Spawning the session against a saved machine and attaching to it. |
dvv_control acquire |
under 1 ms | Taking the control lease. dvv_status measures about the same. |
dvv_screen at scale: 0.25 |
25 ms | A 112 KB image of the whole 1920x1080 framebuffer. |
dvv_screen at full scale |
70 ms | A 1.5 MB image of the same frame. |
| One observe then act cycle | 19 ms | Look, act, settle. About 52 actions per second. |
dvv_type at wpm: 12000 |
447 chars/s | Typing throughput, one key event pair per Unicode code point. |
The full table, with the surrounding notes, is in the project README.
Choosing the settings a loop should run at
Scale the screen, and the picture shrinks with it
Use scale: 0.25 unless you need to read small text. That is 25 ms and 112 KB instead of 70 ms and 1.5 MB, and the loop still has everything a decision needs on it. When you do need a region, ask for one at native resolution rather than asking for an image somebody will resize, because then the scale factor is one you did not choose and cannot invert.
dvv_screen {"limbId": "...", "form": "full", "scale": 0.25}
dvv_click {"limbId": "...", "x": 700, "y": 400, "generation": 1}
dvv_wait {"limbId": "...", "until": "screen-stable", "quietMs": 750}
dvv_screen {"limbId": "...", "form": "damage-crop"} # only what changed
damage-crop is the cheapest useful answer and the one to use after an action: it reads only what moved.
Raise wpm for throughput, not for speed
wpm is a correctness control rather than a politeness one. Neither wire carries an acknowledgement, so a machine that drops characters under a fast synthetic type does it silently. Raise it when you want throughput, and keep it moderate when the machine is remote enough that the input path is the one under test.
Ask for the mirror only where you need one
dvv_open with perceive: true attaches a framebuffer mirror, which is what dvv_screen reads. It is off by default, and for good arithmetic: eight 4K mirrors cost 264 MB before anything is decoded. Open mirrors on the limbs you will look at, and leave terminal limbs without one, because a terminal limb answers with text.
Ten machines are ten loops
Every open machine is its own limb with its own lease. There is no shared queue and no shared lock, so an agent runs as many independent loops as it has machines, and a group runs one action on every member at once, concurrently, rather than in a loop.
dvv_group_open {"hostIds": ["<a>", "<b>", "<c>", "<d>", "<e>",
"<f>", "<g>", "<h>", "<i>", "<j>"],
"perceive": true}
dvv_group_run {"groupId": "...",
"action": "run",
"arguments": {"command": "uptime", "timeoutMs": 5000}}
dvv_group_close {"groupId": "..."}
One agent holding two desktops and typing a different sum into a calculator on each finished both in 0.95 seconds. Every member starts before any finishes, and one member failing is reported for that member alone without stopping the others.
Guide 7 has the group calls in full, including growing and shrinking a set.
Put the numbers to work
-
The four call loop
Read the library, open a machine, take the wheel, look, act, look again. Every argument in the manifest, with the four fences the plane enforces.
-
An agent that reads a screen and acts
The situation this is measured for, and the calls that carry it out end to end.