Latency benchmarks
The measurements below have no recorded benchmark date, CLI/model versions, or raw result artifact. They are historical reported values, not measurements of the current release. Use the method to compare your own flows before sizing a CI timeout. Runtime depends on how a test is written and the state of the system under test. The measurements below describe one login flow and a set of preset actions; they are not a latency guarantee for other suites. Step caching can avoid repeated model calls for stable steps.Step latency
The following historical ranges describe work that required a fresh model
resolution, evaluation, or generation. Cached steps can avoid repeated model
work on later runs:
Momentic vs. Playwright
We benchmarked Momentic against Playwright using the same login flow on Practice Test Automation’s practice login page. In the reported P50 results, cached Momentic steps took 212ms longer in total than the comparable Playwright steps. The first run with four fresh AI completions took 26,379ms. Results depend on the browser, model, app, and runner. This benchmark covers the login steps shown below. It does not measure AI assertions, visual comparisons, or other flows. The final checks also differ: the Momentic example searches page HTML for a substring, while the Playwright example waits for visible text.Method
We built a Momentic test and an equivalent Playwright script that perform the same login flow. We measured three execution modes:- Steps only measures only the time spent executing steps.
- End-to-end includes Momentic’s fixed bootstrap and result-upload time. The Playwright measurement includes CLI initialization but no result upload.
- First-run disables caching and includes four fresh AI completions.
Results
All values are P50 milliseconds.
The benchmark sources:
Parallelism and sharding
By default, the CLI runs tests one at a time on a single machine. Use parallelism, sharding, or both to reduce the wall-clock time of a large suite.Parallelism on one machine
--parallel <n> runs n tests at once, each in its own browser instance. Raise
it until the runner’s CPU or memory is saturated.
Sharding across machines
--shard-count splits the suite across runners, and --shard-index selects the
current runner’s shard. Merge the outputs with momentic results merge before
uploading.
--parallel within each runner.
Make sure step caching is saving in CI. Cached steps
can avoid repeated model work, but application readiness and network requests
still affect runtime. See
Saving caches for setup.
Mobile test suites
Mobile runs use the same--parallel, --shard-count, and --shard-index
flags through momentic-mobile, with a few platform constraints:
--parallel AUTOsaturates the current shard. On remote emulators, each test gets its own session and organization quotas are enforced server-side.- Local Android runs need a distinct AVD per parallel test. Multiple tests
cannot share the same
--local-avd-id. - Local iOS runs should use
--parallel 1because concurrent sessions share a device-driver manifest.
momentic-mobile results merge
before uploading. See the
momentic-mobile run reference
for every flag.