Skip to main content
How long a Momentic step takes depends on the step type and whether the step cache is warm. Measure your own flows when setting CI timeouts. This page keeps an earlier login-flow benchmark for its method and shows the flags for running tests concurrently.

Latency benchmarks

The measurements below have no recorded benchmark date, CLI/model versions, or raw result artifact. They are historical reported values, not measurements of the current release. Use the method to compare your own flows before sizing a CI timeout. Runtime depends on how a test is written and the state of the system under test. The measurements below describe one login flow and a set of preset actions; they are not a latency guarantee for other suites. Step caching can avoid repeated model calls for stable steps.

Step latency

The following historical ranges describe work that required a fresh model resolution, evaluation, or generation. Cached steps can avoid repeated model work on later runs:

Momentic vs. Playwright

We benchmarked Momentic against Playwright using the same login flow on Practice Test Automation’s practice login page. In the reported P50 results, cached Momentic steps took 212ms longer in total than the comparable Playwright steps. The first run with four fresh AI completions took 26,379ms. Results depend on the browser, model, app, and runner. This benchmark covers the login steps shown below. It does not measure AI assertions, visual comparisons, or other flows. The final checks also differ: the Momentic example searches page HTML for a substring, while the Playwright example waits for visible text.

Method

We built a Momentic test and an equivalent Playwright script that perform the same login flow. We measured three execution modes:
  • Steps only measures only the time spent executing steps.
  • End-to-end includes Momentic’s fixed bootstrap and result-upload time. The Playwright measurement includes CLI initialization but no result upload.
  • First-run disables caching and includes four fresh AI completions.
Measurements ran on an M3 Max MacBook Pro with 36GB RAM running macOS Sonoma.

Results

All values are P50 milliseconds. The benchmark sources:

Parallelism and sharding

By default, the CLI runs tests one at a time on a single machine. Use parallelism, sharding, or both to reduce the wall-clock time of a large suite.

Parallelism on one machine

--parallel <n> runs n tests at once, each in its own browser instance. Raise it until the runner’s CPU or memory is saturated.

Sharding across machines

--shard-count splits the suite across runners, and --shard-index selects the current runner’s shard. Merge the outputs with momentic results merge before uploading.
You can shard across CI runners and use --parallel within each runner.
Four shards with --parallel 4 can run up to 16 tests at once. See the GitHub Actions guide for a CI matrix and merge workflow, and the run reference for every flag.
Make sure step caching is saving in CI. Cached steps can avoid repeated model work, but application readiness and network requests still affect runtime. See Saving caches for setup.

Mobile test suites

Mobile runs use the same --parallel, --shard-count, and --shard-index flags through momentic-mobile, with a few platform constraints:
  • --parallel AUTO saturates the current shard. On remote emulators, each test gets its own session and organization quotas are enforced server-side.
  • Local Android runs need a distinct AVD per parallel test. Multiple tests cannot share the same --local-avd-id.
  • Local iOS runs should use --parallel 1 because concurrent sessions share a device-driver manifest.
Merge shards with momentic-mobile results merge before uploading. See the momentic-mobile run reference for every flag.