---
title: "agent-qa vs agent-device"
description: "Compare agent-qa vs Callstack agent-device for mobile QA: device control, replay suites, YAML expectations, behavioral memory, MCP workflows, and run evidence."
canonical_url: "https://vostride.com/product-comparison/agent-qa-vs-agent-device"
md_url: "https://vostride.com/product-comparison/agent-qa-vs-agent-device.md"
last_updated: "2026-10-02T00:00:00.000Z"
---

> [Documentation index](https://vostride.com/docs.md) · [Site index](https://vostride.com/llms.txt)
> Follow the Markdown links to read individual pages; the indexes list the available documentation.

# agent-qa vs agent-device

> Compare agent-qa vs Callstack agent-device for mobile QA: device control, replay suites, YAML expectations, behavioral memory, MCP workflows, and run evidence.

An agent-device alternative for tests that retain context.

Callstack's agent-device gives agents device control and repeatable replay workflows. agent-qa organizes verification around natural-language test contracts and product memory shared across runs.

Try agent-qa, a source-available QA runtime that turns expected behavior into repository-owned tests. Use CLI, MCP, and Skills to run a journey, inspect its evidence, and carry relevant product observations into the next change.

## The choice for recurring QA

Choose agent-qa for a recurring QA suite built around YAML intent and curated behavioral memory. Choose agent-device when direct device control, saved command flows, and programmable automation are central. It already supports verification and CI, so the decision is about the test model you want to own.

Sources reviewed 2026-10-02 by Vostride.

> This comparison is based on publicly available information. Product capabilities and pricing can change; verify details with each vendor before making a purchase decision.

## Capability comparison

- **CLI and MCP for coding agents.** agent-qa: Yes; agent-device: Yes. agent-device exposes its device runtime through CLI, MCP, and a typed Node.js API. agent-qa's CLI and MCP center on test execution and result inspection. Both let coding agents verify changes through structured tools. Sources: [agent-device: platforms, APIs, and MIT license](https://github.com/callstack/agent-device), [agent-qa: agent workflow and CI](https://vostride.com/docs/agent-qa/guides/coding-agent-workflow.md).

- **Natural-language test contracts.** agent-qa: Yes; agent-device: Partial. agent-device lets an agent explore an app and save command flows. agent-qa stores natural-language steps and expected outcomes as the test definition. Choose whether repeatability should follow a saved action sequence or a declared journey. Sources: [Callstack agent-device: product overview](https://agent-device.dev/), [agent-device: replay, suites, and Maestro YAML](https://oss.callstack.com/agent-device/docs/replay-e2e), [agent-qa: natural-language test contract](https://vostride.com/docs/agent-qa/guides/first-test.md).

- **Repository files and CI suites.** agent-qa: Yes; agent-device: Yes. agent-device has a suite runner for .ad scripts with retries, artifacts, and JUnit reporting. agent-qa runs YAML tests and suites with CI output. Both can preserve release checks as files in your repository. Sources: [agent-device: replay, suites, and Maestro YAML](https://oss.callstack.com/agent-device/docs/replay-e2e), [agent-qa: agent workflow and CI](https://vostride.com/docs/agent-qa/guides/coding-agent-workflow.md).

- **Curated behavioral memory.** agent-qa: Yes; agent-device: Partial. Saved agent-device scripts retain actions for replay. agent-qa additionally stores behavioral observations in product, suite, and test Markdown files, retrieves relevant context, and curates it after runs. The cited replay docs do not establish that same memory lifecycle. Sources: [agent-device: replay, suites, and Maestro YAML](https://oss.callstack.com/agent-device/docs/replay-e2e), [agent-qa: behavioral memory](https://vostride.com/docs/agent-qa/memory.md).

- **Native mobile execution.** agent-qa: Yes; agent-device: Yes. agent-device supports Android and iOS through its device backends. agent-qa uses configured Appium targets. Compare provisioning, app-state handling, and the devices your team needs; native-mobile support is shared rather than unique to either product. Sources: [agent-device: platforms, APIs, and MIT license](https://github.com/callstack/agent-device), [agent-qa: Appium mobile setup](https://vostride.com/docs/agent-qa/guides/mobile-testing.md).

- **Browser verification.** agent-qa: Yes; agent-device: Partial. agent-device documents basic web support through agent-browser within its session and replay system. agent-qa treats web as a configured QA target. The partial rating reflects agent-device's stated basic scope, not an absence of browser support. Sources: [agent-device: platforms, APIs, and MIT license](https://github.com/callstack/agent-device), [agent-qa: natural-language test contract](https://vostride.com/docs/agent-qa/guides/first-test.md).

- **Inspectable failure evidence.** agent-qa: Yes; agent-device: Yes. agent-device captures screenshots, recordings, logs, and network diagnostics. agent-qa associates evidence with test-step outcomes and exposes it to a run inspector and coding-agent tools. Assess which view makes a failed acceptance condition easiest to explain. Sources: [Callstack agent-device: product overview](https://agent-device.dev/), [agent-qa: recorded evidence](https://vostride.com/docs/agent-qa/guides/recorded-evidence.md), [agent-qa: agent workflow and CI](https://vostride.com/docs/agent-qa/guides/coding-agent-workflow.md).

- **Published source.** agent-qa: Yes; agent-device: Yes. agent-device is MIT licensed. agent-qa is source available under FSL-1.1-ALv2, with a future Apache 2.0 license. Evaluate those different permissions alongside the operational fit, especially when embedding a runtime in another product. Sources: [agent-device: platforms, APIs, and MIT license](https://github.com/callstack/agent-device), [agent-qa: license terms](https://vostride.com/license.md).

## Choose the contract your QA suite should preserve

### Keep the expected behavior explicit

Write a journey and the outcome that makes it correct. agent-qa's YAML format gives reviewers a stable statement of intent to assess alongside the application change.

### Retain observations across journeys

Use memory scoped to a product, suite, or test so later runs can retrieve useful context. The files remain reviewable, and current application evidence still governs the result.

### Organize diagnosis around test results

agent-qa links the declared test, step evidence, and follow-up run in one workspace. Use that structure when the team needs a repeatable acceptance-testing process around its coding agents.

## Verdict

agent-qa fits teams that want recurring verification expressed as intent, with reviewable product memory. agent-device fits teams that want a programmable device toolkit and repeatable command flows; both can contribute to a release process.

> Note: Support depth varies by device and platform. Partial indicates a different abstraction or the documented scope, not that custom orchestration is impossible. No comparative performance benchmark was run.

## Evaluate agent-qa against agent-device

1. Choose one app journey with a defined starting state and a precise acceptance condition.
2. Save an agent-device replay and author an agent-qa YAML test that checks the same behavior on the same build.
3. Break the acceptance condition and inspect each result, its artifacts, and the CI status before considering a passing run sufficient.
4. Change the UI without changing behavior, then compare maintenance effort, memory updates, device setup, and repeated-run cost.

[Set up your coding agent](https://vostride.com/docs/agent-qa/agent-quickstart.md) · [Evaluation guide](https://vostride.com/docs/agent-qa/guides/evaluating-agent-qa.md) · [Recorded QA evidence](https://vostride.com/docs/agent-qa/guides/recorded-evidence.md)

## Frequently asked questions

### Which agent-device project is compared here?

This page covers Callstack's agent-device at agent-device.dev and github.com/callstack/agent-device. It is a device automation and verification toolkit, distinct from Vostride's agent-qa runtime.

### Does agent-device already run E2E tests in CI?

Yes. Its replay and suite tooling supports repeatable checks, artifacts, retries, and JUnit results. agent-qa offers a different contract based on natural-language YAML and behavioral memory. CI support alone is not a reason to switch.

### Can agent-device work with YAML?

Its replay documentation includes Maestro YAML support and export with explicit compatibility limits. Those flows are not agent-qa tests. Treat any move as a translation of actions and expected results, with a fresh check of device setup and app state.

### Is agent-qa built on agent-device?

agent-qa's documented native mobile setup uses Appium. The separate tester-army/e2e framework uses agent-device as its mobile engine. These are distinct architectures; do not assume that scripts or device configurations transfer between them.

### When is agent-device the better fit?

Choose it for direct device interaction, replay-oriented automation, or a programmable toolkit with MIT licensing. Evaluate agent-qa when a suite of natural-language expectations and retained product context is the main deliverable. Confirm platform-specific capabilities for either tool.

## Sources

- [Callstack agent-device: product overview](https://agent-device.dev/)
- [agent-device: platforms, APIs, and MIT license](https://github.com/callstack/agent-device)
- [agent-device: replay, suites, and Maestro YAML](https://oss.callstack.com/agent-device/docs/replay-e2e)
- [tester-army/e2e: agent-device mobile engine](https://github.com/tester-army/e2e)
- [agent-qa: natural-language test contract](https://vostride.com/docs/agent-qa/guides/first-test.md)
- [agent-qa: agent workflow and CI](https://vostride.com/docs/agent-qa/guides/coding-agent-workflow.md)
- [agent-qa: behavioral memory](https://vostride.com/docs/agent-qa/memory.md)
- [agent-qa: Appium mobile setup](https://vostride.com/docs/agent-qa/guides/mobile-testing.md)
- [agent-qa: recorded evidence](https://vostride.com/docs/agent-qa/guides/recorded-evidence.md)
- [agent-qa: license terms](https://vostride.com/license.md)

## Compare other approaches

- [agent-qa vs Argent](https://vostride.com/product-comparison/agent-qa-vs-argent.md): Compare Software Mansion's Argent with agent-qa for mobile debugging, recorded flows, recurring QA, and retained product knowledge.
- [agent-qa vs TesterArmy e2e](https://vostride.com/product-comparison/agent-qa-vs-tester-army-e2e.md): Compare the tester-army/e2e framework with agent-qa: TypeScript versus YAML, agent workflows, cached execution, and behavioral memory.
- [agent-qa vs Appium](https://vostride.com/product-comparison/agent-qa-vs-appium.md): Compare agent-qa with Appium: natural-language mobile E2E testing with memory versus programmatic mobile automation.
- [agent-qa vs Maestro](https://vostride.com/product-comparison/agent-qa-vs-maestro.md): Compare agent-qa with Maestro: AI-driven natural-language testing with memory versus declarative YAML mobile flows.
