> agent-qa scores 100% on iOS and 99.1% on Android in AppControlBench
Click the card to flip it
agent-qa is an open-source self-improving QA agent. We have tested it on hundreds of web and mobile apps, and on Software Mansion's AppControlBench, where an agent operates real Bluesky and Element apps from plain-language requests, it completed all 60 iOS tasks and scored 54.5 of 55 on Android, one attempt per task.
| Category | Benchmark | agent-qagpt-6-sol | argentqwen3.8-max (high) | agent-deviceqwen3.8-max (high) | MidsceneDoubao Seed 2.1 Turbo | No toolqwen3.8-max (high) |
|---|---|---|---|---|---|---|
| iOS | Completion60 tasks | 100.0% | 98.3% | 92.5% | 97.5% | 61.7% |
| Pass rate60 tasks | 100.0% | 96.7% | 88.3% | 96.7% | 50.0% | |
| iOS by app | Bluesky30 tasks | 100.0% | 98.3% | 93.3% | 96.7% | 51.7% |
| Element30 tasks | 100.0% | 98.3% | 91.7% | 98.3% | 71.7% | |
| iOS by task | Navigation39 tasks | 100.0% | 98.7% | 97.4% | 100.0% | 60.3% |
| Interaction13 tasks | 100.0% | 96.2% | 84.6% | 88.5% | 61.5% | |
| Posting8 tasks | 100.0% | 100.0% | 81.3% | 100.0% | 68.8% | |
| Android | Completion55 tasks | 99.1% | — | — | — | — |
Completion counts a success as 1, a partial as 0.5 and a fail as 0; pass rate counts full successes only. The argent, agent-device and No tool columns are each the best published model for that setup, qwen3.8-max (high) in all three, from the AppControlBench leaderboard as of 2026-10-04, recomputed per app and per task from its per-task verdicts. Midscene's figures come from its own report, which rounds its completion to 98%, and agent-qa's from our runs: neither is on the leaderboard, and models, judges and iOS versions differ between the columns. The leaderboard has no Android results yet.
The problem
When a coding agent writes the code and its tests, both carry the same assumptions. If it misread the requirement, the tests encode the mistake, and coding agents are very good at convincing themselves it works. Asking the same agent to test itself adversarially is slow, flaky and does not scale: it is the same model grading its own homework. QA needs its own agent.
Our approach
We rebuilt QA for the agentic era. agent-qa tests on real browsers and devices and learns from every run.
The harness
Battle-tested frameworks, Playwright and Appium, are the kernel that executes every action. A separate verifier must agree each step is done, and every run is recorded.
memidex
A memory layer for QA: plain-text facts and procedures you review in pull requests.
Self-improvement loop
After each run, agent-qa suggests what it should remember. It keeps a lesson only when runs back it up, and if you mark a run as passed or failed, your call wins over the agent's.
AppControlBench
Software Mansion's AppControlBench is 60 plain-language tasks in real Bluesky and Element apps, such as in the Book Club room, reply to one of the existing messages with the text 'sounds good'
. A vision judge grades the final screen: success counts 1, partial 0.5, fail 0. agent-qa completed all 60 iOS tasks and scored 54.5 of 55 on Android, one attempt per task; the benchmark's Android port, still in review, supports 55 of the tasks. Every run is below.
memidex and the self-improvement loop were switched off for this benchmark, and every task started in a fresh workspace with an empty memory. agent-qa ran on gpt-6-sol and the judge was gpt-6-sol (high). Each task's context added the same task-agnostic paragraph to the benchmark's preamble, on ending at a screen that shows the requested result. One iOS task, element-17, was rerun after a low-battery safeguard stopped it.
iOS runs
iPhone 17 Pro simulator, iOS 27.0; Element 1.11.40, Bluesky 1.122.0. Each run opens on agent-qa's own record of it: every step, what the agent planned, its screenshots and the verifier's reasoning. The benchmark trajectory tab shows what the judge graded.
agent-qa's own verdict can differ from the judge's. On bsky-13 it failed the run as an application error, because Bluesky's chat showed Whoops! could not resolve iss did
; that is the screen the benchmark's solution ends on, so the judge graded it a success.
- Completion
- 100%
- Success / partial / fail
- 60 / 0 / 0
- Median time
- 1m 14s
- Median actions
- 2.5
Android runs
Pixel 6 emulator (AVD), Android 13 (API 33); Element X 26.08.0, Bluesky 1.122.0. Each run opens on agent-qa's own record of it: every step, what the agent planned, its screenshots and the verifier's reasoning. The benchmark trajectory tab shows what the judge graded.
- Completion
- 99.1%
- Success / partial / fail
- 54 / 1 / 0
- Median time
- 43s
- Median actions
- 3