Skip to main content
agent-qa is #1 on AndroidWorld benchmark

> agent-qa scores 100% on iOS and 99.1% on Android in AppControlBench

#1

AppControlBench

> agent-qa

Overall
99.6% 114.5/115
iOS
100% 60/60
Android
99.1% 54.5/55

Click the card to flip it

agent-qa is an open-source self-improving QA agent. We have tested it on hundreds of web and mobile apps, and on Software Mansion's AppControlBench, where an agent operates real Bluesky and Element apps from plain-language requests, it completed all 60 iOS tasks and scored 54.5 of 55 on Android, one attempt per task.

AppControlBench completion by agent
CategoryBenchmarkagent-qagpt-6-solargentqwen3.8-max (high)agent-deviceqwen3.8-max (high)MidsceneDoubao Seed 2.1 TurboNo toolqwen3.8-max (high)
iOSCompletion60 tasks100.0%98.3%92.5%97.5%61.7%
Pass rate60 tasks100.0%96.7%88.3%96.7%50.0%
iOS by appBluesky30 tasks100.0%98.3%93.3%96.7%51.7%
Element30 tasks100.0%98.3%91.7%98.3%71.7%
iOS by taskNavigation39 tasks100.0%98.7%97.4%100.0%60.3%
Interaction13 tasks100.0%96.2%84.6%88.5%61.5%
Posting8 tasks100.0%100.0%81.3%100.0%68.8%
AndroidCompletion55 tasks99.1%————

Completion counts a success as 1, a partial as 0.5 and a fail as 0; pass rate counts full successes only. The argent, agent-device and No tool columns are each the best published model for that setup, qwen3.8-max (high) in all three, from the AppControlBench leaderboard as of 2026-10-04, recomputed per app and per task from its per-task verdicts. Midscene's figures come from its own report, which rounds its completion to 98%, and agent-qa's from our runs: neither is on the leaderboard, and models, judges and iOS versions differ between the columns. The leaderboard has no Android results yet.

The problem

When a coding agent writes the code and its tests, both carry the same assumptions. If it misread the requirement, the tests encode the mistake, and coding agents are very good at convincing themselves it works. Asking the same agent to test itself adversarially is slow, flaky and does not scale: it is the same model grading its own homework. QA needs its own agent.

Our approach

We rebuilt QA for the agentic era. agent-qa tests on real browsers and devices and learns from every run.

The harness

Battle-tested frameworks, Playwright and Appium, are the kernel that executes every action. A separate verifier must agree each step is done, and every run is recorded.

memidex

A memory layer for QA: plain-text facts and procedures you review in pull requests.

Self-improvement loop

After each run, agent-qa suggests what it should remember. It keeps a lesson only when runs back it up, and if you mark a run as passed or failed, your call wins over the agent's.

AppControlBench

Software Mansion's AppControlBench is 60 plain-language tasks in real Bluesky and Element apps, such as in the Book Club room, reply to one of the existing messages with the text 'sounds good'. A vision judge grades the final screen: success counts 1, partial 0.5, fail 0. agent-qa completed all 60 iOS tasks and scored 54.5 of 55 on Android, one attempt per task; the benchmark's Android port, still in review, supports 55 of the tasks. Every run is below.

memidex and the self-improvement loop were switched off for this benchmark, and every task started in a fresh workspace with an empty memory. agent-qa ran on gpt-6-sol and the judge was gpt-6-sol (high). Each task's context added the same task-agnostic paragraph to the benchmark's preamble, on ending at a screen that shows the requested result. One iOS task, element-17, was rerun after a low-battery safeguard stopped it.

iOS runs

iPhone 17 Pro simulator, iOS 27.0; Element 1.11.40, Bluesky 1.122.0. Each run opens on agent-qa's own record of it: every step, what the agent planned, its screenshots and the verifier's reasoning. The benchmark trajectory tab shows what the judge graded.

agent-qa's own verdict can differ from the judge's. On bsky-13 it failed the run as an application error, because Bluesky's chat showed Whoops! could not resolve iss did; that is the screen the benchmark's solution ends on, so the judge graded it a success.

Completion
100%
Success / partial / fail
60 / 0 / 0
Median time
1m 14s
Median actions
2.5

Android runs

Pixel 6 emulator (AVD), Android 13 (API 33); Element X 26.08.0, Bluesky 1.122.0. Each run opens on agent-qa's own record of it: every step, what the agent planned, its screenshots and the verifier's reasoning. The benchmark trajectory tab shows what the judge graded.

Completion
99.1%
Success / partial / fail
54 / 1 / 0
Median time
43s
Median actions
3

Give your coding agent a QA agent.