Skip to main content

> agent-qa is #1 on AndroidWorld benchmark

agent-qa is an open-source self-improving QA agent. We have tested it on hundreds of web and mobile apps, and we are happy to announce that it has saturated the AndroidWorld benchmark.

AndroidWorld leaderboardsuccess rate
  1. agent-qagpt-6-sol

    100.0%
  2. ArtemisGemini 3.7 Flash, Gemini Robotics ER 2

    99.1%
  3. AGI-0AGI-0

    97.4%
  4. Midscene.jsGemini-3.5-Flash

    93.1%
  5. DroidRunGPT5, Gemini 2.5 Pro

    91.4%
  6. HumanAndroidWorld paper baseline

    80.0%

The problem

When a coding agent writes the code and its tests, both carry the same assumptions. If it misread the requirement, the tests encode the mistake, and coding agents are very good at convincing themselves it works. Asking the same agent to test itself adversarially is slow, flaky and does not scale: it is the same model grading its own homework. QA needs its own agent.

Our approach

We rebuilt QA for the agentic era. agent-qa tests on real browsers and devices and learns from every run.

The harness

Battle-tested frameworks, Playwright and Appium, are the kernel that executes every action. A separate verifier must agree each step is done, and every run is recorded.

memidex

A memory layer for QA: plain-text facts and procedures you review in pull requests.

Self-improvement loop

After each run, agent-qa suggests what it should remember. It keeps a lesson only when runs back it up, and if you mark a run as passed or failed, your call wins over the agent's.

AndroidWorld

Google's AndroidWorld benchmark is 116 tasks across 20 real Android apps, scored by its own evaluators. agent-qa passed 116 of 116, beating Google's Artemis (99.1%). Every run is below.

memidex and the self-improvement loop were switched off for this benchmark. Every task started in a fresh workspace with an empty memory, and nothing learned in one run carried over to the next.

Runs

Give your coding agent a QA agent.