ABCDEFGHIJKLMNOPQRSTUVWXYZAAABACAD
1
⚠️ Community-submitted results. No independent verification is performed. Scores are self-reported and may be inaccurate. Use at your own risk.
2
RankRelease DateResult SourceModel TypeOpen?Model SizeModelScreen
Representation
Success Rate
(pass@1)
Number
of trials
Success Rate
(pass@k)
Trajectory
submissions
Note
3
108/2026FluizAIAI agent-gpt-4o / gpt-5.6-solScreenshot + A11y tree1001
https://android-world.fluiz.ai/
Code is not open-sourced yet.
4
208/2026ArtemisAI agent-
Gemini 3.7 Flash, Gemini Robotics ER 2
Screenshot + A11y tree99.11Open-source project available at https://github.com/google/artemis
5
310/2025AGI-0AI agent-AGI-0Screenshot97.41n/a
10/14/2025. Changed from model --> agent, based on description in blog post. https://www.theagi.company/blog/android-world
6
302/2026MobileUseAgentAI agent-Seed1.8-GUIScreenshot97.41
7
33/2026FinalrunAI agent-
Gemini 3 flash, Gemini 3 flash lite
Screenshot + A11y tree97.41
https://drive.google.com/file/d/1NITEkiYb3itSq46I8BYsmzefEddzYbqm/view?usp=sharing
04/2026: Project is open-source at https://github.com/final-run/finalrun-agent
03/2026: Updated results: https://blogs.finalrun.app/finalrun-achieves-the-new-1-sota-mobile-agent-eclipsing-benchmarks-at-97-4
08/2025: The technical doc/approach can be found here: https://github.com/final-run/finalrun-android-world-benchmark
8
610/2025
askui AndroidVisionAgent
AI agent-
askui AndroidVisionAgent, Claude 4.5 Sonnet + Claude 4.0 Sonnet
Screenshot94.81n/a
9
601/2026AutoDeviceAI agent-
gemini 3 pro + sonnet 4.5
Screenshot94.81
https://autodevice.io/benchmark/
Agent code is available on https://github.com/autodevice/android_world
10
806/2026Midscene.jsAI agent-Gemini-3.5-FlashScreenshot93.11
pass@2 95.69, pass@3 97.41
Trajectory submissions: https://midscenejs.com/android-world-benchmark-report
11
910/2025DroidRunAI agent-GPT5, Gemini 2.5 ProScreenshot + A11y tree91.41Trajectories
[10/6/2025]: 78.4 --> 91.4; details here
[09/19/2025]: Updated score from 63.0 --> 78.4. For this run, we used GPT-5 for reasoning in combination with Gemini-2.5-Pro for acting, while continuing to rely on a hybrid method of the Accessibility (A11y) tree plus corresponding screenshots to provide both structural and visual context to the agent.

12
99/2025mobile-useAI agent-
Llama 4-scout, Gemini 2.5 pro, GPT-5 nano
Screenshot + A11y tree91.41
https://minitap.ai/benchmark
[12/16/2025] 84.5% -> 91.4%
[10/1/2025]: 77.6% -> 84.5%
[09/12/2025]: Updated score from 74.1% --> 77.6%
[08/19/2025]: Initial trajectory labeling issue (MarkorCreateNoteFromClipboard) has been corrected by the authors. They provide Discord support for setup issues.
[08/18/2025]: Repo has been reported broken by some users (not independently verified). Some trajectories (e.g., MarkorCreateNoteFromClipboard, BrowserMaze) are labeled as successful by the authors but contain incorrect actions
13
1112/2025Agent-ViscoAI agent--Screenshot88.81
https://work.aliyun.com/alimail/openLinks/downloadMimeMetaDiskBigAttach?id=netdiskid%3Av001%3Afile%3ADzzzzzzNqZy%3BG8EhE7UV7U9nk2t2O6opysa%2BgRf5Zm6eTqXqtVApm1nDD8W%2BxnLAgdDsJsOSlxfXTx44NU7qRM%2BkJMuIA0smltxGWspXGX2Ox228yqp07rNX%2FIm2QehcGJBgHwUs0f81cz2sX76WNxvxnhpZnqmxiYxHNHRd2hybrPLi2DRQ8bM5ORhI%2BtGG%2B7n%2BcxdY1Ji0
Research paper and codes will be open-source in early months.
14
1210/2025Surfer 2AI agent-o3 + holo1.5-72bScreenshot87.11pass@3 93.1
https://hcompai.github.io/android-world-traces/
Agent is based on https://github.com/hcompai/surfer-h-cli - same architecture, but with modifications to prompt and action space
15
1310/2025gbox.aiAI agent-
Sonnet 4.5 + Sonnet 4
Screenshot86.21
https://github.com/babelcloud/android_world_benchmark/tree/main/GBOX/trajectory
GBOX uses Claude code as the agent with GBOX mcp. Link to report -> https://github.com/babelcloud/android_world_benchmark/blob/main/GBOX/report.md
16
148/2025AutoGLM-MobileModel9BAutoGLM-MobileScreenshot + A11y tree80.21
autoglm-mobile-aw-ckpt.zip
[9/26/2025]: Updated score from 75.8 --> 80.2 This is an updated version of AutoGLM-Mobile. In this version, we added more pretraining data and trained the VLM at full image resolution.
17
159/2025LX-GUIAgentAI agent-LX-GUIAgentScreenshot + A11y tree79.31
androidworld-trajectories.zip
[9/23/2025]: Updated score from 75.0 --> 79.3. Added trajectories.
18
1612/2025AgentProgAI agent-
Gemini-2.5-Pro+UI-TARS-1.5
Screenshot78.01
AgentProg_AndroidWorld.zip
Paper: https://arxiv.org/pdf/2512.10371; Code: https://github.com/MobileLLM/AgentProg
19
179/2025K²-AgentAI agent72B + 7B
Qwen2.5-VL-72B + Qwen2.5-VL-7B
Screenshot76.71
https://github.com/k2-agent/k2-agent/blob/main/androidworld-trajectories.zip
K²-Agent, a hierarchical approach that self-evolves a Qwen2.5-VL-72B for high-level planning and post-trains a Qwen2.5-VL-7B for low-level execution.
Our GitHub repository, which contains a detailed description of our approach, can be found here: https://github.com/k2-agent/k2-agent.
20
171/2026MAI-UIModel235BMAI-UI-235B-A22BScreenshot76.7
21
199/2025MobileUse-v2AI agent32BHammer-UI-32BScreenshot75.01
MobileUse-v2-aw-ckpt.zip
Make further post-training based on the GUI-Owl-32b model. Optimize the memory and knowledge module of the MobileUse framework.
22
208/2025Mobile-Agent-v3AI agent32BGUI-Owl-32BScreenshot73.31https://github.com/X-PLUG/MobileAgent/tree/main
23
201/2026MAI-UIModel32BMAI-UI-32BScreenshot73.3
24
221/2026MAI-UIModel8BMAI-UI-8BScreenshot70.7
25
2310/2025
Gemini 2.5 Computer Use
Model-
Gemini 2.5 Computer Use
Screenshot69.71
26
246/2025JT-GUIAgent-V2AI agent-JT-GUIAgent-V2Screenshot67.21
27
258/2025GUI-Owl-7BModel7BGUI-Owl-7BScreenshot66.41https://github.com/X-PLUG/MobileAgent/tree/main
28
268/2025UI-VenusModel72BUI-Venus-Navi-72BScreenshot65.91
https://github.com/inclusionAI/UI-Venus/blob/main/vis_androidworld/UI-Venus-androidworld.zip
https://huggingface.co/inclusionAI/UI-Venus-Navi-72B
29
2707/2025MobileUseAI agent72BQwen2.5-VL-72BScreenshot62.91
30
2805/2025Seed1.5-VLModel20.BSeed1.5-VLScreenshot + A11y tree62.11
31
296/2025JT-GUIAgent-V1AI agent-JT-GUIAgent-V1Screenshot60.01
32
303/2025V-Droid PaperAI agent8BV-Droid (Llama8B)A11y tree59.51--Training data consists of apps and tasks from the AndroidWorld benchmark. Code.
33
314/2025Agent S2AI agent-Agent S2Screenshot54.31-
34
328/2025UI-VenusModel7BVenus-Navi-7BScreenshot49.11
https://github.com/inclusionAI/UI-Venus/blob/main/vis_androidworld/UI-Venus-androidworld.zip
https://huggingface.co/inclusionAI/UI-Venus-Navi-7B
35
321/2026MAI-UIModel2BMAI-UI-2BScreenshot49.1
36
3405/2025GUI-ExplorerAI agent-GPT-4oScreenshot + A11y tree47.41
37
354/2025AndroidGenAI agent-GPT-4oA11y tree46.81
38
361/2025UI-TARSModel72BUI-TARSScreenshot46.61-
39
3712/2024Aria-UIModel-GPT-4o + Aria-UIScreenshot44.81--
40
384/2025ScaleTrackModel8BScaleTrack-7BA11y tree44.01
41
381/2025UGroundModel-GPT-4o + UGroundScreenshot44.01-TrajectoriesCode for reproduction
42
406/2025Mirage-1AI agent-GPT-4oScreenshot42.21With Mirage-1-O; uses OS-Atlas grounder
43
4112/2024Ponder & PressAI agent-GPT-4oScreenshot34.51--Code is not yet open-sourced
44
4205/2024AndroidWorldAI agent-GPT-4 TurboA11y tree30.61--
45
436/2025GUI-Critic-R1Model7BQwen-2.5-VL-7BScreenshot + A11y tree27.61
46
4305/2024EcoAgentAI agent-
GPT-4o, OS-Atlas-Pro 4B, Qwen2-VL-2B-Instruct
Screenshot27.61
47
451/2025InfiGUIAgentModel2B
Qwen2-VL-2B (fine-tuned)
Screenshot9.01--
48
10/2024OSCARAI agent-GPT-4oScreenshot161.6 (k=4)-Code will be open-source upon publication.
49
50
Human Performance
51
05/2024AndroidWorld-Human80.03
52
53
Comment here or email crawles@gmail.com to submit your work! Please attach how should the data entry look like.
54
55
Definitions
56
Model
A relatively simple prompt involving one LLM /VLM call
57
AI agent
A multi-agent architecture involving several LLM calls and a protocol to coordinate the various agents or an LLM wrapped into an advanced agent with memory, subgoal planning, etc.
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100