Google’s Android Bench benchmark just got a major refresh. New models, new framework, and a very awkward result: Google’s own Gemini still isn’t winning.
Android Bench measures how well LLMs handle 100 Android development tasks. It launched back in March, but today’s update adds eight new heavy-hitters: Claude Fable 5, Claude Sonnet 5, Claude Opus 4.8, GLM 5.2, Kimi K2.7 Code, MiniMax M3, Qwen 3.7 Plus, and Qwen 3.7 Max.
The new leaderboard isn’t great news for Google. Gemini 3.1 Pro sits in fifth place, behind GPT 5.4, Claude Sonnet 5, and Claude Fable 5. Fable 5 leads the pack with 84.5% accuracy. Gemini’s not even close.
But there’s a trade-off. Fable 5 and GPT 5.5 are expensive — they burn through over $130 in tokens for the full benchmark. Gemini 3.1 Pro costs $87. And Gemini 3.5 Flash, supposedly the cheap option, ended up being the most expensive at $165 because it took 28 hours to finish.
That’s a problem for Google as it pushes deeper into agentic development. The company would obviously prefer Android developers use Google’s tools. The performance gap might explain why Google’s reportedly been buying app source code from developers for AI training.
Google’s also switching Android Bench to the Harbor framework, making it easier for developers to run their own tests and submit results. The GitHub repo has been updated with the new dataset and instructions.
The historical data stays online in an archive. Google re-ran all previous tests with Harbor to establish a new baseline, so scores shifted a bit even for the same tests.
