Google updates Android Bench with new LLMs, but Gemini still lags behind
Google has released a significant update to Android Bench, its benchmark for measuring how well large language models handle Android app development tasks. The update adds eight new models and switches to a more accessible testing framework, but the headline finding is that Google's own Gemini models continue to trail rivals from OpenAI and Anthropic. This matters because Google is increasingly steering its own projects towards agentic development and would prefer developers to adopt its tools, making the performance gap awkward for the company.
Across the 100-task suite, Gemini 3.1 Pro now sits in fifth place, behind GPT 5.4, Claude Sonnet 5 and Claude Fable 5, with Fable 5 leading at 84.5 per cent accuracy. Performance comes at a cost, however: Fable 5 and GPT 5.5 each burn through more than $130 in tokens over the 10-run benchmark, while Gemini 3.1 Pro is cheaper at $87, and the supposedly economical Gemini 3.5 Flash proved the most expensive at $165 per run due to a 28-hour runtime. Google has adopted the Harbor framework to make it easier for developers to run, evaluate and share results, re-ran all previous tests to establish a new baseline, and is inviting the community to submit their own tasks for possible inclusion via the updated Android Bench GitHub.
- Google's Gemini still trails OpenAI and Anthropic on its own Android Bench.
- Claude Fable 5 leads with 84.5% accuracy but high cost.
- New Harbor framework lets developers run and submit their own tests.