Google just put the latest AI models through a brutal coding test — here’s how they did

September 18, 2026:

Google just put the latest AI models through a brutal coding test — here’s how they did

What you need to know

  • Google has launched Android Bench 2.0 to test AI models on complex Android development tasks that can take days.
  • The new benchmark includes tasks like upgrading dependencies, adding major features, and building Android apps from scratch.
  • GPT-6 Astra currently leads Google’s new benchmark with a 28% pass rate, while Gemini 3.8 Flash scored just 8%.

Google has announced Android Bench 2.0, an updated version of its benchmark for evaluating how well large language models (LLMs) and AI agents handle complex Android development tasks.

Earlier this year, Google introduced the first version of Android Bench to measure how AI models perform on real-world Android development work. The company has now updated the benchmark with Android Bench 2.0, which is designed to evaluate models and agents against more complex tasks that better reflect actual software development.

Source link