Close Menu
DroidfeatsDroidfeats

    Subscribe to Updates

    Join our most passionate readers to get instant access to tech tips as they arrive!

    What's Hot
    Google Launches Android Bench 2.0 With Long-Horizon Coding Tasks

    Google Launches Android Bench 2.0 With Long-Horizon Coding Tasks

    September 18, 2026
    OpenAI Copyright Lawsuit Deepens as Publishers Demand Royalties

    OpenAI Copyright Lawsuit Deepens as Publishers Demand Royalties

    September 17, 2026
    Bill Gates Urges US and China to Cooperate on AI Risks

    Bill Gates Urges US and China to Cooperate on AI Risks

    September 17, 2026
    Facebook X (Twitter) Instagram
    • About
    • Privacy Policy
    • DMCA
    • Team
    • Get In Touch
    Facebook X (Twitter) Instagram Pinterest RSS
    DroidfeatsDroidfeats
    • Home
    • News
    • Apps
    • Tips
    • VPN
    • #TheBest
      • Get GCam Ports
      • USB Drivers
      • Get Magisk
      • Get Play Store
      • Get ADB binaries
    • LoL Tier List
    Best Deals
    DroidfeatsDroidfeats
    You are at:Home»News»Google Launches Android Bench 2.0 With Long-Horizon Coding Tasks
    News

    Google Launches Android Bench 2.0 With Long-Horizon Coding Tasks

    Saeed Ashif AhmedBy Saeed Ashif AhmedSeptember 18, 2026No Comments4 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Reddit Email

    Google released Android Bench 2.0 on Wednesday, introducing multi-day engineering tests and continuous scoring metrics to measure how artificial intelligence models handle complex, real-world application development.

    Key points

    • Google released Android Bench 2.0 to evaluate AI coding models on multi-day software engineering challenges.
    • The benchmark abandons binary pass-fail grading for a continuous completion rate based on functionality and visual fidelity.
    • OpenAI's GPT-6 Astra leads the new leaderboard with a 28.0% pass rate, followed by Claude Fable 5.1 at 22.7%.
    • No evaluated model achieved a 100% pass rate when porting cross-platform applications directly to Android.
    • The framework tests complete agent harnesses including OpenAI Codex and Google's Antigravity SDK alongside the underlying models.
    Google Launches Android Bench 2.0 With Long-Horizon Coding Tasks
    Android code running inside a terminal environment during an automated agent evaluation test. Image: Android Headlines

    The revised benchmark shifts away from small, isolated bug fixes and instead challenges AI models with assignments that typically take human engineers several days to a full week to finish. As software teams increasingly rely on autonomous tools to modify large codebases, the update provides a standardized measurement of how effectively these systems manage sustained technical labor on the Android operating system.

    Continuous Scoring Replaces Pass-Fail Grades

    The original benchmark assessed model output using binary pass or fail criteria. Under that framework, an agent that completed most of an extensive migration but missed a single edge condition received a zero, obscuring partial success on complex assignments.

    On multi-day engineering tasks, binary pass or fail grading doesn’t capture the full picture. For example, an agent might refactor 40 screens to Jetpack Compose, set up database tables, and pass 90% of requirements, but fail a single edge

    case assertion. Binary scoring rates this run as 0%.

    According to the Android Bench 2.0 announcement authored by Matthew McCullough, VP of Product Management for Android Developer, the system now calculates a continuous completion rate. The metric measures functional execution, visual fidelity, and the absence of regressions, while applying penalties if a model deviates from explicit instructions or structural constraints.

    Long-Horizon Tasks Lower Top AI Scores

    The transition to multi-day task profiles reduced overall success rates across all leading models. While frontier AI tools frequently scored above 90% on previous incremental bug-fix benchmarks, the highest-performing models struggled to complete the expanded challenges.

    ModelFramework HarnessPass Rate
    OpenAI GPT-6 AstraOpenAI Codex28.0%
    Anthropic Claude Fable 5.1Claude-Code22.7%
    Gemini 3.8 FlashAntigravity SDKSub-20%

    OpenAI’s GPT-6 Astra took the top position on the updated leaderboard with a 28.0% pass rate. Anthropic’s Claude Fable 5.1 followed in second place with 22.7%. AI systems performed reliably on deterministic assignments, such as converting Java files to Kotlin or swapping network layers from Retrofit to Ktor. Models also managed feature creation from clean slates. Difficulties emerged during architectural refactoring, runtime verification, handling unreleased libraries, and porting cross-platform codebases to native Android, where completion rates peaked at 80% and zero models reached 100%.

    Evaluating Developer Harnesses and Agent Workflows

    Android Bench 2.0 expands testing beyond standalone foundational models to measure complete developer agents. Google paired models with their respective operational harnesses, including OpenAI Codex, Claude-Code, the Antigravity SDK, Kimi-Code, and Qwen-Coder, to measure real-world performance.

    • Evaluates end-to-end agentic workflows rather than isolated raw API prompts.
    • Measures token consumption and context window management across multi-turn sessions.
    • Assesses developer environment integration, build-system repairs, and automated terminal commands.

    The updated framework builds on the original Android Bench introduced in March 2026 and incorporates the sandboxed testing environments standardized by Google under the Harbor framework in July 2026. The testing suite continues to run models against isolated virtual devices to verify visual and operational fidelity during grading runs.

    Related coverage

    • Android 17 QPR2 Beta 4 Adds New Notification Intelligence

    Source

    Originally reported by Android Headlines.

    “Google’s Android Bench 2.0 Replaces Pass/Fail Grades for Real-World Coding Tests”

    Read the original report ↗
    Android Bench 2.0 Android Development Artificial Intelligence Google Software Development
    Add as Preferred on Google Follow on Flipboard
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email Copy Link
    Previous ArticleOpenAI Copyright Lawsuit Deepens as Publishers Demand Royalties
    Saeed Ashif Ahmed
    • Website
    • Facebook
    • X (Twitter)
    • LinkedIn

    Saeed Ashif Ahmed is passionate about emerging technologies and their potential to create a fairer world. A car enthusiast and civil engineer, he loves cricket and cherishes his alma mater, Navodaya Vidyalaya (JNV).

    Related Posts

    OpenAI Copyright Lawsuit Deepens as Publishers Demand Royalties
    News

    OpenAI Copyright Lawsuit Deepens as Publishers Demand Royalties

    September 17, 2026
    Bill Gates Urges US and China to Cooperate on AI Risks
    News

    Bill Gates Urges US and China to Cooperate on AI Risks

    September 17, 2026
    Sam Altman Backs Public Fear Over AI Risks and Control
    News

    Sam Altman Backs Public Fear Over AI Risks and Control

    September 17, 2026
    Add A Comment
    Leave A Reply Cancel Reply

    Top Posts
    MX Player Custom Codec [AC3, DTS, MLP, TrueHD, and more]

    MX Player Custom Codec [AC3, DTS, MLP, TrueHD, and more]

    May 1, 2026223K Views
    Do a Barrel Roll 20 Times on Google & other search games (2026)

    Do a Barrel Roll 20 Times on Google & other search games (2026)

    May 1, 202658,611 Views
    47 best root apps for Android devices in 2025 (NEW LIST – Updated)

    47 best root apps for Android devices in 2025 (NEW LIST – Updated)

    April 16, 202530,541 Views
    Stay In Touch
    • Facebook
    • Twitter
    • Instagram
    • Pinterest
    • Reddit
    • Telegram
    Latest News
    OpenAI Copyright Lawsuit Deepens as Publishers Demand Royalties News

    OpenAI Copyright Lawsuit Deepens as Publishers Demand Royalties

    By Joyce de CastroSeptember 17, 2026
    Bill Gates Urges US and China to Cooperate on AI Risks News

    Bill Gates Urges US and China to Cooperate on AI Risks

    By Palumbo AngelaSeptember 17, 2026
    Sam Altman Backs Public Fear Over AI Risks and Control News

    Sam Altman Backs Public Fear Over AI Risks and Control

    By Joyce de CastroSeptember 17, 2026

    Subscribe to Updates

    Join our most passionate readers to get instant access to tech tips as they arrive!

    Most Popular
    MX Player Custom Codec [AC3, DTS, MLP, TrueHD, and more]

    MX Player Custom Codec [AC3, DTS, MLP, TrueHD, and more]

    May 1, 2026223K Views
    Do a Barrel Roll 20 Times on Google & other search games (2026)

    Do a Barrel Roll 20 Times on Google & other search games (2026)

    May 1, 202658,611 Views
    47 best root apps for Android devices in 2025 (NEW LIST – Updated)

    47 best root apps for Android devices in 2025 (NEW LIST – Updated)

    April 16, 202530,541 Views
    Our Picks
    Google Launches Android Bench 2.0 With Long-Horizon Coding Tasks

    Google Launches Android Bench 2.0 With Long-Horizon Coding Tasks

    September 18, 2026
    OpenAI Copyright Lawsuit Deepens as Publishers Demand Royalties

    OpenAI Copyright Lawsuit Deepens as Publishers Demand Royalties

    September 17, 2026
    Bill Gates Urges US and China to Cooperate on AI Risks

    Bill Gates Urges US and China to Cooperate on AI Risks

    September 17, 2026

    Subscribe to Updates

    Join our most passionate readers to get instant access to tech tips as they arrive!

    Facebook X (Twitter) Instagram Pinterest
    • About Us
    • Privacy Policy
    • Terms
    • Jobs
    • Contact
    © 2026 Droidfeats. All rights reserved.

    Type above and press Enter to search. Press Esc to cancel.

    Ad Blocker Enabled!
    Ad Blocker Enabled!
    Our website is made possible by displaying online advertisements to our visitors. Please support us by disabling your Ad Blocker.