DEV Community

Artem Garazha
Artem Garazha

Posted on

Vibe coding on Android: problems nobody asked for

What the research actually says about AI-generated Android code, with numbers, studies, and production incidents.


Andrej Karpathy coined "vibe coding" in 2025. You describe an app in natural language, AI generates the code, you don't read every line, you iterate until it works.

For web prototypes, fine. For Android, it works right up until your first production deploy. Then the crash reports start rolling in and you're reading every line anyway, just at 3am with an incident channel pinging.

Android isn't a language plus a framework. It's Kotlin/Java, Gradle with multi-module configs, Activity/Fragment lifecycles, coroutines, Jetpack Compose recomposition, Hilt, Room, Navigation Component, WorkManager, and API versions that change faster than LLM training cycles. When an AI doesn't understand this stack, it generates code that looks right and behaves wrong. That's worse than code that won't compile. The compiler at least tells you something's broken.


1. API hallucinations: code from a parallel universe

The most obvious and most frequent problem. AI models train on data frozen at a cutoff date. The Android SDK, Jetpack libraries, and Kotlin update every few months. The model generates code using APIs that no longer exist, or never existed.

Real examples from dev.to and production:

  • AsyncTask — removed in Android API 33. AI keeps generating it.
  • TestCoroutineDispatcher — removed from coroutines-test 1.8+.
  • ContextualFlowRow — deprecated in Compose 1.8.
  • android:screenOrientation="portrait" — illegal on screens ≥600dp under Android 16.
  • Navigation 2 patterns in a Navigation 3 project.
  • LiveData instead of StateFlow in Compose projects.

Developer Vikas Sahani, creator of AndroJack MCP, documented this: "My AI assistants (Cursor, Claude, Windsurf) were exceptionally fast at generating code, but they were consistently generating the wrong code. They suggested Navigation 2 patterns for a Navigation 3 project, hallucinated Gradle coordinates that didn't exist on Maven, used APIs removed from Android years ago. This wasn't a failure of the models. It was a failure of grounding."

Scale of the problem, per USENIX Security 2025:

  • 2.23 million generated code samples, 16 models analyzed.
  • Open-source models hallucinate package names in 21.7% of cases on average.
  • Commercial models: 5.2%. That's still roughly 1 in 20 dependencies.
  • 20% of all AI-suggested packages are hallucinations (Snyk data).

2. Slopsquatting: when hallucination becomes an attack vector

This is the part that scares me. When AI hallucinates a package name that doesn't exist (aws-helper-sdk, express-rate-guard, huggingface-cli), someone can register that name with malicious code. And because AI hallucinates the same names consistently (43% of hallucinated names repeat across 10 separate queries, per USENIX Security 2025), the attacker knows exactly what to squat.

Researchers classified the hallucinations:

  • Pure fabrication (51%): plausible but invented. express-rate-guard.
  • Conflation of two real names (38%): react-codeshift (jscodeshift + react-codemod). Dangerous because they sound semantically correct.
  • Typo of a real package (13%): lodahs instead of lodash.

For Android this matters because Gradle pulls dependencies from Maven Central and Google Maven. An attacker can publish a malicious artifact under a hallucinated name and wait for an AI to suggest it. Developer Zhenya Sedoy built SafeDroid, a GitHub Action that validates Gradle dependencies before merge, specifically because "AI sometimes confidently suggests a package name that doesn't exist."


3. Coroutines and lifecycle: leaks that don't crash, but kill

This is the Android-specific stuff that vibe coding is worst at. Praveen Yadati, researching AI-generated Kotlin/Compose code, identified five major failure patterns.

Wrong coroutine scope

AI loves GlobalScope or forgets to cancel coroutines entirely. A leaked coroutine doesn't crash. It just runs in the background forever, draining battery and creating bugs you won't find until users report that their phones are hot and their battery is dead by lunch.

Force-unwrap (!!) operators

Kotlin is designed for null safety. AI responds by sprinkling !! everywhere to make the code compile. Every !! is a NullPointerException waiting for production. I've watched this happen: the PR looks clean, the build passes, and six hours later Crashlytics lights up.

Recomposition issues in Compose

AI rarely optimizes for Compose performance. It generates lambdas that trigger unnecessary recompositions. If your UI feels janky, check for unstable parameters being passed to composables. The AI won't flag this.

Missing loading, error, and empty states

AI builds the happy path. It almost never adds loaders, empty states, or error handling unless you explicitly ask. I now add "Include loading, empty, and error states in the UI" to every prompt. Learned that one the hard way.

Room migrations

AI generates @Database(version = 1) with no migration strategy. Add a column, crash. @AutoMigrations is almost never included.


4. What the numbers say

This isn't anecdotal. 2025–2026 produced a lot of research.

Columbia DAPLab (January 2026): Analyzed 15+ applications built through 5 top AI agents (Claude, Cline, Cursor, V0, Replit). Nine critical failure patterns:

  1. UI Grounding Mismatch — agent doesn't understand spatial requests.
  2. State Management Failures — state loss during refactoring.
  3. Business Logic Mismatch — code runs without errors but produces wrong output.
  4. Data Management Errors — ID confusion (firestore_id vs jira_id).
  5. API & External Service Integration — hardcoded fake keys instead of requesting real ones.
  6. Security Vulnerabilities — mixing admin and user roles.
  7. Repeated Code — duplication instead of abstraction.
  8. Codebase Awareness — more files means more errors.
  9. Exception & Error Handling — suppressing errors instead of reporting them.

DAPLab's headline: agents prioritize runnable code over correct code and suppress errors instead of reporting them.

CodeRabbit (December 2025): 470 open-source PRs analyzed:

  • AI code has 1.7× more defects than human code.
  • Logic and correctness errors: 1.75× higher.
  • XSS vulnerabilities: 2.74× higher.
  • Password handling issues: 1.88× higher.
  • Performance issues: nearly 8× more frequent.

Lightrun Survey (2026): 43% of AI-generated code changes required additional debugging after deployment.

Google DORA Report (2025): AI adoption correlates with roughly a 10% increase in code instability.

CloudBees (2026): 81% of enterprises see rising production failures alongside AI code adoption.

Stack Overflow Developer Survey (2025, 49,000 developers):

  • 84% use AI tools.
  • Only 29% trust AI output accuracy, down from 40% the year before.
  • 35% of Stack Overflow visits are now triggered by debugging AI-generated code.

5. Context degradation: when AI forgets the project

This one is specific to Android's multi-file structure. Developer Prabhakaran Thota described the pattern: "I was spending the first 3-4 messages of every session just correcting the AI. When a teammate started using the same tools, same problem. Claude Code, Copilot, Gemini — didn't matter. No project-level context."

Dwayne Parkinson documented his Gemini experience in Android Studio: "I had to wipe our entire conversation history, losing all context. Starting fresh felt like onboarding a new teammate. I fed Gemini all my project files manually and continued the experiment."

An Android project is an AndroidManifest.xml, Gradle scripts, modules, resources, and a navigation graph. Dozens of files that have to agree with each other. An AI that drops context mid-session starts generating code incompatible with what's already there. I've watched it suggest a ViewModel that references a repository that was deleted three prompts ago.


6. Security: what AI doesn't understand

Veracode (testing 100+ AI models):

  • 86% of AI-generated code samples fail XSS defense.
  • 88% are vulnerable to log injection.
  • CWE pass rate stuck at roughly 55% from 2023 to 2025. Models got smarter. They didn't get more secure.

Hardcoded secrets: AI regularly embeds API keys, database passwords, and tokens directly in source code. Per Snyk, 80% of developers admitted to bypassing security policies when using AI tools, and only 10% scan most AI-generated code before shipping.

Read those numbers again. Eight out of ten devs skip their own security policies when an AI writes the code. That's not a tool problem. That's a process problem the tool makes worse.


7. Zero test coverage: a time bomb

Vibe-coded apps almost never include automated tests. AI generates the feature but not the safety net. Every subsequent change, by AI or human, has no guardrails. A fix in one area breaks three others, and nobody finds out until a user complains.

Praveen Yadati recommends adding to every prompt: "and write a unit test for the ViewModel using MockK and Turbine for Flow testing." When you actually do this, the AI-generated tests often reveal broken logic that the implementation was hiding. Funny how that works.


8. Real incidents

Amazon (December 2025 – March 2026): At least four Sev-1 production incidents followed the AI-assisted development mandate, including a 6-hour outage with an estimated 6.3 million lost orders.

Google AI Studio (March 2026): Google launched the ability to vibe code native Android apps directly from AI Studio. The New Stack reported: "You can now build native Kotlin/Compose apps from prompts, including emulator, device testing, and Play Console upload." The same article also warned about code quality, security, and long-term maintainability. Bold move from Google. We'll see how it ages.


9. What to actually do about it

Start with architecture, not code. Before you prompt, write three lines: what the feature does, what data it needs, how the user interacts with it. It takes thirty seconds and prevents the AI from wandering off in random directions.

Generate in layers. Data layer first (model, DAO, repository), then domain, then UI. Not everything at once. When you feed the entire feature to the AI in one prompt, it cuts corners you won't notice until later.

Be explicit about your stack. "Jetpack Compose, Material 3, MVVM, Hilt, Room, StateFlow, Kotlin 2.0." Without this, the AI mixes APIs from different Android eras. You get LiveData next to Compose, deprecated coroutine APIs next to modern ones.

Always request error states and tests. Every prompt should include: "Include loading, empty, and error states. Write unit tests with MockK and Turbine." This alone catches a surprising amount of broken logic.

Review the dangerous parts manually. You don't need to read every line, but you must check data transformations, error handling, and state management. These are where AI makes silent mistakes. A method called mapResponseToDomain is where your bugs are.

Validate every dependency. Before you install an AI-suggested package, check it exists: npm info, pip index versions, Maven Central search. Don't install blind. A hallucinated package name is an invitation for someone to register it with malware.

Use grounding tools. MCP servers like AndroJack connect AI to current Android documentation. Prompts like agents.md control how the AI responds, not what it knows. The models are frozen in time. Your tooling shouldn't be.

Treat AI like a very fast junior dev. Praveen Yadati put it well: "Think of it less as replacing your skills and more as hiring a very fast, very eager, slightly overconfident junior developer. You still need to be the tech lead." If you wouldn't merge a junior's PR without reading it, don't merge AI code without reading it either.


10. Where we actually are

I've shipped AI-generated Android code. Some of it is still running. Some of it caused incidents I'd rather not think about.

The tool is a multiplier. A developer who knows the platform uses it to move 3–4× faster on boilerplate. A developer who doesn't understand Android gets buried under code that compiles and ships incorrect behavior.

The numbers don't leave much room for debate: 43% of AI code needs post-deployment fixes. 81% of enterprises see rising incidents. AI code security hasn't improved in three years. The problem isn't the AI. It's developers shipping code they didn't read.

If you don't know what the model generated, you can't own what it deploys.


Sources: Columbia DAPLab (2026), CodeRabbit State of AI vs Human Code Generation Report (2025), USENIX Security (2025), Lightrun Developer Survey (2026), Google DORA Report (2025), CloudBees AI Code Production Failures Survey (2026), Veracode GenAI Code Security Report, Stack Overflow Developer Survey (2025), Snyk Secure Adoption in the GenAI Era, Praveen Yadati / Medium, Vikas Sahani / dev.to, Zhenya Sedoy / dev.to, Modall (2026).

Top comments (1)

Collapse
 
mythex profile image
Mythex •

"Runnable over correct" is the pattern we fight most too (I'm building Mythex, an AI app builder). The guardrail that helped most: the agent isn't allowed to call something done without evidence, meaning a passing health check plus a screenshot of the running app. A claim with no proof attached gets checked again.

Two Android versions of the same idea that fit your list:

  • Turn "include loading, empty and error states" from a prompt into a check: a Compose UI test (or a Paparazzi/Roborazzi screenshot test) that renders each screen in all three states. The prompt asks for them; the test proves they exist.
  • For slopsquatting, Gradle's dependency verification (gradle/verification-metadata.xml with checksums) makes an unknown artifact fail the build, instead of relying on someone catching it in review.

Did the Room migration problem show up for you in production, or did you catch it before release?