DEV Community

Robin for Capawesome

Posted on Originally published at capawesome.io

How to Use Gemini Nano in a Capacitor App

Originally published on the Capawesome blog.

Gemini Nano is Google's on-device language model for Android, and a Capacitor app can prompt it without an API key, a server or a per-request bill. The Capacitor LLM plugin runs it through AICore, the Android system service that manages the model on supported phones. This guide to Gemini Nano in a Capacitor app walks through the device list, the availability check, the model download with a determinate progress bar, text generation and streaming, and the quota and foreground rules that AICore enforces. The plugin is part of Capawesome Insiders, a paid subscription. If you only need summaries, proofreading or rewriting on Android, the free Capacitor ML Kit GenAI plugins cover those tasks, and a later section compares the two options.

Thousands of teams ship faster with Capawesome Cloud Native Builds and Live Updates

Key takeaways:

  • AICore stores one shared copy of Gemini Nano per device, so the model can already be available before your app downloads anything.
  • The Capacitor LLM plugin's getAvailability() never rejects and reports four statuses on Android: available, downloadable, downloading and unavailable.
  • During downloadModel(), the downloadProgress event reports a value between 0 and 1, so the download can drive a determinate progress bar.
  • AICore answers too many requests with a BUSY quota error that calls for exponential backoff, and allows inference only while your app is the top foreground app.
  • Prompt API devices include the Pixel 9, Pixel 10, Galaxy S26 and OnePlus 15, and each needs API level 26 or higher and a locked bootloader.

How Gemini Nano runs

Gemini Nano runs inside AICore, which Google describes as "an Android system service that enables on-device execution of GenAI foundation models". Your app never ships the model, because AICore handles its distribution and its updates. Apps reach it through Google's ML Kit GenAI APIs, which come in two kinds:

  • Prompt API: takes free-form prompts. This is the API the Capacitor LLM plugin calls on Android.
  • Feature-specific APIs: Summarization, Proofreading, Rewriting and Image Description each handle one fixed task. A sixth API covers speech recognition.

Google offers the Prompt API and the four feature-specific APIs as beta, "not subject to any SLA or deprecation policy", and Speech Recognition as alpha. The Prompt API moved from alpha to beta on January 28, 2026.

Because the APIs run on the device, Google states that input, inference and output data are processed locally and that the features keep working without a reliable internet connection. Prompts and responses stay on the device. Google's ML Kit GenAI terms still require you to inform users about Google's processing of metrics data, so cover that in your privacy policy. According to Google's July 2026 Android Developers blog post, Gemini Nano now runs on over 140 million devices.

The Capacitor LLM plugin exposes the Prompt API to your web code through one TypeScript interface. It is part of Capawesome Insiders, a paid subscription. To install the Capacitor LLM plugin, please refer to the Installation section in the plugin documentation.

Supported devices

Gemini Nano runs only on the devices Google lists for each ML Kit GenAI API, and the Prompt API device list is the one that applies to the Capacitor LLM plugin. It names phones and tablets from 14 brands and groups them by the Nano version each device runs:

Gemini Nano version Example devices
nano-v2 OnePlus 13, Xiaomi 15, Galaxy Z Fold7, Honor Magic 7, realme GT 7 Pro
nano-v3 Pixel 9 and Pixel 10 (including Pro, Pro XL and Pro Fold), Galaxy S26 series, OnePlus 15, OPPO Find X9, vivo X300
nano-v4 Pixel 11 (including Pro, Pro XL and Pro Fold), Galaxy Z Flip8, Galaxy Z Fold8

The version affects the output. Google notes that "different versions of Gemini Nano may return different output from the same prompt", so a prompt tuned on a Pixel 10 needs a second look on a nano-v2 phone such as the Xiaomi 15.

Every device on the list also has to meet the requirements Google repeats on each API page. It must run Android API level 26 or higher, and the APIs are not supported on devices with an unlocked bootloader. Google names MediaTek Dimensity, Qualcomm Snapdragon and Google Tensor as the supported chip platforms.

On iOS, the same plugin talks to Apple Intelligence instead, and How to Use Apple Intelligence in a Capacitor App covers that side.

Check availability

The Capacitor LLM plugin reports whether Gemini Nano is ready with getAvailability(), a method that never rejects. On the web, or anywhere else without a system model, it resolves with unavailable, so you can call it unconditionally at startup. Android reports four of the plugin's statuses:

Status Meaning on Android What the screen shows
available The model is downloaded and ready. The feature.
downloadable The device supports Gemini Nano, but the model is not on it yet. A download button.
downloading A download is running. A waiting indicator.
unavailable Gemini Nano cannot run on this device right now. Nothing, or a fallback.

The following function turns the status into one of four UI states and keeps it current. It reads the status once, then subscribes to the availabilityChange event, which fires when the status changes, for example when a download completes:

import { Llm } from '@capawesome-team/capacitor-llm';
import type { AvailabilityStatus } from '@capawesome-team/capacitor-llm';

type AssistantState = 'ready' | 'needs-download' | 'downloading' | 'hidden';

const toAssistantState = (status: AvailabilityStatus): AssistantState => {
  switch (status) {
    case 'available':
      return 'ready';
    case 'downloadable':
      return 'needs-download';
    case 'downloading':
      return 'downloading';
    default:
      return 'hidden';
  }
};

const observeAssistantState = async (onChange: (state: AssistantState) => void) => {
  const { status } = await Llm.getAvailability();
  onChange(toAssistantState(status));
  return Llm.addListener('availabilityChange', event => {
    onChange(toAssistantState(event.status));
  });
};
Enter fullscreen mode Exit fullscreen mode

The plugin watches the status only while at least one listener is attached, so call remove() on the returned handle when the screen closes. The default branch covers unavailable as well as the statuses that only iOS reports.

Download the model

Gemini Nano has to be downloaded to the device before a Capacitor app can use it, but the download may already have happened. AICore keeps one shared copy of the model for all apps on the device, so if another app triggered the download earlier, getAvailability() reports available and your app can skip this step. The four feature-specific ML Kit APIs also fetch a small adapter model per feature on top of that shared base model.

When the status is downloadable, downloadModel() starts the download and resolves once it completes. Along the way, the downloadProgress event reports progress as a value between 0 and 1, which maps directly onto a `` element. Google publishes no download size, and its Firebase documentation says the download time "depends on many factors, including your network", so show the bar instead of a spinner.

Start the download from a clear user action, such as an "Enable reply drafts" button, and offer it before the user reaches the screen that needs the model:

`typescript
import { Llm } from '@capawesome-team/capacitor-llm';

const downloadGeminiNano = async (onProgress: (progress: number) => void) => {
const listener = await Llm.addListener('downloadProgress', event => {
onProgress(event.progress);
});
try {
await Llm.downloadModel();
} finally {
await listener.remove();
}
};

const enableReplyDrafts = async () => {
showProgressBar(0);
try {
await downloadGeminiNano(showProgressBar);
showReplyDrafts();
} catch {
showRetryButton();
}
};
`

The screen can also open in the downloading state. That happens when your app started the download in an earlier session, or when another app started it. Show an indeterminate indicator in that case, and let the availabilityChange listener from the previous section switch the screen to ready once the status turns available.

On a supported phone, unavailable is not always final. After a device is set up or reset, AICore first has to download its configuration, and Google's setup notes say that this "usually takes a few minutes to a few hours" while the device is online. Check the status again on the next app start instead of hiding the feature for good. If downloadModel() itself rejects, the retry button above covers it. Google's advice for network errors during a download is to keep the connection, wait a few minutes and retry.

Generate and stream

Text generation in the Capacitor LLM plugin happens inside a chat, which carries instructions (the system prompt) and the conversation history. The example in this section drafts replies to customer reviews in a shop app, so the owner edits a suggestion instead of starting from an empty field. createChat(...) sets the tone once, and generateText(...) returns a complete draft:

`typescript
import { Llm } from '@capawesome-team/capacitor-llm';

const createReviewReplyChat = async () => {
const { id } = await Llm.createChat({
instructions:
'You draft replies to customer reviews for an online shop. Thank the customer, respond to each point they raise, stay under 80 words and never promise refunds.',
temperature: 0.3,
maxOutputTokens: 200,
});
return id;
};

const draftReply = async (chatId: string, review: string) => {
const { text } = await Llm.generateText({
chatId,
prompt: Draft a reply to this review:\n\n${review},
});
return text;
};
`

The Android API has no native multi-turn sessions, so the plugin keeps the chat history in memory and includes it in every prompt. That history counts toward the input limit of about 4,000 tokens. Create one chat per review, and delete it with deleteChat(...) once the reply is sent.

The other two Android limits apply per chat or per request: temperature must be between 0.0 and 1.0, and maxOutputTokens can be at most 4096. A low temperature such as 0.3 gives more predictable drafts. Only one generation runs per chat at a time, and a second call rejects with GENERATION_IN_PROGRESS, so disable the button while a draft is being written.

A full draft takes several seconds to generate. Google's May 2025 benchmark measured a decode speed of 11 tokens per second on a Pixel 9 Pro, so a reply of a few sentences takes several seconds. streamText(...) shows the draft while Gemini Nano writes it. The plugin emits a textChunk event per piece, tagged with its chatId, and cancelGeneration(...) backs a discard button:

`typescript
import { Llm } from '@capawesome-team/capacitor-llm';

const streamReply = async (chatId: string, review: string, onDraft: (draft: string) => void) => {
let draft = '';
const listener = await Llm.addListener('textChunk', event => {
if (event.chatId === chatId) {
draft += event.text;
onDraft(draft);
}
});
try {
const { text } = await Llm.streamText({
chatId,
prompt: Draft a reply to this review:\n\n${review},
});
return text;
} catch (error) {
if ((error as { code?: string }).code === 'GENERATION_CANCELED') {
return null;
}
throw error;
} finally {
await listener.remove();
}
};

const discardReply = async (chatId: string) => {
await Llm.cancelGeneration({ chatId });
};
`

A canceled generation rejects with GENERATION_CANCELED, which the example turns into null. On Android, cancellation is best-effort, and a few more chunks may arrive after the user taps discard, so have the UI ignore updates for a draft it has already thrown away.

Quota and foreground rules

AICore decides how often and when your app may run Gemini Nano, and two rules from Google's ML Kit GenAI overview apply to every Capacitor app that uses it. The first one is a quota:

> AICore enforces an inference quota per app. Making too many GenAI API requests in a short period will result in an ErrorCode.BUSY response. When receiving such an error, consider using exponential backoff to retry the request. Also, ErrorCode.PER_APP_BATTERY_USE_QUOTA_EXCEEDED can be returned if an app exceeds a long-duration quota (e.g. daily quota).

The second one limits inference to the foreground, as the background usage section of the same page states:

> GenAI API inference is permitted only when the app is the top foreground application. Using the API when the app is not in the foreground, including using a foreground service, will result in an ErrorCode.BACKGROUND_USE_BLOCKED response.

A Capacitor app runs its web code inside the app's own activity, so a generation that starts while the user looks at the screen meets the second rule. A retry timer, a queued request or a push handler that fires after the user switched apps does not.

The Capacitor LLM plugin does not pass ML Kit's error code or its retry delay for quota errors to JavaScript. A quota or background rejection arrives as GENERATION_FAILED with the platform's message, the same code the plugin uses for an exceeded context window and for content blocked by safety guardrails. Matching on the message text is fragile, so treat GENERATION_FAILED as retryable with exponential backoff, cap the attempts, and stop retrying once the app is in the background. The helper below uses the official App plugin to check the app state:

`typescript
import { App } from '@capacitor/app';
import { Llm } from '@capawesome-team/capacitor-llm';

const MAX_ATTEMPTS = 4;

const delay = (milliseconds: number) => new Promise(resolve => setTimeout(resolve, milliseconds));

const generateWithBackoff = async (chatId: string, prompt: string) => {
for (let attempt = 1; ; attempt++) {
try {
const { text } = await Llm.generateText({ chatId, prompt });
return text;
} catch (error) {
const isRetryable = (error as { code?: string }).code === 'GENERATION_FAILED';
if (!isRetryable || attempt === MAX_ATTEMPTS) {
throw error;
}
await delay(1000 * 2 ** attempt);
const { isActive } = await App.getState();
if (!isActive) {
throw error;
}
}
}
};
`

The helper waits two, four and eight seconds between attempts and gives up after the fourth. Neither a daily quota nor a chat whose history outgrew the input limit recovers within that window, so show a "Try again" button when the helper throws and create a new chat for the next attempt. If your screen queues several drafts, pause the queue when the appStateChange event reports isActive: false. For streams, Google's error reference recommends removing an interrupted result from the UI, so clear the partial draft when streamText(...) rejects.

Testing on devices

Testing Gemini Nano in a Capacitor app requires a phone from the Prompt API list, ideally one per Nano version you target. Frequent test runs in a short period can trigger the BUSY quota error from the previous section. Google's AICore Developer Preview program, which you join through the aicore-experimental Google group and the AICore testing program on Google Play, adds a "Bypass quota limits" toggle to the AICore settings. The toggle exists only on Pixel phones and applies to every app on the device that uses AICore, so switch it off again before you test how your backoff behaves.

The other thing to test is the SDK version. The plugin pulls in Google's com.google.mlkit:genai-prompt dependency with a default of 1.0.0-beta2, while Google's latest release is 1.0.0-beta4 from July 21, 2026. Google's release notes list compatibility fixes for Gemini Nano v4 in that release, so if the Pixel 11 or the Galaxy Z Flip8 and Z Fold8 are in your target group, test generation and streaming on one of them before you ship. The $mlkitGenaiPromptVersion variable in android/variables.gradle overrides the version. Treat that change like any other dependency upgrade and run your tests against it.

ML Kit GenAI plugins

The Capacitor ML Kit project maintains six free, open-source GenAI plugins that call Google's ML Kit GenAI APIs directly on Android, one per API:

Like the rest of Capacitor ML Kit, these plugins are unofficial and not affiliated with Google. Their lifecycle mirrors the one above under different names. checkFeatureStatus(...) returns AVAILABLE, DOWNLOADABLE, DOWNLOADING or UNAVAILABLE, but unlike getAvailability() it rejects on iOS and the web, so check the platform first. downloadFeature(...) fetches the feature, and its downloadProgress event carries only totalBytesDownloaded. Without the total size, no percentage is possible, so show an indeterminate bar with a byte counter. For the four feature-specific APIs the explicit download is optional, since Google states that the first inference request triggers it as well. Calling it first is what gives you progress events. These APIs also have their own device list, which includes phones that are missing from the Prompt list.

Google's rule of thumb is to "use Prompt API when you need more customization and flexibility, and use the feature-specific APIs for standard tasks that don't require complex logic." The Capacitor LLM plugin adds chats, textChunk streaming, cancellation and iOS support on top of the Prompt API. Capacitor ML Kit 8.2.0 shows the check, download and summarize flow of the free plugins with code.

FAQ

Do I need to download Gemini Nano before using it?

Yes, the model has to be on the device, but it may already be there. AICore stores one shared copy for all apps, so if another app downloaded it, getAvailability() reports available. Otherwise the status is downloadable, and downloadModel() fetches the model while the downloadProgress event reports a value between 0 and 1.

Does Gemini Nano work offline?

Yes, once the model is on the device. Google states that input, inference and output data are processed locally and that functionality "remains the same without reliable internet connection". The model download and AICore's configuration after a device setup need a connection.

Which devices support Gemini Nano in a Capacitor app?

Gemini Nano runs on the devices from Google's Prompt API device list, such as the Pixel 9, Pixel 10 Pro, Pixel 11, Galaxy S26 Ultra, OnePlus 15 and Xiaomi 15. Each one needs Android API level 26 or higher and a locked bootloader. Check getAvailability() at runtime instead of maintaining your own list.

Why does a generation reject with GENERATION_FAILED?

The Capacitor LLM plugin uses GENERATION_FAILED for an exceeded context window, for content blocked by safety guardrails, and for AICore's quota and foreground rules. The error message carries the platform's reason. Retry with exponential backoff while the app is in the foreground, and create a new chat if the retries keep failing.

Can I run Gemini Nano in the background?

No. Google permits GenAI inference only while the app is the top foreground application, and a foreground service does not change that. Start generations from user actions on screen, and move work that arrives in the background, such as a push notification, to the next time the user opens the app.

Is the Capacitor LLM plugin free?

No. It is part of the Capawesome Insiders subscription, which covers every Insiders plugin. Gemini Nano itself costs nothing per request, and the six Capacitor ML Kit GenAI plugins are free and open source.

Related posts

Stay in the loop

Google changed the ML Kit GenAI APIs several times in 2026, from the Prompt API beta in January to the Gemini Nano v4 fixes in July. Plugin releases that follow those changes, and guides like this one, go out in the Capawesome newsletter first.

Subscribe to the Capawesome Newsletter

Conclusion

Pick the plugin by the shape of your feature. If a summary, a proofread or a rewrite on Android covers it, start with the free ML Kit GenAI plugin for that task. If you need free-form prompts, multi-turn chats or the same code on iOS, build on the Capacitor LLM plugin, and show the feature on every Android device where getAvailability() does not report unavailable, with a fallback everywhere else. The chat and streaming code from the review example runs unchanged on iPhones with Apple Intelligence, as How to Use Apple Intelligence in a Capacitor App shows.

If you have any questions, join us on the Capawesome Discord server. To stay updated on the latest news, subscribe to the Capawesome newsletter.

Top comments (0)