DEV Community

Baur
Baur

Posted on

BYOK: Why I Let Users Bring Their Own LLM Key Instead of Running My Own API

There is a fairly predictable moment when an indie developer adds an LLM to a product.

The prototype works. The model produces useful output. The API call costs a few cents.

Then you start thinking about what happens when there are 100 users. Or 1,000.

The architectural question becomes less about whether the LLM works and more about who pays for every token.

There are two obvious approaches.

The first is the traditional SaaS model: users talk to my backend, my backend calls OpenAI, Anthropic, or another provider, and I pay the inference bill. I can hide the complexity from users and potentially charge more than my variable cost.

The second is BYOK — Bring Your Own Key. Users provide their own API key, and the desktop application calls the LLM provider directly.

For my desktop application, I chose the second approach.

It isn't universally better. It moves several problems rather than eliminating them. But for a small developer building a desktop application, the tradeoff can make a lot of sense.

What BYOK actually means

In this architecture, there is no application API that proxies LLM requests.

The simplified flow looks like this:

┌──────────────────── Desktop application ────────────────────┐
│                                                             │
│  Audio ──> local whisper.cpp transcription                  │
│                         │                                   │
│                         v                                   │
│                 application logic                            │
│                         │                                   │
│                         v                                   │
│                 provider abstraction                         │
│                         │                                   │
└─────────────────────────┼───────────────────────────────────┘
                          │
                          │ HTTPS
                          v
                 ┌─────────────────┐
                 │ OpenAI /        │
                 │ Anthropic /     │
                 │ OpenRouter      │
                 └─────────────────┘
Enter fullscreen mode Exit fullscreen mode

The important part is what is not in the diagram:

Desktop app ──X──> my backend ──> LLM provider
Enter fullscreen mode Exit fullscreen mode

There is no intermediary API operated by me.

The user enters an API key into the desktop application. The application stores that key in the operating system's secure credential storage and retrieves it when an LLM request is needed.

The key isn't uploaded to my servers.

That's an important distinction because "stored locally" and "stored securely" are not automatically the same thing.

Don't put API keys in ordinary application storage

An API key shouldn't normally live in ~/.config/my-app/config.json, localStorage, or an unencrypted SQLite database.

A desktop application has access to platform credential facilities specifically designed for secrets.

On macOS, that means Keychain. On Windows, an application can use Windows Credential Manager. On Linux, Secret Service API. Tauri provides mechanisms for interacting with platform-native secure storage, and Rust libraries like keyring can also provide an abstraction over the relevant credential stores.

Conceptually, the application does this:

// save_secret("openai", api_key)
// load_secret("openai")
let key = load_secret("openai")?;
let client = ProviderClient::new(key);
client.generate(request).await?;
Enter fullscreen mode Exit fullscreen mode

The important property isn't the exact API call. It's the storage boundary.

The API key belongs to the user and remains under the operating system's credential-management layer.

I also avoid logging the key, including indirectly through request-debugging output. This is a surprisingly easy mistake to make:

println!("Request headers: {:?}", headers);
Enter fullscreen mode Exit fullscreen mode

If the authorization header is part of headers, you've just created a credential leak in a log file. A safer approach is to explicitly redact sensitive fields before logging:

fn log_request(method: &str, url: &str) {
    println!("{method} {url}");
}
Enter fullscreen mode Exit fullscreen mode

Why I chose BYOK

The biggest reason is simple: I don't want my business model to depend on predicting another company's token costs.

With a conventional backend architecture, every active user creates some amount of variable infrastructure cost. The application might charge $X per month, while actual LLM usage depends on how frequently the application is used, which model the user selects, prompt size, response size, retries, context size, and provider pricing changes.

You can either absorb the variability into your margin or introduce usage limits. Neither is particularly attractive for a small product when the core application can already run without a backend.

With BYOK, the economic relationship is different — the user pays the provider for LLM usage and pays me for the application, rather than paying me for both infrastructure and LLM usage combined.

That doesn't make the product free to operate. There are still costs for distribution, documentation, development, payment processing, and support. But LLM inference isn't a variable cost I have to estimate for every customer.

I also don't have to build usage infrastructure

A hosted LLM API usually creates another set of engineering questions: how many requests has this user made, what's their monthly token allowance, what happens when they reach it, how do I prevent abuse, how do I rate-limit requests, what happens when provider pricing changes.

Those are perfectly reasonable problems for a SaaS company to solve. I just didn't want to build them for this application.

With BYOK, the provider already handles the user's account, authentication, quotas, billing, and provider-side rate limits. The application mostly needs to handle provider errors correctly.

Users get to choose their provider

The application isn't tightly coupled to one model vendor. I can expose a small provider abstraction:

#[async_trait]
trait LlmProvider {
    async fn generate(
        &self,
        request: LlmRequest,
    ) -> Result<LlmResponse, LlmError>;
}
Enter fullscreen mode Exit fullscreen mode

Then implementations handle different APIs:

struct OpenAiProvider { api_key: SecretString }
struct AnthropicProvider { api_key: SecretString }
struct OpenRouterProvider { api_key: SecretString }
Enter fullscreen mode Exit fullscreen mode

The rest of the application doesn't need to know how each provider authenticates or structures its HTTP request:

match provider {
    Provider::OpenAI => openai.generate(request).await,
    Provider::Anthropic => anthropic.generate(request).await,
    Provider::OpenRouter => openrouter.generate(request).await,
}
Enter fullscreen mode Exit fullscreen mode

The real implementation needs more abstraction than this, particularly around model identifiers, streaming, errors, and token usage differences between provider APIs. But keeping that boundary explicit has been useful — I can add another provider without changing the parts of the application that consume generated text.

The downside: onboarding gets worse

BYOK isn't a free lunch. The biggest UX disadvantage is obvious: users have to obtain an API key.

A hosted service can give someone a text field for email and password and make everything else invisible. BYOK introduces another sequence: create an account with a model provider, find the API-key section, create a key, copy it, open the desktop application, paste it, make sure the selected model is available, make sure the account has sufficient credit.

That's substantially more friction.

It also creates a support category that doesn't exist in the same way with a hosted API: "My key doesn't work." That sentence can mean many things — invalid key, no balance, unavailable model, rate limit, context limit, temporary provider outage.

So error messages need to be much more useful than LLM request failed. I try to preserve the provider's meaningful error category while avoiding exposing credentials or unnecessary implementation details:

Authentication failed.
Check that your API key is valid for this provider.
Enter fullscreen mode Exit fullscreen mode

is much more useful than HTTP 401. And:

Provider rate limit reached.
Try again later or check your provider account.
Enter fullscreen mode Exit fullscreen mode

is more actionable than Request failed.

I gave up margin in exchange for cost transparency

If I operated the LLM API myself, I could potentially buy inference wholesale and charge users a higher price. BYOK removes that possibility — the user pays the model provider directly.

So instead of trying to hide token economics behind a subscription, I treat the LLM cost as an explicit part of the architecture. The application can show which provider and model are configured, and the user can see that requests are sent directly to that provider. There is no mysterious internal "AI allowance" whose cost I have to recover through an opaque pricing model.

The tradeoff, roughly:

Hosted API:
    simpler onboarding
    centralized control
    potential inference margin
    developer pays variable LLM costs

BYOK:
    more onboarding friction
    less centralized control
    no inference margin
    user pays variable LLM costs
Enter fullscreen mode Exit fullscreen mode

Neither side is universally correct. The important thing is deciding which set of problems you actually want to own.

Tauri makes the boundary particularly useful

For a desktop application, the architecture can remain relatively small. The Rust side handles things that shouldn't be implemented casually in the frontend, including access to operating-system facilities and request orchestration. The UI just asks the Rust layer for an operation:

Frontend
   │
   │ invoke()
   v
Tauri command
   │
   ├── retrieve credential
   ├── construct provider request
   └── perform HTTPS request
             │
             v
        LLM provider
Enter fullscreen mode Exit fullscreen mode

I don't need to turn the frontend into a credential-management system, and I don't need a backend solely to hide a provider key — because there is no application-owned provider key in the first place.

If the architecture were Frontend -> my backend -> provider, my backend would need credentials for the provider. With BYOK (Desktop -> provider), the credential belongs to the user.

Keep provider-specific code behind one interface

The provider abstraction has another advantage: it prevents provider-specific details from leaking through the application. The application constructs an internal request:

struct LlmRequest {
    model: String,
    system_prompt: String,
    messages: Vec<Message>,
    max_tokens: Option<u32>,
}
Enter fullscreen mode Exit fullscreen mode

Each provider implementation converts that into its own request format. That gives one place to deal with authentication headers, endpoint URLs, request schemas, response parsing, streaming, error formats, and model identifiers — and it makes testing easier, since the application can test its behavior against a mock LlmProvider without making real API calls.

What I would do differently for a larger SaaS

I wouldn't assume BYOK is automatically the right architecture at scale. For a larger product, a managed backend may be preferable — centralized API access makes it easier to provide a consistent experience, enforce product-level limits, manage abuse, aggregate billing, and control which models are available. It also creates infrastructure and security responsibilities that a BYOK desktop application can avoid.

The architecture should follow the product's actual requirements rather than treating BYOK as a philosophy.

For my particular situation the calculation was simpler: I already had a native desktop application, some processing happens locally, and the application needs occasional LLM requests. I didn't need a persistent application backend merely to forward those requests. So I decided not to introduce one.

The boundary I ended up with

The resulting architecture is deliberately boring: audio goes through local whisper.cpp transcription, application logic and a provider abstraction sit on top, the key comes from the OS keychain, and the request goes straight to the LLM provider.

The key never needs to pass through my infrastructure. The audio transcription doesn't need my infrastructure either. And the LLM provider relationship belongs to the person using the application.

For an indie developer, that's a meaningful reduction in operational surface area. The price is more friction at onboarding and more responsibility for explaining provider configuration.

For me, that tradeoff is preferable to operating an LLM proxy whose main job is to turn my users' unpredictable token consumption into my unpredictable infrastructure bill.


If you're curious about the architecture in production, the resulting app is SyntaxCue Pro.

Top comments (0)