<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Rio Consulting</title>
    <description>The latest articles on DEV Community by Rio Consulting (@rio_consulting).</description>
    <link>https://dev.to/rio_consulting</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4165082%2Fd62e6fbf-63bf-4157-9b25-0a10db22af34.png</url>
      <title>DEV Community: Rio Consulting</title>
      <link>https://dev.to/rio_consulting</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/rio_consulting"/>
    <language>en</language>
    <item>
      <title>Private LLM Options: Local, Cloud, or Confidential?</title>
      <dc:creator>Rio Consulting</dc:creator>
      <pubDate>Tue, 06 Oct 2026 01:22:01 +0000</pubDate>
      <link>https://dev.to/rio_consulting/private-llm-options-local-cloud-or-confidential-4h1p</link>
      <guid>https://dev.to/rio_consulting/private-llm-options-local-cloud-or-confidential-4h1p</guid>
      <description>&lt;p&gt;You want a private LLM, and the obvious move is to buy a box and run it yourself. Before you do, it's worth understanding what that hardware really costs, what happens when it's out of date in two years and you're still paying it off, and what the other options are. There are three routes: your own hardware, a cloud provider under contract, or confidential inference on hardware that can't read the prompt in the first place.&lt;/p&gt;

&lt;p&gt;This post walks through your options for Private LLMs, what each costs, and a deeper look into AI privacy when you're sharing the agent with the rest of your team.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What does "private" mean for an LLM?&lt;/li&gt;
&lt;li&gt;Option 1: a local LLM on your own hardware&lt;/li&gt;
&lt;li&gt;Option 2: a cloud provider with a contract&lt;/li&gt;
&lt;li&gt;Option 3: confidential inference with SayGM&lt;/li&gt;
&lt;li&gt;How do the three compare?&lt;/li&gt;
&lt;li&gt;The other side: private memory in shared agents&lt;/li&gt;
&lt;li&gt;Common questions about private LLMs&lt;/li&gt;
&lt;li&gt;Which one should you pick?&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What does "private" mean for an LLM?
&lt;/h2&gt;

&lt;p&gt;The term Private LLM / AI is used across various contexts&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Physical privacy&lt;/strong&gt; means the prompt never leaves a machine you control. &lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Contractual privacy&lt;/strong&gt; means it leaves, but the provider has agreed in writing not to train on it, share it, or keep it past a set window. &lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Technical privacy&lt;/strong&gt; means it leaves, but it's processed inside hardware that the operator can't see into, and you can check that yourself.&lt;/p&gt;

&lt;p&gt;Who is able to read the prompt, and is the thing stopping them a wall, a promise, or a chip?&lt;/p&gt;

&lt;h2&gt;
  
  
  Option 1: a local LLM on your own hardware
&lt;/h2&gt;

&lt;p&gt;A local LLM is the cleanest answer to the privacy question. Nothing leaves the building, there's no provider to trust, and no terms of service to reread every time they change.&lt;/p&gt;

&lt;p&gt;A current hardware roundup puts an RTX 5090 at $1,999 MSRP (USD) with 32 GB of VRAM, managing roughly 25 to 30 tokens per second on a 70B model, and a 128 GB Mac Studio at $3,500 to $4,000 for similar speeds (&lt;a href="https://fungies.io/best-hardware-local-llms-2026-3/" rel="noopener noreferrer"&gt;fungies.io hardware guide&lt;/a&gt;). That's a perfectly good setup for one person.&lt;/p&gt;

&lt;p&gt;Tradeoffs are:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Concurrency.&lt;/strong&gt; One card serving 25 tokens per second is one conversation. An agent that fans out into ten parallel calls is queueing on that same card.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Your Wife (or Husband).&lt;/strong&gt; Will they believe that you really need a kitted out gaming PC for "work"?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Model ceiling.&lt;/strong&gt; A self hosted LLM means open-weight models, usually quantised to fit. They're good, and getting better fast, but the largest frontier models aren't available to download at any price.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Flexibility.&lt;/strong&gt; One year down the line, if the latest models are just outside of your reach you'll need to buy a whole new machine again. Realistically, you'll be stuck with this hardware for 4+ years. At least it will still run Crysis when the spouse isn't looking.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Your time.&lt;/strong&gt; Drivers, serving stack, updates, the box that falls over on a Saturday. &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So local is the most private option and often the cheapest per token once the hardware is paid off. It is a large upfront cost and is likely to get outdated soon.&lt;/p&gt;

&lt;h2&gt;
  
  
  Option 2: a cloud provider with a contract
&lt;/h2&gt;

&lt;p&gt;This is where most privacy-first enterprises land. The big three clouds all sell access to frontier models under enterprise terms that are a long way from pasting things into a consumer chat app. The privacy here is contractual, so the detail is in the documents. &lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AWS Bedrock&lt;/strong&gt; states that model providers don't have access to customer prompts and completions. Its &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/data-retention.html" rel="noopener noreferrer"&gt;data retention documentation&lt;/a&gt; describes a zero data retention mode where "no request or response data is written to durable storage by AWS or shared with the model provider". The same page also describes a review mode that applies to some of the newest frontier models, where prompts and completions are "retained within the AWS boundary for up to 30 days and may be reviewed by AWS". Which mode you get depends on the model you pick.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Microsoft Azure&lt;/strong&gt; is similarly direct. Its &lt;a href="https://learn.microsoft.com/en-us/azure/foundry/responsible-ai/openai/data-privacy" rel="noopener noreferrer"&gt;data privacy page&lt;/a&gt; says prompts and completions "are NOT available to OpenAI or other providers", and are not used to train foundation models without your permission. By default there is an abuse monitoring store where prompts can be held for human review. Managed customers can apply to have that modified, and you can confirm it's off by checking for a &lt;code&gt;"ContentLogging": "false"&lt;/code&gt; attribute on the resource.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Google Cloud&lt;/strong&gt; documents its own &lt;a href="https://docs.cloud.google.com/vertex-ai/generative-ai/docs/vertex-ai-zero-data-retention" rel="noopener noreferrer"&gt;zero data retention path for Vertex AI&lt;/a&gt;, which again is something you configure and request rather than something you get by default.&lt;/p&gt;

&lt;p&gt;The provider's staff and systems are technically able to see the prompt; and the thing stopping them is policy. For a lot of teams that's fine. For source code, client documents or anything you'd lose sleep over, it's worth knowing that's the shape of the guarantee.&lt;/p&gt;

&lt;p&gt;This approach also comes with a costly enterprise contract to pay. Best for hospitals and other large public organisations who need someone to point the finger at.&lt;/p&gt;

&lt;h2&gt;
  
  
  Option 3: a private LLM gateway
&lt;/h2&gt;

&lt;p&gt;The third option is newer and sits between the other two. A private LLM gateway is an API you call like any cloud endpoint, except the gateway runs inside a trusted execution environment (TEE). The short version of a TEE is that the chip encrypts the memory of whatever is running inside it, so the host machine and the people operating it can't read what's in there. You rent the hardware like cloud, but the privacy is enforced by the chip, closer to local.&lt;/p&gt;

&lt;p&gt;For example, &lt;a href="https://saygm.com/privacy" rel="noopener noreferrer"&gt;SayGM&lt;/a&gt; runs its gateway in an Intel TDX confidential VM and explain very well &lt;a href="https://saygm.com/blog/trusted-execution-environments-explained-for-ai-teams" rel="noopener noreferrer"&gt;how trusted execution environments work for AI teams&lt;/a&gt; if you want the mechanics.&lt;/p&gt;

&lt;p&gt;There are two different guarantees on a gateway like this and it matters not to blur them:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Confidential models.&lt;/strong&gt; Open-weight models run inside the enclave itself. The prompt is only decrypted inside the hardware running the model, so the gateway can't read it, and neither can the host.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Frontier models.&lt;/strong&gt; The route through the gateway is sealed, but if Anthropic or OpenAI are serving your models they will still see the prompts - though there are often options to mask out personal information.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The benefits of private LLM Gateways are:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;No upfront cost, or expensive contracts&lt;/li&gt;
&lt;li&gt;Pay per use&lt;/li&gt;
&lt;li&gt;Hardware-level verification that your data isn't being logged&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;However, they aren't able to offer truly private enterprise access to frontier models like Amazon Bedrock can.&lt;/p&gt;

&lt;h2&gt;
  
  
  The other side: private memory in shared agents
&lt;/h2&gt;

&lt;p&gt;Everything above is about protecting a prompt on its way to a model. Once you give an agent memory and let more than one person use it, now you have to control leaking internally within users.&lt;/p&gt;

&lt;p&gt;A shared agent with memory means that anybody with access can attempt to coax information that other people had shared out of the agent. Say you set up a fully private AI Agent for your company, but then the agent begins sharing private conversations from the HR team to the staff it concerns.&lt;/p&gt;

&lt;p&gt;The research on this is not reassuring. A 2026 benchmark from ServiceNow AI Research and Mila, &lt;a href="https://arxiv.org/abs/2607.05318" rel="noopener noreferrer"&gt;PiSAs&lt;/a&gt;, tested multi-user agent systems and found at least one privacy violation in over 75% of test runs, with shared memory making things markedly worse. Separately, the &lt;a href="https://arxiv.org/abs/2502.13172" rel="noopener noreferrer"&gt;MEXTRA paper&lt;/a&gt; showed that private interactions stored in agent memory can be pulled out with black-box prompts alone.&lt;/p&gt;

&lt;p&gt;You can run the most private LLM in the world and still have your agent tell the intern what the founder asked it last week. A few things that help:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Scope memory per user by default, and make shared memory something a person opts a specific note into.&lt;/li&gt;
&lt;li&gt;Enforce access in the storage layer. A line in the system prompt saying "don't share private data" is the approach the benchmark found insufficient.&lt;/li&gt;
&lt;li&gt;Redact before you write. Secrets and personal data that never reach memory can't be recalled from it.&lt;/li&gt;
&lt;li&gt;Treat anything in shared memory as readable by everyone who can talk to the agent, because in practice it is.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Common questions about private LLMs
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Is a local LLM completely private?
&lt;/h3&gt;

&lt;p&gt;The inference is. Whether the whole setup is depends on what else the machine talks to: telemetry in the tools around it, web search plugins, and any memory store your agent writes to.&lt;/p&gt;

&lt;h3&gt;
  
  
  Do cloud providers train on my prompts?
&lt;/h3&gt;

&lt;p&gt;Under enterprise API terms, generally not. Azure says so in writing in the page linked above. Retention is the more useful thing to check, since prompts can be stored for abuse review even when they aren't used for training.&lt;/p&gt;

&lt;h3&gt;
  
  
  What is the cheapest private LLM option?
&lt;/h3&gt;

&lt;p&gt;At low volume, a pay-per-token confidential tier, because there's no hardware to buy. At high, steady volume on a single model, owned hardware can win once it's paid off, if you don't count your own time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which one should you pick?
&lt;/h2&gt;

&lt;p&gt;If it's just you and the workload is steady, a local LLM is hard to argue with. If you need the top frontier models with an SLA and a contract your legal team already understands, go cloud and read the retention pages properly. If you're a small team running agents over data you'd rather nobody could read, and you'd like to verify that instead of trusting it, a private LLM gateway is the option built for that.&lt;/p&gt;

&lt;p&gt;Whichever private LLM route you take, sort the memory question out before the agent goes multi-user. That's the leak people find second.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>privacy</category>
    </item>
  </channel>
</rss>
