<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: John</title>
    <description>The latest articles on DEV Community by John (@john_182319291).</description>
    <link>https://dev.to/john_182319291</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4063265%2F912320f7-d3d8-4432-876b-a7e372adb01a.jpg</url>
      <title>DEV Community: John</title>
      <link>https://dev.to/john_182319291</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/john_182319291"/>
    <language>en</language>
    <item>
      <title>Self-Hosted WordPress vs WordPress.com Business vs WP Engine and Kinsta: The Real 3-Year Cost for a Two-Person Startup</title>
      <dc:creator>John</dc:creator>
      <pubDate>Thu, 17 Sep 2026 07:04:41 +0000</pubDate>
      <link>https://dev.to/john_182319291/self-hosted-wordpress-vs-wordpresscom-business-vs-wp-engine-and-kinsta-the-real-3-year-cost-for-a-5248</link>
      <guid>https://dev.to/john_182319291/self-hosted-wordpress-vs-wordpresscom-business-vs-wp-engine-and-kinsta-the-real-3-year-cost-for-a-5248</guid>
      <description>&lt;p&gt;For most two-person startups, self-hosted WordPress has the lowest cash cost over three years. It only has the lowest total cost if one founder can spend a small, steady amount of time each month on updates, backups and security, and does not bill that time at a consultant's rate. WordPress.com Business and managed hosts such as WP Engine or Kinsta cost more in fees but take most of that maintenance work away. The deciding figures are not the hosting fees. They are the renewal price of premium plugins, the cost of an incident, and how much each founder's hour is worth.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The technical founder with spare evenings&lt;/strong&gt; (for example, a developer building a SaaS landing page and blog): self-host on a small VPS or a personal cloud server, because the hosting bill stays small and the maintenance is familiar work.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The non-technical pair selling a service&lt;/strong&gt; (for example, two consultants who need a credible site and a contact form): use WordPress.com Business, because a fixed yearly fee costs less than learning server administration.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The WooCommerce shop with real revenue at stake&lt;/strong&gt; (for example, a two-person direct-to-consumer brand): use a managed host like Kinsta or WP Engine, because staging, backups and support cost less than one bad outage during a sale.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The content-heavy startup expecting traffic spikes&lt;/strong&gt; (for example, a media newsletter that gets shared widely): self-host behind a CDN or pick a managed plan with generous visit limits, because per-visit plan limits are where managed hosting costs rise.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The bootstrapped team planning to pivot&lt;/strong&gt; (for example, founders still testing product-market fit): self-host with a lean free plugin stack, because leaving is cheap and nothing is locked into a plan tier.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The team handling customer data under strict rules&lt;/strong&gt; (for example, a health or legal tech startup): self-host where you control the hosting location, because data location and access logs matter more than saving admin hours.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The central tradeoff is simple: self-hosting spends your time to save money, and managed WordPress spends money to save your time.&lt;/p&gt;




&lt;h2&gt;
  
  
  Table of contents
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;How much does WordPress really cost a two-person startup over 3 years?&lt;/li&gt;
&lt;li&gt;What goes into the total cost beyond the hosting bill&lt;/li&gt;
&lt;li&gt;What does self-hosting WordPress cost in year one, year two and year three?&lt;/li&gt;
&lt;li&gt;How WordPress.com Business pricing and plan limits play out over three years&lt;/li&gt;
&lt;li&gt;What do WP Engine and Kinsta charge once you hit visit and storage limits?&lt;/li&gt;
&lt;li&gt;Premium plugins and themes: the renewal costs that grow each year&lt;/li&gt;
&lt;li&gt;How many hours a month does self-hosted WordPress maintenance actually take?&lt;/li&gt;
&lt;li&gt;Backups, staging and disaster recovery across the three options&lt;/li&gt;
&lt;li&gt;Is self-hosted WordPress secure enough without a managed host's protection?&lt;/li&gt;
&lt;li&gt;What happens to the cost when traffic grows or WooCommerce is added?&lt;/li&gt;
&lt;li&gt;Where to run self-hosted WordPress: VPS, home server or personal cloud server&lt;/li&gt;
&lt;li&gt;Data sovereignty and hosting location for a small WordPress site&lt;/li&gt;
&lt;li&gt;How hard is it to switch between self-hosted and managed WordPress later?&lt;/li&gt;
&lt;li&gt;Which option fits your startup: recommendations by profile&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  How much does WordPress really cost a two-person startup over 3 years?
&lt;/h2&gt;

&lt;p&gt;WordPress itself costs nothing. The software is released under the GPLv2 licence, and you can download it from wordpress.org free of charge. The real cost is everything around it, and over 36 months those extra costs add up to far more than the first invoice suggests.&lt;/p&gt;

&lt;p&gt;A fair comparison needs three kinds of cost. &lt;strong&gt;Cash&lt;/strong&gt; is what goes on the card: hosting, domains, premium plugins and backup storage. &lt;strong&gt;Time&lt;/strong&gt; is the hours a founder spends on updates, fixing things and restores. &lt;strong&gt;Risk&lt;/strong&gt; is the expected cost of downtime, a hacked site or lost orders. Pricing pages only show the first kind. For a two-person team, the second and third often decide the answer.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Option&lt;/th&gt;
&lt;th&gt;What the bill covers&lt;/th&gt;
&lt;th&gt;What you still handle&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Self-hosted on a VPS&lt;/td&gt;
&lt;td&gt;Server, domain, any paid plugins, off-site backup storage&lt;/td&gt;
&lt;td&gt;Server updates, WordPress updates, security, backups, restores&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;WordPress.com Business&lt;/td&gt;
&lt;td&gt;Hosting, updates, backups, security for the platform&lt;/td&gt;
&lt;td&gt;Plugin choices, content, anything a plan limit blocks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;WP Engine&lt;/td&gt;
&lt;td&gt;Managed hosting, staging, backups, support&lt;/td&gt;
&lt;td&gt;Plugin licences, overage if you pass the visit limits&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kinsta&lt;/td&gt;
&lt;td&gt;Managed hosting, staging, backups, CDN, support&lt;/td&gt;
&lt;td&gt;Plugin licences, add-ons, moving to a higher tier as you grow&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The pattern is the same across all four. Self-hosting has the smallest fixed bill and the largest time cost. Managed plans reverse that. Over three years, renewals and growth decide which side comes out ahead. Plugin renewals and plan upgrades happen every year, while the time you spend learning to run a server mostly happens in year one.&lt;/p&gt;




&lt;h2&gt;
  
  
  What goes into the total cost beyond the hosting bill
&lt;/h2&gt;

&lt;p&gt;The hosting bill is the easy part to see. Most of the three-year total sits in smaller costs that only show up once the site is live. Each one is small on its own, but together they are what make a cheap-looking setup expensive.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Domain renewal&lt;/strong&gt;: you pay it every year whatever you choose, and some registrars charge far more to renew than they did the first year.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;TLS certificates&lt;/strong&gt;: Let's Encrypt certificates are free but expire after 90 days, so a self-hosted setup needs automatic renewal that someone checks is working.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Transactional email&lt;/strong&gt;: WordPress sends email through &lt;code&gt;wp_mail()&lt;/code&gt;, which uses PHP mail by default and often lands in spam, so most teams add an SMTP plugin and a sending service such as Amazon SES or Mailgun.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Off-site backups&lt;/strong&gt;: a backup kept on the same server does not protect you, so you need storage somewhere else, such as Backblaze B2 or an S3 bucket, plus a plugin like UpdraftPlus to send backups there.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CDN and DNS&lt;/strong&gt;: Cloudflare's free plan covers many small sites, but image-heavy pages or WooCommerce checkouts can push you towards paid features.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Uptime monitoring&lt;/strong&gt;: if nobody is watching, the first person to notice an outage is usually a customer.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Founder hours&lt;/strong&gt;: every update, failed restore and plugin conflict costs time, and time is what a two-person team has least of.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Where you run the site changes which of these you manage yourself. A self-managed VPS, a home server and a NAS all leave TLS, networking and backups to you. Yundera is a managed Personal Cloud Server, built on CasaOS, that runs self-hosted apps as Docker containers on a server dedicated to the user. On that kind of setup, public HTTPS access comes with the subdomain, but email, off-site backups and plugin upkeep are still your job.&lt;/p&gt;




&lt;h2&gt;
  
  
  What does self-hosting WordPress cost in year one, year two and year three?
&lt;/h2&gt;

&lt;p&gt;Self-hosting costs change shape over three years. Year one has most of the learning and setup work. Years two and three are cheaper in time, but plugin renewals and growing storage start to show up on the bill.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Cost line&lt;/th&gt;
&lt;th&gt;Year one&lt;/th&gt;
&lt;th&gt;Years two and three&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Server&lt;/td&gt;
&lt;td&gt;Small VPS, home server, NAS or a platform such as Yundera, often the smallest size that runs PHP and MySQL&lt;/td&gt;
&lt;td&gt;Same size unless traffic or WooCommerce needs more memory&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Setup time&lt;/td&gt;
&lt;td&gt;Heaviest: web server, database, TLS, SMTP, backups, caching&lt;/td&gt;
&lt;td&gt;Light: occasional rebuild or migration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Software upkeep&lt;/td&gt;
&lt;td&gt;WordPress applies minor updates automatically, but a person must test the major releases that come out a few times a year&lt;/td&gt;
&lt;td&gt;Same routine, plus PHP 8.x version upgrades that can break older plugins&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Operating system&lt;/td&gt;
&lt;td&gt;Choose an Ubuntu LTS release to get 5 years of standard support&lt;/td&gt;
&lt;td&gt;Plan one OS upgrade if you started late in a release's life&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Premium plugins&lt;/td&gt;
&lt;td&gt;First-year prices, often discounted&lt;/td&gt;
&lt;td&gt;Full renewal prices on every licence you kept&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Backup storage&lt;/td&gt;
&lt;td&gt;Small, a few snapshots&lt;/td&gt;
&lt;td&gt;Grows with each upload and every retained copy&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Incidents&lt;/td&gt;
&lt;td&gt;Most likely, while the setup is still new&lt;/td&gt;
&lt;td&gt;Less frequent, but each one takes longer if nobody has practised a restore&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The mistake teams make is judging the whole cost on year one alone. In year one, self-hosting looks expensive in hours and cheap in cash. By year three, you spend far fewer hours, but the cash cost has quietly risen through plugin renewals and storage. That third year is the right number to hold up against a managed plan's renewal price, not the introductory one.&lt;/p&gt;

&lt;p&gt;One routine keeps the time cost predictable. Set a fixed monthly maintenance window, put every update through a staging copy first, and run a timed test restore once a quarter.&lt;/p&gt;




&lt;h2&gt;
  
  
  How WordPress.com Business pricing and plan limits play out over three years
&lt;/h2&gt;

&lt;p&gt;WordPress.com Business is the only WordPress.com plan most startups need to think about, because it is the first plan where you can install your own plugins and themes. That access is what makes it a real alternative to self-hosting. The three-year cost depends less on the price shown on the page and more on how you pay and what the plan does not include.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Billing term&lt;/strong&gt;: monthly billing costs the most per month, while annual and multi-year terms cost less per month but require payment upfront. For a startup short on cash, paying for 36 months at once can be a hard call.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Introductory pricing&lt;/strong&gt;: the first-term price is often discounted, and renewals charge the standard rate. Build your three-year total on the renewal price, not the checkout price.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bundled domain&lt;/strong&gt;: an annual plan includes a domain registration for the first year only. From year two, you pay the domain renewal separately.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Developer access&lt;/strong&gt;: Business includes SFTP, SSH and database access, so a technical founder can still debug problems. You do not get root access to the server, so you cannot install system packages or tune PHP beyond what the platform lets you change.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Plugins you still pay for&lt;/strong&gt;: the plan pays for hosting, not for premium plugin licences, so your SEO, forms and membership renewals cost the same as they would on a VPS.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Store features&lt;/strong&gt;: a serious WooCommerce shop may push you onto a higher Commerce tier, which moves you into a different price band.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This plan works best as a fixed, predictable cost. It works worst when you need something the platform does not allow. That usually means leaving for Kinsta, WP Engine or a VPS, and paying for the migration on top.&lt;/p&gt;




&lt;h2&gt;
  
  
  What do WP Engine and Kinsta charge once you hit visit and storage limits?
&lt;/h2&gt;

&lt;p&gt;WP Engine and Kinsta both sell plans in tiers. Each tier sets a limit on sites, visits and storage. The entry price is only an accurate guide to your three-year cost if the site stays inside those limits. Once you go over, you either pay overage fees or move up a tier, and the next tier is usually a big jump rather than a small step.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Limit&lt;/th&gt;
&lt;th&gt;How it raises the bill&lt;/th&gt;
&lt;th&gt;What to watch&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Monthly visits&lt;/td&gt;
&lt;td&gt;Overage fees on each block of extra visits, or a forced move to a higher tier&lt;/td&gt;
&lt;td&gt;Both hosts count unique visitors over a 24-hour window, and bot traffic can count too&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Storage&lt;/td&gt;
&lt;td&gt;Paid add-on or higher tier once media and backups fill the quota&lt;/td&gt;
&lt;td&gt;WooCommerce product images and uncompressed uploads grow fastest&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Bandwidth or CDN usage&lt;/td&gt;
&lt;td&gt;Charges or tier changes on plans that measure data transfer&lt;/td&gt;
&lt;td&gt;Large downloads, video and podcast files served from the same site&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Number of sites&lt;/td&gt;
&lt;td&gt;A second site, such as a docs or app marketing site, can need a bigger plan&lt;/td&gt;
&lt;td&gt;Staging copies usually do not count, but separate production sites do&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Add-ons&lt;/td&gt;
&lt;td&gt;Extra fees for features like extra backups, security add-ons or additional PHP workers&lt;/td&gt;
&lt;td&gt;Features you enable once and then forget are billed every month&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For a two-person startup, the danger is a success spike, not steady growth. A launch post that goes viral or a mention in a newsletter can push one month over the visit limit. On a self-hosted server behind Cloudflare, the same spike usually costs nothing extra as long as caching holds.&lt;/p&gt;

&lt;p&gt;Before you commit to a year, look at the analytics from your current site. Estimate your visits in year three, including bots, and price the tier you will need then, not the tier you need today.&lt;/p&gt;




&lt;h2&gt;
  
  
  Premium plugins and themes: the renewal costs that grow each year
&lt;/h2&gt;

&lt;p&gt;Premium plugins are the part of the WordPress budget people forget to check. They cost the same whether you self-host, use WordPress.com Business or pay Kinsta, so the hosting choice does not change this line. What does change is how many plugins you keep adding over three years. Most are sold as yearly licences, and each one renews on its own date.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Licence expiry is not a shutdown&lt;/strong&gt;: plugins are GPL, so the code keeps running after the licence lapses, but updates and support stop. Running old code on a public site is a security debt that someone eventually has to pay.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Renewal versus first-year price&lt;/strong&gt;: many vendors discount the first year, so your year-two cost for Gravity Forms, WP Rocket or Yoast SEO Premium can be higher than what you paid at checkout.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Site-count tiers&lt;/strong&gt;: a single-site licence covers one production site. Adding a second site, such as a documentation site, can force every licence up to the multi-site tier at the same time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Page builders and themes&lt;/strong&gt;: Elementor Pro or a premium theme ties your layouts to that vendor, so dropping the licence later means rebuilding pages instead of simply cancelling.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;WooCommerce extensions&lt;/strong&gt;: subscriptions, bookings and payment gateway add-ons are often sold separately, and a shop can end up with more of these licences than every other plugin combined.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Free alternatives&lt;/strong&gt;: Rank Math or Yoast free, Contact Form 7 and a well-configured server cache cover many needs at no licence cost, as long as you accept less vendor support.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Run &lt;code&gt;wp plugin list&lt;/code&gt; once a quarter. For each paid plugin, write down its renewal date and what you would use instead. If nobody on the team can name a reason to keep a plugin, cancel it before it renews.&lt;/p&gt;




&lt;h2&gt;
  
  
  How many hours a month does self-hosted WordPress maintenance actually take?
&lt;/h2&gt;

&lt;p&gt;Nobody can give you an honest single number, because the hours depend on your setup more than on WordPress. Instead, count the tasks and time each one yourself over your first three months. After that, the monthly total usually settles into a predictable routine with occasional spikes.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Core and plugin updates&lt;/strong&gt;: running &lt;code&gt;wp core update&lt;/code&gt; and &lt;code&gt;wp plugin update --all&lt;/code&gt; takes minutes. Most of the time goes on testing checkout, forms and login on a staging copy before you push to production.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Server patching&lt;/strong&gt;: &lt;code&gt;apt upgrade&lt;/code&gt; and occasional reboots on a VPS, home server or NAS. The OS layer is less of your job on platforms that package WordPress as a Docker container, such as Yundera, although you still update the app itself.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Backup checks&lt;/strong&gt;: a backup job that reports success is not proof. Export a test with &lt;code&gt;wp db export&lt;/code&gt;, restore it somewhere else, and time the whole process.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Security review&lt;/strong&gt;: look at failed logins, unknown admin users and file changes in &lt;code&gt;wp-content&lt;/code&gt;. Each check is short, but it needs to happen on schedule.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Performance work&lt;/strong&gt;: clear caches, compress images and look for slow queries when pages get sluggish. This work comes in bursts rather than every month.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Incident response&lt;/strong&gt;: a plugin conflict, a full disk or an expired certificate. These are rare but take the longest, and they always happen at a bad time.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;To turn hours into money, multiply the monthly hours by the value of a founder's hour. Use the rate that founder could bill clients, or the value of product work they are not doing. Add a buffer for incidents. Compare that figure with the difference between your hosting bill and a managed plan. In a two-person team, one founder usually ends up owning this work, so price it at their rate.&lt;/p&gt;




&lt;h2&gt;
  
  
  Backups, staging and disaster recovery across the three options
&lt;/h2&gt;

&lt;p&gt;Backups only matter on the day you need to restore one, so compare the three options by what a restore actually involves. For WooCommerce, also ask how many orders you could lose between the last backup and the moment the site broke.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Recovery need&lt;/th&gt;
&lt;th&gt;Self-hosted WordPress&lt;/th&gt;
&lt;th&gt;WordPress.com Business, WP Engine, Kinsta&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Daily backups&lt;/td&gt;
&lt;td&gt;You set it up yourself with UpdraftPlus, a cron job or server snapshots&lt;/td&gt;
&lt;td&gt;Included and automatic, with retention set by the plan&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Off-site copy&lt;/td&gt;
&lt;td&gt;Your job: follow the 3-2-1 rule with at least one copy on separate storage&lt;/td&gt;
&lt;td&gt;Stored by the host, and you need an extra export if you want a copy outside the host&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Staging site&lt;/td&gt;
&lt;td&gt;Manual clone of files and database, plus search-and-replace for URLs with &lt;code&gt;wp search-replace&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;One-click staging, then push to live when ready&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Restore speed&lt;/td&gt;
&lt;td&gt;Depends on how often you practise it and how big &lt;code&gt;wp-content/uploads&lt;/code&gt; is&lt;/td&gt;
&lt;td&gt;A button in the dashboard, usually quick for small sites&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;WooCommerce orders&lt;/td&gt;
&lt;td&gt;A nightly backup can lose a full day of orders on restore&lt;/td&gt;
&lt;td&gt;Same risk unless the plan offers more frequent backups&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Full host failure&lt;/td&gt;
&lt;td&gt;You rebuild on another server from your off-site copy&lt;/td&gt;
&lt;td&gt;You wait for the provider, or restore elsewhere from your own export&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Managed hosts are better at the everyday cases: you break something with an update and roll back in minutes. Self-hosting can be better at the rare disaster, but only if you have been keeping an independent copy of the database and &lt;code&gt;wp-content&lt;/code&gt;. Many managed customers never make that copy, and they discover the gap when an account is suspended or a billing problem locks them out.&lt;/p&gt;

&lt;p&gt;Whichever option you choose, keep &lt;code&gt;wp-config.php&lt;/code&gt; settings, the latest database export and the uploads folder somewhere your hosting provider cannot reach. Test restoring from it twice a year.&lt;/p&gt;




&lt;h2&gt;
  
  
  Is self-hosted WordPress secure enough without a managed host's protection?
&lt;/h2&gt;

&lt;p&gt;Yes, if you treat security as a routine rather than a product. Most WordPress compromises come through outdated plugins, weak admin passwords and abandoned themes, not through WordPress core. A managed host reduces some of that risk, but it cannot protect a site that runs a vulnerable plugin you chose to install.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Plugin auto-updates&lt;/strong&gt;: since WordPress 5.5 you can switch on auto-updates for each plugin. Turn them on for low-risk plugins and keep manual, tested updates for WooCommerce and payment gateways.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Admin access&lt;/strong&gt;: use unique passwords and two-factor authentication through a plugin such as WP 2FA or Wordfence. Keep administrator accounts to the two founders only.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hardening in &lt;code&gt;wp-config.php&lt;/code&gt;&lt;/strong&gt;: set &lt;code&gt;DISALLOW_FILE_EDIT&lt;/code&gt; to &lt;code&gt;true&lt;/code&gt; so a stolen login cannot edit PHP files from the dashboard. Keep the database credentials out of any public repository.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Attack surface&lt;/strong&gt;: block &lt;code&gt;xmlrpc.php&lt;/code&gt; if nothing uses it, rate-limit &lt;code&gt;wp-login.php&lt;/code&gt; with fail2ban or your firewall, and delete unused themes and plugins instead of just deactivating them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Edge filtering&lt;/strong&gt;: Cloudflare in front of the origin hides its IP and absorbs a lot of automated traffic. Its free plan includes basic protections, and managed rulesets need a paid tier.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Detection&lt;/strong&gt;: file-integrity scans and alerts when a new admin account appears. Without them, a quiet compromise can sit unnoticed for months.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What managed hosts really add is people and speed. Kinsta and WP Engine run server-level firewalls, patch PHP for you, and some clean up malware if a site is infected. On a self-hosted server, cleanup is your job, and it is the most expensive hour in this whole cost model. Put a realistic incident allowance into your three-year budget rather than assuming it will not happen.&lt;/p&gt;




&lt;h2&gt;
  
  
  What happens to the cost when traffic grows or WooCommerce is added?
&lt;/h2&gt;

&lt;p&gt;A brochure site with a blog is mostly static pages. A cache can serve those pages without running PHP, so extra traffic costs very little. WooCommerce changes that, because carts, checkouts and account pages are different for every visitor and cannot be served from a shared page cache.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Uncacheable pages&lt;/strong&gt;: &lt;code&gt;/cart/&lt;/code&gt;, &lt;code&gt;/checkout/&lt;/code&gt; and &lt;code&gt;/my-account/&lt;/code&gt; run PHP and query the database for every visitor. On a VPS that means more memory. On Kinsta or WP Engine it can mean more PHP workers or a higher tier.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Object caching&lt;/strong&gt;: Redis cuts repeated database queries on dynamic pages. Self-hosted, it is one more service to run and keep an eye on. On managed plans, it is sometimes a paid add-on.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Order storage&lt;/strong&gt;: WooCommerce 8.2 made High-Performance Order Storage the default for new stores. Older stores that still keep orders in &lt;code&gt;wp_posts&lt;/code&gt; should migrate before order volume makes the database slow.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Scheduled tasks&lt;/strong&gt;: the default &lt;code&gt;wp-cron&lt;/code&gt; only runs when someone visits the site, which is unreliable for subscription renewals. Set &lt;code&gt;DISABLE_WP_CRON&lt;/code&gt; and call &lt;code&gt;wp cron event run --due-now&lt;/code&gt; from a real system cron instead.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Payment compliance&lt;/strong&gt;: hosted payment fields from Stripe or PayPal keep card data off your server and usually put you in the simplest PCI DSS questionnaire, SAQ A. Collecting card numbers yourself makes compliance far more work.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Extension licences&lt;/strong&gt;: shipping, tax and subscription plugins add renewal lines, often several at once.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The pattern over three years is clear. Traffic growth by itself barely moves a self-hosted budget, but it pushes managed plans up through visit limits. Adding WooCommerce raises the cost on both sides. For a store, the cost of downtime now includes lost sales, which makes paying for managed support easier to justify.&lt;/p&gt;




&lt;h2&gt;
  
  
  Where to run self-hosted WordPress: VPS, home server or personal cloud server
&lt;/h2&gt;

&lt;p&gt;Once you decide to self-host, the next choice is where the server lives. That choice decides how much of the stack you manage below WordPress itself. It is the second-largest time factor after your plugin list.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Option&lt;/th&gt;
&lt;th&gt;What you manage&lt;/th&gt;
&lt;th&gt;Main tradeoff&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;VPS (Hetzner, DigitalOcean, OVHcloud)&lt;/td&gt;
&lt;td&gt;OS, web server, PHP, database, TLS, firewall, backups&lt;/td&gt;
&lt;td&gt;Full control and a public IP, but every layer of the stack is your job&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Home server&lt;/td&gt;
&lt;td&gt;Hardware, OS, networking, power, plus everything a VPS needs&lt;/td&gt;
&lt;td&gt;Hardware you already own, but CGNAT, blocked ports 80 and 443, and slow home upload speeds can rule it out&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;NAS (Synology, QNAP)&lt;/td&gt;
&lt;td&gt;Container or package setup, router forwarding, DNS, TLS&lt;/td&gt;
&lt;td&gt;Convenient if you already own one, but a public store on the same box as company files mixes risks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Managed personal cloud server&lt;/td&gt;
&lt;td&gt;WordPress, plugins, content, off-site backups&lt;/td&gt;
&lt;td&gt;Dedicated server with one-click app installs and HTTPS handled, less control over the underlying system&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Shared hosting&lt;/td&gt;
&lt;td&gt;WordPress and plugins only&lt;/td&gt;
&lt;td&gt;Low effort, but noisy neighbours and limited PHP settings&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For the first four options, running WordPress in containers is the most portable approach. The official &lt;code&gt;wordpress&lt;/code&gt; Docker image plus a &lt;code&gt;mariadb&lt;/code&gt; container, with &lt;code&gt;wp-content&lt;/code&gt; and the database on named volumes, can be moved between a VPS, a NAS and a dedicated server without reinstalling anything.&lt;/p&gt;

&lt;p&gt;Pick based on who owns the pager. If one founder is comfortable with Linux, a VPS is the most flexible choice. If neither wants to manage networking and certificates, choose an option that handles those layers. Otherwise, the hours you save on hosting fees will go into troubleshooting DNS and TLS. A home server works for staging but is a risky place for a store that needs to stay online.&lt;/p&gt;




&lt;h2&gt;
  
  
  Data sovereignty and hosting location for a small WordPress site
&lt;/h2&gt;

&lt;p&gt;Even a small WordPress site handles personal data: contact form entries, comment IP addresses in &lt;code&gt;wp_comments&lt;/code&gt;, newsletter signups and, with WooCommerce, full names and addresses on every order. If you serve customers in the EU, the GDPR requires you to know where that data is stored and who can access it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Advantages of controlling the hosting location:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Known jurisdiction&lt;/strong&gt;: you choose the country the server runs in, instead of reading it from a provider's list of subprocessors.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Shorter processor chain&lt;/strong&gt;: fewer companies to cover with a data processing agreement under GDPR Article 28.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Direct access to logs&lt;/strong&gt;: web server and database logs stay under your control, and you can hand them over during an audit or a breach investigation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Exit without permission&lt;/strong&gt;: a full copy of the database and &lt;code&gt;wp-content&lt;/code&gt; is always yours, whatever happens to a billing account.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Checklist before you pick a host:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Region choice&lt;/strong&gt;: Kinsta and WP Engine let you choose a data centre region on their cloud providers. Check whether your WordPress.com plan gives you any say over location.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Third-party calls&lt;/strong&gt;: Gravatar, remotely loaded Google Fonts and embedded analytics send visitor data elsewhere. A Munich court fined a site owner in 2022 for loading Google Fonts remotely, so host fonts locally.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Email path&lt;/strong&gt;: order and form emails go through your SMTP provider, which becomes a processor too.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Backup location&lt;/strong&gt;: an off-site bucket in another country moves the data just as much as the server does.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Retention&lt;/strong&gt;: delete old form entries and inactive customer accounts on a schedule rather than keeping them forever.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Self-hosting does not make you compliant on its own. What it gives you is a shorter, clearer list of places your data goes, and a short list is easier to audit.&lt;/p&gt;




&lt;h2&gt;
  
  
  How hard is it to switch between self-hosted and managed WordPress later?
&lt;/h2&gt;

&lt;p&gt;It is easier than with most platforms, because WordPress is the same software everywhere. A site is a database plus a &lt;code&gt;wp-content&lt;/code&gt; folder, and a migration moves both. The difficulty comes from what each host adds on top of WordPress and from anything the site does while the move is in progress.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Migration tools&lt;/strong&gt;: Duplicator, All-in-One WP Migration and Migrate Guru package the site into one archive. WP Engine and Kinsta also offer their own migration tools or assisted moves, which lowers the cost of moving onto them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;URLs and serialized data&lt;/strong&gt;: if the domain or path changes, use &lt;code&gt;wp search-replace&lt;/code&gt; rather than raw SQL. It handles serialized PHP arrays that a plain find-and-replace would corrupt.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Host-specific code&lt;/strong&gt;: managed hosts add their own must-use plugins and caching layers in &lt;code&gt;wp-content/mu-plugins&lt;/code&gt;. Remove them after moving out, or they will fail silently or conflict with your own cache.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;DNS cutover&lt;/strong&gt;: lower the DNS record TTL to 300 seconds a day before the move, so visitors reach the new server within minutes instead of hours.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;WooCommerce orders&lt;/strong&gt;: orders placed on the old server after the export are lost. Schedule a short maintenance window or freeze checkout during the final sync.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Email and integrations&lt;/strong&gt;: SPF and DKIM records, payment webhooks and API keys stored in &lt;code&gt;wp-config.php&lt;/code&gt; all need checking on the new host.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For a two-person startup, this means the first hosting decision is not permanent. A site with a few plugins can move in an afternoon. A store with years of orders and many extensions needs a staged rehearsal. The practical lesson for your three-year budget: avoid host-only features you cannot rebuild elsewhere, and leaving stays cheap.&lt;/p&gt;




&lt;h2&gt;
  
  
  Which option fits your startup: recommendations by profile
&lt;/h2&gt;

&lt;p&gt;Use your team's skills, your revenue risk and your expected growth to choose. The hosting price alone should not decide it.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Profile&lt;/th&gt;
&lt;th&gt;Recommendation&lt;/th&gt;
&lt;th&gt;Main reason&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Developer founder, marketing site and blog&lt;/td&gt;
&lt;td&gt;Self-host on a VPS&lt;/td&gt;
&lt;td&gt;Low cash cost, and the upkeep uses skills the founder already has&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Two non-technical founders, service business&lt;/td&gt;
&lt;td&gt;WordPress.com Business&lt;/td&gt;
&lt;td&gt;Fixed yearly cost, no server work&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;WooCommerce store with daily orders&lt;/td&gt;
&lt;td&gt;Kinsta or WP Engine&lt;/td&gt;
&lt;td&gt;Staging, backups and support are worth more than one outage&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Content site expecting viral spikes&lt;/td&gt;
&lt;td&gt;Self-host behind Cloudflare&lt;/td&gt;
&lt;td&gt;Traffic spikes do not trigger visit overages&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;EU startup collecting customer data&lt;/td&gt;
&lt;td&gt;Self-host in a region you choose&lt;/td&gt;
&lt;td&gt;A shorter list of data processors, and logs you control&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pre-product-market-fit team likely to pivot&lt;/td&gt;
&lt;td&gt;Self-host with free plugins only&lt;/td&gt;
&lt;td&gt;No yearly commitments to cancel&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agency-style team running several client sites&lt;/td&gt;
&lt;td&gt;Managed multi-site plan&lt;/td&gt;
&lt;td&gt;One dashboard and support contract across every site&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Next steps:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If you self-host:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Deploy the &lt;code&gt;wordpress&lt;/code&gt; and &lt;code&gt;mariadb&lt;/code&gt; containers with named volumes.&lt;/li&gt;
&lt;li&gt;Set up off-site backups and test one restore before launch.&lt;/li&gt;
&lt;li&gt;Put a fixed monthly maintenance window in both founders' calendars.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If you choose WordPress.com Business:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Price the three-year total using renewal rates, not the introductory price.&lt;/li&gt;
&lt;li&gt;Check that every plugin you need is allowed on the plan.&lt;/li&gt;
&lt;li&gt;Schedule a regular full export stored outside WordPress.com.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If you choose Kinsta or WP Engine:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Estimate year-three visits, including bots, and pick that tier.&lt;/li&gt;
&lt;li&gt;List every add-on and its monthly price.&lt;/li&gt;
&lt;li&gt;Keep your own off-site copy of the database and &lt;code&gt;wp-content&lt;/code&gt;.&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>wordpress</category>
      <category>selfhosted</category>
      <category>startup</category>
      <category>hosting</category>
    </item>
    <item>
      <title>Netdata Alerts in Week One: Which to Keep, Tune or Silence, and How to Route Them to Slack, Discord and ntfy</title>
      <dc:creator>John</dc:creator>
      <pubDate>Tue, 15 Sep 2026 07:06:26 +0000</pubDate>
      <link>https://dev.to/john_182319291/netdata-alerts-in-week-one-which-to-keep-tune-or-silence-and-how-to-route-them-to-slack-discord-3000</link>
      <guid>https://dev.to/john_182319291/netdata-alerts-in-week-one-which-to-keep-tune-or-silence-and-how-to-route-them-to-slack-discord-3000</guid>
      <description>&lt;p&gt;Keep Netdata's stock alerts for disk space, memory pressure, OOM kills and failed systemd units exactly as they ship. In the first week, raise the thresholds or lengthen the delays on the CPU, network error and TCP reset alerts, and silence the ones for hardware or services you don't run. Send CRITICAL alerts to ntfy so they reach a phone, and send WARNING alerts to a Slack or Discord channel that someone reads once a day. That leaves a small team with a handful of real alerts a week, not dozens.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Solo on-call founder with two VPS nodes, e.g. a SaaS API plus its database box&lt;/strong&gt;: route CRITICAL to ntfy and WARNING to Discord, because one phone and one channel is all two people can realistically watch.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Team that already lives in Slack, e.g. a two-person agency using Slack with clients&lt;/strong&gt;: use a private Slack channel for WARNING and ntfy for CRITICAL, because paging through the same app you chat in gets muted within days.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Docker-heavy single host, e.g. twelve containers on one mini PC&lt;/strong&gt;: tune the per-container cgroup CPU and memory alerts first, because they produce most of the first-week noise on busy hosts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Parent-child streaming setup, e.g. one parent collecting from four children&lt;/strong&gt;: run health checks and notifications on the parent only, so one config decides what reaches you and each alert isn't sent several times.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Production database owner, e.g. a single PostgreSQL primary&lt;/strong&gt;: keep the connection, disk and replication alerts at stock or tighter, because those failures cost data, not just uptime.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Staging or hobby node, e.g. a preview environment rebuilt weekly&lt;/strong&gt;: set &lt;code&gt;to: silent&lt;/code&gt; for most alerts on it and keep the dashboard, because alerts on a box you routinely destroy teach you to ignore alerts everywhere.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The central tradeoff is simple: every alert you silence is an outage you might hear about from a customer, and every alert you keep but ignore trains you to miss the one that matters.&lt;/p&gt;




&lt;h2&gt;
  
  
  Table of contents
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Why do Netdata's stock alerts fire so often in the first week?&lt;/li&gt;
&lt;li&gt;How Netdata's health configuration is structured on disk&lt;/li&gt;
&lt;li&gt;Which stock alerts should you keep exactly as shipped?&lt;/li&gt;
&lt;li&gt;Which alerts are worth tuning rather than deleting?&lt;/li&gt;
&lt;li&gt;Which alerts can a small team silence outright?&lt;/li&gt;
&lt;li&gt;How do you change thresholds, delays and hysteresis without losing them on upgrade?&lt;/li&gt;
&lt;li&gt;Anomaly detection versus fixed thresholds for a two-person rotation&lt;/li&gt;
&lt;li&gt;How do you send Netdata alerts to Slack, Discord and ntfy?&lt;/li&gt;
&lt;li&gt;Splitting pages from notices with roles and severities&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Why do Netdata's stock alerts fire so often in the first week?
&lt;/h2&gt;

&lt;p&gt;Netdata ships with hundreds of alert definitions. They are written to be useful on any Linux machine, so they aren't tuned to yours. A new install turns on almost all of them at once, before you've seen what normal looks like on your hosts. Most first-week noise has a few predictable causes.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Auto-discovery attaches alerts to everything it finds&lt;/strong&gt;: every collector that detects a service, disk, interface or container brings its own alerts. You never explicitly asked for any of them, so it isn't obvious where a notification comes from.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Templates multiply across instances&lt;/strong&gt;: one definition that watches network interfaces runs separately against &lt;code&gt;eth0&lt;/code&gt;, &lt;code&gt;docker0&lt;/code&gt; and every &lt;code&gt;veth&lt;/code&gt; pair Docker creates. A host with twenty containers can raise the same warning many times over for traffic that is perfectly normal.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Thresholds assume a generic server&lt;/strong&gt;: the stock &lt;code&gt;10min_cpu_usage&lt;/code&gt; alert warns when average CPU over ten minutes crosses 85% and goes critical above 95%. A small VPS that runs a nightly build or backup crosses that line on schedule, every night.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Internet-facing noise looks like failure&lt;/strong&gt;: the TCP reset and dropped-packet alerts react to port scanners and bots hitting a public IP. On a box exposed to the internet, that is background activity, not a sign of trouble.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Short-lived containers churn state&lt;/strong&gt;: containers that start, stop and get rebuilt during deploys make their per-container alerts appear, change state and get removed. Each transition can produce a notification.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of this means the defaults are wrong. It means the first week is when you find out which alerts match how your hosts actually behave, and that is exactly what the next few sections sort through.&lt;/p&gt;




&lt;h2&gt;
  
  
  How Netdata's health configuration is structured on disk
&lt;/h2&gt;

&lt;p&gt;Before you change any alert, you need to know which files Netdata reads and which files a package upgrade overwrites. The layout differs slightly between install types. On a native package install, stock files live under &lt;code&gt;/usr/lib/netdata/conf.d/&lt;/code&gt; and your changes go under &lt;code&gt;/etc/netdata/&lt;/code&gt;. A static install from the kickstart script puts both trees under &lt;code&gt;/opt/netdata/&lt;/code&gt;. In the official Docker image, only &lt;code&gt;/etc/netdata&lt;/code&gt; is worth mounting as a volume.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Stock alert definitions&lt;/strong&gt;: &lt;code&gt;/usr/lib/netdata/conf.d/health.d/&lt;/code&gt; holds one &lt;code&gt;.conf&lt;/code&gt; file per area, for example &lt;code&gt;cpu.conf&lt;/code&gt;, &lt;code&gt;ram.conf&lt;/code&gt;, &lt;code&gt;disks.conf&lt;/code&gt; and &lt;code&gt;tcp_resets.conf&lt;/code&gt;. Upgrades replace these files, so any edit you make there disappears.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Your overrides&lt;/strong&gt;: &lt;code&gt;/etc/netdata/health.d/&lt;/code&gt; is where your versions go. A file here with the same name as a stock file replaces that stock file completely. Netdata does not merge the two line by line, so copy the whole file before you edit it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The edit-config helper&lt;/strong&gt;: run &lt;code&gt;sudo ./edit-config health.d/cpu.conf&lt;/code&gt; from inside &lt;code&gt;/etc/netdata&lt;/code&gt;. It copies the stock file into place and opens it, so you never start from a blank file by accident.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Alarm versus template&lt;/strong&gt;: each definition starts with either &lt;code&gt;alarm:&lt;/code&gt;, which is tied to one specific chart, or &lt;code&gt;template:&lt;/code&gt;, which applies to every chart of a context such as every disk or every container. Most stock noise comes from templates, because a single definition runs against every instance.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Notification settings&lt;/strong&gt;: &lt;code&gt;/etc/netdata/health_alarm_notify.conf&lt;/code&gt; holds webhooks, recipients and roles, apart from the alert logic.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The global switch&lt;/strong&gt;: the &lt;code&gt;[health]&lt;/code&gt; section of &lt;code&gt;netdata.conf&lt;/code&gt; turns the health engine on or off for the whole node.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;After editing, run &lt;code&gt;netdatacli reload-health&lt;/code&gt; to apply the change without restarting the agent.&lt;/p&gt;




&lt;h2&gt;
  
  
  Which stock alerts should you keep exactly as shipped?
&lt;/h2&gt;

&lt;p&gt;Keep an alert as shipped when the condition it catches ends in data loss, a crash, or a service that stays down until someone steps in. A small team can't afford to learn about these from a customer, and the stock thresholds already leave room before real damage. You can check the exact names on your nodes in the dashboard's Alerts tab or with &lt;code&gt;curl localhost:19999/api/v1/alarms?all&lt;/code&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;disk_space_usage&lt;/code&gt; and &lt;code&gt;disk_inode_usage&lt;/code&gt;&lt;/strong&gt;: a full filesystem breaks databases, log writers and Docker image pulls all at once. Running out of inodes produces the same "no space left" errors while &lt;code&gt;df -h&lt;/code&gt; still shows free space, which is exactly why a separate alert for it is worth having.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;oom_kill&lt;/code&gt;&lt;/strong&gt;: this fires when the kernel has already killed a process to free memory. It is not a prediction, so every notification means something on the host just died.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;ram_available&lt;/code&gt;&lt;/strong&gt;: this looks at memory the system can actually reclaim, not raw usage. That makes it a better early warning than &lt;code&gt;ram_in_use&lt;/code&gt; on hosts where the page cache fills RAM by design.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;systemd_service_unit_failed_state&lt;/code&gt;&lt;/strong&gt;: a unit that has crashed and hit its restart limit won't recover on its own. On a host without containers, this is often the only sign a background worker has stopped.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;docker_container_unhealthy&lt;/code&gt;&lt;/strong&gt;: this goes off only for containers you gave a &lt;code&gt;HEALTHCHECK&lt;/code&gt;, so it is quiet by default and precise when it does fire.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;postgres_total_connection_utilization&lt;/code&gt;&lt;/strong&gt;: running out of connections looks like an application outage even though the database itself is healthy.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If one of these turns out noisy, look at the host before you blame the threshold. On a disk that sits at 88% for a week, the fix is cleanup or more storage, not a quieter alert.&lt;/p&gt;




&lt;h2&gt;
  
  
  Which alerts are worth tuning rather than deleting?
&lt;/h2&gt;

&lt;p&gt;Tune an alert when the condition it tracks does matter, but the stock threshold or time window doesn't fit how your hosts behave. Delete an alert only when the condition never matters to you. The difference is simple: a CPU pegged at 100% for two hours is a real problem, while a CPU at 90% for twelve minutes during a nightly build is not. The fix is to change when the alert fires, not whether it exists.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Alert&lt;/th&gt;
&lt;th&gt;Why it is noisy at stock settings&lt;/th&gt;
&lt;th&gt;What to change&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;10min_cpu_usage&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Scheduled builds, backups and image pulls push the average above the line every day&lt;/td&gt;
&lt;td&gt;Widen the &lt;code&gt;lookup&lt;/code&gt; window to 30 minutes, or add &lt;code&gt;delay: up 15m&lt;/code&gt; so short spikes never notify&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;cgroup_10min_cpu_usage&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Fires for each container, so one busy deploy produces several warnings&lt;/td&gt;
&lt;td&gt;Keep CRITICAL, and match only the containers that serve users with a &lt;code&gt;chart labels&lt;/code&gt; filter&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;cgroup_ram_in_use&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Containers with tight memory limits sit near them by design, especially JVM and Node services&lt;/td&gt;
&lt;td&gt;Raise the WARNING level for those services, and keep &lt;code&gt;oom_kill&lt;/code&gt; as the hard signal&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;10min_disk_backlog&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Backup and &lt;code&gt;docker system prune&lt;/code&gt; runs queue I/O for minutes at a time&lt;/td&gt;
&lt;td&gt;Lengthen the up delay so it fires only when the backlog lasts longer than your longest backup&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;1m_ipv4_tcp_resets_sent&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Port scans against a public IP send resets all day&lt;/td&gt;
&lt;td&gt;Keep it as a low-priority notice, never a page, and raise the threshold once you know the baseline&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;inbound_packets_dropped_ratio&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Docker &lt;code&gt;veth&lt;/code&gt; interfaces drop packets during container churn&lt;/td&gt;
&lt;td&gt;Limit it to physical interfaces with a &lt;code&gt;chart labels&lt;/code&gt; match on the device name&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Before you change anything, write down the value each alert showed when it fired during the week. Set the new threshold from those notes, not from a guess.&lt;/p&gt;




&lt;h2&gt;
  
  
  Which alerts can a small team silence outright?
&lt;/h2&gt;

&lt;p&gt;Silence an alert when no value it could reach would make you change anything. There are two ways to do it, and the difference matters. Setting &lt;code&gt;to: silent&lt;/code&gt; in the definition keeps the alert running, so its state still shows on the dashboard, but nothing gets sent. Disabling it by name in the &lt;code&gt;[health]&lt;/code&gt; section of &lt;code&gt;netdata.conf&lt;/code&gt; stops it from being evaluated at all. For a small team, &lt;code&gt;to: silent&lt;/code&gt; is usually the better choice, because you can still look back at the history after an incident.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;1m_received_traffic_overflow&lt;/code&gt; and &lt;code&gt;1m_sent_traffic_overflow&lt;/code&gt;&lt;/strong&gt;: these compare traffic to the interface's link speed. Virtual NICs on cloud VPS plans often report a speed that has nothing to do with your real bandwidth cap, so the percentage is meaningless.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;load_average_1&lt;/code&gt;&lt;/strong&gt;: one-minute load on a small VPS jumps with every cron job. &lt;code&gt;10min_cpu_usage&lt;/code&gt;, once tuned, already covers sustained saturation, so this adds noise without adding a signal.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;system_clock_sync_state&lt;/code&gt;&lt;/strong&gt;: on some virtual machines and inside certain container setups, the host or hypervisor manages time. The agent can report "unsynchronised" indefinitely while the clock is actually correct.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;used_swap&lt;/code&gt;&lt;/strong&gt;: if you set up swap or zram on purpose so a 2 GB box can absorb spikes, using swap is the plan, not a failure. &lt;code&gt;oom_kill&lt;/code&gt; still catches the case where that plan runs out.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Web log redirect and bad-request ratios&lt;/strong&gt;: on a public site, crawlers and vulnerability scanners generate 3xx and 4xx responses constantly. Your application's error tracking is a better place to catch real client errors.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Every alert on disposable staging nodes&lt;/strong&gt;: match them with a host label and silence the lot, keeping only &lt;code&gt;disk_space_usage&lt;/code&gt;, so a full disk doesn't block the next deploy.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Keep a short comment above each change explaining why you silenced it. Six months from now, that line is the only record anyone will have.&lt;/p&gt;




&lt;h2&gt;
  
  
  How do you change thresholds, delays and hysteresis without losing them on upgrade?
&lt;/h2&gt;

&lt;p&gt;Copying a whole stock file into &lt;code&gt;/etc/netdata/health.d/&lt;/code&gt; works, but it freezes that file in time. When a later release fixes a stock definition in &lt;code&gt;cpu.conf&lt;/code&gt;, your copy keeps the old version forever. A cleaner pattern is to leave stock files untouched, write your own definitions under new names, and turn off only the stock alerts you replaced.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Put overrides in one file&lt;/strong&gt;: create something like &lt;code&gt;/etc/netdata/health.d/local-overrides.conf&lt;/code&gt; and give each definition a new name, for example &lt;code&gt;team_10min_cpu_usage&lt;/code&gt;. The name never matches a stock file, so upgrades can't shadow it or be shadowed by it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Disable only the originals&lt;/strong&gt;: in &lt;code&gt;netdata.conf&lt;/code&gt;, set &lt;code&gt;enabled alarms = !10min_cpu_usage !cgroup_ram_in_use *&lt;/code&gt; under &lt;code&gt;[health]&lt;/code&gt;. The trailing &lt;code&gt;*&lt;/code&gt; keeps every other stock alert active and upgradeable.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Widen the window with &lt;code&gt;lookup&lt;/code&gt;&lt;/strong&gt;: &lt;code&gt;lookup: average -30m unaligned of user,system&lt;/code&gt; averages user and system CPU over 30 minutes, so a ten-minute spike can't push it over the line.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Add hysteresis inside &lt;code&gt;warn&lt;/code&gt;&lt;/strong&gt;: &lt;code&gt;warn: $this &amp;gt; (($status &amp;gt;= $WARNING) ? (75) : (90))&lt;/code&gt; raises at 90% but clears only below 75%. That stops an alert flapping around a single value.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hold notifications with &lt;code&gt;delay&lt;/code&gt;&lt;/strong&gt;: &lt;code&gt;delay: up 10m down 15m multiplier 1.5 max 1h&lt;/code&gt; waits ten minutes before telling you, waits fifteen before sending the all-clear, and stretches both delays when an alert keeps flipping, up to one hour.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Version the directory&lt;/strong&gt;: keep &lt;code&gt;/etc/netdata&lt;/code&gt; in a git repository. After each upgrade, &lt;code&gt;diff&lt;/code&gt; the stock definitions you replaced against your own versions, so upstream fixes don't pass unnoticed.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Check a change by watching the renamed alert appear in &lt;code&gt;api/v1/alarms?all&lt;/code&gt; before you trust it.&lt;/p&gt;




&lt;h2&gt;
  
  
  Anomaly detection versus fixed thresholds for a two-person rotation
&lt;/h2&gt;

&lt;p&gt;Netdata trains small machine learning models for each metric, right on the agent, and marks every one-second sample as anomalous or normal. The result shows up as an anomaly rate. You can write an alert on it with a &lt;code&gt;lookup&lt;/code&gt; that uses the &lt;code&gt;anomaly-bit&lt;/code&gt; option, for example the average anomaly rate across a chart over the last 10 minutes. The question is whether that signal should ever wake someone up.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;Fixed thresholds&lt;/th&gt;
&lt;th&gt;Anomaly rate alerts&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;What triggers it&lt;/td&gt;
&lt;td&gt;A value crosses a number you picked, such as disk above 90%&lt;/td&gt;
&lt;td&gt;A metric behaves unlike its own recent history&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;First-week behaviour&lt;/td&gt;
&lt;td&gt;Noisy until tuned, but the noise is predictable&lt;/td&gt;
&lt;td&gt;Unreliable while models are still training on only a few days of data&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Catches well&lt;/td&gt;
&lt;td&gt;Hard limits: full disks, exhausted connections, OOM kills&lt;/td&gt;
&lt;td&gt;Unusual patterns with no obvious limit, such as a quiet API suddenly doing steady writes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Blind spot&lt;/td&gt;
&lt;td&gt;Anything you didn't think to set a threshold for&lt;/td&gt;
&lt;td&gt;Slow drift, because a disk filling 1% a day just becomes the new normal&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Readable at 3 a.m.&lt;/td&gt;
&lt;td&gt;Yes: "disk at 94%" tells you what to do&lt;/td&gt;
&lt;td&gt;Rarely: "anomaly rate 40% on 12 charts" tells you to go and look&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Resource cost&lt;/td&gt;
&lt;td&gt;Negligible&lt;/td&gt;
&lt;td&gt;Extra CPU on every node, which is why many parent-child setups turn off &lt;code&gt;[ml]&lt;/code&gt; on children and let the parent train&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For two people sharing on-call, page only on fixed thresholds. Send anomaly rate alerts, if you enable them, to the daily notice channel. Use them as a prompt to open the dashboard, not as proof something is broken. If an anomaly alert keeps pointing at a real problem, turn that into a fixed-threshold alert you can act on.&lt;/p&gt;




&lt;h2&gt;
  
  
  How do you send Netdata alerts to Slack, Discord and ntfy?
&lt;/h2&gt;

&lt;p&gt;The agent sends notifications through &lt;code&gt;alarm-notify.sh&lt;/code&gt;, and every setting it reads lives in &lt;code&gt;health_alarm_notify.conf&lt;/code&gt;. Open it with &lt;code&gt;sudo ./edit-config health_alarm_notify.conf&lt;/code&gt; from &lt;code&gt;/etc/netdata&lt;/code&gt;. It is a long file, but each service needs only a switch, a destination and a default recipient.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Slack&lt;/strong&gt;: create an incoming webhook for your workspace, then set &lt;code&gt;SEND_SLACK="YES"&lt;/code&gt;, paste the URL into &lt;code&gt;SLACK_WEBHOOK_URL&lt;/code&gt;, and set &lt;code&gt;DEFAULT_RECIPIENT_SLACK&lt;/code&gt; to a channel such as &lt;code&gt;#ops-alerts&lt;/code&gt;. Each Slack webhook is tied to a single channel, so plan on one webhook for each channel you want to post to.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Discord&lt;/strong&gt;: in the channel settings, create a webhook under Integrations. Set &lt;code&gt;SEND_DISCORD="YES"&lt;/code&gt;, put the URL in &lt;code&gt;DISCORD_WEBHOOK_URL&lt;/code&gt;, and set &lt;code&gt;DEFAULT_RECIPIENT_DISCORD&lt;/code&gt; to the channel name. Discord is often the free option for a small team that doesn't pay for Slack.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ntfy&lt;/strong&gt;: set &lt;code&gt;SEND_NTFY="YES"&lt;/code&gt; and put the full topic URL in &lt;code&gt;DEFAULT_RECIPIENT_NTFY&lt;/code&gt;, for example &lt;code&gt;https://ntfy.sh/&lt;/code&gt; followed by your topic. On the public ntfy.sh server, anyone who knows a topic name can subscribe to it, so choose a long random name or run your own ntfy server and set &lt;code&gt;NTFY_ACCESS_TOKEN&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Email as a fallback&lt;/strong&gt;: &lt;code&gt;SEND_EMAIL&lt;/code&gt; depends on a working &lt;code&gt;sendmail&lt;/code&gt; on the host. Most small VPS setups don't have one configured, so turn it off rather than letting messages fail without anyone noticing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Test before trusting&lt;/strong&gt;: switch to the &lt;code&gt;netdata&lt;/code&gt; user with &lt;code&gt;sudo su -s /bin/bash netdata&lt;/code&gt;, then run &lt;code&gt;/usr/libexec/netdata/plugins.d/alarm-notify.sh test&lt;/code&gt;. It sends a WARNING, a CRITICAL and a CLEAR to every destination you enabled.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If the test works from a terminal but real alerts never arrive, check that the &lt;code&gt;netdata&lt;/code&gt; user can reach the internet through your firewall.&lt;/p&gt;




&lt;h2&gt;
  
  
  Splitting pages from notices with roles and severities
&lt;/h2&gt;

&lt;p&gt;Every alert definition has a &lt;code&gt;to:&lt;/code&gt; line naming a role, and most stock alerts use &lt;code&gt;to: sysadmin&lt;/code&gt;. In &lt;code&gt;health_alarm_notify.conf&lt;/code&gt;, each role maps to recipients for each service. Put those two ideas together and you get a two-tier setup: a short list of alerts that can wake a person up, and everything else collected for a daily look.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Role mappings&lt;/strong&gt;: &lt;code&gt;role_recipients_slack[sysadmin]="#ops-alerts"&lt;/code&gt; sends every sysadmin alert to that channel. Roles you don't map fall back to the &lt;code&gt;DEFAULT_RECIPIENT_&lt;/code&gt; value for each service, so a role you forgot to map still reaches someone.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The &lt;code&gt;|critical&lt;/code&gt; modifier&lt;/strong&gt;: add it to a recipient, as in &lt;code&gt;role_recipients_ntfy[sysadmin]="https://ntfy.sh/&amp;lt;topic&amp;gt;|critical"&lt;/code&gt;, and that destination receives only CRITICAL alerts and the CLEAR that follows. WARNING alerts still reach Slack or Discord, but they never buzz a phone.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A dedicated pager role&lt;/strong&gt;: in your override file, set &lt;code&gt;to: pager&lt;/code&gt; on the handful of alerts that justify waking someone, such as &lt;code&gt;disk_space_usage&lt;/code&gt;, &lt;code&gt;oom_kill&lt;/code&gt; and database connection exhaustion. Map &lt;code&gt;pager&lt;/code&gt; to ntfy only, and a threshold edit elsewhere can never quietly add to the page list.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep a written record in chat&lt;/strong&gt;: map &lt;code&gt;pager&lt;/code&gt; to Slack or Discord as well as ntfy. The phone push gets attention, and the channel keeps a timestamped history both of you can scroll back through the next morning.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Per-person routing&lt;/strong&gt;: with two people, give each their own ntfy topic and alternate which one sits in &lt;code&gt;role_recipients_ntfy[pager]&lt;/code&gt; week by week. That is a simple on-call rotation without a paging SaaS.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Database and web roles&lt;/strong&gt;: stock database alerts use roles like &lt;code&gt;dba&lt;/code&gt;. Map them explicitly, or they fall through to your defaults and page you with the wrong priority.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Run &lt;code&gt;alarm-notify.sh test "pager"&lt;/code&gt; after every mapping change.&lt;/p&gt;

</description>
      <category>netdata</category>
      <category>alerting</category>
      <category>sre</category>
      <category>devops</category>
    </item>
    <item>
      <title>rclone mount vs rclone bisync after Dropbox: the real latency, RAM and conflict numbers</title>
      <dc:creator>John</dc:creator>
      <pubDate>Sat, 12 Sep 2026 07:06:51 +0000</pubDate>
      <link>https://dev.to/john_182319291/rclone-mount-vs-rclone-bisync-after-dropbox-the-real-latency-ram-and-conflict-numbers-1jlk</link>
      <guid>https://dev.to/john_182319291/rclone-mount-vs-rclone-bisync-after-dropbox-the-real-latency-ram-and-conflict-numbers-1jlk</guid>
      <description>&lt;p&gt;Neither one alone reproduces the Dropbox desktop client, and the honest answer is that you run both. An &lt;code&gt;rclone mount&lt;/code&gt; with &lt;code&gt;--vfs-cache-mode full&lt;/code&gt; gives you a folder that looks complete in your file manager, but every first read of an uncached file costs a network round trip, and nothing is available offline until you have opened it. &lt;code&gt;rclone bisync&lt;/code&gt; gives you genuine local copies with local disk latency and true offline access, at the cost of a scheduled run, a mandatory first &lt;code&gt;--resync&lt;/code&gt;, and conflict files you have to resolve yourself. Split your data: bisync the working set you edit daily, mount the archive you only read.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TL;DR by reader profile&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Solo dev with one laptop and a home server, 40 GB of code and notes:&lt;/strong&gt; run &lt;code&gt;bisync&lt;/code&gt; on a 5 minute timer for the working set, because local disk latency and offline access matter more than seeing every remote change instantly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Photographer or media hoarder with 2 TB on object storage:&lt;/strong&gt; run a single &lt;code&gt;mount&lt;/code&gt; with &lt;code&gt;--vfs-cache-mode full&lt;/code&gt; and a capped cache, because copying 2 TB to every device defeats the point of moving off the hosted plan.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Developer who builds inside the synced folder:&lt;/strong&gt; never point a compiler, a &lt;code&gt;node_modules&lt;/code&gt; tree or a Git repository at a mount, because per file stat and open calls turn a 20 second build into minutes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Someone with two laptops and a phone editing the same files:&lt;/strong&gt; bisync each laptop against the remote and expose the server folder over &lt;code&gt;rclone serve webdav&lt;/code&gt;, because bisync is two way between exactly two ends, not a mesh.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Anyone on a metered or flaky connection:&lt;/strong&gt; choose bisync with &lt;code&gt;--check-access&lt;/code&gt; and a low &lt;code&gt;--transfers&lt;/code&gt; value, because a mount reacts to a dropped link by returning I/O errors mid save.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Team of one who wants the sharing links back:&lt;/strong&gt; accept that neither tool replaces them, and plan a second front end for public links rather than stretching rclone into the job.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The central tradeoff is simple: a mount buys you unlimited apparent capacity and pays for it in latency and fragility, while bisync buys you local speed and offline safety and pays for it in disk space, scheduling and conflict resolution.&lt;/p&gt;




&lt;h2&gt;
  
  
  Table of contents
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;What does an rclone mount actually replace when the Dropbox client is gone?&lt;/li&gt;
&lt;li&gt;How does rclone bisync differ from a mount in day to day use?&lt;/li&gt;
&lt;li&gt;Which VFS cache mode should you run: off, minimal, writes or full?&lt;/li&gt;
&lt;li&gt;How much RAM and disk does an rclone mount really consume?&lt;/li&gt;
&lt;li&gt;What latency do open, stat and write calls show against a local disk baseline?&lt;/li&gt;
&lt;li&gt;How fast do remote changes appear, and what do dir-cache-time and poll-interval control?&lt;/li&gt;
&lt;li&gt;How do you size the VFS cache without filling the disk?&lt;/li&gt;
&lt;li&gt;Conflicts, renames and deletes: what bisync does and what the first resync costs&lt;/li&gt;
&lt;li&gt;What breaks when the network drops or the provider rate limits you?&lt;/li&gt;
&lt;li&gt;Running the mount under systemd and recovering it after a reboot&lt;/li&gt;
&lt;li&gt;Where should the mount run: home server, NAS, VPS or managed Personal Cloud Server?&lt;/li&gt;
&lt;li&gt;How do you reach the same folder from a phone or a second laptop?&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  What does an rclone mount actually replace when the Dropbox client is gone?
&lt;/h2&gt;

&lt;p&gt;A mount replaces the appearance of the Dropbox folder, not its behaviour. &lt;code&gt;rclone mount remote:files ~/Cloud --vfs-cache-mode full&lt;/code&gt; presents a FUSE filesystem where &lt;code&gt;ls&lt;/code&gt; returns the full listing within milliseconds once the directory cache is warm, so your file manager, your editor's open dialog and &lt;code&gt;find&lt;/code&gt; all behave normally. What changes is what happens underneath each file you touch.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Directory listings become remote metadata calls:&lt;/strong&gt; the first &lt;code&gt;ls&lt;/code&gt; on a cold directory hits the provider API, and rclone holds that listing for the &lt;code&gt;--dir-cache-time&lt;/code&gt; window, which defaults to 5 minutes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;File contents arrive on demand:&lt;/strong&gt; opening a 400 MB video streams it over the network rather than reading it from disk, and with &lt;code&gt;--vfs-cache-mode full&lt;/code&gt; the bytes land in the cache directory only after that first read completes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Writes go through a local staging file:&lt;/strong&gt; in &lt;code&gt;full&lt;/code&gt; or &lt;code&gt;writes&lt;/code&gt; mode your application writes to the cache, the file closes, then rclone uploads it, so a save that returns instantly can still be in flight seconds later.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Offline access disappears by default:&lt;/strong&gt; the Dropbox client kept every non selective file on disk, while a mount shows you names you cannot open when the link is down.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Selective sync has no equivalent:&lt;/strong&gt; you approximate it with &lt;code&gt;--exclude&lt;/code&gt; filters or a second mount, not with a checkbox.&lt;/p&gt;

&lt;p&gt;The practical consequence is that a mount is excellent for an archive you read occasionally and poor for a directory you compile in. Treat it as a network drive with a good cache, which is what it is.&lt;/p&gt;




&lt;h2&gt;
  
  
  How does rclone bisync differ from a mount in day to day use?
&lt;/h2&gt;

&lt;p&gt;Bisync is a scheduled job, not a filesystem. You run &lt;code&gt;rclone bisync ~/Cloud remote:files&lt;/code&gt; from cron or a systemd timer, it compares both sides against a stored listing pair in &lt;code&gt;~/.cache/rclone/bisync/&lt;/code&gt;, and it copies the differences in both directions. Between runs, your local directory is an ordinary folder on an ordinary disk.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Behaviour&lt;/th&gt;
&lt;th&gt;rclone mount&lt;/th&gt;
&lt;th&gt;rclone bisync&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Where files live&lt;/td&gt;
&lt;td&gt;Cache directory, populated on first read&lt;/td&gt;
&lt;td&gt;Full local copy of everything in scope&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Offline access&lt;/td&gt;
&lt;td&gt;Only what is already cached&lt;/td&gt;
&lt;td&gt;Every file, always readable and writable&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Change propagation&lt;/td&gt;
&lt;td&gt;Within the &lt;code&gt;--dir-cache-time&lt;/code&gt; window, 5 minutes by default&lt;/td&gt;
&lt;td&gt;Only when the next scheduled run fires&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Disk required&lt;/td&gt;
&lt;td&gt;Whatever you cap the cache at&lt;/td&gt;
&lt;td&gt;At least the size of the synced tree&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Failure mode&lt;/td&gt;
&lt;td&gt;I/O errors reaching the application mid operation&lt;/td&gt;
&lt;td&gt;A run exits non zero and leaves both sides untouched&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Conflicting edits&lt;/td&gt;
&lt;td&gt;Last writer wins silently&lt;/td&gt;
&lt;td&gt;Conflict files written to both sides&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The scheduling interval is the knob that defines the experience. A 5 minute timer means a file you save on the server appears on the laptop within 5 minutes and a deletion propagates on the same clock. A 60 minute timer halves your API calls and doubles the window in which the two sides can diverge.&lt;/p&gt;

&lt;p&gt;The other practical difference is scope. A mount covers a whole remote cheaply because it stores almost nothing. Bisync covers only what you are prepared to store twice, which is why the working set versus archive split is the decision that shapes everything else.&lt;/p&gt;




&lt;h2&gt;
  
  
  Which VFS cache mode should you run: off, minimal, writes or full?
&lt;/h2&gt;

&lt;p&gt;There are four modes and only two of them are realistic for a folder you actually work in.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;--vfs-cache-mode off&lt;/code&gt;:&lt;/strong&gt; files are read and written straight through, and any application that opens a file for both reading and writing fails, which rules out most editors, SQLite databases and Office formats.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;--vfs-cache-mode minimal&lt;/code&gt;:&lt;/strong&gt; files opened read write are staged on disk, everything else streams, so &lt;code&gt;vim&lt;/code&gt; saves work but a seek backwards in a large file still costs a fresh range request.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;--vfs-cache-mode writes&lt;/code&gt;:&lt;/strong&gt; all writes go to the cache first and uploads happen on close, while reads stream from the remote every time, which suits a machine with a small disk and an append heavy workload such as a log or backup target.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;--vfs-cache-mode full&lt;/code&gt;:&lt;/strong&gt; reads and writes both use the cache, sparse chunks are kept, and a second open of the same file costs local disk latency instead of a round trip, which is the only mode that feels remotely like the old desktop client.&lt;/p&gt;

&lt;p&gt;Run &lt;code&gt;full&lt;/code&gt; unless the disk under the mount cannot spare the space. That disk is the real constraint, so decide where the mount lives before you tune the flags: a home server with a spare SSD, a NAS, a self managed VPS with a small volume, or a managed host. Yundera is a managed Personal Cloud Server, built on CasaOS, that runs self-hosted apps as Docker containers on a server dedicated to the user. Whichever you pick, budget cache space on the same box that runs &lt;code&gt;rclone mount&lt;/code&gt;, because the cache is local by definition and cannot be pushed onto the remote.&lt;/p&gt;




&lt;h2&gt;
  
  
  How much RAM and disk does an rclone mount really consume?
&lt;/h2&gt;

&lt;p&gt;Memory use is not a single number, it is a formula with three terms you control.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A base process footprint:&lt;/strong&gt; a single mount process with nothing open is small enough that it is not the term you should worry about, and you can confirm yours with &lt;code&gt;ps -o rss= -C rclone&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Read buffers, one per open file:&lt;/strong&gt; &lt;code&gt;--buffer-size&lt;/code&gt; defaults to 16 MiB and is allocated per file handle in use, so ten files open simultaneously is on the order of 160 MiB of buffers before anything else. Set &lt;code&gt;--buffer-size 0&lt;/code&gt; on a small box and accept more round trips.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cached directory metadata:&lt;/strong&gt; rclone keeps listings in memory for the &lt;code&gt;--dir-cache-time&lt;/code&gt; window, and the cost scales with the number of objects you have walked, not their size, so a &lt;code&gt;find&lt;/code&gt; across 500,000 files is what turns a quiet mount into a hungry one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Upload concurrency:&lt;/strong&gt; &lt;code&gt;--transfers&lt;/code&gt; defaults to 4, and each in flight upload of a large file holds chunk buffers whose size depends on the backend, which is why a bulk copy into the mount is the peak you should size for.&lt;/p&gt;

&lt;p&gt;Disk is simpler and more dangerous. With &lt;code&gt;--vfs-cache-mode full&lt;/code&gt; the cache directory, by default under &lt;code&gt;~/.cache/rclone/vfs/&lt;/code&gt;, grows to hold every file you have read or written until an eviction rule removes it. Read one 40 GB video and you have written 40 GB locally. Two limits matter: &lt;code&gt;--vfs-cache-max-size&lt;/code&gt; caps total bytes, and &lt;code&gt;--vfs-cache-max-age&lt;/code&gt; defaults to 1 hour and expires idle entries. Leave both unset and the cache is bounded only by the filesystem it sits on.&lt;/p&gt;




&lt;h2&gt;
  
  
  What latency do open, stat and write calls show against a local disk baseline?
&lt;/h2&gt;

&lt;p&gt;Measure your own numbers rather than trusting anyone's table, because the dominant term is the round trip time between your server and the provider region. Get the baseline first with &lt;code&gt;ping&lt;/code&gt; to the endpoint host, then run &lt;code&gt;strace -c -f ls -l&lt;/code&gt; in a cold directory on the mount and the same command in a bisync folder. The difference you see is the entire argument of this article.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cold &lt;code&gt;stat&lt;/code&gt; on a mount:&lt;/strong&gt; costs at least one round trip if the parent directory is not cached, which is why a single &lt;code&gt;ls -l&lt;/code&gt; on 2,000 files can take seconds while the same command on local disk returns in milliseconds.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Warm &lt;code&gt;stat&lt;/code&gt; on a mount:&lt;/strong&gt; served from kernel attribute cache and rclone memory, where &lt;code&gt;--attr-timeout&lt;/code&gt; defaults to 1s and controls how long the kernel trusts what it was told.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;First &lt;code&gt;open&lt;/code&gt; and read:&lt;/strong&gt; one round trip plus transfer time for the requested range, so latency scales with file size on a streaming read and with RTT on a small file.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Second &lt;code&gt;open&lt;/code&gt; in &lt;code&gt;full&lt;/code&gt; mode:&lt;/strong&gt; local disk latency, because the bytes are already in the cache directory.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;write&lt;/code&gt; followed by &lt;code&gt;close&lt;/code&gt;:&lt;/strong&gt; the write returns at local speed, then &lt;code&gt;close&lt;/code&gt; blocks until the upload starts, so a save of a 2 MB file feels instant and a save of a 2 GB file does not.&lt;/p&gt;

&lt;p&gt;Small files are where mounts lose decisively. A tree of 10,000 source files under 4 KB each turns into 10,000 metadata operations, and no cache mode fixes the first pass. A bisync folder pays that cost once per run, in the background, with &lt;code&gt;--checkers&lt;/code&gt; running them in parallel.&lt;/p&gt;




&lt;h2&gt;
  
  
  How fast do remote changes appear, and what do dir-cache-time and poll-interval control?
&lt;/h2&gt;

&lt;p&gt;Two independent timers decide when a file written elsewhere shows up in your mount, and confusing them is the most common reason people think the mount is broken.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;--dir-cache-time&lt;/code&gt;, default 5m0s:&lt;/strong&gt; the maximum age of a cached directory listing. When it expires, the next access re-lists that directory from the remote. Raise it to 1000h and you cut metadata calls to almost nothing, at the price of never noticing outside changes on your own.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;--poll-interval&lt;/code&gt;, default 1m0s:&lt;/strong&gt; how often rclone asks the backend for a change feed and invalidates only the affected directories. This is the mechanism that makes a long &lt;code&gt;--dir-cache-time&lt;/code&gt; safe.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Backend support is the catch:&lt;/strong&gt; polling depends on the backend implementing change notification. Google Drive, Dropbox and OneDrive do. Plain S3, Backblaze B2 and SFTP do not, and on those the only thing that refreshes a listing is &lt;code&gt;--dir-cache-time&lt;/code&gt; expiry.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Manual invalidation as the escape hatch:&lt;/strong&gt; start the mount with &lt;code&gt;--rc&lt;/code&gt; and run &lt;code&gt;rclone rc vfs/refresh recursive=true&lt;/code&gt; or &lt;code&gt;rclone rc vfs/forget dir=path/to/folder&lt;/code&gt; to force a re-list immediately.&lt;/p&gt;

&lt;p&gt;The practical recipe for a polling capable backend is &lt;code&gt;--dir-cache-time 1000h --poll-interval 15s&lt;/code&gt;, which gives you near instant propagation and almost no idle API traffic. On a non polling backend such as B2, keep &lt;code&gt;--dir-cache-time&lt;/code&gt; between 1m and 5m if you need freshness, and accept the listing calls, or leave it long and drive refreshes from whatever process writes the files.&lt;/p&gt;

&lt;p&gt;None of this applies to bisync, where propagation is exactly your timer interval and nothing else.&lt;/p&gt;




&lt;h2&gt;
  
  
  How do you size the VFS cache without filling the disk?
&lt;/h2&gt;

&lt;p&gt;Start from the largest single file you will ever write, not from your total data. The cache must hold a complete copy of any file being written, so &lt;code&gt;--vfs-cache-max-size&lt;/code&gt; is a target for eviction, not a hard ceiling: an in use or not yet uploaded file is never evicted, and the directory can exceed the cap while an upload is in flight.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Workload&lt;/th&gt;
&lt;th&gt;Settings to start with&lt;/th&gt;
&lt;th&gt;What you give up&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Archive you mostly read, 500 GB SSD&lt;/td&gt;
&lt;td&gt;&lt;code&gt;--vfs-cache-max-size 100G --vfs-cache-max-age 720h&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;100 GB of disk permanently committed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Small VPS with a 40 GB volume&lt;/td&gt;
&lt;td&gt;&lt;code&gt;--vfs-cache-max-size 8G --vfs-cache-max-age 24h&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Repeat reads of large files go back to the network&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Write heavy target, backups or logs&lt;/td&gt;
&lt;td&gt;&lt;code&gt;--vfs-cache-mode writes --vfs-cache-max-age 1h&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Every read is a fresh download&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Laptop with limited space&lt;/td&gt;
&lt;td&gt;&lt;code&gt;--vfs-cache-max-size 4G --vfs-cache-poll-interval 30s&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;More frequent eviction churn on the disk&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Large file editing, video or disk images&lt;/td&gt;
&lt;td&gt;Cache at least 2x the biggest file&lt;/td&gt;
&lt;td&gt;The cap is advisory during writes&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Eviction is least recently used and runs on &lt;code&gt;--vfs-cache-poll-interval&lt;/code&gt;, which defaults to 1m0s. &lt;code&gt;--vfs-cache-max-age&lt;/code&gt; defaults to 1h0m0s, so out of the box a file you read this morning is gone by lunchtime. Put the cache somewhere you control with &lt;code&gt;--cache-dir /var/cache/rclone&lt;/code&gt; rather than leaving it under &lt;code&gt;~/.cache&lt;/code&gt;, and alert on that path at 85 percent full. Also watch &lt;code&gt;--vfs-write-back&lt;/code&gt;, default 5s, which is how long a closed file waits before upload starts.&lt;/p&gt;




&lt;h2&gt;
  
  
  Conflicts, renames and deletes: what bisync does and what the first resync costs
&lt;/h2&gt;

&lt;p&gt;The first run is not optional and it is not symmetric. &lt;code&gt;rclone bisync ~/Cloud remote:files --resync&lt;/code&gt; establishes the baseline listings that every later run compares against, and by default path1 wins for files that differ on both sides. Run it on a copy first, or with &lt;code&gt;--dry-run&lt;/code&gt;, because a careless &lt;code&gt;--resync&lt;/code&gt; is the one command in this article that can overwrite good data.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Conflicts are renamed, not merged:&lt;/strong&gt; when both sides changed since the last run, the default &lt;code&gt;--conflict-resolve none&lt;/code&gt; keeps both and appends &lt;code&gt;..path1&lt;/code&gt; and &lt;code&gt;..path2&lt;/code&gt; to the filenames, leaving you to pick. Set &lt;code&gt;--conflict-resolve newer&lt;/code&gt; if you would rather it decide, and &lt;code&gt;--conflict-loser delete&lt;/code&gt; if you do not want the loser kept.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Deletes propagate, with a brake:&lt;/strong&gt; &lt;code&gt;--max-delete&lt;/code&gt; defaults to 50, meaning a run aborts if more than 50 percent of files on either side would be deleted, which is what saves you when a drive fails to mount.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Renames cost a full transfer:&lt;/strong&gt; a renamed directory is seen as deletions plus new files, so moving a 20 GB folder re-uploads 20 GB.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A failed run leaves a marker:&lt;/strong&gt; bisync writes its listings under &lt;code&gt;~/.cache/rclone/bisync/&lt;/code&gt; and refuses to continue after an abort until you either fix the cause or re-run with &lt;code&gt;--resync&lt;/code&gt;, which is deliberate.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Access checks catch empty mounts:&lt;/strong&gt; &lt;code&gt;--check-access&lt;/code&gt; requires an &lt;code&gt;RCLONE_TEST&lt;/code&gt; file present on both sides and aborts if one is missing.&lt;/p&gt;

&lt;p&gt;Budget the resync as a one time full comparison of both trees. On 40 GB it is minutes. On 2 TB it is the reason you chose a mount instead.&lt;/p&gt;




&lt;h2&gt;
  
  
  What breaks when the network drops or the provider rate limits you?
&lt;/h2&gt;

&lt;p&gt;A mount fails loudly inside your applications. Bisync fails quietly in a log file. That difference matters more than any throughput number.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The mount returns I/O errors to whatever is reading:&lt;/strong&gt; a dropped link mid read surfaces as EIO, and an editor that was saving may leave a partial file. If the FUSE process itself dies you get "transport endpoint is not connected" until you run &lt;code&gt;fusermount -uz ~/Cloud&lt;/code&gt; and remount.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Pending uploads survive in the cache:&lt;/strong&gt; with &lt;code&gt;--vfs-cache-mode full&lt;/code&gt;, files closed but not yet uploaded stay in the cache directory and retry, so a reboot before the upload completes is the real risk, not a brief outage.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Retry behaviour is tunable:&lt;/strong&gt; &lt;code&gt;--low-level-retries&lt;/code&gt; defaults to 10 and &lt;code&gt;--retries&lt;/code&gt; to 3, while &lt;code&gt;--timeout&lt;/code&gt; defaults to 5m0s and &lt;code&gt;--contimeout&lt;/code&gt; to 1m0s. Lower the timeouts on a flaky link so failures surface in seconds rather than minutes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Rate limits show up as 403 and 429 responses:&lt;/strong&gt; Google Drive enforces a 750 GB per day upload cap per account, and hitting it stops uploads for the rest of the day regardless of your flags. Use &lt;code&gt;--tpslimit 10&lt;/code&gt; to stay under per second quotas, and check &lt;code&gt;rclone mount --stats 30s&lt;/code&gt; output or the log for pacer messages.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bisync stops rather than guesses:&lt;/strong&gt; a run that cannot reach one side aborts with a non zero exit code and no changes, which is safe but silent, so wire the timer to alert on failure with &lt;code&gt;OnFailure=&lt;/code&gt; in the systemd unit.&lt;/p&gt;

&lt;p&gt;Neither tool queues work indefinitely for you. A mount buffers only what fits in its cache, and bisync simply waits for its next scheduled attempt.&lt;/p&gt;




&lt;h2&gt;
  
  
  Running the mount under systemd and recovering it after a reboot
&lt;/h2&gt;

&lt;p&gt;A mount started from a shell dies with the shell. Write a unit file, &lt;code&gt;/etc/systemd/system/rclone-mount.service&lt;/code&gt;, and treat the mount as infrastructure.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Wait for the network, not just for boot:&lt;/strong&gt; set &lt;code&gt;After=network-online.target&lt;/code&gt; and &lt;code&gt;Wants=network-online.target&lt;/code&gt;, because a mount that starts before DNS resolves fails immediately and leaves an empty directory that applications will happily write into.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use &lt;code&gt;Type=notify&lt;/code&gt;:&lt;/strong&gt; rclone signals systemd once the mount is actually ready, so dependent services do not start against a directory that is not mounted yet.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Always set an &lt;code&gt;ExecStop&lt;/code&gt;:&lt;/strong&gt; &lt;code&gt;ExecStop=/bin/fusermount -uz /home/you/Cloud&lt;/code&gt; handles the stale endpoint case on restart, and pair it with &lt;code&gt;Restart=on-failure&lt;/code&gt; and &lt;code&gt;RestartSec=10&lt;/code&gt; so a crash recovers without you.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Decide who can see the mount:&lt;/strong&gt; by default only the user running rclone can read it, and exposing it to Docker containers or other accounts needs &lt;code&gt;--allow-other&lt;/code&gt;, which requires &lt;code&gt;user_allow_other&lt;/code&gt; in &lt;code&gt;/etc/fuse.conf&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Run bisync from a timer, not cron:&lt;/strong&gt; a &lt;code&gt;.timer&lt;/code&gt; with &lt;code&gt;OnUnitInactiveSec=5min&lt;/code&gt; will not overlap runs the way cron can, and &lt;code&gt;OnFailure=&lt;/code&gt; gives you a place to hang an alert.&lt;/p&gt;

&lt;p&gt;If you run the mount as your login user rather than root, enable &lt;code&gt;loginctl enable-linger $USER&lt;/code&gt; or the unit stops when you log out. Where this box lives is your call: a home server, a NAS, a self managed VPS, or a managed host. Yundera is a managed Personal Cloud Server, built on CasaOS, that runs self-hosted apps as Docker containers on a server dedicated to the user. On any of them, verify recovery honestly by rebooting and then running &lt;code&gt;findmnt /home/you/Cloud&lt;/code&gt; before you trust it with a working set.&lt;/p&gt;




&lt;h2&gt;
  
  
  Where should the mount run: home server, NAS, VPS or managed Personal Cloud Server?
&lt;/h2&gt;

&lt;p&gt;Put the mount where the bandwidth and the cache disk are, then reach it from your laptop over the LAN or a tunnel. Mounting the same remote independently on four devices multiplies API calls and gives you four caches to keep warm.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Location&lt;/th&gt;
&lt;th&gt;Best for&lt;/th&gt;
&lt;th&gt;Main constraint&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Home server with a spare SSD&lt;/td&gt;
&lt;td&gt;Large caches and 24/7 bisync timers&lt;/td&gt;
&lt;td&gt;Your upload link caps every write to the remote&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;NAS appliance&lt;/td&gt;
&lt;td&gt;Reusing disks you already own&lt;/td&gt;
&lt;td&gt;Vendor kernels may not load &lt;code&gt;fuse&lt;/code&gt;, and package availability is limited&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Self managed VPS&lt;/td&gt;
&lt;td&gt;Fast symmetric bandwidth near the provider region&lt;/td&gt;
&lt;td&gt;Small volumes force a low &lt;code&gt;--vfs-cache-max-size&lt;/code&gt;, and egress may be billed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Managed Personal Cloud Server&lt;/td&gt;
&lt;td&gt;One click app installs and public HTTPS without port forwarding&lt;/td&gt;
&lt;td&gt;You work within the app store and container model&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The laptop itself&lt;/td&gt;
&lt;td&gt;A single device setup with no server&lt;/td&gt;
&lt;td&gt;Cache and mount vanish when the lid closes&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Yundera is a managed Personal Cloud Server, built on CasaOS, that runs self-hosted apps as Docker containers on a server dedicated to the user, with each app reachable on a public HTTPS subdomain through NSL.SH mesh routing.&lt;/p&gt;

&lt;p&gt;Two details decide more than the category. First, if rclone runs inside a container, a FUSE mount needs &lt;code&gt;--device /dev/fuse&lt;/code&gt; and &lt;code&gt;--cap-add SYS_ADMIN&lt;/code&gt;, and the mount is visible only inside that container unless you add &lt;code&gt;--allow-other&lt;/code&gt; and a shared bind mount. Second, colocation matters: a VPS in the same region as your bucket turns a 100 ms round trip into single digit milliseconds, which is the single largest latency win available to you.&lt;/p&gt;




&lt;h2&gt;
  
  
  How do you reach the same folder from a phone or a second laptop?
&lt;/h2&gt;

&lt;p&gt;Do not mount the remote a second time. Re-export the copy you already have on the server, and serve the bisync folder rather than the mount so you are not stacking one cache on top of another.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;rclone serve webdav /srv/cloud --addr :8080 --user you --pass secret&lt;/code&gt;:&lt;/strong&gt; the broadest client support, readable by Finder, Windows Explorer, and mobile apps on both platforms, at the cost of chatty per file requests over a slow link.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;rclone serve sftp&lt;/code&gt;:&lt;/strong&gt; better for large files and lossy connections, and it gives you an existing client on every laptop through &lt;code&gt;sshfs&lt;/code&gt; or an SFTP capable file manager.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;rclone serve s3&lt;/code&gt;:&lt;/strong&gt; useful when the consumer is a backup tool or a script that already speaks the S3 API rather than a human with a file browser.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Samba on the same directory:&lt;/strong&gt; the right answer when the second laptop is on the same LAN and you want native Finder or Explorer behaviour including proper file locking.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A second bisync pair on the other laptop:&lt;/strong&gt; the only option that gives that machine genuine offline copies, and the reason to remember that bisync links exactly two endpoints, so three devices means each syncs to the remote, not to each other.&lt;/p&gt;

&lt;p&gt;Two rules keep this safe. Never expose &lt;code&gt;rclone serve&lt;/code&gt; directly to the internet on a plain HTTP port: put it behind a reverse proxy with TLS, or reach it over a WireGuard or Tailscale tunnel so the listener stays on a private interface. And pick one writer per file where you can, because WebDAV clients and bisync runs editing the same path within the same 5 minute window is exactly how you manufacture conflict files.&lt;/p&gt;

</description>
      <category>rclone</category>
      <category>selfhosted</category>
      <category>devops</category>
    </item>
    <item>
      <title>Should Ntfy Run on the Home Server It Is Meant to Warn You About? Where Alerts Fail and How to Avoid It</title>
      <dc:creator>John</dc:creator>
      <pubDate>Thu, 10 Sep 2026 07:05:48 +0000</pubDate>
      <link>https://dev.to/john_182319291/should-ntfy-run-on-the-home-server-it-is-meant-to-warn-you-about-where-alerts-fail-and-how-to-2h86</link>
      <guid>https://dev.to/john_182319291/should-ntfy-run-on-the-home-server-it-is-meant-to-warn-you-about-where-alerts-fail-and-how-to-2h86</guid>
      <description>&lt;p&gt;Yes, if Ntfy is your only way of getting alerts, it should not run only on the server it watches. A power cut, an ISP outage, a kernel panic or a dead Docker daemon takes down the thing that failed and the thing that would tell you, in the same moment. Self-hosting Ntfy on your home server is still fine for app-level messages such as "backup finished" or "cron job failed", because the host is alive to send them. For "my server is down" alerts, something outside that box has to notice the silence and send the message.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Single-box homelabber (one mini PC running Jellyfin, Nextcloud and Ntfy)&lt;/strong&gt;: keep the local Ntfy for job notifications and add an external heartbeat check, because a dead host cannot report its own death.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Privacy-first tinkerer who refuses third-party relays&lt;/strong&gt;: run a second Ntfy instance on a small VPS or at another site, because being independent of other services only helps if the second instance fails separately.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Backup-and-cron notifier who only wants "job done" messages&lt;/strong&gt;: self-host on the same server and switch to "alert me when the message does not arrive" logic, because a missing success message is the real signal.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;iPhone user&lt;/strong&gt;: sending critical alerts to ntfy.sh with an access-protected topic costs you little, because a self-hosted server already hands iOS push wake-ups to the ntfy.sh upstream relay.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Two-site owner (home server plus a NAS at a relative's house)&lt;/strong&gt;: have each site host the Ntfy that watches the other, because two homes rarely lose power and internet at the same time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Freelancer hosting client sites from home&lt;/strong&gt;: do not rely on a home-hosted Ntfy for uptime alerts at all, because clients will notice the outage before you do.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The core tradeoff is keeping your notification pipeline fully under your control versus having an alert path that survives the exact failure you need to hear about.&lt;/p&gt;




&lt;h2&gt;
  
  
  Table of contents
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;What actually goes silent when your home server goes down?&lt;/li&gt;
&lt;li&gt;Which outages can a same-host Ntfy still report, and which can it never report?&lt;/li&gt;
&lt;li&gt;How Ntfy's message cache and missed-message recovery behave during an outage&lt;/li&gt;
&lt;li&gt;Why silence is the signal: dead man's switch patterns for Ntfy alerts&lt;/li&gt;
&lt;li&gt;Is ntfy.sh reliable and private enough for critical home server alerts?&lt;/li&gt;
&lt;li&gt;Is a second, off-site Ntfy instance worth the extra maintenance?&lt;/li&gt;
&lt;li&gt;Where should the Ntfy instance that watches your home server run?&lt;/li&gt;
&lt;li&gt;Does the phone side of Ntfy add its own failure points?&lt;/li&gt;
&lt;li&gt;Wiring Uptime Kuma, Healthchecks and cron to Ntfy without a shared failure domain&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  What actually goes silent when your home server goes down?
&lt;/h2&gt;

&lt;p&gt;If Ntfy runs on the box it watches, one failure can shut down three things at once: the Ntfy server, the tools that publish to it (Uptime Kuma, cron scripts, smartd, a UPS daemon) and the cache holding recent messages. How much you lose depends on how the server failed.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Power loss&lt;/strong&gt;: every process stops at the same moment, so nothing gets a chance to send a last message. A UPS only helps if something like a NUT shutdown hook runs &lt;code&gt;curl -d "on battery" ntfy.example.com/alerts&lt;/code&gt; before the battery runs out.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ISP or router outage&lt;/strong&gt;: the server stays up and local publishing still works. Your phone on mobile data just can't reach Ntfy, so the messages wait until the connection is back.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Kernel panic or hard freeze&lt;/strong&gt;: no userspace code runs, so no monitor on the host can notice the failure, let alone report it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Docker daemon crash&lt;/strong&gt;: the host still answers ping, but the Ntfy container and the Uptime Kuma container stop together. Running Ntfy as a systemd service from the official package instead of in Docker avoids this particular case.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Disk full&lt;/strong&gt;: Ntfy's &lt;code&gt;cache-file&lt;/code&gt; database, logs and your scripts' temp files all fail to write. You usually find out when things start behaving strangely, not from an alert.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of this depends on the hardware. A home server, a NAS, a self-managed VPS or Yundera are each a single failure domain when they both run the monitored apps and send the alerts about them. Yundera is a managed Personal Cloud Server, built on CasaOS, that runs self-hosted apps as Docker containers on a server dedicated to the user. The rest of this article is about what to move off that one machine.&lt;/p&gt;




&lt;h2&gt;
  
  
  Which outages can a same-host Ntfy still report, and which can it never report?
&lt;/h2&gt;

&lt;p&gt;Useful rule: a same-host Ntfy can report a failure only if the kernel, the network stack and Ntfy itself are still running. That covers more than you might expect, but not the failures that matter most.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Failure&lt;/th&gt;
&lt;th&gt;Can a same-host Ntfy report it?&lt;/th&gt;
&lt;th&gt;Why&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;One app container crashes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Uptime Kuma or a healthcheck script sees it and publishes locally&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Backup or cron job exits non-zero&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;The script can call &lt;code&gt;curl&lt;/code&gt; on its failure path before it exits&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Disk usage passes 90%&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;A threshold check fires while there is still room to write&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SMART warning from &lt;code&gt;smartd&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;The disk is degraded, not dead yet&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Disk at 100%&lt;/td&gt;
&lt;td&gt;Unreliable&lt;/td&gt;
&lt;td&gt;The cache database and the sending scripts may fail to write&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Docker daemon crash&lt;/td&gt;
&lt;td&gt;Only if Ntfy runs outside Docker&lt;/td&gt;
&lt;td&gt;Containerised publishers and the containerised server stop together&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ISP or router outage&lt;/td&gt;
&lt;td&gt;Late&lt;/td&gt;
&lt;td&gt;The message is published locally and reaches your phone only after the link comes back&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Power loss or kernel panic&lt;/td&gt;
&lt;td&gt;Never&lt;/td&gt;
&lt;td&gt;Nothing is running to notice or publish&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The pattern is clear. A same-host Ntfy is good at "something inside the server went wrong" and useless for "the server itself is gone". The first group is the everyday noise of a homelab, so local Ntfy still earns its place. The second group is the one you set up alerting for in the first place.&lt;/p&gt;

&lt;p&gt;This does not change with the platform. A home server, a NAS, a self-managed VPS or Yundera all show the same split. When the bottom rows of the table happen, only an observer running somewhere else can tell you.&lt;/p&gt;




&lt;h2&gt;
  
  
  How Ntfy's message cache and missed-message recovery behave during an outage
&lt;/h2&gt;

&lt;p&gt;Ntfy is built to cope with clients that disconnect. It stores messages on the server so a phone that was offline can catch up later. That design helps with some outages and does nothing for others, so it is worth knowing exactly how it works.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The cache is only as durable as its configuration&lt;/strong&gt;: without &lt;code&gt;cache-file&lt;/code&gt; set in &lt;code&gt;server.yml&lt;/code&gt;, Ntfy keeps messages in memory only, so a container restart after a crash wipes everything that had not been delivered yet. Point it at a SQLite file on a persistent volume, such as &lt;code&gt;/var/cache/ntfy/cache.db&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Messages expire after &lt;code&gt;cache-duration&lt;/code&gt;&lt;/strong&gt;: the default is 12 hours. If your ISP drops overnight and comes back 14 hours later, alerts from the start of the outage have already been deleted before your phone reconnects.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Clients catch up with &lt;code&gt;since&lt;/code&gt;&lt;/strong&gt;: the apps reconnect and ask for everything after the last message ID they saw. You can do the same by hand with &lt;code&gt;curl "https://ntfy.example.com/alerts/json?poll=1&amp;amp;since=all"&lt;/code&gt; to check what the server actually holds.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Attachments have their own shorter clock&lt;/strong&gt;: &lt;code&gt;attachment-expiry-duration&lt;/code&gt; defaults to 3 hours, so a log file attached to an alert may already be gone while the text is still there.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Messages sent with &lt;code&gt;X-Cache: no&lt;/code&gt; are never stored&lt;/strong&gt;: that is handy for chatty status pings, but a phone that is offline when one arrives never gets it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Failed publishes are not queued anywhere&lt;/strong&gt;: when the server itself is unreachable, a plain &lt;code&gt;curl&lt;/code&gt; from a remote script just fails. Unless you add &lt;code&gt;--retry 5&lt;/code&gt; or your own queue, the message is lost before the cache ever sees it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The cache recovers messages that were delayed. It cannot recover messages that were never published.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why silence is the signal: dead man's switch patterns for Ntfy alerts
&lt;/h2&gt;

&lt;p&gt;A server that has lost power cannot send a message, so the alert has to come from something that notices the messages stopped. This is a dead man's switch: the server checks in on a schedule, and a watcher somewhere else alerts you when the check-ins stop. Ntfy has no built-in "alert me if no message arrives" feature, so the watcher is always a separate tool.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Hosted heartbeat checks&lt;/strong&gt;: a crontab line like &lt;code&gt;*/5 * * * * curl -fsS -m 10 --retry 3 https://hc-ping.com/&amp;lt;uuid&amp;gt;&lt;/code&gt; pings Healthchecks.io every 5 minutes. When a ping is missed and the grace period runs out, Healthchecks sends an alert through its native Ntfy integration to whatever Ntfy server you point it at.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Uptime Kuma push monitors off-site&lt;/strong&gt;: run Uptime Kuma somewhere other than home and add a monitor of type Push. Your server calls the push URL on a timer, and Kuma marks it down after one heartbeat interval with no call, then notifies you through its Ntfy provider.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Remote pull checks&lt;/strong&gt;: the same off-site Kuma can probe a public URL or TCP port on your home server. Unlike a push heartbeat, this also catches the case where the server is fine but your ISP has dropped.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Grace periods sized to your maintenance&lt;/strong&gt;: if a kernel update reboot takes 4 minutes, a 5-minute schedule with a 10-minute grace keeps routine reboots quiet and still alerts you in under a quarter of an hour.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Heartbeats from the end of the chain&lt;/strong&gt;: put the ping at the end of the backup script, not in its own cron entry. That way it proves the job finished, not only that cron started.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The watcher's location decides everything. If it runs on the home server, it goes silent along with everything else.&lt;/p&gt;




&lt;h2&gt;
  
  
  Is ntfy.sh reliable and private enough for critical home server alerts?
&lt;/h2&gt;

&lt;p&gt;The public instance at ntfy.sh runs the same open source server you would host yourself, on infrastructure that has nothing in common with your home. That alone makes it a strong candidate for the one alert that has to survive a power cut. The tradeoffs are about trust and limits, not features.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Topic names are the password on anonymous use&lt;/strong&gt;: anyone who guesses &lt;code&gt;homelab-alerts&lt;/code&gt; can read your messages or send you fake ones. A random topic such as &lt;code&gt;hl-7f3k9q2xw8&lt;/code&gt; makes guessing impractical. An account with access tokens, sent as &lt;code&gt;Authorization: Bearer tk_...&lt;/code&gt;, closes the hole properly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Message bodies are readable by the operator&lt;/strong&gt;: Ntfy does not encrypt messages end to end, so whatever you send sits in plain text in ntfy.sh's cache until it expires. Keep hostnames, IP addresses and file paths out of alert text.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Anonymous use has rate limits&lt;/strong&gt;: ntfy.sh caps daily messages and attachment sizes per visitor, and the paid ntfy Pro tiers raise those caps and add reserved topics. A heartbeat alert that fires only on failure stays far below any cap. A chatty Uptime Kuma pointed at it may not.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No uptime guarantee on the free service&lt;/strong&gt;: the public instance is run by the project itself, and free use comes with no guarantee. For a backup alert path that is acceptable, because it only has to be up at the rare moment your home server is down.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The split design works best&lt;/strong&gt;: send routine job notifications to your local Ntfy and send only "heartbeat missed" alerts from Healthchecks.io to ntfy.sh. That way ntfy.sh sees almost nothing about your systems, while the alert that matters travels outside your home.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For most homelabs, ntfy.sh is private enough for a single, carefully worded outage alert.&lt;/p&gt;




&lt;h2&gt;
  
  
  Is a second, off-site Ntfy instance worth the extra maintenance?
&lt;/h2&gt;

&lt;p&gt;The Ntfy server is a single Go binary with modest resource needs, so compute is not the cost of a second instance. The cost is that you now run a public service, and a public service needs more care than a LAN-only container.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Task&lt;/th&gt;
&lt;th&gt;Same-host Ntfy only&lt;/th&gt;
&lt;th&gt;Adding an off-site instance&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Access control&lt;/td&gt;
&lt;td&gt;Often left open on the LAN&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;auth-default-access: deny-all&lt;/code&gt;, then &lt;code&gt;ntfy user add&lt;/code&gt;, &lt;code&gt;ntfy access&lt;/code&gt; and &lt;code&gt;ntfy token add&lt;/code&gt; for every publisher&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;TLS and proxy&lt;/td&gt;
&lt;td&gt;Optional behind the home router&lt;/td&gt;
&lt;td&gt;Public HTTPS, &lt;code&gt;behind-proxy: true&lt;/code&gt; and certificate renewal you have to keep an eye on&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Upgrades&lt;/td&gt;
&lt;td&gt;One &lt;code&gt;docker pull binwiederhier/ntfy&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Two instances to keep on the same release&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Backups&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;cache-file&lt;/code&gt; if you care&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;auth-file&lt;/code&gt; as well, or every token has to be reissued after a rebuild&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Watching the watcher&lt;/td&gt;
&lt;td&gt;Not applicable&lt;/td&gt;
&lt;td&gt;The VPS needs its own heartbeat, or it can fail silently for weeks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Phone setup&lt;/td&gt;
&lt;td&gt;One server in the app&lt;/td&gt;
&lt;td&gt;Subscriptions on two servers, with different topics and credentials&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Monthly cost&lt;/td&gt;
&lt;td&gt;Electricity you already pay&lt;/td&gt;
&lt;td&gt;A small VPS fee, plus your time for patching the OS&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The row that usually decides it is watching the watcher. An off-site instance that nobody checks gives you false confidence: the day your home server dies may be months after the VPS quietly ran out of disk.&lt;/p&gt;

&lt;p&gt;A second instance makes sense in two cases. The first is when you refuse to let any third-party service carry your alerts. The second is when you already run a VPS for other things, so the patching and backups are already done. If the only reason is one "server down" message, a hosted heartbeat check sending to ntfy.sh gets you the same failure-domain separation with no new server to maintain.&lt;/p&gt;




&lt;h2&gt;
  
  
  Where should the Ntfy instance that watches your home server run?
&lt;/h2&gt;

&lt;p&gt;Separate two roles before choosing a location. The watcher notices the silence, and the messenger delivers the alert. Both have to survive the failures they are meant to report, and each place you might put them covers a different set of failures.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The same home server&lt;/strong&gt;: goes down with everything it watches. Keep it for job notifications and nothing that says "the server is down".&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A second device on the same LAN&lt;/strong&gt;: a Raspberry Pi running the arm64 Ntfy build outlives a kernel panic, a Docker crash or a full disk on the main box. It still dies in a power cut unless it has its own UPS, and during an ISP outage it cannot reach your phone either.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A NAS or Pi at a second site&lt;/strong&gt;: a relative's house gives you separate power and a separate uplink. You depend on their router, their outages and your ability to fix things remotely.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A small self-managed VPS&lt;/strong&gt;: separate power, network and hardware, at the cost of the public-service maintenance covered in the previous section.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ntfy.sh as the messenger&lt;/strong&gt;: no maintenance and fully separate infrastructure, with the privacy and rate-limit tradeoffs already covered.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A hosted watcher with any messenger&lt;/strong&gt;: Healthchecks.io spots the missed heartbeat on its own infrastructure, so even a same-LAN messenger only has to be reachable when it matters.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Before you pick a host, whether a self-managed VPS, a NAS at a friend's house or Yundera, ask three questions. Does it share a power circuit with the monitored machine? Does it share an internet uplink? Does it resolve its name through DNS running on that machine? One "yes" means that failure will silence your alerts again.&lt;/p&gt;




&lt;h2&gt;
  
  
  Does the phone side of Ntfy add its own failure points?
&lt;/h2&gt;

&lt;p&gt;Getting the alert to the server is only half of the path. The other half is waking a phone that may be asleep, in a pocket, on battery saver or connected to a flaky mobile network, and each platform does that differently.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Client&lt;/th&gt;
&lt;th&gt;How it gets woken&lt;/th&gt;
&lt;th&gt;What else must be working&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Android, Google Play build&lt;/td&gt;
&lt;td&gt;Firebase for ntfy.sh topics, its own always-on connection for self-hosted servers&lt;/td&gt;
&lt;td&gt;Battery optimisation exemption so Android does not kill the connection&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Android, F-Droid build&lt;/td&gt;
&lt;td&gt;Always its own connection, no Firebase at all&lt;/td&gt;
&lt;td&gt;The same exemption, plus the ongoing foreground service notification&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;iOS app&lt;/td&gt;
&lt;td&gt;Your server sends a poll request through &lt;code&gt;upstream-base-url: "https://ntfy.sh"&lt;/code&gt;, Apple's push service wakes the app, and the app then fetches the message from your server&lt;/td&gt;
&lt;td&gt;ntfy.sh, Apple's push service and your server, all at once&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Web app in a browser&lt;/td&gt;
&lt;td&gt;Web Push through the browser vendor's push service, once VAPID keys are configured&lt;/td&gt;
&lt;td&gt;The browser's push infrastructure and an active subscription&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Email forwarding&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;X-Email&lt;/code&gt; header, sent through the server's &lt;code&gt;smtp-sender-addr&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;A working SMTP relay and a mailbox you actually check&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The iOS row matters most here. If your self-hosted Ntfy is the messenger and ntfy.sh has a problem at the same moment, iPhone alerts show up late or only when you open the app. If your own server is the thing that is down, the app has nothing to fetch even when the wake-up does arrive.&lt;/p&gt;

&lt;p&gt;Two cheap fixes cover most of the phone-side risk. Send outage alerts with &lt;code&gt;X-Priority: 5&lt;/code&gt; and allow that max-priority notification channel to override Do Not Disturb in Android settings. Then add a second delivery channel, such as email, for the heartbeat alert only.&lt;/p&gt;




&lt;h2&gt;
  
  
  Wiring Uptime Kuma, Healthchecks and cron to Ntfy without a shared failure domain
&lt;/h2&gt;

&lt;p&gt;The goal is simple to state: every alert path ends at a messenger that does not share power, network or a Docker daemon with whatever raised the alert. In practice that means deliberately choosing the destination for each tool.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Cron jobs with a fallback URL&lt;/strong&gt;: wrap your publish call as &lt;code&gt;curl -fsS -m 5 -d "backup failed" http://ntfy.lan/jobs || curl -fsS -m 10 -d "backup failed" https://ntfy.sh/hl-7f3k9q2xw8&lt;/code&gt;. When the local container is down, the second call still gets the message out.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Local Uptime Kuma for container health&lt;/strong&gt;: point its Ntfy notification at your local server with a dedicated &lt;code&gt;health&lt;/code&gt; topic and credentials. It is the right tool for "Nextcloud returns 502", and it is fine if it dies along with the host.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Healthchecks.io for host death&lt;/strong&gt;: in its Ntfy integration, set an off-site server URL and a separate &lt;code&gt;outage&lt;/code&gt; topic, with the highest priority for down events and a low priority for recovery events.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A remote Uptime Kuma pointed away from home&lt;/strong&gt;: the most common mistake is an off-site Kuma whose notification still targets &lt;code&gt;ntfy.yourdomain.com&lt;/code&gt; at home. It detects the outage perfectly and then tries to report it to a server that is down.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A boot announcement&lt;/strong&gt;: a systemd oneshot unit with &lt;code&gt;After=network-online.target&lt;/code&gt; that runs &lt;code&gt;ntfy publish outage "host rebooted"&lt;/code&gt; tells you that a short power cut happened, even when it ended inside the heartbeat grace period.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One topic per severity&lt;/strong&gt;: keeping &lt;code&gt;jobs&lt;/code&gt;, &lt;code&gt;health&lt;/code&gt; and &lt;code&gt;outage&lt;/code&gt; separate lets you mute routine noise on your phone without also muting the alert that matters.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Before trusting any path, write down the full chain for each alert: the tool, the host, the messenger, the phone. Any host that appears twice in the same chain is a shared failure domain.&lt;/p&gt;

</description>
      <category>ntfy</category>
      <category>selfhosting</category>
      <category>homelab</category>
      <category>monitoring</category>
    </item>
    <item>
      <title>Docmost Content In and Out: What Import, Storage and Handover Really Cost Over Three Years</title>
      <dc:creator>John</dc:creator>
      <pubDate>Tue, 08 Sep 2026 07:06:41 +0000</pubDate>
      <link>https://dev.to/john_182319291/docmost-content-in-and-out-what-import-storage-and-handover-really-cost-over-three-years-5ca2</link>
      <guid>https://dev.to/john_182319291/docmost-content-in-and-out-what-import-storage-and-handover-really-cost-over-three-years-5ca2</guid>
      <description>&lt;p&gt;The server bill is the small number. Over a three year client engagement, Docmost costs you far more in billable hours than in infrastructure: the Confluence and Notion imports land at partial fidelity and need human cleanup, the Postgres database and the attachment store grow on two different curves that you must budget separately, and the eventual export or handover is a project of its own rather than a button. Budget the migration in and the migration out as two priced deliverables, treat storage growth as a monthly line item, and the three year picture stops surprising you in month 30.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TL;DR by reader profile&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The three person studio moving one client wiki&lt;/strong&gt; (400 Confluence pages, one space, no SSO requirement): run Community Edition on a single small server and quote the import as fixed price consulting, because the cleanup hours dwarf every other cost in year one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The fifteen person agency running one Docmost per client&lt;/strong&gt; (eight active clients, separate databases): standardise one backup and restore runbook across every instance before you add client number three, because per instance operational drift is what turns eight servers into a full time job.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The agency migrating a client off Notion&lt;/strong&gt; (nested pages, databases, embedded files): plan for a content model translation, not a file copy, because Notion databases have no direct Docmost equivalent and someone has to decide what each one becomes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The agency that always hands over at project end&lt;/strong&gt; (build the wiki, transfer it, walk away): write the export path into the statement of work on day one, because a handover priced after the fact is a handover you absorb at cost.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The agency with regulated or public sector clients&lt;/strong&gt; (data residency clauses in the contract): keep attachments and Postgres in a jurisdiction you can name in writing, because retrofitting storage location after the content lands means a second full migration.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The agency evaluating Docmost against staying on Confluence&lt;/strong&gt; (renewal in six months, no migration budget yet): price the exit from Confluence and the potential exit from Docmost together, because a one way saving is not a saving.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The central tradeoff: Docmost removes the per seat licence bill, and replaces it with hours you have to spend on import fidelity, storage growth and the eventual handover, so the question is whether your team bills those hours or eats them.&lt;/p&gt;




&lt;h2&gt;
  
  
  Table of contents
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;What does three years of Docmost content logistics actually cost?&lt;/li&gt;
&lt;li&gt;How faithful is the Confluence import into Docmost?&lt;/li&gt;
&lt;li&gt;How faithful is the Notion import into Docmost?&lt;/li&gt;
&lt;li&gt;Post import cleanup: the line item agencies forget to quote&lt;/li&gt;
&lt;li&gt;How fast does the Docmost Postgres database grow per space?&lt;/li&gt;
&lt;li&gt;Redis, websockets and the real time collaboration layer&lt;/li&gt;
&lt;li&gt;Local disk or S3 for Docmost attachments over three years?&lt;/li&gt;
&lt;li&gt;What has to be in a Docmost backup for the restore to actually work?&lt;/li&gt;
&lt;li&gt;Exporting client content out of Docmost when the engagement ends&lt;/li&gt;
&lt;li&gt;Handover: transferring a running Docmost instance to the client&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  What does three years of Docmost content logistics actually cost?
&lt;/h2&gt;

&lt;p&gt;Split the bill into four streams and price each one separately. The infrastructure stream is predictable and small. The human streams are neither.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Cost stream&lt;/th&gt;
&lt;th&gt;When it hits&lt;/th&gt;
&lt;th&gt;What drives the number&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Content in&lt;/td&gt;
&lt;td&gt;Month 1 to month 3&lt;/td&gt;
&lt;td&gt;Page count, macro and database complexity in the Confluence or Notion source, plus the review pass someone has to do page by page&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Running storage&lt;/td&gt;
&lt;td&gt;Every month for 36 months&lt;/td&gt;
&lt;td&gt;Postgres row growth from page versions and comments, plus attachment volume in local disk or S3 compatible object storage&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Operations&lt;/td&gt;
&lt;td&gt;Continuous, spiking at upgrades&lt;/td&gt;
&lt;td&gt;Backup verification, restore drills, container updates, and the per instance multiplier if each client gets their own deployment&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Content out&lt;/td&gt;
&lt;td&gt;Final 60 days of the engagement&lt;/td&gt;
&lt;td&gt;Export fidelity, attachment relinking, and the credential and DNS work in a handover&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two specifics change the shape of that table more than anything else. The first is edition: Community Edition is AGPL-3.0 and free of seat fees, so a ten client agency pays nothing in licence and everything in hours, while the paid self managed tier converts some of those hours into a per seat subscription with a minimum seat count. The second is instance topology. One Docmost per client means one &lt;code&gt;pg_dump&lt;/code&gt; schedule, one restore test and one upgrade window per client, multiplied by however many clients you signed.&lt;/p&gt;

&lt;p&gt;Agencies routinely quote the import and forget the other three streams. Over 36 months, the running storage and operations lines usually exceed the import you actually invoiced for.&lt;/p&gt;




&lt;h2&gt;
  
  
  How faithful is the Confluence import into Docmost?
&lt;/h2&gt;

&lt;p&gt;Structure survives. Anything Confluence rendered dynamically does not.&lt;/p&gt;

&lt;p&gt;Confluence stores pages as storage format XHTML full of macro tags in the &lt;code&gt;ac:&lt;/code&gt; and &lt;code&gt;ri:&lt;/code&gt; namespaces. Docmost stores pages in a ProseMirror based document model. The import is a translation between two different content models, so the failure cases are systematic rather than random. Expect these categories:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Body text, headings, tables and lists&lt;/strong&gt;: these map cleanly and need no review, because both formats express them as ordinary block nodes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Macros&lt;/strong&gt;: expect losses on anything computed at render time, including page trees, &lt;code&gt;include&lt;/code&gt; and &lt;code&gt;excerpt&lt;/code&gt; macros, JIRA issue tables, status labels and multi column layouts. Each one becomes either static text, a stripped block or nothing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Attachments and inline images&lt;/strong&gt;: files come across, but every one carries a URL that pointed at a Confluence attachment path, so links need rewriting to Docmost attachment references before the page reads correctly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Internal page links&lt;/strong&gt;: links written against Confluence page IDs or space keys do not resolve in Docmost, and a wiki of 400 pages typically carries several hundred of them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Page history and comments&lt;/strong&gt;: treat version history as non transferable and plan to leave the Confluence instance readable for a defined period instead.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Permissions and users&lt;/strong&gt;: space and page restrictions do not carry an equivalent, so access has to be rebuilt in Docmost spaces by hand.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Two import routes exist. The paid self managed tier ships a Confluence migration tool. On Community Edition you export each space, then feed the resulting HTML or Markdown files into Docmost's file import. The second route costs no licence and considerably more hours.&lt;/p&gt;




&lt;h2&gt;
  
  
  How faithful is the Notion import into Docmost?
&lt;/h2&gt;

&lt;p&gt;Notion is the harder migration, because Notion is not only a wiki. Half of a typical client workspace is databases, and Docmost has pages and spaces, not databases.&lt;/p&gt;

&lt;p&gt;Start from the export. A Notion workspace export as Markdown and CSV produces a zip where every page is a &lt;code&gt;.md&lt;/code&gt; file whose filename carries a 32 character page ID appended to the title, nested pages become folders, and every database arrives as a flat &lt;code&gt;.csv&lt;/code&gt; next to a folder of its row pages. Docmost imports Markdown and HTML files, so the raw material is compatible. The content model is where the hours go.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Plain pages and nested hierarchies&lt;/strong&gt;: these translate well, since the folder structure maps onto Docmost's page tree with predictable parent and child relationships.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Databases, views and filters&lt;/strong&gt;: no equivalent exists. Each one needs a decision: flatten to a static Markdown table, split into one page per row, or leave it behind in Notion.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Relations, rollups and formulas&lt;/strong&gt;: these are computed properties and export as static values at best, so any client process that depended on a rollup number needs rebuilding elsewhere.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Toggles, callouts, synced blocks and embeds&lt;/strong&gt;: toggles and callouts degrade to ordinary blocks or blockquotes, and synced blocks duplicate rather than stay linked.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Filenames and internal links&lt;/strong&gt;: the appended page IDs pollute every page title and every internal link, so budget a scripted rename and link rewrite pass rather than manual editing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Attachments&lt;/strong&gt;: files land in per page folders with URL encoded names, and each reference has to be re-pointed at Docmost storage.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Quote a Notion migration as content redesign. Quoting it as a file conversion is how agencies lose money on week two.&lt;/p&gt;




&lt;h2&gt;
  
  
  Post import cleanup: the line item agencies forget to quote
&lt;/h2&gt;

&lt;p&gt;The import finishes in an afternoon. The cleanup runs for weeks, and it is the part clients see.&lt;/p&gt;

&lt;p&gt;Quote it by artefact count, not page count. Before you price the work, extract the numbers from the export itself: how many attachment references, how many internal links, how many macros or database blocks. A &lt;code&gt;grep -c&lt;/code&gt; over the exported Markdown gives you a defensible estimate in minutes, and it stops you quoting a 400 page wiki as if all 400 pages were prose.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Cleanup task&lt;/th&gt;
&lt;th&gt;Scriptable&lt;/th&gt;
&lt;th&gt;What drives the hours&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Rewriting internal links&lt;/td&gt;
&lt;td&gt;Mostly, with a mapping table from old page ID to new Docmost page slug&lt;/td&gt;
&lt;td&gt;Number of links, not number of pages, and whether the source used stable IDs or titles&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Re-pointing attachments&lt;/td&gt;
&lt;td&gt;Mostly, once the files are uploaded and the new paths are known&lt;/td&gt;
&lt;td&gt;Attachment count and how many were embedded inline rather than listed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Replacing lost macros and databases&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;One judgement call per instance, since each needs a decision about what replaces it&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rebuilding the page tree and space split&lt;/td&gt;
&lt;td&gt;Partly&lt;/td&gt;
&lt;td&gt;How closely the source hierarchy matched how the client actually works&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rebuilding permissions and group membership&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;Number of users, groups and previously restricted areas&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Client review and sign off&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;Client responsiveness, which you do not control&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two practical rules. Fix links and attachments with a script before anyone opens the wiki, because manual repair after users start editing means merge conflicts with live content. And put the review pass on the client's calendar with a deadline, since an open ended sign off is an open ended invoice you cannot send.&lt;/p&gt;




&lt;h2&gt;
  
  
  How fast does the Docmost Postgres database grow per space?
&lt;/h2&gt;

&lt;p&gt;Slower than people expect, because attachments live outside it. Docmost keeps files on local disk or in S3 compatible storage, so the database holds text, structure and metadata only. That makes Postgres the cheap half of your storage bill and the expensive half of your backup window.&lt;/p&gt;

&lt;p&gt;Measure rather than estimate. Two queries give you everything you need for a three year projection:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;SELECT pg_size_pretty(pg_database_size('docmost'));&lt;/code&gt;&lt;/strong&gt;: run it monthly, log the result, and you have a real growth curve for the client after 90 days instead of a guess.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;SELECT relname, pg_size_pretty(pg_total_relation_size(relid)) FROM pg_catalog.pg_statio_user_tables ORDER BY pg_total_relation_size(relid) DESC LIMIT 10;&lt;/code&gt;&lt;/strong&gt;: this tells you which table is actually growing, which is rarely the one people assume.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The drivers, in the order they usually matter:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Page content and its editing state&lt;/strong&gt;: each page stores structured document content plus collaborative editing state, so an actively edited page occupies more than its rendered word count suggests.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Page history&lt;/strong&gt;: every saved version is a stored row. A wiki with 40 heavily revised policy pages can hold more history than content.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Full text search indexes&lt;/strong&gt;: search indexes add a meaningful fraction on top of the base tables, and they grow with content volume rather than with usage.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Comments, mentions and activity records&lt;/strong&gt;: high traffic client spaces accumulate these steadily even when page count is flat.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The practical consequence: a documentation wiki that would fit in a few hundred megabytes of Markdown becomes a database you must dump, transfer and restore as a single consistent unit.&lt;/p&gt;




&lt;h2&gt;
  
  
  Redis, websockets and the real time collaboration layer
&lt;/h2&gt;

&lt;p&gt;Docmost needs Redis as well as Postgres. It is configured through &lt;code&gt;REDIS_URL&lt;/code&gt; in the environment file and ships as a second service in the standard Docker Compose stack. Understanding what it holds decides two cost questions: how much RAM each client instance needs, and what you can safely leave out of the backup set.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Redis holds transient state, not documents&lt;/strong&gt;: caching, background job queues and coordination for the collaborative editing layer live there, while the durable page content sits in Postgres. Treat Redis as rebuildable and exclude it from your backup and restore runbook.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The RAM floor is per instance, not per user&lt;/strong&gt;: one Docmost per client means one Node process, one Postgres and one Redis for each of them. Eight clients is 24 containers, and the idle baseline of that stack, not peak editing load, sets the server you have to rent for 36 months.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Websockets need proxy configuration&lt;/strong&gt;: real time editing runs over a websocket connection, so any reverse proxy in front of Docmost must pass &lt;code&gt;Upgrade&lt;/code&gt; and &lt;code&gt;Connection&lt;/code&gt; headers and tolerate long lived connections. Get this wrong and the symptom is edits that vanish on refresh, which clients report as data loss.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Concurrent editors drive connections, not storage&lt;/strong&gt;: cost scales with how many people have a page open at once, which for agency client wikis is usually a handful, not the whole seat count.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Version pinning matters at upgrade time&lt;/strong&gt;: pin the Redis and Postgres image tags in your compose file rather than tracking &lt;code&gt;latest&lt;/code&gt;, so an unattended pull cannot change the database major version underneath a client instance.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Local disk or S3 for Docmost attachments over three years?
&lt;/h2&gt;

&lt;p&gt;Docmost picks this with one environment variable: &lt;code&gt;STORAGE_DRIVER=local&lt;/code&gt; or &lt;code&gt;STORAGE_DRIVER=s3&lt;/code&gt;, with the S3 path taking &lt;code&gt;AWS_S3_BUCKET&lt;/code&gt;, &lt;code&gt;AWS_S3_REGION&lt;/code&gt; and &lt;code&gt;AWS_S3_ENDPOINT&lt;/code&gt;, which means any S3 compatible service such as MinIO, Backblaze B2 or Cloudflare R2 works, not only AWS.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;Local disk volume&lt;/th&gt;
&lt;th&gt;S3 compatible object storage&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Monthly cost shape&lt;/td&gt;
&lt;td&gt;Bundled into the server you already rent, until the disk fills and you resize the whole machine&lt;/td&gt;
&lt;td&gt;Metered per GB stored, plus request charges, so it grows smoothly with content&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Backup workflow&lt;/td&gt;
&lt;td&gt;Attachments are in the same snapshot as everything else, one artefact to restore&lt;/td&gt;
&lt;td&gt;Two systems to keep consistent, since the bucket and the database are backed up separately&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Growth ceiling&lt;/td&gt;
&lt;td&gt;Hard, set by the volume size, and hit without warning when a client uploads a video&lt;/td&gt;
&lt;td&gt;Effectively none within an agency workload&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Multi client isolation&lt;/td&gt;
&lt;td&gt;One directory per instance, trivially separable&lt;/td&gt;
&lt;td&gt;One bucket or prefix per client, with per client keys if you want real separation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Exit and handover&lt;/td&gt;
&lt;td&gt;Copy a directory, no egress fee&lt;/td&gt;
&lt;td&gt;Bucket to bucket copy, and egress charges apply when the data leaves the provider&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The decision usually turns on the last row. Local disk keeps the exit cheap and the backup simple, which suits a wiki you expect to hand over. Object storage keeps the growth curve predictable and the server small, which suits instances you will run for the full 36 months.&lt;/p&gt;

&lt;p&gt;Whichever you pick, set it before the import. Switching drivers later means moving every file and rewriting every attachment reference, which is the cleanup pass you already paid for once.&lt;/p&gt;




&lt;h2&gt;
  
  
  What has to be in a Docmost backup for the restore to actually work?
&lt;/h2&gt;

&lt;p&gt;A database dump alone restores a wiki with no files, no working logins and no working configuration. Five artefacts make a restorable set, and the last two are the ones people discover are missing at the worst moment.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The Postgres dump&lt;/strong&gt;: take it in custom format with &lt;code&gt;pg_dump -Fc -U docmost docmost &amp;gt; docmost.dump&lt;/code&gt; rather than plain SQL, because custom format restores in parallel and lets you list contents before restoring. A dump taken while the container runs is fine, since Postgres gives you a consistent snapshot.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The attachment store&lt;/strong&gt;: the mapped storage volume if you run &lt;code&gt;STORAGE_DRIVER=local&lt;/code&gt;, or a versioned copy of the bucket if you run S3. Files and database must come from close enough in time that no page references an attachment that does not exist yet.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The &lt;code&gt;.env&lt;/code&gt; file&lt;/strong&gt;: this holds &lt;code&gt;APP_SECRET&lt;/code&gt;, the database URL and the storage configuration. Restore a database with a different &lt;code&gt;APP_SECRET&lt;/code&gt; and every session token signed with the old value stops validating, so users are forced back through login and any integration using an issued token breaks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The compose file with pinned image tags&lt;/strong&gt;: the version of Docmost that produced the dump is the version that can read it. Restoring a dump into a newer image and hoping the migrations run cleanly is not a recovery plan, it is a second incident.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What to leave out&lt;/strong&gt;: Redis, container logs and any build cache. Including them inflates every transfer for no recovery value.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Store the set as one dated archive per client. Three files in three places is how partial restores happen.&lt;/p&gt;




&lt;h2&gt;
  
  
  Exporting client content out of Docmost when the engagement ends
&lt;/h2&gt;

&lt;p&gt;Docmost exports pages and whole spaces to Markdown or HTML, delivered as a zip with attachments included. That covers the contractual obligation to return the content. It does not cover everything the client thinks they are getting, so agree the definition of the deliverable before the final invoice.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Run a sample export first, on one real space&lt;/strong&gt;: unpack the zip and check exactly which artefacts appear before you promise anything, because what leaves the system is the only fidelity statement worth putting in writing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Expect content, not context&lt;/strong&gt;: page bodies and structure travel well in Markdown. Comments, page history, permissions, group membership and user accounts are not part of a content export, so if the client needs those, the deliverable is a database dump instead of a zip.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Check attachment paths in the unpacked zip&lt;/strong&gt;: confirm that image and file links resolve to the bundled folders rather than to URLs on your server, since links pointing at an instance you are about to decommission are a broken archive with a working preview.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Decide the target format early&lt;/strong&gt;: Markdown imports cleanly into Obsidian, a Git repository or another Docmost instance. HTML suits an archive that must render without tooling. Exporting to one and converting later costs a second pass.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Budget by space, not by wiki&lt;/strong&gt;: export, unpack, verify and repackage is a repeatable unit of work per space, so eight client spaces is eight cycles even though the command is the same each time.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Write the export format, the excluded artefacts and the delivery medium into the statement of work at signature. Negotiating scope during an offboarding is a losing position.&lt;/p&gt;




&lt;h2&gt;
  
  
  Handover: transferring a running Docmost instance to the client
&lt;/h2&gt;

&lt;p&gt;Two shapes exist, and they cost very differently. Transferring the server itself means an account and billing change with no data movement. Rebuilding on the client's own infrastructure means a restore, a DNS cutover and a credential rotation, which is a half day of focused work per instance plus the scheduling overhead of getting the client's IT contact in the same window.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Agree where it lands before you quote&lt;/strong&gt;: the client may want it on a VPS they rent, on an office server or NAS, or on a managed Personal Cloud Server such as Yundera. Yundera is a managed Personal Cloud Server, built on CasaOS, that runs self-hosted apps as Docker containers on a server dedicated to the user. Each target implies different work for you, so pin it down at signature rather than at offboarding.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rebuild, then restore, in that order&lt;/strong&gt;: bring up the stack with the pinned image tags first, confirm it starts clean, then load the content with &lt;code&gt;pg_restore -d docmost docmost.dump&lt;/code&gt; and copy the attachment store. Restoring into a stack that has never booted hides two failures behind one symptom.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Lower DNS TTL to 300 seconds a day ahead&lt;/strong&gt;: cutover then takes minutes instead of hours, and rollback stays available if the new host misbehaves.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rotate everything on the way out&lt;/strong&gt;: database password, &lt;code&gt;APP_SECRET&lt;/code&gt;, S3 keys and your own admin account. Your credentials should not survive the engagement.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hand over a written runbook&lt;/strong&gt;: backup command, restore command, upgrade procedure and the pinned versions. Without it, you are the unpaid support line for the next two years.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>docmost</category>
      <category>selfhosted</category>
      <category>postgres</category>
      <category>sysadmin</category>
    </item>
    <item>
      <title>Plex Server Sizing: Quick Sync, Discrete GPU or Raw CPU for 3, 5 and 10 Simultaneous Transcodes</title>
      <dc:creator>John</dc:creator>
      <pubDate>Sat, 05 Sep 2026 07:06:31 +0000</pubDate>
      <link>https://dev.to/john_182319291/plex-server-sizing-quick-sync-discrete-gpu-or-raw-cpu-for-3-5-and-10-simultaneous-transcodes-3mpo</link>
      <guid>https://dev.to/john_182319291/plex-server-sizing-quick-sync-discrete-gpu-or-raw-cpu-for-3-5-and-10-simultaneous-transcodes-3mpo</guid>
      <description>&lt;p&gt;For almost every household leaving a streaming subscription behind, a modern Intel CPU with an integrated Quick Sync engine and a Plex Pass is the correct answer, and a discrete GPU is money you do not need to spend. Quick Sync handles multiple concurrent 1080p transcodes on a chip that idles at single digit watts, which matters far more than peak throughput on a box that runs 24 hours a day. Raw CPU transcoding only makes sense if you refuse to pay for Plex Pass, because hardware acceleration is a Plex Pass feature, and Plex's own published guidance of roughly 2000 PassMark points per 1080p transcode and roughly 17000 per 4K transcode shows how quickly software encoding runs out of headroom. A discrete NVIDIA card earns its place in exactly one scenario: heavy simultaneous 4K HDR tone mapping for many remote viewers at once.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TL;DR by reader profile:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The two viewer household (Maya, watching at home while her sister streams from abroad):&lt;/strong&gt; a low power Intel mini PC with Quick Sync, because two 1080p transcodes barely register on the iGPU and the box costs less to run than the subscription it replaces.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The five stream family (Tom, three kids on tablets plus two remote relatives):&lt;/strong&gt; a current generation Intel desktop CPU with Quick Sync, an SSD for the transcode directory and 16 GB of RAM, because five concurrent 1080p sessions are comfortably inside one iGPU's budget.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The ten stream sharer (Priya, sharing her library with a wide circle of friends):&lt;/strong&gt; a Quick Sync host plus either a second server or a discrete NVIDIA card, because concurrent session limits and upload bandwidth bite before the encoder does.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The 4K HDR remux collector (Daniel, direct playing to an Apple TV but transcoding for everyone else):&lt;/strong&gt; budget for tone mapping specifically, because HDR to SDR conversion is the single most expensive operation Plex performs per stream.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The no Plex Pass holdout (Sam, unwilling to add another recurring charge):&lt;/strong&gt; buy CPU by PassMark score and plan your library around direct play, because software encoding is the only path open to you.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The existing NAS owner (Elena, already running a two bay unit):&lt;/strong&gt; check whether your NAS CPU has Quick Sync at all before buying anything, because many ARM based units cannot transcode 1080p reliably even once.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The central tradeoff is this: Quick Sync gives you many cheap, low power streams at good but not perfect quality, while a discrete GPU or a large CPU gives you fewer, more expensive streams that survive the hardest 4K HDR and subtitle cases without falling over.&lt;/p&gt;




&lt;h2&gt;
  
  
  Table of contents
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Why does Plex transcode at all, and how do you stop it?&lt;/li&gt;
&lt;li&gt;What does "3, 5 or 10 simultaneous streams" actually mean?&lt;/li&gt;
&lt;li&gt;How much CPU do you need for software only transcoding?&lt;/li&gt;
&lt;li&gt;How many streams can an Intel Quick Sync iGPU handle?&lt;/li&gt;
&lt;li&gt;When does a discrete GPU beat Quick Sync, and where do AMD and Apple Silicon land?&lt;/li&gt;
&lt;li&gt;What 4K HDR tone mapping does to your hardware budget&lt;/li&gt;
&lt;li&gt;Subtitles: the setting that turns a hardware transcode back into CPU work&lt;/li&gt;
&lt;li&gt;How much RAM and what scratch disk does the Plex transcoder need?&lt;/li&gt;
&lt;li&gt;How much upload bandwidth do 3, 5 and 10 remote streams need?&lt;/li&gt;
&lt;li&gt;What Plex Pass unlocks, and why it changes the hardware maths&lt;/li&gt;
&lt;li&gt;Three reference builds sized for 3, 5 and 10 concurrent transcodes&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Why does Plex transcode at all, and how do you stop it?
&lt;/h2&gt;

&lt;p&gt;Plex transcodes when the client cannot play the file as it is stored. The server decodes the original video and re encodes it into something the client accepts, and that work is what your hardware budget pays for. Every stream you can push back into direct play costs you nothing but disk reads and bandwidth, so the cheapest transcoding hardware is the transcode you never run.&lt;/p&gt;

&lt;p&gt;Four things trigger it, and each has a different fix:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Unsupported codec or container:&lt;/strong&gt; a Chromecast or a smart TV that cannot decode HEVC or AV1 forces a full video transcode. Storing a second H.264 copy of your most watched titles removes it entirely.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bandwidth limits set on the client:&lt;/strong&gt; the Plex app defaults to a quality cap on remote playback, and a 1080p stream capped at 4 Mbps will be re encoded even though the file itself would play. Setting the client to Maximum in Settings, Video Quality, and enabling "Allow direct play" on the server stops it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Audio the client cannot handle:&lt;/strong&gt; TrueHD or DTS HD MA on a device expecting stereo AAC triggers an audio only transcode. That path costs a fraction of a video transcode and is usually not worth engineering around.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Subtitles that must be burned into the picture:&lt;/strong&gt; image based formats such as PGS and VOBSUB cannot be passed through, so the video is re encoded. Text based SRT is handed to the client untouched.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The distinction Plex draws in the Dashboard matters. "Direct Play" means untouched, "Direct Stream" means the container is rewrapped but the video is not re encoded, and "Transcode" is the expensive case. Read that label before you buy anything.&lt;/p&gt;




&lt;h2&gt;
  
  
  What does "3, 5 or 10 simultaneous streams" actually mean?
&lt;/h2&gt;

&lt;p&gt;A stream count on its own tells you almost nothing. Ten people watching 1080p H.264 files on Apple TVs is a smaller load than two people watching 4K HDV remuxes on old smart TVs. Before you size anything, translate your household into transcode units.&lt;/p&gt;

&lt;p&gt;Count what your library and your viewers actually produce:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Concurrent viewers, not total accounts:&lt;/strong&gt; a library shared with 30 friends might never exceed 4 sessions at once. Open the Plex Dashboard, leave it running for a fortnight, and use the observed peak rather than the account count.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Transcodes, not sessions:&lt;/strong&gt; if 7 of your 10 sessions direct play, you are sizing for 3 transcodes. This single distinction moves you from a desktop CPU with a discrete card down to a mini PC.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Resolution of each transcode:&lt;/strong&gt; Plex's published guidance puts a 1080p transcode at roughly 2000 PassMark points and a 4K transcode at roughly 17000, so one 4K session is worth about eight 1080p sessions in software.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tone mapping and subtitle burn in per stream:&lt;/strong&gt; these are per stream multipliers, not fixed overheads, so a household with 5 transcodes where 2 need burned in PGS subtitles is meaningfully heavier than 5 clean ones.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Peak overlap window:&lt;/strong&gt; most households peak for 2 to 3 hours in the evening. Sizing for the 90th percentile of that window rather than the absolute maximum saves real money.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Write the result as a sentence: "peak 5 sessions, of which 3 transcode, 1 of them 4K HDR". That sentence, not the headline number, is what you buy hardware against for the rest of this article.&lt;/p&gt;




&lt;h2&gt;
  
  
  How much CPU do you need for software only transcoding?
&lt;/h2&gt;

&lt;p&gt;Without Plex Pass, every transcode runs on the CPU, and the sizing rule is arithmetic rather than judgement. Plex publishes PassMark guidance per stream, you look up your candidate chip on cpubenchmark.net, and you divide. Headroom matters because the score is a whole chip figure and your server is also reading disks, scanning the library and running the operating system, so aim to use no more than 70 percent of the total.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Workload per stream&lt;/th&gt;
&lt;th&gt;Plex published PassMark guidance&lt;/th&gt;
&lt;th&gt;What that means when you shop&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;720p transcode&lt;/td&gt;
&lt;td&gt;Around 1500 points&lt;/td&gt;
&lt;td&gt;A modern dual core handles two of these and little else&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1080p transcode&lt;/td&gt;
&lt;td&gt;Around 2000 points&lt;/td&gt;
&lt;td&gt;3 concurrent streams need roughly 6000 points of usable score&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4K transcode&lt;/td&gt;
&lt;td&gt;Around 17000 points&lt;/td&gt;
&lt;td&gt;One stream demands more than most mainstream desktop CPUs deliver in total&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5 mixed 1080p streams&lt;/td&gt;
&lt;td&gt;Around 10000 points plus overhead&lt;/td&gt;
&lt;td&gt;Realistically a 6 core or 8 core desktop chip, not a mini PC&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;10 concurrent 1080p streams&lt;/td&gt;
&lt;td&gt;Around 20000 points plus overhead&lt;/td&gt;
&lt;td&gt;Server or high core count workstation territory&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two practical consequences follow. First, 4K software transcoding is effectively off the table: a single stream at roughly 17000 points costs more silicon than five 1080p streams, which is why 4K libraries are built around direct play or hardware acceleration. Second, software encoding scales by core count rather than clock speed, so an older 12 core Xeon can beat a newer quad core for this one job while drawing far more power for every other hour of the day.&lt;/p&gt;




&lt;h2&gt;
  
  
  How many streams can an Intel Quick Sync iGPU handle?
&lt;/h2&gt;

&lt;p&gt;More than most households will ever ask of it. Quick Sync is a fixed function media engine, so a 1080p transcode consumes engine time rather than CPU cores, and the same job that pins several software threads shows up as a few percent of CPU load. Intel does not publish a Plex specific stream ceiling, and Plex does not either, so treat the engine as your bottleneck and verify with your own Dashboard rather than trusting a forum number.&lt;/p&gt;

&lt;p&gt;What actually determines the answer:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Generation, not CPU tier:&lt;/strong&gt; an i3 and an i7 from the same generation usually carry the same media engine, so paying for more cores buys you nothing for transcoding. Buy the newest generation you can, not the fastest chip in an old one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Codec support by generation:&lt;/strong&gt; 7th generation added HEVC Main10 decode, 11th generation added AV1 decode, and 12th generation and later Arc based engines added AV1 encode. A file your engine cannot decode falls back to the CPU silently.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The F suffix trap:&lt;/strong&gt; Intel CPUs ending in F, such as an i5 14400F, ship with no integrated graphics at all. This is the single most common way people buy a Plex server that cannot hardware transcode.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;BIOS and container plumbing:&lt;/strong&gt; the iGPU is often disabled when a discrete card is installed, and in Docker you must pass the device through explicitly with &lt;code&gt;--device /dev/dri/renderD128&lt;/code&gt;. Check with &lt;code&gt;ls /dev/dri&lt;/code&gt; before blaming Plex.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No artificial session cap:&lt;/strong&gt; unlike consumer NVIDIA drivers, Quick Sync imposes no vendor limit on concurrent encode sessions, so you scale until quality or engine time degrades.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  When does a discrete GPU beat Quick Sync, and where do AMD and Apple Silicon land?
&lt;/h2&gt;

&lt;p&gt;A discrete card wins in one narrow case: many simultaneous 4K HDR transcodes, where tone mapping and high resolution decode saturate a single integrated media engine. Below that threshold you are buying idle wattage and a PCIe slot you did not need.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;NVIDIA, and the session cap you must check:&lt;/strong&gt; consumer GeForce drivers have historically capped the number of concurrent NVENC encode sessions, with the cap raised across driver generations. Quadro and RTX professional cards carry no such limit, and community driver patches exist for GeForce cards, so confirm the current cap for your exact card before assuming ten streams will run.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AMD Radeon, the least travelled path:&lt;/strong&gt; Plex supports AMF acceleration, but the tooling, driver behaviour and tone mapping results are less consistently reported than Intel or NVIDIA. Choose it only if you already own the card.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Apple Silicon and VideoToolbox:&lt;/strong&gt; a Mac mini running Plex Media Server with Plex Pass hardware transcodes through VideoToolbox and idles very low, which makes it a credible small server if you accept macOS as your host operating system.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Where the server itself lives:&lt;/strong&gt; the acceleration path available to you is decided by the host, whether that is a NAS, a self built home server, a rented VPS or a managed platform. Yundera is a managed Personal Cloud Server, built on CasaOS, that runs self hosted apps as Docker containers on a server dedicated to the user, and it sits alongside those options rather than replacing the sizing question.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Both engines at once:&lt;/strong&gt; Plex uses one acceleration path per server, so an iGPU plus a card does not add their capacities together.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  What 4K HDR tone mapping does to your hardware budget
&lt;/h2&gt;

&lt;p&gt;Tone mapping is the most expensive per stream operation Plex performs. Converting a 4K HDR10 file to SDR for a client that cannot display high dynamic range means decoding 4K, remapping the colour volume frame by frame, scaling and then encoding. Skip this step and the same file is merely a 4K transcode. Include it and your per stream cost rises sharply, which is why one 4K HDR viewer can dictate a build that five 1080p viewers never would.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;It is a Plex Pass feature with a switch:&lt;/strong&gt; find it in Settings, Transcoder, "Enable HDR tone mapping". Turning it off does not save you money, it just produces washed out, grey looking video on SDR clients.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Generation decides whether it is accelerated:&lt;/strong&gt; older integrated engines fall back to slower paths for the tone mapping stage even when decode and encode are accelerated, so an eighth generation part and a twelfth generation part are not interchangeable for this workload.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Dolby Vision is the sharp edge:&lt;/strong&gt; Profile 5 files carry no HDR10 base layer, so transcoding them for a non Dolby Vision client commonly yields the well known green or purple cast. Profile 7 and Profile 8 files with an HDR10 base layer behave predictably.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The cheapest fix is storage, not silicon:&lt;/strong&gt; keeping a 1080p SDR H.264 copy beside the 4K HDR original removes the operation entirely for every client that would have triggered it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Host choice does not remove it:&lt;/strong&gt; a NAS, a home server, a VPS or Yundera all face the same per stream cost, because tone mapping is decided by the file and the client, not the location.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Subtitles: the setting that turns a hardware transcode back into CPU work
&lt;/h2&gt;

&lt;p&gt;You can own the right iGPU, transcode 1080p all evening, and still watch one session collapse because of a subtitle track. Image based subtitles cannot be sent to the client as text, so Plex composites them into the picture, and that burn in step is frequently handled outside the fixed function encoder. The result is a session that shows as hardware accelerated in the Dashboard while CPU load climbs anyway.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Know which formats force it:&lt;/strong&gt; PGS from Blu-ray sources and VOBSUB from DVD sources are bitmap tracks and always burn in. Styled ASS and SSA tracks often burn in too, because their positioning and effects cannot survive being passed as plain text.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;SRT is the format that costs nothing:&lt;/strong&gt; a plain text track is handed to the client and rendered there, so the video can direct play untouched. A sidecar file named &lt;code&gt;Movie.2019.en.srt&lt;/code&gt; beside the video is the simplest way to get one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Set the client policy deliberately:&lt;/strong&gt; in Settings, Player, the "Burn Subtitles" option accepts Never, Only image formats and Always. Leaving it on Always converts every text track into an expensive transcode for no benefit.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Convert once instead of burning nightly:&lt;/strong&gt; running your PGS tracks through OCR in a tool such as Subtitle Edit produces SRT files that serve every future playback. One conversion replaces hundreds of transcodes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Forced tracks are the hidden repeat offender:&lt;/strong&gt; files with forced foreign dialogue subtitles trigger burn in on titles you assumed were direct playing, so audit your most watched foreign language content first.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Budget one extra transcode unit for every stream that burns subtitles, then work to eliminate them.&lt;/p&gt;




&lt;h2&gt;
  
  
  How much RAM and what scratch disk does the Plex transcoder need?
&lt;/h2&gt;

&lt;p&gt;RAM is rarely the constraint. Plex Media Server itself is modest, and a transcode holds working buffers rather than the whole file, so the memory question is really about how much you can spare for a RAM backed scratch directory. The scratch disk is where careless builds get punished, because every transcoding session writes segment files continuously for as long as someone is watching.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Start at 8 GB and choose 16 GB if you use tmpfs:&lt;/strong&gt; 8 GB runs a small server comfortably, and the second 8 GB exists so you can hand several gigabytes to a RAM disk without starving the operating system's file cache.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Move the transcode directory off the media array:&lt;/strong&gt; set Settings, Transcoder, "Transcoder temporary directory" to an SSD path or to &lt;code&gt;/dev/shm&lt;/code&gt; on Linux. Leaving it on a spinning array makes disks seek between reading source files and writing segments while other users are browsing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Understand what the buffer setting costs you:&lt;/strong&gt; "Transcoder default throttle buffer" defaults to 60 seconds, so each session writes roughly a minute of encoded video ahead of the viewer. Multiply that by your peak transcode count to size the directory.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Watch SSD endurance, not just speed:&lt;/strong&gt; transcoding is a sustained write workload on a drive that is otherwise idle, which is a real argument for tmpfs on a machine with spare RAM.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Put the library database on an SSD regardless:&lt;/strong&gt; &lt;code&gt;com.plexapp.plugins.library.db&lt;/code&gt; drives every browse, scan and search, and it is the difference between a client that loads instantly and one that stalls on the home screen.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Size RAM for the scratch directory you intend to use, then stop.&lt;/p&gt;




&lt;h2&gt;
  
  
  How much upload bandwidth do 3, 5 and 10 remote streams need?
&lt;/h2&gt;

&lt;p&gt;Transcoding capacity is worthless if your upstream cannot carry the result. Remote playback is limited by your upload speed, and most domestic connections are asymmetric, so a household with 500 Mbps down might have 40 Mbps up. Size against Plex's own quality presets, which are the bitrates your viewers will actually select, and add roughly 20 percent of headroom for protocol overhead and anything else using the line.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Concurrent remote streams&lt;/th&gt;
&lt;th&gt;Upload required at the 1080p 8 Mbps preset&lt;/th&gt;
&lt;th&gt;Practical note&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;3 streams&lt;/td&gt;
&lt;td&gt;Around 24 Mbps plus overhead, so 30 Mbps&lt;/td&gt;
&lt;td&gt;Within reach of most fibre and some cable uploads&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5 streams&lt;/td&gt;
&lt;td&gt;Around 40 Mbps plus overhead, so 48 Mbps&lt;/td&gt;
&lt;td&gt;Rules out most VDSL and older cable connections&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;10 streams&lt;/td&gt;
&lt;td&gt;Around 80 Mbps plus overhead, so 96 Mbps&lt;/td&gt;
&lt;td&gt;Symmetric fibre or a hosted server territory&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3 streams at the 20 Mbps 1080p preset&lt;/td&gt;
&lt;td&gt;Around 60 Mbps plus overhead, so 72 Mbps&lt;/td&gt;
&lt;td&gt;Quality preset matters more than stream count&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1 stream of an untranscoded 4K remux&lt;/td&gt;
&lt;td&gt;Frequently 50 Mbps or more, depending on the file&lt;/td&gt;
&lt;td&gt;Direct play saves CPU and spends bandwidth instead&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two controls keep this predictable. Settings, Server, Remote Access carries a "Limit remote stream bitrate" option that caps what any single remote viewer can pull, which protects the connection at the cost of forcing more transcodes. Separately, if direct connection fails and sessions fall back to Plex Relay, throughput is capped at 1 Mbps, or 2 Mbps with Plex Pass, which is why a working port forward or reverse proxy matters more than any encoder.&lt;/p&gt;




&lt;h2&gt;
  
  
  What Plex Pass unlocks, and why it changes the hardware maths
&lt;/h2&gt;

&lt;p&gt;Hardware transcoding is a Plex Pass feature. That single fact reframes the whole build, because the choice is not "iGPU or CPU", it is "pay for a subscription and buy a small machine" against "pay nothing recurring and buy a large one". For a reader who has just cancelled a streaming service, the second option often costs more in the first year and far more in electricity across three.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Hardware accelerated encode and decode:&lt;/strong&gt; without a Pass, Plex ignores Quick Sync, NVENC and VideoToolbox entirely, no matter what silicon you installed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Subscription shape, not a single price:&lt;/strong&gt; Plex Pass is sold as a monthly plan, an annual plan and a one time lifetime purchase, so the comparison against a bigger CPU is a three year total, not a monthly line item.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Server side background work:&lt;/strong&gt; intro and credit detection, chapter thumbnails and loudness analysis run as scheduled tasks on your server. They consume CPU outside playback hours, which is an argument for keeping some general purpose headroom even on a Quick Sync build.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mobile downloads change your peak:&lt;/strong&gt; synced offline copies are transcoded once, at full speed rather than in real time, so a single sync job can briefly load the machine harder than three live streams.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The Pass follows the account, not the machine:&lt;/strong&gt; you can move your server between a NAS, a self built box, a VPS or Yundera without repurchasing, so the hardware decision stays reversible.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Decide the Pass question first. Every sizing number in this article changes depending on the answer.&lt;/p&gt;




&lt;h2&gt;
  
  
  Three reference builds sized for 3, 5 and 10 concurrent transcodes
&lt;/h2&gt;

&lt;p&gt;These three builds assume Plex Pass, 1080p transcodes with SDR sources, and the scratch directory already moved off the media array. Adjust upward if your peak sentence from earlier includes 4K HDR or burned in subtitles.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Target load&lt;/th&gt;
&lt;th&gt;Core hardware&lt;/th&gt;
&lt;th&gt;Why this is the cut off&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;3 concurrent 1080p transcodes&lt;/td&gt;
&lt;td&gt;Intel N100 or N305 mini PC, 16 GB RAM, 500 GB NVMe for the database and scratch&lt;/td&gt;
&lt;td&gt;A twelfth generation media engine covers this load while the whole machine draws less power than a single desktop GPU at idle&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5 concurrent 1080p transcodes&lt;/td&gt;
&lt;td&gt;Intel Core i3 or i5 desktop of a recent generation, non F model, 16 GB RAM, separate NVMe&lt;/td&gt;
&lt;td&gt;Same engine class in a chassis with room for drives and a 3.5 inch bay backplane&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;10 concurrent 1080p transcodes&lt;/td&gt;
&lt;td&gt;Recent Intel Core i5 or i7, non F model, 32 GB RAM, NVMe scratch, symmetric fibre upstream&lt;/td&gt;
&lt;td&gt;Bandwidth and library database contention arrive before the encoder saturates&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3 concurrent 4K HDR transcodes&lt;/td&gt;
&lt;td&gt;Recent Intel Core i5 with Arc based graphics, or an NVIDIA RTX card in a tower&lt;/td&gt;
&lt;td&gt;Tone mapping at 4K is the workload that justifies discrete silicon&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;No Plex Pass, any count&lt;/td&gt;
&lt;td&gt;High core count desktop or used workstation CPU, sized by PassMark&lt;/td&gt;
&lt;td&gt;Software encoding scales by cores, so the chassis, cooling and power bill all grow&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two notes apply to every row. Buy the storage array separately from the compute decision, because drives do not affect transcoding. And validate with &lt;code&gt;intel_gpu_top&lt;/code&gt; or the Plex Dashboard under real load before adding hardware you assumed you needed.&lt;/p&gt;

</description>
      <category>plex</category>
      <category>selfhosted</category>
      <category>homelab</category>
      <category>performance</category>
    </item>
    <item>
      <title>Migrating to Hubs Community Edition: what survives the move, and what you rebuild by hand</title>
      <dc:creator>John</dc:creator>
      <pubDate>Thu, 03 Sep 2026 07:07:11 +0000</pubDate>
      <link>https://dev.to/john_182319291/migrating-to-hubs-community-edition-what-survives-the-move-and-what-you-rebuild-by-hand-1k1l</link>
      <guid>https://dev.to/john_182319291/migrating-to-hubs-community-edition-what-survives-the-move-and-what-you-rebuild-by-hand-1k1l</guid>
      <description>&lt;p&gt;Your files survive the move to Hubs Community Edition. Your rooms do not. Scene GLBs, avatar GLBs and uploaded media are portable binary assets that carry across cleanly, but the database rows that turn those assets into a working room, the short room IDs, the permalinks, the accounts, the per room permissions and the moderation settings are all Reticulum state that a fresh deployment generates from scratch. Plan the migration as an asset transfer plus a manual rebuild of everything that was a link, a login or a setting.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TL;DR by reader profile&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Agency hosting client spaces&lt;/strong&gt; (a five person studio running 12 branded client showcases): migrate the scenes, rebuild the rooms, and budget for reissuing every client facing URL, because the old links are the single largest source of post migration support tickets.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Education team&lt;/strong&gt; (a university department with 40 seminar rooms and one shared campus scene): migrate the handful of scenes that matter and re-create rooms per term, since most seminar rooms were disposable and rebuilding them is faster than reconstructing their state.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Events company&lt;/strong&gt; (a three person team that spins up a venue scene per event): migrate scene sources and avatars, then invest the saved time in the Dialog SFU and TURN setup, because live audio quality decides whether the event works at all.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Solo creator or community organiser&lt;/strong&gt; (one person running a monthly meetup room): expect a single evening of work, mostly re-uploading avatars and re-creating one room, and accept that your bookmarked room link is gone for good.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Enterprise or public sector team&lt;/strong&gt; (an internal comms group with retention obligations): treat the migration as a new system build with its own data map, because self hosting moves recordings, uploads and access logs onto infrastructure you now have to account for.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The tradeoff is simple: binary assets move almost for free, while identity, links and configuration cost you manual work that scales with the number of rooms you kept, not the number of gigabytes you stored.&lt;/p&gt;




&lt;h2&gt;
  
  
  Table of contents
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;What survives the move to Hubs Community Edition, and what does not?&lt;/li&gt;
&lt;li&gt;How Reticulum stores rooms, scenes, avatars and media&lt;/li&gt;
&lt;li&gt;Do your existing room links still work after the move?&lt;/li&gt;
&lt;li&gt;Which parts of a Spoke scene transfer, and which need republishing?&lt;/li&gt;
&lt;li&gt;Custom avatars: what carries over and what you re-upload&lt;/li&gt;
&lt;li&gt;Uploaded files, orphaned assets and the storage directory&lt;/li&gt;
&lt;li&gt;Accounts, permissions and moderation settings you rebuild by hand&lt;/li&gt;
&lt;li&gt;Which third party integrations stop working until you supply your own keys?&lt;/li&gt;
&lt;li&gt;How long does a 50 room migration take, and how do you verify it?&lt;/li&gt;
&lt;li&gt;Where do you run Hubs Community Edition, and what does the stack need?&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  What survives the move to Hubs Community Edition, and what does not?
&lt;/h2&gt;

&lt;p&gt;Split your inventory into two piles before you touch anything. Binary assets survive because they are just files on disk: a published Spoke scene is a glTF 2.0 binary, a custom avatar is a &lt;code&gt;.glb&lt;/code&gt; with a small set of texture overrides, and an uploaded image, PDF or video is the original file plus a generated thumbnail. Everything else is Reticulum state, which means rows in a Postgres database that a new deployment creates fresh and never inherits.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Item&lt;/th&gt;
&lt;th&gt;What survives the move&lt;/th&gt;
&lt;th&gt;What you rebuild by hand&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Published scenes&lt;/td&gt;
&lt;td&gt;The scene GLB and its textures, byte for byte&lt;/td&gt;
&lt;td&gt;The scene entry, its listing, its thumbnail and its owner&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Spoke projects&lt;/td&gt;
&lt;td&gt;The project source and any assets you exported&lt;/td&gt;
&lt;td&gt;The project entry in the editor, plus a republish per scene&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Custom avatars&lt;/td&gt;
&lt;td&gt;The base GLB and override maps&lt;/td&gt;
&lt;td&gt;The avatar record, its name and its assignment to an account&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Room media&lt;/td&gt;
&lt;td&gt;The uploaded originals in the storage directory&lt;/td&gt;
&lt;td&gt;The link between each file and the room it was pinned in&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rooms&lt;/td&gt;
&lt;td&gt;Nothing automatic&lt;/td&gt;
&lt;td&gt;Room name, scene binding, occupancy limit, permissions, pinned objects&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Accounts and links&lt;/td&gt;
&lt;td&gt;Nothing automatic&lt;/td&gt;
&lt;td&gt;Every account, every role, and every short room ID in a shared URL&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The practical consequence is that migration effort tracks your room count, not your storage size. Twelve rooms sharing one 80 MB scene is a longer job than one room referencing 4 GB of uploads, because the 4 GB copies in a single &lt;code&gt;rsync&lt;/code&gt; and the twelve rooms do not.&lt;/p&gt;




&lt;h2&gt;
  
  
  How Reticulum stores rooms, scenes, avatars and media
&lt;/h2&gt;

&lt;p&gt;Reticulum, the Elixir backend behind Hubs, splits your world into two stores that fail independently. Postgres holds the meaning, and a storage directory on disk holds the bytes. A migration that copies one without the other produces either empty rooms or unreachable files, so map both before you plan anything.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Rooms live entirely in Postgres&lt;/strong&gt;: a hub row carries the short room ID used in the URL, the display name, the maximum occupancy, the member permissions bitfield and the binding to a scene. Nothing about a room exists as a file you can copy.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Scenes and avatars are rows that point at files&lt;/strong&gt;: the scene or avatar record holds the name, description, attribution and its own short ID, then references the GLB and thumbnail through file records. Copy the GLB without the row and Hubs has no way to list it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Uploaded media are owned file records&lt;/strong&gt;: each upload gets an ID, a content type and an access key that the delivery URL must present, so the raw file on disk is not addressable unless the matching row exists in the database.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Storage separates temporary from permanent&lt;/strong&gt;: files dropped into a room start in an expiring area and are promoted to permanent storage when someone pins them, which is why unpinned uploads disappear from a copy taken later.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pinned objects are room state, not files&lt;/strong&gt;: position, rotation, scale and the object reference sit in their own rows, so a pinned image restores as a floating object only if both the file and that row survive.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Take &lt;code&gt;pg_dump&lt;/code&gt; and &lt;code&gt;rsync&lt;/code&gt; of the storage path together, from the same moment.&lt;/p&gt;




&lt;h2&gt;
  
  
  Do your existing room links still work after the move?
&lt;/h2&gt;

&lt;p&gt;No. A Hubs room URL is a hostname plus a short room ID plus a decorative slug, and only the first two parts matter. The slug is ignored on resolution, which is why renaming a room never broke its link, but the hostname is now yours and the short ID is regenerated by your own Reticulum instance. Every link you ever pasted into a calendar invite, a client deck or a printed programme resolves to nothing you control.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The hostname is unrecoverable&lt;/strong&gt;: you cannot issue a 301 redirect from a domain you never owned, so any bookmark pointing at the old service is dead and no DNS change on your side fixes it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Short IDs are generated, not carried&lt;/strong&gt;: creating a room on your instance produces a new ID, so the same scene republished twice gives you two different room addresses with identical content.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You can force the ID if you insert it&lt;/strong&gt;: seeding the hub row with a chosen short ID makes the path component match the old link exactly, which turns a broken URL into one that only needs the domain swapped.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Six digit entry codes are per deployment&lt;/strong&gt;: the short code flow people used to jump from a headset browser into a room only knows about rooms in your database, so every previously shared code is invalid.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Embeds break silently&lt;/strong&gt;: an iframe embedded in a client site keeps rendering its container and fails inside, so nobody reports it until a meeting starts.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Budget one communication task per audience, not per room. Twelve client rooms usually means three emails, not twelve.&lt;/p&gt;




&lt;h2&gt;
  
  
  Which parts of a Spoke scene transfer, and which need republishing?
&lt;/h2&gt;

&lt;p&gt;Treat the published scene and the editable project as two separate things, because only one of them is self contained. The published output is a GLB carrying Hubs specific behaviour in the &lt;code&gt;MOZ_hubs_components&lt;/code&gt; glTF extension, so spawn points, waypoints, audio zones, mirrors and the generated navmesh all travel inside the file. The project is a JSON document full of absolute URLs, and those URLs pointed at an asset host that no longer answers.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The published GLB renders immediately&lt;/strong&gt;: drop it into your own instance, bind a room to it, and geometry, materials, baked lighting settings, collision and navigation behave exactly as before, with no re-authoring.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The project file opens with holes&lt;/strong&gt;: your &lt;code&gt;.spoke&lt;/code&gt; project references every imported model, texture, image and audio clip by URL, so reopening it in your own Spoke gives you a scene tree with unresolved assets rather than a working edit.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Third party imports are the worst offenders&lt;/strong&gt;: anything pulled in from Sketchfab or a media search result was referenced rather than embedded, so those nodes need re-importing one by one against your own configured providers.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Scene metadata does not ride along&lt;/strong&gt;: name, description, attribution credits and the scene thumbnail are database fields, so a transferred GLB appears as an untitled entry until you set them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Republishing regenerates derived data&lt;/strong&gt;: a fresh publish rebuilds the navmesh and the thumbnail, which is why a scene you can only supply as a GLB stays frozen at its last published state.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Migrate GLBs for continuity, and re-import projects only for the 2 or 3 scenes you still expect to edit.&lt;/p&gt;




&lt;h2&gt;
  
  
  Custom avatars: what carries over and what you re-upload
&lt;/h2&gt;

&lt;p&gt;Avatars split the same way scenes do, and the derived ones cause the trouble. A fully custom avatar is a single GLB you can copy anywhere. A derived avatar is not a model at all: it is a reference to a parent avatar plus texture overrides, typically a base map, an emissive map, a normal map and an ORM map. Move the child without the parent and it resolves to nothing.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Standalone avatar GLBs transfer intact&lt;/strong&gt;: rig, skin weights, blend shapes and materials all survive, so a bespoke client mascot loads on your instance the moment you re-upload it and give it a name.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Derived avatars need their parent first&lt;/strong&gt;: upload the base avatar, then the overrides, in that order, or the child record points at a parent ID your database has never seen.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The default avatar set is not your data&lt;/strong&gt;: the stock avatars ship with the client and a fresh deployment starts with its own listings, so any curated in house set has to be re-created as avatar listings in the admin panel.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Nobody keeps their selection&lt;/strong&gt;: avatar choice is stored against an account, and since accounts do not migrate, every returning user lands on a default avatar and picks again on first entry.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ready Player Me links keep working&lt;/strong&gt;: avatars pulled in by URL are fetched at use time rather than stored as your records, so that path survives the move untouched.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For a team of 20 with one branded avatar each, expect the uploads to take under an hour and the re-selection support requests to trickle in for a fortnight.&lt;/p&gt;




&lt;h2&gt;
  
  
  Uploaded files, orphaned assets and the storage directory
&lt;/h2&gt;

&lt;p&gt;Open the storage directory and you will find no filenames. Files are written under opaque identifiers with no extension, so a 4 MB blob could be a client logo, a PDF or a thumbnail, and only the matching database row tells you which. This is why a partial migration is so hard to debug: the disk looks healthy, the rooms look empty, and nothing in the file tree explains the gap.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Reconcile counts before you trust the copy&lt;/strong&gt;: run &lt;code&gt;find /storage -type f | wc -l&lt;/code&gt; on both sides and compare it with the number of file rows in Postgres, because a mismatch found now is a five minute fix and a mismatch found later is an archaeology project.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Orphans come in two directions&lt;/strong&gt;: files with no row are dead weight you can delete, while rows with no file are worse, since Hubs renders the object placeholder and then fails to load it in front of a client.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Derivatives inflate your storage estimate&lt;/strong&gt;: every image carries a generated thumbnail alongside the original, so the directory size is not the sum of what people uploaded and copying only the originals leaves visible gaps in room listings.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Large video is served as it was uploaded&lt;/strong&gt;: without a transcoding step in your stack, a 500 MB source file is what every visitor streams, so audit the biggest 20 files before you size bandwidth.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Permissions travel with the copy&lt;/strong&gt;: use &lt;code&gt;rsync -a&lt;/code&gt; and keep ownership consistent with the user your Reticulum container runs as, or uploads succeed and reads fail.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Copy files first, restore the database second, then verify one room per scene.&lt;/p&gt;




&lt;h2&gt;
  
  
  Accounts, permissions and moderation settings you rebuild by hand
&lt;/h2&gt;

&lt;p&gt;There is one piece of good news here: Hubs never stored passwords. Login is a magic link sent to an email address, so nobody needs a reset campaign and you are not migrating credentials. What you are migrating is the absence of everything else, because an account row carries the identity, the admin flag and the ownership of every scene and avatar that person created.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Every account starts again&lt;/strong&gt;: a returning user enters the same email address on your domain and receives a brand new account, with no history, no owned scenes and no saved avatar.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ownership does not follow the asset&lt;/strong&gt;: a scene you re-upload as an admin is owned by the admin account, so the designer who built it can no longer edit or republish it without you reassigning the record.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The first admin comes from configuration, not the database&lt;/strong&gt;: your deployment nominates an admin email before anyone logs in, and that is the only account that can reach the admin panel to promote the next one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Room permissions are per room defaults you set again&lt;/strong&gt;: the toggles governing who may pin objects, spawn media, draw, share a camera or fly are stored on the hub row, so a locked down client showcase reverts to the default permissive setup unless you reapply it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Moderation state is gone&lt;/strong&gt;: closed rooms, kicked users and any ban you applied were tied to identities on the old deployment and have no meaning against new account IDs.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For 20 users across 12 rooms, the account side is self service and costs you nothing, while the permission side is roughly 12 short admin tasks.&lt;/p&gt;




&lt;h2&gt;
  
  
  Which third party integrations stop working until you supply your own keys?
&lt;/h2&gt;

&lt;p&gt;The hosted service quietly held API credentials on your behalf. Your deployment holds none, so several features that felt native to Hubs become inert until you register for accounts and add keys to your configuration. One of them is not optional: without working outbound email, nobody can log in at all, because the magic link never arrives.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Integration&lt;/th&gt;
&lt;th&gt;State on a fresh deployment&lt;/th&gt;
&lt;th&gt;What it needs from you&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Magic link login&lt;/td&gt;
&lt;td&gt;Broken, blocks all sign in&lt;/td&gt;
&lt;td&gt;SMTP credentials or a transactional mail provider, plus SPF and DKIM records&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sketchfab import&lt;/td&gt;
&lt;td&gt;Search returns nothing&lt;/td&gt;
&lt;td&gt;A Sketchfab API token in your Reticulum configuration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GIF and image search&lt;/td&gt;
&lt;td&gt;Empty results panel&lt;/td&gt;
&lt;td&gt;A key per provider, one for GIFs and one for image search&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Video links from streaming sites&lt;/td&gt;
&lt;td&gt;Paste works, playback varies&lt;/td&gt;
&lt;td&gt;A resolver service reachable from your instance, checked per site&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Discord room binding&lt;/td&gt;
&lt;td&gt;Absent&lt;/td&gt;
&lt;td&gt;The separate Hubs Discord bot deployed with its own bot token&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Third party media embeds&lt;/td&gt;
&lt;td&gt;Blocked by the browser&lt;/td&gt;
&lt;td&gt;A working CORS proxy host, included in your content security policy&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two consequences matter for a migration plan. First, test the email path before you announce anything, because every other check depends on being able to sign in. Second, audit your existing scenes for objects that came from a search panel, since those nodes will render as broken references until the matching provider key is in place. A scene with 30 imported models can hide 3 or 4 of these, and they only surface once someone walks up to them.&lt;/p&gt;




&lt;h2&gt;
  
  
  How long does a 50 room migration take, and how do you verify it?
&lt;/h2&gt;

&lt;p&gt;Do not estimate the whole job. Time one room end to end, then multiply, because the per room work is almost perfectly repetitive. Create the room, bind the scene, set the permissions, re-pin the objects, capture the new URL. If that loop takes you 4 minutes, 50 rooms is a little over 3 hours of clicking, and the asset copy underneath it happens once regardless of room count.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Front load the one time work&lt;/strong&gt;: the file copy, the database restore, the email configuration and the provider keys are a single block of effort, and none of it repeats when you add the fiftieth room.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Batch by scene, not by room&lt;/strong&gt;: rooms sharing one scene need that scene republished once, so ordering your list by scene turns 50 tasks into as few as 8 or 10 real pieces of work.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Verify with two browser windows&lt;/strong&gt;: join the same room twice on different profiles, confirm you hear yourself on the second connection, and confirm both avatars move, which tests the SFU and the state channel together.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Watch the network panel on entry&lt;/strong&gt;: a room that looks correct can still be firing failed requests for pinned media, so treat anything other than zero failed asset loads as unfinished.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Test one room on an actual headset&lt;/strong&gt;: desktop entry passes long before the headset path does, and problems here are usually certificate or host related rather than anything to do with your migrated content.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Reserve a quarter of your total budget for the last 3 rooms. The stragglers are always the ones with imported models and pinned video.&lt;/p&gt;




&lt;h2&gt;
  
  
  Where do you run Hubs Community Edition, and what does the stack need?
&lt;/h2&gt;

&lt;p&gt;Hubs is not one container. A working deployment runs the Reticulum backend, the web client, Spoke, the Dialog SFU that carries voice and video, a TURN server for clients behind restrictive networks, Postgres and a storage volume. That shape drives the hosting decision more than CPU or RAM does, because real time media has requirements a typical web app never raises.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;You need several hostnames, not one&lt;/strong&gt;: the app, the asset host, the media server and the CORS proxy each answer on their own name, all on HTTPS, and your content security policy has to list them consistently or the client blocks its own assets.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;UDP is the hard requirement&lt;/strong&gt;: WebRTC media flows over a UDP port range to the SFU, and any host that only exposes TCP 443 forces every participant onto TURN relaying, which changes your bandwidth profile completely.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;TURN is not optional in practice&lt;/strong&gt;: corporate networks routinely block direct paths, so plan for a relay listening on the standard 3478 and 5349 ports even if your own testing never touches it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Postgres and storage are the state&lt;/strong&gt;: everything else can be recreated from images, so your backup job only has to cover the database and the storage directory.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Your options span the usual ground&lt;/strong&gt;: a self managed VPS, a home server, a NAS or a managed personal server. Yundera is a managed Personal Cloud Server, built on CasaOS, that runs self-hosted apps as Docker containers on a server dedicated to the user.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Choose the host on its networking behaviour first. A machine that cannot pass UDP will run every meeting through a relay.&lt;/p&gt;

</description>
      <category>webxr</category>
      <category>selfhosted</category>
      <category>devops</category>
    </item>
    <item>
      <title>Ollama Has No API Authentication: How To Properly Gate Port 11434</title>
      <dc:creator>John</dc:creator>
      <pubDate>Tue, 01 Sep 2026 07:06:06 +0000</pubDate>
      <link>https://dev.to/john_182319291/ollama-has-no-api-authentication-how-to-properly-gate-port-11434-48e3</link>
      <guid>https://dev.to/john_182319291/ollama-has-no-api-authentication-how-to-properly-gate-port-11434-48e3</guid>
      <description>&lt;p&gt;Ollama has no username, no password and no API key. Anything that can reach TCP port 11434 can list your models, run inference on your GPU, pull a 40 GB model onto your disk and delete every model you have, with a single curl command and no credential. The fix is not a setting inside Ollama, because there is not one: you keep the default bind on 127.0.0.1, and if you genuinely need remote access you put a bearer token proxy, an SSH tunnel or a mesh VPN in front of it. Everything else in this article is choosing which of those three, and proving you got it right.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TL;DR by reader profile:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The single laptop user&lt;/strong&gt; (Maya, a writer running llama3.2 on a MacBook for offline drafting): keep the default 127.0.0.1 bind and never touch OLLAMA_HOST, because nothing outside the laptop has any reason to reach the API.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The home GPU box owner&lt;/strong&gt; (Tom, one desktop with an RTX card in the study, phone and tablet elsewhere in the house): use a mesh VPN such as Tailscale or WireGuard, because it gives you remote access without any port ever being open to the network or the router.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The laptop plus workstation developer&lt;/strong&gt; (Priya, coding on a thin laptop against a beefy machine two rooms away): use an SSH tunnel, because it needs zero changes on the Ollama side and dies the moment you close the terminal.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The self-hoster running a web chat UI&lt;/strong&gt; (Ben, Open WebUI on a rented VPS with a public domain): terminate HTTPS at a reverse proxy that checks a bearer token, and never publish 11434 itself, because the browser tier and the model tier need separate trust.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The small team sharing one model server&lt;/strong&gt; (a six person studio pooling one GPU): use a proxy with per user tokens and rate limits, or mutual TLS, because you need to revoke one person without rotating everyone.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Anyone who already set OLLAMA_HOST=0.0.0.0&lt;/strong&gt; (following a blog post to make a client work): stop and audit before hardening, because an exposed endpoint may already have models on disk that you did not pull.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The real tradeoff is this: every method that makes a new client easy to connect also widens the set of machines that can reach your models, and you cannot optimise both at once.&lt;/p&gt;




&lt;h2&gt;
  
  
  Table of contents
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;What is actually listening on port 11434 when Ollama runs?&lt;/li&gt;
&lt;li&gt;What can somebody do with an unauthenticated Ollama API?&lt;/li&gt;
&lt;li&gt;Is binding to 127.0.0.1 enough, and when does that guarantee break?&lt;/li&gt;
&lt;li&gt;How do I check right now whether my Ollama is reachable from outside?&lt;/li&gt;
&lt;li&gt;How do I set OLLAMA_HOST correctly on Linux, macOS and Windows?&lt;/li&gt;
&lt;li&gt;Does OLLAMA_ORIGINS stop a website from talking to my local models?&lt;/li&gt;
&lt;li&gt;Firewall rules or bind address: which layer should you rely on?&lt;/li&gt;
&lt;li&gt;How do I put a bearer token in front of Ollama with a reverse proxy?&lt;/li&gt;
&lt;li&gt;When is mutual TLS worth the extra work?&lt;/li&gt;
&lt;li&gt;SSH tunnels and mesh VPNs: reaching Ollama without opening a port&lt;/li&gt;
&lt;li&gt;How much latency does a proxy add to streamed tokens?&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  What is actually listening on port 11434 when Ollama runs?
&lt;/h2&gt;

&lt;p&gt;When you install Ollama, you get two things: a command line client and a long running HTTP server. The server is the part that matters here. It listens on TCP port 11434 and speaks plain HTTP, not HTTPS, and it answers every request it receives without asking who sent it.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A REST API, not a chat program:&lt;/strong&gt; the &lt;code&gt;ollama run llama3.2&lt;/code&gt; command you type is a thin client that talks to the same HTTP endpoint any other machine would use, so anything curl can send, a stranger can send.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A background service that starts itself:&lt;/strong&gt; on Linux it is usually a systemd unit called &lt;code&gt;ollama.service&lt;/code&gt;, on macOS it is the menu bar app, and on Windows it is a tray process, which means the port is open long after you close the terminal you launched it from.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Full lifecycle control, not just inference:&lt;/strong&gt; the API covers listing, pulling, creating, copying and deleting models, so it is an administrative interface as much as a generation interface.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Direct access to your model directory:&lt;/strong&gt; models live under &lt;code&gt;~/.ollama/models&lt;/code&gt; on Linux and macOS, and the API can add to or remove from that directory without touching the shell.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No credential layer of any kind:&lt;/strong&gt; there is no login, no API key setting and no per client identity, so the only access control you have is which network interfaces the port is bound to.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Where that port lives depends on where you run Ollama: a laptop, a home server, a NAS, a rented VPS, or a managed Personal Cloud Server. Yundera is a managed Personal Cloud Server, built on CasaOS, that runs self-hosted apps as Docker containers on a server dedicated to the user.&lt;/p&gt;




&lt;h2&gt;
  
  
  What can somebody do with an unauthenticated Ollama API?
&lt;/h2&gt;

&lt;p&gt;Treat reachability as full control. There is no read only mode, so anyone who can open a socket to the port holds the same powers you do from your own shell.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Inventory your models:&lt;/strong&gt; a single &lt;code&gt;curl http://your-host:11434/api/tags&lt;/code&gt; returns every model name, size and digest on the machine, which tells an intruder both what you use the box for and how much disk you have committed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Run inference on your hardware for free:&lt;/strong&gt; &lt;code&gt;POST /api/generate&lt;/code&gt; and &lt;code&gt;POST /api/chat&lt;/code&gt; execute on your GPU and your electricity bill, with no rate limit and no quota, and an abuser can point a public chat frontend at your endpoint and let strangers use it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fill your disk on demand:&lt;/strong&gt; &lt;code&gt;POST /api/pull&lt;/code&gt; downloads any model from the registry, and large models run to tens of gigabytes each, so a loop of pull requests is a straightforward way to exhaust storage until the host stops working.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Delete everything you have downloaded:&lt;/strong&gt; &lt;code&gt;DELETE /api/delete&lt;/code&gt; removes a model with one request and no confirmation prompt, so hours of downloads disappear without a trace in any interface you normally watch.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Read what you are currently doing:&lt;/strong&gt; &lt;code&gt;GET /api/ps&lt;/code&gt; lists models loaded in memory right now, which leaks activity patterns even when no prompt content is exposed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Create custom models on your machine:&lt;/strong&gt; &lt;code&gt;POST /api/create&lt;/code&gt; builds a new model entry from a Modelfile, including a system prompt of the attacker's choosing, so a client of yours can be silently served a modified model under a familiar name.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of this requires an exploit or an unpatched version. It is the documented API behaving exactly as designed.&lt;/p&gt;




&lt;h2&gt;
  
  
  Is binding to 127.0.0.1 enough, and when does that guarantee break?
&lt;/h2&gt;

&lt;p&gt;Yes, with one condition: a process bound to 127.0.0.1 is unreachable from any other machine, because the kernel refuses packets arriving on a physical interface for a loopback address. No firewall rule is doing that work. The bind itself is the boundary. The problem is how many ordinary situations quietly move the bind somewhere else.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;You set OLLAMA_HOST to make one client work:&lt;/strong&gt; the usual advice for connecting a phone, a second laptop or a container is &lt;code&gt;OLLAMA_HOST=0.0.0.0:11434&lt;/code&gt;, which replaces a boundary that was airtight with none at all, on every interface at once including any public one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Loopback is per machine, not per user:&lt;/strong&gt; every local account, every browser extension and every background process on that host can reach 127.0.0.1:11434, so on a shared workstation the bind protects you from the network and from nobody else.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Windows Subsystem for Linux crosses the line:&lt;/strong&gt; WSL2 runs on its own virtual network, so a server bound to loopback inside WSL and a server bound to loopback on Windows are two different boundaries, and people usually resolve the confusion by binding wide.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A VPN or mesh interface is still an interface:&lt;/strong&gt; if you bind to 0.0.0.0 expecting only your VPN subnet to reach it, the same socket answers on Ethernet and Wi Fi too.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;IPv6 has its own loopback:&lt;/strong&gt; binding to 127.0.0.1 does not cover &lt;code&gt;::1&lt;/code&gt;, and binding to &lt;code&gt;[::]&lt;/code&gt; covers far more than you probably intend.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The rule to keep: change the bind only when you have already decided what will check the requests.&lt;/p&gt;




&lt;h2&gt;
  
  
  How do I check right now whether my Ollama is reachable from outside?
&lt;/h2&gt;

&lt;p&gt;Do this before you change anything. Five minutes of checking tells you whether you are hardening a closed door or cleaning up after an open one.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Look at the bind address on the host itself:&lt;/strong&gt; run &lt;code&gt;ss -tlnp | grep 11434&lt;/code&gt; on Linux, &lt;code&gt;lsof -nP -i :11434&lt;/code&gt; on macOS, or &lt;code&gt;netstat -ano | findstr 11434&lt;/code&gt; on Windows, and read the left hand address, because &lt;code&gt;127.0.0.1:11434&lt;/code&gt; is safe and &lt;code&gt;0.0.0.0:11434&lt;/code&gt; or &lt;code&gt;[::]:11434&lt;/code&gt; means every interface is answering.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Test from a second device on the same network:&lt;/strong&gt; from a phone or another laptop, run &lt;code&gt;curl -m 5 http://192.168.1.50:11434/api/tags&lt;/code&gt; with your machine's local IP, and a JSON model list coming back proves the LAN can reach it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Test from outside your network:&lt;/strong&gt; turn Wi Fi off on your phone so it uses mobile data, then try the same request against your public IP, because that is the only test that distinguishes a LAN exposure from an internet exposure.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Check the router for a forwarding rule:&lt;/strong&gt; open the port forwarding or virtual server page and look for anything sending 11434 inward, since a rule added months ago for a different purpose can survive a router reboot and a firmware update.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Audit the model directory for things you did not pull:&lt;/strong&gt; list &lt;code&gt;~/.ollama/models&lt;/code&gt; and compare it against what you remember downloading, because an unfamiliar model or a sudden jump in disk usage is the clearest sign someone else has been using the API.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Write the results down. You will repeat this test after every change in the sections that follow.&lt;/p&gt;




&lt;h2&gt;
  
  
  How do I set OLLAMA_HOST correctly on Linux, macOS and Windows?
&lt;/h2&gt;

&lt;p&gt;The variable has to reach the background service, not your terminal. Exporting &lt;code&gt;OLLAMA_HOST&lt;/code&gt; in a shell changes where the client looks, not where the server listens, which is the single most common reason a change appears to do nothing.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Platform&lt;/th&gt;
&lt;th&gt;Where the value belongs&lt;/th&gt;
&lt;th&gt;What makes it take effect&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Linux with systemd&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;sudo systemctl edit ollama.service&lt;/code&gt;, then add &lt;code&gt;Environment="OLLAMA_HOST=127.0.0.1:11434"&lt;/code&gt; under a &lt;code&gt;[Service]&lt;/code&gt; block&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;sudo systemctl daemon-reload&lt;/code&gt; followed by &lt;code&gt;sudo systemctl restart ollama&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;macOS desktop app&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;launchctl setenv OLLAMA_HOST "127.0.0.1:11434"&lt;/code&gt; so the value is visible to apps launched by the session&lt;/td&gt;
&lt;td&gt;Quit Ollama from the menu bar and start it again, since the running process keeps its old environment&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Windows&lt;/td&gt;
&lt;td&gt;User environment variables in the System Properties dialog, one entry named &lt;code&gt;OLLAMA_HOST&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Exit the tray icon completely, then relaunch, because a minimised process is still the old one&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Docker container&lt;/td&gt;
&lt;td&gt;An &lt;code&gt;-e OLLAMA_HOST=...&lt;/code&gt; flag or an &lt;code&gt;environment:&lt;/code&gt; entry, paired with a publish flag such as &lt;code&gt;-p 127.0.0.1:11434:11434&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Recreate the container, since environment variables are fixed at creation&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;After any of these, repeat the two device test from the previous section rather than trusting the config file. The value is only correct if a request from another machine now fails.&lt;/p&gt;

&lt;p&gt;Where you run the server decides which row applies to you: a laptop and a self managed VPS use the platform rows, while a NAS, a home server running containers and a managed Personal Cloud Server such as Yundera all use the Docker row, because apps there run as Docker containers on a server dedicated to the user.&lt;/p&gt;




&lt;h2&gt;
  
  
  Does OLLAMA_ORIGINS stop a website from talking to my local models?
&lt;/h2&gt;

&lt;p&gt;Partly, and only inside a browser. &lt;code&gt;OLLAMA_ORIGINS&lt;/code&gt; sets which web origins get an &lt;code&gt;Access-Control-Allow-Origin&lt;/code&gt; header back. That is a rule the browser enforces on the page, not a rule the server enforces on the request. Understand the difference before you rely on it.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;It does block the ordinary case:&lt;/strong&gt; a random site you visit cannot script a &lt;code&gt;POST&lt;/code&gt; to &lt;code&gt;http://localhost:11434/api/chat&lt;/code&gt;, because a JSON content type triggers a preflight &lt;code&gt;OPTIONS&lt;/code&gt; request first, and a rejected preflight means the real request is never sent.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It does not block anything outside a browser:&lt;/strong&gt; curl, Python, a mobile app and a script on another machine ignore CORS entirely, so an origin allowlist has zero effect on the exposure risks covered earlier.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Simple requests still leave your machine:&lt;/strong&gt; a plain &lt;code&gt;GET&lt;/code&gt; to &lt;code&gt;/api/tags&lt;/code&gt; needs no preflight, so the browser sends it and only blocks the page from reading the answer, which is a weaker guarantee than most people assume.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The default is deliberately local:&lt;/strong&gt; Ollama already permits origins such as &lt;code&gt;http://localhost&lt;/code&gt; and &lt;code&gt;http://127.0.0.1&lt;/code&gt; on any port, plus browser extension origins, which is why a locally served web UI works with no configuration at all.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Setting it to &lt;code&gt;*&lt;/code&gt; removes the only browser side check you had:&lt;/strong&gt; the wildcard is the standard fix suggested when a self hosted frontend on a different host cannot connect, and it lets every page in every tab you open reach the API.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If a remote frontend needs access, name its exact origin, for example &lt;code&gt;OLLAMA_ORIGINS=https://chat.example.com&lt;/code&gt;, rather than opening it to everything.&lt;/p&gt;




&lt;h2&gt;
  
  
  Firewall rules or bind address: which layer should you rely on?
&lt;/h2&gt;

&lt;p&gt;Rely on the bind address as the primary control and treat the firewall as the backstop. A bind is a property of the socket, so it cannot be bypassed by a rule someone forgets to reapply. A firewall is a separate ruleset that has to be correct, loaded and evaluated before the traffic reaches the port.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;What it genuinely stops&lt;/th&gt;
&lt;th&gt;Where it lets you down&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Bind address (&lt;code&gt;OLLAMA_HOST=127.0.0.1:11434&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;Every packet from every other machine, on every interface, with no ruleset to maintain&lt;/td&gt;
&lt;td&gt;Nothing on the local host, and it is one environment variable away from being undone&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Host firewall (&lt;code&gt;sudo ufw deny 11434/tcp&lt;/code&gt;, or Windows Defender Firewall)&lt;/td&gt;
&lt;td&gt;Remote access even when the service is bound to &lt;code&gt;0.0.0.0&lt;/code&gt;, and it survives an accidental config change&lt;/td&gt;
&lt;td&gt;Docker published ports write their own forwarding rules, so a container can be reachable while &lt;code&gt;ufw status&lt;/code&gt; looks correct&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Router or cloud provider firewall&lt;/td&gt;
&lt;td&gt;Inbound traffic from the internet, including scans that find port 11434 within hours of exposure&lt;/td&gt;
&lt;td&gt;Nothing on your LAN, so every other device in the house or office still has full API access&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Bind plus host firewall together&lt;/td&gt;
&lt;td&gt;Both the remote path and a future misconfiguration of either one&lt;/td&gt;
&lt;td&gt;Neither layer identifies who is calling, so a permitted machine still has unlimited control&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Set the bind first, add the deny rule second, then repeat the external test. If your provider gives you a network level firewall, close 11434 there as well and leave it closed permanently.&lt;/p&gt;




&lt;h2&gt;
  
  
  How do I put a bearer token in front of Ollama with a reverse proxy?
&lt;/h2&gt;

&lt;p&gt;The pattern is always the same: Ollama stays bound to 127.0.0.1, a proxy on the same host owns the public port, and the proxy rejects anything without the right &lt;code&gt;Authorization&lt;/code&gt; header. Ollama itself is never modified, and it never learns that authentication exists.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Generate a token with real entropy:&lt;/strong&gt; &lt;code&gt;openssl rand -hex 32&lt;/code&gt; gives you 64 hexadecimal characters, which is far beyond guessable, and it costs nothing compared with a memorable passphrase that a dictionary attack will find.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Make the proxy compare, then forward:&lt;/strong&gt; in Caddy this is a matcher on &lt;code&gt;header Authorization "Bearer &amp;lt;token&amp;gt;"&lt;/code&gt; followed by &lt;code&gt;reverse_proxy 127.0.0.1:11434&lt;/code&gt;, and in nginx it is an &lt;code&gt;if&lt;/code&gt; test on &lt;code&gt;$http_authorization&lt;/code&gt; returning &lt;code&gt;401&lt;/code&gt; before the &lt;code&gt;proxy_pass&lt;/code&gt; line.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Terminate TLS at the proxy:&lt;/strong&gt; a bearer token over plain HTTP is readable by anything between the client and the server, so Caddy's automatic certificates or a certbot managed nginx config are part of the control, not an optional extra.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Turn off response buffering:&lt;/strong&gt; nginx needs &lt;code&gt;proxy_buffering off;&lt;/code&gt; and a long &lt;code&gt;proxy_read_timeout&lt;/code&gt;, otherwise streamed tokens arrive in one block at the end and a slow generation looks like a hung request.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Add a rate limit while you are there:&lt;/strong&gt; capping requests per IP turns a leaked token into a nuisance instead of unlimited free access to your GPU.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Check your clients can send headers first:&lt;/strong&gt; Open WebUI and most OpenAI compatible tools have a field for an API key, while the &lt;code&gt;ollama&lt;/code&gt; command line client sets only a host URL, so CLI users need a tunnel rather than a token.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Test with a deliberately wrong token and confirm you get a 401.&lt;/p&gt;




&lt;h2&gt;
  
  
  When is mutual TLS worth the extra work?
&lt;/h2&gt;

&lt;p&gt;Mutual TLS means the proxy in front of Ollama demands a client certificate before it forwards anything. A bearer token proves knowledge of a string. A client certificate proves possession of a private key that never travels over the wire. That is a real upgrade, and it costs you setup time on every device you own.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;It is worth it when you have several fixed devices:&lt;/strong&gt; issue one certificate per machine from your own certificate authority, and you can revoke the stolen laptop alone instead of rotating a shared token and reconfiguring everything else.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It is worth it when the endpoint must stay on the public internet:&lt;/strong&gt; an unauthenticated scanner gets a TLS handshake failure rather than an HTTP 401, so your endpoint never confirms that anything is listening behind it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It is worth it when tokens keep leaking into places you cannot audit:&lt;/strong&gt; shell history, container environment variables and editor config files all hold plain strings, while a private key can live in the operating system keystore.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It is not worth it for a single user on a single laptop:&lt;/strong&gt; the loopback bind already gives you a stronger boundary than certificates would, with no expiry to manage.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It is not worth it when your clients cannot present certificates:&lt;/strong&gt; many mobile apps and some desktop LLM frontends have no field for a client key, so you would end up running a token path beside the certificate path and inheriting the weaker of the two.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The practical cost is renewal. Certificates expire, commonly at 365 days, and a forgotten renewal takes every client offline at once with an error message that looks nothing like an authentication problem.&lt;/p&gt;




&lt;h2&gt;
  
  
  SSH tunnels and mesh VPNs: reaching Ollama without opening a port
&lt;/h2&gt;

&lt;p&gt;Both approaches leave the bind on 127.0.0.1 and move the network problem somewhere that already has authentication. Nothing about Ollama changes, and port 11434 stays closed to the world in both cases.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The SSH tunnel is the fastest to set up:&lt;/strong&gt; &lt;code&gt;ssh -N -L 11434:127.0.0.1:11434 user@gpu-box&lt;/code&gt; makes the remote API appear on your own loopback, so the &lt;code&gt;ollama&lt;/code&gt; command line client, Open WebUI and every default CORS origin work unchanged with no token anywhere.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The tunnel inherits SSH's authentication:&lt;/strong&gt; key based login means no new secret to manage, and closing the terminal closes the access, which is exactly what you want for occasional use and exactly what you do not want for a service that must stay up.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A mesh VPN suits always on access:&lt;/strong&gt; Tailscale or plain WireGuard give the host a private address, in Tailscale's case inside the 100.64.0.0/10 range, reachable only by devices you have enrolled.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The mesh handles the network problems for you:&lt;/strong&gt; it negotiates through NAT over UDP, WireGuard on port 51820 by default and Tailscale on 41641, so there is no router rule to add and no public IP requirement.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Access control moves to the mesh, not the API:&lt;/strong&gt; every enrolled device gets unrestricted use of the model server unless you write ACLs, so treat enrolment as the sensitive act and remove old phones and laptops when you stop using them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Phones and tablets are where the tunnel loses:&lt;/strong&gt; mobile SSH forwarding is awkward, while a VPN client on the phone makes the endpoint reachable from the sofa with no configuration in the app itself.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  How much latency does a proxy add to streamed tokens?
&lt;/h2&gt;

&lt;p&gt;Less than people fear, and in the wrong place. Token generation is bounded by your GPU or CPU, while a local proxy hop is a copy across loopback. What actually ruins the experience is buffering and connection setup, not throughput. Measure it on your own hardware with &lt;code&gt;curl -N -w "%{time_starttransfer}"&lt;/code&gt; and compare the direct call against the gated one.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Path to the API&lt;/th&gt;
&lt;th&gt;Where the extra delay comes from&lt;/th&gt;
&lt;th&gt;What keeps it small&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Direct to 127.0.0.1:11434&lt;/td&gt;
&lt;td&gt;Nothing beyond model load and generation, so this is your baseline number&lt;/td&gt;
&lt;td&gt;Keep a model resident so the first request does not pay the load cost again&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Local reverse proxy with a bearer token&lt;/td&gt;
&lt;td&gt;One extra loopback hop plus a TLS handshake on each new connection&lt;/td&gt;
&lt;td&gt;Enable HTTP keep alive and reuse connections instead of opening one per request&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SSH tunnel to another machine&lt;/td&gt;
&lt;td&gt;The round trip to that machine, plus SSH encryption on every chunk of the stream&lt;/td&gt;
&lt;td&gt;Run it over a wired LAN where possible, since the network round trip dominates everything else&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mesh VPN across the internet&lt;/td&gt;
&lt;td&gt;The full path to the remote host, and a relay hop if a direct connection cannot be negotiated&lt;/td&gt;
&lt;td&gt;Confirm the connection is direct rather than relayed, because relayed traffic takes a longer route&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The number that matters to a reader is time to first token, not tokens per second. Streaming hides steady state overhead well. It hides nothing about a slow start. If a gated setup feels sluggish, check buffering and connection reuse before you blame the encryption.&lt;/p&gt;

</description>
      <category>ollama</category>
      <category>security</category>
      <category>selfhosted</category>
      <category>llm</category>
    </item>
    <item>
      <title>Immich Sizing Guide: What a 20,000 Photo Google Takeout Import Really Costs in Disk, RAM and Hours</title>
      <dc:creator>John</dc:creator>
      <pubDate>Sat, 29 Aug 2026 07:05:56 +0000</pubDate>
      <link>https://dev.to/john_182319291/immich-sizing-guide-what-a-20000-photo-google-takeout-import-really-costs-in-disk-ram-and-hours-48ic</link>
      <guid>https://dev.to/john_182319291/immich-sizing-guide-what-a-20000-photo-google-takeout-import-really-costs-in-disk-ram-and-hours-48ic</guid>
      <description>&lt;p&gt;Budget roughly 1.35 to 1.6 times your raw photo size on disk, 6 GB of RAM for the full stack with machine learning enabled, 4 CPU cores, and a full weekend of background processing before a 20,000 item library is fully searchable. The import itself is the fast part. Thumbnail generation, video transcoding and machine learning jobs are what actually keep the server busy, and three settings decided in the first hour, the storage template, the machine learning model, and the transcoding policy, are the ones that force a full re-run of every job if you change your mind in month two.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TL;DR by reader profile:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The phone-only archivist, one iPhone and 20,000 photos with almost no video:&lt;/strong&gt; 2 CPU cores, 4 GB RAM and 1.4x your library size is enough, because thumbnails dominate and transcoding barely runs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The family with a decade of camcorder video, 20,000 items where 3,000 are clips:&lt;/strong&gt; go to 4 cores and 8 GB RAM before you import, because video transcoding is the single job that will run for days on weak hardware.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The RAW shooter with a Lightroom folder tree, 20,000 files averaging 25 MB:&lt;/strong&gt; use an external library in read-only mode rather than uploading, so Immich never becomes the only copy of your originals.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The privacy-first switcher leaving Google Photos this week:&lt;/strong&gt; set the storage template and the machine learning model before the first upload, then never touch them again, because both settings rewrite work already done.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The low-power host, a Raspberry Pi 5 or an N100 mini PC:&lt;/strong&gt; disable the smart search model on day one and run facial recognition only, or accept that machine learning jobs will still be queued a week later.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The person who just wants it working tonight:&lt;/strong&gt; import with default settings, leave transcoding on the optimal policy, and treat the first 72 hours as unattended processing time rather than a broken install.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The central tradeoff is that every setting which makes Immich feel fast and searchable later, larger machine learning models, generated previews, transcoded video, costs you disk and hours of CPU during the first week, and changing your mind afterwards means paying that cost a second time.&lt;/p&gt;




&lt;h2&gt;
  
  
  Table of contents
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;How much disk space does Immich actually need for 20,000 photos?&lt;/li&gt;
&lt;li&gt;Why is the Immich database and thumbnail folder bigger than you expected?&lt;/li&gt;
&lt;li&gt;How much RAM does Immich need, and what happens when it runs out?&lt;/li&gt;
&lt;li&gt;How many CPU cores does the first import really use?&lt;/li&gt;
&lt;li&gt;How long does a 20,000 item Google Takeout import take from start to searchable?&lt;/li&gt;
&lt;li&gt;Should you upload into Immich or point it at an external library?&lt;/li&gt;
&lt;li&gt;What does the Google Takeout format break, and how do you fix it before importing?&lt;/li&gt;
&lt;li&gt;Which storage template should you set before the first upload?&lt;/li&gt;
&lt;li&gt;Which machine learning model should you choose on day one?&lt;/li&gt;
&lt;li&gt;What should the video transcoding policy be, and what does it cost you?&lt;/li&gt;
&lt;li&gt;Where should you actually run Immich: NAS, mini PC, home server or hosted?&lt;/li&gt;
&lt;li&gt;Which day-one settings force a full re-run if you change them later?&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  How much disk space does Immich actually need for 20,000 photos?
&lt;/h2&gt;

&lt;p&gt;Start from the size of your Google Takeout, not the photo count. A 20,000 item library from a modern phone is usually somewhere between 60 GB and 120 GB of originals, because a 12 megapixel HEIC frame lands near 2 MB while a single 4K clip can pass 400 MB. Immich stores your originals untouched, then adds derived files on top.&lt;/p&gt;

&lt;p&gt;The multiplier you should plan for is 1.35x to 1.6x the original size, split across four directories under &lt;code&gt;UPLOAD_LOCATION&lt;/code&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;upload/&lt;/code&gt; and &lt;code&gt;library/&lt;/code&gt;, the originals:&lt;/strong&gt; 100 percent of your Takeout size, byte for byte, because Immich never recompresses the file you gave it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;thumbs/&lt;/code&gt;, the previews:&lt;/strong&gt; typically 15 to 25 percent on a photo heavy library, since every asset gets a small thumbnail plus a larger preview image used by the timeline and the detail view.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;encoded-video/&lt;/code&gt;, the transcodes:&lt;/strong&gt; zero if you have no video, but easily 50 to 100 percent of your video bytes again when the default policy re-encodes clips your browser cannot play natively.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The Postgres volume, metadata and machine learning vectors:&lt;/strong&gt; small in absolute terms, usually a few hundred megabytes at this scale, but it grows with embeddings and faces rather than with file size.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Practical floor: if your Takeout unzips to 90 GB, provision 200 GB and do not let the filesystem cross 80 percent during import. Running &lt;code&gt;du -sh&lt;/code&gt; on each subdirectory after the first 1,000 assets gives you a real multiplier for your own library, which beats any generic estimate.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why is the Immich database and thumbnail folder bigger than you expected?
&lt;/h2&gt;

&lt;p&gt;Because Immich does not generate one thumbnail per photo. It generates a small WebP tile for the timeline grid and a much larger preview used whenever you open an asset, and that second file is the one that surprises people. A 2 MB HEIC original can produce a preview of several hundred kilobytes, so the ratio between a phone photo and its derivatives is far worse than it is for a 25 MB RAW file, where the same preview is a rounding error.&lt;/p&gt;

&lt;p&gt;Four things inflate these directories beyond a naive estimate:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Preview resolution is a setting, not a constant:&lt;/strong&gt; the preview size configured in Administration, Settings, Image Settings applies to every asset, so raising it after import means every existing preview is regenerated and the old ones replaced.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Motion photos count twice:&lt;/strong&gt; an iPhone Live Photo or a Samsung motion shot arrives as a still plus a short video, and Immich stores and processes both, which quietly turns 20,000 selected items into more than 20,000 stored assets.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Postgres carries vectors, not just rows:&lt;/strong&gt; smart search stores an embedding per asset and facial recognition stores one per detected face, so a library with many group photos grows the database faster than a library of landscapes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Deleted assets linger for 30 days:&lt;/strong&gt; the trash retention default keeps originals and derivatives on disk until the period expires, so disk usage during a messy first week reflects your mistakes as well as your library.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Check the real split with &lt;code&gt;du -sh thumbs encoded-video&lt;/code&gt; inside &lt;code&gt;UPLOAD_LOCATION&lt;/code&gt; before you conclude anything about total size.&lt;/p&gt;




&lt;h2&gt;
  
  
  How much RAM does Immich need, and what happens when it runs out?
&lt;/h2&gt;

&lt;p&gt;Plan for 6 GB of system RAM for the full stack with machine learning enabled, and 4 GB if you disable smart search. The four containers are not equal: the server and Postgres are steady and modest, Redis is tiny, and &lt;code&gt;immich-machine-learning&lt;/code&gt; is the one that spikes, because it loads a model into memory the first time a job needs it and holds it there.&lt;/p&gt;

&lt;p&gt;What failure looks like is the part nobody warns you about. The kernel OOM killer does not stop Immich politely. It kills the machine learning container mid job, the job returns to the queue, the container restarts, reloads the model, and fails again on the same asset. You get a silent loop that burns CPU for hours and never advances the counter.&lt;/p&gt;

&lt;p&gt;Three signals tell you this is happening:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;docker compose ps&lt;/code&gt; shows a restart count climbing:&lt;/strong&gt; the machine learning container sits at 5 or 10 restarts while the server container sits at 0.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The job queue stalls at the same number:&lt;/strong&gt; smart search stays frozen while thumbnail generation keeps moving, because only one container is dying.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;dmesg | grep -i oom&lt;/code&gt; prints kill lines:&lt;/strong&gt; this is the definitive check, and it names the process that was terminated.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The fix is either more RAM, a smaller model, or fewer concurrent machine learning jobs, set to 1 in the job settings.&lt;/p&gt;

&lt;p&gt;Hosting choice matters here because you cannot add RAM to a box that has none free. A NAS with 4 GB shared with other services, a self-managed VPS, a mini PC and a managed option are all viable starting points. Yundera is a managed Personal Cloud Server, built on CasaOS, that runs self-hosted apps as Docker containers on a server dedicated to the user.&lt;/p&gt;




&lt;h2&gt;
  
  
  How many CPU cores does the first import really use?
&lt;/h2&gt;

&lt;p&gt;All of them, unless you tell Immich otherwise. Immich runs background work as named job queues, and each queue has its own concurrency setting in Administration, Settings, Job Settings. The defaults are tuned for a machine that is doing nothing else, so on a 4 core box the first import will pin every core for hours and make the web interface feel sluggish while you are still uploading.&lt;/p&gt;

&lt;p&gt;The queues do not weigh the same:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Thumbnail Generation:&lt;/strong&gt; the heaviest sustained consumer during a bulk import, because it decodes every original and writes two derivatives, and it scales almost linearly with the concurrency value you set.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Video Transcoding:&lt;/strong&gt; the one queue you should keep at concurrency 1, since FFmpeg already uses multiple threads internally and running two transcodes at once mostly makes both slower.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Smart Search and Face Detection:&lt;/strong&gt; CPU bound unless you have a supported GPU, and these are the queues that decide whether your library is searchable on day three or day seven.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Metadata Extraction:&lt;/strong&gt; cheap per asset, fast to complete, and rarely the bottleneck at 20,000 items.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Four cores is a workable floor. Two cores works but roughly doubles your wall clock time, and a shared vCPU with a burst credit balance will collapse to baseline speed partway through and stay there.&lt;/p&gt;

&lt;p&gt;Watch it live with &lt;code&gt;docker stats&lt;/code&gt;, which shows per container CPU percentage. If &lt;code&gt;immich-server&lt;/code&gt; sits near your core count times 100 percent for hours, that is normal during first import, not a fault. Lower Thumbnail Generation concurrency to 2 if you need the machine responsive for anything else.&lt;/p&gt;




&lt;h2&gt;
  
  
  How long does a 20,000 item Google Takeout import take from start to searchable?
&lt;/h2&gt;

&lt;p&gt;Think in phases, not in one number. The upload finishes long before the library is usable, and the gap between those two moments is where people assume something is broken.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Phase&lt;/th&gt;
&lt;th&gt;What is happening&lt;/th&gt;
&lt;th&gt;Typical shape on 4 cores&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Upload via &lt;code&gt;immich-go&lt;/code&gt; or the CLI&lt;/td&gt;
&lt;td&gt;Files transferred, metadata read, duplicates skipped&lt;/td&gt;
&lt;td&gt;Minutes to a few hours, limited by disk or LAN speed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Metadata extraction&lt;/td&gt;
&lt;td&gt;Dates, GPS, camera fields written to Postgres&lt;/td&gt;
&lt;td&gt;Completes soon after upload, rarely the bottleneck&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Thumbnail generation&lt;/td&gt;
&lt;td&gt;Two derivatives written per asset&lt;/td&gt;
&lt;td&gt;Several hours, and the timeline stays gappy until it ends&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Video transcoding&lt;/td&gt;
&lt;td&gt;Non compatible clips re-encoded by FFmpeg&lt;/td&gt;
&lt;td&gt;Hours to days, entirely driven by how much video you have&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Smart search and face detection&lt;/td&gt;
&lt;td&gt;Embeddings and faces computed per asset&lt;/td&gt;
&lt;td&gt;The long tail, often the last queue still running&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A photo heavy library with under 500 short clips is usually fully processed inside 24 hours on 4 cores. Add 3,000 camcorder clips and the same library can take a long weekend, because transcoding runs at concurrency 1 by design.&lt;/p&gt;

&lt;p&gt;Two practical points. First, upload speed and processing speed are independent: &lt;code&gt;immich-go&lt;/code&gt; can finish at 2 a.m. while the job queues still have 18,000 items pending. Second, the Jobs page in the admin panel shows an active and a waiting count per queue, and the waiting count falling is the only honest progress bar you have.&lt;/p&gt;

&lt;p&gt;Do not judge the install until every queue reads zero. Search results, people grouping and the map view are all incomplete before that point, and re-running them costs the same hours again.&lt;/p&gt;




&lt;h2&gt;
  
  
  Should you upload into Immich or point it at an external library?
&lt;/h2&gt;

&lt;p&gt;Upload if Immich is becoming your primary photo home. Use an external library if you already have a folder tree you edit with other tools and want to keep owning.&lt;/p&gt;

&lt;p&gt;The two paths differ in who controls the files on disk:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Uploaded assets live under &lt;code&gt;UPLOAD_LOCATION&lt;/code&gt; and obey the storage template:&lt;/strong&gt; Immich decides the folder layout and filename, moves files when you change the template, and treats deletion in the app as deletion on disk after the trash period.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;External libraries are read only by design:&lt;/strong&gt; you register a path such as &lt;code&gt;/mnt/photos&lt;/code&gt; in the container, Immich indexes what it finds, generates thumbnails into its own directories, and never writes to or renames your originals.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;External libraries need a rescan to notice changes:&lt;/strong&gt; new files added by Lightroom, Syncthing or a NAS share appear after a scan job runs, either on the configured interval or when you trigger it manually, which is not the instant behaviour uploads give you.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mobile backup always uploads:&lt;/strong&gt; the phone app has no concept of an external path, so a mixed setup is normal, historic archive as an external library and new phone photos as uploads.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Duplicate risk is real if you do both with the same files:&lt;/strong&gt; importing a Takeout with &lt;code&gt;immich-go&lt;/code&gt; and also mounting that same folder externally gives you every asset twice, counted twice in storage reporting.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For a 20,000 file RAW archive averaging 25 MB, the external route is the safer first week choice: your originals stay where your backup script already finds them, and a mistake in Immich cannot rename or move 500 GB of files.&lt;/p&gt;




&lt;h2&gt;
  
  
  What does the Google Takeout format break, and how do you fix it before importing?
&lt;/h2&gt;

&lt;p&gt;Takeout does not hand you a photo library. It hands you archives of files whose metadata has been moved out of the images and into JSON sidecars, and if you import them naively your entire timeline collapses onto the import date.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Dates and GPS live in sidecar JSON, not in the file:&lt;/strong&gt; each media file gets a companion &lt;code&gt;.json&lt;/code&gt; holding &lt;code&gt;photoTakenTime&lt;/code&gt; and location, and Immich's plain upload path does not merge them, which is why so many first imports show 20,000 photos all dated today.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sidecar filenames do not reliably match their media:&lt;/strong&gt; Google truncates long names and appends counters like &lt;code&gt;IMG_1234(1).jpg&lt;/code&gt;, so naive pairing by filename fails on a meaningful slice of any large export.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Edited copies arrive alongside originals:&lt;/strong&gt; a file plus its &lt;code&gt;-edited&lt;/code&gt; variant are two assets, which inflates your item count above what Google Photos showed you.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Albums are folders, not metadata:&lt;/strong&gt; album membership is expressed by directory layout and an album JSON, so a plain folder upload gives you every photo and zero albums.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Live Photos are split:&lt;/strong&gt; the still and the paired video land as separate files that need rejoining, not as one asset.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The fix is tooling, not manual cleanup. &lt;code&gt;immich-go&lt;/code&gt; was written for exactly this: it reads the sidecars, rebuilds dates, recreates albums and pairs motion photos, and it can consume the Takeout zip files directly without you unzipping 100 GB first.&lt;/p&gt;

&lt;p&gt;Import a single archive as a test run of a few hundred assets, confirm the timeline dates look correct, then delete those assets and run the full set. Discovering a date problem after 20,000 items means redoing all of it.&lt;/p&gt;




&lt;h2&gt;
  
  
  Which storage template should you set before the first upload?
&lt;/h2&gt;

&lt;p&gt;Decide this before a single asset lands, because the template controls the on disk path of every uploaded file, and changing it later triggers a Storage Migration job that physically moves all 20,000 files.&lt;/p&gt;

&lt;p&gt;Templates are built from variables such as &lt;code&gt;{{y}}&lt;/code&gt;, &lt;code&gt;{{MM}}&lt;/code&gt;, &lt;code&gt;{{filename}}&lt;/code&gt; and &lt;code&gt;{{ext}}&lt;/code&gt;, set in Administration, Settings, Storage Template. Four sane choices:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Template&lt;/th&gt;
&lt;th&gt;Resulting path shape&lt;/th&gt;
&lt;th&gt;Best for&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Disabled, the default&lt;/td&gt;
&lt;td&gt;Random directory and asset id under the user folder&lt;/td&gt;
&lt;td&gt;People who will never touch the files outside Immich&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;{{y}}/{{MM}}/{{filename}}&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;2019/07/IMG_1234.jpg&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Anyone who wants a browsable archive that survives Immich&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;{{y}}/{{y}}-{{MM}}-{{dd}}/{{filename}}&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;2019/2019-07-14/IMG_1234.jpg&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Event heavy libraries where one day equals one shoot&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;{{album}}/{{filename}}&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;Corsica 2019/IMG_1234.jpg&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Album driven workflows, with the caveat that assets in no album fall back&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The default is genuinely defensible. Random paths never collide and never break when you rename an album. The argument against it is portability: if Immich is your only index, a corrupted database leaves you with a directory of meaningless filenames.&lt;/p&gt;

&lt;p&gt;Two hazards. Filename collisions inside the same folder get a numeric suffix, so &lt;code&gt;{{y}}/{{filename}}&lt;/code&gt; on a phone that resets its counter will produce &lt;code&gt;IMG_0001_1.jpg&lt;/code&gt;. And the template applies to uploaded assets only, never to external libraries, which keep their original paths untouched.&lt;/p&gt;

&lt;p&gt;Pick a year and month layout unless you have a specific reason not to. It is readable, it sorts, and it means a plain file browser can still make sense of your archive years from now.&lt;/p&gt;




&lt;h2&gt;
  
  
  Which machine learning model should you choose on day one?
&lt;/h2&gt;

&lt;p&gt;This is the setting with the harshest change penalty. Smart search stores one embedding per asset, and embeddings from different models are not interchangeable, so switching models invalidates all 20,000 of them and forces a full re-run of the Smart Search queue.&lt;/p&gt;

&lt;p&gt;The choice lives in Administration, Settings, Machine Learning Settings, and the models are pulled from the Immich Hugging Face collection on first use.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The default CLIP model:&lt;/strong&gt; shipped because it fits modest hardware, and on a 4 core box with no GPU it is the only option that finishes a 20,000 item library in hours rather than days.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Larger visual models:&lt;/strong&gt; better at abstract queries like "birthday cake" or "snow on a mountain", at the cost of more RAM in the machine learning container and a longer queue, which is exactly the combination that triggers the OOM loop on a 4 GB host.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Multilingual models:&lt;/strong&gt; the only way to search in a language other than English, noticeably heavier than their English only equivalents, and worth choosing on day one if your household does not think in English.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Facial recognition is separate:&lt;/strong&gt; it runs its own detection and recognition models with their own concurrency, so you can keep faces enabled while leaving smart search off entirely.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Off is a valid answer:&lt;/strong&gt; on a Raspberry Pi 5 or an N100, disabling smart search turns a week of queued jobs into a library that is fully browsable by date and album tonight.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Test with 200 assets before committing. Search for three things you would realistically look for, judge the results, then import the rest. Making that judgement after the full import costs you the entire queue again.&lt;/p&gt;




&lt;h2&gt;
  
  
  What should the video transcoding policy be, and what does it cost you?
&lt;/h2&gt;

&lt;p&gt;Leave it on the default optimal policy unless you have a specific reason not to, then understand exactly what that policy is doing to your disk and your weekend.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Transcode policy&lt;/th&gt;
&lt;th&gt;What it re-encodes&lt;/th&gt;
&lt;th&gt;What it costs you&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Don't transcode&lt;/td&gt;
&lt;td&gt;Nothing&lt;/td&gt;
&lt;td&gt;Zero extra disk, but clips your browser cannot decode simply will not play&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Videos not in an accepted format&lt;/td&gt;
&lt;td&gt;Only unsupported codecs and containers&lt;/td&gt;
&lt;td&gt;The cheapest useful option, ideal for phone footage that is already H.264 in MP4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Videos higher than target resolution or not in an accepted format&lt;/td&gt;
&lt;td&gt;The above, plus anything above the target resolution, 720p by default&lt;/td&gt;
&lt;td&gt;The default, and the reason a 4K library balloons in processing time&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;All videos&lt;/td&gt;
&lt;td&gt;Every clip regardless of source&lt;/td&gt;
&lt;td&gt;Doubles your video storage and can run for days on 3,000 items&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Three things decide the real cost. The target resolution, because raising it to 1080p produces larger transcodes but keeps quality your phone actually shot. The constant rate factor, where the default of 23 trades size against fidelity. And hardware acceleration, which is the difference between hours and days: Immich supports NVENC, QSV, VAAPI and RKMPP, enabled by adding &lt;code&gt;hwaccel.transcoding.yml&lt;/code&gt; to your compose command and selecting the matching API in the settings.&lt;/p&gt;

&lt;p&gt;The trap is that the original is always kept. Transcodes are additional files in &lt;code&gt;encoded-video/&lt;/code&gt;, never replacements, so an aggressive policy costs storage on top of your originals rather than saving any.&lt;/p&gt;

&lt;p&gt;If your video is mostly modern phone footage, switching to "videos not in an accepted format" before the first import can remove most of the transcoding queue entirely.&lt;/p&gt;




&lt;h2&gt;
  
  
  Where should you actually run Immich: NAS, mini PC, home server or hosted?
&lt;/h2&gt;

&lt;p&gt;The hardware question is really a question about video and about who fixes it at 11 p.m. when the phone app stops backing up.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A consumer NAS, Synology or QNAP class:&lt;/strong&gt; attractive because the disks are already there, but many models ship 2 GB or 4 GB of RAM shared with the NAS operating system, and their Celeron and ARM CPUs turn the transcoding queue into a multi day job. Check for a container station and free RAM before assuming it will work.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A mini PC, N100 or similar:&lt;/strong&gt; the sweet spot for a first Immich server. Four cores, upgradable RAM, low idle draw, and Intel Quick Sync on board, which is the single change that collapses video processing time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;An old desktop or a home server:&lt;/strong&gt; the most CPU per euro you will ever get, and the right answer if you already own one. The costs are electricity, noise and the fact that your library now depends on a machine with no redundancy.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A VPS or a managed personal server:&lt;/strong&gt; removes the hardware problem and the port forwarding problem, and moves the storage question into a monthly bill, since 200 GB of block storage is a different line item than a 4 TB disk you buy once. Yundera is a managed Personal Cloud Server, built on CasaOS, where apps are installed from an app store in one click rather than assembled from compose files by hand, and each app is reachable on a public HTTPS subdomain via NSL.SH mesh routing.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Whatever you pick, the deciding number is free RAM, not disk. Disk you can add later. A 4 GB ceiling shared with three other containers is what actually stops a 20,000 item import.&lt;/p&gt;




&lt;h2&gt;
  
  
  Which day-one settings force a full re-run if you change them later?
&lt;/h2&gt;

&lt;p&gt;Some Immich settings are free to change at any time. Five are not, and the difference is whether the change invalidates work already written to disk. On a 20,000 item library, each of these costs you the same hours you spent during the first import.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Setting&lt;/th&gt;
&lt;th&gt;What changing it later triggers&lt;/th&gt;
&lt;th&gt;Safe default for week one&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Storage template&lt;/td&gt;
&lt;td&gt;A Storage Migration job that moves every uploaded file on disk&lt;/td&gt;
&lt;td&gt;Set it before the first upload, then leave it&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Smart search model&lt;/td&gt;
&lt;td&gt;All existing embeddings discarded, full Smart Search queue re-run&lt;/td&gt;
&lt;td&gt;Pick the model you can afford to run, test on 200 assets&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Preview and thumbnail size&lt;/td&gt;
&lt;td&gt;Thumbnail Generation re-runs for every asset, old derivatives replaced&lt;/td&gt;
&lt;td&gt;Accept the default unless you view photos on a 4K display&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Transcode policy or target resolution&lt;/td&gt;
&lt;td&gt;Every clip matching the new rule is re-encoded from the original&lt;/td&gt;
&lt;td&gt;Decide by looking at what codecs your clips actually use&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Facial recognition model&lt;/td&gt;
&lt;td&gt;Face detection and recognition re-run, and people you named can need reassigning&lt;/td&gt;
&lt;td&gt;Enable it once and do not switch models casually&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two settings that look scary and are not: &lt;code&gt;UPLOAD_LOCATION&lt;/code&gt; can be moved if you move the directory and keep the structure intact, and job concurrency can be raised or lowered mid import with no penalty at all.&lt;/p&gt;

&lt;p&gt;The practical rule is simple. Anything that changes how a derived file is generated forces regeneration of every derived file. Anything that changes scheduling or naming of the running system does not.&lt;/p&gt;

&lt;p&gt;Write your five choices down before you import. Reviewing them takes 15 minutes. Discovering one was wrong in month two costs you another weekend of queued jobs.&lt;/p&gt;

</description>
      <category>immich</category>
      <category>selfhosted</category>
      <category>docker</category>
      <category>photography</category>
    </item>
    <item>
      <title>ConvertX Hardware Requirements: How Much CPU, RAM and Disk Before Conversions Start Failing</title>
      <dc:creator>John</dc:creator>
      <pubDate>Thu, 27 Aug 2026 07:06:17 +0000</pubDate>
      <link>https://dev.to/john_182319291/convertx-hardware-requirements-how-much-cpu-ram-and-disk-before-conversions-start-failing-5dal</link>
      <guid>https://dev.to/john_182319291/convertx-hardware-requirements-how-much-cpu-ram-and-disk-before-conversions-start-failing-5dal</guid>
      <description>&lt;p&gt;ConvertX will install and run on a 1 vCPU, 1 GB virtual machine, and it will convert a 2 MB PNG to WebP in a second. That box will also kill a 4K video conversion, a 300 MB PDF and a large EPUB rebuild, usually with no visible error at all. The honest hardware floor for general use is 2 vCPU and 4 GB of RAM with 3 times your largest file free on disk, and the floor for routine video work is 4 vCPU and 8 GB. If your real workload is long video transcodes or 500 MB office documents, and you have no appetite for reading container logs, a hosted converter will fail less than a small self-hosted instance.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TL;DR by reader profile:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The occasional image converter, for example someone flattening HEIC holiday photos to JPEG once a month:&lt;/strong&gt; self-host ConvertX on the smallest tier you have, because single image jobs finish in seconds and memory pressure never builds.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The document and ebook person, for example a researcher pushing DOCX and EPUB through Pandoc and Calibre weekly:&lt;/strong&gt; self-host on 2 vCPU and 4 GB, because LibreOffice and Calibre are single threaded and want headroom rather than cores.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The video converter, for example a parent re-encoding phone clips to H.264 for a TV:&lt;/strong&gt; budget 4 vCPU, 8 GB and a large scratch volume, or accept conversions measured in hours.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The privacy-driven user with sensitive files, for example someone converting bank statements and medical scans:&lt;/strong&gt; self-host even on modest hardware, because the failure mode is a slow job, not a document sitting on a third party server.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The shared household or small team instance, for example five people with one link:&lt;/strong&gt; treat concurrency as the real constraint, because two overlapping jobs on 4 GB will OOM before either one finishes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The reader with no sysadmin appetite at all, for example someone who cannot read &lt;code&gt;docker logs&lt;/code&gt; and does not want to:&lt;/strong&gt; use a hosted converter for anything over 100 MB, because silent failure with no diagnosis is worse than a public upload you understand.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The central tradeoff is this: ConvertX moves your files off other people's servers, but it also moves every timeout, memory limit and disk exhaustion onto a machine you now have to size correctly yourself.&lt;/p&gt;




&lt;h2&gt;
  
  
  Table of contents
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;What happens inside ConvertX when you press Convert&lt;/li&gt;
&lt;li&gt;What are the real hardware floors for ConvertX, tier by tier&lt;/li&gt;
&lt;li&gt;The upload path: where large files die before ConvertX ever sees them&lt;/li&gt;
&lt;li&gt;How do you tell an OOM kill from a timeout in ConvertX&lt;/li&gt;
&lt;li&gt;Why does a ConvertX job finish with no output and no error&lt;/li&gt;
&lt;li&gt;Scratch disk: the hidden multiplier on every ConvertX conversion&lt;/li&gt;
&lt;li&gt;Concurrency in ConvertX: what breaks when two jobs overlap&lt;/li&gt;
&lt;li&gt;Video and FFmpeg in ConvertX: the workload that decides your tier&lt;/li&gt;
&lt;li&gt;Images and ImageMagick in ConvertX: pixel maths, memory limits and policy files&lt;/li&gt;
&lt;li&gt;Documents, ebooks and LibreOffice in ConvertX: slow, single threaded, quietly fragile&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  What happens inside ConvertX when you press Convert
&lt;/h2&gt;

&lt;p&gt;ConvertX is not one converter. It is a Bun web application that receives your upload, writes it to disk, then shells out to whichever command line tool handles that format pair. Understanding that chain is the whole hardware question, because the web layer costs almost nothing and the child process costs everything.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The upload lands on disk first:&lt;/strong&gt; your file is written into the container's data volume before any conversion starts, so a 500 MB video occupies 500 MB before a single frame is encoded.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A converter binary is selected by format pair:&lt;/strong&gt; FFmpeg for audio and video, ImageMagick and libvips for raster images, Inkscape and resvg for vectors, Calibre for ebooks, Pandoc and LibreOffice for documents, Assimp for 3D meshes. Each has its own memory behaviour and none of them know about the others.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The job runs as a child process, not inside the web server:&lt;/strong&gt; ConvertX waits for that process to exit. If the kernel kills it, ConvertX sees a dead process, not an explanation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The output is written beside the input:&lt;/strong&gt; peak disk usage is input plus output plus whatever temporary files the tool creates, which for LibreOffice and Calibre is substantial.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Results persist until cleanup runs:&lt;/strong&gt; &lt;code&gt;AUTO_DELETE_EVERY_N_HOURS&lt;/code&gt; defaults to 24, so a day of conversions accumulates on the volume before anything is reclaimed.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The practical consequence: ConvertX itself idles at a few hundred megabytes. Your 8 GB requirement is not ConvertX, it is FFmpeg encoding 4K, or ImageMagick decompressing a 200 megapixel TIFF into an uncompressed pixel buffer. Size the box for the worst binary you will ever invoke, not for the web interface.&lt;/p&gt;




&lt;h2&gt;
  
  
  What are the real hardware floors for ConvertX, tier by tier
&lt;/h2&gt;

&lt;p&gt;Pick your tier from the heaviest format you actually convert, not from the average one. The table below maps common hardware to what genuinely completes and what starts failing.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tier&lt;/th&gt;
&lt;th&gt;Completes reliably&lt;/th&gt;
&lt;th&gt;Starts failing at&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1 vCPU, 1 GB RAM&lt;/td&gt;
&lt;td&gt;Single images under 20 megapixels, PDF to text, small Markdown and DOCX via Pandoc, audio transcodes&lt;/td&gt;
&lt;td&gt;Any video over roughly 2 minutes, LibreOffice on large spreadsheets, two concurrent jobs of any kind&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2 vCPU, 4 GB RAM&lt;/td&gt;
&lt;td&gt;Batches of 20 photos, EPUB and MOBI via Calibre, 1080p clips under 5 minutes, most office documents&lt;/td&gt;
&lt;td&gt;4K video, TIFF files over 100 megapixels, three or more overlapping users&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4 vCPU, 8 GB RAM&lt;/td&gt;
&lt;td&gt;1080p video at usable speed, 4K short clips, large PDF rasterisation, 2 to 3 concurrent jobs&lt;/td&gt;
&lt;td&gt;Long 4K transcodes measured in hours, batch video, sustained multi user load&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;8 vCPU, 16 GB RAM&lt;/td&gt;
&lt;td&gt;Feature length video, large batch image work, several simultaneous users&lt;/td&gt;
&lt;td&gt;Little in practice, the constraint becomes disk throughput and patience&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Disk matters as much as RAM. Reserve at least 3 times the size of your largest file, and remember that with &lt;code&gt;AUTO_DELETE_EVERY_N_HOURS&lt;/code&gt; at its default of 24, a day of jobs stays resident.&lt;/p&gt;

&lt;p&gt;Where that hardware lives is a separate decision. A self managed VPS, an always on home server, a NAS running Docker and a managed option all reach the same floors. Yundera is a managed Personal Cloud Server, built on CasaOS, that runs self-hosted apps as Docker containers on a server dedicated to the user. Whichever you choose, the tier table above is what decides whether a conversion finishes.&lt;/p&gt;




&lt;h2&gt;
  
  
  The upload path: where large files die before ConvertX ever sees them
&lt;/h2&gt;

&lt;p&gt;Half of the "ConvertX is broken" reports are not ConvertX at all. The file never reached the container. Every layer between your browser and port 3000 has its own size ceiling, and most of them reject silently or with a generic 413.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Nginx caps uploads at 1 MB by default:&lt;/strong&gt; &lt;code&gt;client_max_body_size&lt;/code&gt; starts at &lt;code&gt;1m&lt;/code&gt;. Until you raise it, anything larger returns 413 Request Entity Too Large and ConvertX logs nothing, because no request arrived. Set it to a value above your largest file, or to &lt;code&gt;0&lt;/code&gt; to disable the check.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cloudflare's free proxy caps request bodies at 100 MB:&lt;/strong&gt; if your domain is proxied through Cloudflare on a free plan, a 400 MB video cannot reach your server regardless of your own configuration. Grey clouding the record or using a direct hostname is the only fix.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Proxy read timeouts kill slow uploads:&lt;/strong&gt; nginx defaults &lt;code&gt;proxy_read_timeout&lt;/code&gt; to 60 seconds. A large file over a slow home uplink can exceed that before the body finishes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Your upstream bandwidth is the real limit at home:&lt;/strong&gt; a 10 Mbit upload link moves roughly 1.25 MB per second, so a 500 MB file takes over 6 minutes of held connection before conversion even begins.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Caddy and Traefik behave differently:&lt;/strong&gt; Caddy applies no default body limit, so people migrating from nginx often see the ceiling vanish without understanding why.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Where you run the instance changes which of these you have to configure. A self managed VPS means editing the proxy yourself, a NAS exposes its own reverse proxy settings, and on Yundera each app is reachable on a public HTTPS subdomain via NSL.SH mesh routing. Test with a file at your true maximum size before trusting any of it.&lt;/p&gt;




&lt;h2&gt;
  
  
  How do you tell an OOM kill from a timeout in ConvertX
&lt;/h2&gt;

&lt;p&gt;These two produce almost identical symptoms in the browser: a job that never completes. They need opposite fixes. More RAM does nothing for a timeout, and a longer timeout does nothing for an OOM kill. Diagnose before you resize.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Signal&lt;/th&gt;
&lt;th&gt;Points to an OOM kill&lt;/th&gt;
&lt;th&gt;Points to a timeout&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Exit status in &lt;code&gt;docker logs convertx&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Process terminated by signal 9, container exit code 137&lt;/td&gt;
&lt;td&gt;Process still running, or killed cleanly after a fixed interval&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;docker inspect --format '{{.State.OOMKilled}}' convertx&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Returns &lt;code&gt;true&lt;/code&gt; after the container itself dies&lt;/td&gt;
&lt;td&gt;Returns &lt;code&gt;false&lt;/code&gt;, the container is healthy&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Host kernel log via &lt;code&gt;dmesg -T&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Contains "Out of memory: Killed process" naming ffmpeg, soffice or convert&lt;/td&gt;
&lt;td&gt;Contains nothing at all&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Browser behaviour&lt;/td&gt;
&lt;td&gt;Connection stays open, then the job vanishes from progress&lt;/td&gt;
&lt;td&gt;504 Gateway Timeout from the reverse proxy while the job continues server side&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;docker stats&lt;/code&gt; during the job&lt;/td&gt;
&lt;td&gt;Memory climbs to the limit then the row disappears&lt;/td&gt;
&lt;td&gt;Memory sits flat, CPU pegged at 100 percent on one core&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The pattern worth memorising: an OOM kill is sudden and leaves kernel evidence, a timeout is patient and leaves the converter alive. A frequent surprise is a job that hits both, where FFmpeg keeps encoding for 40 minutes after the proxy already returned 504, so the output file appears later with no matching entry in your browser.&lt;/p&gt;

&lt;p&gt;Add swap before you upgrade the tier. On a 4 GB box, 2 GB of swap turns some ImageMagick kills into slow but successful conversions. It will not save a 4K transcode, and a job that swaps heavily for 20 minutes is usually a signal that the tier is wrong rather than the configuration.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why does a ConvertX job finish with no output and no error
&lt;/h2&gt;

&lt;p&gt;This is the failure mode that drives people back to hosted converters. The job appears in the history, the status does not scream at you, and the download is either missing or a 0 byte file. ConvertX only knows what the child process told it, and command line converters lie about success constantly.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The converter exited 0 but wrote nothing:&lt;/strong&gt; LibreOffice is the worst offender here. It returns success after refusing to convert a document it could not parse, so ConvertX records a finished job pointing at an empty file.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A stale LibreOffice profile lock:&lt;/strong&gt; if a previous conversion was killed mid run, the leftover profile directory blocks the next invocation. The fix is restarting the container, which is why the same document fails five times then works after &lt;code&gt;docker restart convertx&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Missing fonts or unsupported subfeatures:&lt;/strong&gt; a DOCX with an embedded font, or a PDF using a CJK typeface the container lacks, converts to a page of blank boxes rather than an error.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The output was cleaned up before you fetched it:&lt;/strong&gt; with &lt;code&gt;AUTO_DELETE_EVERY_N_HOURS&lt;/code&gt; at 24, a job you started last night and downloaded this evening can be gone.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Disk filled mid write:&lt;/strong&gt; the converter writes a partial file, then fails on the final flush. You get a truncated result and a job marked done.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Your only reliable diagnostic is &lt;code&gt;docker logs convertx&lt;/code&gt;, ideally with &lt;code&gt;--tail 200&lt;/code&gt; right after the failure. If reading container logs to work out why a PDF came back blank sounds like a bad evening, that is a legitimate reason to keep a hosted converter for anything important.&lt;/p&gt;




&lt;h2&gt;
  
  
  Scratch disk: the hidden multiplier on every ConvertX conversion
&lt;/h2&gt;

&lt;p&gt;RAM gets the attention, disk causes the outages. ConvertX needs room for the input, the output and whatever intermediate the converter invents along the way, all at the same moment. That peak is routinely 5 to 10 times the size of the file you uploaded.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Rasterising a PDF is arithmetic, not magic:&lt;/strong&gt; one A4 page at 300 DPI in RGB is about 3508 by 2480 pixels, roughly 26 MB uncompressed. A 100 page document becomes around 2.6 GB of intermediate data before anything is written as PNG or JPEG.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Uncompressed video intermediates are enormous:&lt;/strong&gt; a single 1080p RGB frame is 1920 by 1080 by 3 bytes, about 6.2 MB. At 25 frames per second that is 155 MB of scratch per second of footage if a lossless intermediate is produced.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Calibre and LibreOffice unpack before they convert:&lt;/strong&gt; an EPUB or DOCX is a zip archive, and the expanded working directory plus embedded images can dwarf the original file.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Everything accumulates in the same volume:&lt;/strong&gt; uploads, outputs and &lt;code&gt;mydb.sqlite&lt;/code&gt; share &lt;code&gt;/app/data&lt;/code&gt;. If you did not mount a named volume, all of it lands on the container layer and fills the host root partition instead.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Retention multiplies the total:&lt;/strong&gt; at the default 24 hour cleanup, a day of 20 conversions averaging 200 MB in and 150 MB out holds roughly 7 GB before anything is reclaimed.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Check with &lt;code&gt;df -h&lt;/code&gt; on the host and &lt;code&gt;docker system df&lt;/code&gt; for volume usage, not by guessing from your upload sizes. A full disk produces truncated outputs rather than clean errors, so a 20 GB volume on a video oriented instance is a floor, not a comfortable allocation.&lt;/p&gt;




&lt;h2&gt;
  
  
  Concurrency in ConvertX: what breaks when two jobs overlap
&lt;/h2&gt;

&lt;p&gt;ConvertX has no job queue that serialises heavy work. A batch upload of 12 files, or two people clicking Convert at the same moment, means multiple converter processes competing for the same cores and the same memory ceiling. This is the single most common reason a box that felt adequate suddenly fails.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;FFmpeg claims every core by default:&lt;/strong&gt; with no explicit &lt;code&gt;-threads&lt;/code&gt; value it scales to all available CPUs. On 4 vCPU, two simultaneous encodes request 8 threads across 4 cores, so both run at roughly half speed plus context switching overhead, and neither finishes when you expect.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ImageMagick multiplies memory, not just time:&lt;/strong&gt; each process holds its own pixel buffer. Two 100 megapixel TIFF conversions on a 4 GB box are not one job twice, they are two independent allocations racing the same limit.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;LibreOffice does not like company:&lt;/strong&gt; two &lt;code&gt;soffice&lt;/code&gt; invocations contending for the same user profile directory produce one success and one job that returns nothing usable.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;SQLite serialises writes:&lt;/strong&gt; &lt;code&gt;mydb.sqlite&lt;/code&gt; handles job records fine at household scale, but it is a single writer database, not something to plan a 20 user instance around.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The web interface stays responsive while everything else starves:&lt;/strong&gt; Bun keeps serving pages, so the instance looks healthy while conversions crawl. Health checks tell you nothing here.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The practical control is Docker resource limits. Setting &lt;code&gt;cpus: "3.0"&lt;/code&gt; and &lt;code&gt;mem_limit: 6g&lt;/code&gt; on a 4 vCPU, 8 GB host leaves the operating system breathing room and makes the failure predictable rather than random. If your instance serves more than three people who convert video, plan for one job at a time and tell them so.&lt;/p&gt;




&lt;h2&gt;
  
  
  Video and FFmpeg in ConvertX: the workload that decides your tier
&lt;/h2&gt;

&lt;p&gt;If you never convert video, ignore most of the hardware advice in this article and run ConvertX on whatever you have. If you do, video alone sets your tier, because encoding is the only workload here that saturates every core for a sustained period.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Source material&lt;/th&gt;
&lt;th&gt;On 2 vCPU, 4 GB&lt;/th&gt;
&lt;th&gt;On 4 vCPU, 8 GB&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Audio only, MP3 to FLAC or WAV&lt;/td&gt;
&lt;td&gt;Completes in seconds, no real load&lt;/td&gt;
&lt;td&gt;Identical, the tier is irrelevant&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1080p H.264 clip under 5 minutes&lt;/td&gt;
&lt;td&gt;Completes, slower than realtime, tolerable&lt;/td&gt;
&lt;td&gt;Completes comfortably, cores are the binding constraint&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1080p H.264 to H.265&lt;/td&gt;
&lt;td&gt;Encoding runs far slower than the equivalent x264 job, expect to leave it running&lt;/td&gt;
&lt;td&gt;Usable, still the slowest common conversion&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4K to 1080p, 10 minutes or longer&lt;/td&gt;
&lt;td&gt;Frequently exceeds proxy patience and sometimes memory&lt;/td&gt;
&lt;td&gt;Completes, but plan for an unattended run&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Feature length or batch video&lt;/td&gt;
&lt;td&gt;Not a realistic workload&lt;/td&gt;
&lt;td&gt;Marginal, this is where you want 8 vCPU&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Three facts explain the whole table. A 4K frame carries 8.3 million pixels against 2.1 million at 1080p, so the same encoder does roughly four times the work per frame. The x264 preset ladder from &lt;code&gt;ultrafast&lt;/code&gt; to &lt;code&gt;veryslow&lt;/code&gt; spans an enormous speed range at similar quality targets, and ConvertX exposes this through the &lt;code&gt;FFMPEG_ARGS&lt;/code&gt; environment variable, so setting &lt;code&gt;-preset veryfast&lt;/code&gt; is the cheapest tier upgrade available. Finally, the container has no GPU acceleration unless you explicitly pass a device through, so everything runs on the CPU.&lt;/p&gt;

&lt;p&gt;Set &lt;code&gt;FFMPEG_ARGS&lt;/code&gt; before you buy more vCPU. A preset change costs nothing and often moves a job from abandoned to finished.&lt;/p&gt;




&lt;h2&gt;
  
  
  Images and ImageMagick in ConvertX: pixel maths, memory limits and policy files
&lt;/h2&gt;

&lt;p&gt;Image conversion looks harmless because the files are small. The decoded buffer is what matters, and that number has nothing to do with the size on disk. A 12 MB JPEG and a 12 MB PNG can consume wildly different amounts of RAM.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Work out the buffer, not the file size:&lt;/strong&gt; ImageMagick is normally built at Q16, meaning 2 bytes per channel. RGBA at 16 bits is 8 bytes per pixel, so a 100 megapixel scan needs roughly 800 MB of memory before any processing begins. That alone explains most failures on a 1 GB box.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The pixel cache silently falls back to disk:&lt;/strong&gt; when the memory limit is reached, ImageMagick starts using a disk backed cache rather than failing. The job still completes, hundreds of times slower, which is why one photo occasionally takes 20 minutes while its neighbours take 2 seconds.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;policy.xml&lt;/code&gt; can refuse the conversion outright:&lt;/strong&gt; Debian based images ship with the PDF, PS and EPS coders disabled. The giveaway in the logs is "attempt to perform an operation not allowed by the security policy". That is a policy decision, not a hardware problem, and adding RAM will never fix it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The same file has a cheap path and an expensive one:&lt;/strong&gt; ConvertX also bundles libvips, which streams in tiles rather than loading the full image. Where both can handle a format pair, the vips route uses a fraction of the memory.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Batches compound the buffer, not the file count:&lt;/strong&gt; 20 photos at 24 megapixels is not 20 small jobs, it is repeated 400 MB allocations.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Check a suspicious file with &lt;code&gt;identify -verbose&lt;/code&gt; before blaming the tier.&lt;/p&gt;




&lt;h2&gt;
  
  
  Documents, ebooks and LibreOffice in ConvertX: slow, single threaded, quietly fragile
&lt;/h2&gt;

&lt;p&gt;Document conversion is the workload people assume is trivial and then cannot explain. It rarely triggers an OOM kill on 4 GB. It just takes far longer than the file size suggests, and it does not get faster when you add cores.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Which binary handles the job decides everything:&lt;/strong&gt; Pandoc converting DOCX to Markdown or HTML is a text transformation that finishes almost immediately on any tier. The same DOCX to PDF goes through LibreOffice or XeLaTeX, which is a completely different order of cost.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;LibreOffice pays a cold start on every single job:&lt;/strong&gt; each conversion launches &lt;code&gt;soffice&lt;/code&gt; in headless mode, which means loading an entire office suite before the first page is laid out. That fixed overhead dominates on small documents, so a 40 KB letter and a 4 MB report can take a similar time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Layout is single threaded:&lt;/strong&gt; repagination, table reflow and font metrics run on one core. A 4 vCPU box converts a large spreadsheet no faster than a 1 vCPU box, which is why upgrading the tier for document work is usually wasted money.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;LaTeX routes run multiple passes:&lt;/strong&gt; XeLaTeX processes a document two or three times to resolve the table of contents and cross references, so the wall clock is a multiple of one pass, not one pass.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Calibre scales with embedded images, not page count:&lt;/strong&gt; a 600 page text only EPUB converts easily, while a 90 page illustrated PDF to EPUB decodes every image and needs far more headroom.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The useful move is choosing your target format deliberately. If HTML or Markdown will do, you avoid LibreOffice and LaTeX entirely and turn a minutes long job into a seconds long one.&lt;/p&gt;

</description>
      <category>selfhosted</category>
      <category>docker</category>
      <category>ffmpeg</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Why Your Self-Hosted Stirling-PDF Container OOMs, Hangs or Returns a Broken PDF: 15 Anti-Patterns and Their Fixes</title>
      <dc:creator>John</dc:creator>
      <pubDate>Tue, 25 Aug 2026 07:06:47 +0000</pubDate>
      <link>https://dev.to/john_182319291/why-your-self-hosted-stirling-pdf-container-ooms-hangs-or-returns-a-broken-pdf-15-anti-patterns-4db2</link>
      <guid>https://dev.to/john_182319291/why-your-self-hosted-stirling-pdf-container-ooms-hangs-or-returns-a-broken-pdf-15-anti-patterns-4db2</guid>
      <description>&lt;p&gt;Stirling-PDF is not an unstable application. Almost every crash, hang and quietly mangled output traces back to one of fifteen decisions you made outside the app: an unbounded container, the wrong image variant, a reverse proxy that gives up after 60 seconds, or an OCR call with no language pack behind it. The Java process inside the container claims a default heap of 25 percent of visible RAM, office conversions serialise through a single background LibreOffice process, and every job writes scratch files to disk before it returns a single byte. Correct those four assumptions and most of the remaining catalogue stops happening.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TL;DR by profile:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The weekend homelab tinkerer (one N100 mini PC, roughly ten PDFs a week, merge and split only):&lt;/strong&gt; run the ultra-lite image with a hard 1 GB container limit, because you never touch the two subsystems that consume memory.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The scan hoarder (a decade of paper, 40,000 pages, OCR every one of them):&lt;/strong&gt; run the full image, mount a real tessdata volume with your languages, and process in batches of 50 pages or fewer, because OCR cost scales with page count and DPI, not file count.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The automation builder (an n8n or cron job calling the Stirling-PDF API unattended):&lt;/strong&gt; add your own queue and set explicit timeouts on both sides, because the app accepts every concurrent request you throw at it and holds each one in memory.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The small team sharing one instance (four people behind a VPN, mixed office conversions):&lt;/strong&gt; budget for LibreOffice serialisation and raise proxy timeouts before you raise CPU, because your bottleneck is a single conversion process, not cores.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The NAS owner (a box already running twenty other containers):&lt;/strong&gt; set the memory limit and the temp volume first, because an unbounded Stirling-PDF job is the container most likely to evict your other services.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The central tradeoff is headroom against completion: every limit that stops Stirling-PDF from taking down its host also stops some legitimate large job from ever finishing.&lt;/p&gt;




&lt;h2&gt;
  
  
  Table of contents
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;What is actually failing when Stirling-PDF dies, and how do you confirm it in two minutes?&lt;/li&gt;
&lt;li&gt;Anti-pattern 1 and 2: no container memory limit, and no JVM heap ceiling inside it&lt;/li&gt;
&lt;li&gt;Anti-pattern 3 and 4: running the wrong image variant, then calling a tool it does not contain&lt;/li&gt;
&lt;li&gt;Where does Stirling-PDF write its temporary files, and why does that fill your disk?&lt;/li&gt;
&lt;li&gt;Why does OCR hang, fail outright, or return a file with no selectable text?&lt;/li&gt;
&lt;li&gt;Anti-pattern 9 and 10: treating office conversion as fast, parallel and font-independent&lt;/li&gt;
&lt;li&gt;Why does a large upload fail before Stirling-PDF ever sees the file?&lt;/li&gt;
&lt;li&gt;Anti-pattern 12: reverse proxy timeouts shorter than the job you just started&lt;/li&gt;
&lt;li&gt;Anti-pattern 13: compression and repair settings that return a valid file with wrong content&lt;/li&gt;
&lt;li&gt;Anti-pattern 14: volume permissions, custom fonts and language packs mounted incorrectly&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  What is actually failing when Stirling-PDF dies, and how do you confirm it in two minutes?
&lt;/h2&gt;

&lt;p&gt;Three different failures get reported as "Stirling-PDF crashed", and they need opposite fixes. Run &lt;code&gt;docker inspect stirling-pdf --format '{{.State.ExitCode}}'&lt;/code&gt; and &lt;code&gt;docker logs --tail 200 stirling-pdf&lt;/code&gt; before you change anything. The exit code alone separates two of the three cases.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Exit code 137, container killed by the host:&lt;/strong&gt; the kernel OOM killer reclaimed the process because the container exceeded its cgroup memory limit, or the host ran out of RAM entirely. The application log ends mid-sentence with no stack trace, which is the giveaway. Confirm it with &lt;code&gt;dmesg -T | grep -i oom&lt;/code&gt; on the host or &lt;code&gt;docker inspect&lt;/code&gt; reporting &lt;code&gt;OOMKilled: true&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A &lt;code&gt;java.lang.OutOfMemoryError: Java heap space&lt;/code&gt; stack trace, container still running:&lt;/strong&gt; the JVM hit its own heap ceiling while the container still had free memory. The web UI usually stays reachable and only that one job fails. This is the fix nobody applies, because the container looks healthy.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No error at all, the request just never returns:&lt;/strong&gt; the job is running, but something between your browser and the app gave up first. A 504 after roughly 60 seconds points at the reverse proxy, not at Stirling-PDF.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;HTTP 200 with a file that opens but is wrong:&lt;/strong&gt; text vanished, fonts substituted, images blurred to unreadable. Nothing in the log marks this as an error, because from the application's point of view it succeeded.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The first two are memory problems, the third is a timeout problem, and the fourth is a settings problem. Diagnosing the wrong one costs you a weekend.&lt;/p&gt;




&lt;h2&gt;
  
  
  Anti-pattern 1 and 2: no container memory limit, and no JVM heap ceiling inside it
&lt;/h2&gt;

&lt;p&gt;Docker applies no memory limit by default. A container-aware JVM with no &lt;code&gt;-Xmx&lt;/code&gt; claims a maximum heap of 25 percent of the memory it can see, which is the host's total RAM when you set no limit. Both defaults are wrong for this workload, and they fail in opposite directions.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Configuration&lt;/th&gt;
&lt;th&gt;What the JVM sees&lt;/th&gt;
&lt;th&gt;What breaks&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;No &lt;code&gt;mem_limit&lt;/code&gt;, no heap setting&lt;/td&gt;
&lt;td&gt;25 percent of host RAM as max heap, on 8 GB that is 2 GB&lt;/td&gt;
&lt;td&gt;The host runs out of RAM before the container does, and your other services get evicted first&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;mem_limit: 2g&lt;/code&gt;, no heap setting&lt;/td&gt;
&lt;td&gt;512 MB max heap, calculated from the limit&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;OutOfMemoryError&lt;/code&gt; on a job the machine could easily have handled&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;mem_limit: 2g&lt;/code&gt;, heap raised to 1.5 GB&lt;/td&gt;
&lt;td&gt;1.5 GB heap, 512 MB left for everything else&lt;/td&gt;
&lt;td&gt;Exit 137 the moment a native helper process starts, because Ghostscript and Tesseract live outside the heap&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;mem_limit: 4g&lt;/code&gt;, heap capped at 2 GB&lt;/td&gt;
&lt;td&gt;2 GB heap, 2 GB for native processes and page cache&lt;/td&gt;
&lt;td&gt;Nothing, for a single-user instance&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The subtlety that catches people out: capping the Java heap does not cap the container. LibreOffice, Ghostscript, Tesseract and qpdf are separate native processes. Their memory counts against the cgroup limit but never against &lt;code&gt;-Xmx&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Set both, and leave the native processes at least as much room as the heap. Add &lt;code&gt;JAVA_TOOL_OPTIONS=-XX:MaxRAMPercentage=50&lt;/code&gt; rather than a fixed &lt;code&gt;-Xmx&lt;/code&gt;, so the ratio survives a change to the container limit. Then watch &lt;code&gt;docker stats stirling-pdf&lt;/code&gt; during your largest real job instead of guessing.&lt;/p&gt;




&lt;h2&gt;
  
  
  Anti-pattern 3 and 4: running the wrong image variant, then calling a tool it does not contain
&lt;/h2&gt;

&lt;p&gt;Stirling-PDF ships as more than one image, and they are not interchangeable. Pulling &lt;code&gt;latest&lt;/code&gt; because it sounds current, or &lt;code&gt;ultra-lite&lt;/code&gt; because it sounds efficient, decides which of the 50 plus tools actually work at runtime. The UI still shows every button either way, which is why this fails so confusingly.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;stirlingtools/stirling-pdf:latest-ultra-lite&lt;/code&gt;:&lt;/strong&gt; core PDF manipulation only, so merge, split, rotate, reorder and metadata edits. No OCR, no LibreOffice, no Python. It is the right choice if you genuinely never convert or OCR, and the wrong choice the first time you try.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;stirlingtools/stirling-pdf:latest&lt;/code&gt;:&lt;/strong&gt; the standard build, with OCR and office conversion included. This is the default answer for most self-hosters, and the variant the rest of this article assumes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;stirlingtools/stirling-pdf:latest-fat&lt;/code&gt;:&lt;/strong&gt; everything preinstalled, including the ebook and advanced HTML tooling that the standard image otherwise fetches at container start when you set &lt;code&gt;INSTALL_BOOK_AND_ADVANCED_HTML_OPS=true&lt;/code&gt;. It trades disk for a container that starts ready and needs no outbound network on boot.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The tool you call but did not install:&lt;/strong&gt; the request reaches the backend, the underlying binary is absent, and you get a generic failure rather than "this image cannot do that". Confirm with &lt;code&gt;docker exec stirling-pdf which soffice tesseract&lt;/code&gt; before blaming your configuration.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Where you run it changes how much this bites. On a self-managed VPS, a home server or a NAS you pick the tag yourself in a compose file. Yundera is a managed Personal Cloud Server, built on CasaOS, that runs self-hosted apps as Docker containers on a server dedicated to the user, so the variant arrives already chosen by the packaged app. Either way, verify which binaries exist before you design a workflow around them.&lt;/p&gt;




&lt;h2&gt;
  
  
  Where does Stirling-PDF write its temporary files, and why does that fill your disk?
&lt;/h2&gt;

&lt;p&gt;Every non-trivial operation is a file-on-disk pipeline, not an in-memory transform. The container writes scratch files under &lt;code&gt;/tmp/stirling-pdf&lt;/code&gt;, hands them to Ghostscript, Tesseract or LibreOffice, and only then streams a result back to your browser. If you never mounted anything at that path, all of it lands on the container's writable overlay layer, on the same filesystem as your Docker root.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The multiplication nobody budgets for:&lt;/strong&gt; an OCR pass rasterises every page, writes the image set, writes an intermediate PDF, then writes the output. Peak disk use is several times the input size, and it happens before you see any progress at all.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Failed jobs do not always clean up:&lt;/strong&gt; a container killed at exit 137, or a request abandoned when you closed the tab, leaves its scratch files behind. Check with &lt;code&gt;docker exec stirling-pdf du -sh /tmp/stirling-pdf&lt;/code&gt; after a week of real use.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The overlay layer is the worst possible target:&lt;/strong&gt; it is slow, it counts against your Docker storage pool, and it disappears on &lt;code&gt;docker compose down&lt;/code&gt;, taking any recoverable partial output with it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Two fixes, different tradeoffs:&lt;/strong&gt; mount a tmpfs sized at 1 GB or 2 GB for speed, accepting that a job larger than the tmpfs fails outright, or bind mount a real directory for capacity, accepting slower rasterisation on spinning disks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cleanup is configurable:&lt;/strong&gt; the temp file management block in &lt;code&gt;/configs/settings.yml&lt;/code&gt; controls the cleanup interval, the maximum age of scratch files and whether a sweep runs at startup. Set it once rather than adding a cron job.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This matters more on constrained storage than on a roomy VPS, whether that is a NAS volume, a home server SSD or a Yundera instance.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why does OCR hang, fail outright, or return a file with no selectable text?
&lt;/h2&gt;

&lt;p&gt;OCR is the most expensive thing this application does, and it is driven by option combinations most people pick at random. Three failures dominate: no language data present, the wrong OCR mode for the input, and a page count the container cannot chew through before something upstream gives up.&lt;/p&gt;

&lt;p&gt;Language data comes first. Run &lt;code&gt;docker exec stirling-pdf tesseract --list-langs&lt;/code&gt;. If your language is missing, OCR fails or produces nonsense, and mounting a volume at &lt;code&gt;/usr/share/tessdata&lt;/code&gt; with the traineddata files you need is the fix, not a setting change.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;OCR option&lt;/th&gt;
&lt;th&gt;What it actually does&lt;/th&gt;
&lt;th&gt;Where it bites&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Force OCR&lt;/td&gt;
&lt;td&gt;Rasterises every page, then layers recognised text over the image&lt;/td&gt;
&lt;td&gt;Destroys existing vector text and searchable layers, and inflates output size on documents that were already digital&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Skip text&lt;/td&gt;
&lt;td&gt;Leaves pages that already contain a text layer untouched&lt;/td&gt;
&lt;td&gt;Returns a file with no new selectable text on mixed documents, which reads as a silent failure&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Clean and deskew preprocessing&lt;/td&gt;
&lt;td&gt;Runs image correction before recognition&lt;/td&gt;
&lt;td&gt;Adds a full extra image pass per page, so processing time rises on exactly the large scans you were already struggling with&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sidecar output&lt;/td&gt;
&lt;td&gt;Writes recognised text to a separate file alongside the PDF&lt;/td&gt;
&lt;td&gt;Doubles the scratch files per job, which matters if your temp mount is a small tmpfs&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Cost scales with pages and resolution, not with file count. A 400 page scan at 600 DPI is a fundamentally different job from 40 pages at 200 DPI, even at similar file sizes. Batch large documents into chunks of 50 pages or fewer, run them one at a time, and treat any OCR job over a few minutes as something that needs a raised timeout rather than a retry.&lt;/p&gt;




&lt;h2&gt;
  
  
  Anti-pattern 9 and 10: treating office conversion as fast, parallel and font-independent
&lt;/h2&gt;

&lt;p&gt;Converting DOCX, XLSX or PPTX to PDF does not happen in Java. Stirling-PDF hands the file to a headless LibreOffice process inside the container. That single detail explains both of this section's anti-patterns.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The first conversion after a container start is the slow one:&lt;/strong&gt; the headless &lt;code&gt;soffice&lt;/code&gt; process has to initialise before it can do any work. Measure your second conversion, not your first, or you will size the machine against a number you will never see again.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Conversions serialise, so more cores do not help:&lt;/strong&gt; requests queue behind one conversion backend. Two users submitting large presentations at the same time do not each get half the speed, the second one waits. If office conversion is your bottleneck, raise your reverse proxy timeout before you add vCPUs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A wedged &lt;code&gt;soffice&lt;/code&gt; process blocks everything behind it:&lt;/strong&gt; one malformed document can leave the backend stuck, and every later conversion times out while the rest of the application stays perfectly healthy. Check with &lt;code&gt;docker exec stirling-pdf ps aux | grep soffice&lt;/code&gt;, and restart the container rather than debugging the document.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Missing fonts corrupt layout silently:&lt;/strong&gt; the container ships a limited font set. A document written in Calibri or Cambria gets substituted glyphs, so line breaks move, tables overflow and page counts change. The output is a valid PDF that does not match the original, and nothing in the log calls this an error.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fix fonts by mounting them, not by hoping:&lt;/strong&gt; bind mount your TTF files into a directory under &lt;code&gt;/usr/share/fonts&lt;/code&gt;, then confirm with &lt;code&gt;docker exec stirling-pdf fc-list | wc -l&lt;/code&gt;. Carlito and Caladea are the metric-compatible stand-ins for Calibri and Cambria, and Liberation Sans covers Arial.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Why does a large upload fail before Stirling-PDF ever sees the file?
&lt;/h2&gt;

&lt;p&gt;A 413 error, or a progress bar that reaches 100 percent and then dies, is almost never the application. Every layer between your browser and the container can refuse a body, and each one has a different default. The container logs stay empty, which is the clue: if Stirling-PDF had rejected the file, it would have said so.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Default body limit&lt;/th&gt;
&lt;th&gt;What you see&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;nginx or nginx proxy manager&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;client_max_body_size&lt;/code&gt; defaults to 1 MB&lt;/td&gt;
&lt;td&gt;HTTP 413 within seconds, no entry in the Stirling-PDF log at all&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Caddy or Traefik&lt;/td&gt;
&lt;td&gt;No request body limit by default&lt;/td&gt;
&lt;td&gt;Nothing, these two are rarely the culprit unless you added a buffering middleware yourself&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Apache httpd&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;LimitRequestBody&lt;/code&gt; defaults to unlimited&lt;/td&gt;
&lt;td&gt;Nothing, unless a distribution config or a hardening guide set it for you&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cloudflare proxied hostname&lt;/td&gt;
&lt;td&gt;100 MB per request on the free plan&lt;/td&gt;
&lt;td&gt;HTTP 413 from Cloudflare's edge, with a Cloudflare branded error page rather than your own&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Stirling-PDF itself&lt;/td&gt;
&lt;td&gt;A multipart upload limit in &lt;code&gt;/configs/settings.yml&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;A clean application-level rejection that does appear in the log&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Diagnose it by bypassing the chain. Send the same file directly to the container port with &lt;code&gt;curl -F "fileInput=@big.pdf" http://localhost:8080/...&lt;/code&gt; from the host. If that works and the browser does not, the problem is in front of the app, and no amount of container tuning will fix it.&lt;/p&gt;

&lt;p&gt;Raise the limits in order, from the outermost layer inward, and raise them to the same value. A proxy that accepts 500 MB in front of an app that accepts 50 MB just moves the failure one hop later, after the user has already spent the upload time.&lt;/p&gt;




&lt;h2&gt;
  
  
  Anti-pattern 12: reverse proxy timeouts shorter than the job you just started
&lt;/h2&gt;

&lt;p&gt;A 504 after roughly a minute, on a job you know takes longer, is a proxy default and nothing else. The container never stopped working. It is still rasterising pages while your browser shows an error page.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;nginx gives up after 60 seconds:&lt;/strong&gt; &lt;code&gt;proxy_read_timeout&lt;/code&gt; and &lt;code&gt;proxy_send_timeout&lt;/code&gt; both default to 60s, which is shorter than almost any real OCR or large office conversion. Raise them to 300s or 600s in the location block for Stirling-PDF, not globally.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cloudflare cuts the origin off at 100 seconds:&lt;/strong&gt; a proxied hostname on the free plan returns error 524 when the origin has not responded in time, and you cannot raise that on the free plan. Long jobs behind an orange cloud need a different path to the app, such as a VPN or a direct hostname.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Caddy and Traefik do not impose a short default here:&lt;/strong&gt; if you use either and still see a timeout, look at the client, at Cloudflare, or at an explicit middleware you configured yourself.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The retry is what actually kills the container:&lt;/strong&gt; the abandoned job keeps running and keeps its memory. Clicking the button again starts a second copy of the same work, so peak memory doubles, and the container that was surviving one job gets killed at exit 137 running two. This is why "it broke worse when I tried again" is such a common report.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A timeout is not a substitute for batching:&lt;/strong&gt; raising the ceiling to 600s makes a long job possible, but a job that needs 600 seconds should be split. Confirm the real duration by timing it against the container port directly, with the proxy out of the picture, before you decide which number to set.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Anti-pattern 13: compression and repair settings that return a valid file with wrong content
&lt;/h2&gt;

&lt;p&gt;This is the failure class with no error message. The request returns 200, the PDF opens, and the damage only surfaces weeks later when someone needs the original. Compression, flattening, repair and sanitisation all rewrite the document structure, and several of them discard data by design.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Aggressive compression downsamples images, permanently:&lt;/strong&gt; the optimisation levels run from light restructuring to heavy image resampling. On a scanned invoice, a high level can drop resolution below the point where the text is legible, and there is no undo. Test one representative document at each level before you apply anything to a batch.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Targeting an output size lets the tool decide how much to destroy:&lt;/strong&gt; asking for a specific final size hands the algorithm permission to resample as far as it needs to. Set a level you have tested instead of a size you hope for.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Flattening converts form fields into static content:&lt;/strong&gt; an interactive form becomes a picture of a form. Entered data is preserved visually, but the fields are gone, and any workflow that reads field values downstream stops working.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Repair rewrites the file structure and can drop what it does not understand:&lt;/strong&gt; bookmarks, annotations, attachments and tagging are the usual casualties. Repair is for files that genuinely fail to open, not a routine step.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Any rewrite invalidates a digital signature:&lt;/strong&gt; compress, repair, flatten, sanitise or even a metadata edit breaks the cryptographic seal. If a document is signed, it leaves the pipeline untouched.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Verify instead of trusting the 200:&lt;/strong&gt; compare page counts and extracted text between input and output with &lt;code&gt;pdfinfo&lt;/code&gt; and &lt;code&gt;pdftotext&lt;/code&gt; from poppler-utils, and run &lt;code&gt;qpdf --check output.pdf&lt;/code&gt;. Keep the original file until that comparison passes.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Anti-pattern 14: volume permissions, custom fonts and language packs mounted incorrectly
&lt;/h2&gt;

&lt;p&gt;Two mount mistakes account for most of the "I followed the guide and it still does not work" reports. Both look correct in the compose file.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Mounting an empty host directory over a populated container path hides what was already there:&lt;/strong&gt; bind mount an empty folder at &lt;code&gt;/usr/share/tessdata&lt;/code&gt; and the traineddata files shipped in the image vanish, including English. OCR that worked before your fix stops working after it. Copy the existing contents out first with &lt;code&gt;docker cp stirling-pdf:/usr/share/tessdata ./tessdata&lt;/code&gt;, then mount that directory back.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The same trap applies to fonts:&lt;/strong&gt; mounting your font collection at &lt;code&gt;/usr/share/fonts&lt;/code&gt; replaces the container's entire font tree. Mount a subdirectory such as &lt;code&gt;/usr/share/fonts/truetype/custom&lt;/code&gt; so the bundled families survive alongside yours.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Docker creates missing bind mount sources as root:&lt;/strong&gt; if the host path does not exist when the container starts, Docker makes it owned by &lt;code&gt;root:root&lt;/code&gt;. An app running under a non-root user then fails to write, and configuration changes made in the UI silently do not persist across a restart.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Check the identity before you chase the permission:&lt;/strong&gt; run &lt;code&gt;docker exec stirling-pdf id&lt;/code&gt; to see which UID the process actually uses, then &lt;code&gt;docker exec stirling-pdf ls -ln /configs&lt;/code&gt; to compare it against the ownership of the mounted directory. The numbers either match or they do not, and that answers the question in five seconds.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Set &lt;code&gt;PUID&lt;/code&gt; and &lt;code&gt;PGID&lt;/code&gt; to your own account, not to 0:&lt;/strong&gt; running as root to make a permission error go away leaves every file the app writes owned by root on your host, which becomes your problem the first time you try to back up or move the data directory.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>selfhosted</category>
      <category>docker</category>
      <category>pdf</category>
    </item>
    <item>
      <title>Hubs Community Edition backup and restore: what npm run backup captures, what it misses, and how long recovery takes</title>
      <dc:creator>John</dc:creator>
      <pubDate>Sat, 22 Aug 2026 07:06:26 +0000</pubDate>
      <link>https://dev.to/john_182319291/hubs-community-edition-backup-and-restore-what-npm-run-backup-captures-what-it-misses-and-how-13of</link>
      <guid>https://dev.to/john_182319291/hubs-community-edition-backup-and-restore-what-npm-run-backup-captures-what-it-misses-and-how-13of</guid>
      <description>&lt;p&gt;The bundled scripts in Hubs Community Edition back up two things: the PostgreSQL database and the Reticulum file store that holds uploaded GLB scenes, avatars, thumbnails and room media. They do not back up your &lt;code&gt;hcce.yaml&lt;/code&gt;, your TLS material, your DNS records, your Kubernetes cluster or your object storage credentials, and if you point Hubs at an external database instead of the bundled Postgres pod, the scripts fall back to backing up the Reticulum files only. That means &lt;code&gt;npm run restore-backup&lt;/code&gt; is not a disaster recovery plan on its own: it is the last step of one, and everything before it is a rebuild you have to be able to perform from your own notes. Plan for a restore that is a fresh install plus a data load, and time it once so you know the real number instead of a hoped-for one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TL;DR by reader profile:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Small agency hosting client worlds&lt;/strong&gt; (a four person studio running six separate client instances): keep the scripts, but wrap them in per instance archives with your &lt;code&gt;hcce.yaml&lt;/code&gt; and secrets stored beside each archive, because a client asking for their world back is a per instance restore and not a cluster restore.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Solo operator running one instance&lt;/strong&gt; (a single lecturer hosting seminar rooms on one VPS): the built in backup plus a nightly filesystem snapshot of the node is enough, because your recovery target is the whole box and not a selective extraction.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Team using an external managed Postgres&lt;/strong&gt; (a studio on a hosted database with automated point in time recovery): treat the two halves separately, because the scripts stop covering the database the moment it leaves the bundled pod and your restore must line up two independent timelines.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Team with heavy scene libraries&lt;/strong&gt; (a studio with hundreds of client uploaded GLB files): budget your restore window from file store size and disk throughput, because the database restores quickly and the assets dominate the clock.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Team facing a domain or host migration&lt;/strong&gt; (a move to a new provider or a new client domain): rehearse the restore on a scratch host first, because room URLs, the Reticulum host configuration and certificate issuance all have to be corrected before anyone can join a room.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The central tradeoff: the built in scripts are simple and cover the data that actually matters, but they assume you can rebuild the surrounding installation by hand, so the effort you save at backup time is effort you pay back, under pressure, at restore time.&lt;/p&gt;




&lt;h2&gt;
  
  
  Table of contents
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;What does npm run backup actually capture in Hubs Community Edition?&lt;/li&gt;
&lt;li&gt;What does the Hubs backup script silently miss?&lt;/li&gt;
&lt;li&gt;Where does Hubs store rooms, avatars and uploaded GLB scenes on disk?&lt;/li&gt;
&lt;li&gt;How long does a full Hubs restore actually take?&lt;/li&gt;
&lt;li&gt;How do you run npm run restore-backup without breaking existing room URLs?&lt;/li&gt;
&lt;li&gt;Should you back up the Hubs Postgres database separately from the Reticulum file store?&lt;/li&gt;
&lt;li&gt;What changes when you run Hubs against an external database instead of the bundled pgsql pod?&lt;/li&gt;
&lt;li&gt;Which Hubs secrets, certificates and configuration files must be saved outside the backup archive?&lt;/li&gt;
&lt;li&gt;How much disk space does a year of Hubs backups need?&lt;/li&gt;
&lt;li&gt;Snapshots, scripts or object storage: which backup strategy fits a Hubs instance?&lt;/li&gt;
&lt;li&gt;How do you test a Hubs restore before you actually need it?&lt;/li&gt;
&lt;li&gt;What breaks when you restore Hubs onto a different domain or a different host?&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  What does npm run backup actually capture in Hubs Community Edition?
&lt;/h2&gt;

&lt;p&gt;Run &lt;code&gt;npm run backup&lt;/code&gt; from your Hubs Community Edition checkout and you get a single timestamped archive, named in the &lt;code&gt;data_backup_1234567890123&lt;/code&gt; pattern, containing two payloads and nothing else.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The PostgreSQL database, dumped from the bundled pgsql pod&lt;/strong&gt;: this is the record of every room, its title, its permissions, its owner account, its scene assignment and the metadata rows that point at uploaded files. Lose it and your GLB files still exist on disk but no longer belong to any room.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The Reticulum file store&lt;/strong&gt;: the uploaded assets themselves, meaning GLB scenes exported from Spoke, custom avatars, room thumbnails, images, audio and video that people dropped into rooms. This is the part that grows without limit and dominates the archive size.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Nothing about your cluster or your configuration&lt;/strong&gt;: &lt;code&gt;hcce.yaml&lt;/code&gt;, your Kubernetes state, your certificates and your DNS live outside the archive and are your responsibility to version elsewhere.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Nothing about the state of the services at the moment of the dump&lt;/strong&gt;: the archive is a data snapshot, not a running system image, so a restore always lands on a freshly installed instance.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The matching command is &lt;code&gt;npm run restore-backup data_backup_1234567890123&lt;/code&gt;, and calling &lt;code&gt;npm run restore-backup&lt;/code&gt; with no argument restores the most recent archive it finds. Where the instance runs shapes the rest of the plan: a self managed VPS, a home server, a NAS and Yundera are all viable homes for the surrounding stack. Yundera is a managed Personal Cloud Server, built on CasaOS, that runs self-hosted apps as Docker containers on a server dedicated to the user. Whichever you pick, the archive stays the same two payloads.&lt;/p&gt;




&lt;h2&gt;
  
  
  What does the Hubs backup script silently miss?
&lt;/h2&gt;

&lt;p&gt;The word to notice is silently. The script does not warn you about the gaps, it simply finishes, prints a path and leaves you feeling covered. Six things are outside it.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Your &lt;code&gt;hcce.yaml&lt;/code&gt;&lt;/strong&gt;: this file carries your domain, your admin email, your subdomain layout and the deployment options that make the instance yours. Restore the data onto a default configuration and you get a working Hubs that is not your Hubs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Secrets and keys&lt;/strong&gt;: the Reticulum secret key base, the Perms keypair used to sign room join tokens, the cookie signing salt and any OAuth or email credentials. Regenerating the Perms keypair invalidates tokens that clients already hold, so sessions and pending invites break.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;TLS certificates and the cert-manager state&lt;/strong&gt;: a fresh install reissues through Let's Encrypt, which means the restore is gated on DNS pointing at the new host and on rate limits, not on your archive.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The k3s cluster itself&lt;/strong&gt;: node configuration, storage class, persistent volume definitions and any manual &lt;code&gt;kubectl&lt;/code&gt; edits you made months ago and forgot.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;External object storage and CDN configuration&lt;/strong&gt;: if you moved uploads off local disk, the bucket credentials and the bucket contents are not in the archive at all.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Email, TURN and third party settings&lt;/strong&gt;: SMTP relay credentials and Coturn configuration decide whether people can receive invites and connect through restrictive networks.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Test the gap cheaply. Run &lt;code&gt;npm run backup&lt;/code&gt;, extract the archive to a scratch directory and list what came out. Everything you expected but cannot find is now an item on a second backup job, ideally a git repository holding &lt;code&gt;hcce.yaml&lt;/code&gt; and an encrypted secrets file.&lt;/p&gt;




&lt;h2&gt;
  
  
  Where does Hubs store rooms, avatars and uploaded GLB scenes on disk?
&lt;/h2&gt;

&lt;p&gt;Two places, and knowing which is which decides what you can recover selectively later.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;PostgreSQL holds the structure, not the bytes&lt;/strong&gt;: the Reticulum schema carries tables such as &lt;code&gt;hubs&lt;/code&gt; for rooms, &lt;code&gt;accounts&lt;/code&gt; for identities, &lt;code&gt;scenes&lt;/code&gt; and &lt;code&gt;avatars&lt;/code&gt; for published assets, and &lt;code&gt;owned_files&lt;/code&gt; for every upload. A row in &lt;code&gt;owned_files&lt;/code&gt; records the owner, the content type and the identifier, while the actual file sits elsewhere.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The Reticulum storage volume holds the bytes&lt;/strong&gt;: GLB scenes exported from Spoke, avatar GLBs, room thumbnails, images, PDFs, audio and video are written into a fanned out directory tree, each upload stored as a data file paired with a small metadata file. The names are identifiers, not human readable titles, which is why a file only means something when the database row that points at it survives too.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Derived media is cached, not authored&lt;/strong&gt;: image resizing and video transcoding produce output that can be regenerated, so it is worth excluding from a tight recovery target and rebuilding after the fact.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The volume is a Kubernetes persistent volume claim&lt;/strong&gt;: run &lt;code&gt;kubectl get pvc&lt;/code&gt; in your Hubs namespace to see the claim and its size, and &lt;code&gt;kubectl exec&lt;/code&gt; into the Reticulum pod to measure the tree with &lt;code&gt;du -sh&lt;/code&gt; before you plan any window.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Where that volume physically lives depends on the host you chose, whether a self managed VPS, a home server, a NAS or Yundera. The split matters because the database is small and fast to move, while the storage volume is the part that grows every time a client uploads a scene.&lt;/p&gt;




&lt;h2&gt;
  
  
  How long does a full Hubs restore actually take?
&lt;/h2&gt;

&lt;p&gt;Nobody can hand you a single number, because the archive is dominated by whatever your clients uploaded. What you can do is decompose the clock into five stages, measure each once on your own hardware, and turn the total into a promise you can actually keep.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Restore stage&lt;/th&gt;
&lt;th&gt;What dominates the clock&lt;/th&gt;
&lt;th&gt;How to shorten it&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Provision the host and install k3s&lt;/td&gt;
&lt;td&gt;Provider provisioning time plus package downloads&lt;/td&gt;
&lt;td&gt;Keep a prebuilt image or a scripted node setup instead of typing it fresh&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Deploy Hubs from &lt;code&gt;hcce.yaml&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Container image pulls for Reticulum, Dialog, the client and Postgres&lt;/td&gt;
&lt;td&gt;Pre pull images on the standby host, or host them in a local registry&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Issue TLS and point DNS&lt;/td&gt;
&lt;td&gt;DNS propagation plus Let's Encrypt issuance, both outside your control&lt;/td&gt;
&lt;td&gt;Lower the record TTL before a planned migration, not during an incident&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Load the database dump&lt;/td&gt;
&lt;td&gt;Row count in &lt;code&gt;hubs&lt;/code&gt;, &lt;code&gt;accounts&lt;/code&gt; and &lt;code&gt;owned_files&lt;/code&gt;, usually the smallest stage&lt;/td&gt;
&lt;td&gt;Nothing needed, this is rarely the bottleneck&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Copy back the Reticulum file store&lt;/td&gt;
&lt;td&gt;Archive size divided by disk and network throughput, the largest stage by far&lt;/td&gt;
&lt;td&gt;Restore assets from storage that is already close to the new host&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Measure it with &lt;code&gt;time npm run restore-backup data_backup_1234567890123&lt;/code&gt; on a scratch instance, then add the stages the script does not cover. Two rules follow. First, the file store sets your recovery time objective, so track its growth with &lt;code&gt;du -sh&lt;/code&gt; monthly. Second, quote clients a window built from a real rehearsal, not from the script's runtime alone.&lt;/p&gt;




&lt;h2&gt;
  
  
  How do you run npm run restore-backup without breaking existing room URLs?
&lt;/h2&gt;

&lt;p&gt;A Hubs room URL is built from the room identifier stored in the &lt;code&gt;hubs&lt;/code&gt; table, so the URLs survive a restore for free as long as the database rows come back unchanged and the domain in front of them stays the same. What breaks them is everything you do around the restore.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Restore onto a fresh install, never onto a used one&lt;/strong&gt;: bring up Hubs from &lt;code&gt;hcce.yaml&lt;/code&gt;, then restore before anyone creates a room or publishes a scene. New activity writes rows that the incoming dump will collide with or overwrite, and the reconciliation is manual.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep the domain identical at restore time&lt;/strong&gt;: put the same value in &lt;code&gt;hcce.yaml&lt;/code&gt; that the archive was taken under, restore, verify, and only then perform a domain change as a separate, deliberate step.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Quiesce the services first&lt;/strong&gt;: scale the Reticulum deployment to zero with &lt;code&gt;kubectl scale&lt;/code&gt;, run the restore, then scale back. Restoring underneath a running Reticulum means live processes hold state that no longer matches the database.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Name the archive explicitly&lt;/strong&gt;: run &lt;code&gt;npm run restore-backup data_backup_1234567890123&lt;/code&gt; rather than the bare &lt;code&gt;npm run restore-backup&lt;/code&gt;, because the no argument form picks the most recent archive it finds and that is rarely what you want during an incident.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Verify before announcing&lt;/strong&gt;: open one known room URL, one published scene and one custom avatar. Three checks catch the common failure, which is a database restored without its matching files.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The host underneath, a self managed VPS, a home server, a NAS or Yundera, does not change this sequence, only how quickly you can stand up the replacement instance.&lt;/p&gt;




&lt;h2&gt;
  
  
  Should you back up the Hubs Postgres database separately from the Reticulum file store?
&lt;/h2&gt;

&lt;p&gt;For a single instance you can run &lt;code&gt;npm run backup&lt;/code&gt; on a schedule and stop thinking about it. For anything larger, split the two, because they have opposite shapes: the database is small and changes every time someone creates a room, while the file store is large and mostly append only.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;One combined archive&lt;/th&gt;
&lt;th&gt;Database and files backed up separately&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Sensible cadence&lt;/td&gt;
&lt;td&gt;Whatever the slowest half tolerates, usually nightly&lt;/td&gt;
&lt;td&gt;Database hourly, files daily or on change&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Storage growth&lt;/td&gt;
&lt;td&gt;Every run copies the full file store again&lt;/td&gt;
&lt;td&gt;Files deduplicate through incremental tooling, dumps stay small&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tooling&lt;/td&gt;
&lt;td&gt;The bundled scripts only&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;pg_dump&lt;/code&gt; plus restic, borg or rsync against the storage volume&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Recovery granularity&lt;/td&gt;
&lt;td&gt;All or nothing, back to the archive time&lt;/td&gt;
&lt;td&gt;Roll the database to one point, the files to another&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Operational cost&lt;/td&gt;
&lt;td&gt;One command, one cron entry, one thing to forget&lt;/td&gt;
&lt;td&gt;Two jobs, two retention policies, two things to monitor&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The practical rule is the loss window. If your clients publish scenes weekly but create rooms daily, an archive that repeats 40 GB of unchanged GLB files every night buys you nothing except a bigger bill. Dump the database hourly, back the storage volume up incrementally, and keep one full combined archive weekly as the known good fallback that needs no tooling to read.&lt;/p&gt;

&lt;p&gt;Split backups carry one obligation: always restore the file store first and the database second, so no row can reference an asset that is not on disk yet. Write that ordering into the runbook, because it is the step people improvise wrongly under pressure.&lt;/p&gt;




&lt;h2&gt;
  
  
  What changes when you run Hubs against an external database instead of the bundled pgsql pod?
&lt;/h2&gt;

&lt;p&gt;One thing changes and it is the important one: the backup scripts stop covering half your data. With an external database configured, &lt;code&gt;npm run backup&lt;/code&gt; and &lt;code&gt;npm run restore-backup&lt;/code&gt; handle the Reticulum files only. Nobody prints a warning. You simply get a smaller archive.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;You now own the database backup outright&lt;/strong&gt;: either your provider's automated snapshots and point in time recovery, or your own scheduled &lt;code&gt;pg_dump&lt;/code&gt; against the Hubs database. Whichever you pick, it belongs in a runbook next to the file archive, not in your head.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You have two timelines to reconcile&lt;/strong&gt;: a database recovered to 02:00 and a file archive taken at 03:00 leaves an hour of &lt;code&gt;owned_files&lt;/code&gt; rows pointing at assets the file store does not have. Restore the files first, then roll the database forward to a point at or before the file archive, never after.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Credentials become a recovery dependency&lt;/strong&gt;: the connection string, user and password live in your &lt;code&gt;hcce.yaml&lt;/code&gt; and secrets, so losing those means the data survives while access to it does not.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Migrations still run on boot&lt;/strong&gt;: Reticulum applies its schema migrations when it starts, so a database restored from an older instance can be upgraded by a newer deployment without you asking. Restore onto the same version you backed up from and upgrade deliberately afterwards.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The bundled pod stays simpler for small instances&lt;/strong&gt;: one archive, one timeline, one restore command, at the cost of the operational features a managed database gives you.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Choose the split deliberately, then write down which system owns which half.&lt;/p&gt;




&lt;h2&gt;
  
  
  Which Hubs secrets, certificates and configuration files must be saved outside the backup archive?
&lt;/h2&gt;

&lt;p&gt;Sort every item into two buckets before you build the vault: things that can be regenerated on the new host, and things that change behaviour for users if they are regenerated. Only the second bucket is a real backup obligation.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Item&lt;/th&gt;
&lt;th&gt;Preserve or regenerate&lt;/th&gt;
&lt;th&gt;Where it belongs&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;hcce.yaml&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Preserve, it defines your domain, subdomains and admin email&lt;/td&gt;
&lt;td&gt;A private git repository, committed on every change&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Perms keypair and Reticulum secret key base&lt;/td&gt;
&lt;td&gt;Preserve, regenerating invalidates issued room tokens and signed cookies&lt;/td&gt;
&lt;td&gt;Encrypted secrets file, sops with age or your password manager&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;External database credentials&lt;/td&gt;
&lt;td&gt;Preserve, the data is useless without access to it&lt;/td&gt;
&lt;td&gt;Same encrypted store as the keys, never in the same archive as the dump&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SMTP, OAuth and Coturn credentials&lt;/td&gt;
&lt;td&gt;Preserve, they come from third parties and cannot be recreated locally&lt;/td&gt;
&lt;td&gt;Encrypted store, with the provider account noted beside them&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;TLS certificates&lt;/td&gt;
&lt;td&gt;Regenerate, cert-manager reissues after DNS points at the new host&lt;/td&gt;
&lt;td&gt;Nothing to keep, but record the issuer and DNS provider steps&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Derived Kubernetes secrets&lt;/td&gt;
&lt;td&gt;Regenerate from the two entries above during deployment&lt;/td&gt;
&lt;td&gt;Nothing to keep, they are outputs&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Capture the live state once with &lt;code&gt;kubectl get secret -n hubs -o yaml&lt;/code&gt; so you can see exactly what your instance holds rather than trusting this list. Encrypt that export immediately, because it contains plaintext values.&lt;/p&gt;

&lt;p&gt;One discipline makes the whole thing work: the encrypted secrets store must be recoverable without the Hubs instance. If your only copy of the vault key lives on the server you are restoring, you do not have a backup, you have a circular dependency.&lt;/p&gt;




&lt;h2&gt;
  
  
  How much disk space does a year of Hubs backups need?
&lt;/h2&gt;

&lt;p&gt;Work it out from two measurements rather than a guess. Get the current storage volume size with &lt;code&gt;du -sh&lt;/code&gt; inside the Reticulum pod, then take two archives a month apart and subtract to get your monthly growth. Everything else is arithmetic on those numbers.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Full copies multiply, so count them first&lt;/strong&gt;: a grandfather, father, son policy of 7 daily, 4 weekly and 12 monthly archives is 23 retained copies. With the bundled script, each one contains the whole file store, so the year costs roughly 23 times your current storage size plus the accumulated growth. This is the number that surprises agencies.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Incremental tooling changes the multiplier, not the base&lt;/strong&gt;: because GLB scenes and avatars are written once and rarely modified, deduplicating backups with restic or borg store close to one copy plus the new uploads, so 23 restore points cost far less than 23 full archives.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Database dumps are rounding error&lt;/strong&gt;: &lt;code&gt;pg_dump&lt;/code&gt; output for room, account and &lt;code&gt;owned_files&lt;/code&gt; rows compresses well and stays small next to the assets, which is why hourly database backups are affordable while hourly file archives are not.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Derived media inflates the total without adding value&lt;/strong&gt;: transcoded video and resized images can be regenerated, so excluding them from the archive shrinks every copy for the whole year.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Off site copies double whatever you decided&lt;/strong&gt;: the 3-2-1 rule means one primary, one local secondary and one remote, so budget the total twice, not once.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Set retention against a real number. A client that publishes a scene weekly needs different depth than one that uploaded once and never returned.&lt;/p&gt;




&lt;h2&gt;
  
  
  Snapshots, scripts or object storage: which backup strategy fits a Hubs instance?
&lt;/h2&gt;

&lt;p&gt;Three approaches exist and they solve different failures. Pick by asking what you expect to lose: the whole machine, the data inside it, or the provider account holding both.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Provider or filesystem snapshots of the node&lt;/strong&gt;: these capture the entire k3s host, so recovery is a rollback rather than a rebuild, and your &lt;code&gt;hcce.yaml&lt;/code&gt;, secrets and certificates come back with it. They fail you in two ways: they usually live in the same provider account as the server, and they restore everything or nothing, which is useless when one client wants one room recovered.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The bundled scripts alone&lt;/strong&gt;: &lt;code&gt;npm run backup&lt;/code&gt; is portable, understandable and independent of any provider feature, which makes it the archive you can still read in three years. It covers the two data payloads and nothing around them, so it needs a companion job for configuration and keys.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Incremental backups to object storage&lt;/strong&gt;: restic or borg pushing the Reticulum storage volume and &lt;code&gt;pg_dump&lt;/code&gt; output to an S3 compatible target such as MinIO or Backblaze B2 gives deduplication, encryption at rest and an off site copy in one tool. The cost is a second system to monitor, plus a restore that now depends on network throughput.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The hybrid most agencies land on&lt;/strong&gt;: nightly snapshots for fast whole host rollback, weekly bundled archives as the provider independent fallback, and continuous incremental pushes off site for the 3-2-1 requirement.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The choice is not about which tool is best. It is about which of the three failures you consider most likely, and whether your recovery needs to be selective.&lt;/p&gt;




&lt;h2&gt;
  
  
  How do you test a Hubs restore before you actually need it?
&lt;/h2&gt;

&lt;p&gt;Build a scratch instance on a throwaway subdomain, restore last night's archive into it, and check five things by hand. Repeat quarterly and after every version upgrade, because an untested archive is a hypothesis.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Use a separate domain, never the live one&lt;/strong&gt;: set a value such as &lt;code&gt;hubs-drtest.example.com&lt;/code&gt; in the scratch &lt;code&gt;hcce.yaml&lt;/code&gt;. Restoring production data onto production DNS during a drill is how a test becomes an outage.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Join a room and check presence works&lt;/strong&gt;: open a restored room URL in two browsers and confirm audio and video connect. This exercises Dialog and Coturn, which a database check alone never touches.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Load a published GLB scene, not just the room list&lt;/strong&gt;: rooms that open with a missing environment are the signature of a database restored without its matching file store, and the room index looks perfectly healthy in that state.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Load a custom avatar and a room thumbnail&lt;/strong&gt;: avatars and thumbnails are separate &lt;code&gt;owned_files&lt;/code&gt; rows, so they fail independently of scenes and prove the file tree copied completely rather than partially.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Time each stage and write the number down&lt;/strong&gt;: run the drill with a stopwatch, record the total, and give clients that figure rather than an estimate. The number moves as the file store grows.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Finish by destroying the scratch instance and any DNS records it created, so nobody stumbles into stale rooms months later. Keep a one page log of every drill: the date, the archive name, the elapsed time and what broke. Three entries in that log tell you more about your recovery position than any amount of configuration review.&lt;/p&gt;




&lt;h2&gt;
  
  
  What breaks when you restore Hubs onto a different domain or a different host?
&lt;/h2&gt;

&lt;p&gt;A different host alone is uneventful: the archive is portable and the deployment is rebuilt from &lt;code&gt;hcce.yaml&lt;/code&gt;, so a move between a self managed VPS, a home server, a NAS or Yundera changes little beyond throughput. A different domain is the one that hurts, because the old hostname is written into places the restore script does not touch.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Stored absolute URLs&lt;/strong&gt;: rows created under the old domain can carry fully qualified asset and scene URLs. After a domain change, query the database with &lt;code&gt;psql&lt;/code&gt; for the old hostname before you assume the restore was clean, and update what you find.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Links your clients already shared&lt;/strong&gt;: room URLs printed in emails, calendar invites and slide decks point at the old host. Keep the old domain resolving with a redirect for at least one full booking cycle rather than a fixed number of days.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OAuth and email&lt;/strong&gt;: third party callback URLs and the sending domain in your SMTP configuration are registered against the old hostname and reject the new one until you update both providers.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;TURN and media&lt;/strong&gt;: Dialog and Coturn advertise addresses derived from your deployment configuration, so audio and video can fail while the room list looks perfectly healthy.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Certificate issuance order&lt;/strong&gt;: DNS must point at the new host before cert-manager can complete a challenge, so the sequence is DNS first, deploy second, restore third.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Do the domain change as its own step, after a same domain restore has been verified. Two changes at once means any failure gives you two suspects instead of one.&lt;/p&gt;

</description>
      <category>hubs</category>
      <category>selfhosting</category>
      <category>devops</category>
      <category>backup</category>
    </item>
  </channel>
</rss>
