<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: EniyiSunucum</title>
    <description>The latest articles on DEV Community by EniyiSunucum (@eniyisunucum).</description>
    <link>https://dev.to/eniyisunucum</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4128830%2Fffb3d3f8-e637-4ae5-80f0-abce55799033.png</url>
      <title>DEV Community: EniyiSunucum</title>
      <link>https://dev.to/eniyisunucum</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/eniyisunucum"/>
    <language>en</language>
    <item>
      <title>Stop Hosting Client WordPress Sites in One Shared Account</title>
      <dc:creator>EniyiSunucum</dc:creator>
      <pubDate>Sat, 19 Sep 2026 17:20:59 +0000</pubDate>
      <link>https://dev.to/eniyisunucum/stop-hosting-client-wordpress-sites-in-one-shared-account-385b</link>
      <guid>https://dev.to/eniyisunucum/stop-hosting-client-wordpress-sites-in-one-shared-account-385b</guid>
      <description>&lt;p&gt;Hosting ten client sites is not the same problem as hosting one site ten times. The difference is not storage capacity or the number of domains a control panel accepts. It is the size of the failure domain: which clients are affected when one password leaks, one plugin is compromised, one PHP worker pool is exhausted, or one backup cannot be restored.&lt;/p&gt;

&lt;p&gt;Disclosure: This technical guide was prepared with AI assistance and reviewed for accuracy and clarity by the EniyiSunucum team.&lt;/p&gt;

&lt;p&gt;Small agencies often begin with several WordPress installations inside one shared hosting account. That can be convenient while the portfolio is small, but the model becomes harder to defend as sites accumulate revenue, personal data, email, and contractual uptime expectations. A safer design gives each client a clear security boundary, a measurable resource budget, and an independently testable recovery path.&lt;/p&gt;

&lt;p&gt;This guide explains how to build that operating model without assuming that every agency needs to manage a fleet of virtual machines.&lt;/p&gt;

&lt;h2&gt;
  
  
  Begin with the failure domain
&lt;/h2&gt;

&lt;p&gt;A failure domain is the set of systems that can be disrupted by one event. If every client site uses the same hosting login, filesystem owner, backup job, and PHP limits, the account itself is the failure domain.&lt;/p&gt;

&lt;p&gt;That design creates several forms of coupling:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A compromised administrator credential can expose every site.&lt;/li&gt;
&lt;li&gt;A vulnerable plugin may be able to read or modify neighboring files.&lt;/li&gt;
&lt;li&gt;One traffic spike can consume the account's shared workers or I/O allowance.&lt;/li&gt;
&lt;li&gt;A full disk can interrupt unrelated databases, logs, sessions, and email.&lt;/li&gt;
&lt;li&gt;Restoring one client may require changing a backup that contains all clients.&lt;/li&gt;
&lt;li&gt;Giving a contractor access to one project may expose the rest of the portfolio.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The first architectural question is therefore not “How many websites fit?” It is “Which clients can fail together, and why?”&lt;/p&gt;

&lt;p&gt;For most agencies, the practical target is one hosting account per client or per independently managed project. This does not guarantee perfect isolation—the provider and platform still matter—but it creates a much cleaner administrative, filesystem, credential, and recovery boundary than placing every site under one account.&lt;/p&gt;

&lt;h2&gt;
  
  
  Give every client an explicit identity boundary
&lt;/h2&gt;

&lt;p&gt;Each client environment should have its own control-panel account, system user where the platform exposes one, SFTP credentials, databases, and application secrets. Shared administrator credentials should be the exception, not the workflow.&lt;/p&gt;

&lt;p&gt;A useful minimum standard is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a separate hosting account for each client;&lt;/li&gt;
&lt;li&gt;a unique WordPress administrator identity for each human operator;&lt;/li&gt;
&lt;li&gt;a unique database user and password for each site;&lt;/li&gt;
&lt;li&gt;separate SFTP or deployment credentials;&lt;/li&gt;
&lt;li&gt;no password reuse between production, staging, email, and the registrar;&lt;/li&gt;
&lt;li&gt;multi-factor authentication wherever the control plane supports it;&lt;/li&gt;
&lt;li&gt;an access-removal checklist for staff and contractors.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This model improves incident response. If credentials for one customer are exposed, the agency can rotate that customer's secrets without scheduling a portfolio-wide outage. Audit trails also become more meaningful because actions can be attributed to a person or a project instead of a shared login.&lt;/p&gt;

&lt;h2&gt;
  
  
  Treat PHP capacity as a queueing problem
&lt;/h2&gt;

&lt;p&gt;WordPress performance is frequently described as a CPU or cache problem. Under concurrency, it is often a queueing problem.&lt;/p&gt;

&lt;p&gt;Every uncached dynamic request occupies a PHP worker until application code, database queries, remote API calls, and filesystem operations finish. When all workers are busy, additional requests wait. A site can therefore report modest average CPU utilization while users experience long tail latency because the worker pool is saturated or blocked.&lt;/p&gt;

&lt;p&gt;Track at least these signals per site:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;concurrent PHP workers or entry processes;&lt;/li&gt;
&lt;li&gt;request rate and cache-hit ratio;&lt;/li&gt;
&lt;li&gt;p50, p95, and p99 response time;&lt;/li&gt;
&lt;li&gt;slow PHP transactions;&lt;/li&gt;
&lt;li&gt;database query latency and connection pressure;&lt;/li&gt;
&lt;li&gt;CPU throttling, memory pressure, and I/O wait;&lt;/li&gt;
&lt;li&gt;the duration and frequency of WordPress cron jobs;&lt;/li&gt;
&lt;li&gt;external service timeouts.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Avoid sizing from page views alone. Two sites with the same monthly traffic can have radically different resource profiles. A mostly cached publication may serve large volumes efficiently, while a membership site with authenticated sessions, searches, checkout requests, or personalized dashboards can generate dynamic work on nearly every request.&lt;/p&gt;

&lt;p&gt;When comparing a &lt;a href="https://eniyisunucum.com/en/wordpress-hosting" rel="noopener noreferrer"&gt;WordPress hosting environment&lt;/a&gt;, evaluate PHP limits, web-server behavior, database access, backup tooling, and observability together. A long feature list is less useful than knowing how the platform behaves when the site's uncached concurrency rises.&lt;/p&gt;

&lt;h2&gt;
  
  
  Allocate a resource budget per site
&lt;/h2&gt;

&lt;p&gt;An account boundary is operationally useful only when resource consumption can also be observed and controlled. Define a budget for each client based on normal load, peak load, deployment behavior, and business importance.&lt;/p&gt;

&lt;p&gt;The budget may include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;CPU time or virtual cores;&lt;/li&gt;
&lt;li&gt;memory;&lt;/li&gt;
&lt;li&gt;concurrent processes;&lt;/li&gt;
&lt;li&gt;PHP workers or entry processes;&lt;/li&gt;
&lt;li&gt;disk throughput and IOPS;&lt;/li&gt;
&lt;li&gt;database size and query volume;&lt;/li&gt;
&lt;li&gt;storage growth;&lt;/li&gt;
&lt;li&gt;backup size and execution window;&lt;/li&gt;
&lt;li&gt;outbound email or API activity.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The objective is not to keep every graph low. It is to prevent one tenant from silently consuming the headroom needed by others. Record a healthy baseline and define a trigger for investigation—for example, repeated worker exhaustion, sustained memory pressure, a backup window that overlaps business traffic, or a p95 response time outside the client's service objective.&lt;/p&gt;

&lt;p&gt;Resource limits should produce an escalation path, not a surprise. Decide in advance whether a constrained site will be optimized, moved to a higher account tier, separated onto a VPS or VDS, or redesigned to reduce synchronous work.&lt;/p&gt;

&lt;h2&gt;
  
  
  Do not confuse WordPress Multisite with hosting isolation
&lt;/h2&gt;

&lt;p&gt;WordPress Multisite is an application feature for managing a network of related sites. It is not a substitute for isolating unrelated customers.&lt;/p&gt;

&lt;p&gt;Sites in a Multisite network share a WordPress core, a database structure, plugins, themes, and a super-administrator boundary. That can be appropriate for one organization operating regional sites, a university managing departments, or a product with centrally governed tenants. It is usually a poor boundary for unrelated agency clients who need independent plugin policies, maintenance windows, credentials, billing, exports, and exit plans.&lt;/p&gt;

&lt;p&gt;Use Multisite when shared governance is intentional. Use separate client environments when ownership and operational responsibility must remain distinct.&lt;/p&gt;

&lt;h2&gt;
  
  
  Make restoration the unit of backup quality
&lt;/h2&gt;

&lt;p&gt;A backup is not proven because a dashboard displays a green check. It is proven when the team can restore the required data within the promised recovery time.&lt;/p&gt;

&lt;p&gt;Define two values for every client:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Recovery Point Objective (RPO): how much recent data can be lost.&lt;/li&gt;
&lt;li&gt;Recovery Time Objective (RTO): how long the service can remain unavailable.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then verify that the backup design can meet them. Check database and filesystem consistency, retention, encryption, storage location, access controls, and whether the backup system is independent of the production account. A backup that an attacker can delete with the same compromised credential is a weak last line of defense.&lt;/p&gt;

&lt;p&gt;Run scheduled restore tests to a separate environment. The test should verify more than the home page:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;restore the database and files;&lt;/li&gt;
&lt;li&gt;replace environment-specific URLs or secrets safely;&lt;/li&gt;
&lt;li&gt;confirm administrator access;&lt;/li&gt;
&lt;li&gt;test forms, checkout, search, and authenticated workflows;&lt;/li&gt;
&lt;li&gt;inspect media and generated files;&lt;/li&gt;
&lt;li&gt;validate scheduled jobs and outbound integrations;&lt;/li&gt;
&lt;li&gt;record the elapsed time and any manual steps.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The result becomes evidence for capacity planning. If restoring a growing media library already exceeds the client's RTO, the agency has discovered the problem before an incident.&lt;/p&gt;

&lt;h2&gt;
  
  
  Separate staging from production
&lt;/h2&gt;

&lt;p&gt;Staging should reduce deployment risk without becoming another path into production.&lt;/p&gt;

&lt;p&gt;Use separate credentials and, where possible, a separate account or isolated environment. Do not copy live customer data into staging by default. When realistic data is required, remove or mask personal information and payment-related fields. Disable public indexing, transactional email, payment callbacks, analytics pollution, and scheduled jobs that should run only in production.&lt;/p&gt;

&lt;p&gt;A disciplined deployment path is small and repeatable:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;source-controlled themes, plugins, and configuration;&lt;/li&gt;
&lt;li&gt;documented database migrations;&lt;/li&gt;
&lt;li&gt;a pre-deployment backup or snapshot appropriate to the change;&lt;/li&gt;
&lt;li&gt;a health check after deployment;&lt;/li&gt;
&lt;li&gt;a rollback decision with a clear time limit.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Avoid using production as the first place to discover whether a plugin update changes the database, PHP requirement, cron schedule, or cache behavior.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep DNS, email, and certificates outside the blast radius
&lt;/h2&gt;

&lt;p&gt;Website recovery frequently fails because the team focuses only on WordPress files and the database.&lt;/p&gt;

&lt;p&gt;Document ownership and access for the domain registrar, DNS provider, certificate automation, transactional email service, and mailbox provider. Use role-based access where possible. The agency should know which assets belong to the client, which are managed on the client's behalf, and how control will be transferred at the end of the relationship.&lt;/p&gt;

&lt;p&gt;Separating authoritative DNS and domain registration from the hosting account can preserve a recovery route when the hosting control plane is unavailable. Likewise, keeping critical operational email independent of the affected site helps the team receive alerts and coordinate during an outage.&lt;/p&gt;

&lt;h2&gt;
  
  
  Choose the operating model deliberately
&lt;/h2&gt;

&lt;p&gt;Agencies generally choose among three models.&lt;/p&gt;

&lt;h3&gt;
  
  
  Individual retail hosting accounts
&lt;/h3&gt;

&lt;p&gt;This offers strong commercial separation but can create administrative overhead as the portfolio grows. It works well when clients pay providers directly or require ownership of their own account from day one.&lt;/p&gt;

&lt;h3&gt;
  
  
  Reseller hosting
&lt;/h3&gt;

&lt;p&gt;A reseller platform can give an agency one administrative view while preserving separate client accounts. It is a sensible middle ground when the agency wants account-level isolation and standard hosting operations without managing the operating system.&lt;/p&gt;

&lt;p&gt;When evaluating a &lt;a href="https://eniyisunucum.com/en/reseller-hosting" rel="noopener noreferrer"&gt;reseller hosting model&lt;/a&gt;, compare the number of accounts, per-site resource limits, backup and restore workflow, delegation controls, supported runtimes, and the provider's escalation process. “Unlimited” storage or traffic should never be interpreted as unlimited CPU, memory, workers, or operational support.&lt;/p&gt;

&lt;h3&gt;
  
  
  VPS or VDS
&lt;/h3&gt;

&lt;p&gt;A virtual server provides more control over the web stack, monitoring, networking, and automation. It also transfers more responsibility to the agency: operating-system updates, security hardening, backups, service monitoring, capacity planning, and recovery testing. It is appropriate when the team can operate that layer reliably or has a managed service covering it.&lt;/p&gt;

&lt;p&gt;The most expensive model is not automatically the safest. The correct choice is the least complex platform that still meets the required isolation, observability, and recovery objectives.&lt;/p&gt;

&lt;h2&gt;
  
  
  Build an exit path for every client
&lt;/h2&gt;

&lt;p&gt;Client portability is part of good infrastructure design. Maintain an inventory of domains, applications, databases, mailboxes, certificates, third-party integrations, licenses, and backup locations. Record who owns each asset and which credentials are needed for transfer.&lt;/p&gt;

&lt;p&gt;A clean handover package can include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a current filesystem archive;&lt;/li&gt;
&lt;li&gt;a consistent database export;&lt;/li&gt;
&lt;li&gt;DNS records and TTL values;&lt;/li&gt;
&lt;li&gt;runtime and PHP extension requirements;&lt;/li&gt;
&lt;li&gt;scheduled tasks;&lt;/li&gt;
&lt;li&gt;external service dependencies;&lt;/li&gt;
&lt;li&gt;license ownership;&lt;/li&gt;
&lt;li&gt;a restore or migration runbook;&lt;/li&gt;
&lt;li&gt;confirmation that agency-only credentials were removed.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If a site cannot be exported and restored without tribal knowledge, the environment is not yet operationally mature.&lt;/p&gt;

&lt;h2&gt;
  
  
  A practical migration sequence
&lt;/h2&gt;

&lt;p&gt;An agency moving away from one large shared account should avoid a single portfolio-wide cutover. Migrate in controlled batches.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Inventory every site, domain, database, mailbox, job, and integration.&lt;/li&gt;
&lt;li&gt;Rank clients by business impact, technical complexity, and current risk.&lt;/li&gt;
&lt;li&gt;Establish separate destination accounts and unique credentials.&lt;/li&gt;
&lt;li&gt;Test backup restoration before changing DNS.&lt;/li&gt;
&lt;li&gt;Lower DNS TTL only when the migration schedule justifies it.&lt;/li&gt;
&lt;li&gt;Synchronize files and databases using a documented cutover method.&lt;/li&gt;
&lt;li&gt;Validate dynamic workflows, not only cached pages.&lt;/li&gt;
&lt;li&gt;Monitor errors, latency, email, and background jobs after cutover.&lt;/li&gt;
&lt;li&gt;Keep the old environment read-only for an agreed rollback window.&lt;/li&gt;
&lt;li&gt;Revoke obsolete credentials and securely retire residual data.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Start with a representative low-risk client, improve the runbook, and then move more critical sites. The goal is repeatability, not speed on the first migration.&lt;/p&gt;

&lt;h2&gt;
  
  
  Measure the result
&lt;/h2&gt;

&lt;p&gt;The new architecture is successful when incidents become smaller, diagnosis becomes faster, and recovery becomes predictable.&lt;/p&gt;

&lt;p&gt;Useful portfolio-level measures include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;percentage of clients in separate accounts;&lt;/li&gt;
&lt;li&gt;percentage with tested restores in the last quarter;&lt;/li&gt;
&lt;li&gt;median and worst-case restore time;&lt;/li&gt;
&lt;li&gt;number of shared credentials;&lt;/li&gt;
&lt;li&gt;sites repeatedly exceeding resource budgets;&lt;/li&gt;
&lt;li&gt;patch and update lead time;&lt;/li&gt;
&lt;li&gt;failed backup or cron jobs;&lt;/li&gt;
&lt;li&gt;incidents that crossed a client boundary.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These metrics turn “better hosting” into an operating system the agency can review and improve.&lt;/p&gt;

&lt;h2&gt;
  
  
  Isolation is a business feature
&lt;/h2&gt;

&lt;p&gt;Separating client sites is not only a security improvement. It clarifies ownership, limits incident scope, makes billing and capacity visible, simplifies access changes, and gives every customer a credible recovery story.&lt;/p&gt;

&lt;p&gt;Begin with the failure domain. Give each client an identity and resource boundary. Test restoration instead of trusting backup status. Keep staging, DNS, email, and credentials out of the production blast radius. Then select individual hosting, reseller hosting, or a managed virtual server according to the team's real operating capability.&lt;/p&gt;

&lt;p&gt;That approach scales an agency more safely than counting how many domains still fit inside one account.&lt;/p&gt;

</description>
      <category>wordpress</category>
      <category>webdev</category>
      <category>security</category>
      <category>devops</category>
    </item>
    <item>
      <title>Benchmarking a Ryzen VDS Without Fooling Yourself: CPU Scheduling, Storage Latency, and Reproducible Tests</title>
      <dc:creator>EniyiSunucum</dc:creator>
      <pubDate>Sat, 19 Sep 2026 16:33:17 +0000</pubDate>
      <link>https://dev.to/eniyisunucum/benchmarking-a-ryzen-vds-without-fooling-yourself-cpu-scheduling-storage-latency-and-1jem</link>
      <guid>https://dev.to/eniyisunucum/benchmarking-a-ryzen-vds-without-fooling-yourself-cpu-scheduling-storage-latency-and-1jem</guid>
      <description>&lt;p&gt;A virtual server benchmark is easy to run and surprisingly hard to interpret. A single attractive score can be caused by a short turbo window, a warm cache, an idle host, or a storage queue that does not resemble the application. A disappointing score can be equally misleading: perhaps another guest briefly competed for CPU time, package updates were running, or the test measured throughput when the workload actually depends on latency.&lt;/p&gt;

&lt;p&gt;This article describes a small, repeatable method for evaluating a Linux VDS on KVM. It is not a ranking system and it does not produce one universal “performance number.” The goal is to separate three questions:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Is the guest receiving predictable CPU time?&lt;/li&gt;
&lt;li&gt;How does one guest vCPU behave, and how well do several vCPUs scale?&lt;/li&gt;
&lt;li&gt;What latency does storage deliver at both low and moderate queue depth?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The method deliberately favors evidence that can be retained and reviewed: metadata, warm-up rules, repeated runs, latency percentiles, and raw output.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with a test contract
&lt;/h2&gt;

&lt;p&gt;Before installing a benchmark, write down what decision the result must support. A web worker, a compilation runner, and a nightly database report stress different parts of a system. A &lt;a href="https://eniyisunucum.com/en/ryzen-vds-server" rel="noopener noreferrer"&gt;Ryzen VDS configuration table&lt;/a&gt; can serve as one example of the source profile: record the listed vCPU count, memory, and storage allocation as test inputs, not as evidence of the outcome. Then add the workload’s runtime, concurrency, and I/O pattern before choosing any synthetic test.&lt;/p&gt;

&lt;p&gt;Record the environment at the beginning of every test session:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;date&lt;/span&gt; &lt;span class="nt"&gt;-u&lt;/span&gt;
&lt;span class="nb"&gt;uname&lt;/span&gt; &lt;span class="nt"&gt;-a&lt;/span&gt;
lscpu
&lt;span class="nb"&gt;nproc
&lt;/span&gt;free &lt;span class="nt"&gt;-h&lt;/span&gt;
lsblk &lt;span class="nt"&gt;-o&lt;/span&gt; NAME,TYPE,SIZE,FSTYPE,MOUNTPOINTS
findmnt &lt;span class="nt"&gt;-no&lt;/span&gt; SOURCE,FSTYPE,OPTIONS /
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Also record the VDS plan, vCPU count, memory, guest kernel, filesystem, virtualization-visible CPU model, benchmark tool versions, test-file path and size, and the UTC start and end times. Do not infer the physical host topology from &lt;code&gt;lscpu&lt;/code&gt; inside the guest. KVM exposes a virtual topology, and a provider may map or migrate vCPUs without making the underlying layout observable.&lt;/p&gt;

&lt;p&gt;Choose the measurement window in advance. Avoid backups, package upgrades, log rotation, and application deployments unless those activities are intentionally part of the test. For a production guest, set explicit load and latency stop conditions. A benchmark that harms the workload it is meant to evaluate has failed operationally even if its data is technically valid.&lt;/p&gt;

&lt;h2&gt;
  
  
  Observe scheduling before measuring CPU
&lt;/h2&gt;

&lt;p&gt;In a KVM guest, a runnable vCPU still needs a host CPU thread on which to run. Linux reports time when the guest wanted to run but the hypervisor scheduled something else as &lt;code&gt;%steal&lt;/code&gt;. Capture a quiet baseline, then capture the same counters while the CPU test runs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;mkdir&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; results
mpstat &lt;span class="nt"&gt;-P&lt;/span&gt; ALL 1 60 | &lt;span class="nb"&gt;tee &lt;/span&gt;results/mpstat-baseline.txt
vmstat 1 60 | &lt;span class="nb"&gt;tee &lt;/span&gt;results/vmstat-baseline.txt
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;mpstat -P ALL&lt;/code&gt; matters because an aggregate can hide one unstable vCPU. Save per-vCPU &lt;code&gt;%usr&lt;/code&gt;, &lt;code&gt;%sys&lt;/code&gt;, &lt;code&gt;%iowait&lt;/code&gt;, and &lt;code&gt;%steal&lt;/code&gt;. &lt;code&gt;vmstat&lt;/code&gt; adds runnable tasks (&lt;code&gt;r&lt;/code&gt;), blocked tasks (&lt;code&gt;b&lt;/code&gt;), context switches, and a second view of CPU state.&lt;/p&gt;

&lt;p&gt;Steal time is evidence, not a verdict. One non-zero sample may be harmless, and low steal does not prove that the CPU is dedicated. Look for coincidence: did throughput fall in the same seconds that steal rose? Does the pattern repeat in several runs or at different times of day? A sustained or recurrent correlation is more useful than a screenshot of one peak.&lt;/p&gt;

&lt;p&gt;Do not interpret &lt;code&gt;%iowait&lt;/code&gt; as disk latency. It is CPU accounting time during which a CPU was idle while I/O was outstanding. A system can have slow I/O and little iowait when other runnable work keeps the CPU busy. Storage latency must be measured from the I/O request and observed at the block layer.&lt;/p&gt;

&lt;h2&gt;
  
  
  Separate single-vCPU speed from scaling
&lt;/h2&gt;

&lt;p&gt;A multi-threaded score mixes at least two properties: per-vCPU execution speed and the scheduler’s ability to run several vCPUs concurrently. Measure them separately. &lt;code&gt;sysbench&lt;/code&gt; is convenient, but its version and parameters must be retained because scores from different builds or prime limits are not directly comparable.&lt;/p&gt;

&lt;p&gt;First warm the code path, then run one worker pinned to one guest vCPU:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;taskset &lt;span class="nt"&gt;-c&lt;/span&gt; 0 sysbench cpu &lt;span class="nt"&gt;--threads&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;1 &lt;span class="nt"&gt;--time&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;20 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--cpu-max-prime&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;20000 run &lt;span class="o"&gt;&amp;gt;&lt;/span&gt;/dev/null

taskset &lt;span class="nt"&gt;-c&lt;/span&gt; 0 sysbench cpu &lt;span class="nt"&gt;--threads&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;1 &lt;span class="nt"&gt;--time&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;60 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--cpu-max-prime&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;20000 run | &lt;span class="nb"&gt;tee &lt;/span&gt;results/cpu-1t-run01.txt
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;taskset&lt;/code&gt; prevents the process from moving among guest vCPUs, which removes one source of variation. It does not pin the workload to a physical Ryzen core on the host; only the hypervisor operator can make that guarantee. Repeat the timed run at least five times. Rotate the selected guest CPU in a separate experiment if you want to detect an unusually noisy vCPU, but do not silently mix those results into the primary series.&lt;/p&gt;

&lt;p&gt;Then test the intended concurrency. For a four-vCPU guest:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;sysbench cpu &lt;span class="nt"&gt;--threads&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;4 &lt;span class="nt"&gt;--time&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;20 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--cpu-max-prime&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;20000 run &lt;span class="o"&gt;&amp;gt;&lt;/span&gt;/dev/null

sysbench cpu &lt;span class="nt"&gt;--threads&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;4 &lt;span class="nt"&gt;--time&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;60 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--cpu-max-prime&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;20000 run | &lt;span class="nb"&gt;tee &lt;/span&gt;results/cpu-4t-run01.txt
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run &lt;code&gt;mpstat -P ALL 1&lt;/code&gt; in another shell during both tests. Report events per second and total events, plus the median, minimum, and maximum across repetitions. Scaling efficiency can be expressed as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;efficiency = multi-thread throughput / (single-thread throughput × worker count)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Efficiency below 100% is normal because of scheduling, shared caches, memory bandwidth, and benchmark overhead. The useful signal is reproducibility. Stable single-thread results with volatile multi-thread results and matching steal spikes suggest scheduling contention. Stable results at both thread counts make that explanation less likely.&lt;/p&gt;

&lt;p&gt;Do not compare a 30-second result on one server with a 10-minute result on another. Boost behavior, thermal state, host power policy, and contention can change over time. Keep duration, worker count, affinity policy, tool version, and warm-up identical.&lt;/p&gt;

&lt;h2&gt;
  
  
  Measure storage without turning a read test into a write test
&lt;/h2&gt;

&lt;p&gt;Use &lt;code&gt;fio&lt;/code&gt; only against a dedicated test file that already exists and was provisioned during a maintenance window. Never point a casual benchmark at a database file, a mounted block device, or an unknown path. Verify the target before every run:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;realpath&lt;/span&gt; /srv/benchmark/fio-test.bin
&lt;span class="nb"&gt;stat&lt;/span&gt; /srv/benchmark/fio-test.bin
findmnt &lt;span class="nt"&gt;-T&lt;/span&gt; /srv/benchmark/fio-test.bin
fio &lt;span class="nt"&gt;--version&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The following low-queue-depth test is read-only. &lt;code&gt;--readonly&lt;/code&gt; is a safety check that rejects write or trim workloads; it is worth keeping even when &lt;code&gt;--rw=randread&lt;/code&gt; is already specified.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;fio &lt;span class="nt"&gt;--name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;randread-4k-q1 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--filename&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;/srv/benchmark/fio-test.bin &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--readonly&lt;/span&gt; &lt;span class="nt"&gt;--rw&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;randread &lt;span class="nt"&gt;--bs&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;4k &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--ioengine&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;libaio &lt;span class="nt"&gt;--direct&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;1 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--iodepth&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;1 &lt;span class="nt"&gt;--numjobs&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;1 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--time_based&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;1 &lt;span class="nt"&gt;--runtime&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;90 &lt;span class="nt"&gt;--ramp_time&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;20 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--randrepeat&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;1 &lt;span class="nt"&gt;--randseed&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;20260919 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--lat_percentiles&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;1 &lt;span class="nt"&gt;--percentile_list&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;50:95:99:99.9 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--group_reporting&lt;/span&gt; &lt;span class="nt"&gt;--output-format&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;json &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--output&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;results/fio-4k-q1-run01.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;--direct=1&lt;/code&gt; reduces page cache effects, but it does not bypass every cache in the storage path. &lt;code&gt;--ramp_time=20&lt;/code&gt; gives the system a warm-up interval before statistics are collected. A fixed random seed makes the access sequence repeatable; that improves controlled comparison but can favor an upstream cache on later runs. State explicitly whether the objective is cold behavior, warmed steady state, or both. Do not claim “raw disk” performance from a virtual guest.&lt;/p&gt;

&lt;p&gt;Low queue depth is useful for latency-sensitive operations. Add a separate test at a queue depth and job count that resemble the application—for example QD32 for a deliberately concurrent workload—but do not replace QD1 with it. High concurrency can produce impressive IOPS while individual requests wait longer.&lt;/p&gt;

&lt;p&gt;During each &lt;code&gt;fio&lt;/code&gt; run, observe the guest block layer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;iostat &lt;span class="nt"&gt;-xz&lt;/span&gt; 1 120 | &lt;span class="nb"&gt;tee &lt;/span&gt;results/iostat-fio-q1-run01.txt
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In extended &lt;code&gt;iostat&lt;/code&gt; output, &lt;code&gt;r_await&lt;/code&gt; is the average time for read requests, including queue and service time. &lt;code&gt;aqu-sz&lt;/code&gt; shows the average queue length. &lt;code&gt;%util&lt;/code&gt; can be informative, but in a virtual or parallel storage stack it is not a universal saturation gauge. Interpret these fields together with &lt;code&gt;fio&lt;/code&gt; latency and IOPS, not in isolation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prefer percentiles to averages
&lt;/h2&gt;

&lt;p&gt;An average conceals the tail. If 99 reads finish in 0.5 ms and one takes 100 ms, the mean does not describe the pause that a request-sensitive application experiences. Retain at least p50, p95, p99, and preferably p99.9 latency from &lt;code&gt;fio&lt;/code&gt;. Keep the units visible: depending on the output and version, latency values may be represented in nanoseconds or microseconds.&lt;/p&gt;

&lt;p&gt;Percentiles within a run answer “how were individual requests distributed?” Repetitions answer a different question: “how stable is the environment across runs?” Do not compute a persuasive-looking p99 from only five aggregate scores. For each scenario, report the median result across at least five runs, along with its minimum and maximum. Preserve the within-run p95 and p99 from every raw JSON file.&lt;/p&gt;

&lt;p&gt;Run the series in more than one time window if neighbor activity is part of the risk being evaluated. Use the same sequence, or alternate scenario order to prevent every QD32 test from always inheriting the warmest cache. Note every deviation rather than deleting an inconvenient run. Exclude a run only by a rule defined before testing, such as an OS update process appearing in the log.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep a reproducible evidence bundle
&lt;/h2&gt;

&lt;p&gt;A useful result directory is understandable months later:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;results/
├── metadata.txt
├── mpstat-baseline.txt
├── cpu-1t-run01.txt
├── cpu-4t-run01.txt
├── fio-4k-q1-run01.json
├── iostat-fio-q1-run01.txt
└── notes.md
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Store exact commands in &lt;code&gt;notes.md&lt;/code&gt;. Hash the raw files after the session so later processing cannot silently change the evidence:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sha256sum &lt;/span&gt;results/&lt;span class="k"&gt;*&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; results/SHA256SUMS
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When comparing two VDS configurations, change one factor at a time and use the same guest image, kernel, filesystem, test-file size, duration, and tool versions. Raw data should accompany summary tables. A chart without the command and source output is an illustration, not a reproducible benchmark.&lt;/p&gt;

&lt;h2&gt;
  
  
  Turn observations into decisions
&lt;/h2&gt;

&lt;p&gt;Avoid universal pass/fail numbers. Define thresholds from the application’s service objective, then use patterns such as these:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Observation&lt;/th&gt;
&lt;th&gt;Plausible interpretation&lt;/th&gt;
&lt;th&gt;Next check&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Single-thread throughput is stable; multi-thread throughput varies with &lt;code&gt;%steal&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Host scheduling contention is affecting parallel work&lt;/td&gt;
&lt;td&gt;Repeat in another time window and correlate per-vCPU steal with each run&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CPU results are stable; QD1 p99 read latency is high&lt;/td&gt;
&lt;td&gt;Latency-sensitive storage work may be at risk&lt;/td&gt;
&lt;td&gt;Compare with an application trace and inspect &lt;code&gt;r_await&lt;/code&gt; and queue depth&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;QD1 latency is acceptable; QD32 IOPS is high but p99 grows sharply&lt;/td&gt;
&lt;td&gt;Storage benefits from concurrency at a tail-latency cost&lt;/td&gt;
&lt;td&gt;Cap application concurrency and test the intended queue depth&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;fio&lt;/code&gt; latency and &lt;code&gt;iostat&lt;/code&gt; &lt;code&gt;r_await&lt;/code&gt; rise together&lt;/td&gt;
&lt;td&gt;Delay is visible through the guest block path&lt;/td&gt;
&lt;td&gt;Check queue growth, other guest I/O, and time-of-day repetition&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;iostat&lt;/code&gt; is busy but the benchmark remains stable&lt;/td&gt;
&lt;td&gt;Another workload may share the device without yet harming the test&lt;/td&gt;
&lt;td&gt;Identify the process and repeat in a controlled window&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Results drift after warm-up&lt;/td&gt;
&lt;td&gt;Boost, caching, throttling, or background activity may dominate short tests&lt;/td&gt;
&lt;td&gt;Extend runtime and retain a time series instead of one total&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The final decision should name the workload and its tolerance: for example, “p99 4 KiB read latency must remain below the application’s measured limit while four CPU workers meet the batch deadline.” That is defensible. “This VDS scored X” is not.&lt;/p&gt;

&lt;h2&gt;
  
  
  A benchmark as a controlled argument
&lt;/h2&gt;

&lt;p&gt;A credible VDS benchmark does not need dozens of tools. It needs a declared question, a safe target, observable scheduling, separate single- and multi-vCPU tests, storage tests at realistic queue depths, warm-up, repetitions, and raw evidence. On KVM, &lt;code&gt;%steal&lt;/code&gt; helps explain missing CPU time; it does not explain everything. For storage, &lt;code&gt;await&lt;/code&gt; and average latency provide context, while p95 and p99 reveal the pauses users are more likely to notice.&lt;/p&gt;

&lt;p&gt;The most valuable outcome is not the highest number. It is a result that another engineer can reproduce, challenge, and connect to a real workload.&lt;/p&gt;

&lt;h2&gt;
  
  
  Türkçe paket karşılaştırma kaynakları
&lt;/h2&gt;

&lt;p&gt;Bu rehber EniyiSunucum tarafından yayımlanmıştır. Yukarıdaki ölçüm yöntemini uygularken seçtiğiniz paketin vCPU, RAM ve depolama sınırlarını test notlarınıza ekleyin. Sonuçları farklı paketlerle karşılaştırırken test süresini, eşzamanlı iş sayısını ve veri boyutunu sabit tutun.&lt;/p&gt;

&lt;p&gt;Türkçe ürün tablolarımızda sanal sunucu kaynaklarını &lt;a href="https://eniyisunucum.com/vds-sunucu" rel="noopener noreferrer"&gt;vds satın al&lt;/a&gt; ve &lt;a href="https://eniyisunucum.com/vps-sunucu" rel="noopener noreferrer"&gt;vps satın al&lt;/a&gt; sayfalarından inceleyebilirsiniz. AMD Ryzen işlemcili seçenekler için &lt;a href="https://eniyisunucum.com/ryzen-vds-sunucu" rel="noopener noreferrer"&gt;ryzen vds kirala&lt;/a&gt; sayfamızdaki paket kaynaklarını karşılaştırabilirsiniz. Bu bağlantılar şirketimizin kendi ürün sayfalarıdır; yayınlanan kaynak miktarları, uygulamanız için ölçülmüş performans sonucu anlamına gelmez.&lt;/p&gt;

&lt;h2&gt;
  
  
  Disclosure
&lt;/h2&gt;

&lt;p&gt;This article was drafted with AI assistance. Commands, tool behavior, and safety assumptions should be checked against the versions installed on the target system and validated in a non-production environment before use.&lt;/p&gt;

</description>
      <category>linux</category>
      <category>devops</category>
      <category>performance</category>
      <category>infrastructure</category>
    </item>
    <item>
      <title>Your XenForo Forum Is Slow: A Measurement-First PHP-FPM and MySQL Runbook</title>
      <dc:creator>EniyiSunucum</dc:creator>
      <pubDate>Fri, 18 Sep 2026 12:45:53 +0000</pubDate>
      <link>https://dev.to/eniyisunucum/your-xenforo-forum-is-slow-a-measurement-first-php-fpm-and-mysql-runbook-2e9k</link>
      <guid>https://dev.to/eniyisunucum/your-xenforo-forum-is-slow-a-measurement-first-php-fpm-and-mysql-runbook-2e9k</guid>
      <description>&lt;p&gt;A slow forum is rarely fixed by changing one PHP setting or buying a larger server blindly. The useful question is not &lt;em&gt;Is the server slow?&lt;/em&gt; but &lt;em&gt;Which queue grows when users feel the slowdown?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;This runbook separates four common causes of XenForo latency: CPU contention, PHP-FPM saturation, MySQL work, and storage delay. Every command below is read-only unless a step explicitly says otherwise. Capture a normal window and an incident window with the same commands; isolated screenshots are weak evidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Define the symptom before touching the stack
&lt;/h2&gt;

&lt;p&gt;Record one UTC interval when the forum is healthy and one when it is slow. For each interval, keep:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the affected URL class: thread view, search, login, posting, attachment upload, or admin task&lt;/li&gt;
&lt;li&gt;application response time, preferably p50, p95, and p99 rather than only an average&lt;/li&gt;
&lt;li&gt;HTTP status codes and upstream response time from the web-server access log&lt;/li&gt;
&lt;li&gt;concurrent requests and logged-in user count&lt;/li&gt;
&lt;li&gt;scheduled jobs, backups, imports, search indexing, and add-on maintenance running at that time.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Use one incident directory so timestamps line up:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;incident&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"/var/tmp/xf-&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;date&lt;/span&gt; &lt;span class="nt"&gt;-u&lt;/span&gt; +%Y%m%dT%H%M%SZ&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="nt"&gt;-m&lt;/span&gt; 700 &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$incident&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;span class="nb"&gt;date&lt;/span&gt; &lt;span class="nt"&gt;-u&lt;/span&gt; &lt;span class="nt"&gt;--iso-8601&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;seconds | &lt;span class="nb"&gt;tee&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$incident&lt;/span&gt;&lt;span class="s2"&gt;/start.txt"&lt;/span&gt;
&lt;span class="nb"&gt;uname&lt;/span&gt; &lt;span class="nt"&gt;-a&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$incident&lt;/span&gt;&lt;span class="s2"&gt;/kernel.txt"&lt;/span&gt;
&lt;span class="nb"&gt;uptime&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$incident&lt;/span&gt;&lt;span class="s2"&gt;/uptime.txt"&lt;/span&gt;
free &lt;span class="nt"&gt;-m&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$incident&lt;/span&gt;&lt;span class="s2"&gt;/memory.txt"&lt;/span&gt;
&lt;span class="nb"&gt;df&lt;/span&gt; &lt;span class="nt"&gt;-hT&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$incident&lt;/span&gt;&lt;span class="s2"&gt;/filesystems.txt"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Do not publish raw logs before removing IP addresses, usernames, query strings, cookies, tokens, and private paths.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Decide whether the wait begins at CPU or storage
&lt;/h2&gt;

&lt;p&gt;Collect CPU, run-queue, pressure-stall, process, and block-device data during the same 60-second window:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;LC_ALL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;C mpstat &lt;span class="nt"&gt;-P&lt;/span&gt; ALL 1 60 &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$incident&lt;/span&gt;&lt;span class="s2"&gt;/mpstat.txt"&lt;/span&gt; &amp;amp;
&lt;span class="nv"&gt;LC_ALL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;C vmstat &lt;span class="nt"&gt;-w&lt;/span&gt; &lt;span class="nt"&gt;-t&lt;/span&gt; 1 60 &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$incident&lt;/span&gt;&lt;span class="s2"&gt;/vmstat.txt"&lt;/span&gt; &amp;amp;
&lt;span class="nv"&gt;LC_ALL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;C pidstat &lt;span class="nt"&gt;-u&lt;/span&gt; &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="nt"&gt;-w&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; ALL 1 60 &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$incident&lt;/span&gt;&lt;span class="s2"&gt;/pidstat.txt"&lt;/span&gt; &amp;amp;
&lt;span class="nv"&gt;LC_ALL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;C iostat &lt;span class="nt"&gt;-xz&lt;/span&gt; &lt;span class="nt"&gt;-t&lt;/span&gt; 1 60 &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$incident&lt;/span&gt;&lt;span class="s2"&gt;/iostat.txt"&lt;/span&gt; &amp;amp;
&lt;span class="nb"&gt;wait&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Interpret combinations, not single columns:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A sustained run queue above the available vCPU count, high CPU pressure, and low idle time points to guest-side CPU saturation.&lt;/li&gt;
&lt;li&gt;Rising steal time that coincides with application latency can indicate hypervisor scheduling delay. A zero value does not prove that the host is perfect, but a correlated rise is useful evidence.&lt;/li&gt;
&lt;li&gt;Rising read or write latency, queue depth, and I/O pressure points toward the storage path. Do not treat device utilization alone as a universal saturation signal on virtual disks or modern SSDs.&lt;/li&gt;
&lt;li&gt;High iowait is a clue, not a diagnosis. Correlate it with device latency, pressure, and the processes issuing I/O.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If &lt;code&gt;/proc/pressure/cpu&lt;/code&gt; and &lt;code&gt;/proc/pressure/io&lt;/code&gt; exist, sample them once per second. PSI tells you how much time runnable work is stalled, which is often closer to the user-visible problem than a five-minute load average.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Check PHP-FPM as a queue, not as a memory formula
&lt;/h2&gt;

&lt;p&gt;A busy forum can be slow while the host still has free CPU because requests are waiting for PHP workers. Enable the PHP-FPM status endpoint only on a private management path protected by an allowlist or local socket. Never expose it publicly.&lt;/p&gt;

&lt;p&gt;During the incident, compare:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;active processes versus the configured maximum&lt;/li&gt;
&lt;li&gt;max-active-processes and max-children-reached counters&lt;/li&gt;
&lt;li&gt;listen-queue depth and its historical maximum&lt;/li&gt;
&lt;li&gt;slow-request records from the FPM slow log&lt;/li&gt;
&lt;li&gt;resident memory per worker under realistic traffic.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If &lt;code&gt;max children reached&lt;/code&gt; increases while the CPU is not saturated, requests are queuing at FPM. Raising the worker limit may help only when memory and database capacity exist. Estimate from observed RSS, leave room for MySQL, the kernel page cache, the web server, and traffic spikes, then change one variable in a maintenance window.&lt;/p&gt;

&lt;p&gt;If workers are busy for a long time, adding more can amplify the real bottleneck. Use the slow log to find whether workers wait on database queries, remote APIs, filesystem operations, or add-on code.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Separate MySQL concurrency from slow individual queries
&lt;/h2&gt;

&lt;p&gt;Start with current state:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SHOW&lt;/span&gt; &lt;span class="k"&gt;GLOBAL&lt;/span&gt; &lt;span class="n"&gt;STATUS&lt;/span&gt; &lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;Variable_name&lt;/span&gt; &lt;span class="k"&gt;IN&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="s1"&gt;'Threads_connected'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'Threads_running'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="s1"&gt;'Created_tmp_disk_tables'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'Created_tmp_tables'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="s1"&gt;'Innodb_buffer_pool_reads'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'Innodb_buffer_pool_read_requests'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="s1"&gt;'Innodb_row_lock_time'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'Innodb_row_lock_waits'&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="k"&gt;SHOW&lt;/span&gt; &lt;span class="k"&gt;FULL&lt;/span&gt; &lt;span class="n"&gt;PROCESSLIST&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;SHOW&lt;/span&gt; &lt;span class="n"&gt;ENGINE&lt;/span&gt; &lt;span class="n"&gt;INNODB&lt;/span&gt; &lt;span class="n"&gt;STATUS&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Capture these at the same time as the Linux metrics. A large connection count is not automatically bad; &lt;code&gt;Threads_running&lt;/code&gt;, lock waits, and query latency are more useful. Likewise, a buffer-pool miss counter is cumulative, so compare deltas over a fixed interval.&lt;/p&gt;

&lt;p&gt;If Performance Schema is enabled, identify statement digests with high total latency, high rows examined, or repeated executions. Inspect representative queries with &lt;code&gt;EXPLAIN&lt;/code&gt; on a staging copy or during a controlled diagnostic window. Do not add indexes from intuition alone: an index can improve reads while increasing write cost and storage.&lt;/p&gt;

&lt;p&gt;A temporary slow-query-log window can be valuable, but enabling it changes server state and may write sensitive query data. Plan it, protect the log, set a short observation period, and disable it after collection.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Test XenForo-specific suspects one at a time
&lt;/h2&gt;

&lt;p&gt;Infrastructure graphs cannot identify every application cause. Common forum-side suspects include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;an add-on executing work on every page view&lt;/li&gt;
&lt;li&gt;search or indexing tasks overlapping peak traffic&lt;/li&gt;
&lt;li&gt;attachment or image processing bursts&lt;/li&gt;
&lt;li&gt;scheduled jobs, email queues, imports, and sitemap generation&lt;/li&gt;
&lt;li&gt;external API calls inside request handling&lt;/li&gt;
&lt;li&gt;session, permission, or template work made expensive by an add-on&lt;/li&gt;
&lt;li&gt;a large thread or search path that produces unusually heavy queries.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Reproduce the slow URL with the same account permissions in staging. Disable one suspect add-on at a time there, keep caches and dataset size comparable, and record the before-and-after request trace. Clearing every cache, restarting every service, and changing multiple limits simultaneously destroys the evidence you need.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Use a decision table
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Evidence during the same interval&lt;/th&gt;
&lt;th&gt;First hypothesis&lt;/th&gt;
&lt;th&gt;Next check&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;FPM queue grows; CPU and MySQL are calm&lt;/td&gt;
&lt;td&gt;Worker capacity or blocked PHP calls&lt;/td&gt;
&lt;td&gt;FPM slow log and per-worker RSS&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;Threads_running&lt;/code&gt; and lock waits rise&lt;/td&gt;
&lt;td&gt;Database contention&lt;/td&gt;
&lt;td&gt;statement digests, InnoDB status, transaction scope&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CPU PSI and run queue rise; steal stays low&lt;/td&gt;
&lt;td&gt;Guest CPU saturation&lt;/td&gt;
&lt;td&gt;hot processes, add-ons, request mix&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Steal rises with latency; guest CPU is not full&lt;/td&gt;
&lt;td&gt;Hypervisor scheduling pressure&lt;/td&gt;
&lt;td&gt;repeatable timestamps for the provider&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;I/O PSI and device latency rise together&lt;/td&gt;
&lt;td&gt;Storage-path contention&lt;/td&gt;
&lt;td&gt;per-process I/O, backup/index tasks, provider evidence&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Only one route is slow&lt;/td&gt;
&lt;td&gt;Application/query path&lt;/td&gt;
&lt;td&gt;request trace and query digest&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The table is a prioritization tool, not proof by itself. Repeat the observation and look for the same correlation before migrating or resizing.&lt;/p&gt;

&lt;h2&gt;
  
  
  7. Scale only after the bottleneck has a name
&lt;/h2&gt;

&lt;p&gt;More vCPU helps a CPU-bound workload only when useful work can run in parallel. Faster single-core performance can matter for serial PHP execution, but it will not fix a locked database transaction. More RAM can reduce disk reads when the working set does not fit, but it will not repair an unindexed query. NVMe can reduce storage latency, but it will not clear an FPM listen queue caused by a remote API.&lt;/p&gt;

&lt;p&gt;When comparing shared hosting, VPS, or VDS, carry your evidence into the decision: sustained CPU demand, measured memory working set, database size, p95 storage latency, peak concurrent PHP workers, backup window, and growth margin. The product label is less important than verified resource boundaries and an upgrade path.&lt;/p&gt;

&lt;h2&gt;
  
  
  A practical handoff bundle
&lt;/h2&gt;

&lt;p&gt;A useful incident bundle contains UTC start and end times, the affected URLs, application percentiles, web-server upstream timings, FPM status and slow traces, MySQL snapshots, Linux CPU/I/O evidence, exact package versions, and a short list of changes made. Add checksums after collection so later edits are visible.&lt;/p&gt;

&lt;p&gt;That bundle turns “the forum feels slow” into a claim another engineer or provider can test.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Disclosure:&lt;/strong&gt; This article was prepared by the EniyiSunucum infrastructure team, which provides server and hosting services. Readers evaluating managed environments can compare &lt;a href="https://eniyisunucum.com/en/xenforo-hosting" rel="noopener noreferrer"&gt;XenForo hosting options&lt;/a&gt;, but the diagnostic method above is provider-neutral and contains no unpublished benchmark claims. AI assistance was used for structure and language editing; commands, interpretations, and links were reviewed before publication.&lt;/p&gt;

</description>
      <category>php</category>
      <category>performance</category>
      <category>linux</category>
      <category>database</category>
    </item>
    <item>
      <title>CPU Looks Fine but the VDS Is Slow: A Linux Incident Runbook</title>
      <dc:creator>EniyiSunucum</dc:creator>
      <pubDate>Thu, 17 Sep 2026 00:17:29 +0000</pubDate>
      <link>https://dev.to/eniyisunucum/cpu-looks-fine-but-the-vds-is-slow-a-linux-incident-runbook-3iha</link>
      <guid>https://dev.to/eniyisunucum/cpu-looks-fine-but-the-vds-is-slow-a-linux-incident-runbook-3iha</guid>
      <description>&lt;p&gt;A user reports that an application is slow or intermittently unreachable. The monitoring dashboard shows 20% CPU, memory still available, and no obvious outage. It is tempting to answer, “the server looks healthy.”&lt;/p&gt;

&lt;p&gt;That answer is usually premature.&lt;/p&gt;

&lt;p&gt;A single CPU graph cannot tell you whether a Linux guest is waiting on storage, reclaiming memory, losing packets, stalling inside a cgroup, contending for hypervisor time, or following a degraded network path. The useful question is not “is CPU high?” but &lt;strong&gt;where did the request spend its time?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This runbook provides a repeatable way to answer that question on a Linux virtual server. It starts with low-risk evidence, separates guest-level faults from upstream faults, and produces a compact incident package you can hand to an application team or infrastructure provider.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Safety note:&lt;/strong&gt; The first stages below are read-only. Run load generators such as fio or iperf3 only during an approved window, against endpoints you control, and with explicit rate limits. Never benchmark a production filesystem or a third-party host without permission.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  1. Freeze the incident boundary
&lt;/h2&gt;

&lt;p&gt;Before collecting metrics, write down four facts:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The exact start and end time in UTC.&lt;/li&gt;
&lt;li&gt;The affected hostname, IP, port, and protocol.&lt;/li&gt;
&lt;li&gt;Whether all clients were affected or only one location/ISP.&lt;/li&gt;
&lt;li&gt;One concrete symptom: timeout, connection reset, slow TTFB, packet loss, or high application latency.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Capture a small identity bundle:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;date&lt;/span&gt; &lt;span class="nt"&gt;-u&lt;/span&gt; &lt;span class="nt"&gt;-Is&lt;/span&gt;
hostnamectl
&lt;span class="nb"&gt;uname&lt;/span&gt; &lt;span class="nt"&gt;-a&lt;/span&gt;
&lt;span class="nb"&gt;uptime
who&lt;/span&gt; &lt;span class="nt"&gt;-b&lt;/span&gt;
journalctl &lt;span class="nt"&gt;--list-boots&lt;/span&gt; | &lt;span class="nb"&gt;head&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This prevents a common failure: comparing a user report from 14:05 with metrics from 14:40, after the system has already recovered.&lt;/p&gt;

&lt;p&gt;If the problem is HTTP, record timings from both the server and an external client:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;time &lt;/span&gt;curl &lt;span class="nt"&gt;-sS&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; /dev/null https://example.com/
curl &lt;span class="nt"&gt;-sS&lt;/span&gt; &lt;span class="nt"&gt;-D&lt;/span&gt; - &lt;span class="nt"&gt;-o&lt;/span&gt; /dev/null https://example.com/
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Repeat the same request against localhost or the private service address when possible. Fast locally but slow externally points away from the application process and toward the proxy, firewall, network, or client path.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Read CPU as a scheduler, not a percentage
&lt;/h2&gt;

&lt;p&gt;Start with a short time series rather than a single snapshot:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;vmstat 1 10
mpstat &lt;span class="nt"&gt;-P&lt;/span&gt; ALL 1 10
pidstat &lt;span class="nt"&gt;-u&lt;/span&gt; &lt;span class="nt"&gt;-w&lt;/span&gt; 1 10
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;These commands are provided by procps and sysstat on most distributions.&lt;/p&gt;

&lt;p&gt;Pay attention to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;r&lt;/strong&gt; in vmstat: runnable tasks waiting for CPU.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;b&lt;/strong&gt;: tasks blocked, often on I/O.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;wa&lt;/strong&gt;: time waiting for I/O.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;st&lt;/strong&gt;: steal time, when the virtual CPU was ready but the hypervisor scheduled something else.&lt;/li&gt;
&lt;li&gt;Context switches and migrations in pidstat.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Low aggregate CPU does not rule out saturation. One single-threaded worker can max one vCPU while the average across eight vCPUs remains near 12.5%. Per-CPU output from mpstat exposes that pattern.&lt;/p&gt;

&lt;p&gt;Sustained steal time that correlates with latency is strong evidence of host-side contention, but one sample is not proof. Capture several intervals and correlate them with application timing.&lt;/p&gt;

&lt;p&gt;Linux Pressure Stall Information adds another useful view:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;cat&lt;/span&gt; /proc/pressure/cpu
&lt;span class="nb"&gt;cat&lt;/span&gt; /proc/pressure/memory
&lt;span class="nb"&gt;cat&lt;/span&gt; /proc/pressure/io
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The “some” and “full” lines describe how much time tasks were delayed because a resource was unavailable. PSI is often more informative than utilization because it measures waiting experienced by workloads.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Distinguish available memory from memory pressure
&lt;/h2&gt;

&lt;p&gt;Linux deliberately uses free RAM for cache, so the “free” column alone is not an alarm. Look for reclaim, swap activity, and allocation failures:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;free &lt;span class="nt"&gt;-h&lt;/span&gt;
vmstat 1 10
&lt;span class="nb"&gt;cat&lt;/span&gt; /proc/pressure/memory
ps &lt;span class="nt"&gt;-eo&lt;/span&gt; pid,ppid,comm,%mem,rss,vsz &lt;span class="nt"&gt;--sort&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nt"&gt;-rss&lt;/span&gt; | &lt;span class="nb"&gt;head&lt;/span&gt; &lt;span class="nt"&gt;-20&lt;/span&gt;
journalctl &lt;span class="nt"&gt;-k&lt;/span&gt; &lt;span class="nt"&gt;--since&lt;/span&gt; &lt;span class="s1"&gt;'-30 min'&lt;/span&gt; | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-Ei&lt;/span&gt; &lt;span class="s1"&gt;'oom|out of memory|memory cgroup'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Interpret the evidence together:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Continuous &lt;strong&gt;si/so&lt;/strong&gt; in vmstat indicates active swapping.&lt;/li&gt;
&lt;li&gt;Rising memory PSI means tasks are stalling during reclaim.&lt;/li&gt;
&lt;li&gt;OOM or memory-cgroup messages explain sudden process restarts.&lt;/li&gt;
&lt;li&gt;A large page cache with low swap activity is usually normal.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Inside containers, host memory may look comfortable while a service is hitting its cgroup limit. On cgroup v2 systems, inspect the relevant scope under /sys/fs/cgroup and compare memory.current with memory.max. Do not assume the VM total is the service limit.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Check storage latency, not only disk usage
&lt;/h2&gt;

&lt;p&gt;A filesystem can have plenty of free space and still respond slowly. First rule out capacity and inode exhaustion:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;df&lt;/span&gt; &lt;span class="nt"&gt;-hT&lt;/span&gt;
&lt;span class="nb"&gt;df&lt;/span&gt; &lt;span class="nt"&gt;-i&lt;/span&gt;
lsblk &lt;span class="nt"&gt;-o&lt;/span&gt; NAME,TYPE,SIZE,FSTYPE,MOUNTPOINTS
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then sample the block layer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;iostat &lt;span class="nt"&gt;-xz&lt;/span&gt; 1 10
pidstat &lt;span class="nt"&gt;-d&lt;/span&gt; 1 10
&lt;span class="nb"&gt;cat&lt;/span&gt; /proc/pressure/io
journalctl &lt;span class="nt"&gt;-k&lt;/span&gt; &lt;span class="nt"&gt;--since&lt;/span&gt; &lt;span class="s1"&gt;'-30 min'&lt;/span&gt; | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-Ei&lt;/span&gt; &lt;span class="s1"&gt;'I/O error|timeout|reset|nvme|blk_update'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Useful indicators include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;await&lt;/strong&gt;: average time for an I/O request to complete.&lt;/li&gt;
&lt;li&gt;Queue size: sustained queues suggest the device cannot keep up.&lt;/li&gt;
&lt;li&gt;I/O PSI: application-visible time lost waiting on storage.&lt;/li&gt;
&lt;li&gt;Per-process reads and writes from pidstat.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Treat %util carefully on modern virtual and parallel storage. A high value can be meaningful, but a low value does not guarantee low latency. The combination of await, queueing, PSI, and application timing is stronger than any one number.&lt;/p&gt;

&lt;p&gt;Avoid immediately running fio. A benchmark can turn a partial incident into a full outage and can contaminate the very evidence you are trying to preserve. If a controlled test is necessary, use a disposable file, direct I/O where appropriate, a fixed size, a runtime limit, and an agreed bandwidth or IOPS cap.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Inspect sockets and the network stack
&lt;/h2&gt;

&lt;p&gt;Begin at the guest boundary:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ss &lt;span class="nt"&gt;-s&lt;/span&gt;
ss &lt;span class="nt"&gt;-lntup&lt;/span&gt;
ss &lt;span class="nt"&gt;-tin&lt;/span&gt;
ip &lt;span class="nt"&gt;-s&lt;/span&gt; &lt;span class="nb"&gt;link
&lt;/span&gt;nstat &lt;span class="nt"&gt;-az&lt;/span&gt;
sar &lt;span class="nt"&gt;-n&lt;/span&gt; DEV,TCP,ETCP 1 10
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Look for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Retransmissions increasing during the incident.&lt;/li&gt;
&lt;li&gt;Receive or transmit drops on the interface.&lt;/li&gt;
&lt;li&gt;A growing listen or SYN backlog.&lt;/li&gt;
&lt;li&gt;Many sockets stuck in SYN-SENT, SYN-RECV, or CLOSE-WAIT.&lt;/li&gt;
&lt;li&gt;TCP RTT and retransmission fields in ss -tin.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If the interface counters are clean but the application times out, inspect each layer of the request path: resolver, firewall, reverse proxy, application listener, and upstream dependency.&lt;/p&gt;

&lt;p&gt;For DNS, compare authoritative and recursive answers:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;dig example.com A +noall +answer +stats
dig @1.1.1.1 example.com A +noall +answer +stats
resolvectl query example.com
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A slow resolver can make an otherwise healthy service feel randomly slow, especially when caches expire.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Test the path without over-reading ICMP
&lt;/h2&gt;

&lt;p&gt;From a client that actually experienced the problem, collect a bounded path sample:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;mtr &lt;span class="nt"&gt;-rwzc&lt;/span&gt; 50 server.example.com
tracepath server.example.com
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run the reverse direction when you control both endpoints. Internet paths are often asymmetric.&lt;/p&gt;

&lt;p&gt;Do not diagnose packet loss from one intermediate hop alone. Routers may rate-limit or deprioritize ICMP while continuing to forward application traffic normally. Loss becomes meaningful when it begins at a hop and persists through later hops, especially to the destination, and when it correlates with TCP or application symptoms.&lt;/p&gt;

&lt;p&gt;Use iperf3 only between systems you control. Start the server on one endpoint, choose a modest bandwidth cap for UDP tests, and stop if the link or production workload degrades. Throughput is not a substitute for application timing; it is one controlled data point.&lt;/p&gt;

&lt;h2&gt;
  
  
  7. Add external routing evidence
&lt;/h2&gt;

&lt;p&gt;When the guest looks healthy but multiple remote networks fail, check the route from outside your infrastructure:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Is the prefix visible from several public route collectors?&lt;/li&gt;
&lt;li&gt;Did the origin ASN change?&lt;/li&gt;
&lt;li&gt;Is the announcement covered by a valid RPKI Route Origin Authorization?&lt;/li&gt;
&lt;li&gt;Did the AS path change near the incident time?&lt;/li&gt;
&lt;li&gt;Is the problem limited to one region or upstream?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A route visible from one collector is not proof of global reachability. Compare more than one vantage point and record timestamps. This is especially important for regional hosting, where an upstream or peering issue may affect one country while local monitoring stays green.&lt;/p&gt;

&lt;h2&gt;
  
  
  8. Decide which boundary owns the next action
&lt;/h2&gt;

&lt;p&gt;Use the evidence to route the incident:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Evidence&lt;/th&gt;
&lt;th&gt;Most likely next owner&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;One process or one vCPU saturated&lt;/td&gt;
&lt;td&gt;Application/service team&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Swap, memory PSI, or cgroup limit&lt;/td&gt;
&lt;td&gt;Guest configuration/application team&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;High await, queueing, and I/O PSI&lt;/td&gt;
&lt;td&gt;Storage or infrastructure team&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sustained steal time with normal guest load&lt;/td&gt;
&lt;td&gt;Virtualization provider&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Interface drops or retransmits from the guest&lt;/td&gt;
&lt;td&gt;Guest networking or provider edge&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Clean guest, path degradation across several clients&lt;/td&gt;
&lt;td&gt;Network/upstream provider&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Local service fast, public hostname slow&lt;/td&gt;
&lt;td&gt;DNS, proxy, firewall, or network path&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This table is not a verdict. It tells you where the next discriminating test belongs.&lt;/p&gt;

&lt;h2&gt;
  
  
  9. Build an evidence package a provider can act on
&lt;/h2&gt;

&lt;p&gt;A useful escalation is short and reproducible. Include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;VM identifier and affected public IP.&lt;/li&gt;
&lt;li&gt;UTC incident window.&lt;/li&gt;
&lt;li&gt;Source locations or ISPs affected.&lt;/li&gt;
&lt;li&gt;Exact destination IP, port, and protocol.&lt;/li&gt;
&lt;li&gt;Ten to sixty seconds of vmstat, mpstat, iostat, PSI, and interface counters.&lt;/li&gt;
&lt;li&gt;Application timing from inside and outside.&lt;/li&gt;
&lt;li&gt;MTR from both directions when available.&lt;/li&gt;
&lt;li&gt;Public routing observations and timestamps.&lt;/li&gt;
&lt;li&gt;A clear request, such as “please check host scheduling for this VM between 14:02 and 14:09 UTC.”&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Avoid screenshots without axes, “the network is slow,” or a five-megabyte log dump with no timeline. The goal is to let another engineer test a specific hypothesis.&lt;/p&gt;

&lt;h2&gt;
  
  
  10. Keep the runbook lightweight
&lt;/h2&gt;

&lt;p&gt;The best incident kit is the one already installed and practiced. A small baseline of sysstat, curl, dig, mtr, and journal access covers a large share of Linux performance incidents. Add persistent monitoring for PSI, steal time, disk latency, retransmissions, and application percentiles so the next investigation begins with history instead of guesswork.&lt;/p&gt;

&lt;p&gt;A &lt;a href="https://eniyisunucum.com/en/vds-server" rel="noopener noreferrer"&gt;VDS&lt;/a&gt; plan should make the CPU allocation, memory limit, storage type, network policy, and support boundary explicit. Those details determine which signals you can observe and which evidence your provider must supply.&lt;/p&gt;

&lt;h2&gt;
  
  
  Final principle
&lt;/h2&gt;

&lt;p&gt;When CPU looks fine, do not jump straight to a benchmark or a provider ticket. Follow the request across boundaries:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;application → scheduler → memory → storage → socket → guest interface → network path → routing&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;At each boundary, collect one time-stamped signal that can falsify a hypothesis. That turns “the server is slow” into an incident another engineer can reproduce and resolve.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Disclosure: I work on Linux-based virtual server infrastructure and AS215068 at &lt;a href="https://eniyisunucum.com/en/" rel="noopener noreferrer"&gt;EniyiSunucum&lt;/a&gt;. This runbook is vendor-neutral and the commands apply to standard Linux environments.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>devops</category>
      <category>linux</category>
      <category>performance</category>
      <category>networking</category>
    </item>
  </channel>
</rss>
