<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Aditya Sharma</title>
    <description>The latest articles on DEV Community by Aditya Sharma (@aditya_d_sharma).</description>
    <link>https://dev.to/aditya_d_sharma</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4051470%2Fe252149c-c45f-483e-935c-5356bb60e731.png</url>
      <title>DEV Community: Aditya Sharma</title>
      <link>https://dev.to/aditya_d_sharma</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/aditya_d_sharma"/>
    <language>en</language>
    <item>
      <title>What Is ASLR? How Does Randomizing Memory Addresses Stop Exploitation?</title>
      <dc:creator>Aditya Sharma</dc:creator>
      <pubDate>Sun, 27 Sep 2026 13:45:52 +0000</pubDate>
      <link>https://dev.to/aditya_d_sharma/what-is-aslr-how-does-randomizing-memory-addresses-stop-exploitation-509b</link>
      <guid>https://dev.to/aditya_d_sharma/what-is-aslr-how-does-randomizing-memory-addresses-stop-exploitation-509b</guid>
      <description>&lt;p&gt;What if the memory address an attacker needs today isn't the same address tomorrow?&lt;/p&gt;

&lt;p&gt;That question is the core idea behind Address Space Layout Randomization. Before getting into how it works, it helps to understand the problem it addresses.&lt;/p&gt;

&lt;p&gt;Many exploitation techniques rely on knowing where things are in memory. If you can corrupt memory in a running process, that corruption is only useful if you can direct it somewhere meaningful. Overwrite the right data, redirect execution to the right address, or interfere with the right memory region. Without knowing where those things are, the attacker has a harder time turning a vulnerability into reliable controlled behavior.&lt;/p&gt;

&lt;p&gt;ASLR doesn't fix the underlying vulnerability. The buggy code still exists. What ASLR does is make the memory layout unpredictable, introducing uncertainty that makes reliable exploitation significantly harder.&lt;/p&gt;




&lt;h2&gt;
  
  
  Understanding Virtual Memory First
&lt;/h2&gt;

&lt;p&gt;To understand ASLR, you need a clear picture of what "address space" means and what ASLR is actually randomizing.&lt;/p&gt;

&lt;p&gt;When a program runs, it doesn't directly address physical RAM. The operating system gives it a virtual address space, a range of addresses the process can use as if it had access to a large, contiguous block of memory. The operating system and hardware (through the memory management unit) translate those virtual addresses to actual physical memory locations at runtime.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Process
   |
   v
Virtual Address Space
   |
   v
Virtual-to-Physical Mapping (page table)
   |
   v
Physical Memory
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each process has its own virtual address space. Two different processes can use the same virtual addresses while mapping to entirely different physical memory. This is why one process can't accidentally read another's memory: the same virtual address means different physical memory in different processes.&lt;/p&gt;

&lt;p&gt;ASLR works within this virtual address space layer. It changes where different regions of a process's virtual address space are placed each time the program loads.&lt;/p&gt;




&lt;h2&gt;
  
  
  A Process's Memory Layout
&lt;/h2&gt;

&lt;p&gt;A running process typically has several distinct memory regions, each serving a different purpose:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;High Addresses
------------------
|    Stack       |   Local variables, return state, function call frames
------------------
|     ...        |
------------------
| Shared Libs    |   Dynamically linked code (libc, etc.)
------------------
|     Heap       |   Dynamically allocated memory
------------------
| Data / BSS     |   Global/static variables
------------------
|    Code        |   Program instructions (text segment)
------------------
Low Addresses
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The exact layout depends on the operating system, architecture, executable format, linker, and loader. This is a conceptual model, not a universal specification. But the general principle holds: different types of data and code live in different regions of the virtual address space.&lt;/p&gt;

&lt;p&gt;Without ASLR, these regions tend to load at fixed or predictable base addresses. The code segment starts at the same place every time. The stack begins at the same place. Shared libraries map to the same addresses. An attacker who studies the binary or observes memory once can often predict exactly where things will be in future runs.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why Predictable Addresses Help Attackers
&lt;/h2&gt;

&lt;p&gt;Consider a memory corruption vulnerability, a type confusion bug, an out-of-bounds write, a buffer overflow. The vulnerability lets an attacker write data somewhere in memory. That's a starting point, not an end goal. The question is what that write can accomplish.&lt;/p&gt;

&lt;p&gt;Exploitation techniques often need to redirect execution, corrupt data structures at specific locations, or reference existing code that does something useful. All of these require knowing the addresses involved.&lt;/p&gt;

&lt;p&gt;Without ASLR:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="go"&gt;Run 1 → libc at address 0xf7400000
Run 2 → libc at address 0xf7400000
Run 3 → libc at address 0xf7400000
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The attacker can work out those addresses through static analysis, running the application locally, or reading public information about the system. Once known, they stay known.&lt;/p&gt;

&lt;p&gt;With ASLR:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="go"&gt;Run 1 → libc at address 0xf7312000
Run 2 → libc at address 0xb76a1000
Run 3 → libc at address 0xf4a23000
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The attacker's knowledge of where libc was in one execution doesn't help for the next. They face a new problem: before exploiting the vulnerability, they need to discover the current layout.&lt;/p&gt;




&lt;h2&gt;
  
  
  What ASLR Actually Randomizes
&lt;/h2&gt;

&lt;p&gt;ASLR randomizes the base addresses of major process memory regions when the process loads. Depending on the operating system, configuration, and whether the executable is built with appropriate support, this can include:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The stack.&lt;/strong&gt; Where the stack begins in virtual memory changes between runs. This affects the addresses of local variables, saved return state, and other stack-based structures.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The heap.&lt;/strong&gt; Where the dynamic memory allocator starts its managed region changes between runs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Shared libraries.&lt;/strong&gt; Dynamically linked libraries are mapped into the process address space at load time. With ASLR, their base addresses are randomized, changing the addresses of all the code and data they contain.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Memory-mapped regions.&lt;/strong&gt; Files or other resources mapped into memory also receive randomized addresses.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The main executable.&lt;/strong&gt; This only receives randomization if the executable is built as a Position Independent Executable (PIE), discussed next.&lt;/p&gt;




&lt;h2&gt;
  
  
  PIE: Randomizing the Executable Itself
&lt;/h2&gt;

&lt;p&gt;A normal compiled executable often assumes it will be loaded at a specific base address. Instructions in the binary may use hardcoded addresses relative to where the program expects to be. If the operating system tries to load it at a different address, the program breaks.&lt;/p&gt;

&lt;p&gt;A Position Independent Executable is built differently. The compiler and linker generate code that works correctly regardless of where it is loaded in virtual memory. Instead of hardcoded absolute addresses, the code uses position-relative addressing. The loader can place the executable at any address and everything still functions.&lt;/p&gt;

&lt;p&gt;This distinction matters:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ASLR = the OS mechanism that randomizes load addresses at runtime

PIE = the executable property that allows the main binary's base address to be randomized
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;ASLR can randomize shared libraries, the stack, and the heap without PIE. But without PIE, the main executable's code and data often remain at a predictable base address. That gives an attacker a fixed anchor point in an otherwise randomized address space.&lt;/p&gt;

&lt;p&gt;Modern build toolchains default to PIE for executables on many platforms, but this isn't universal. Older software, software built with older toolchains, or software that explicitly disables PIE may lack it.&lt;/p&gt;




&lt;h2&gt;
  
  
  Entropy: How Much Randomization Is Enough?
&lt;/h2&gt;

&lt;p&gt;Randomization is only as strong as the range of possible positions. If the stack can only land in one of a small number of locations, an attacker might attempt every possibility, a strategy sometimes called brute-forcing the layout.&lt;/p&gt;

&lt;p&gt;The number of possible positions is described by the entropy: how many bits of randomness are applied to each region. With n bits of entropy, there are 2^n possible base addresses.&lt;/p&gt;

&lt;p&gt;On 32-bit systems, the total virtual address space is limited (4 gigabytes, or 32 bits). Allocating space for code, libraries, stack, and heap leaves relatively little room for randomization. A region might have only 16 bits of entropy for its base address, meaning 65,536 possible positions. Still an obstacle, but potentially feasible to brute-force in certain scenarios.&lt;/p&gt;

&lt;p&gt;On 64-bit systems, the virtual address space is vastly larger. More bits are available for randomization. A shared library might have 28 or more bits of entropy in its base address, which means over 268 million possible positions. Brute-forcing this is impractical under normal circumstances.&lt;/p&gt;

&lt;p&gt;This is a primary reason why 64-bit systems generally offer meaningfully stronger ASLR protection than their 32-bit predecessors.&lt;/p&gt;




&lt;h2&gt;
  
  
  Information Leaks: ASLR's Biggest Practical Weakness
&lt;/h2&gt;

&lt;p&gt;ASLR works by making addresses unknown. The obvious way to defeat it is to make them known.&lt;/p&gt;

&lt;p&gt;An information disclosure vulnerability is a bug that reveals memory contents the attacker shouldn't be able to read. If one of those memory contents is a pointer to a known structure in a randomized region, the attacker now knows where that region is loaded. The randomization still happened; the secrecy it provided didn't last.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ASLR randomizes layout
        ↓
Addresses are unknown
        ↓
Information disclosure vulnerability reveals a pointer
        ↓
Attacker calculates region's base address
        ↓
Layout becomes known for this execution
        ↓
ASLR's protection is reduced
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is why modern exploitation often involves chaining multiple vulnerabilities. The first step is finding and using an information leak to defeat ASLR. The second step is using that knowledge to make the actual corruption or redirection work.&lt;/p&gt;

&lt;p&gt;ASLR doesn't become worthless when information leaks exist, but it stops being the sole obstacle. The attacker's job gets harder overall, but not impossible.&lt;/p&gt;




&lt;h2&gt;
  
  
  ASLR and the Stack
&lt;/h2&gt;

&lt;p&gt;The stack holds function call frames, local variables, and the saved return state that lets a function know where to go when it finishes. In terms of exploitation, the stack has historically been a target because stack-based buffer overflows can overwrite saved return state, potentially redirecting execution.&lt;/p&gt;

&lt;p&gt;When the stack is randomized, the addresses of local variables and saved state change between runs. An attacker who wants to overwrite a specific location on the stack, or who wants to jump to a specific stack address, cannot rely on a hardcoded value.&lt;/p&gt;

&lt;p&gt;Stack canaries are a related but distinct protection. A canary is a value placed between local variables and saved return state. Before a function returns, the runtime checks that the canary hasn't changed. If it has, something overwrote it, and execution is terminated rather than redirected. This detects certain stack corruption independently of whether the attacker knows the stack's location.&lt;/p&gt;

&lt;p&gt;ASLR and stack canaries address different problems. ASLR makes the stack's location unpredictable. Canaries detect that the stack has been corrupted. Both can be present simultaneously, and modern systems typically deploy both.&lt;/p&gt;




&lt;h2&gt;
  
  
  ASLR and the Heap
&lt;/h2&gt;

&lt;p&gt;The heap is where dynamically allocated memory lives. Every object allocated with &lt;code&gt;malloc&lt;/code&gt; in C, or &lt;code&gt;new&lt;/code&gt; in C++, lands somewhere on the heap. The heap allocator manages free and used regions, tracking metadata about each allocation.&lt;/p&gt;

&lt;p&gt;Predictable heap addresses matter when an attacker wants to corrupt a specific heap object or craft a structure at a known location. Heap randomization changes where the heap begins and how the allocator distributes memory, introducing uncertainty about where specific objects will land.&lt;/p&gt;

&lt;p&gt;The heap is more complex than the stack in how its randomization interacts with exploitation, because heap layout depends on the allocation sequence throughout the program's lifetime, not just on a base address. But randomizing the heap's starting location is still a meaningful obstacle.&lt;/p&gt;




&lt;h2&gt;
  
  
  DEP/NX: A Different Problem
&lt;/h2&gt;

&lt;p&gt;ASLR and Data Execution Prevention (DEP), also called NX (No-eXecute) on systems that use that terminology, are often mentioned together. They solve different problems.&lt;/p&gt;

&lt;p&gt;ASLR answers the question: "Where are the things I need to reach?"&lt;/p&gt;

&lt;p&gt;DEP/NX answers the question: "Can I execute code in this memory region?"&lt;/p&gt;

&lt;p&gt;With DEP/NX, memory pages that hold data (the stack, the heap, most data regions) are marked non-executable at the hardware level. A process that tries to execute code from a non-executable page triggers a fault. This directly counteracts a classic technique of placing executable code into a data region and then redirecting execution there.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Without DEP/NX:
Attacker places code in stack/heap
Redirects execution there
Code runs

With DEP/NX:
Attacker places code in stack/heap
Redirects execution there
Hardware fault: memory is not executable
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;ASLR and DEP/NX are complementary. DEP/NX makes it harder to inject new executable code. ASLR makes the addresses of existing executable code harder to predict. Together they force attackers to look for other approaches, and those approaches face their own obstacles.&lt;/p&gt;




&lt;h2&gt;
  
  
  A Conceptual Scenario
&lt;/h2&gt;

&lt;p&gt;Consider a fictional application with a memory corruption vulnerability. Without any mitigations, an attacker who can trigger the vulnerability might be able to redirect execution by overwriting a return address, pointing it at useful existing code at a fixed, known address.&lt;/p&gt;

&lt;p&gt;With ASLR and PIE enabled, the same vulnerability exists. The same corruption is possible. But the useful code isn't at a predictable address. Every run, the layout is different. To reliably exploit the bug, the attacker now needs:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;A way to discover the current layout before or during exploitation.&lt;/li&gt;
&lt;li&gt;Enough time and attempts to use that information before the target changes.
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Application
     ↓
Memory Corruption Vulnerability
     ↓
Need to redirect execution somewhere useful
     ↓
ASLR: useful addresses are unknown
     ↓
Need an information leak to discover them
     ↓
Exploitation becomes a multi-step problem
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The underlying vulnerability is unchanged. What changed is the difficulty of turning it into reliable controlled behavior. That is what ASLR is designed to do.&lt;/p&gt;




&lt;h2&gt;
  
  
  How Much Protection Does ASLR Actually Provide?
&lt;/h2&gt;

&lt;p&gt;ASLR is a mitigation. It raises the cost and complexity of exploitation. It is not an elimination of exploitation risk.&lt;/p&gt;

&lt;p&gt;The practical protection depends on entropy (how many possible positions exist), whether PIE is enabled for the executable, whether the operating system randomizes all relevant regions, and whether information leaks are present that can reveal runtime addresses.&lt;/p&gt;

&lt;p&gt;On modern 64-bit operating systems with PIE-enabled executables and no information disclosure vulnerabilities, ASLR is a significant barrier. An attacker without a way to learn the runtime layout faces an enormous range of possible addresses.&lt;/p&gt;

&lt;p&gt;On older 32-bit systems, or systems with limited entropy, or applications compiled without PIE, the protection is weaker.&lt;/p&gt;

&lt;p&gt;ASLR is also not identical across platforms. Linux, Windows, and macOS all implement address randomization, but the entropy applied, the regions randomized, the loader behavior, and the interaction with executable formats differ. Comparing them requires looking at specific implementation details rather than assuming uniformity.&lt;/p&gt;




&lt;h2&gt;
  
  
  Defense in Depth
&lt;/h2&gt;

&lt;p&gt;ASLR is one layer in a broader defense model:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Memory-Safe Code
      +
ASLR (randomized layout)
      +
PIE (randomized executable base)
      +
DEP/NX (non-executable data)
      +
Stack Canaries (stack corruption detection)
      +
Control-Flow Integrity (restrict execution targets)
      +
Sandboxing (limit process capabilities)
      =
Significantly higher exploitation difficulty
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No single mitigation is expected to be impenetrable. The goal is layering defenses so that exploitation requires overcoming multiple independent obstacles. Defeating one mitigation leaves others intact.&lt;/p&gt;

&lt;p&gt;For developers, this means compiling executables with PIE and modern hardening flags, keeping toolchains and operating systems updated to benefit from improved mitigations, and not treating ASLR as a substitute for writing memory-safe code. A vulnerability that ASLR makes harder to exploit today might become more exploitable as techniques evolve.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Core Idea
&lt;/h2&gt;

&lt;p&gt;A memory address is only useful to an attacker if they can depend on it being there.&lt;/p&gt;

&lt;p&gt;ASLR attacks that dependability. The vulnerability can still exist. The memory can still be corruptible. But the address an attacker would need to make that corruption meaningful changes with every run.&lt;/p&gt;

&lt;p&gt;Turning a bug into a reliable exploit requires more than finding a flaw. It requires knowing the terrain. ASLR changes the terrain.&lt;/p&gt;

</description>
      <category>cybersecurity</category>
      <category>infosec</category>
      <category>security</category>
    </item>
    <item>
      <title>What Is Privilege Escalation? How Can a Low-Privilege User Become Root?</title>
      <dc:creator>Aditya Sharma</dc:creator>
      <pubDate>Sat, 26 Sep 2026 10:46:56 +0000</pubDate>
      <link>https://dev.to/aditya_d_sharma/what-is-privilege-escalation-how-can-a-low-privilege-user-become-root-1la3</link>
      <guid>https://dev.to/aditya_d_sharma/what-is-privilege-escalation-how-can-a-low-privilege-user-become-root-1la3</guid>
      <description>&lt;p&gt;You already have access to the machine. The problem is that you're only supposed to have access to some of it.&lt;/p&gt;

&lt;p&gt;You can read your own files. Run programs. Use network services. But you can't read other users' files, modify system configuration, install software globally, or change how the kernel behaves. Those actions require privileges you don't have.&lt;/p&gt;

&lt;p&gt;Privilege escalation is what happens when that boundary breaks. An attacker who starts with limited access ends up with more authority than they were supposed to have, sometimes far more. Understanding how that happens requires understanding what privileges are and why operating systems enforce them in the first place.&lt;/p&gt;




&lt;h2&gt;
  
  
  What a Privilege Actually Is
&lt;/h2&gt;

&lt;p&gt;A privilege is a permission to take an action the system doesn't allow by default.&lt;/p&gt;

&lt;p&gt;Ordinary processes run under user accounts with restricted permissions. They can interact with resources their owner can access, and nothing else. A privileged account can do things that cross that boundary: install software, read any file on the system, manage other users, change system-wide configuration, load kernel modules.&lt;/p&gt;

&lt;p&gt;Every operating system separates these levels of access because trust is not binary. Not every user should be able to do everything. Not every program needs access to every resource. Privilege separation is the mechanism that enforces this.&lt;/p&gt;

&lt;p&gt;Privilege escalation occurs when an actor that was supposed to operate at one privilege level ends up operating at a higher one without being authorized to do so.&lt;/p&gt;




&lt;h2&gt;
  
  
  Vertical vs Horizontal Escalation
&lt;/h2&gt;

&lt;p&gt;It helps to distinguish two kinds of privilege escalation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Vertical escalation&lt;/strong&gt; means moving up the privilege hierarchy:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Normal user account
        ↓
Administrator or root
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A normal user becomes able to act with the authority of a privileged account. This is the most recognized form and the main focus of this article.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Horizontal escalation&lt;/strong&gt; means moving sideways across accounts at the same privilege level:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User A's account
        ↓
User B's data and resources
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;User A can access files, sessions, or data belonging to User B, even though both are ordinary users. This is equally a security failure, but the mechanism and impact are different.&lt;/p&gt;

&lt;p&gt;Both matter. A horizontal escalation can expose sensitive data belonging to another user. A vertical escalation can expose the entire system.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why Privilege Separation Exists
&lt;/h2&gt;

&lt;p&gt;The principle of least privilege is one of the oldest ideas in security: give each identity and process only the permissions it needs to do its job, and nothing more.&lt;/p&gt;

&lt;p&gt;Operating systems implement this through account hierarchies:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Guest / Unauthenticated
          ↓
Standard User
          ↓
Privileged Service / Application
          ↓
Administrator / Root
          ↓
Kernel
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each level has more authority than the one above it, and that authority is intentionally restricted. A standard user account cannot modify system binaries because that would allow any compromised user process to affect every other user and every program on the machine. A service that only needs to read log files shouldn't have write access to configuration files.&lt;/p&gt;

&lt;p&gt;When privilege separation works correctly, a compromise at one level doesn't automatically compromise everything else. When it fails, a foothold at a low level can become control of the entire system.&lt;/p&gt;




&lt;h2&gt;
  
  
  How Privilege Escalation Happens
&lt;/h2&gt;

&lt;p&gt;Privilege escalation is not a single technique. It is a category of outcomes that can result from very different root causes.&lt;/p&gt;

&lt;h3&gt;
  
  
  Misconfigurations
&lt;/h3&gt;

&lt;p&gt;Perhaps the most common source. Someone configures a service, file, or resource with broader permissions than intended. A service that runs as root to simplify setup. A directory that is writable by all users. A scheduled task that runs a script in a user-writable location. None of these are software bugs. They are administrative mistakes, but their effect is the same: a path from low privilege to high.&lt;/p&gt;

&lt;h3&gt;
  
  
  Vulnerable Privileged Programs
&lt;/h3&gt;

&lt;p&gt;Some programs legitimately need elevated privileges to do their job. A program that manages network interfaces needs permissions an ordinary user doesn't have. A program that changes file ownership needs permissions an ordinary user doesn't have.&lt;/p&gt;

&lt;p&gt;If one of these privileged programs contains a security flaw, it can become a path to privilege escalation. An attacker who can manipulate the program's behavior through its inputs, its environment, or a race condition can potentially cause it to perform operations on their behalf that they couldn't perform directly.&lt;/p&gt;

&lt;h3&gt;
  
  
  SUID Binaries on Linux
&lt;/h3&gt;

&lt;p&gt;Linux has a mechanism called the setuid bit (SUID). When a program with SUID set is executed, it runs with the effective user identity of its owner rather than the user who executed it. Programs like &lt;code&gt;passwd&lt;/code&gt; use this mechanism: a normal user needs to run &lt;code&gt;passwd&lt;/code&gt; to change their password, but the actual operation requires writing to system files that only root can modify.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User executes SUID program
        ↓
Process runs with elevated effective privileges
        ↓
Program performs privileged operation safely
        ↓
Operation completes, privileges drop
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is legitimate and necessary. The problem arises when a SUID program has a security flaw. If the program can be manipulated into performing unintended privileged operations, it becomes a bridge across the privilege boundary:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User finds SUID program with a flaw
        ↓
Manipulates program behavior
        ↓
Program performs unintended privileged action
        ↓
User gains unauthorized elevated access
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The distinction matters: SUID is not inherently dangerous, but every SUID binary is a privileged component that needs to be secure against misuse.&lt;/p&gt;

&lt;h3&gt;
  
  
  Linux Capabilities
&lt;/h3&gt;

&lt;p&gt;Modern Linux systems can assign specific privileged operations to processes without giving them full root access. These are called capabilities.&lt;/p&gt;

&lt;p&gt;Instead of a process being either ordinary or fully root, capabilities allow finer control. A process might have the capability to bind to privileged network ports without having the capability to read every file on the system. A process might have the capability to change its own group without having broad administrative access.&lt;/p&gt;

&lt;p&gt;The security concern is that capabilities, when incorrectly assigned, can still create privilege escalation paths. A process with the capability to load kernel modules has enormous power over the system, arguably more than many root operations. Capabilities reduce the blast radius of privilege in well-designed systems, but only when assigned thoughtfully.&lt;/p&gt;

&lt;h3&gt;
  
  
  Windows Services and Access Tokens
&lt;/h3&gt;

&lt;p&gt;On Windows, every process runs within a security context defined by an access token. The token carries the identity of the account, its group memberships, and the privileges assigned to it.&lt;/p&gt;

&lt;p&gt;Standard users have limited tokens. Administrators have broader tokens. The SYSTEM account, which many Windows services run under, typically has more access than even a local administrator to system internals.&lt;/p&gt;

&lt;p&gt;Privilege escalation on Windows often involves getting code to run within a more privileged security context:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A service running as SYSTEM that processes user-supplied data insecurely&lt;/li&gt;
&lt;li&gt;A scheduled task running with elevated privileges from a location a normal user can modify&lt;/li&gt;
&lt;li&gt;An administrative tool that can be manipulated into performing privileged operations
The underlying pattern is the same as on Linux: find a privileged component and find a way to influence what it does.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Kernel Vulnerabilities
&lt;/h3&gt;

&lt;p&gt;The kernel sits at the bottom of the privilege hierarchy. It manages everything below the software abstraction: memory, process scheduling, hardware access, the enforcement of all the access controls discussed above.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User Space Process
        ↓
System Call Interface
        ↓
Kernel
        ↓
Hardware
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A vulnerability in the kernel is particularly significant because the kernel's trust boundary is the strongest boundary in the system. A process that successfully exploits a kernel vulnerability can potentially gain control over the kernel's execution, which means control over the entire machine. All the user-space privilege separation becomes irrelevant once the attacker is operating at the kernel level.&lt;/p&gt;

&lt;h3&gt;
  
  
  Credentials and Configuration
&lt;/h3&gt;

&lt;p&gt;Not every privilege escalation involves a software flaw. An attacker who finds credentials for a privileged account, whether stored in a configuration file, a script, an environment variable, or a database, has effectively escalated their privileges through credential theft.&lt;/p&gt;

&lt;p&gt;Similarly, finding that a root cron job processes attacker-writable data, or that a privileged service reads configuration from an attacker-writable location, achieves the same result through a configuration weakness rather than a code vulnerability.&lt;/p&gt;




&lt;h2&gt;
  
  
  A Conceptual Privilege Escalation Chain
&lt;/h2&gt;

&lt;p&gt;To make this concrete without being operational, here is a generic scenario.&lt;/p&gt;

&lt;p&gt;A system has two components: a low-privilege application that accepts user input and a privileged service that processes data from the application. The separation is intentional. The privileged service is supposed to do one specific thing, safely.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Low-privilege user
        ↓
Low-privilege application
        ↓
Data passed to privileged service
        ↓
Security flaw in privileged service
        ↓
Service performs unintended privileged operation
        ↓
Attacker influences outcome
        ↓
Effective escalation to higher privilege level
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The attacker didn't break the privilege boundary directly. They found a component that legitimately crossed it and made it do something unintended. This pattern, using a privileged component as a bridge, is central to how privilege escalation works across many different contexts.&lt;/p&gt;




&lt;h2&gt;
  
  
  Vulnerabilities vs Misconfigurations vs Credentials
&lt;/h2&gt;

&lt;p&gt;These three root causes produce similar outcomes but require different responses.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A software vulnerability&lt;/strong&gt; means a privileged program contains a flaw that allows unintended behavior. The fix is patching the software.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A misconfiguration&lt;/strong&gt; means a permission or trust relationship is set up incorrectly. The fix is correcting the configuration and understanding why the mistake was made.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A credential issue&lt;/strong&gt; means an attacker obtained credentials belonging to a more privileged account. The fix involves credential rotation, better storage practices, and understanding how the credentials were exposed.&lt;/p&gt;

&lt;p&gt;Treating a misconfiguration like a software vulnerability, or vice versa, leads to incomplete responses. A patch doesn't fix a misconfigured service. Fixing a configuration doesn't prevent the same code flaw from being exploited through a different path.&lt;/p&gt;




&lt;h2&gt;
  
  
  How Defenders Reduce the Risk
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Least privilege as a design principle.&lt;/strong&gt; Every user, service, and process should have only the permissions it needs. This limits the damage that any single compromised component can cause. A service running as a standard user with narrow permissions provides a much smaller bridge for privilege escalation than a service running as root.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Patch management.&lt;/strong&gt; Privileged programs that contain vulnerabilities need to be updated. Kernel vulnerabilities in particular need rapid response.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Audit SUID binaries and capabilities.&lt;/strong&gt; Every SUID binary and every process with elevated capabilities is a potential pivot point. Know what exists on each system and whether those elevated permissions are actually necessary.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Secure service configuration.&lt;/strong&gt; Services running with elevated privileges should read from, write to, and process data from locations that unprivileged users cannot influence.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Credential protection.&lt;/strong&gt; Passwords, tokens, API keys, and other credentials for privileged accounts should not appear in scripts, configuration files, environment variables, or log output where lower-privileged processes can read them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Monitor privileged actions.&lt;/strong&gt; Sudden changes in account privilege levels, unexpected processes running under privileged identities, or unusual access to sensitive system resources can be early indicators of privilege escalation attempts.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Isolation.&lt;/strong&gt; Containers, virtual machines, and sandboxing mechanisms add additional layers of isolation, so a compromised component has fewer paths to privileged resources.&lt;/p&gt;

&lt;p&gt;Least privilege doesn't eliminate privilege escalation as a possibility. A service needs some privileges to function, and those privileges can be abused. But minimizing privileges minimizes the available attack surface and the potential impact of any successful escalation.&lt;/p&gt;




&lt;h2&gt;
  
  
  Detecting Privilege Escalation
&lt;/h2&gt;

&lt;p&gt;On a running system, defenders look for signs that privilege boundaries have been crossed:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;New administrative accounts that weren't created through an expected provisioning process&lt;/li&gt;
&lt;li&gt;Processes suddenly running under privileged accounts or with unexpected tokens&lt;/li&gt;
&lt;li&gt;Unexpected modifications to files only privileged processes should touch&lt;/li&gt;
&lt;li&gt;System logs showing privilege changes without corresponding administrative activity&lt;/li&gt;
&lt;li&gt;Unusual access to system directories, shadow password files, service credentials, or kernel interfaces&lt;/li&gt;
&lt;li&gt;Security tools reporting that a process is operating with elevated effective privileges it shouldn't have
Privilege escalation typically doesn't announce itself. An attacker who has successfully escalated will often try to maintain persistence quietly. The window for detection is often before the escalation completes: catching the reconnaissance, the probing of privileged services, or the early stages of exploitation.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Privilege Escalation vs Initial Access
&lt;/h2&gt;

&lt;p&gt;It is worth being explicit about where privilege escalation fits in the broader picture.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Initial access&lt;/strong&gt; is how an attacker gains any foothold on a system: exploiting a public-facing service, using stolen credentials, social engineering, a vulnerable client application. At this point, the attacker may have very limited access.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Privilege escalation&lt;/strong&gt; is what happens after. The attacker has some access and wants more.&lt;/p&gt;

&lt;p&gt;The reason this distinction matters is that an initial access with limited privileges is not necessarily a contained incident. A normal user account on a system is often a starting point, not a final position. Defenders who focus entirely on preventing initial access sometimes underestimate the importance of limiting what an attacker can do once they're in.&lt;/p&gt;

&lt;p&gt;A system where a compromised low-privilege account can quickly become root is more dangerous than a system where a compromised low-privilege account remains constrained. The goal of privilege separation is that limited access stays limited.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Deeper Lesson
&lt;/h2&gt;

&lt;p&gt;Operating systems make a promise: different accounts get different levels of authority, and those boundaries hold. Privilege escalation is what happens when that promise breaks.&lt;/p&gt;

&lt;p&gt;The break can come from a software flaw in a privileged component. It can come from a configuration mistake that makes a trust relationship too broad. It can come from credentials that were stored where they shouldn't have been.&lt;/p&gt;

&lt;p&gt;In every case, the mechanism is the same: a path from a lower privilege level to a higher one that shouldn't exist.&lt;/p&gt;

&lt;p&gt;The attacker's job is to find that path. The defender's job is to make sure as few paths as possible exist, and that any path that does exist is difficult to traverse without being detected.&lt;/p&gt;

&lt;p&gt;The security goal isn't just keeping attackers out. It's ensuring that wherever they enter, they cannot reach the parts of the system that matter most.&lt;/p&gt;

</description>
      <category>cybersecurity</category>
      <category>infosec</category>
      <category>linux</category>
      <category>security</category>
    </item>
    <item>
      <title>What Is Bootstrap Aggregation? How Does Bagging Make Machine Learning Models More Robust?</title>
      <dc:creator>Aditya Sharma</dc:creator>
      <pubDate>Fri, 25 Sep 2026 05:30:43 +0000</pubDate>
      <link>https://dev.to/aditya_d_sharma/what-is-bootstrap-aggregation-how-does-bagging-make-machine-learning-models-more-robust-8fm</link>
      <guid>https://dev.to/aditya_d_sharma/what-is-bootstrap-aggregation-how-does-bagging-make-machine-learning-models-more-robust-8fm</guid>
      <description>&lt;p&gt;A single ML model is just one opinion. It learned from one arrangement of training data, made specific decisions about what patterns matter, and internalized the noise in that particular dataset alongside the signal.&lt;/p&gt;

&lt;p&gt;What if, instead of trusting that one opinion, you trained dozens of models on slightly different versions of the same data and took a vote?&lt;/p&gt;

&lt;p&gt;That's the core idea behind Bootstrap Aggregation, usually called Bagging. It doesn't try to build a perfect model. It builds many imperfect models with deliberate variation between them, then combines their predictions into something more stable than any individual model could produce.&lt;/p&gt;




&lt;h2&gt;
  
  
  What "Bootstrap" Actually Means
&lt;/h2&gt;

&lt;p&gt;Before the machine learning part, there's a statistical concept worth understanding.&lt;/p&gt;

&lt;p&gt;A bootstrap sample is a new dataset created by sampling with replacement from the original dataset. The resulting sample has the same number of rows as the original, but it's not identical. Because each observation is drawn independently with replacement, some rows from the original dataset will appear multiple times, and others won't appear at all.&lt;/p&gt;

&lt;p&gt;Here's a small example. Say your original training set has five examples:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Original: [A, B, C, D, E]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A bootstrap sample of size 5 might look like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Bootstrap 1: [A, A, C, D, D]   (B and E omitted; A and D appear twice)
Bootstrap 2: [B, C, C, E, E]   (A and D omitted; C and E appear twice)
Bootstrap 3: [A, B, B, D, E]   (C omitted; B appears twice)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each bootstrap sample represents a slightly different picture of the same underlying data. The statistical theory behind this goes back to work on estimating properties of a distribution from a sample, but for our purposes the key point is this: by sampling with replacement, you get genuine variation between samples without needing more data.&lt;/p&gt;

&lt;p&gt;For a dataset with n observations, any given bootstrap sample will omit roughly 1/e of the original observations on average, which works out to about 36.8%. Those omitted observations become useful later.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Bagging Pipeline
&lt;/h2&gt;

&lt;p&gt;The full process looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Original Training Dataset
          |
          +----------+----------+----------+
          |          |          |          |
     Bootstrap 1  Bootstrap 2  Bootstrap 3  ...
          |          |          |          |
       Model 1    Model 2    Model 3    ...
          |          |          |          |
          +----------+----------+----------+
                         |
                    Aggregation
                         |
                  Final Prediction
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Bootstrap sampling:&lt;/strong&gt; Create k bootstrap samples from the original training set. Each sample has the same number of rows as the original but is different due to sampling with replacement.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Train independently:&lt;/strong&gt; Train one model on each bootstrap sample. These models are trained completely independently of each other. This is important: they don't communicate, don't share gradients, and don't adjust based on each other's performance.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Generate predictions:&lt;/strong&gt; For a new input, run it through all k models and collect their predictions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Aggregate:&lt;/strong&gt; Combine those predictions into one.&lt;/p&gt;

&lt;p&gt;For regression, aggregation typically means averaging the numerical outputs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Final prediction = (pred_1 + pred_2 + ... + pred_k) / k
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For classification, it typically means majority voting:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Final prediction = most common class among (pred_1, pred_2, ..., pred_k)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Why Does Bagging Actually Work?
&lt;/h2&gt;

&lt;p&gt;This is the important part. Bagging isn't magic. It works because of a specific property of aggregating independent, diverse predictions.&lt;/p&gt;

&lt;p&gt;Start with the concept of &lt;strong&gt;variance&lt;/strong&gt; in the context of a model. A high-variance model is sensitive to the specific training data it saw. Train it on one dataset, and it learns one set of patterns. Train it on a slightly different dataset, and it might make quite different predictions. Decision trees with no depth limit are a classic example: they'll fit their training data nearly perfectly but their predictions can shift dramatically if the training data changes slightly.&lt;/p&gt;

&lt;p&gt;Now consider what happens when you average multiple high-variance predictions.&lt;/p&gt;

&lt;p&gt;If the errors made by individual models are uncorrelated (or at least not perfectly correlated), their random mistakes tend to cancel out when averaged. One model overestimates on a particular region of the input space; another underestimates. Their average is closer to the truth than either individual prediction.&lt;/p&gt;

&lt;p&gt;Formally, if you have k models each with variance sigma squared and zero covariance between their errors, the variance of their average is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Var(average) = sigma^2 / k
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;As k increases, the variance of the ensemble prediction shrinks. This is the same reason that averaging many noisy measurements of a quantity gives you a better estimate than relying on a single noisy measurement.&lt;/p&gt;

&lt;p&gt;The important caveat is the "uncorrelated" part. If all your models make the same mistakes because they're too similar to each other, averaging them doesn't help much. The formula for the average variance when models have pairwise correlation rho is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Var(average) = rho * sigma^2 + (1 - rho) * sigma^2 / k
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When rho approaches 1 (models are identical), the variance reduction from k models disappears. This is why the bootstrap sampling step matters: it deliberately creates variation between the training sets, which creates variation between the models, which reduces correlation between their errors.&lt;/p&gt;

&lt;p&gt;Bagging primarily reduces variance. It does not reliably reduce bias. If all your models are systematically wrong in the same direction, averaging them will still give you a systematically wrong answer.&lt;/p&gt;




&lt;h2&gt;
  
  
  A Concrete Example
&lt;/h2&gt;

&lt;p&gt;Say you're classifying whether an email is spam or not, and you train five small decision trees using bagging. For a particular email, they each make a prediction:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Model 1: Spam
Model 2: Not Spam
Model 3: Spam
Model 4: Spam
Model 5: Not Spam
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Majority vote: Spam appears three times, Not Spam appears twice. Final prediction: Spam.&lt;/p&gt;

&lt;p&gt;For a regression example, say you're predicting house prices and five models return:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Model 1: $412,000
Model 2: $395,000
Model 3: $428,000
Model 4: $407,000
Model 5: $418,000
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Average: ($412,000 + $395,000 + $428,000 + $407,000 + $418,000) / 5 = $412,000&lt;/p&gt;

&lt;p&gt;Any one of these models might be somewhat off. Their average tends to be closer to the true value than most of the individual predictions, assuming the individual errors aren't all biased in the same direction.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why Not Just Train the Same Model Repeatedly?
&lt;/h2&gt;

&lt;p&gt;The obvious question: why go through the trouble of bootstrap sampling? Why not just train the same model multiple times on the same data?&lt;/p&gt;

&lt;p&gt;If your model is deterministic and you train it on the same dataset, you'll get the same model every time. Averaging five copies of the same model gives you exactly that model, with no variance reduction at all.&lt;/p&gt;

&lt;p&gt;Bootstrap sampling is what creates the necessary diversity. Each sample contains a different subset of the training examples, with different duplications. A model trained on Bootstrap 1 never sees some of the original examples and sees others multiple times. A model trained on Bootstrap 2 has a different combination of repeated and missing examples. The result is k genuinely different models that learned slightly different things.&lt;/p&gt;

&lt;p&gt;A secondary benefit: because each model is trained independently, the training process can be parallelized. Train k models simultaneously across k machines or k CPU cores. This is not possible with methods where models must be trained sequentially.&lt;/p&gt;




&lt;h2&gt;
  
  
  Out-of-Bag Samples
&lt;/h2&gt;

&lt;p&gt;Remember that each bootstrap sample omits roughly 36.8% of the original training examples. Those omitted examples are called out-of-bag (OOB) samples for that particular model.&lt;/p&gt;

&lt;p&gt;Since the model was never trained on its out-of-bag examples, those examples can be used to evaluate the model's performance on data it hasn't seen. For each original training example, it was out-of-bag for some fraction of the k models. You can collect predictions from those models (the ones that didn't train on this example) and use them to estimate the model's generalization error.&lt;/p&gt;

&lt;p&gt;Aggregating these OOB predictions across the entire training set gives you an OOB error estimate: an estimate of how well the ensemble performs on unseen data, without needing to hold out a separate validation set.&lt;/p&gt;

&lt;p&gt;This is a useful property. In situations where you want to use as much data as possible for training, OOB evaluation lets you get a generalization estimate without sacrificing any training examples to a held-out set.&lt;/p&gt;

&lt;p&gt;OOB error is not always equivalent to k-fold cross-validation or other evaluation strategies, and the two can diverge depending on the dataset and model. But it's a practically useful tool that comes essentially for free when you're already doing bagging.&lt;/p&gt;




&lt;h2&gt;
  
  
  Bagging and Random Forests
&lt;/h2&gt;

&lt;p&gt;Random Forest is one of the most successful applications of the bagging idea, but it's not identical to bagging.&lt;/p&gt;

&lt;p&gt;Standard bagging: take bootstrap samples, train full models independently on each, aggregate predictions.&lt;/p&gt;

&lt;p&gt;Random Forest: also uses bootstrap samples and independent training, but adds one more source of randomness. At each split point in each decision tree, instead of considering all available features, it considers only a random subset of features. This additional randomization makes the trees even more different from each other.&lt;/p&gt;

&lt;p&gt;Why does this help? In standard bagging with decision trees, if there's one very strong predictor in the dataset, most of the trees will use that predictor near the top of each tree. The resulting trees will be more correlated with each other than bootstrap sampling alone would suggest, because they're all making similar early splits. Restricting feature selection at each split breaks this correlation and lets other predictors contribute meaningfully across different trees.&lt;/p&gt;

&lt;p&gt;Random Forest therefore benefits from two distinct sources of diversity: bootstrap sampling (different training examples) and random feature selection (different features considered at each split). The result is typically lower correlation between trees and better ensemble performance than standard bagging with decision trees.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Bagging:
  Bootstrap samples + independent training + aggregation

Random Forest:
  Bootstrap samples + independent training + random feature subsets + aggregation
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Bagging vs. Boosting
&lt;/h2&gt;

&lt;p&gt;Boosting is another ensemble method, but it operates on entirely different principles.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Bagging&lt;/th&gt;
&lt;th&gt;Boosting&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Training order&lt;/td&gt;
&lt;td&gt;Independent, parallel&lt;/td&gt;
&lt;td&gt;Sequential, each model trained on results of previous&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Data weighting&lt;/td&gt;
&lt;td&gt;Each model sees a bootstrap sample&lt;/td&gt;
&lt;td&gt;Examples weighted by difficulty; hard examples emphasized&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Primary effect&lt;/td&gt;
&lt;td&gt;Reduces variance&lt;/td&gt;
&lt;td&gt;Reduces bias (and variance in some formulations)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model relationship&lt;/td&gt;
&lt;td&gt;Each model stands alone&lt;/td&gt;
&lt;td&gt;Each model corrects residuals of the previous ensemble&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Examples&lt;/td&gt;
&lt;td&gt;Random Forest&lt;/td&gt;
&lt;td&gt;AdaBoost, Gradient Boosting, XGBoost&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;In boosting, model 2 is trained on what model 1 got wrong. Model 3 is trained on what the combination of model 1 and model 2 got wrong. Each new model is guided by the failures of the previous ones, which makes the ensemble increasingly accurate on hard examples.&lt;/p&gt;

&lt;p&gt;Bagging makes models independent on purpose. Boosting makes models dependent on purpose.&lt;/p&gt;

&lt;p&gt;Both can produce strong ensembles, but they're solving different problems. Bagging is most useful when your base model is high-variance and unstable. Boosting can sometimes help when your base model is too simple (high-bias), though this depends significantly on the specific boosting algorithm and configuration.&lt;/p&gt;




&lt;h2&gt;
  
  
  When Bagging Helps (and When It Doesn't)
&lt;/h2&gt;

&lt;p&gt;Bagging is most effective when the base model is a high-variance, low-bias learner. Deep decision trees are the canonical example: they overfit individual training sets aggressively, and bagging provides substantial stability gains.&lt;/p&gt;

&lt;p&gt;If your base model is already stable (low-variance), bagging provides little benefit. A linear regression model trained on reasonable data will produce nearly the same predictions regardless of small variations in the training set. Averaging many nearly identical linear models doesn't give you much.&lt;/p&gt;

&lt;p&gt;Bagging also doesn't fix bias. If your model consistently underestimates a quantity because it lacks the capacity to fit the true relationship, averaging more such models will still underestimate.&lt;/p&gt;

&lt;p&gt;The conditions where bagging provides meaningful improvement:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;High-variance base models&lt;/li&gt;
&lt;li&gt;Datasets where different bootstrap samples genuinely change model behavior&lt;/li&gt;
&lt;li&gt;Sufficient diversity between the resulting models
The conditions where bagging helps less:&lt;/li&gt;
&lt;li&gt;Already-stable base models&lt;/li&gt;
&lt;li&gt;Highly correlated base models despite bootstrap sampling&lt;/li&gt;
&lt;li&gt;Problems where bias is the primary issue&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  The Deeper Intuition
&lt;/h2&gt;

&lt;p&gt;Training one model is like asking one person for directions. They might know the area well, or they might give you confidently wrong advice based on a bad experience they had once.&lt;/p&gt;

&lt;p&gt;Asking ten people who've each explored slightly different parts of the area gives you more coverage. Their individual answers might vary, but their collective answer tends to be more reliable than any single person's, especially if they each have genuine knowledge rather than all reading from the same map.&lt;/p&gt;

&lt;p&gt;Bagging does the same thing with models. It doesn't make any individual model better. It creates genuine variation through bootstrap sampling, exploits the statistical property that averaging independent noisy estimates reduces variance, and produces an ensemble whose stability comes from the diversity of its components.&lt;/p&gt;

&lt;p&gt;The robustness isn't in any single tree or any single decision. It emerges from the agreement across models that each learned from a slightly different slice of the data.&lt;/p&gt;

</description>
      <category>algorithms</category>
      <category>datascience</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>What Is a Polyglot File? How Can One File Be Valid as Multiple File Formats?</title>
      <dc:creator>Aditya Sharma</dc:creator>
      <pubDate>Thu, 24 Sep 2026 06:02:21 +0000</pubDate>
      <link>https://dev.to/aditya_d_sharma/what-is-a-polyglot-file-how-can-one-file-be-valid-as-multiple-file-formats-ppd</link>
      <guid>https://dev.to/aditya_d_sharma/what-is-a-polyglot-file-how-can-one-file-be-valid-as-multiple-file-formats-ppd</guid>
      <description>&lt;p&gt;You upload one file. The upload validator inspects the bytes and says: "That's a JPEG image. Safe to accept."&lt;/p&gt;

&lt;p&gt;Later, a different component in the same system opens the exact same bytes. It says: "That's a ZIP archive."&lt;/p&gt;

&lt;p&gt;Neither component is broken. Neither is guessing. Both are applying real parsing logic to the same byte sequence, and both are finding something valid.&lt;/p&gt;

&lt;p&gt;The question is: how can one sequence of bytes legitimately satisfy the syntax rules of two completely different file formats?&lt;/p&gt;




&lt;h2&gt;
  
  
  A File Format Is a Grammar for Bytes
&lt;/h2&gt;

&lt;p&gt;Before getting into polyglots, it helps to be precise about what a file format actually is.&lt;/p&gt;

&lt;p&gt;A file format is a specification that describes how bytes should be structured and interpreted. It defines things like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Where the file begins (often a recognizable sequence of bytes called a magic number or file signature)&lt;/li&gt;
&lt;li&gt;How the header is structured&lt;/li&gt;
&lt;li&gt;Where the actual content data lives&lt;/li&gt;
&lt;li&gt;How size, offset, and length values are encoded&lt;/li&gt;
&lt;li&gt;Where the file ends, or what markers signal the end of meaningful content
A PDF begins with &lt;code&gt;%PDF-&lt;/code&gt;. A JPEG begins with &lt;code&gt;FF D8 FF&lt;/code&gt;. A ZIP typically begins with &lt;code&gt;PK&lt;/code&gt; (the initials of Phil Katz). These are not conventions the parser ignores. They are required structural elements that a conforming parser expects to find.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A filename extension like &lt;code&gt;.jpg&lt;/code&gt; or &lt;code&gt;.pdf&lt;/code&gt; is just a label. It tells the operating system which application to open the file with, but it doesn't affect the bytes themselves. A parser that receives raw bytes typically doesn't look at the filename. It looks at the content.&lt;/p&gt;




&lt;h2&gt;
  
  
  What a Polyglot File Actually Is
&lt;/h2&gt;

&lt;p&gt;A polyglot file is a byte sequence that simultaneously satisfies the structural requirements of more than one file format.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Same byte sequence
        ↓
 ┌──────────────────────┐
 ↓                      ↓
Parser A              Parser B
 ↓                      ↓
Valid format A        Valid format B
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The file has not been corrupted. The bytes have been deliberately (or sometimes accidentally) arranged so that two different parsers, each applying its own grammar, both find a valid structure.&lt;/p&gt;

&lt;p&gt;This is different from a renamed file. Renaming &lt;code&gt;malware.exe&lt;/code&gt; to &lt;code&gt;photo.jpg&lt;/code&gt; is extension spoofing. The bytes inside don't satisfy any JPEG parser. A robust validator that actually parses the content will reject it. A polyglot is more subtle: the bytes genuinely satisfy more than one parser's expectations.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why Two Formats Can Coexist in One File
&lt;/h2&gt;

&lt;p&gt;Different file formats make different structural assumptions. Those differences create space for coexistence.&lt;/p&gt;

&lt;p&gt;Some parsers care only about the beginning of the file. If the magic bytes match, they accept the file. They may not validate every subsequent byte.&lt;/p&gt;

&lt;p&gt;Some parsers care about structures at specific offsets. A format might require its header at byte 0, then allow arbitrary content until a particular marker appears.&lt;/p&gt;

&lt;p&gt;Some formats define what they are interested in and treat anything else as ignorable. A parser that encounters bytes it does not recognize in a region it considers padding or metadata may simply skip past them.&lt;/p&gt;

&lt;p&gt;Some formats place their critical structures near the end of the file. ZIP archives, for example, locate their central directory at the end. A parser reading a ZIP starts near the end-of-central-directory signature rather than at byte 0. This means the beginning of the file is largely irrelevant to a ZIP parser.&lt;/p&gt;

&lt;p&gt;These structural properties create regions where two formats can coexist without conflict.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Byte offset 0
│
├── Format A header (magic bytes, required header)
│
├── Region ignored by Format A
│   ↕
│   Format B-compatible structures live here
│
├── Shared data
│
└── Format A/B terminal structures
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The exact possibilities depend entirely on the specific formats involved. A JPEG parser and a ZIP parser make very different demands on the same byte sequence, and those demands happen to be compatible in ways that allow a single file to satisfy both.&lt;/p&gt;




&lt;h2&gt;
  
  
  Magic Bytes Are Not the Same as Full Validation
&lt;/h2&gt;

&lt;p&gt;A common validation approach checks the first few bytes of a file against known signatures. If the file starts with &lt;code&gt;FF D8 FF&lt;/code&gt;, it's probably a JPEG. Accept it.&lt;/p&gt;

&lt;p&gt;This answers one question: "Does the file begin like format X?" That is not the same as "Does the entire file conform to format X?"&lt;/p&gt;

&lt;p&gt;A file can begin with valid JPEG magic bytes and still contain structures elsewhere that another parser will successfully interpret as a different format. A thorough validator parses the entire file using the expected format's grammar: validating headers, checking that declared lengths match actual content, ensuring all structures fall within expected boundaries, and rejecting anything that doesn't conform. That is a fundamentally different operation from reading the first four bytes.&lt;/p&gt;




&lt;h2&gt;
  
  
  When Parser Disagreement Becomes a Security Problem
&lt;/h2&gt;

&lt;p&gt;A polyglot file by itself is not inherently malicious. The security concern arises when different components in the same system interpret the same bytes differently, and those components sit at different points in a security decision.&lt;/p&gt;

&lt;p&gt;Consider a typical file upload flow:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User uploads file
        ↓
Upload validator (checks magic bytes, MIME type, extension)
        ↓
"Looks like a JPEG. Accepted."
        ↓
File stored on disk
        ↓
Another component processes the same file
        ↓
This component finds a different valid structure in the bytes
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The upload validator made a decision about what the file is. A downstream component made a different decision about the same bytes. The security boundary was crossed using the conclusion from the first parser, but the second parser operates by different rules.&lt;/p&gt;

&lt;p&gt;This is the key insight: a security check is only meaningful if the component enforcing the check and the component eventually consuming the file agree on what the file is.&lt;/p&gt;




&lt;h2&gt;
  
  
  Polyglots and File Upload Validation
&lt;/h2&gt;

&lt;p&gt;Security-sensitive file upload handling typically involves several layers of validation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Extension checks&lt;/strong&gt; verify that the filename ends with an expected suffix. These are easy to spoof and should not be the primary defense.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;MIME type checks&lt;/strong&gt; look at the &lt;code&gt;Content-Type&lt;/code&gt; header in the upload request. This is client-supplied data and equally easy to manipulate.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Magic byte checks&lt;/strong&gt; inspect the first bytes of the file. More reliable than the above, but as discussed, insufficient on their own.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Full format parsing&lt;/strong&gt; validates the entire file structure using a library that implements the expected format's grammar. This is more robust but depends on the library being used correctly and on the format being parsed matching what downstream components will use.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Transcoding and re-encoding&lt;/strong&gt; take a different approach: instead of validating the uploaded file, the system decodes it and immediately re-encodes it into a clean, known-good representation. If someone uploads an image, the server converts it to a pixel buffer and writes a fresh image file from scratch. Any structures in the original file that didn't belong to the image data are discarded in the process. This approach is strong precisely because it doesn't need to understand everything the attacker might have embedded.&lt;/p&gt;

&lt;p&gt;The limitation with any single-stage validation is that it captures only one interpretation of the bytes. The component doing the validation may reach a different conclusion than the component eventually consuming the file.&lt;/p&gt;




&lt;h2&gt;
  
  
  Polyglot Files vs. Related Concepts
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Extension spoofing:&lt;/strong&gt; A file with a misleading name. The bytes don't satisfy the spoofed format. A parser-level check catches it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;MIME-type spoofing:&lt;/strong&gt; A file uploaded with a false &lt;code&gt;Content-Type&lt;/code&gt; header. Also client-supplied and trivially manipulated.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Parser differential:&lt;/strong&gt; Two components receiving the same input and interpreting it differently. A polyglot can create parser differentials, but parser differentials can occur without a polyglot. An ambiguous HTTP header or a malformed request can also cause two components to disagree without involving a file format at all.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Content-type confusion:&lt;/strong&gt; A broader term for situations where a component makes incorrect assumptions about what type of content it is handling. A polyglot can produce content-type confusion, but the term covers more than just file formats.&lt;/p&gt;

&lt;p&gt;A polyglot is specifically about the underlying byte sequence satisfying multiple format grammars simultaneously. That is a stronger property than simply misleading a label.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why the Processing Pipeline Is the Real Boundary
&lt;/h2&gt;

&lt;p&gt;Modern frameworks often provide upload validation utilities, file-type detection libraries, and secure storage helpers. The security properties of the entire system still depend on how the whole pipeline behaves together.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Upload validator
        ↓
Storage
        ↓
Image processing library
        ↓
Web server serving the file
        ↓
Browser parsing the response
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each step has its own notion of what the file is. If the validator uses one parser and a downstream component uses another, there is potential for disagreement. A file can pass the validator's check while still containing structures that a later component interprets differently. Security testing that covers only the upload endpoint misses this entirely.&lt;/p&gt;




&lt;h2&gt;
  
  
  Defenses That Follow From the Mechanism
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Do not rely on file extensions or MIME types from the client.&lt;/strong&gt; Both are trivially manipulated by the sender and tell you nothing about the actual bytes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Validate using the correct parser for the expected format.&lt;/strong&gt; The parser used at validation should match the parser that will ultimately consume the file. A mismatch between those two creates exactly the kind of gap a polyglot exploits.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;For image uploads, consider re-encoding.&lt;/strong&gt; Decoding to a pixel buffer and writing a fresh image from scratch discards any embedded structures that don't belong to the image data. This sidesteps validation entirely in favor of producing a known-clean output.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Reject ambiguous or malformed content.&lt;/strong&gt; A file that is partially valid under one format should not pass a lenient check silently.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Store uploads outside executable or directly web-accessible paths.&lt;/strong&gt; Even if validation is incomplete, a file that cannot be executed or served directly limits the impact.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Test the full processing chain.&lt;/strong&gt; Each component that touches the file is a potential point where interpretation diverges from what earlier stages assumed. Testing only the upload endpoint misses this entirely.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Deeper Lesson
&lt;/h2&gt;

&lt;p&gt;A file does not carry its own authoritative interpretation.&lt;/p&gt;

&lt;p&gt;Meaning is assigned by parsers, and parsers differ. When a byte sequence passes through multiple components that each apply their own grammar, the same bytes can mean different things at different points in the system.&lt;/p&gt;

&lt;p&gt;This is not a flaw in any particular file format or library. It is a property of how parsing works. Different grammars, applied to the same bytes, can reach different conclusions.&lt;/p&gt;

&lt;p&gt;The security model collapses when a decision made by one parser is trusted by components that operate with a different parser. The validation check confirmed one interpretation. The eventual consumer operated on another.&lt;/p&gt;

&lt;p&gt;That is what a polyglot exploits. Not that validators are broken. Not that parsers are buggy. The file is read in one way at the checkpoint, and in another way once it is past it.&lt;/p&gt;

</description>
      <category>computerscience</category>
      <category>security</category>
      <category>software</category>
    </item>
    <item>
      <title>What Is Web Cache Deception? How Can an Attacker Make a Cache Store Private Data?</title>
      <dc:creator>Aditya Sharma</dc:creator>
      <pubDate>Wed, 23 Sep 2026 06:14:07 +0000</pubDate>
      <link>https://dev.to/aditya_d_sharma/what-is-web-cache-deception-how-can-an-attacker-make-a-cache-store-private-data-36dn</link>
      <guid>https://dev.to/aditya_d_sharma/what-is-web-cache-deception-how-can-an-attacker-make-a-cache-store-private-data-36dn</guid>
      <description>&lt;p&gt;A user authenticates. They request a page that only they should see. The application checks their session, confirms they are authorized, and returns their private data. No authorization check was skipped. No SQL was injected. The application behaved correctly.&lt;/p&gt;

&lt;p&gt;But somewhere between the user's browser and the origin server, a cache stored that response. Later, someone else requested the same cache key. The cache returned the previous user's private page.&lt;/p&gt;

&lt;p&gt;The application didn't leak the data. The cache did.&lt;/p&gt;

&lt;p&gt;This is the core question: how can a system designed to make websites faster accidentally become a place where private data is stored and served to the wrong person?&lt;/p&gt;




&lt;h2&gt;
  
  
  How Web Caches Work
&lt;/h2&gt;

&lt;p&gt;When a user requests a web page, the request doesn't always travel directly to the origin server. Most production applications sit behind one or more intermediaries: CDNs, reverse proxies, load balancers. These components can serve responses from a local store rather than forwarding every request upstream.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Browser
  ↓
CDN / Reverse Proxy / Cache
  ↓
Origin Server
  ↓
Application
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A cache stores a response associated with a cache key, typically derived from the request URL and sometimes other request attributes. When a subsequent request arrives with a matching key, the cache can return the stored response without contacting the origin. This reduces latency and origin load.&lt;/p&gt;

&lt;p&gt;Caching is a performance mechanism. The security problem arises when a cache stores a response that should not have been stored, or serves it to a request that should not receive it.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Interpretation Mismatch
&lt;/h2&gt;

&lt;p&gt;Modern web applications often route requests flexibly. A path like &lt;code&gt;/account/profile&lt;/code&gt; might be handled by the application's routing logic, which checks the session cookie, fetches the relevant data, and returns a personalized response. The URL is not a path to a file. It is a route to application behavior.&lt;/p&gt;

&lt;p&gt;Caches, however, often make cacheability decisions using different signals: URL path patterns, file extensions, explicit cache-control headers, or cache rules configured by an administrator. Those rules may or may not agree with what the application actually does.&lt;/p&gt;

&lt;p&gt;The vulnerability class called web cache deception exploits the gap between these two interpretations.&lt;/p&gt;

&lt;p&gt;Consider:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;/account/profile
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The application routes this to the authenticated profile handler and returns private content.&lt;/p&gt;

&lt;p&gt;Now consider:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;/account/profile/style.css
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Whether this URL reaches the same handler depends entirely on how the application's routing is configured. Many frameworks will match the most specific route and fall back to the parent route when nothing more specific is found. Some serve the same profile content regardless of a trailing path segment they don't recognize. Others return 404.&lt;/p&gt;

&lt;p&gt;If the application returns the same private profile response for &lt;code&gt;/account/profile/style.css&lt;/code&gt; because its routing treats the unknown suffix as a parameter or falls through to the profile handler, then the origin's behavior is unchanged. The user gets their private data.&lt;/p&gt;

&lt;p&gt;But the cache may see a &lt;code&gt;.css&lt;/code&gt; extension and classify the response as a static resource, eligible for caching. The origin said "here is private user data." The cache heard "here is a stylesheet. I'll remember this."&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Origin interpretation:   /account/profile/style.css → private profile data
Cache interpretation:    /account/profile/style.css → static CSS → cacheable
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Same URL. Different meanings to different components.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Attack Sequence
&lt;/h2&gt;

&lt;p&gt;The conceptual attack flow has several steps:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Victim authenticates
    ↓
Victim visits deceptive URL
    ↓
Origin returns private response (correctly authorized)
    ↓
Cache stores response under that URL's cache key
    ↓
Attacker requests the same URL
    ↓
Cache returns victim's private response
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For this to succeed, specific conditions need to align. The application must return the same private content for the crafted URL as for the original one. The cache must consider the crafted URL cacheable when it would not cache the original URL. The cache must not include session cookies or other user-specific attributes in its cache key for that request. And the response headers must not instruct the cache to treat the response as private.&lt;/p&gt;

&lt;p&gt;These conditions are not guaranteed to exist. Whether a given system is vulnerable depends on the application's routing, the cache's configuration, the response headers, and how those components interact. Cache deception is a configuration and design problem, not an inherent flaw in any particular technology.&lt;/p&gt;




&lt;h2&gt;
  
  
  Cache Keys, Cacheability, and Why Authentication Does Not Save the Victim
&lt;/h2&gt;

&lt;p&gt;To understand why this matters at a deeper level, it helps to understand how caches make decisions.&lt;/p&gt;

&lt;p&gt;A cache key identifies which stored response to return for a given request. The key is often the URL, but many caches include or exclude other request attributes based on the &lt;code&gt;Vary&lt;/code&gt; response header or configuration rules. If a session cookie is not included in the cache key, two requests with different sessions but the same URL may receive the same cached response.&lt;/p&gt;

&lt;p&gt;Cacheability is a separate question. A cache determines whether a response can be stored based on the HTTP method, response status code, &lt;code&gt;Cache-Control&lt;/code&gt; directives, and sometimes configured rules that override headers.&lt;/p&gt;

&lt;p&gt;The dangerous scenario is when the cache key does not include user identity, the cache considers the response cacheable based on the URL pattern, and the application returns user-specific content for that URL.&lt;/p&gt;

&lt;p&gt;It is important to be precise about authentication here. Cache deception does not bypass the application's authentication. The origin correctly authenticates the victim and correctly generates their private response. The authorization logic works. The problem occurs afterward: the cache observes a response already authorized by the origin and stores it, then serves that stored response to the next matching request without involving the application at all.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"The application authorized this response for this user."

vs.

"The cache should have stored this response for anyone who asks."
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;These are different questions. The application answered the first correctly. The cache answered the second incorrectly.&lt;/p&gt;




&lt;h2&gt;
  
  
  HTTP Headers and Cache Control
&lt;/h2&gt;

&lt;p&gt;HTTP provides mechanisms for controlling caching behavior through the &lt;code&gt;Cache-Control&lt;/code&gt; response header.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;Cache-Control: private&lt;/code&gt; tells shared caches not to store the response. &lt;code&gt;Cache-Control: no-store&lt;/code&gt; tells caches not to store it at all. The &lt;code&gt;Vary&lt;/code&gt; header specifies which request headers should differentiate cached responses: &lt;code&gt;Vary: Cookie&lt;/code&gt; tells the cache to treat requests with different cookies as separate cache entries. RFC 7234 also specifies that caches should not store responses to requests carrying an &lt;code&gt;Authorization&lt;/code&gt; header unless the response explicitly permits it.&lt;/p&gt;

&lt;p&gt;These mechanisms work when correctly implemented and when the cache actually respects them. CDN cache rules that override response headers, misconfigured cache exceptions, or rules that prioritize URL matching over response headers can undermine what the application intends. The headers are necessary but not sufficient on their own.&lt;/p&gt;




&lt;h2&gt;
  
  
  Cache Deception vs. Cache Poisoning
&lt;/h2&gt;

&lt;p&gt;These two vulnerability classes involve caches behaving in unintended ways, but the direction of the problem is different.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cache deception:&lt;/strong&gt; A victim's private, authorized response gets stored under a cacheable-looking URL. A subsequent request retrieves that private response from the cache.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cache poisoning:&lt;/strong&gt; An attacker causes a malicious or undesirable response to be stored in the cache, so legitimate users receive it.&lt;/p&gt;

&lt;p&gt;In cache deception, the attacker benefits from something the victim received. In cache poisoning, the attacker puts something into the cache that other users receive. The mechanism involves caches in both cases, but the goal and the exploit structure are different.&lt;/p&gt;

&lt;p&gt;Normal caching is neither: a public response is stored and reused, which is exactly what was intended.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why Frameworks and CDNs Do Not Automatically Solve This
&lt;/h2&gt;

&lt;p&gt;Modern frameworks often provide sensible defaults for authenticated routes. Many CDNs allow fine-grained cache rules. But security depends on the entire configuration working correctly together.&lt;/p&gt;

&lt;p&gt;A framework that sets &lt;code&gt;Cache-Control: private&lt;/code&gt; on authenticated responses does not help if a CDN rule overrides cache headers for URLs ending in &lt;code&gt;.css&lt;/code&gt;. An application that routes unknown suffixes to dynamic handlers does not help if it does not set appropriate headers for those responses. A CDN configured to cache static assets does not distinguish between a real stylesheet and a URL that merely resembles one if no other signal is present.&lt;/p&gt;

&lt;p&gt;The interaction between application routing, response headers, and cache configuration determines actual behavior. Assuming any single layer handles this correctly without testing the combination is where security gaps appear.&lt;/p&gt;




&lt;h2&gt;
  
  
  Defenses
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Set appropriate Cache-Control headers on private responses.&lt;/strong&gt; &lt;code&gt;Cache-Control: no-store&lt;/code&gt; or &lt;code&gt;Cache-Control: private&lt;/code&gt; tells shared caches not to store the response. This is the most direct defense and must be present on every response carrying user-specific content.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Ensure cache rules and application routing agree.&lt;/strong&gt; If the cache caches all URLs with &lt;code&gt;.css&lt;/code&gt; extensions, the application should not serve private content for those URLs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Do not treat file extension alone as evidence of cacheability.&lt;/strong&gt; URL pattern-based cache decisions without consulting response headers can override application intent.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Test URL normalization and routing differences.&lt;/strong&gt; Check whether the application serves the same content for variations of a private URL: unknown path segments, trailing content, encoded characters, and unusual extensions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Include user-identifying attributes in the cache key where appropriate.&lt;/strong&gt; If personalized content must be cached, the cache key must reflect the attributes that differentiate responses between users.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Inspect actual CDN and reverse proxy behavior.&lt;/strong&gt; Test the full request path under production-like conditions. Configuration and runtime behavior can differ.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Include cache behavior in security testing.&lt;/strong&gt; Testing the application in isolation does not reveal how caches interact with it.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Deeper Lesson
&lt;/h2&gt;

&lt;p&gt;Authorization is not the only layer that determines whether a response reaches the right person.&lt;/p&gt;

&lt;p&gt;An application can correctly authenticate a user, correctly generate a private response, and correctly close the connection. If the response was stored by a cache in between, the application's authorization decision no longer governs what happens next.&lt;/p&gt;

&lt;p&gt;The security boundary for a web response includes every component that stores, transforms, routes, or reuses that response. A cache is not passive. It makes decisions, and those decisions can diverge from what the application intended.&lt;/p&gt;

&lt;p&gt;Web cache deception exists in that gap. The application authorized a response. The cache decided to share it.&lt;/p&gt;

</description>
      <category>cybersecurity</category>
      <category>infosec</category>
      <category>security</category>
    </item>
    <item>
      <title>What Is Zip Slip? How Can a ZIP File Write Files Outside Its Extraction Folder?</title>
      <dc:creator>Aditya Sharma</dc:creator>
      <pubDate>Tue, 22 Sep 2026 05:40:04 +0000</pubDate>
      <link>https://dev.to/aditya_d_sharma/what-is-zip-slip-how-can-a-zip-file-write-files-outside-its-extraction-folder-283h</link>
      <guid>https://dev.to/aditya_d_sharma/what-is-zip-slip-how-can-a-zip-file-write-files-outside-its-extraction-folder-283h</guid>
      <description>&lt;p&gt;A user uploads a ZIP file. The server extracts it into a temporary directory. Nothing seems wrong. The files are written. The process completes cleanly.&lt;/p&gt;

&lt;p&gt;But one of those files didn't land in the temporary directory. It landed somewhere else on the filesystem entirely. Somewhere it shouldn't have reached.&lt;/p&gt;

&lt;p&gt;The ZIP file didn't execute any code. No memory was corrupted. The extraction library worked exactly as designed. The problem was that the application trusted a filename inside the archive to be safe, and used it directly to construct a filesystem path.&lt;/p&gt;

&lt;p&gt;That's Zip Slip.&lt;/p&gt;




&lt;h2&gt;
  
  
  How ZIP Archives Store File Paths
&lt;/h2&gt;

&lt;p&gt;A ZIP archive is not just a collection of compressed file contents. Each entry in the archive has metadata: a filename, compression method, timestamps, and other attributes. That filename can contain a path, not just a bare name.&lt;/p&gt;

&lt;p&gt;When you create a ZIP containing &lt;code&gt;reports/2024/summary.csv&lt;/code&gt;, the archive entry stores &lt;code&gt;reports/2024/summary.csv&lt;/code&gt; as the entry name. Extraction software reads this name and uses it to reconstruct the directory structure at the destination.&lt;/p&gt;

&lt;p&gt;The critical point is that the archive itself defines the filename. The extraction software reads whatever string is stored in the entry metadata and uses it as a filesystem path. If that string contains &lt;code&gt;../&lt;/code&gt; sequences, the path can escape the intended extraction directory.&lt;/p&gt;




&lt;h2&gt;
  
  
  How Path Traversal Works in This Context
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;../&lt;/code&gt; is a standard filesystem path component meaning "go up one directory level." It is not a special attack payload. It is valid in any operating system's path resolution rules.&lt;/p&gt;

&lt;p&gt;So if an archive entry's name is &lt;code&gt;../important.conf&lt;/code&gt;, and the extraction directory is &lt;code&gt;/tmp/upload/&lt;/code&gt;, the extracted destination becomes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Extraction directory:    /tmp/upload/
Entry name:              ../important.conf
Joined path:             /tmp/upload/../important.conf
Resolved path:           /tmp/important.conf
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The file lands one level above the extraction directory. Extend this further:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Entry name:              ../../etc/cron.d/job
Joined path:             /tmp/upload/../../etc/cron.d/job
Resolved path:           /etc/cron.d/job
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each &lt;code&gt;../&lt;/code&gt; walks up one level. Enough of them can reach anywhere on the filesystem from the extraction root.&lt;/p&gt;

&lt;p&gt;The normal flow versus the vulnerable flow:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Safe archive entry
      ↓
Extraction directory + entry name
      ↓
Resolved path inside extraction directory
      ↓
File written safely

Malicious archive entry (../../etc/cron.d/job)
      ↓
Extraction directory + entry name
      ↓
Resolved path escapes extraction directory
      ↓
File written to arbitrary location
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Why Checking the Filename Alone Is Not Enough
&lt;/h2&gt;

&lt;p&gt;A naive defense might look for &lt;code&gt;../&lt;/code&gt; in the entry name and reject any entry that contains it. That catches the obvious case. But it doesn't catch everything.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;URL encoding and other representations.&lt;/strong&gt; Some extraction code may decode entry names before writing. An entry named &lt;code&gt;..%2F&lt;/code&gt; or &lt;code&gt;%2e%2e/&lt;/code&gt; may be decoded into &lt;code&gt;../&lt;/code&gt; after any string-level check has already run. The check sees a safe string; the filesystem sees a traversal.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Absolute paths.&lt;/strong&gt; An archive entry can also store an absolute path like &lt;code&gt;/etc/passwd&lt;/code&gt; rather than a relative one. If the extraction code joins the extraction directory with an absolute path on some platforms, or simply uses the path as-is, the result skips the extraction directory entirely.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Mixed separators.&lt;/strong&gt; On Windows, both &lt;code&gt;/&lt;/code&gt; and &lt;code&gt;\&lt;/code&gt; are valid path separators. An entry named &lt;code&gt;..\config\app.ini&lt;/code&gt; may pass a check that only looks for &lt;code&gt;../&lt;/code&gt; while still traversing upward.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Normalization after joining.&lt;/strong&gt; The important check is not whether the entry name looks suspicious. The important check is whether the fully resolved destination path still falls inside the intended extraction directory. These are different questions.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Role of Path Canonicalization
&lt;/h2&gt;

&lt;p&gt;Canonicalization is the process of resolving a path to its definitive form: expanding &lt;code&gt;..&lt;/code&gt; components, resolving symlinks, collapsing redundant separators, and producing an absolute path.&lt;/p&gt;

&lt;p&gt;The correct approach is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;entry_name&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;read&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="n"&gt;archive&lt;/span&gt;
&lt;span class="n"&gt;destination&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;extraction_dir&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;entry_name&lt;/span&gt;
&lt;span class="n"&gt;canonical&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;canonicalize&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;destination&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;canonical&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;startswith&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;extraction_dir&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;reject&lt;/span&gt; &lt;span class="n"&gt;entry&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This check happens after resolution, not before. It doesn't matter what the entry name looks like. What matters is where the file would actually end up.&lt;/p&gt;

&lt;p&gt;The naive approach that fails:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;entry_name&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;read&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="n"&gt;archive&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;..&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;entry_name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;reject&lt;/span&gt;
&lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;write&lt;/span&gt; &lt;span class="n"&gt;to&lt;/span&gt; &lt;span class="n"&gt;extraction_dir&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;entry_name&lt;/span&gt;   &lt;span class="c1"&gt;# still vulnerable
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Checking the entry name in isolation cannot account for how the operating system will interpret the joined path. The filesystem is the authority on where a path points. Any security check that doesn't consult the filesystem's own resolution is working with incomplete information.&lt;/p&gt;




&lt;h2&gt;
  
  
  Symlinks Complicate Extraction Further
&lt;/h2&gt;

&lt;p&gt;Zip Slip becomes more complex when the archive contains symbolic links rather than regular files.&lt;/p&gt;

&lt;p&gt;A symlink is a filesystem object that points to another path. If an archive contains an entry that creates a symlink at &lt;code&gt;extract/link&lt;/code&gt; pointing to &lt;code&gt;/etc&lt;/code&gt;, and then a second entry that writes a file to &lt;code&gt;extract/link/cron.d/job&lt;/code&gt;, the second write follows the symlink. The file ends up at &lt;code&gt;/etc/cron.d/job&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Step 1: Extract symlink
  archive entry "link" → symlink to /etc
  written to /tmp/upload/link → points to /etc

Step 2: Extract file using symlink path
  archive entry "link/cron.d/job"
  resolved: /tmp/upload/link/cron.d/job
  followed: /etc/cron.d/job
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both entries individually may pass a path traversal check. The first creates a symlink inside the extraction directory. The second writes a file to a path inside the extraction directory. But the combination results in a write outside it.&lt;/p&gt;

&lt;p&gt;Safe extraction requires either rejecting symlinks from untrusted archives, or resolving intermediate symlinks before checking whether the final destination is inside the intended extraction root.&lt;/p&gt;




&lt;h2&gt;
  
  
  Where Zip Slip Appears
&lt;/h2&gt;

&lt;p&gt;The vulnerability isn't specific to ZIP. Any archive format that stores filenames with paths can carry the same problem: tar, jar, war, ear, apk, gem. The extraction logic, not the format, determines the exposure.&lt;/p&gt;

&lt;p&gt;Web applications that accept archive uploads are the most obvious case. A user uploads a ZIP, the server extracts it, and if the extraction code doesn't validate entry paths, the archive controls where files land.&lt;/p&gt;

&lt;p&gt;CI/CD pipelines and build systems unpack artifacts and dependencies regularly. Package managers in several ecosystems have historically been affected. Backup and restore systems often extract with elevated privileges, which increases the reach of any malicious entries. Any automated pipeline that accepts archives from external sources and extracts them programmatically is in scope.&lt;/p&gt;




&lt;h2&gt;
  
  
  Impact Depends on What the Process Can Reach
&lt;/h2&gt;

&lt;p&gt;Zip Slip enables arbitrary file write within what the extraction process can access. A web application running as a low-privilege user has limited reach. A build pipeline with broader system access might allow overwriting configuration files, startup scripts, or application code. An extraction process running as root can reach anything writable on the filesystem.&lt;/p&gt;

&lt;p&gt;The vulnerability enables writing, not reading, and does not by itself provide code execution. But overwriting a configuration file, an application startup script, or a cron job can be a path to execution depending on the system. Whether an existing file can be overwritten also depends on filesystem permissions and other constraints. The impact is contextual.&lt;/p&gt;




&lt;h2&gt;
  
  
  Zip Slip vs Ordinary Path Traversal
&lt;/h2&gt;

&lt;p&gt;Classic path traversal attacks manipulate URL parameters or form fields to cause a server to read a file outside its intended scope: &lt;code&gt;GET /files?name=../../etc/passwd&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Zip Slip is a path traversal in archive extraction. The traversal string doesn't come from a URL parameter or an HTTP header. It comes from the filename stored inside an archive entry.&lt;/p&gt;

&lt;p&gt;The ZIP format itself is not executing anything and is not inherently malicious. The vulnerability is in the extraction logic that treats the archive-controlled filename as a safe basis for constructing a filesystem path. The archive is just the delivery mechanism for a filename the application should not have trusted.&lt;/p&gt;




&lt;h2&gt;
  
  
  Defenses
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Resolve before checking.&lt;/strong&gt; Compute the canonical absolute path of the destination before writing any file. Verify that it starts with the canonical absolute path of the extraction directory. Reject any entry whose resolved destination falls outside.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Handle symlinks explicitly.&lt;/strong&gt; Decide whether your extraction logic should allow symlinks at all. If symlinks are necessary, resolve intermediate symlinks during path verification, not just the final path.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use a well-maintained extraction library.&lt;/strong&gt; Many languages and ecosystems now have extraction libraries that perform these checks by default. Prefer those over manual path construction.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Do not rely on entry name inspection.&lt;/strong&gt; Checking whether an entry name "looks safe" is insufficient. The resolution happens at the filesystem level, and the check must happen there too.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Treat extraction as a privilege boundary.&lt;/strong&gt; If the extraction process runs with elevated privileges, the blast radius of a successful Zip Slip is larger. Run extraction with the minimum necessary permissions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Test with adversarial archives.&lt;/strong&gt; A test suite that only extracts well-formed archives will not catch this. Explicitly test with entries containing &lt;code&gt;../&lt;/code&gt; sequences, absolute paths, null bytes in filenames, and symlinks pointing outside the extraction root.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Technical Takeaway
&lt;/h2&gt;

&lt;p&gt;Archive filenames are attacker-controlled input. An archive produced by a user, a third-party system, or an external package source defines its own entry names, and those names can contain any string the format allows.&lt;/p&gt;

&lt;p&gt;The code that converts those names into filesystem paths is the security boundary. Not the archive format. Not the filename extension. Not a string check on the entry name. The security boundary is the moment the application decides where to write the file, and whether it verified that decision against the filesystem's own resolution before acting on it.&lt;/p&gt;

&lt;p&gt;A ZIP file that writes outside its extraction folder is not doing anything the ZIP format prohibits. It is doing exactly what the extraction code allows.&lt;/p&gt;

</description>
      <category>cybersecurity</category>
      <category>infosec</category>
      <category>security</category>
    </item>
    <item>
      <title>What Is Rowhammer? How Can Repeated Memory Access Flip Bits in RAM?</title>
      <dc:creator>Aditya Sharma</dc:creator>
      <pubDate>Mon, 21 Sep 2026 05:12:07 +0000</pubDate>
      <link>https://dev.to/aditya_d_sharma/what-is-rowhammer-how-can-repeated-memory-access-flip-bits-in-ram-1l1d</link>
      <guid>https://dev.to/aditya_d_sharma/what-is-rowhammer-how-can-repeated-memory-access-flip-bits-in-ram-1l1d</guid>
      <description>&lt;p&gt;You are not allowed to modify this memory.&lt;/p&gt;

&lt;p&gt;The operating system says no. The CPU's memory protection says no. The page table marks the region as belonging to a different process.&lt;/p&gt;

&lt;p&gt;So you don't write to it.&lt;/p&gt;

&lt;p&gt;Instead, you repeatedly access the memory next to it. Thousands of times per second. You stay entirely within your own allocated region.&lt;/p&gt;

&lt;p&gt;And eventually, a bit in the protected memory changes.&lt;/p&gt;

&lt;p&gt;No permission was violated. No buffer was overflowed. No kernel bug was triggered. The hardware did it.&lt;/p&gt;

&lt;p&gt;That's Rowhammer.&lt;/p&gt;




&lt;h2&gt;
  
  
  How DRAM Actually Stores a Bit
&lt;/h2&gt;

&lt;p&gt;To understand why this happens, you need a basic picture of how DRAM works at the physical level.&lt;/p&gt;

&lt;p&gt;Each bit in DRAM is stored in a cell that consists of a capacitor and an access transistor. The capacitor holds charge. Charge present represents one binary state; charge absent represents the other. The transistor connects the capacitor to a bitline, a wire that runs through the memory array and is used during read and write operations.&lt;/p&gt;

&lt;p&gt;Cells are arranged in a grid. Wordlines run horizontally, connecting the transistors of all cells in a row. Bitlines run vertically. To access a row, the memory controller asserts the wordline, which opens all the transistors in that row simultaneously, connecting every cell's capacitor to its corresponding bitline. Sense amplifiers at the end of each bitline then detect and amplify the tiny charge signal.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;         Bitline   Bitline   Bitline
            |         |         |
Row A ──[T]─C─    ──[T]─C─    ──[T]─C─
            |         |         |
Row B ──[T]─C─    ──[T]─C─    ──[T]─C─
            |         |         |
Row C ──[T]─C─    ──[T]─C─    ──[T]─C─

T = access transistor, C = capacitor
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This physical arrangement is what makes Rowhammer possible. Rows are physically close to each other. The cells are packed tightly.&lt;/p&gt;




&lt;h2&gt;
  
  
  What Happens When You Activate a Row
&lt;/h2&gt;

&lt;p&gt;When the memory controller accesses a DRAM row, it activates the entire row: the wordline is asserted, all the transistors open, and the capacitors interact with the bitlines and sense amplifiers. This is not a subtle operation at the hardware level.&lt;/p&gt;

&lt;p&gt;From software, you see addresses and bytes. From the hardware perspective:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Virtual Address
       ↓
Page Table
       ↓
Physical Address
       ↓
Memory Controller
       ↓
DRAM Bank and Row
       ↓
Row Activation / Sense Amplifiers
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The gap between what software sees and what hardware does is where Rowhammer lives.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Hammering Mechanism
&lt;/h2&gt;

&lt;p&gt;DRAM cells are densely packed. When a wordline is asserted and a row is activated, the electrical disturbance is not perfectly contained. Neighboring rows can experience small amounts of interference.&lt;/p&gt;

&lt;p&gt;In normal usage, this doesn't matter. A row is activated occasionally, and DRAM refresh operations periodically restore charge to cells before they drift too much.&lt;/p&gt;

&lt;p&gt;But if an attacker repeatedly activates a row in a tight loop, hundreds of thousands of times per second, the cumulative electrical disturbance in neighboring rows can grow. The capacitors in those adjacent rows get disturbed more than the refresh mechanism expected. In some DRAM chips and configurations, this can cause a bit in a neighboring row to flip from its correct value to the opposite state.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Victim Row
─────────────────────────────────────
    ↑ electrical disturbance
─────────────────────────────────────

Hammer Row A       Hammer Row B
█████████████      █████████████
activate           activate
activate           activate
activate           activate
...                ...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Double-sided hammering, where rows on both sides of the victim row are repeatedly activated, concentrates the disturbance from two directions. This can increase the chance of inducing a flip in a susceptible system compared to hammering one side.&lt;/p&gt;

&lt;p&gt;Susceptibility varies. Not every DRAM module experiences this, and behavior depends on the DRAM design, manufacturing process, access pattern, timing, temperature, and what mitigations are present.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why Software Isolation Doesn't Stop It
&lt;/h2&gt;

&lt;p&gt;This is the counterintuitive part.&lt;/p&gt;

&lt;p&gt;The operating system creates an abstraction: each process has a virtual address space, and the page table ensures that accesses to one process's memory don't reach another process's pages. When Process A tries to read Process B's memory, the CPU checks the page table and refuses.&lt;/p&gt;

&lt;p&gt;Rowhammer doesn't ask the CPU to access another process's memory. It stays within its own legal pages.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Process A (attacker)

Legal pages
████████████████
       ↓
   repeated activation
       ↓
Physical DRAM disturbance
       ↓
Nearby physical cells
       ↓
████████████████
   Victim memory (different process or OS)
       ↓
possible bit flip
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The CPU's page-table permission check is bypassed not because it was defeated, but because the attacker never triggered it. The hardware underneath the abstraction is doing something the abstraction wasn't designed to account for.&lt;/p&gt;




&lt;h2&gt;
  
  
  From a Bit Flip to a Security Consequence
&lt;/h2&gt;

&lt;p&gt;A random bit flip in an unused region of memory is just data corruption. Annoying, not dangerous.&lt;/p&gt;

&lt;p&gt;The security question is whether an attacker can cause a bit to flip in a location that matters: a page table entry, permission bits on a memory page, metadata used by the OS to track ownership, or values the program trusts as authoritative.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Physical disturbance
       ↓
Bit flip
       ↓
Security-sensitive location
       ↓
Corrupted metadata or permission structure
       ↓
Potential isolation violation
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For this to go from bit flip to exploit, several things need to align: the flip has to happen in the right location, the attacker needs some ability to influence physical memory placement, the corrupted value has to be security-sensitive, and the system has to act on it before detecting the inconsistency.&lt;/p&gt;

&lt;p&gt;This makes Rowhammer exploitation significantly harder than pointing it at a target and waiting. But researchers have demonstrated that, in certain configurations, this chain can be constructed.&lt;/p&gt;




&lt;h2&gt;
  
  
  Rowhammer vs. Software Memory Corruption
&lt;/h2&gt;

&lt;p&gt;It's worth being precise about the distinction.&lt;/p&gt;

&lt;p&gt;A buffer overflow is a software bug. The program writes past the end of a buffer because bounds checking failed or was absent.&lt;/p&gt;

&lt;p&gt;A use-after-free is a software lifetime bug. The program accesses memory it no longer owns.&lt;/p&gt;

&lt;p&gt;Rowhammer is neither. The attacker performs memory accesses that are within their permissions. No application bug is involved. The disturbance occurs in the physical DRAM subsystem, not in a software buffer.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Buffer overflow:
program writes outside intended bounds (software bug)

Use-after-free:
program accesses freed memory (software bug)

Rowhammer:
program performs legitimate accesses
→ hardware disturbance
→ bit flip in different physical location
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This means traditional software mitigations like address-space layout randomization, stack canaries, or safe languages don't address the underlying phenomenon. The vulnerability exists at a layer below them.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why Mitigations Are Difficult
&lt;/h2&gt;

&lt;p&gt;Several mitigations exist, and each operates at a different layer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Increased refresh rates.&lt;/strong&gt; Refreshing DRAM cells more frequently reduces the window during which a cell can drift. This trades power and bandwidth for resilience.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Targeted Row Refresh (TRR).&lt;/strong&gt; The memory controller or DRAM monitors row activation frequency. When a row crosses a threshold, neighboring rows are proactively refreshed. TRR implementations vary, and some have gaps under certain access patterns.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;ECC memory.&lt;/strong&gt; Error-correcting code memory can detect and correct certain bit errors. For single-bit errors, most ECC schemes can correct the flip transparently. But ECC doesn't prevent the flip from occurring, and multi-bit or unusual error patterns may fall outside what the scheme can correct.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;OS-level strategies.&lt;/strong&gt; Careful physical memory allocation can reduce the chance that attacker pages are physically adjacent to sensitive structures. This changes what a flip can reach but doesn't prevent the underlying disturbance.&lt;/p&gt;

&lt;p&gt;No single mitigation addresses all configurations. Rowhammer is partly a hardware problem and partly an architectural one, and the mitigations reflect that.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Deeper Lesson
&lt;/h2&gt;

&lt;p&gt;Security boundaries in software depend on the layers underneath behaving as expected.&lt;/p&gt;

&lt;p&gt;The operating system's isolation model assumes that one process cannot influence another's memory without going through software primitives that the OS controls. That assumption is correct at the software layer.&lt;/p&gt;

&lt;p&gt;Rowhammer reveals that the physical implementation of memory can introduce a side channel that crosses that boundary without touching any of the software mechanisms guarding it.&lt;/p&gt;

&lt;p&gt;This isn't a failure of the page table, the CPU, or the operating system. It's a reminder that software abstractions are built on physical systems, and those physical systems have their own properties.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;The attacker can access Row A.

The victim's data is in Row B.

The attacker cannot write Row B.

But repeated activation of Row A
can disturb nearby DRAM cells.

If a bit in Row B flips,
software observes a state it never intentionally created.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Rowhammer doesn't break the rule that you can't write someone else's memory.&lt;/p&gt;

&lt;p&gt;It attacks the hardware underneath the rule.&lt;/p&gt;

</description>
      <category>cybersecurity</category>
      <category>hardware</category>
      <category>infosec</category>
      <category>security</category>
    </item>
    <item>
      <title>What Is DOM Clobbering? How Can HTML Elements Change What JavaScript Thinks a Variable Is?</title>
      <dc:creator>Aditya Sharma</dc:creator>
      <pubDate>Sun, 20 Sep 2026 06:56:42 +0000</pubDate>
      <link>https://dev.to/aditya_d_sharma/what-is-dom-clobbering-how-can-html-elements-change-what-javascript-thinks-a-variable-is-1077</link>
      <guid>https://dev.to/aditya_d_sharma/what-is-dom-clobbering-how-can-html-elements-change-what-javascript-thinks-a-variable-is-1077</guid>
      <description>&lt;p&gt;You didn't inject a script tag. You didn't modify any JavaScript. You didn't find a way to execute code directly.&lt;/p&gt;

&lt;p&gt;You added an HTML element.&lt;/p&gt;

&lt;p&gt;And somewhere in the application's JavaScript, a value that the developer assumed would be their own configuration object suddenly resolved to something else entirely.&lt;/p&gt;

&lt;p&gt;No new code ran. The existing JavaScript behaved differently because the environment it ran in had changed.&lt;/p&gt;

&lt;p&gt;That's the strange thing about DOM Clobbering. The browser isn't misbehaving. The HTML is valid. JavaScript is doing a perfectly normal property lookup. The problem is an assumption the developer made about what that lookup could return.&lt;/p&gt;




&lt;h2&gt;
  
  
  What "Clobbering" Actually Means
&lt;/h2&gt;

&lt;p&gt;The DOM (Document Object Model) isn't just a tree of elements you manipulate with JavaScript. It participates in name resolution in ways that aren't always obvious.&lt;/p&gt;

&lt;p&gt;Browsers expose certain HTML elements through named properties. When an element has an &lt;code&gt;id&lt;/code&gt;, the browser may make it accessible through the &lt;code&gt;window&lt;/code&gt; object's named property collection. When certain elements like forms or iframes use &lt;code&gt;name&lt;/code&gt; attributes, similar behavior applies.&lt;/p&gt;

&lt;p&gt;"Clobbering" describes what happens when attacker-controlled HTML introduces a name that collides with a property the application's JavaScript expects to control.&lt;/p&gt;

&lt;p&gt;The application reaches for &lt;code&gt;window.config&lt;/code&gt; expecting its own configuration object. But a DOM element with &lt;code&gt;id="config"&lt;/code&gt; has been introduced. The lookup resolves to something the developer didn't put there.&lt;/p&gt;

&lt;p&gt;DOM Clobbering isn't simply "every id becomes a global." The behavior is more nuanced than that. It depends on the element type, the property being accessed, which object is being queried, and specific browser named-property rules. But when the conditions align, an HTML element can change what a JavaScript name resolves to.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Browser Behavior Behind It
&lt;/h2&gt;

&lt;p&gt;Browsers expose what the HTML specification calls "named properties" on &lt;code&gt;window&lt;/code&gt; and on certain DOM objects. This is a deliberate feature, not a bug. It exists for historical reasons and backward compatibility.&lt;/p&gt;

&lt;p&gt;When you access &lt;code&gt;window.someIdentifier&lt;/code&gt;, the browser doesn't only look at JavaScript variables you defined. It also checks whether any HTML element with a matching &lt;code&gt;id&lt;/code&gt; exists on the page, among other named-property resolution rules.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="nb"&gt;window&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;config&lt;/span&gt;

&lt;span class="nx"&gt;Browser&lt;/span&gt; &lt;span class="nx"&gt;checks&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
&lt;span class="err"&gt;→&lt;/span&gt; &lt;span class="nx"&gt;JavaScript&lt;/span&gt; &lt;span class="nx"&gt;variable&lt;/span&gt; &lt;span class="nx"&gt;named&lt;/span&gt; &lt;span class="nx"&gt;config&lt;/span&gt;&lt;span class="p"&gt;?&lt;/span&gt;
&lt;span class="err"&gt;→&lt;/span&gt; &lt;span class="nx"&gt;Named&lt;/span&gt; &lt;span class="nx"&gt;property&lt;/span&gt; &lt;span class="nx"&gt;collision&lt;/span&gt; &lt;span class="kd"&gt;with&lt;/span&gt; &lt;span class="nx"&gt;a&lt;/span&gt; &lt;span class="nx"&gt;DOM&lt;/span&gt; &lt;span class="nx"&gt;element&lt;/span&gt;&lt;span class="p"&gt;?&lt;/span&gt;
&lt;span class="err"&gt;→&lt;/span&gt; &lt;span class="nx"&gt;Something&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="nx"&gt;the&lt;/span&gt; &lt;span class="nx"&gt;named&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="nx"&gt;property&lt;/span&gt; &lt;span class="nx"&gt;chain&lt;/span&gt;&lt;span class="p"&gt;?&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If an attacker can inject HTML into a page, they may be able to introduce an element whose &lt;code&gt;id&lt;/code&gt; or &lt;code&gt;name&lt;/code&gt; matches something the application's JavaScript relies on.&lt;/p&gt;

&lt;p&gt;The browser then has two candidates for what that name means. Depending on the specifics, the DOM-introduced element may win.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why &lt;code&gt;id&lt;/code&gt; and &lt;code&gt;name&lt;/code&gt; Matter Differently
&lt;/h2&gt;

&lt;p&gt;Not all HTML attributes participate equally. Elements with &lt;code&gt;id&lt;/code&gt; attributes can be reached through &lt;code&gt;window&lt;/code&gt; named properties in many browsers, but the behavior isn't universal across all contexts. Certain elements have additional named-property behavior: forms and iframes are the most commonly relevant. A &lt;code&gt;&amp;lt;form name="something"&amp;gt;&lt;/code&gt; creates a named property accessible through &lt;code&gt;document&lt;/code&gt;. An iframe with a &lt;code&gt;name&lt;/code&gt; attribute can be accessible through &lt;code&gt;window&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;This is also why the element type matters, not just the attribute. A &lt;code&gt;&amp;lt;div id="config"&amp;gt;&lt;/code&gt; and a &lt;code&gt;&amp;lt;form id="config"&amp;gt;&lt;/code&gt; both introduce the same name, but what you can do with the resulting named property differs.&lt;/p&gt;




&lt;h2&gt;
  
  
  A Concrete Example
&lt;/h2&gt;

&lt;p&gt;Consider an application that initializes itself like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;config&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;window&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;config&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;config&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;config&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;endpoint&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;config&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;endpoint&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;/api/data&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The developer expects &lt;code&gt;window.config&lt;/code&gt; to be a JavaScript object they defined elsewhere. It has an &lt;code&gt;endpoint&lt;/code&gt; property pointing to their API.&lt;/p&gt;

&lt;p&gt;Now suppose an attacker can inject HTML into the page before this code runs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight html"&gt;&lt;code&gt;&lt;span class="nt"&gt;&amp;lt;div&lt;/span&gt; &lt;span class="na"&gt;id=&lt;/span&gt;&lt;span class="s"&gt;"config"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&amp;lt;/div&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In browsers where &lt;code&gt;id&lt;/code&gt; attributes participate in &lt;code&gt;window&lt;/code&gt; named property resolution, &lt;code&gt;window.config&lt;/code&gt; no longer necessarily returns the developer's JavaScript object. It may return the DOM element.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="nx"&gt;Application&lt;/span&gt; &lt;span class="nx"&gt;assumption&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
&lt;span class="nb"&gt;window&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;config&lt;/span&gt; &lt;span class="err"&gt;→&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nl"&gt;endpoint&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;https://api.example.com&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="nx"&gt;After&lt;/span&gt; &lt;span class="nx"&gt;HTML&lt;/span&gt; &lt;span class="nx"&gt;injection&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
&lt;span class="nb"&gt;window&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;config&lt;/span&gt; &lt;span class="err"&gt;→&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;div&lt;/span&gt; &lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;config&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;element&lt;/span&gt;
&lt;span class="nx"&gt;config&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;endpoint&lt;/span&gt; &lt;span class="err"&gt;→&lt;/span&gt; &lt;span class="nf"&gt;undefined &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;or&lt;/span&gt; &lt;span class="nx"&gt;something&lt;/span&gt; &lt;span class="nx"&gt;unexpected&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The fetch doesn't happen. Or it happens with an unexpected value. Depending on what the application does with this, the consequences range from a broken feature to something more significant.&lt;/p&gt;

&lt;p&gt;The attacker never touched the JavaScript. They changed what the JavaScript operated on.&lt;/p&gt;




&lt;h2&gt;
  
  
  Forms, Iframes, and Nested Property Access
&lt;/h2&gt;

&lt;p&gt;With a plain &lt;code&gt;&amp;lt;div id="config"&amp;gt;&lt;/code&gt;, &lt;code&gt;window.config&lt;/code&gt; resolves to the element, and property access like &lt;code&gt;config.endpoint&lt;/code&gt; returns &lt;code&gt;undefined&lt;/code&gt;. Not always dangerous on its own.&lt;/p&gt;

&lt;p&gt;But certain element types allow nested named properties. A &lt;code&gt;&amp;lt;form&amp;gt;&lt;/code&gt; element named &lt;code&gt;config&lt;/code&gt; with a child &lt;code&gt;&amp;lt;input name="endpoint"&amp;gt;&lt;/code&gt; can make &lt;code&gt;config.endpoint&lt;/code&gt; resolve to the input element rather than &lt;code&gt;undefined&lt;/code&gt;. This happens because form elements expose their named child inputs as named properties.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight html"&gt;&lt;code&gt;&lt;span class="nt"&gt;&amp;lt;form&lt;/span&gt; &lt;span class="na"&gt;id=&lt;/span&gt;&lt;span class="s"&gt;"config"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;input&lt;/span&gt; &lt;span class="na"&gt;name=&lt;/span&gt;&lt;span class="s"&gt;"endpoint"&lt;/span&gt; &lt;span class="na"&gt;value=&lt;/span&gt;&lt;span class="s"&gt;"..."&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;/form&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="nb"&gt;window&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;config&lt;/span&gt; &lt;span class="err"&gt;→&lt;/span&gt; &lt;span class="nx"&gt;the&lt;/span&gt; &lt;span class="nx"&gt;form&lt;/span&gt; &lt;span class="nx"&gt;element&lt;/span&gt;
&lt;span class="nb"&gt;window&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;config&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;endpoint&lt;/span&gt; &lt;span class="err"&gt;→&lt;/span&gt; &lt;span class="nx"&gt;the&lt;/span&gt; &lt;span class="nx"&gt;input&lt;/span&gt; &lt;span class="nx"&gt;element&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is why clobbering research often involves structured HTML rather than a single element. The goal is to satisfy the application's property-access pattern, not just the top-level lookup. This doesn't let an attacker construct arbitrary JavaScript objects, but it can produce enough structure to satisfy a conditional check and feed unexpected values into subsequent logic.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why &lt;code&gt;document.getElementById()&lt;/code&gt; Is Different
&lt;/h2&gt;

&lt;p&gt;There's an important distinction between implicit named property resolution and explicit DOM queries.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Implicit - subject to named property resolution&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;config&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;window&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;config&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;// Explicit - clearly queries the DOM&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;el&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;document&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getElementById&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;config&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;These look similar but behave differently. &lt;code&gt;document.getElementById("config")&lt;/code&gt; is an explicit API call. You get a DOM element, and you know you're getting a DOM element. The code is clearly asking for an element.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;window.config&lt;/code&gt; is a named property lookup. The developer may intend it to resolve to a JavaScript object they defined. The lookup doesn't make that distinction for you.&lt;/p&gt;

&lt;p&gt;DOM Clobbering exploits the gap between the developer's assumption and the browser's actual resolution behavior. Explicit APIs make the intent and the result clearer. They don't prevent you from making other mistakes, but they avoid one specific class of implicit name collision.&lt;/p&gt;




&lt;h2&gt;
  
  
  DOM Clobbering vs. XSS
&lt;/h2&gt;

&lt;p&gt;These are different mechanisms and shouldn't be conflated.&lt;/p&gt;

&lt;p&gt;XSS means attacker-controlled input becomes executable JavaScript. A script runs that the attacker introduced.&lt;/p&gt;

&lt;p&gt;DOM Clobbering means attacker-controlled HTML changes the DOM environment in a way that influences existing JavaScript. The attacker's code doesn't execute. The application's own code executes against a different environment than it expected.&lt;/p&gt;

&lt;p&gt;The attacker may not be able to run JavaScript at all. Perhaps the application has a strict Content Security Policy that blocks inline scripts and untrusted script sources. DOM Clobbering doesn't require script execution. It requires only the ability to introduce certain HTML.&lt;/p&gt;

&lt;p&gt;This makes it a relevant technique in scenarios where traditional script injection is blocked but HTML injection is still possible.&lt;/p&gt;




&lt;h2&gt;
  
  
  How It Becomes a Security Problem
&lt;/h2&gt;

&lt;p&gt;DOM Clobbering by itself isn't a vulnerability. It becomes one when the clobbered value influences something security-sensitive.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Attacker can influence page HTML
        ↓
Application relies on implicit named property resolution
        ↓
A security-sensitive lookup resolves to attacker-influenced value
        ↓
Application trusts the resolved value without type-checking
        ↓
Unexpected behavior
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;What matters is what the application does with the unexpected value: URL construction, configuration, resource loading, security checks, application initialization. In code paths that don't touch anything sensitive, the practical impact may be minimal. In code paths that do, the consequences scale with what the clobbered value controls.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why Frameworks Don't Automatically Solve It
&lt;/h2&gt;

&lt;p&gt;React, Vue, and similar frameworks run inside a browser. Browser semantics don't change because a framework is in use. If application code, or a library the application depends on, performs named property resolution on &lt;code&gt;window&lt;/code&gt; or other browser-exposed objects, the underlying behavior still applies. Framework abstraction doesn't change how the browser resolves named properties.&lt;/p&gt;




&lt;h2&gt;
  
  
  Defenses
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Avoid implicit named property resolution for security-sensitive values.&lt;/strong&gt; Don't assume &lt;code&gt;window.someIdentifier&lt;/code&gt; must refer to an application-defined object.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Keep configuration in explicitly controlled scope.&lt;/strong&gt; A JavaScript module that exports a configuration object is not subject to named property collision in the same way a &lt;code&gt;window&lt;/code&gt;-level reference is.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Validate type and structure, not just existence.&lt;/strong&gt; Checking &lt;code&gt;if (config)&lt;/code&gt; tells you the value is truthy. A DOM element is truthy. Verify that the value has the expected type and the properties you actually need before using it in sensitive operations.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Sanitize attacker-controlled HTML.&lt;/strong&gt; If the application accepts user-supplied HTML, an appropriate sanitizer can strip &lt;code&gt;id&lt;/code&gt; and &lt;code&gt;name&lt;/code&gt; attributes that could participate in clobbering.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prefer explicit DOM queries when querying the DOM.&lt;/strong&gt; &lt;code&gt;document.getElementById("config")&lt;/code&gt; clearly asks for an element and returns one. &lt;code&gt;window.config&lt;/code&gt; makes an implicit assumption about what that name resolves to.&lt;/p&gt;

&lt;p&gt;CSP and Trusted Types address related script-injection risks but don't fix the core issue: an unsafe assumption about named property resolution. The fix is removing that assumption and controlling attacker-influenced HTML at the source.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Deeper Lesson
&lt;/h2&gt;

&lt;p&gt;Nothing in a DOM Clobbering scenario is technically broken.&lt;/p&gt;

&lt;p&gt;The browser exposes named properties exactly as specified. The HTML is valid markup. The DOM is constructed the way browsers construct DOMs. JavaScript performs a normal property lookup and gets back a value.&lt;/p&gt;

&lt;p&gt;The problem is a developer assumption: that the lookup can only return what the developer put there.&lt;/p&gt;

&lt;p&gt;That assumption is wrong. The DOM is a shared environment. HTML that an attacker introduces participates in that environment. Browser name resolution doesn't distinguish between developer-controlled and attacker-controlled elements.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;HTML is not just markup.
HTML creates the DOM.
The DOM participates in browser name resolution.
JavaScript operates in that environment.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When application code assumes a name belongs entirely to the application's own JavaScript, it's assuming something the browser doesn't guarantee.&lt;/p&gt;

&lt;p&gt;The attacker didn't change the JavaScript. They changed the environment the JavaScript was running in. The JavaScript did exactly what it was written to do.&lt;/p&gt;

</description>
      <category>cybersecurity</category>
      <category>html</category>
      <category>javascript</category>
      <category>security</category>
    </item>
    <item>
      <title>What Is a Parser Differential? How Can the Same Input Mean Different Things to Different Systems?</title>
      <dc:creator>Aditya Sharma</dc:creator>
      <pubDate>Sat, 19 Sep 2026 14:18:22 +0000</pubDate>
      <link>https://dev.to/aditya_d_sharma/what-is-a-parser-differential-how-can-the-same-input-mean-different-things-to-different-systems-p4g</link>
      <guid>https://dev.to/aditya_d_sharma/what-is-a-parser-differential-how-can-the-same-input-mean-different-things-to-different-systems-p4g</guid>
      <description>&lt;p&gt;A request arrives at a web application. Before it reaches the application, it passes through a security gateway. The gateway reads the request, checks it for anything suspicious, and decides it is safe. The request continues.&lt;/p&gt;

&lt;p&gt;The application receives what looks like the same request. But it interprets one part of it differently from how the gateway did.&lt;/p&gt;

&lt;p&gt;The application does something the gateway didn't expect.&lt;/p&gt;

&lt;p&gt;No one modified the request between the gateway and the application. The gateway made a correct decision based on how it understood the input. The application made a correct decision based on how it understood the input. Both components behaved exactly as designed.&lt;/p&gt;

&lt;p&gt;The problem wasn't that either component was wrong. The problem was that they disagreed.&lt;/p&gt;




&lt;h2&gt;
  
  
  What Is a Parser Differential?
&lt;/h2&gt;

&lt;p&gt;Every component that handles data has a parser: logic that reads raw input and turns it into structured meaning. Parsers make decisions about where values begin and end, how encoding should be handled, which characters are special, what to do with unexpected input.&lt;/p&gt;

&lt;p&gt;Different parsers follow different rules. They may be written by different teams, implement different versions of a specification, or handle edge cases in different ways. That's normal and usually harmless.&lt;/p&gt;

&lt;p&gt;It becomes dangerous when two parsers disagree across a security boundary.&lt;/p&gt;

&lt;p&gt;A security boundary is the point where one component decides whether input is safe to pass to another component. If the component making that decision parses the input differently from the component that will eventually consume it, the security decision may have been made against a representation that doesn't match what the downstream system will see.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Same Input
    ↓
┌──────────────────┬───────────────────┐
↓                  ↓
Security           Application
Component          Component
↓                  ↓
"Safe"             Different meaning
                   ↓
             Unexpected behavior
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The dangerous part: both components receive the same bytes. The disagreement is in how those bytes are interpreted.&lt;/p&gt;




&lt;h2&gt;
  
  
  Where the Disagreement Comes From
&lt;/h2&gt;

&lt;p&gt;Parsers disagree for several reasons, most of them mundane.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Encoding and decoding.&lt;/strong&gt; URLs and other data formats can represent characters through encoding schemes, and different components may decode those representations at different stages. A percent-encoded sequence like &lt;code&gt;%2F&lt;/code&gt; represents a forward slash. A security filter that inspects the raw encoded form sees &lt;code&gt;%2F&lt;/code&gt; as a literal string. An application that decodes before routing sees a &lt;code&gt;/&lt;/code&gt;. A path like &lt;code&gt;/safe%2F../admin&lt;/code&gt; might look safe to a filter examining the encoded form while resolving to &lt;code&gt;/admin&lt;/code&gt; for an application that decodes first.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Path normalization.&lt;/strong&gt; Paths often contain redundant segments. &lt;code&gt;/../&lt;/code&gt; refers to the parent directory. &lt;code&gt;/a/b/../c&lt;/code&gt; is equivalent to &lt;code&gt;/a/c&lt;/code&gt;. Components may normalize these differently, or not at all. A security filter that sees &lt;code&gt;/api/v1/../private&lt;/code&gt; and treats it as a path to &lt;code&gt;/api/v1/&lt;/code&gt; may be passing something an application resolves differently.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Duplicate parameters.&lt;/strong&gt; A URL query string can contain the same parameter name twice: &lt;code&gt;?id=1&amp;amp;id=2&lt;/code&gt;. Different frameworks and parsers handle this in different ways: some use the first value, some use the last, some combine them, some reject the request entirely. If a security filter validates one value while the application selects another, the filtered value was never the one that mattered.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Whitespace and delimiters.&lt;/strong&gt; HTTP headers, content-type values, and structured data fields have rules about how whitespace works. Parsers that are lenient about extra spaces or unusual characters can interpret boundaries differently from stricter parsers, which creates opportunities for one component to see a different field structure than another.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Malformed input.&lt;/strong&gt; When input doesn't follow the specification, parsers have to decide what to do. Many try to recover rather than reject. Two parsers following different recovery strategies can arrive at very different results from the same malformed bytes.&lt;/p&gt;




&lt;h2&gt;
  
  
  A Simple Example
&lt;/h2&gt;

&lt;p&gt;Consider a security filter that checks the path in a URL before passing it to an application.&lt;/p&gt;

&lt;p&gt;The filter receives:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="nf"&gt;GET&lt;/span&gt; &lt;span class="nn"&gt;/api/v1/%2e%2e/admin&lt;/span&gt; &lt;span class="k"&gt;HTTP&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="m"&gt;1.1&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The filter parses the path as &lt;code&gt;/api/v1/%2e%2e/admin&lt;/code&gt;. It checks this string against a blocklist. The path starts with &lt;code&gt;/api/v1/&lt;/code&gt;, which is an allowed prefix. It passes.&lt;/p&gt;

&lt;p&gt;The application receives the same request. But the application decodes percent-encoded characters before routing. It resolves &lt;code&gt;%2e%2e&lt;/code&gt; to &lt;code&gt;..&lt;/code&gt;, giving it &lt;code&gt;/api/v1/../admin&lt;/code&gt;, which normalizes to &lt;code&gt;/admin&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The filter checked one path. The application routed a different one.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Request:  GET /api/v1/%2e%2e/admin

Filter sees:     /api/v1/%2e%2e/admin  → allowed prefix → pass

Application:     /api/v1/%2e%2e/admin
                 decode percent-encoding
                 /api/v1/../admin
                 normalize
                 /admin  → serves protected resource
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Neither component made an error by its own rules. The filter correctly identified that the string started with an allowed prefix. The application correctly decoded and normalized the path. The problem is that they were operating on different representations of the same bytes.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why Security Filters Are Especially Exposed
&lt;/h2&gt;

&lt;p&gt;A WAF, API gateway, reverse proxy, or authentication middleware is in the business of making a decision about input before that input reaches the application. The security decision is meaningful only if the validator and the application agree on what the input means.&lt;/p&gt;

&lt;p&gt;When they don't, the validator is effectively checking one thing and the application is processing another. The validation still happened. The check was real. But it wasn't checking the value that actually matters.&lt;/p&gt;

&lt;p&gt;This is why "the input was validated" is a weaker guarantee than it sounds. Validated against which representation? Using which normalization rules? Compared to what the application will see?&lt;/p&gt;

&lt;p&gt;A security check is only meaningful when it is made against the same interpretation the eventual consumer will use.&lt;/p&gt;




&lt;h2&gt;
  
  
  Parser Differential vs. Related Problems
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;HTTP request smuggling&lt;/strong&gt; is a specific class of parser differential: disagreement between a front-end proxy and a back-end server about where one HTTP request ends and the next begins. Not every parser differential is request smuggling.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Path traversal&lt;/strong&gt; describes an outcome, not a mechanism. A parser differential in path normalization is one mechanism that can produce it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Canonicalization issues&lt;/strong&gt; are directly related. When different components reduce the same input to different standard forms, the result is a parser differential.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input validation&lt;/strong&gt; can fail silently here. A validator may be correct about the representation it examined while the application consumes a different one entirely.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why "Just Validate the Input" Isn't Always Enough
&lt;/h2&gt;

&lt;p&gt;The instinct when hearing about parser differentials is to say "just validate more carefully." But the issue isn't the quality of the validation. It's the sequence.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Raw input
↓
Validator parses and normalizes
↓
Makes security decision
↓
Application re-parses the original input
↓
Interprets it differently
↓
Acts on a representation the validator never examined
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The safer model is one where the security decision is made against the same canonical representation the application will consume: the form that results after the relevant decoding, normalization, and transformation steps have already been applied. The goal isn't just "parse before validating" as a rule of thumb. It's ensuring that the security component and the eventual consumer are working from the same interpretation of the input. If they aren't, validation is checking the wrong thing regardless of how carefully it's done.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Raw input
↓
Decode and normalize (same rules the application applies)
↓
Canonical representation
↓
Validate against this
↓
Application consumes the same representation
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Where This Appears in Modern Systems
&lt;/h2&gt;

&lt;p&gt;Parser differentials are an architectural problem, not a product-specific one. They appear wherever multiple components process the same input in sequence: reverse proxies inspecting paths before routing, API gateways validating parameters before forwarding, WAF rules applied to raw HTTP before the application framework parses request parameters, authentication middleware checking tokens before handing off to application code.&lt;/p&gt;

&lt;p&gt;The more independently those components are developed and configured, the more likely a parsing disagreement exists somewhere. Browser and server interpretation differences are also a real category: security tools that analyze traffic from the browser's perspective can miss what the server actually receives.&lt;/p&gt;




&lt;h2&gt;
  
  
  Defenses That Actually Follow From the Mechanism
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Parse before validating.&lt;/strong&gt; Make security decisions against the representation the application will use, not an earlier form of the input.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Normalize consistently.&lt;/strong&gt; If normalization happens, apply it once before validation, using the same rules the application would apply.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Reject ambiguous input.&lt;/strong&gt; Don't let individual components silently recover from malformed input in different ways. Inconsistent recovery is a common source of parser disagreements.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Minimize parsing layers.&lt;/strong&gt; Each component that independently parses the same input is another opportunity for disagreement. When security decisions need to be made, fewer intermediate parsers mean fewer chances for interpretation to diverge.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Test across the full path.&lt;/strong&gt; Testing the WAF in isolation tells you what the WAF thinks about the input. Testing the application in isolation tells you what the application thinks. Neither test tells you how they compare. Security testing should cover the combination.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Core Insight
&lt;/h2&gt;

&lt;p&gt;A security component can behave perfectly and still miss something, if what it analyzes isn't what the eventual consumer processes.&lt;/p&gt;

&lt;p&gt;Input doesn't have a single canonical meaning that every system agrees on. Meaning is assigned by parsers, and parsers differ. The security boundary only protects against what it actually understands.&lt;/p&gt;

&lt;p&gt;The most important gap to watch for in a system isn't always between trusted and untrusted input. Sometimes it's between two trusted components that receive the same bytes and arrive at different conclusions about what those bytes say.&lt;/p&gt;

</description>
      <category>cybersecurity</category>
      <category>infosec</category>
      <category>security</category>
    </item>
    <item>
      <title>What Is eBPF? How Can Linux Safely Run Programs Inside the Kernel?</title>
      <dc:creator>Aditya Sharma</dc:creator>
      <pubDate>Fri, 18 Sep 2026 05:11:14 +0000</pubDate>
      <link>https://dev.to/aditya_d_sharma/what-is-ebpf-how-can-linux-safely-run-programs-inside-the-kernel-1gao</link>
      <guid>https://dev.to/aditya_d_sharma/what-is-ebpf-how-can-linux-safely-run-programs-inside-the-kernel-1gao</guid>
      <description>&lt;p&gt;Linux can take a program written in user space, analyze it, compile it, and run it inside the kernel.&lt;/p&gt;

&lt;p&gt;That sentence deserves a second read.&lt;/p&gt;

&lt;p&gt;User space and kernel space are supposed to be separated. Kernel code runs with enormous privilege. A bug in kernel code can crash the system, corrupt memory, or break security boundaries in ways that a bug in an ordinary application cannot. The whole point of the user space / kernel space divide is to keep untrusted code far away from that level of the system.&lt;/p&gt;

&lt;p&gt;So how is Linux allowing code supplied from outside the kernel to run inside it?&lt;/p&gt;

&lt;p&gt;The answer is eBPF, and the interesting part isn't the feature. It's how Linux keeps that boundary from collapsing entirely.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Boundary That eBPF Crosses
&lt;/h2&gt;

&lt;p&gt;To understand why this is unusual, it helps to be clear about what the boundary normally looks like.&lt;/p&gt;

&lt;p&gt;When an application wants to do something privileged, it asks the kernel through a system call. The application doesn't execute inside the kernel. It makes a request, the kernel handles it, and control returns to user space. The separation is deliberate.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User Space
    ↓
System Call
    ↓
Kernel
    ↓
Hardware / Resources
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Kernel code itself is trusted by definition. Operating system developers wrote it. It was compiled into the kernel image. It went through code review and extensive testing. When the kernel runs something, it isn't asking permission.&lt;/p&gt;

&lt;p&gt;An eBPF program comes from somewhere else entirely. A developer writes it, compiles it to eBPF bytecode, and asks the kernel to load it. That request goes through a very different path than anything the kernel normally trusts.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User Space
    ↓
eBPF Program Submitted
    ↓
Kernel Verifier
    ↓
JIT Compilation
    ↓
Attached to Kernel Hook
    ↓
Executes When Event Fires
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The pipeline between "submitted" and "executes" is where the security model lives. Skip any step and the whole thing breaks.&lt;/p&gt;




&lt;h2&gt;
  
  
  What eBPF Actually Is
&lt;/h2&gt;

&lt;p&gt;eBPF (extended Berkeley Packet Filter, though the name is now largely historical) is a kernel subsystem that allows sandboxed programs to run at specific points inside the Linux kernel. These attachment points include system calls, network events, kernel functions, scheduler events, and more.&lt;/p&gt;

&lt;p&gt;The important thing isn't the list. It's the constraint: an eBPF program doesn't execute wherever it wants. The kernel decides where it can attach and what context it receives there. The program runs when that specific event fires, with access to only what that context provides.&lt;/p&gt;

&lt;p&gt;That controlled attachment is the first layer of restriction. But it's not the important one.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Verifier: Where Linux Decides Whether to Trust the Program
&lt;/h2&gt;

&lt;p&gt;Before an eBPF program runs a single instruction, the kernel's verifier analyzes it.&lt;/p&gt;

&lt;p&gt;The verifier is a static analysis engine. It reads the eBPF bytecode and analyzes possible execution paths before the program runs, tracking things like register states, pointer types, and memory-access bounds at each point. It doesn't execute the program. It reasons about what the program could do across all reachable paths and determines whether those possibilities satisfy a defined set of safety constraints.&lt;/p&gt;

&lt;p&gt;This includes:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Control flow.&lt;/strong&gt; The verifier checks that execution can't jump to arbitrary instructions, that every path terminates, and that the program doesn't loop indefinitely. eBPF programs must have bounded execution.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Memory access.&lt;/strong&gt; The verifier tracks what each register holds at every point. It knows whether a register contains a valid pointer, what kind, what memory region it addresses, and what bounds apply. If the program tries to dereference a pointer the verifier can't confirm is in-bounds, the program is rejected.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Pointer types.&lt;/strong&gt; A pointer into an eBPF map has different rules from a pointer to packet data. The verifier enforces these distinctions. A program can't take an arbitrary integer, cast it to a pointer, and dereference it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Helper calls.&lt;/strong&gt; eBPF programs can't call arbitrary kernel functions. Only a defined set of helper functions per program type is allowed. The verifier checks every call.&lt;/p&gt;

&lt;p&gt;To make pointer tracking concrete: a networking program might load with a register pointing to packet data and another holding the packet length. The verifier asks whether every memory access using that pointer stays within the available data on every possible execution path. If it can't establish that, the program is rejected before it runs.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;r1 → pointer to packet data
r2 → packet length

        ↓

Verifier asks:

Is r1 a valid pointer to packet data?
Is this offset within bounds for all code paths?
Can this access become invalid on any reachable path?

        ↓

All constraints satisfied → accepted
Any constraint fails → rejected
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The security model isn't "trust the developer." It's closer to: "the program must satisfy a defined set of safety constraints before it touches the kernel."&lt;/p&gt;




&lt;h2&gt;
  
  
  Helpers: A Controlled Interface Into the Kernel
&lt;/h2&gt;

&lt;p&gt;Instead of letting eBPF programs call arbitrary kernel functions, the kernel exposes a set of approved helper functions for each program type.&lt;/p&gt;

&lt;p&gt;A networking eBPF program might have helpers for reading packet data, modifying it, or redirecting it. A tracing program might have helpers for recording events. All of these are controlled interfaces. The eBPF program can't reach past them into arbitrary kernel internals.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;eBPF Program
     ↓
Helper Function
     ↓
Kernel-managed operation
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Helpers are defined per program type. A networking hook doesn't get the helpers meant for a tracing hook. The scope is intentionally narrow.&lt;/p&gt;




&lt;h2&gt;
  
  
  Maps: Shared State Between Kernel and User Space
&lt;/h2&gt;

&lt;p&gt;eBPF programs use maps: structured key-value stores accessible by both the in-kernel eBPF program and user-space applications. This is how security monitors and observability tools get data out of kernel-attached programs without needing a persistent kernel connection. The program writes events into the map; user space reads them.&lt;/p&gt;




&lt;h2&gt;
  
  
  JIT Compilation: Performance, Not Security
&lt;/h2&gt;

&lt;p&gt;Once verified, the kernel can JIT-compile the eBPF bytecode into native machine instructions. This matters because eBPF programs can run in extremely performance-sensitive paths like network packet processing.&lt;/p&gt;

&lt;p&gt;JIT compilation is not the security step. The verifier is. By the time JIT runs, the kernel has already decided the program is acceptable. Compilation is an optimization, not a trust decision.&lt;/p&gt;




&lt;h2&gt;
  
  
  eBPF vs. Kernel Modules
&lt;/h2&gt;

&lt;p&gt;A kernel module is the traditional way to extend the kernel. A loaded module executes with kernel privileges and can interact directly with kernel APIs and memory, giving it a much broader capability set than a verified eBPF program. Kernel modules are powerful, but the trust model is simple: you either trust the module completely or you don't load it.&lt;/p&gt;

&lt;p&gt;eBPF operates differently. The kernel doesn't trust the submitted program. It verifies it. The program executes in a restricted environment with controlled helpers, bounded execution, and verified memory access. The restriction is the point.&lt;/p&gt;

&lt;p&gt;This doesn't mean eBPF programs are less capable at the specific things they're designed to do. It means the capability they have is defined and constrained rather than open-ended.&lt;/p&gt;




&lt;h2&gt;
  
  
  eBPF As a Security Tool (And a Security Boundary)
&lt;/h2&gt;

&lt;p&gt;Security teams use eBPF because it can observe the system at a layer that application-level logging can't reach. A system-call monitoring tool built with eBPF can attach to kernel instrumentation points and observe system-call activity across processes, with coverage determined by the attachment points used and the permissions the tool holds. That kind of visibility is genuinely difficult to achieve from user space.&lt;/p&gt;

&lt;p&gt;The flip side: because eBPF operates inside the kernel, vulnerabilities in the verifier, JIT compiler, or helper implementations are kernel-level vulnerabilities. Bugs in the verifier have historically allowed privilege escalation. The security benefit comes with a significant responsibility: the components that implement eBPF's safety model are themselves security-critical.&lt;/p&gt;

&lt;p&gt;"eBPF has a verifier" and "eBPF can never have security vulnerabilities" are not the same statement.&lt;/p&gt;




&lt;h2&gt;
  
  
  Who Can Load eBPF Programs?
&lt;/h2&gt;

&lt;p&gt;Not everyone. Loading eBPF programs requires appropriate Linux capabilities, and the specific requirements depend on the program type, kernel version, and system configuration. The security model isn't just the verifier: it's capability controls plus verification plus restricted execution plus controlled helpers, all together. Removing any layer changes the security properties.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Answer to the Original Question
&lt;/h2&gt;

&lt;p&gt;How can Linux allow a program supplied from user space to execute inside the kernel without simply giving that program unrestricted access?&lt;/p&gt;

&lt;p&gt;By making the boundary programmable but not removing it.&lt;/p&gt;

&lt;p&gt;The kernel doesn't simply trust the submitted program. It analyzes the bytecode and requires it to satisfy a set of safety constraints before allowing it to execute anywhere near kernel memory. The verifier enforces those constraints. The helper interface constrains what the program can do even after verification. The attachment model constrains where it runs. Capability requirements constrain who can load it.&lt;/p&gt;

&lt;p&gt;What eBPF offers isn't unrestricted kernel access dressed up with a safety label. It's a constrained execution model where the kernel decides what "safe enough for kernel execution" means, and then enforces that definition before a single instruction runs.&lt;/p&gt;

&lt;p&gt;The boundary still exists. eBPF just made it programmable.&lt;/p&gt;

</description>
      <category>architecture</category>
      <category>linux</category>
      <category>programming</category>
      <category>security</category>
    </item>
    <item>
      <title>What Is Subdomain Takeover? How Can an Abandoned Subdomain Become Dangerous?</title>
      <dc:creator>Aditya Sharma</dc:creator>
      <pubDate>Thu, 17 Sep 2026 05:12:22 +0000</pubDate>
      <link>https://dev.to/aditya_d_sharma/what-is-subdomain-takeover-how-can-an-abandoned-subdomain-become-dangerous-2j21</link>
      <guid>https://dev.to/aditya_d_sharma/what-is-subdomain-takeover-how-can-an-abandoned-subdomain-become-dangerous-2j21</guid>
      <description>&lt;p&gt;A company owns company.com. That much is clear. Their registrar shows it. Their DNS is configured by them. Nobody disputes it.&lt;/p&gt;

&lt;p&gt;But one of their subdomains, blog.company.com, starts serving content the company didn't put there.&lt;/p&gt;

&lt;p&gt;The company didn't lose their domain. The DNS record is still theirs. So what happened?&lt;/p&gt;




&lt;h2&gt;
  
  
  Two Different Things Called "Ownership"
&lt;/h2&gt;

&lt;p&gt;When an organization creates a subdomain like blog.company.com, they control the DNS record for it. That record typically points somewhere: a cloud hosting platform, a SaaS documentation tool, a CDN distribution, a third-party service.&lt;/p&gt;

&lt;p&gt;Here's the part that trips people up: controlling the DNS record and controlling the resource it points to are not the same thing.&lt;/p&gt;

&lt;p&gt;The DNS record says: "When you ask for blog.company.com, go here."&lt;/p&gt;

&lt;p&gt;But "here" is often infrastructure owned and operated by an external provider. The organization doesn't own that infrastructure. They have a project or account on it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;company.com (organization owns this)
    ↓
blog.company.com DNS record (organization controls this)
    ↓
CNAME → company-blog.hosting-platform.example
    ↓
Hosted resource (external provider owns and manages this)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That separation is normal and useful. The issue appears when one side of that relationship goes away.&lt;/p&gt;




&lt;h2&gt;
  
  
  What a Dangling DNS Record Is
&lt;/h2&gt;

&lt;p&gt;Organizations create infrastructure constantly: marketing pages, documentation sites, staging environments, developer tools, short-lived campaign sites. Resources get provisioned and later deleted. Teams change. Projects wind down.&lt;/p&gt;

&lt;p&gt;When an external resource is deleted, the DNS record that pointed to it often doesn't get cleaned up. Nobody removes it because nobody thinks to. The original team moved on. The infrastructure list wasn't updated.&lt;/p&gt;

&lt;p&gt;The DNS record still exists. It still points to where the resource used to be. But the resource itself is gone.&lt;/p&gt;

&lt;p&gt;This is a dangling DNS record: a valid hostname pointing to an external destination where the referenced resource no longer exists.&lt;/p&gt;

&lt;p&gt;A dangling record by itself isn't necessarily dangerous. The next question is whether that vacancy can be filled by someone else.&lt;/p&gt;




&lt;h2&gt;
  
  
  When It Becomes a Takeover
&lt;/h2&gt;

&lt;p&gt;Some external providers allow new users to claim a resource at the same identifier that a previous account used. This is often by design. Platforms need to recycle project names, subdomains, or endpoints so they can be reused.&lt;/p&gt;

&lt;p&gt;If that's possible, an attacker can do the following:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Find a subdomain that still points to an external platform.&lt;/li&gt;
&lt;li&gt;Discover that the resource on that platform is no longer active.&lt;/li&gt;
&lt;li&gt;Create a new account on that platform and claim the same identifier.&lt;/li&gt;
&lt;li&gt;The organization's DNS record now points to the attacker's resource.
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User visits blog.company.com
       ↓
DNS resolves via CNAME
       ↓
Hosting platform
       ↓
Attacker's claimed resource
       ↓
Attacker-controlled content served under company's subdomain
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The attacker didn't touch company.com. They didn't modify the DNS record. They simply occupied a vacancy that the DNS record kept pointing at.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why That Subdomain Context Matters
&lt;/h2&gt;

&lt;p&gt;Content served under blog.company.com inherits the visual and reputational context of the organization. A visitor navigating there sees the company's domain in the address bar. It looks legitimate. The TLS certificate may resolve without errors because it's issued for a valid hostname.&lt;/p&gt;

&lt;p&gt;Depending on what the subdomain was previously used for and how the organization's other applications are configured, consequences can range from reputational damage and phishing risk to more significant issues in some architectures.&lt;/p&gt;

&lt;p&gt;That said, the exact impact isn't fixed. It depends on what the subdomain was used for, how cookies are scoped, whether other services trust that hostname, and what the attacker actually does with it. Subdomain takeover doesn't automatically break session security or compromise parent domain applications. But serving attacker-controlled content under a trusted hostname is still a meaningful security failure.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Ownership Boundary Is the Core Idea
&lt;/h2&gt;

&lt;p&gt;It's worth being precise here because the setup is counterintuitive.&lt;/p&gt;

&lt;p&gt;The organization did not lose its domain. company.com is still registered in their name. Their DNS is still under their control. Nothing was stolen.&lt;/p&gt;

&lt;p&gt;What happened is that DNS ownership and resource ownership have separate lifecycles, and only one of them was maintained.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Who controls company.com?
    → The organization.

Who controls the DNS record for blog.company.com?
    → The organization.

Who controls the resource that record points to?
    → The external provider manages the platform.
      The organization controlled a resource on it.
      That resource no longer exists.
      The provider may allow someone else to claim it.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The organization's authority stops at the DNS record. It doesn't extend into the external provider's resource namespace. When the resource was deleted, that external slot opened up.&lt;/p&gt;




&lt;h2&gt;
  
  
  How This Differs From Similar Problems
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;DNS rebinding&lt;/strong&gt; works differently. It involves changing DNS resolution over time, causing a client to connect to a different IP address than it resolved initially. The manipulation happens through DNS TTL and timing. Subdomain takeover is a static configuration problem, not a dynamic resolution problem.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Domain hijacking&lt;/strong&gt; is about unauthorized control of the domain registration itself or the DNS authority. The attacker gains the ability to change DNS records for the domain. In a subdomain takeover, the organization retains full control of their DNS. The attacker never touches it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Dangling DNS&lt;/strong&gt; is the underlying condition. Subdomain takeover is the security consequence when an attacker can claim the external resource that the dangling record points to.&lt;/p&gt;




&lt;h2&gt;
  
  
  Defenses That Follow From the Mechanism
&lt;/h2&gt;

&lt;p&gt;Because the problem is a configuration drift between DNS records and external resources, the defenses are about keeping those two things synchronized.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Remove DNS records when external resources are deleted.&lt;/strong&gt; This is the most direct fix. When a project is decommissioned, the DNS record should be decommissioned with it. The challenge is that this requires DNS and infrastructure changes to be coordinated rather than handled by separate teams independently.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Inventory subdomains and track what they point to.&lt;/strong&gt; Many organizations don't have a clear map of which subdomains exist and which external services they depend on. Without that visibility, stale records are easy to miss.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Include DNS in infrastructure lifecycle processes.&lt;/strong&gt; Provisioning a new resource and creating a DNS record should be a paired operation. Deleting a resource and removing its DNS record should be equally paired.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Monitor for dangling references.&lt;/strong&gt; Organizations can periodically verify that DNS records resolve to resources they still control. A record that points to an unclaimed external resource is a signal worth acting on.&lt;/p&gt;

&lt;p&gt;Monitoring catches drift, but the more fundamental fix is keeping DNS and external resource lifecycles synchronized from the start.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Mental Model
&lt;/h2&gt;

&lt;p&gt;A domain name is an address. An address tells you where to go. It doesn't say anything about who currently controls the destination.&lt;/p&gt;

&lt;p&gt;When an organization creates a subdomain that points to an external service, they're establishing a relationship between a hostname they own and a resource they control on someone else's infrastructure. Both sides of that relationship need to stay maintained.&lt;/p&gt;

&lt;p&gt;If the external resource is deleted and the DNS record isn't, the address still exists. The destination is just vacant. And in some cases, vacant can be claimed.&lt;/p&gt;

&lt;p&gt;The company still owns the domain. The DNS record is still valid. The problem is what that record has been pointing to.&lt;/p&gt;

&lt;p&gt;A hostname can continue looking legitimate long after the infrastructure behind it changed hands.&lt;/p&gt;

</description>
      <category>cybersecurity</category>
      <category>infosec</category>
      <category>security</category>
    </item>
    <item>
      <title>What Is a Confused Deputy? How Can a Trusted Program Be Tricked Into Using Its Own Permissions?</title>
      <dc:creator>Aditya Sharma</dc:creator>
      <pubDate>Wed, 16 Sep 2026 07:00:27 +0000</pubDate>
      <link>https://dev.to/aditya_d_sharma/what-is-a-confused-deputy-how-can-a-trusted-program-be-tricked-into-using-its-own-permissions-4em4</link>
      <guid>https://dev.to/aditya_d_sharma/what-is-a-confused-deputy-how-can-a-trusted-program-be-tricked-into-using-its-own-permissions-4em4</guid>
      <description>&lt;p&gt;You don't have access to the file. The rules are clear about that. Your account can't read it.&lt;/p&gt;

&lt;p&gt;But there's a service running on the same system that can. It's trusted. It has elevated permissions. And it accepts requests.&lt;/p&gt;

&lt;p&gt;So instead of reading the file yourself, you ask the service to read it for you.&lt;/p&gt;

&lt;p&gt;The service checks whether it can access the file. It can. So it does.&lt;/p&gt;

&lt;p&gt;You just read something you weren't allowed to read. And the service did exactly what it was designed to do.&lt;/p&gt;




&lt;h2&gt;
  
  
  What Just Happened
&lt;/h2&gt;

&lt;p&gt;The attacker didn't steal any permissions. Nothing was compromised. No vulnerability was exploited in the traditional sense.&lt;/p&gt;

&lt;p&gt;What happened is simpler and, in a way, more interesting: a trusted component used its own authority to fulfill a request that the caller wasn't authorized to make.&lt;/p&gt;

&lt;p&gt;The deputy didn't get hacked. The deputy got confused.&lt;/p&gt;

&lt;p&gt;In security terms, a &lt;strong&gt;deputy&lt;/strong&gt; is a component that acts on behalf of others. A file service, a cloud function, a privileged API, a system daemon. These components often hold more authority than their callers, because they perform operations users can't or shouldn't perform directly.&lt;/p&gt;

&lt;p&gt;A deputy becomes &lt;strong&gt;confused&lt;/strong&gt; when it can't distinguish between what it is allowed to do and what the caller is allowed to ask it to do. It applies its own authority when it should be applying the caller's.&lt;/p&gt;

&lt;p&gt;Those are two very different things.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Central Problem: Two Separate Authority Questions
&lt;/h2&gt;

&lt;p&gt;When a caller sends a request to a privileged service, there are two questions the service needs to answer:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What can I (the service) access?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What is this caller allowed to access through me?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A secure design answers the second question. A confused deputy only answers the first.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Caller: "Give me resource X"
         ↓
Trusted service
         ↓
"Can I access resource X?" → Yes
         ↓
Returns resource X
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The service validates its own authority. It never validates the caller's authority. So the caller's permissions are irrelevant. Whatever the service can access, the caller can access through it.&lt;/p&gt;

&lt;p&gt;The correct version of that flow looks different:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Caller: "Give me resource X"
         ↓
Trusted service
         ↓
"Is this caller allowed to access resource X?" → ?
         ↓
Allow or deny based on that answer
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The service's own permissions are a prerequisite for performing the operation. They are not authorization for the caller to request it.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why This Happens in Real Systems
&lt;/h2&gt;

&lt;p&gt;Confused deputy problems appear naturally in systems that delegate work.&lt;/p&gt;

&lt;p&gt;Delegation is useful. A privileged service often needs broader permissions than individual users because it operates across many users' contexts. A file processing service needs to read many files. A cloud function needs to call many downstream APIs. A logging daemon needs write access to system directories.&lt;/p&gt;

&lt;p&gt;The architecture isn't the problem. The problem is the security question that gets lost inside it.&lt;/p&gt;

&lt;p&gt;When a service acts on behalf of a user, whose authority should govern what the service is allowed to do? The service's authority tells you what it is technically capable of. It says nothing about what this particular caller is entitled to cause.&lt;/p&gt;

&lt;p&gt;If the service never asks the second question, its entire permission set becomes accessible to anyone who can send it a request.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Same Pattern in Cloud Systems
&lt;/h2&gt;

&lt;p&gt;This isn't a problem specific to operating systems or classic privileged programs. The same structure appears constantly in modern cloud architectures.&lt;/p&gt;

&lt;p&gt;A user can call Service A. Service A has an IAM role that allows it to access storage buckets, query databases, invoke other services. The role exists because Service A legitimately needs those permissions to function.&lt;/p&gt;

&lt;p&gt;If Service A accepts an arbitrary resource identifier from the user and blindly uses its IAM role to access whatever was requested, the user can reach anything Service A can reach:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User (limited permissions)
     ↓
Service A
     ↓
Service A's IAM role (broader permissions)
     ↓
Resource the user couldn't access directly
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The IAM role isn't the problem. The role is doing exactly what IAM roles are supposed to do. The problem is that Service A treats its own IAM authority as authorization for any caller, rather than enforcing what each caller is actually entitled to request.&lt;/p&gt;




&lt;h2&gt;
  
  
  How This Differs From Similar Vulnerabilities
&lt;/h2&gt;

&lt;p&gt;It's worth being precise here, because confused deputy gets conflated with a few other patterns.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SSRF&lt;/strong&gt; is about influencing where a server makes a network request. The attacker's goal is often to reach internal infrastructure. A confused deputy is a different question: the attacker isn't just redirecting a request, they're causing a trusted component to apply its own authority on their behalf. There can be overlap in real systems, but the underlying mechanism is distinct.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;IDOR&lt;/strong&gt; is about accessing an object through a predictable identifier that isn't properly access-controlled. The resource is exposed directly. A confused deputy involves an intermediary that holds the authority, not a directly accessible resource.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Privilege escalation&lt;/strong&gt; is a broad category. Confused deputy can lead to unauthorized access, but the mechanism isn't about exploiting a flaw to gain new permissions. The deputy already has the permissions. The problem is that it applies them without checking whether the caller deserves the result.&lt;/p&gt;




&lt;h2&gt;
  
  
  Defenses That Follow From the Mechanism
&lt;/h2&gt;

&lt;p&gt;Because the core issue is authority confusion, the defenses are about making the authority distinction explicit.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Enforce authorization at the right boundary.&lt;/strong&gt; The service must evaluate whether the caller is allowed to perform the requested operation, not just whether the service itself can perform it. These are different checks and both are required.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Carry caller context through delegation chains.&lt;/strong&gt; When a service acts on behalf of a user, the user's identity and permissions need to be part of the authorization decision, not discarded at the service boundary. A request that passes through multiple layers should not lose track of whose request it originally was.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Constrain what callers can request.&lt;/strong&gt; A service that can access a hundred resources doesn't need to expose all of them to every caller. Scope the operations each caller can invoke rather than exposing the service's full capability set.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Apply least privilege to the service itself.&lt;/strong&gt; A service should only hold the permissions it actually needs. This doesn't eliminate confused deputy problems, but it limits what an attacker can access if the deputy does get confused.&lt;/p&gt;

&lt;p&gt;None of these individually close every confused deputy scenario. The fundamental requirement is that authorization decisions account for the caller's authority, not just the deputy's.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Core Insight
&lt;/h2&gt;

&lt;p&gt;The intuitive assumption is that a trusted component is safe because it's trusted. Its permissions are legitimate. Its operations are intentional.&lt;/p&gt;

&lt;p&gt;But trust in a component is not the same as trust in every request that component receives. A service can be completely legitimate and still be the mechanism through which unauthorized access happens. It doesn't need to be compromised. It just needs to fail at one question.&lt;/p&gt;

&lt;p&gt;The most important mental model from this article:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What the deputy can do and what the caller can cause the deputy to do are two separate things. A confused deputy forgets the difference.&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>cybersecurity</category>
      <category>infosec</category>
      <category>security</category>
    </item>
  </channel>
</rss>
