<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Suchitra Shankar</title>
    <description>The latest articles on DEV Community by Suchitra Shankar (@suchitra_shankar).</description>
    <link>https://dev.to/suchitra_shankar</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F2758132%2F0e872736-abed-401f-a83f-d8020a6bf8c4.png</url>
      <title>DEV Community: Suchitra Shankar</title>
      <link>https://dev.to/suchitra_shankar</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/suchitra_shankar"/>
    <language>en</language>
    <item>
      <title>Tiered or Leveled? Why Not Both; Building Amethyst.</title>
      <dc:creator>Suchitra Shankar</dc:creator>
      <pubDate>Thu, 11 Jun 2026 17:31:36 +0000</pubDate>
      <link>https://dev.to/suchitra_shankar/tiered-or-leveled-why-not-both-building-amethyst-4g50</link>
      <guid>https://dev.to/suchitra_shankar/tiered-or-leveled-why-not-both-building-amethyst-4g50</guid>
      <description>&lt;p&gt;What if your storage engine could "understand" the incoming workload and switch its compaction strategy accordingly at runtime?&lt;/p&gt;

&lt;p&gt;When building a database engine, you are typically forced to choose between optimizing for reads or optimizing for writes. Yet real workloads rarely stay the same.&lt;/p&gt;

&lt;p&gt;So why do we still treat compaction as a static decision in a world where almost every other system component adapts at runtime?&lt;/p&gt;

&lt;p&gt;We decided we didn't want to pick.&lt;/p&gt;

&lt;p&gt;Enter &lt;strong&gt;Amethyst&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Amethyst is a lightweight, adaptive Log-Structured Merge (LSM) tree prototype built to dynamically switch compaction behavior at the segment level based on observed runtime workloads.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. What &lt;em&gt;is&lt;/em&gt; an LSM?
&lt;/h2&gt;

&lt;p&gt;Amethyst is built on a Log-Structured Merge tree (LSM tree). To understand our approach, let's look at the diagram below, which shows the architecture of a typical LSM-based storage engine.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Ffdm9md8zc5ssp4yjdsnj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Ffdm9md8zc5ssp4yjdsnj.png" alt="Figure 1. Anatomy of an LSM-based storage engine. Writes are buffered in the in-memory MemTable (and appended to the WAL for durability) before being flushed to disk as immutable SSTables. Background compaction periodically merges these SSTables to bound file count and read latency — the component Amethyst makes adaptive." width="800" height="281"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;At a high level, an LSM tree achieves high write throughput by converting random writes into sequential disk operations. Incoming updates are first buffered in memory inside a &lt;strong&gt;MemTable&lt;/strong&gt;. To ensure durability, every update is also appended to a &lt;strong&gt;Write-Ahead Log (WAL)&lt;/strong&gt; before being acknowledged.&lt;/p&gt;

&lt;p&gt;Once the MemTable reaches a size threshold, it is flushed to disk as an immutable &lt;strong&gt;Sorted String Table (SSTable)&lt;/strong&gt;. Over time, these SSTables accumulate across the storage hierarchy.&lt;/p&gt;

&lt;p&gt;To prevent the number of files from growing indefinitely and to keep read performance under control, the storage engine periodically merges SSTables in the background through a process known as &lt;strong&gt;compaction&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;This compaction process is the heart of every LSM tree, and it is precisely the component Amethyst seeks to make adaptive.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Compaction Policies
&lt;/h2&gt;

&lt;p&gt;Most storage engines rely on a single compaction policy chosen when the database is initialized.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fnfzc9sbvlluoz3ceefww.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fnfzc9sbvlluoz3ceefww.png" alt="Leveled compaction keeps non-overlapping SSTables within each level, minimizing read amplification at the cost of repeated rewrites (high write amplification). " width="800" height="489"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The first option is &lt;strong&gt;Leveled Compaction&lt;/strong&gt;. Here, the storage hierarchy is divided into levels, each with a fixed capacity. SSTables within a level are kept non-overlapping, allowing lookups to quickly narrow down where a key might exist. This significantly reduces read amplification and improves lookup performance.&lt;/p&gt;

&lt;p&gt;The downside is write amplification. To maintain these non-overlapping levels, the database must repeatedly rewrite data as it moves through the storage hierarchy. The result is excellent read performance at the cost of additional write overhead.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fev9ln23lhrp7wgz9oywr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fev9ln23lhrp7wgz9oywr.png" alt="Tiered compaction lets similarly-sized SSTables accumulate before merging, slashing write amplification but forcing reads to search multiple overlapping runs." width="800" height="586"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The second option is &lt;strong&gt;Tiered Compaction&lt;/strong&gt;. Instead of enforcing strict non-overlapping levels, multiple SSTables are allowed to coexist within the same level. The database waits until a collection of similarly sized SSTables accumulates before merging them into a larger run.&lt;/p&gt;

&lt;p&gt;This dramatically reduces write amplification because compactions occur less frequently and data is rewritten fewer times. However, the tradeoff appears during reads. Since multiple runs may contain overlapping key ranges, lookups often need to search several SSTables before finding the requested key.&lt;/p&gt;

&lt;p&gt;For decades, storage engines have forced developers to make a permanent compromise between these two approaches.&lt;/p&gt;

&lt;p&gt;The problem is that real-world workloads are rarely static.&lt;/p&gt;

&lt;p&gt;A database might spend the morning ingesting large batches of data, making a write-optimized tiered layout ideal. Later, that same database might transition to serving read-heavy analytical queries, where a leveled layout would provide significantly better lookup performance.&lt;/p&gt;

&lt;p&gt;Static compaction policies cannot adapt to these workload shifts. Once a strategy is chosen, the database remains locked into that decision, even when workload characteristics change dramatically.&lt;/p&gt;

&lt;p&gt;This raises an obvious question: why should every part of the database be forced to use the same compaction strategy?&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Observations
&lt;/h2&gt;

&lt;p&gt;If you look across a modern software stack, systems are adapting continuously.&lt;/p&gt;

&lt;p&gt;TCP congestion control dynamically adjusts its window in response to network conditions. Dynamic schedulers rebalance resources as demand shifts. Adaptive caching policies such as ARC continuously balance recency against frequency, while autoscaling systems provision infrastructure based on real-time load signals.&lt;/p&gt;

&lt;p&gt;Yet, for some reason, compaction policy in most storage engines remains entirely static.&lt;/p&gt;

&lt;p&gt;Once a database chooses between tiered and leveled compaction, that decision often remains fixed for the lifetime of the system, regardless of how dramatically the workload changes.&lt;/p&gt;

&lt;p&gt;Back in 2018, database expert Mark Callaghan posed a simple but compelling question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Tiered or leveled compaction, why not both via adaptive compaction?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That question became the foundation for Amethyst.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Building Amethyst
&lt;/h2&gt;

&lt;p&gt;The observation was simple: workloads change, but compaction policies do not.&lt;/p&gt;

&lt;p&gt;Turning that idea into a working storage engine was significantly harder.&lt;/p&gt;

&lt;p&gt;An adaptive compaction system must answer three fundamental questions:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;What should adapt?&lt;/li&gt;
&lt;li&gt;How should workload shifts be detected?&lt;/li&gt;
&lt;li&gt;How can adaptation occur without introducing more overhead than it removes?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Existing approaches often operate at a coarse granularity, making decisions for an entire database or an entire level at once. While this simplifies management, it also makes adaptation expensive and slow to react.&lt;/p&gt;

&lt;p&gt;We wanted to explore a different direction.&lt;/p&gt;

&lt;p&gt;Rather than treating the database as a single unit, Amethyst treats individual segments as independently adaptable entities. Each segment continuously observes its own workload characteristics and can make localized compaction decisions without requiring the rest of the system to change.&lt;/p&gt;

&lt;p&gt;This design allows different regions of the same LSM tree to exhibit different behaviors simultaneously. Segments experiencing heavy write traffic can remain optimized for ingestion, while segments serving read-heavy workloads can gradually migrate toward layouts optimized for lookup performance.&lt;/p&gt;

&lt;p&gt;In other words, different parts of the same database can follow different compaction strategies at the same time.&lt;/p&gt;

&lt;p&gt;Once we committed to segment-level adaptation, the rest of the architecture began to fall into place.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F406vlyeyl43cpcfnbid4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F406vlyeyl43cpcfnbid4.png" alt="amethyst arch" width="800" height="544"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Design Decisions
&lt;/h2&gt;

&lt;p&gt;When we sat down to architect Amethyst, we had to ensure our adaptation mechanics wouldn't introduce more overhead than they solved.&lt;/p&gt;

&lt;p&gt;The biggest problem we noticed with prior adaptive systems is that they try to change the state of the entire database, or even an entire level, all at once. Amethyst shifts this granularity by making the SSTable, along with its associated metadata, the fundamental unit of adaptation. We call this unit a &lt;strong&gt;segment&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F11w26e9e8vgphoeguaw0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F11w26e9e8vgphoeguaw0.png" alt="segment" width="800" height="499"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;By giving each segment its own metadata tracking structure, including &lt;strong&gt;MinKey&lt;/strong&gt;, &lt;strong&gt;MaxKey&lt;/strong&gt;, &lt;strong&gt;Strategy Tags&lt;/strong&gt;, and physical &lt;strong&gt;I/O counters&lt;/strong&gt;, the database can make localized decisions instead of applying a single policy globally.&lt;/p&gt;

&lt;p&gt;However, runtime adaptation inherently risks thrashing: burning CPU cycles and disk I/O by constantly rewriting data as workloads fluctuate. We mitigated this by designing our finite state machine (FSM) with asymmetric trigger thresholds governed by Exponential Moving Averages (EMAs).&lt;/p&gt;

&lt;p&gt;Write pressure is sudden and tends to degrade performance if left unaddressed, so the FSM aggressively transitions a segment to &lt;strong&gt;Tiered Compaction&lt;/strong&gt; after just ten logged write operations. Conversely, read patterns are highly volatile. To prevent short-lived read spikes from triggering an expensive leveled merge, the FSM requires a sustained signal of 500 read operations before committing to a state change.&lt;/p&gt;

&lt;p&gt;We also designed the compaction background thread to perform localized multi-way merges. Only the flagged segment and its immediate overlapping neighbors are read into memory and rewritten. Because SSTables are immutable, background ingestion and user read queries can continue concurrently without interruption.&lt;/p&gt;

&lt;p&gt;For this initial proof of concept, we implemented both &lt;strong&gt;Tiered&lt;/strong&gt; and &lt;strong&gt;Leveled Compaction&lt;/strong&gt; ourselves within the same Go codebase. This eliminated confounding variables and ensured that we were measuring the architectural merit of the adaptation logic itself, rather than differences between independent implementations.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. The math
&lt;/h2&gt;

&lt;p&gt;Adaptive systems are only useful if they can correctly identify when conditions have changed.&lt;/p&gt;

&lt;p&gt;This immediately raises a difficult question: how does a storage engine know when a workload has shifted from being write-heavy to read-heavy, or vice versa?&lt;/p&gt;

&lt;p&gt;A naive approach would be to continuously scan historical workload data and perform expensive analysis. While this may produce accurate classifications, it also introduces significant runtime overhead. Since one of the design goals of Amethyst was to remain lightweight, we needed a simpler mechanism.&lt;/p&gt;

&lt;p&gt;We ultimately settled on tracking a small set of workload signals and smoothing them using Exponential Moving Averages (EMAs). Unlike simple averages, EMAs place greater weight on recent observations while still retaining some historical context. This allowed the system to react to workload shifts without overreacting to short-lived spikes.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fwy8fwqrmb89hiwoepvo3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fwy8fwqrmb89hiwoepvo3.png" alt="FORMULA FOR EMA" width="738" height="145"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;For each segment, Amethyst continuously tracks metrics: read frequency, write frequency, and compaction activity. These metrics are updated incrementally as operations occur, allowing the system to maintain an evolving view of segment behavior with minimal overhead.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F7zdszn0j4bj138bswvbq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F7zdszn0j4bj138bswvbq.png" alt="The per-segment finite state machine with asymmetric, EMA-driven thresholds. Write pressure transitions a segment to tiered compaction aggressively (after 10 logged writes), while a leveled transition requires a sustained 500 read operations : hysteresis that prevents short-lived spikes from triggering expensive merges and thrashing" width="800" height="172"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The decision engine then evaluates these smoothed metrics against predefined thresholds. Segments that exhibit sustained write-heavy behavior are gradually steered toward tiered compaction, while segments experiencing sustained read pressure are moved toward leveled compaction.&lt;/p&gt;

&lt;p&gt;An important design consideration was avoiding oscillation. If policy changes occur too aggressively, segments can repeatedly switch between strategies in response to transient workload fluctuations. This not only wastes resources but can also degrade overall performance.&lt;/p&gt;

&lt;p&gt;To address this, Amethyst introduces hysteresis into its decision process. Policy transitions require persistent evidence that a workload has changed, rather than reacting immediately to every short-term variation. In practice, this significantly improves stability and prevents unnecessary switching.&lt;/p&gt;

&lt;p&gt;The result is a lightweight adaptation mechanism that requires no offline training, forecasting, or complex optimization models. Instead, the system continuously observes its own workload and makes local decisions using a small amount of runtime state.&lt;/p&gt;

&lt;p&gt;The mathematics behind Amethyst is intentionally simple. The objective was never to build the most sophisticated prediction model possible. The objective was to determine whether lightweight workload estimation alone could enable effective adaptive compaction.&lt;/p&gt;

&lt;p&gt;The results suggest that, at least for the workloads we evaluated, the answer is yes.&lt;/p&gt;

&lt;h2&gt;
  
  
  7. Results
&lt;/h2&gt;

&lt;p&gt;The primary goal of Amethyst was not to outperform every existing compaction strategy in every workload. If a workload remains perfectly stable, a carefully chosen static policy will often perform extremely well.&lt;/p&gt;

&lt;p&gt;Instead, our goal was to determine whether a storage engine could adapt to changing workload characteristics at runtime and benefit from doing so.&lt;/p&gt;

&lt;p&gt;To evaluate this, we benchmarked Amethyst against static &lt;strong&gt;Leveled&lt;/strong&gt; and &lt;strong&gt;Tiered&lt;/strong&gt; compaction configurations across a variety of workload mixes. We focused particularly on workloads whose read/write ratios changed over time, since these are precisely the scenarios where a single static policy becomes difficult to justify.&lt;/p&gt;

&lt;p&gt;The results were encouraging.&lt;/p&gt;

&lt;p&gt;Across the evaluated workloads, Amethyst achieved throughput improvements of up to &lt;strong&gt;2.2×&lt;/strong&gt; compared to static compaction strategies.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F7hbew4sadks8n8ilpesh.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F7hbew4sadks8n8ilpesh.png" alt="results table" width="800" height="476"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;More importantly, these improvements appeared in workloads that experienced meaningful shifts in behavior during execution. During write-heavy phases, the system naturally favored tiered-style behavior to minimize write amplification. As read intensity increased, segments gradually transitioned toward leveled layouts, improving lookup efficiency and reducing read amplification.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fq4sdw3tyw579yulvz825.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fq4sdw3tyw579yulvz825.png" alt="ra &amp;amp; wa graph" width="800" height="488"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This adaptive behavior allowed the system to avoid committing permanently to either extreme.&lt;/p&gt;

&lt;p&gt;Perhaps the most interesting result was that adaptation could be achieved using relatively lightweight runtime signals. Amethyst does not require offline tuning, workload forecasting, or complex analytical optimization models. Instead, it relies on continuously observed workload statistics and local segment-level decisions.&lt;/p&gt;

&lt;p&gt;Of course, the gains were not universal. In workloads that remained highly stable, static policies often remained competitive. This was expected. Amethyst was designed specifically for environments where workload characteristics evolve over time, and the results suggest that adaptive compaction is most beneficial in exactly those situations.&lt;/p&gt;

&lt;p&gt;Overall, the evaluation provided evidence that segment-level runtime policy selection is both practical and effective. While there is significant room for improvement in future versions, the results validate the central hypothesis that motivated the project:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A storage engine does not necessarily need to commit to a single compaction strategy for its entire lifetime.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  8. Things We Got Wrong (And What's Next)
&lt;/h2&gt;

&lt;p&gt;Like most research projects, Amethyst did not emerge fully formed. Many of our initial assumptions turned out to be incomplete, naive, or simply incorrect.&lt;/p&gt;

&lt;p&gt;Full disclosure: we built this prototype entirely on our own, without a formal research advisor. That meant a lot of early trial and error. When we hit a wall trying to figure out the right benchmarking metrics, Mark Callaghan's insights were a huge help in getting us unstuck. But for the overarching design, we were largely figuring things out as we went.&lt;/p&gt;

&lt;p&gt;Since we were building this architecture completely from the ground up, we had to make several deliberate trade-offs. Here is a look at what we left out of the first version, and exactly what we plan to tackle next.&lt;/p&gt;

&lt;h3&gt;
  
  
  Space Amplification
&lt;/h3&gt;

&lt;p&gt;Right now, our space amplification is quite high at &lt;strong&gt;8.57×&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;While this would be unacceptable for a production database, it is not a fundamental flaw of adaptive compaction. Garbage collection was intentionally left out of the prototype to keep our benchmarking clean and isolate the behavior of the adaptation FSM.&lt;/p&gt;

&lt;p&gt;Now that we have validated the read and write performance improvements, one of our immediate priorities is implementing a proper background reclamation thread to bring space amplification down to more practical levels.&lt;/p&gt;

&lt;h3&gt;
  
  
  Workload Characterization
&lt;/h3&gt;

&lt;p&gt;We relied on simple Exponential Moving Averages (EMAs) to detect workload shifts. The approach is lightweight, easy to reason about, and performed surprisingly well during evaluation.&lt;/p&gt;

&lt;p&gt;However, collecting and maintaining workload statistics still consumes CPU cycles. As we move toward more realistic workloads, reducing the overhead of metric collection will become increasingly important.&lt;/p&gt;

&lt;p&gt;We are also interested in exploring more predictive workload characterization techniques, including probabilistic models that can anticipate workload shifts instead of simply reacting to them.&lt;/p&gt;

&lt;h3&gt;
  
  
  Parallel Compaction
&lt;/h3&gt;

&lt;p&gt;Amethyst currently executes compaction using a single background thread.&lt;/p&gt;

&lt;p&gt;This simplified implementation helped us isolate the behavior of the adaptation logic, but it also leaves significant performance on the table. Modern storage engines aggressively parallelize compaction work, and Amethyst should be no exception.&lt;/p&gt;

&lt;p&gt;Parallel compaction is one of the first major architectural improvements we plan to investigate in future versions.&lt;/p&gt;

&lt;h3&gt;
  
  
  Bloom Filters
&lt;/h3&gt;

&lt;p&gt;One deliberate omission from V1 was Bloom filters.&lt;/p&gt;

&lt;p&gt;We wanted to measure how much read amplification could be reduced through adaptive compaction alone, without relying on standard read optimizations to mask underlying inefficiencies.&lt;/p&gt;

&lt;p&gt;Now that the adaptive baseline has been established, the next step is integrating dynamic, Monkey-style Bloom filter allocation on a per-segment basis. This will allow us to further improve lookup performance while preserving the adaptive behavior of the system.&lt;/p&gt;

&lt;p&gt;Amethyst v1 was designed as a proof of concept rather than a production-ready database. The goal was to answer a simple question: can lightweight runtime workload estimation make adaptive compaction practical?&lt;/p&gt;

&lt;p&gt;The answer appears to be yes.&lt;/p&gt;

&lt;p&gt;The next challenge is turning that proof of concept into a more complete storage engine.&lt;/p&gt;

&lt;h2&gt;
  
  
  9. Open Questions
&lt;/h2&gt;

&lt;p&gt;While the initial results are encouraging, Amethyst is far from a finished system. In many ways, the most interesting questions only emerged after the prototype was built and evaluated.&lt;/p&gt;

&lt;h3&gt;
  
  
  Granularity
&lt;/h3&gt;

&lt;p&gt;One obvious question is granularity.&lt;/p&gt;

&lt;p&gt;Amethyst currently performs adaptation at the segment level, but there is no guarantee that this is the optimal unit of decision making. Could adaptation become even more fine-grained? Or would the additional complexity outweigh any potential gains?&lt;/p&gt;

&lt;h3&gt;
  
  
  Workload Characterization
&lt;/h3&gt;

&lt;p&gt;Another open problem is workload characterization.&lt;/p&gt;

&lt;p&gt;The current system relies on a relatively small set of runtime signals combined with lightweight smoothing techniques. This was a deliberate design choice to keep overhead low. However, it remains unclear which workload characteristics are truly the most predictive.&lt;/p&gt;

&lt;p&gt;Are there additional signals we should be collecting? Are there better ways to distinguish temporary workload fluctuations from genuine behavioral shifts?&lt;/p&gt;

&lt;h3&gt;
  
  
  Decision Making
&lt;/h3&gt;

&lt;p&gt;The decision engine itself also leaves considerable room for exploration.&lt;/p&gt;

&lt;p&gt;Amethyst currently relies on threshold-based policy transitions. While simple and effective, this approach may not always produce optimal decisions. More sophisticated probabilistic models, reinforcement learning approaches, or analytical cost models could potentially improve adaptation quality, though likely at the cost of additional complexity and runtime overhead.&lt;/p&gt;

&lt;h3&gt;
  
  
  Scaling the System
&lt;/h3&gt;

&lt;p&gt;There are also important systems questions that remain unanswered.&lt;/p&gt;

&lt;p&gt;How does adaptive compaction behave under significantly larger datasets? How does it interact with production features such as Bloom filters, background garbage collection, and parallel compaction threads? Do the benefits persist at scale, or do new bottlenecks emerge?&lt;/p&gt;

&lt;h3&gt;
  
  
  Beyond Compaction
&lt;/h3&gt;

&lt;p&gt;Finally, there is a broader question that motivated the project from the beginning.&lt;/p&gt;

&lt;p&gt;Most modern systems continuously adapt to changing conditions, yet storage engine compaction policies are still largely configured as static choices. If adaptation proves practical at the segment level, are there other storage engine decisions that should become adaptive as well?&lt;/p&gt;

&lt;p&gt;We do not yet have answers to these questions.&lt;/p&gt;

&lt;p&gt;However, they represent the directions we are most excited to explore as Amethyst continues to evolve.&lt;/p&gt;

&lt;h2&gt;
  
  
  10. Closing Thoughts
&lt;/h2&gt;

&lt;p&gt;Amethyst began with a simple question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Why do storage engines force us to choose a single compaction strategy?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The deeper we explored LSM trees, the more that choice seemed at odds with the reality of modern workloads. Systems adapt to changing conditions at nearly every layer of the stack, yet compaction policy often remains a static decision made long before the workload is fully understood.&lt;/p&gt;

&lt;p&gt;This project was our attempt to challenge that assumption.&lt;/p&gt;

&lt;p&gt;The result was &lt;strong&gt;Amethyst&lt;/strong&gt;: a lightweight adaptive LSM-tree prototype capable of selecting compaction behavior at the segment level using runtime workload signals. While the system is still far from production-ready, the results demonstrate that adaptive compaction can be practical, lightweight, and effective under changing workload conditions.&lt;/p&gt;

&lt;p&gt;More importantly, the project reinforced a lesson that applies far beyond storage engines. Many systems problems are not constrained by a lack of solutions, but by assumptions that have gone unquestioned for years. Sometimes progress comes not from inventing something entirely new, but from revisiting an old tradeoff and asking whether it is still necessary.&lt;/p&gt;

&lt;p&gt;We do not believe Amethyst is the final answer to adaptive compaction. However, we hope it contributes to the broader discussion around adaptive storage systems and encourages others to continue exploring the space.&lt;/p&gt;

&lt;p&gt;If you would like to learn more, the paper, source code, and presentation materials are linked below. We would be happy to hear your thoughts, questions, criticisms, or ideas for future directions.&lt;/p&gt;




&lt;h2&gt;
  
  
  Resources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Paper:&lt;/strong&gt; &lt;a href="https://dl.acm.org/doi/10.1145/3802514.3812603" rel="noopener noreferrer"&gt;Adaptive Compaction in LSM-Based Storage Systems&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Code:&lt;/strong&gt; &lt;a href="https://github.com/amethystdb/amethyst_1.0" rel="noopener noreferrer"&gt;Amethyst GitHub Repository&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Contact
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Suchitra Shankar:&lt;/strong&gt; &lt;a href="mailto:suchitrashankar07@gmail.com"&gt;suchitrashankar07@gmail.com&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Nilin Rose:&lt;/strong&gt; &lt;a href="mailto:nilinrose1@gmail.com"&gt;nilinrose1@gmail.com&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>database</category>
      <category>buildinpublic</category>
      <category>programming</category>
      <category>showdev</category>
    </item>
  </channel>
</rss>
