<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Utsab Ghoshal</title>
    <description>The latest articles on DEV Community by Utsab Ghoshal (@ugtech).</description>
    <link>https://dev.to/ugtech</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4143965%2F1cc964c0-7283-4e97-9817-8fba46708f57.png</url>
      <title>DEV Community: Utsab Ghoshal</title>
      <link>https://dev.to/ugtech</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ugtech"/>
    <language>en</language>
    <item>
      <title>UCRF: Exploring Version Provenance for Concurrency Control and Selective Recovery</title>
      <dc:creator>Utsab Ghoshal</dc:creator>
      <pubDate>Mon, 28 Sep 2026 08:08:39 +0000</pubDate>
      <link>https://dev.to/ugtech/ucrf-exploring-version-provenance-for-concurrency-control-and-selective-recovery-20jm</link>
      <guid>https://dev.to/ugtech/ucrf-exploring-version-provenance-for-concurrency-control-and-selective-recovery-20jm</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Research status:&lt;/strong&gt; Ongoing&lt;br&gt;
&lt;strong&gt;Current milestone:&lt;/strong&gt; v0.39&lt;br&gt;
&lt;strong&gt;Type:&lt;/strong&gt; Experimental research prototype&lt;br&gt;
&lt;strong&gt;Repository:&lt;/strong&gt; &lt;a href="https://github.com/UtsabGhoshal/UCRF-research-prototype" rel="noopener noreferrer"&gt;UCRF-research-prototype on GitHub&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;What happens when a transaction fails in the middle of a highly concurrent database workload?&lt;/p&gt;

&lt;p&gt;The obvious answer is often: recover the transaction, replay the affected region, or roll back a larger dependency set.&lt;/p&gt;

&lt;p&gt;But there is a deeper question:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can the database know exactly which operations or versions actually caused a failure, and recover only what is causally necessary?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That question led me to build &lt;strong&gt;UCRF — Unified Concurrency and Recovery Framework&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;UCRF is an ongoing research project exploring whether &lt;strong&gt;version-level provenance&lt;/strong&gt; can act as a shared dependency representation for two problems that are often considered separately:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Concurrency control and serialization validation&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Selective recovery after failures&lt;/strong&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This is not a claim that UCRF is already a superior replacement for existing database systems.&lt;/p&gt;

&lt;p&gt;It is an investigation.&lt;/p&gt;

&lt;p&gt;And some of the most interesting results so far came from experiments that &lt;strong&gt;did not produce the novelty I initially expected&lt;/strong&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  1. Why I Started Exploring This Problem
&lt;/h2&gt;

&lt;p&gt;Concurrency control and recovery are usually discussed as separate database subsystems.&lt;/p&gt;

&lt;p&gt;Concurrency control asks:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Can these transactions safely commit while preserving the required consistency guarantees?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Recovery asks:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;If something fails, what needs to be undone, replayed, or reconstructed?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Both problems, however, depend on relationships between operations.&lt;/p&gt;

&lt;p&gt;Consider:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Transaction A
    |
    | writes X
    v
Transaction B
    |
    | reads X
    v
Transaction C

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If &lt;code&gt;B&lt;/code&gt; depends on a version produced by &lt;code&gt;A&lt;/code&gt;, and &lt;code&gt;C&lt;/code&gt; subsequently depends on &lt;code&gt;B&lt;/code&gt;, then a failure affecting &lt;code&gt;A&lt;/code&gt; may have consequences beyond &lt;code&gt;A&lt;/code&gt; itself.&lt;/p&gt;

&lt;p&gt;That made me wonder:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Could the same dependency information used during concurrency validation also help determine the minimum recovery set?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That became the starting point for UCRF.&lt;/p&gt;




&lt;h1&gt;
  
  
  2. UCRF Is Still Ongoing Research
&lt;/h1&gt;

&lt;p&gt;Before going further, there is an important clarification.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;UCRF is not a finished database engine.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;It is an experimental reference implementation used to investigate ideas around:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;serialization validation&lt;/li&gt;
&lt;li&gt;dependency tracking&lt;/li&gt;
&lt;li&gt;MVCC&lt;/li&gt;
&lt;li&gt;version provenance&lt;/li&gt;
&lt;li&gt;selective recovery&lt;/li&gt;
&lt;li&gt;WAL and checkpoints&lt;/li&gt;
&lt;li&gt;recovery frontiers&lt;/li&gt;
&lt;li&gt;dependency-aware replay&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The implementation has evolved through many experimental milestones, from early transaction-level dependency graphs to the current v0.39 provenance-based architecture.&lt;/p&gt;

&lt;p&gt;The repository is therefore part of the research process itself.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/UtsabGhoshal/UCRF-research-prototype" rel="noopener noreferrer"&gt;Explore the complete UCRF research repository → GitHub&lt;/a&gt;&lt;/p&gt;




&lt;h1&gt;
  
  
  3. The Evolution of UCRF
&lt;/h1&gt;

&lt;p&gt;UCRF did not begin with version provenance.&lt;/p&gt;

&lt;p&gt;It evolved through a sequence of increasingly difficult experiments.&lt;/p&gt;

&lt;p&gt;A simplified view of that evolution is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;v0.16
  ↓
Transaction-level serialization validation
  ↓
v0.17–v0.18
  ↓
MVCC + predicates + ranges
  ↓
v0.19–v0.25
  ↓
Selective / dependency-aware recovery
  ↓
v0.26–v0.34
  ↓
Stronger serialization + WAL + checkpoints
  ↓
v0.35
  ↓
Baseline comparisons
  ↓
v0.36–v0.37
  ↓
Recovery-frontier investigation
  ↓
v0.38
  ↓
Version-level provenance recovery
  ↓
v0.39
  ↓
Version provenance integrated with
serialization validation + recovery

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each stage answered one question while usually exposing another.&lt;/p&gt;

&lt;p&gt;That was actually one of the most useful parts of the project.&lt;/p&gt;




&lt;h1&gt;
  
  
  4. First Problem: Transaction-Level Recovery Was Too Coarse
&lt;/h1&gt;

&lt;p&gt;The initial recovery model operated primarily at the transaction/dependency level.&lt;/p&gt;

&lt;p&gt;Suppose we have:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;T1 → T2 → T3

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If &lt;code&gt;T1&lt;/code&gt; fails, a conservative recovery mechanism may need to consider:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;T1
T2
T3

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;But that does not necessarily mean every operation performed by those transactions needs to be replayed.&lt;/p&gt;

&lt;p&gt;A transaction may contain dozens or hundreds of operations.&lt;/p&gt;

&lt;p&gt;Only a small subset may actually participate in the causal chain leading to the affected state.&lt;/p&gt;

&lt;p&gt;That raised the first major research question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Can recovery operate at a finer granularity than transactions?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  5. The Recovery Frontier Was Not Enough
&lt;/h1&gt;

&lt;p&gt;The next idea was to identify a &lt;strong&gt;recovery frontier&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Instead of replaying everything reachable through a dependency graph, perhaps we could stop at a durable checkpoint or previously established recovery boundary.&lt;/p&gt;

&lt;p&gt;This produced encouraging reductions in logical recovery work.&lt;/p&gt;

&lt;p&gt;But then came an important research lesson:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;An optimization that looks novel is not necessarily novel.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;In v0.35 and v0.36, I compared UCRF-style recovery against stronger dependency-closure baselines.&lt;/p&gt;

&lt;p&gt;The experiments showed that generic causal dependency closure could already reproduce much of the apparent benefit.&lt;/p&gt;

&lt;p&gt;That led to an explicit novelty test in v0.37.&lt;/p&gt;




&lt;h1&gt;
  
  
  6. A Negative Result That Changed the Direction
&lt;/h1&gt;

&lt;p&gt;The v0.37 experiment compared:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;UCRF frontier-aware recovery&lt;/li&gt;
&lt;li&gt;a strong checkpoint-bounded dependency-closure baseline&lt;/li&gt;
&lt;li&gt;an independent minimality oracle&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The evaluation included:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;5,000 randomized DAG/recovery cases&lt;/li&gt;
&lt;li&gt;exhaustive small DAGs up to 5 vertices&lt;/li&gt;
&lt;li&gt;1,098 graphs&lt;/li&gt;
&lt;li&gt;27,362 recovery cases&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The result:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;UCRF frontier-aware recovery
            =
checkpoint-bounded closure
            =
independent oracle

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;under the tested graph model.&lt;/p&gt;

&lt;p&gt;There were:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;0 oracle mismatches&lt;/li&gt;
&lt;li&gt;0 unsafe pruning cases&lt;/li&gt;
&lt;li&gt;0 non-minimal recovery cases&lt;/li&gt;
&lt;li&gt;0 strict improvements over the strong baseline&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That was not the result I was hoping for.&lt;/p&gt;

&lt;p&gt;But it was useful.&lt;/p&gt;

&lt;p&gt;It showed that I should not claim novelty simply because an optimization produced fewer replay operations.&lt;/p&gt;

&lt;p&gt;Instead, I needed to look for a more fundamental representation.&lt;/p&gt;




&lt;h1&gt;
  
  
  7. Moving From Transactions to Versions
&lt;/h1&gt;

&lt;p&gt;That led to the next question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What if the dependency unit is not a transaction, but a version?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Modern MVCC systems already reason in terms of versions.&lt;/p&gt;

&lt;p&gt;Instead of simply saying:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;T1 writes X
T2 reads X

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;we can represent the relationship more precisely:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;T1 / operation 0
        |
        | produces
        v
      X:v1
        |
        | consumed by
        v
T2 / operation 1

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The dependency is now tied to the specific version.&lt;/p&gt;

&lt;p&gt;This is the basis of the provenance model explored in UCRF.&lt;/p&gt;




&lt;h1&gt;
  
  
  8. The UCRF Provenance Graph
&lt;/h1&gt;

&lt;p&gt;The current architecture can be represented conceptually as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                Accepted Operations
                       |
                       v
             Version Provenance Graph
                    /       \
                   /         \
                  v           v
        Serialization       Recovery
         Validation         Analysis

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A simplified provenance chain might look like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A1 writes X:v1
        |
        v
B1 reads X:v1
        |
        v
C1 consumes information derived from B1

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important idea is that the dependency is represented through &lt;strong&gt;versions and their producers/consumers&lt;/strong&gt;, rather than only through transaction-level edges.&lt;/p&gt;

&lt;p&gt;This provides a common representation that can potentially be queried for different purposes.&lt;/p&gt;




&lt;h1&gt;
  
  
  9. Using the Same Provenance for Concurrency Control
&lt;/h1&gt;

&lt;p&gt;This is where v0.39 becomes particularly interesting.&lt;/p&gt;

&lt;p&gt;Earlier versions maintained several mechanisms:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;read/write conflicts&lt;/li&gt;
&lt;li&gt;predicate conflicts&lt;/li&gt;
&lt;li&gt;range conflicts&lt;/li&gt;
&lt;li&gt;serialization edges&lt;/li&gt;
&lt;li&gt;dependency relationships&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Version provenance is now integrated directly into the serialization/commit-validation path.&lt;/p&gt;

&lt;p&gt;The conceptual flow is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Operation
    |
    v
Version produced / consumed
    |
    v
Provenance relationship
    |
    +------&amp;gt; Serialization validation
    |
    +------&amp;gt; Recovery analysis

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When exact provenance is available, it can provide a precise relationship.&lt;/p&gt;

&lt;p&gt;When exact provenance is unavailable, the older conflict/range/predicate mechanisms can act as conservative fallback mechanisms.&lt;/p&gt;

&lt;p&gt;This was necessary because a real system cannot simply assume that perfect provenance will always be available.&lt;/p&gt;




&lt;h1&gt;
  
  
  10. Why MVCC Matters
&lt;/h1&gt;

&lt;p&gt;Version provenance naturally fits an MVCC-style model.&lt;/p&gt;

&lt;p&gt;Instead of treating a data item as simply:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;X

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;we can reason about:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;X:v1
X:v2
X:v3

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Different transactions may observe different versions.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;X:v1
  |
  +----&amp;gt; T1 reads
  |
  +----&amp;gt; T2 reads

X:v2
  |
  +----&amp;gt; T3 reads

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This gives the recovery system more information about what was actually observed.&lt;/p&gt;

&lt;p&gt;It also provides a much finer-grained dependency representation than simply saying:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;T1 depends on T2

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The actual question becomes:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Which version did the operation observe, who produced that version, and what later state depended on it?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  11. Using Provenance for Recovery
&lt;/h1&gt;

&lt;p&gt;Suppose a version is invalidated or a producing operation must be recovered.&lt;/p&gt;

&lt;p&gt;A transaction-level recovery mechanism might begin with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;T1
 ↓
T2
 ↓
T3
 ↓
T4

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A provenance-aware mechanism can instead reason about the specific chain:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;v1
 ↓
operation B1
 ↓
v2
 ↓
operation C4

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This can potentially reduce unnecessary recovery work.&lt;/p&gt;

&lt;p&gt;The important distinction is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The goal is not simply to recover fewer transactions. The goal is to identify the smallest causally necessary set of operations or versions under the model.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That distinction matters when evaluating whether a recovery optimization is actually meaningful.&lt;/p&gt;




&lt;h1&gt;
  
  
  12. Testing the Idea
&lt;/h1&gt;

&lt;p&gt;The project has been developed through increasingly adversarial testing rather than a single benchmark.&lt;/p&gt;

&lt;p&gt;The current v0.39 evaluation includes a randomized set of &lt;strong&gt;5,000 histories&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;For those tested histories, the current implementation reported:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;5,000 randomized histories

0 accepted non-serializable histories
0 serialization-oracle mismatches
0 provenance-state consistency failures

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;An adversarial mutual-dependency cycle was also rejected.&lt;/p&gt;

&lt;p&gt;The repository additionally contains the cumulative v0.39 development test artifact reporting &lt;strong&gt;66 passing tests&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;These numbers should be interpreted carefully.&lt;/p&gt;

&lt;p&gt;They demonstrate behavior of the current &lt;strong&gt;reference model and tested workloads&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;They do not prove correctness for arbitrary SQL workloads, production database engines, or every possible concurrency pattern.&lt;/p&gt;




&lt;h1&gt;
  
  
  13. What Went Wrong During Development
&lt;/h1&gt;

&lt;p&gt;One of the most interesting parts of this project has been debugging the model itself.&lt;/p&gt;

&lt;p&gt;Several problems appeared while integrating provenance into the system.&lt;/p&gt;

&lt;h3&gt;
  
  
  Forward provenance references
&lt;/h3&gt;

&lt;p&gt;Some operations initially attempted to reference versions that had not yet been registered.&lt;/p&gt;

&lt;p&gt;The solution was a two-phase graph construction process.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Phase 1
Register operations and versions

        ↓

Phase 2
Resolve provenance relationships

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Incremental recovery reconstruction
&lt;/h3&gt;

&lt;p&gt;Recovery could reconstruct part of the dependency graph without synchronizing all registered operations.&lt;/p&gt;

&lt;p&gt;The solution was to explicitly synchronize provenance state during recovery reconstruction.&lt;/p&gt;

&lt;h3&gt;
  
  
  Reused logical write IDs
&lt;/h3&gt;

&lt;p&gt;Different transactions could accidentally produce logically identical version identifiers.&lt;/p&gt;

&lt;p&gt;That created ambiguity.&lt;/p&gt;

&lt;p&gt;The solution was to make version identities producer-qualified, for example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;T3-o0|k2

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Missing exact provenance
&lt;/h3&gt;

&lt;p&gt;Not every dependency can necessarily be represented through exact read-from provenance.&lt;/p&gt;

&lt;p&gt;Instead of assuming perfect information, UCRF retains conservative conflict/range/predicate mechanisms as fallback.&lt;/p&gt;

&lt;h3&gt;
  
  
  Metadata overhead
&lt;/h3&gt;

&lt;p&gt;This remains one of the biggest unresolved questions.&lt;/p&gt;

&lt;p&gt;Tracking more provenance means storing more metadata.&lt;/p&gt;

&lt;p&gt;A theoretical reduction in recovery work is not automatically useful if the metadata required to achieve it becomes too expensive.&lt;/p&gt;




&lt;h1&gt;
  
  
  14. What I Learned From the Negative Results
&lt;/h1&gt;

&lt;p&gt;The most important lesson from this project is probably this:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;A result that does not support your hypothesis is still a research result.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Earlier in the project, it would have been easy to say:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“UCRF can perform selective recovery more efficiently.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;But the stronger experiments forced a more precise conclusion.&lt;/p&gt;

&lt;p&gt;Some of the recovery improvements could already be reproduced by conventional dependency closure.&lt;/p&gt;

&lt;p&gt;That eliminated one possible novelty claim.&lt;/p&gt;

&lt;p&gt;Instead of hiding that result, I kept it.&lt;/p&gt;

&lt;p&gt;It changed the direction of the project.&lt;/p&gt;

&lt;p&gt;That eventually led toward version provenance.&lt;/p&gt;




&lt;h1&gt;
  
  
  15. Relationship to Existing Techniques
&lt;/h1&gt;

&lt;p&gt;UCRF does not exist in isolation.&lt;/p&gt;

&lt;p&gt;There are already well-established techniques involving:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;two-phase locking&lt;/li&gt;
&lt;li&gt;optimistic concurrency control&lt;/li&gt;
&lt;li&gt;MVCC&lt;/li&gt;
&lt;li&gt;snapshot isolation&lt;/li&gt;
&lt;li&gt;serializable snapshot isolation&lt;/li&gt;
&lt;li&gt;dependency-based serializability certification&lt;/li&gt;
&lt;li&gt;serialization graphs&lt;/li&gt;
&lt;li&gt;write-ahead logging&lt;/li&gt;
&lt;li&gt;checkpoints&lt;/li&gt;
&lt;li&gt;dependency-aware recovery&lt;/li&gt;
&lt;li&gt;version provenance&lt;/li&gt;
&lt;li&gt;speculative execution and recovery&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In particular, dependency-based serialization certification and version/dependency-aware recovery are established research areas.&lt;/p&gt;

&lt;p&gt;Therefore, the current research question is deliberately narrower:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Can one version-provenance representation be practically useful as a shared basis for both serialization certification and operation-level causal recovery, while providing a favorable metadata/validation/recovery trade-off?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is a question I am still investigating.&lt;/p&gt;




&lt;h1&gt;
  
  
  16. What Could Be Novel?
&lt;/h1&gt;

&lt;p&gt;At this stage, I do &lt;strong&gt;not&lt;/strong&gt; consider UCRF's novelty proven.&lt;/p&gt;

&lt;p&gt;That is an important distinction.&lt;/p&gt;

&lt;p&gt;The potentially interesting part is not simply:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“I used a dependency graph.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Dependency graphs are obviously not new.&lt;/p&gt;

&lt;p&gt;Nor is it:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“I performed selective recovery.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That also has substantial prior art.&lt;/p&gt;

&lt;p&gt;Nor is it:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“I used version provenance.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Version-level provenance and dependency tracking have already been explored in database research.&lt;/p&gt;

&lt;p&gt;The more specific research direction is the possibility of using a &lt;strong&gt;shared version-provenance representation across both concurrency certification and selective recovery&lt;/strong&gt;, while measuring the resulting trade-offs.&lt;/p&gt;

&lt;p&gt;Whether that constitutes a genuinely new contribution remains an open research question.&lt;/p&gt;

&lt;p&gt;A serious evaluation would require stronger comparisons against appropriate existing techniques and implementations.&lt;/p&gt;




&lt;h1&gt;
  
  
  17. What UCRF Has Demonstrated So Far
&lt;/h1&gt;

&lt;p&gt;The project has nevertheless produced several useful findings.&lt;/p&gt;

&lt;h3&gt;
  
  
  Serialization validation
&lt;/h3&gt;

&lt;p&gt;The reference model has been tested against independent serializability oracles across randomized workloads.&lt;/p&gt;

&lt;h3&gt;
  
  
  MVCC modeling
&lt;/h3&gt;

&lt;p&gt;The project contains an executable model for version chains, snapshots, visibility, and version publication.&lt;/p&gt;

&lt;h3&gt;
  
  
  Recovery locality
&lt;/h3&gt;

&lt;p&gt;Several experiments demonstrated that selective dependency-based recovery can be substantially smaller than replaying an entire affected region under the tested logical workloads.&lt;/p&gt;

&lt;h3&gt;
  
  
  WAL and checkpoints
&lt;/h3&gt;

&lt;p&gt;The project evolved to include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;prepare/commit/abort logging&lt;/li&gt;
&lt;li&gt;durable-prefix modeling&lt;/li&gt;
&lt;li&gt;checkpoints&lt;/li&gt;
&lt;li&gt;selective replay&lt;/li&gt;
&lt;li&gt;WAL compaction&lt;/li&gt;
&lt;li&gt;checksummed logical WAL records&lt;/li&gt;
&lt;li&gt;corruption and torn-tail handling&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Negative research results
&lt;/h3&gt;

&lt;p&gt;Perhaps most importantly, the project has explicitly tested and eliminated at least one apparent source of novelty.&lt;/p&gt;

&lt;p&gt;That makes the later hypotheses more focused.&lt;/p&gt;




&lt;h1&gt;
  
  
  18. Current Limitations
&lt;/h1&gt;

&lt;p&gt;There are significant limitations.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. It is a reference model
&lt;/h3&gt;

&lt;p&gt;UCRF is not currently a production database engine.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Metadata overhead
&lt;/h3&gt;

&lt;p&gt;Version provenance requires additional metadata.&lt;/p&gt;

&lt;p&gt;The cost of capturing, persisting, indexing, and garbage-collecting that metadata remains unresolved.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Logical recovery ≠ wall-clock performance
&lt;/h3&gt;

&lt;p&gt;A reduction in the number of replayed logical operations does not automatically mean lower real-world recovery latency.&lt;/p&gt;

&lt;p&gt;Disk I/O, cache behavior, synchronization, logging, CPU overhead, and metadata maintenance all matter.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Simplified workload model
&lt;/h3&gt;

&lt;p&gt;The current implementation does not represent the full complexity of arbitrary SQL workloads.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Predicate and range semantics
&lt;/h3&gt;

&lt;p&gt;Predicate/range tracking remains a model rather than a complete implementation of every database indexing and predicate-evaluation behavior.&lt;/p&gt;

&lt;h3&gt;
  
  
  6. Physical storage integration
&lt;/h3&gt;

&lt;p&gt;The WAL and checkpoint mechanisms are reference abstractions rather than a complete filesystem-crash-safe storage engine.&lt;/p&gt;

&lt;h3&gt;
  
  
  7. Stronger baselines are still required
&lt;/h3&gt;

&lt;p&gt;A meaningful performance study would need comparisons against appropriate existing implementations rather than only Python-level reference models.&lt;/p&gt;




&lt;h1&gt;
  
  
  19. What I Want to Investigate Next
&lt;/h1&gt;

&lt;p&gt;The research is not finished.&lt;/p&gt;

&lt;p&gt;The next stages I am interested in include:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Version provenance
       ↓
Metadata cost model
       ↓
Stronger concurrency-control baselines
       ↓
Stronger recovery baselines
       ↓
Larger workload models
       ↓
Physical storage considerations
       ↓
Performance evaluation

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Some of the questions I still want to answer are:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;How much metadata does provenance actually require?&lt;/li&gt;
&lt;li&gt;Can provenance be compressed?&lt;/li&gt;
&lt;li&gt;Can old provenance safely be garbage-collected?&lt;/li&gt;
&lt;li&gt;How does the approach compare with SSI/SSN-style certification?&lt;/li&gt;
&lt;li&gt;What happens under larger dependency graphs?&lt;/li&gt;
&lt;li&gt;What happens with complex predicates and indexes?&lt;/li&gt;
&lt;li&gt;Can provenance tracking remain efficient under high concurrency?&lt;/li&gt;
&lt;li&gt;Does logical recovery reduction translate into actual recovery-latency reduction?&lt;/li&gt;
&lt;li&gt;Can the model be integrated into a real database engine?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These questions are more important to me now than simply reaching a higher version number.&lt;/p&gt;




&lt;h1&gt;
  
  
  20. Why I Am Publishing This While It Is Still Ongoing
&lt;/h1&gt;

&lt;p&gt;I initially thought I should wait until the research was “finished.”&lt;/p&gt;

&lt;p&gt;But database research is rarely a straight line from:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Idea → Implementation → Proof → Publication

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The actual process looks more like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Idea
 ↓
Prototype
 ↓
Experiment
 ↓
Unexpected result
 ↓
New hypothesis
 ↓
Stronger baseline
 ↓
Negative result
 ↓
Refined model
 ↓
New experiment
 ↓
Repeat

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is exactly what happened with UCRF.&lt;/p&gt;

&lt;p&gt;Publishing the project while it is still evolving also makes the development history visible.&lt;/p&gt;

&lt;p&gt;The repository contains the implementation, experiments, research notes, benchmarks, and milestone history.&lt;/p&gt;

&lt;p&gt;The article explains the reasoning behind that evolution.&lt;/p&gt;




&lt;h1&gt;
  
  
  21. The Research Repository
&lt;/h1&gt;

&lt;p&gt;The complete implementation and research artifacts are available here:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/UtsabGhoshal/UCRF-research-prototype" rel="noopener noreferrer"&gt;UCRF-research-prototype — GitHub&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The repository is the authoritative place for the current implementation and experimental artifacts.&lt;/p&gt;

&lt;p&gt;The current public milestone is &lt;strong&gt;v0.39&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;If the research changes, the repository will change with it.&lt;/p&gt;

&lt;p&gt;That is intentional.&lt;/p&gt;




&lt;h1&gt;
  
  
  22. Conclusion
&lt;/h1&gt;

&lt;p&gt;UCRF started with a relatively simple question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Can recovery be more selective?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That question led through transaction-level dependency graphs, MVCC, ranges, WAL, checkpoints, dependency closure, recovery frontiers, and eventually version provenance.&lt;/p&gt;

&lt;p&gt;Along the way, some ideas looked promising but did not survive stronger comparison.&lt;/p&gt;

&lt;p&gt;That was useful.&lt;/p&gt;

&lt;p&gt;The current direction is therefore more specific:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Can version-level provenance provide a shared representation for both concurrency-control validation and selective causal recovery?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I don't have a final answer yet.&lt;/p&gt;

&lt;p&gt;And that is precisely why this is still a research project.&lt;/p&gt;

&lt;p&gt;For now, UCRF is an experimental framework for exploring the question, documenting the results, testing against stronger baselines, and discovering where the idea holds up — and where it does not.&lt;/p&gt;

&lt;p&gt;The research continues.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/UtsabGhoshal/UCRF-research-prototype" rel="noopener noreferrer"&gt;Follow the implementation and experiments on GitHub → UCRF-research-prototype&lt;/a&gt;&lt;/p&gt;

</description>
      <category>database</category>
      <category>algorithms</category>
      <category>computerscience</category>
      <category>ai</category>
    </item>
  </channel>
</rss>
