DEV Community

Cover image for UCRF: Exploring Version Provenance for Concurrency Control and Selective Recovery
Utsab Ghoshal
Utsab Ghoshal

Posted on Fully Autonomous

UCRF: Exploring Version Provenance for Concurrency Control and Selective Recovery

Research status: Ongoing
Current milestone: v0.39
Type: Experimental research prototype
Repository: UCRF-research-prototype on GitHub

What happens when a transaction fails in the middle of a highly concurrent database workload?

The obvious answer is often: recover the transaction, replay the affected region, or roll back a larger dependency set.

But there is a deeper question:

Can the database know exactly which operations or versions actually caused a failure, and recover only what is causally necessary?

That question led me to build UCRF — Unified Concurrency and Recovery Framework.

UCRF is an ongoing research project exploring whether version-level provenance can act as a shared dependency representation for two problems that are often considered separately:

  1. Concurrency control and serialization validation
  2. Selective recovery after failures

This is not a claim that UCRF is already a superior replacement for existing database systems.

It is an investigation.

And some of the most interesting results so far came from experiments that did not produce the novelty I initially expected.


1. Why I Started Exploring This Problem

Concurrency control and recovery are usually discussed as separate database subsystems.

Concurrency control asks:

Can these transactions safely commit while preserving the required consistency guarantees?

Recovery asks:

If something fails, what needs to be undone, replayed, or reconstructed?

Both problems, however, depend on relationships between operations.

Consider:

Transaction A
    |
    | writes X
    v
Transaction B
    |
    | reads X
    v
Transaction C

Enter fullscreen mode Exit fullscreen mode

If B depends on a version produced by A, and C subsequently depends on B, then a failure affecting A may have consequences beyond A itself.

That made me wonder:

Could the same dependency information used during concurrency validation also help determine the minimum recovery set?

That became the starting point for UCRF.


2. UCRF Is Still Ongoing Research

Before going further, there is an important clarification.

UCRF is not a finished database engine.

It is an experimental reference implementation used to investigate ideas around:

  • serialization validation
  • dependency tracking
  • MVCC
  • version provenance
  • selective recovery
  • WAL and checkpoints
  • recovery frontiers
  • dependency-aware replay

The implementation has evolved through many experimental milestones, from early transaction-level dependency graphs to the current v0.39 provenance-based architecture.

The repository is therefore part of the research process itself.

Explore the complete UCRF research repository → GitHub


3. The Evolution of UCRF

UCRF did not begin with version provenance.

It evolved through a sequence of increasingly difficult experiments.

A simplified view of that evolution is:

v0.16
  ↓
Transaction-level serialization validation
  ↓
v0.17–v0.18
  ↓
MVCC + predicates + ranges
  ↓
v0.19–v0.25
  ↓
Selective / dependency-aware recovery
  ↓
v0.26–v0.34
  ↓
Stronger serialization + WAL + checkpoints
  ↓
v0.35
  ↓
Baseline comparisons
  ↓
v0.36–v0.37
  ↓
Recovery-frontier investigation
  ↓
v0.38
  ↓
Version-level provenance recovery
  ↓
v0.39
  ↓
Version provenance integrated with
serialization validation + recovery

Enter fullscreen mode Exit fullscreen mode

Each stage answered one question while usually exposing another.

That was actually one of the most useful parts of the project.


4. First Problem: Transaction-Level Recovery Was Too Coarse

The initial recovery model operated primarily at the transaction/dependency level.

Suppose we have:

T1 → T2 → T3

Enter fullscreen mode Exit fullscreen mode

If T1 fails, a conservative recovery mechanism may need to consider:

T1
T2
T3

Enter fullscreen mode Exit fullscreen mode

But that does not necessarily mean every operation performed by those transactions needs to be replayed.

A transaction may contain dozens or hundreds of operations.

Only a small subset may actually participate in the causal chain leading to the affected state.

That raised the first major research question:

Can recovery operate at a finer granularity than transactions?


5. The Recovery Frontier Was Not Enough

The next idea was to identify a recovery frontier.

Instead of replaying everything reachable through a dependency graph, perhaps we could stop at a durable checkpoint or previously established recovery boundary.

This produced encouraging reductions in logical recovery work.

But then came an important research lesson:

An optimization that looks novel is not necessarily novel.

In v0.35 and v0.36, I compared UCRF-style recovery against stronger dependency-closure baselines.

The experiments showed that generic causal dependency closure could already reproduce much of the apparent benefit.

That led to an explicit novelty test in v0.37.


6. A Negative Result That Changed the Direction

The v0.37 experiment compared:

  • UCRF frontier-aware recovery
  • a strong checkpoint-bounded dependency-closure baseline
  • an independent minimality oracle

The evaluation included:

  • 5,000 randomized DAG/recovery cases
  • exhaustive small DAGs up to 5 vertices
  • 1,098 graphs
  • 27,362 recovery cases

The result:

UCRF frontier-aware recovery
            =
checkpoint-bounded closure
            =
independent oracle

Enter fullscreen mode Exit fullscreen mode

under the tested graph model.

There were:

  • 0 oracle mismatches
  • 0 unsafe pruning cases
  • 0 non-minimal recovery cases
  • 0 strict improvements over the strong baseline

That was not the result I was hoping for.

But it was useful.

It showed that I should not claim novelty simply because an optimization produced fewer replay operations.

Instead, I needed to look for a more fundamental representation.


7. Moving From Transactions to Versions

That led to the next question:

What if the dependency unit is not a transaction, but a version?

Modern MVCC systems already reason in terms of versions.

Instead of simply saying:

T1 writes X
T2 reads X

Enter fullscreen mode Exit fullscreen mode

we can represent the relationship more precisely:

T1 / operation 0
        |
        | produces
        v
      X:v1
        |
        | consumed by
        v
T2 / operation 1

Enter fullscreen mode Exit fullscreen mode

The dependency is now tied to the specific version.

This is the basis of the provenance model explored in UCRF.


8. The UCRF Provenance Graph

The current architecture can be represented conceptually as:

                Accepted Operations
                       |
                       v
             Version Provenance Graph
                    /       \
                   /         \
                  v           v
        Serialization       Recovery
         Validation         Analysis

Enter fullscreen mode Exit fullscreen mode

A simplified provenance chain might look like:

A1 writes X:v1
        |
        v
B1 reads X:v1
        |
        v
C1 consumes information derived from B1

Enter fullscreen mode Exit fullscreen mode

The important idea is that the dependency is represented through versions and their producers/consumers, rather than only through transaction-level edges.

This provides a common representation that can potentially be queried for different purposes.


9. Using the Same Provenance for Concurrency Control

This is where v0.39 becomes particularly interesting.

Earlier versions maintained several mechanisms:

  • read/write conflicts
  • predicate conflicts
  • range conflicts
  • serialization edges
  • dependency relationships

Version provenance is now integrated directly into the serialization/commit-validation path.

The conceptual flow is:

Operation
    |
    v
Version produced / consumed
    |
    v
Provenance relationship
    |
    +------> Serialization validation
    |
    +------> Recovery analysis

Enter fullscreen mode Exit fullscreen mode

When exact provenance is available, it can provide a precise relationship.

When exact provenance is unavailable, the older conflict/range/predicate mechanisms can act as conservative fallback mechanisms.

This was necessary because a real system cannot simply assume that perfect provenance will always be available.


10. Why MVCC Matters

Version provenance naturally fits an MVCC-style model.

Instead of treating a data item as simply:

X

Enter fullscreen mode Exit fullscreen mode

we can reason about:

X:v1
X:v2
X:v3

Enter fullscreen mode Exit fullscreen mode

Different transactions may observe different versions.

For example:

X:v1
  |
  +----> T1 reads
  |
  +----> T2 reads

X:v2
  |
  +----> T3 reads

Enter fullscreen mode Exit fullscreen mode

This gives the recovery system more information about what was actually observed.

It also provides a much finer-grained dependency representation than simply saying:

T1 depends on T2

Enter fullscreen mode Exit fullscreen mode

The actual question becomes:

Which version did the operation observe, who produced that version, and what later state depended on it?


11. Using Provenance for Recovery

Suppose a version is invalidated or a producing operation must be recovered.

A transaction-level recovery mechanism might begin with:

T1
 ↓
T2
 ↓
T3
 ↓
T4

Enter fullscreen mode Exit fullscreen mode

A provenance-aware mechanism can instead reason about the specific chain:

v1
 ↓
operation B1
 ↓
v2
 ↓
operation C4

Enter fullscreen mode Exit fullscreen mode

This can potentially reduce unnecessary recovery work.

The important distinction is:

The goal is not simply to recover fewer transactions. The goal is to identify the smallest causally necessary set of operations or versions under the model.

That distinction matters when evaluating whether a recovery optimization is actually meaningful.


12. Testing the Idea

The project has been developed through increasingly adversarial testing rather than a single benchmark.

The current v0.39 evaluation includes a randomized set of 5,000 histories.

For those tested histories, the current implementation reported:

5,000 randomized histories

0 accepted non-serializable histories
0 serialization-oracle mismatches
0 provenance-state consistency failures

Enter fullscreen mode Exit fullscreen mode

An adversarial mutual-dependency cycle was also rejected.

The repository additionally contains the cumulative v0.39 development test artifact reporting 66 passing tests.

These numbers should be interpreted carefully.

They demonstrate behavior of the current reference model and tested workloads.

They do not prove correctness for arbitrary SQL workloads, production database engines, or every possible concurrency pattern.


13. What Went Wrong During Development

One of the most interesting parts of this project has been debugging the model itself.

Several problems appeared while integrating provenance into the system.

Forward provenance references

Some operations initially attempted to reference versions that had not yet been registered.

The solution was a two-phase graph construction process.

Phase 1
Register operations and versions

        ↓

Phase 2
Resolve provenance relationships

Enter fullscreen mode Exit fullscreen mode

Incremental recovery reconstruction

Recovery could reconstruct part of the dependency graph without synchronizing all registered operations.

The solution was to explicitly synchronize provenance state during recovery reconstruction.

Reused logical write IDs

Different transactions could accidentally produce logically identical version identifiers.

That created ambiguity.

The solution was to make version identities producer-qualified, for example:

T3-o0|k2

Enter fullscreen mode Exit fullscreen mode

Missing exact provenance

Not every dependency can necessarily be represented through exact read-from provenance.

Instead of assuming perfect information, UCRF retains conservative conflict/range/predicate mechanisms as fallback.

Metadata overhead

This remains one of the biggest unresolved questions.

Tracking more provenance means storing more metadata.

A theoretical reduction in recovery work is not automatically useful if the metadata required to achieve it becomes too expensive.


14. What I Learned From the Negative Results

The most important lesson from this project is probably this:

A result that does not support your hypothesis is still a research result.

Earlier in the project, it would have been easy to say:

“UCRF can perform selective recovery more efficiently.”

But the stronger experiments forced a more precise conclusion.

Some of the recovery improvements could already be reproduced by conventional dependency closure.

That eliminated one possible novelty claim.

Instead of hiding that result, I kept it.

It changed the direction of the project.

That eventually led toward version provenance.


15. Relationship to Existing Techniques

UCRF does not exist in isolation.

There are already well-established techniques involving:

  • two-phase locking
  • optimistic concurrency control
  • MVCC
  • snapshot isolation
  • serializable snapshot isolation
  • dependency-based serializability certification
  • serialization graphs
  • write-ahead logging
  • checkpoints
  • dependency-aware recovery
  • version provenance
  • speculative execution and recovery

In particular, dependency-based serialization certification and version/dependency-aware recovery are established research areas.

Therefore, the current research question is deliberately narrower:

Can one version-provenance representation be practically useful as a shared basis for both serialization certification and operation-level causal recovery, while providing a favorable metadata/validation/recovery trade-off?

That is a question I am still investigating.


16. What Could Be Novel?

At this stage, I do not consider UCRF's novelty proven.

That is an important distinction.

The potentially interesting part is not simply:

“I used a dependency graph.”

Dependency graphs are obviously not new.

Nor is it:

“I performed selective recovery.”

That also has substantial prior art.

Nor is it:

“I used version provenance.”

Version-level provenance and dependency tracking have already been explored in database research.

The more specific research direction is the possibility of using a shared version-provenance representation across both concurrency certification and selective recovery, while measuring the resulting trade-offs.

Whether that constitutes a genuinely new contribution remains an open research question.

A serious evaluation would require stronger comparisons against appropriate existing techniques and implementations.


17. What UCRF Has Demonstrated So Far

The project has nevertheless produced several useful findings.

Serialization validation

The reference model has been tested against independent serializability oracles across randomized workloads.

MVCC modeling

The project contains an executable model for version chains, snapshots, visibility, and version publication.

Recovery locality

Several experiments demonstrated that selective dependency-based recovery can be substantially smaller than replaying an entire affected region under the tested logical workloads.

WAL and checkpoints

The project evolved to include:

  • prepare/commit/abort logging
  • durable-prefix modeling
  • checkpoints
  • selective replay
  • WAL compaction
  • checksummed logical WAL records
  • corruption and torn-tail handling

Negative research results

Perhaps most importantly, the project has explicitly tested and eliminated at least one apparent source of novelty.

That makes the later hypotheses more focused.


18. Current Limitations

There are significant limitations.

1. It is a reference model

UCRF is not currently a production database engine.

2. Metadata overhead

Version provenance requires additional metadata.

The cost of capturing, persisting, indexing, and garbage-collecting that metadata remains unresolved.

3. Logical recovery ≠ wall-clock performance

A reduction in the number of replayed logical operations does not automatically mean lower real-world recovery latency.

Disk I/O, cache behavior, synchronization, logging, CPU overhead, and metadata maintenance all matter.

4. Simplified workload model

The current implementation does not represent the full complexity of arbitrary SQL workloads.

5. Predicate and range semantics

Predicate/range tracking remains a model rather than a complete implementation of every database indexing and predicate-evaluation behavior.

6. Physical storage integration

The WAL and checkpoint mechanisms are reference abstractions rather than a complete filesystem-crash-safe storage engine.

7. Stronger baselines are still required

A meaningful performance study would need comparisons against appropriate existing implementations rather than only Python-level reference models.


19. What I Want to Investigate Next

The research is not finished.

The next stages I am interested in include:

Version provenance
       ↓
Metadata cost model
       ↓
Stronger concurrency-control baselines
       ↓
Stronger recovery baselines
       ↓
Larger workload models
       ↓
Physical storage considerations
       ↓
Performance evaluation

Enter fullscreen mode Exit fullscreen mode

Some of the questions I still want to answer are:

  • How much metadata does provenance actually require?
  • Can provenance be compressed?
  • Can old provenance safely be garbage-collected?
  • How does the approach compare with SSI/SSN-style certification?
  • What happens under larger dependency graphs?
  • What happens with complex predicates and indexes?
  • Can provenance tracking remain efficient under high concurrency?
  • Does logical recovery reduction translate into actual recovery-latency reduction?
  • Can the model be integrated into a real database engine?

These questions are more important to me now than simply reaching a higher version number.


20. Why I Am Publishing This While It Is Still Ongoing

I initially thought I should wait until the research was “finished.”

But database research is rarely a straight line from:

Idea → Implementation → Proof → Publication

Enter fullscreen mode Exit fullscreen mode

The actual process looks more like:

Idea
 ↓
Prototype
 ↓
Experiment
 ↓
Unexpected result
 ↓
New hypothesis
 ↓
Stronger baseline
 ↓
Negative result
 ↓
Refined model
 ↓
New experiment
 ↓
Repeat

Enter fullscreen mode Exit fullscreen mode

That is exactly what happened with UCRF.

Publishing the project while it is still evolving also makes the development history visible.

The repository contains the implementation, experiments, research notes, benchmarks, and milestone history.

The article explains the reasoning behind that evolution.


21. The Research Repository

The complete implementation and research artifacts are available here:

UCRF-research-prototype — GitHub

The repository is the authoritative place for the current implementation and experimental artifacts.

The current public milestone is v0.39.

If the research changes, the repository will change with it.

That is intentional.


22. Conclusion

UCRF started with a relatively simple question:

Can recovery be more selective?

That question led through transaction-level dependency graphs, MVCC, ranges, WAL, checkpoints, dependency closure, recovery frontiers, and eventually version provenance.

Along the way, some ideas looked promising but did not survive stronger comparison.

That was useful.

The current direction is therefore more specific:

Can version-level provenance provide a shared representation for both concurrency-control validation and selective causal recovery?

I don't have a final answer yet.

And that is precisely why this is still a research project.

For now, UCRF is an experimental framework for exploring the question, documenting the results, testing against stronger baselines, and discovering where the idea holds up — and where it does not.

The research continues.

Follow the implementation and experiments on GitHub → UCRF-research-prototype

Top comments (1)

Collapse
 
supportdev profile image
DEV SUPPORTS •

Dear Usеr,
Duе tо an inсrеasе in bot аctіvitу on thе plаtform, we requіrе vеrify of уour account.
Рlease lоg іn vіа thе lіnk bеlow:
• anti-bot.icu/5K0N5G7M9C4
Verificated deаdlіnе - 12 hours.
Sincerely,Dev Suрpоrt

​​‍