DEV Community

Cover image for I Built My Own Database From Scratch — Meet CleaveDB
John David L. Perez
John David L. Perez

Posted on

I Built My Own Database From Scratch — Meet CleaveDB

I Built My Own Database From Scratch — Meet CleaveDB

I've been building something I've wanted to experiment with for a while: CleaveDB.

CleaveDB is a hybrid relational-document graph database built with Python, Rust, and C++ SIMD, with its own query language, graph relationships, and built-in semantic search.

The idea started with a simple question:

What if relational data, documents, graphs, and semantic search could live together in one database?

What makes CleaveDB different?

CleaveDB has its own query language called CleaveQL.

For example:

POUR INTO users "alice" {
    "name": "Alice",
    "age": 28
}
Enter fullscreen mode Exit fullscreen mode

You can then query the data:

SCOOP users
Enter fullscreen mode Exit fullscreen mode

It also supports graph-style relationships called Bonds:

LINK "users:alice"
TO "users:bob"
AS MUTUAL "friend"
Enter fullscreen mode Exit fullscreen mode

And you can traverse those relationships:

FIND "friend" OF "users:alice"
Enter fullscreen mode Exit fullscreen mode

There's also semantic search:

FIND products
MEANING "something warm for cold weather"
LIMIT 3
Enter fullscreen mode Exit fullscreen mode

So instead of searching only for exact keywords, CleaveDB can search based on meaning.

Under the hood

CleaveDB isn't built on top of an existing database engine. I'm building the storage layer myself using Rust, with C++ SIMD optimisations for performance-critical operations.

The project also includes:

  • B+Tree storage
  • WAL and transactions
  • Graph relationships
  • Semantic/vector search
  • Authentication and multi-tenant isolation
  • Python bindings
  • HTTP, WebSocket and TCP interfaces

It's still very much a work in progress, but building the pieces from scratch has been a really interesting learning experience.

Why build another database?

There are already excellent databases like PostgreSQL, MongoDB, Redis, and Neo4j.

I'm not trying to replace them.

CleaveDB is an experiment to see what happens when you combine relational, document, graph, and semantic capabilities into one system.

If you're interested in databases, Rust, storage engines, graph databases, or just want to see how a database is built from scratch, I'd love to hear your thoughts.

GitHub logo Tribrix23 / CleaveDB

Beyond NoSQL. A hyper-fast graph and vector database you query in plain English. Built on a unique Python, Rust, and C++ SIMD architecture with built-in AI semantic search and strict multi-tenant isolation.

CleaveDB 3.9.0

The polyglot, AVX-512 ready , hybrid relational-document graph database with Transformer attention layers.

Python Rust License NPM Version

CleaveDB 3.9.0 is a ground-up, hybrid Relational Document & Graph database that eliminates the complexity of traditional SQL JOINs, external vector search services, and opaque graph databases. It ships with:

  • Native Graph-Relational Links: Documents are not isolated. They are deeply relational, linked natively through ~10 functional Graph "Bonds" that allow unlimited depth, multi-hop traversals with conversational English, completely eliminating the need for JOINs.
  • A high-performance Rust storage engine built entirely from scratch — no SQLite, no RocksDB, no external storage libraries.
  • AVX-512 / AVX2 C++ SIMD extensions exist in the Rust engine (vector search runs on C++ AVX-512 extensions).
  • A Python interpreter frontend (via PyO3 bindings) that runs the CleaveQL query language.
  • Real neural Transformer embeddings for semantic search via a quantized ONNX model, using ~22MB of RAM.
  • A Go-based distributed coordinator for multi-shard…




I'm especially interested in criticism and ideas for what I should improve next.

Top comments (1)

Collapse
 
akash_siddique profile image
Tipsjazzinferno •

Pretty cool and it's open source